datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
financial-news-multisource
Multi-Source Financial & General News
🚀 57.1 MILLION ROWS OF NEWS CONTENT — one unified corpus for market-aware AI/ML
I combined 24 public news datasets (many small on their own) into one consistent, ready-to-use layer so you don’t have to wrangle them yourself. Everything is normalized to a minimal schema (date, text, extra_fields) and shipped as Parquet shards per subset—streamable, DuckDB-friendly, and built with a trading date policy (this can be edited if folks see other… See the full description on the dataset page: https://huggingface.co/datasets/idleengine/financial-news-multisource.financial-news-multisource
Multi-Source Financial & General News
🚀 57.1 MILLION ROWS OF NEWS CONTENT — one unified corpus for market-aware AI/ML
I combined 24 public news datasets (many small on their own) into one consistent, ready-to-use layer so you don’t have to wrangle them yourself. Everything is normalized to a minimal schema (date, text, extra_fields) and shipped as Parquet shards per subset—streamable, DuckDB-friendly, and built with a trading date policy (this can be edited if folks see other use… See the full description on the dataset page: https://huggingface.co/datasets/Brianferrell787/financial-news-multisource.multisource-eeg-motorimagerymultisource-membench
Multi-Source Memory Benchmark
Status — public release.
A diagnostic testbed for selective question-answering (ANSWER / SKIP) over conflicting multi-source personal memory.
Each persona has five evidence streams projected from a single latent event table with known, controlled per-source distortions (bias direction, dropout rate, granularity), allowing methods to be measured against the latent ground truth rather than against any single source.
The benchmark accompanies the… See the full description on the dataset page: https://huggingface.co/datasets/ytc1997/multisource-membench.AgentFEM-MultiSource-Heat-2D
AgentFEM · Multi-source heat conduction
Two smooth volumetric heat sources in a rectangular conducting plate. Learn how source placement, spread, intensity, geometry and conductivity shape the temperature field.
256 independently solved parameter sets · full meshes and fields · 192 / 32 / 32 split · CC BY 4.0
Built with AgentFEM. A small, reproducible engineering dataset for surrogate learning, field prediction and numerical-method experiments.
Physical problem… See the full description on the dataset page: https://huggingface.co/datasets/HaomingLuo/AgentFEM-MultiSource-Heat-2D.multisource_textv2.5-box-precise-and-multi-source-suplement
v2.5 box precise and multi source suplement
64,061 public images, from the final deduplicated trainv2.5final1 corpus. 7,574 retained MyData_Fire images and all their annotations are excluded because their source metadata says Private. The ongoing training corpus is preserved separately.
A further 12,025 images from sources with undocumented redistribution rights are held outside this public release.
Separate annotation variants
Configuration
Categories
COCO… See the full description on the dataset page: https://huggingface.co/datasets/fireviewer/v2.5-box-precise-and-multi-source-suplement.multisource-memory-benchmark
Multi-Source Memory Benchmark
Status — anonymous artefact for double-blind review (NeurIPS 2026 Evaluations & Datasets Track).
Author identities, organisations, and funders are intentionally withheld until the review period concludes.
A diagnostic testbed for selective question-answering (ANSWER / SKIP) over conflicting multi-source personal memory.
Each persona has five evidence streams projected from a single latent event table with known, controlled per-source distortions… See the full description on the dataset page: https://huggingface.co/datasets/anon-neuripsed26/multisource-memory-benchmark.multisource_tok_nobookMulti-Source-Video-Captioning
Multi-source Video Captioning (MSVC) Dataset Card
Dataset details
Dataset type:
MSVC is a set of collected video captioning data. It is constructed to ensure a robust and thorough evaluation of Video-LLMs' video-captioning capabilities.
Dataset detail:
MSVC is introduced to address limitations in existing video caption benchmarks, MSVC samples a total of 1,500 videos with human-annotated captions from MSVD, MSRVTT, and VATEX, ensuring diverse scenarios and domains.… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/Multi-Source-Video-Captioning.skin-assessment-images-multisource-provenance-preservedA source-stratified package of 20,153 distinct skin-condition images in 16 target classes. The archive Skin_Assessment_Multi_Source_Provenance_Preserved.zip contains the image tree and package metadata, including manifest.csv, class_counts.csv, source_class_counts.csv, and source_license_metadata.csv.
The archive was built by exact SHA-256 deduplication: 20,153 included image files have distinct hashes. Its manifest contains 21,052 source rows, including 899 excluded exact-duplicate records.… See the full description on the dataset page: https://huggingface.co/datasets/Cuebot/skin-assessment-images-multisource-provenance-preserved.sonarvision-multisource-v6
🌊 KADAL Multi-Source Side-Scan Sonar Dataset (v6)
SIH 2026 | SIH26057 (Ministry of Earth Sciences / NIOT) — AI-Powered Automated Underwater Marine Debris and Anomaly Detection System using Side-Scan Sonar Imagery.
Official Project Source Code & Documentation:🔗 https://github.com/Dinoman67/sonarvision(Refer to the GitHub repository for preprocessing scripts, augmentations, model weights, edge dashboard, and optional supporting features such as C2 (NMEA/KML) exports… See the full description on the dataset page: https://huggingface.co/datasets/Dinoman1221/sonarvision-multisource-v6.flan_source_race_high_Write_a_multi_choice_question_options_given__67iran_cpi_and_inflation_multisource
شاخص بهای مصرفکننده و تورم ایران (چندمرجعی)
سریهای رسمیِ شاخص قیمت مصرفکننده (CPI) و نرخ تورم ایران از دو مرجعِ رسمی: مرکز آمار
ایران (ماهانه) و بانک مرکزی (سالانه، سریِ تاریخیِ بلند). هر مرجع سریِ مستقلِ خود را دارد و در
کنارشان یک نمای یکپارچه ارائه میشود.
پوشش: بانک مرکزی سالانه ۱۳۱۵–۱۴۰۱ · مرکز آمار ماهانه ۱۳۹۰–۱۴۰۵
سطح: ملی · واحد: شاخص و درصد
شاخصها
شاخص
مرجع
تناوب
شاخص قیمت مصرفکننده
مرکز آمار (پایهٔ ۱۴۰۰) · بانک مرکزی (پایهٔ ۱۳۹۵)… See the full description on the dataset page: https://huggingface.co/datasets/Farmaanaa/iran_cpi_and_inflation_multisource.multisource-spatial-point-predictiondatasets for KDD24 Self-consistent Deep Geometric Learning for Heterogeneous Multi-source Spatial Point Data Prediction
pii-detection-multisource-en-saudi-arabic
PII Detection Multisource EN + Saudi/Arabic
284,619 English examples. 2,088,335 labelled spans. 31 entity types. One label space.
Four public PII datasets, merged into a single schema, plus Saudi and Arabic
coverage that none of them had, plus material for two failure modes that matter
when you run redaction in production.
Built for OnKith, a privacy first voice assistant that
transcribes speech and strips personal information on the device itself, before
anything is allowed to… See the full description on the dataset page: https://huggingface.co/datasets/KhalidAlharbi377/pii-detection-multisource-en-saudi-arabic.code-multi-line-infilling-benchmark
Dataset Summary
This dataset is used to evaluate Multi-Line fill in the middle code completion capabilities of a system.
The dataset is derived from SWE-Bench dataset.
Evaluation is performed by stiching the generated middle portion, with the other patch and passing into the SWE Evaluation harness, which runs unit test verification and calculate Pass@1.
Data Instances
In addition to the fields already calculated by SWE-Bench dataset, this dataset contains five… See the full description on the dataset page: https://huggingface.co/datasets/sourcegraph/code-multi-line-infilling-benchmark.buzz_sources_033_processed_unified_multi_newsGreatPlains-Multisource-2000-2024
Great Plains 8-day Multisource NDVI–Climate Time Series (2000–2024)
1. 数据集概述 (Dataset Summary)
本数据集以 美国南部大平原草原区 为研究区域,范围约为:
经度:105°W–95°W
纬度:32°N–40°N
该区域为典型干旱敏感区,植被以草原为主,对降水异常和干旱事件高度敏感。
本数据集整合了 2000–2024 年间的多源观测,包含:
MODIS Terra NDVI(8 日合成)
CHIRPS 日尺度降水
ERA5-Land 日尺度表层土壤含水量、2 m 气温、潜在蒸散
以 NDVI 时间步为主轴构建的 8 日对齐多变量时间序列(统一建模输入)
适合于:
干旱监测和评估
NDVI 与气候驱动因子的时滞/响应分析
时序预测任务:ARIMA、多变量 LSTM、Encoder–Decoder 等
多源气象–遥感数据融合研究
2. 数据来源 (Data Sources)… See the full description on the dataset page: https://huggingface.co/datasets/TaoDerong/GreatPlains-Multisource-2000-2024.lemonseed-multisource-cogen
lemonseed-multisource-cogen
LemonSeed — multi-source teacher-co-gen training.
Contents
intelligent_multisource_train.jsonl (464 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
LLM-teacher co-generated instruction/chat data for LemonSeed fine-tuning.
crux-mds-multi_news-sourceartificial-intelligence-multi-source-datasetko-multisource-retrieval-dataset
구성
여러 공개 출처를 하나의 스키마로 통합한 한국어 query–passage 페어 데이터셋입니다.
stage1, stage2 두 subset. 각각 train 100,000 / valid 10,000행 (총 220,000행, parquet, 151MB)
필드
설명
dataset
원 출처 식별자 (18종)
query
질의문 (2~2,020자)
passage
대응 문서 (10~2,050자)
출처
AI 허브 — 기계독해 지식검색 일반상식 도서자료 행정문서 뉴스기사 금융법률 숫자연산 표정보 추상요약 이벤트 (AI허브_ 접두사)
KLUE — klue-mrc klue-nli klue-sts
기타 — kakao-nli 공공데이터포털-deepqa LGNLP
참고
출처별로 페어 성격이 다르므로 용도에 맞게 필터링을 권장합니다.
klue-nli klue-sts… See the full description on the dataset page: https://huggingface.co/datasets/FronyAI/ko-multisource-retrieval-dataset.summarization_wikilingua_multisource_chatgpt_segmentationflan_source_task611_mutual_multi_turn_dialogue_291high_quality_data_10k_multisourcemultisource-esco-set
MultiSource-ESCO-Skills: A Unified Dataset for Skill Extraction
This dataset aggregates data from multiple sources—course descriptions, CV content, and job descriptions—all linked to ESCO skills. It is designed to help researchers and practitioners develop and fine-tune NLP models (e.g., BERT or SentenceTransformer-based models) for automated skill extraction.
Dataset Overview
Name: MultiSource-ESCO-Skills
Sources:
Course Content: Educational course materials
CV Content:… See the full description on the dataset page: https://huggingface.co/datasets/Boanerges/multisource-esco-set.deltalora-memory-multisource-seed42
Delta-LoRA Memory Multisource Seed42
This repository contains a mixed memory-training corpus built for Delta-LoRA long-context and online-memory experiments.
Files:
deltalora_memory_multisource_large_seed42.jsonl: full mixed corpus
deltalora_memory_multisource_large_seed42.jsonl.summary.json: summary for the full corpus
deltalora_memory_multisource_6k_seed42.jsonl: stratified 6k subset for quick training
deltalora_memory_multisource_6k_seed42.jsonl.summary.json: summary for the 6k… See the full description on the dataset page: https://huggingface.co/datasets/huaXiaKyrie/deltalora-memory-multisource-seed42.flan_source_race_middle_Write_a_multi_choice_question_options_given__116lemonseed-multisource-cogen-corrective
lemonseed-multisource-cogen-corrective
LemonSeed — multi-source corrective teacher-co-gen training (v3).
Contents
intelligent_multisource_corrective_train_v3.jsonl (522 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
LLM-teacher co-generated instruction/chat data for LemonSeed fine-tuning.
