datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LLaVA-OneVision-1.5-Mid-Training-85M
🚀 LLaVA-One-Vision-1.5-Mid-Training-85M Dataset is being uploaded 🚀
Upload Status
All Completed: ImageNet-21k、LAIONCN、DataComp-1B、Zero250M、COYO700M、SA-1B、MINT、Obelics
📜 Cite
If you find LLaVA-One-Vision-1.5-Mid-Training-85M useful in your research, please consider to cite the following related papers:
@misc{an2025llavaonevision15fullyopenframework,
title={LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training}… See the full description on the dataset page: https://huggingface.co/datasets/mvp-lab/LLaVA-OneVision-1.5-Mid-Training-85M.stack-v3-train
🥞 The Stack v3
What is it?
What is being released
How to download and use it
Dataset statistics
Dataset structure
Dataset creation
Considerations for using the data
Additional information
What is it?
The Stack v3 is the largest, most up-to-date open dataset of source code, crawled directly from GitHub and built to pre-train code LLMs with full-repository context. It is the successor to The Stack v2 and, like its predecessor, is released to make the training… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceCode/stack-v3-train.mmlu_no_trainThis dataset contains a copy of the cais/mmlu HF dataset but without the auxiliary_train split that takes a long time to generate again each time when loading multiple subsets of the dataset.
Please visit https://huggingface.co/datasets/cais/mmlu for more information on the MMLU dataset.
VideoChat-Flash-Training-Data
🦜 VideoChat-Flash-Training-Data
This repos contains all annotaions and most videos for training VideoChat-Flash.
📕 How to use the LongVid data?
For video_dir like longvid_subset/coin_grounding_10k_zip, you need to concat this dir to a zip file as follows:
cat ego4dhcap_eventunderstanding_2k_zip/* > ego4dhcap_eventunderstanding_2k.zip
✏️ Citation
@article{li2024videochatflash,
title={VideoChat-Flash: Hierarchical Compression for Long-Context… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/VideoChat-Flash-Training-Data.Lucie-Training-Dataset
Lucie Training Dataset Card
The Lucie Training Dataset is a curated collection of text data
in English, French, German, Spanish and Italian culled from a variety of sources including: web data, video subtitles, academic papers,
digital books, newspapers, and magazines, some of which were processed by Optical Character Recognition (OCR). It also contains samples of diverse programming languages.
The Lucie Training Dataset was used to pretrain Lucie-7B,
a foundation LLM with… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Lucie-Training-Dataset.GenFusion_Training_DataUrbanVerse-Training-Scenes
UrbanVerse Training Scenes (Urban Cousins)
A collection of ready-to-simulate urban 3D scenes in OpenUSD
for NVIDIA Isaac Sim / Isaac Lab, released by the
VAIL-UCLA lab. Each scene is a self-contained
USD stage with all of its materials and textures, so it can be opened and
simulated directly.
The scenes are generated with UrbanVerse — Scaling Urban Simulation by
Watching City-Tour Videos (Liu et al., ICLR 2026,
arXiv:2510.15018,
project page) — whose UrbanVerse-Gen
pipeline… See the full description on the dataset page: https://huggingface.co/datasets/UCLA-VAIL/UrbanVerse-Training-Scenes.NatureLM-audio-training
Dataset card for NatureLM-audio-training
Overview
NatureLM-audio-training is a large and diverse audio-language dataset designed for training bioacoustic models that can generate a natural language answer to a natural language query on a reference bioacoustic audio recording.
For example, for an in-the-wild audio recording of a bird species, a relevant query might be "What is the common name for the focal species in the audio?" to which an audio-language model trained… See the full description on the dataset page: https://huggingface.co/datasets/EarthSpeciesProject/NatureLM-audio-training.comma_v0.1_training_dataset
Comma v0.1 dataset
This repository contains the dataset used to train Comma v0.1-1T and Comma v0.1-2T.
It is a slightly modified and consolidated version of the Common Pile v0.1 "filtered" data.
If you are looknig for the raw Common Pile v0.1 data, please see this collection.
You can learn more about Common Pile in our paper.
Mixing rates and token counts
The Comma v0.1 models were trained in two stages, a "main" stage and a "cooldown" stage.
During each stage, we… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/comma_v0.1_training_dataset.hf-training-corpus
HF Training Corpus
Bulk-scraped multimodal training corpus: ~92,000 Hugging Face datasets streamed,
normalized and stored as per-dataset Parquet files under data/.
Every modality is captured: text, images (PNG bytes), audio (mono 16-bit WAV bytes),
video (bytes, capped) and tabular/timeseries (serialized to text).
Pipeline
Live crawl from a 92,458-ID corpus list (see
https://github.com/shadyuwugurl/hf-datasets-trees/blob/main/datasets.txt)
streaming=True reads… See the full description on the dataset page: https://huggingface.co/datasets/MC7ever/hf-training-corpus.4DThinker-Training-Data
4DThinker Training Data
This repository contains the training data for 4DThinker, a framework that enables VLMs to "think with 4D" through dynamic latent mental imagery, built upon SpatialVID and DSR_Suite-Data.
Data Structure
data/
├── dift_data.jsonl # DIFT training data (~38K samples)
├── 4drl_data_filtered.jsonl # 4DRL training data (~37K samples)
└── processed_data/ # Video frames & mask overlays
├── <video_id>/
│ ├── frames/… See the full description on the dataset page: https://huggingface.co/datasets/jankin123/4DThinker-Training-Data.SmolDataEnvs-harbor-train
📊 SmolDataEnvs: Harbor (train)
5.5K+ RL tasks for hill-climbing small models in code and data science.
A 2B model on these tasks. Left: what it optimises. Right: 144 held-out tasks it never trains on.
Two runs over the same 5,000 tasks: shuffled against a curriculum ordered easiest to hardest.
The training suite: 5,000 hands-on data-analysis tasks. Each one drops an agent into a sandbox with a
real dataset and a question, and asks it to explore the data… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/SmolDataEnvs-harbor-train.Nemotron-Post-Training-Dataset-v1
Nemotron-Post-Training-Dataset-v1 Release
This dataset is a compilation of SFT data that supports improvements of math, code, stem, general reasoning, and tool calling capabilities of the original Llama instruct model Llama-3.3-Nemotron-Super-49B-v1.5.
Llama-3.3-Nemotron-Super-49B-v1.5 is an LLM which is a derivative of Meta Llama-3.3-70B-Instruct (AKA the reference model).
Llama-3.3-Nemotron-Super-49B-v1.5 offers a great tradeoff between model accuracy and efficiency. Efficiency… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Post-Training-Dataset-v1.Beta-Pre-Train-Corpus
Reactive AI / Beta Pre-Train Corpus
Pre-training corpus for RxT-Beta models, created from public & open datasets. Includes high-quality english and polish web crawl data, mathematic and scientific subsets,
and code in different programming languages.
2k subsets are filtered for 1024-2048 tokens, except MegaMath Web Pro and GitHub Code subsets, that were filtered for 512-2048 tokens
Subsets & original datasets
FineWeb-Edu
fineweb-edu-s100 (51.3M examples) - 50% of… See the full description on the dataset page: https://huggingface.co/datasets/ReactiveAI/Beta-Pre-Train-Corpus.multilingual-embeddings-pre-training-curated
📚 Collection | 📝 Multilingual Blog | 📝 English Blog
Contrastive Multilingual Pre-Training
2.16B query–document pairs across eight languages, plus cross-lingual pairs
mDenseOn |
mLateOn |
DenseOn |
LateOn |
PyLate |
FastPlaid
🎯 TL;DR: The multilingual contrastive pre-training corpus used to train mDenseOn and mLateOn. It extends our curated English data recipe (embeddings-pre-training-curated) to French, German, Italian, Spanish, Portuguese, Swedish, Norwegian, and Arabic… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/multilingual-embeddings-pre-training-curated.RoboJudge-Train
RoboJudge Data
This repository hosts the video assets and canonical metadata for evaluating
Physical Adherence (PA) and Instruction Alignment (IA) in generated
embodied-manipulation videos.
Current RoboJudge release
Use robojudge_release/ for the paper release:
Path
Contents
robojudge_release/train/physical_adherence.json
12,351 PA training records
robojudge_release/train/instruction_alignment.json
11,520 IA training records… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFriends/RoboJudge-Train.Bitext-customer-support-llm-chatbot-training-dataset
Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-customer-support-llm-chatbot-training-dataset.TrainingData_Stage3
AnchorSR Stage3 · metric-v1.0
直接选择 Small / Large
配置
训练题数
用途
small
1,000,000
先验证答案监督/先验恢复,按新版 Large 联合分布抽样
large
89,801,853
筛选后的完整训练集合,包含 Small 全部样本
from datasets import load_dataset
data = load_dataset('AnchorSR/TrainingData_Stage3', 'small', # 或 large
revision='metric-v1.0', streaming=True)
这是对 scaling-v1.0 的语义筛选与统一任务分类,不是增加新数据源。
Large 从 89,828,269 题保留 89,801,853 题,隔离 26,416 题。
旧标签 scaling-v1.0 / video-v1.0 / large-v1.0… See the full description on the dataset page: https://huggingface.co/datasets/AnchorSR/TrainingData_Stage3.abc_130k_v3_trainThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
14
],
"names": [
"left_arm_joint_1",
"left_arm_joint_2",
"left_arm_joint_3",
"left_arm_joint_4",
"left_arm_joint_5"… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/abc_130k_v3_train.IllusionChar_train
IllusionChar — Training Set
Dataset summary
This repository contains the training split of IllusionChar, the optical character recognition (OCR) component of Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions. The task is to transcribe a hidden, case-sensitive alphanumeric sequence from an illusory image, or return No illusion when no sequence is embedded.
Sequences contain 3–5 characters drawn from digits, uppercase Latin letters, and… See the full description on the dataset page: https://huggingface.co/datasets/VQA-Illusion/IllusionChar_train.Ezaris-Training-Sets
Ezaris-Training-Sets
Complete training data, checkpoints, code and provenance for the Ezaris-1B program (formerly LUNA-1B) — a 1.2B-parameter Llama-style model trained from scratch by ASTERIZER.
Layout
Path
Contents
pretrain_raw_datasets/
Raw pretraining sources (SmolLM-corpus, fineweb-edu, OpenWebMath, Finemath, Wiki, code)
pretrain_cleaned_datasets/
Cleaned / deduplicated / tokenized pretrain sets + phase-2 exact & near-dedup outputs… See the full description on the dataset page: https://huggingface.co/datasets/ASTERIZER/Ezaris-Training-Sets.csgo-training-cache
CSGO training cache
Precomputed feature-only training samples, continuously converted from four source datasets:
peih/csgo-3090-v4-overflow-20260923
peih/csgo-4090-v4-20260921
peih/csgo-4090-v4-overflow-20260923
peih/csgo-dual3090-v4-backup
This dataset is still growing. manifest.json is the current verified snapshot; manifests/vNNNNNN.json contains immutable historical snapshots. Read only listed shards at their pinned revision. New uploads become available after their… See the full description on the dataset page: https://huggingface.co/datasets/peih/csgo-training-cache.ImageNet1K-trainmapping:
n01440764 tench, Tinca tinca
n01443537 goldfish, Carassius auratus
n01484850 great white shark, white shark, man-eater, man-eating shark, Carcharodon carcharias
n01491361 tiger shark, Galeocerdo cuvieri
n01494475 hammerhead, hammerhead shark
n01496331 electric ray, crampfish, numbfish, torpedo
n01498041 stingray
n01514668 cock
n01514859 hen
n01518878 ostrich, Struthio camelus
n01530575 brambling, Fringilla montifringilla
n01531178 goldfinch, Carduelis carduelis
n01532829 house finch… See the full description on the dataset page: https://huggingface.co/datasets/mrm8488/ImageNet1K-train.embeddings-pre-training-curated
Embeddings pre-training curated data
This dataset is the English subset of lightonai/embeddings-pre-training, assembled to reproduce the English data recipe described in the mGTE technical report (Zhang et al., 2024).
The mGTE paper describes the data sources used to train the GTE family of multilingual text embedding and reranking models, but does not release the data itself. This dataset is our reconstruction of the English portion of that recipe, curated as part of a research… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-pre-training-curated.tmax-train4k-benchflow4,000 tasks sampled stratified by domain and complexity from allenai/TMax-15K (ODC-BY; Tmax paper arXiv:2606.23321), converted with bench tasks migrate --remove-legacy. Training corpus for PostTrain Arena hosted runs; never an evaluation set.
OmniRet-train
OmniRet training dataset
OmniRet-train is the training-data release for
OmniRet, a unified retrieval model for
text, image, video, and audio. This card documents the released snapshot for
researchers training or analyzing OmniRet.
Dataset summary
The release contains 6,405,109 query rows and 7,119,841 candidate rows from 30
datasets. It covers 15 retrieval directions across text (T), image (I), video
(V), and audio (A). The OmniRet paper reports this corpus as… See the full description on the dataset page: https://huggingface.co/datasets/chuonghm/OmniRet-train.CryoLithe-training-datasetThe training Dataset for CryoLithe Models
The dataset contains selected tilt series, tilt angles, and corresponding cryo-CARE+IsoNet and Icecream reconstructions using odd/even pairs. For EMPIAR-11058
Icecream reconstructions were obtained by splitting across angles.
Whenever available, we also provide dose-fractionated tilt series.
Dataset format:
Files ending with '.rawtlt' or '.tlt' correspond to the tilt angles.
Files ending with '_corrected.mrc' correspond to cryo-CARE+IsoNet… See the full description on the dataset page: https://huggingface.co/datasets/sada-group/CryoLithe-training-dataset.Nemotron-Post-Training-Dataset-v2
Nemotron-Post-Training-Dataset-v2 Release
Data Overview
This dataset adds to NVIDIA’s post-training dataset releases with an extension of SFT and RL data into five target languages: Spanish, French, German, Italian and Japanese. The data supports improvements of math, code, general reasoning, and instruction following capabilities of the NVIDIA-Nemotron-Nano-9B-v2-Base, in support of release of NVIDIA-Nemotron-Nano-8B-v2-Reasoning.
NVIDIA-Nemotron-Nano-9B is a family of… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Post-Training-Dataset-v2.train-bn
Dataset Card for "train-bn"
More Information needed
LEMAS-Dataset-train
Overview
This dataset is part of LEMAS-Project (lemas-project.github.io/LEMAS-Project).
It contains a large-scale training set (150k+ hours) and a curated evaluation set
(500 utterances per language) covering 10 languages, all with word-level alignment.
Fields
key: unique utterance identifier; the first two characters indicate the language ID
audio: relative path to the MP3 audio file (in the eval set, this key is renamed to "file_name" for compatibility with the viewer)… See the full description on the dataset page: https://huggingface.co/datasets/LEMAS-Project/LEMAS-Dataset-train.
