datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pretraining_v1-omega_booksettin-pretraining-data
Ettin Pre-training Data
Phase 1 of 3: Diverse pre-training data mixture (1.7T tokens) used to train the Ettin model suite.
This dataset contains the pre-training phase data used to train all Ettin encoder and decoder models. The data is provided in MDS format ready for use with Composer and the ModernBERT training repository.
📊 Data Composition
Data Source
Tokens (B)
Percentage
Description
DCLM
837.2
49.1%
High-quality web crawl data
CC Head
356.6… See the full description on the dataset page: https://huggingface.co/datasets/jhu-clsp/ettin-pretraining-data.pretraining_v1-omegaMolmoAct-Pretraining-Mixture
MolmoAct - Pretraining Mixture
Data Mixture used for MolmoAct Pretraining. Contains a subset of OXE formulated as Action Reasoning Data along with auxiliary robot data and link to Multimodal Web data.
MolmoAct is a fully open-source action reasoning model for robotic manipulation developed by the Allen Institute for AI. MolmoAct is trained on a subset of OXE and MolmoAct Dataset, a dataset with 10k high-quality trajectories of a single-arm Franka robot performing 93 unique… See the full description on the dataset page: https://huggingface.co/datasets/allenai/MolmoAct-Pretraining-Mixture.multilingual-embeddings-pre-training-curated
📚 Collection | 📝 Multilingual Blog | 📝 English Blog
Contrastive Multilingual Pre-Training
2.16B query–document pairs across eight languages, plus cross-lingual pairs
mDenseOn |
mLateOn |
DenseOn |
LateOn |
PyLate |
FastPlaid
🎯 TL;DR: The multilingual contrastive pre-training corpus used to train mDenseOn and mLateOn. It extends our curated English data recipe (embeddings-pre-training-curated) to French, German, Italian, Spanish, Portuguese, Swedish, Norwegian, and Arabic… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/multilingual-embeddings-pre-training-curated.Pretraining-V1
Indic TTS Unified v1
A large-scale, unified collection of speech data for text-to-speech (TTS) and speech research. This dataset consolidates 17 distinct source datasets into a single, schema-normalized resource covering Indian / South Asian languages, plus major European, African, MENA, and Central Asian languages, with over 13.7 million utterances and 26,000+ hours of audio.
All audio is resampled to 24 kHz mono. Every row follows an identical schema regardless of source… See the full description on the dataset page: https://huggingface.co/datasets/projectkaira/Pretraining-V1.Nemotron-Pretraining-Specialized-v1
Nemotron-Pre-Training-Dataset-v2.1
Dataset Description
The Nemotron-Pre-Training-Dataset-v2.1 extends the previously released Nemotron pretraining datasets with refreshed, higher-quality, and more diverse data across math, code, English Common Crawl, and large-scale synthetic corpora. Designed for the NVIDIA Nemotron 3 family of LLMs, the dataset introduces new Common Crawl code extraction, 2.5T new English web tokens… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Specialized-v1.embeddings-pre-training
Overview
This large-scale dataset is designed for pre-training state-of-the-art text embedding models. Its goal is to reproduce and build upon the data recipe described in the mGTE technical report (Zhang et al., 2024), which details the data sources used to train the GTE family of embedding models but does not release the data itself.
We assembled this dataset as part of a research effort to understand how data composition affects retrieval model quality. Our experiments… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-pre-training.mmBERT-pretraining-data-chunk1
mmBERT Training Data (Ready-to-Use)
Complete Training Dataset: Pre-randomized and ready-to-use multilingual training data (3T tokens) for encoder model pre-training.
This dataset is part of the complete, pre-shuffled training data used to train the mmBERT encoder models. Unlike the individual phase datasets, this version is ready for immediate use but the mixture cannot be modified easily. The data is provided in decompressed MDS format ready for use with ModernBERT's Composer… See the full description on the dataset page: https://huggingface.co/datasets/orionweller/mmBERT-pretraining-data-chunk1.embeddings-pre-training-curated
Embeddings pre-training curated data
This dataset is the English subset of lightonai/embeddings-pre-training, assembled to reproduce the English data recipe described in the mGTE technical report (Zhang et al., 2024).
The mGTE paper describes the data sources used to train the GTE family of multilingual text embedding and reranking models, but does not release the data itself. This dataset is our reconstruction of the English portion of that recipe, curated as part of a research… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-pre-training-curated.Nemotron-Pretraining-Code-v2
Nemotron-Pre-Training-Dataset-v2.1
Dataset Description
The Nemotron-Pre-Training-Dataset-v2.1 extends the previously released Nemotron pretraining datasets with refreshed, higher-quality, and more diverse data across math, code, English Common Crawl, and large-scale synthetic corpora. Designed for the NVIDIA Nemotron 3 family of LLMs, the dataset introduces new Common Crawl code extraction, 2.5T new English web tokens… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Code-v2.nobody-pretraining-shards
Nobody 3.2B — Cybersecurity MoE Pretraining Corpus (v1.0)
Overview
This repository hosts the official, pre-tokenized, and quality-filtered pretraining shards for Nobody 3.2B — Cybersecurity MoE, a foundation model specialized in defensive systems engineering, vulnerability analysis, reverse engineering, and threat intelligence.
The corpus is constructed using a deterministic 10-stage CPU data pipeline and distributed across 12 permanent worker partitions.… See the full description on the dataset page: https://huggingface.co/datasets/mandiyaaman/nobody-pretraining-shards.control-pretraining-datasets-smoke
geodesic-research/control-pretraining-datasets-smoke
Auto-generated by dataset-builder.
Each config below is a separate dataset produced from a versioned YAML build
config. Load with:
from datasets import load_dataset
ds = load_dataset("geodesic-research/control-pretraining-datasets-smoke", "<config_name>", revision="<commit-sha>")
Pin revision= to the specific commit SHA you want; without it, you get the
current HEAD of the dataset repo, which may change when the builder… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/control-pretraining-datasets-smoke.Nemotron-Pretraining-Specialized-v1.2
Nemotron-Pretraining-Specialized-v1.2
Dataset Description:
The Nemotron-Pretraining-Specialized-v1.2 dataset is part of the Nemotron Pretraining Data collection of pretraining datasets. Designed for the NVIDIA Nemotron 3 family of LLMs, this dataset contains a collection of synthetic datasets aimed to improve LLM capabilities on factual recall, moral scenarios, and diverse generative and multiple choice questions.
Note: These are new datasets, not replacements.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Specialized-v1.2.mmBERT-pretraining-data-chunk0
mmBERT Training Data (Ready-to-Use)
Complete Training Dataset: Pre-randomized and ready-to-use multilingual training data (3T tokens) for encoder model pre-training.
This dataset is part of the complete, pre-shuffled training data used to train the mmBERT encoder models. Unlike the individual phase datasets, this version is ready for immediate use but the mixture cannot be modified easily. The data is provided in decompressed MDS format ready for use with ModernBERT's Composer… See the full description on the dataset page: https://huggingface.co/datasets/orionweller/mmBERT-pretraining-data-chunk0.Nemotron-Pretraining-Code-v1
Nemotron-Pre-Training-Dataset-v1 Release
Data Overview
This pretraining dataset, for generative AI model training, preserves high-value math and code while enriching it with diverse multilingual Q&A, fueling the next generation of intelligent, globally-capable models.
This dataset supports NVIDIA Nemotron Nano 2, a family of large language models (LLMs) that consists of the NVIDIA-Nemotron-Nano-9B-v2, NVIDIA-Nemotron-Nano-9B-v2-Base, and NVIDIA-Nemotron-Nano-12B-v2-Base… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Code-v1.carbon-pretraining-corpus
🧬 Carbon Pretraining Corpus
Description
173M DNA & RNA sequences · 1.1 trillion nucleotides — the DNA pretraining mixture used to train Carbon, a genomic foundation model.
This dataset is a collection of data sources intended for training genomic foundation models, such as Carbon. It contains DNA and RNA sequences spanning eukaryote and prokaryote species.
Across the four main configs it totals 1.1 T DNA base pairs (180B tokens with Carbon's 6-mer tokenizer). A… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceBio/carbon-pretraining-corpus.Jp-OCR-Pretrainingthe_stack_v2_python_repos_pretraining_dataset_imported_context-datasetNemotron-Pretraining-Specialized-v1.1
Nemotron-Pretraining-Specialized-v1.1
Dataset Description:
The Nemotron-Pretraining-Specialized-v1.1 dataset is part of the Nemotron Pretraining Data collection of pretraining datasets. Designed for the NVIDIA Nemotron 3 family of LLMs, this dataset contains a collection of synthetic datasets aimed to improve LLM capabilities in code concepts and algorithms, formal logic, economics, and multiple choice questions. The code concepts dataset is an instance of a general… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Specialized-v1.1.Nemotron-Pretraining-SFT-v1
Nemotron-Pre-Training-Dataset-v1 Release
Data Overview
This pretraining dataset, for generative AI model training, preserves high-value math and code while enriching it with diverse multilingual Q&A, fueling the next generation of intelligent, globally-capable models.
This dataset supports NVIDIA Nemotron Nano 2, a family of large language models (LLMs) that consists of the NVIDIA-Nemotron-Nano-9B-v2, NVIDIA-Nemotron-Nano-9B-v2-Base, and NVIDIA-Nemotron-Nano-12B-v2-Base… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-SFT-v1.mmBERT-pretraining-data-chunk2
mmBERT Training Data (Ready-to-Use)
Complete Training Dataset: Pre-randomized and ready-to-use multilingual training data (3T tokens) for encoder model pre-training.
This dataset is part of the complete, pre-shuffled training data used to train the mmBERT encoder models. Unlike the individual phase datasets, this version is ready for immediate use but the mixture cannot be modified easily. The data is provided in decompressed MDS format ready for use with ModernBERT's Composer… See the full description on the dataset page: https://huggingface.co/datasets/orionweller/mmBERT-pretraining-data-chunk2.math-pretraining-corpusaudio_pretraining_sono_synthetic
FORGE3D SLAT training data (flat layout)
Two directories only:
tars/ — every tar, flat. docs/ — manifests, CSVs, notes (former nested paths flattened with / → __).
docs/MANIFEST.csv — one row per object: id, tar, has_slat, has_white_cond, has_recgen_image, recgen_image_tar, recgen_views, is_recgen, ... Use it to select what to download.
What is in tars/
tars
content
used for SLAT training
shard_*.tar, inc_*.tar (not inc_slat2*/inc_rgslat*)… See the full description on the dataset page: https://huggingface.co/datasets/Ronaldo-GOAT/audio_pretraining_sono_synthetic.text_audio_pretraining_for_GAN
bw_jh_dataset — Robot 3D Point-Track Dataset (droid / hrdexdb / robocasa)
A unified multi-source robot manipulation dataset with 3D point tracks of the robot gripper and arm, calibrated multi-view RGB video, language instructions, and success labels. Three splits:
Split
Episodes
Size
Views
Source
droid
27,615 (+~1,050 in droid-11)
~280 GB
2 exterior ZED views
DROID (real, Franka)
hrdexdb
1,601
~998 GB
22 calibrated cameras
HRDexDB (real, xArm6 + dexterous hands)… See the full description on the dataset page: https://huggingface.co/datasets/rooty2020/text_audio_pretraining_for_GAN.Vn-OCR-Pretrainingwikipedia_chunked
Dataset Card for "wikipedia_chunked"
More Information needed
nemotron-pretraining-specialized-collection
Nemotron Pretraining Specialized Collection
This repository is a convenience collection of the NVIDIA Nemotron Pretraining Specialized releases. It preserves each original subset as a separate Hugging Face configuration, so consumers can select a single domain-focused subset without visiting multiple source repositories.
Contents
Source release
Configurations included
nvidia/Nemotron-Pretraining-Specialized-v1
Wiki Rewrite, Math Textbooks, STEM SFT… See the full description on the dataset page: https://huggingface.co/datasets/CrowdMind/nemotron-pretraining-specialized-collection.ts-icl-pretraining-corpus
TS-ICL Pretraining Corpus (community reconstruction)
A unified, cleaned reconstruction of the univariate pretraining corpus described in
Table 5 of TS-ICL: A Flexible Time-Indexed Foundation Model for Time Series via
In-Context Learning (Le Naour, Nabil & Petralia, EDF R&D; arXiv:2606.05878). The TS-ICL
authors did not release their pretraining data pipeline, so this corpus is rebuilt from the
named upstream sources (LOTSA, Chronos, and the TempoPFN synthetic generators) and… See the full description on the dataset page: https://huggingface.co/datasets/JuaAI/ts-icl-pretraining-corpus.Nemotron-Pretraining-Code-v3
Nemotron-Pretraining-Code-v3
Dataset Description:
The Nemotron-Pretraining-Code-v3 dataset is part of the Nemotron Pretraining Data collection of pretraining datasets. Designed for the NVIDIA Nemotron 3 family of LLMs, this dataset is intended to improve the coding capabilities of LLMs.
The Nemotron-Pretraining-Code-v3 dataset contains the metadata corresponding to the raw source-code update to our Nemotron-Pretraining-Code-v2 and Nemotron-Pretraining-Code-v1… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Code-v3.
