Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01applied-ai-018 /pretraining_v1-omega_bookstabular100M<n<1B25 likes489k downloads2y agoHugging Face02jhu-clsp /ettin-pretraining-data Ettin Pre-training Data Phase 1 of 3: Diverse pre-training data mixture (1.7T tokens) used to train the Ettin model suite. This dataset contains the pre-training phase data used to train all Ettin encoder and decoder models. The data is provided in MDS format ready for use with Composer and the ModernBERT training repository. 📊 Data Composition Data Source Tokens (B) Percentage Description DCLM 837.2 49.1% High-quality web crawl data CC Head 356.6… See the full description on the dataset page: https://huggingface.co/datasets/jhu-clsp/ettin-pretraining-data.text-generation10 likes63k downloads1y agoHugging Face03applied-ai-018 /pretraining_v1-omega5 likes27k downloads2y agoHugging Face04allenai /MolmoAct-Pretraining-Mixture MolmoAct - Pretraining Mixture Data Mixture used for MolmoAct Pretraining. Contains a subset of OXE formulated as Action Reasoning Data along with auxiliary robot data and link to Multimodal Web data. MolmoAct is a fully open-source action reasoning model for robotic manipulation developed by the Allen Institute for AI. MolmoAct is trained on a subset of OXE and MolmoAct Dataset, a dataset with 10k high-quality trajectories of a single-arm Franka robot performing 93 unique… See the full description on the dataset page: https://huggingface.co/datasets/allenai/MolmoAct-Pretraining-Mixture.imagerobotics10M<n<100M14 likes20k downloads1y agoHugging Face05lightonai /multilingual-embeddings-pre-training-curated 📚 Collection | 📝 Multilingual Blog | 📝 English Blog Contrastive Multilingual Pre-Training 2.16B query–document pairs across eight languages, plus cross-lingual pairs mDenseOn | mLateOn | DenseOn | LateOn | PyLate | FastPlaid 🎯 TL;DR: The multilingual contrastive pre-training corpus used to train mDenseOn and mLateOn. It extends our curated English data recipe (embeddings-pre-training-curated) to French, German, Italian, Spanish, Portuguese, Swedish, Norwegian, and Arabic… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/multilingual-embeddings-pre-training-curated.texttext-retrieval1B<n<10B6 likes10k downloads2mo agoHugging Face06projectkaira /Pretraining-V1 Indic TTS Unified v1 A large-scale, unified collection of speech data for text-to-speech (TTS) and speech research. This dataset consolidates 17 distinct source datasets into a single, schema-normalized resource covering Indian / South Asian languages, plus major European, African, MENA, and Central Asian languages, with over 13.7 million utterances and 26,000+ hours of audio. All audio is resampled to 24 kHz mono. Every row follows an identical schema regardless of source… See the full description on the dataset page: https://huggingface.co/datasets/projectkaira/Pretraining-V1.audiotext-to-speech10M<n<100M0 likes10k downloads2mo agoHugging Face07nvidia /Nemotron-Pretraining-Specialized-v1 Nemotron-Pre-Training-Dataset-v2.1 Dataset Description The Nemotron-Pre-Training-Dataset-v2.1 extends the previously released Nemotron pretraining datasets with refreshed, higher-quality, and more diverse data across math, code, English Common Crawl, and large-scale synthetic corpora. Designed for the NVIDIA Nemotron 3 family of LLMs, the dataset introduces new Common Crawl code extraction, 2.5T new English web tokens… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Specialized-v1.texttext-generation10M<n<100M88 likes6.2k downloads10mo agoHugging Face08lightonai /embeddings-pre-training Overview This large-scale dataset is designed for pre-training state-of-the-art text embedding models. Its goal is to reproduce and build upon the data recipe described in the mGTE technical report (Zhang et al., 2024), which details the data sources used to train the GTE family of embedding models but does not release the data itself. We assembled this dataset as part of a research effort to understand how data composition affects retrieval model quality. Our experiments… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-pre-training.51 likes5.1k downloads1mo agoHugging Face09orionweller /mmBERT-pretraining-data-chunk1 mmBERT Training Data (Ready-to-Use) Complete Training Dataset: Pre-randomized and ready-to-use multilingual training data (3T tokens) for encoder model pre-training. This dataset is part of the complete, pre-shuffled training data used to train the mmBERT encoder models. Unlike the individual phase datasets, this version is ready for immediate use but the mixture cannot be modified easily. The data is provided in decompressed MDS format ready for use with ModernBERT's Composer… See the full description on the dataset page: https://huggingface.co/datasets/orionweller/mmBERT-pretraining-data-chunk1.fill-mask0 likes5k downloads1y agoHugging Face10lightonai /embeddings-pre-training-curated Embeddings pre-training curated data This dataset is the English subset of lightonai/embeddings-pre-training, assembled to reproduce the English data recipe described in the mGTE technical report (Zhang et al., 2024). The mGTE paper describes the data sources used to train the GTE family of multilingual text embedding and reranking models, but does not release the data itself. This dataset is our reconstruction of the English portion of that recipe, curated as part of a research… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-pre-training-curated.100M<n<1B15 likes4.2k downloads1mo agoHugging Face11nvidia /Nemotron-Pretraining-Code-v2gated Nemotron-Pre-Training-Dataset-v2.1 Dataset Description The Nemotron-Pre-Training-Dataset-v2.1 extends the previously released Nemotron pretraining datasets with refreshed, higher-quality, and more diverse data across math, code, English Common Crawl, and large-scale synthetic corpora. Designed for the NVIDIA Nemotron 3 family of LLMs, the dataset introduces new Common Crawl code extraction, 2.5T new English web tokens… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Code-v2.texttext-generation100M<n<1B135 likes4k downloads10mo agoHugging Face12mandiyaaman /nobody-pretraining-shards Nobody 3.2B — Cybersecurity MoE Pretraining Corpus (v1.0) Overview This repository hosts the official, pre-tokenized, and quality-filtered pretraining shards for Nobody 3.2B — Cybersecurity MoE, a foundation model specialized in defensive systems engineering, vulnerability analysis, reverse engineering, and threat intelligence. The corpus is constructed using a deterministic 10-stage CPU data pipeline and distributed across 12 permanent worker partitions.… See the full description on the dataset page: https://huggingface.co/datasets/mandiyaaman/nobody-pretraining-shards.tabulartext-generationn<1K0 likes2.9k downloads4d agoHugging Face13geodesic-research /control-pretraining-datasets-smoke geodesic-research/control-pretraining-datasets-smoke Auto-generated by dataset-builder. Each config below is a separate dataset produced from a versioned YAML build config. Load with: from datasets import load_dataset ds = load_dataset("geodesic-research/control-pretraining-datasets-smoke", "<config_name>", revision="<commit-sha>") Pin revision= to the specific commit SHA you want; without it, you get the current HEAD of the dataset repo, which may change when the builder… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/control-pretraining-datasets-smoke.text10K<n<100K0 likes2.7k downloads29d agoHugging Face14nvidia /Nemotron-Pretraining-Specialized-v1.2 Nemotron-Pretraining-Specialized-v1.2 Dataset Description: The Nemotron-Pretraining-Specialized-v1.2 dataset is part of the Nemotron Pretraining Data collection of pretraining datasets. Designed for the NVIDIA Nemotron 3 family of LLMs, this dataset contains a collection of synthetic datasets aimed to improve LLM capabilities on factual recall, moral scenarios, and diverse generative and multiple choice questions. Note: These are new datasets, not replacements.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Specialized-v1.2.texttext-generation100M<n<1B17 likes2.7k downloads4mo agoHugging Face15orionweller /mmBERT-pretraining-data-chunk0 mmBERT Training Data (Ready-to-Use) Complete Training Dataset: Pre-randomized and ready-to-use multilingual training data (3T tokens) for encoder model pre-training. This dataset is part of the complete, pre-shuffled training data used to train the mmBERT encoder models. Unlike the individual phase datasets, this version is ready for immediate use but the mixture cannot be modified easily. The data is provided in decompressed MDS format ready for use with ModernBERT's Composer… See the full description on the dataset page: https://huggingface.co/datasets/orionweller/mmBERT-pretraining-data-chunk0.fill-mask0 likes2.6k downloads1y agoHugging Face16nvidia /Nemotron-Pretraining-Code-v1gated Nemotron-Pre-Training-Dataset-v1 Release Data Overview This pretraining dataset, for generative AI model training, preserves high-value math and code while enriching it with diverse multilingual Q&A, fueling the next generation of intelligent, globally-capable models. This dataset supports NVIDIA Nemotron Nano 2, a family of large language models (LLMs) that consists of the NVIDIA-Nemotron-Nano-9B-v2, NVIDIA-Nemotron-Nano-9B-v2-Base, and NVIDIA-Nemotron-Nano-12B-v2-Base… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Code-v1.texttext-generation100M<n<1B82 likes2.6k downloads10mo agoHugging Face17HuggingFaceBio /carbon-pretraining-corpus 🧬 Carbon Pretraining Corpus Description 173M DNA & RNA sequences · 1.1 trillion nucleotides — the DNA pretraining mixture used to train Carbon, a genomic foundation model. This dataset is a collection of data sources intended for training genomic foundation models, such as Carbon. It contains DNA and RNA sequences spanning eukaryote and prokaryote species. Across the four main configs it totals 1.1 T DNA base pairs (180B tokens with Carbon's 6-mer tokenizer). A… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceBio/carbon-pretraining-corpus.tabulartext-generation100M<n<1B30 likes2.4k downloads4mo agoHugging Face18VuSnow /Jp-OCR-Pretrainingimageimage-to-text10K<n<100K2 likes2.4k downloads1y agoHugging Face19tingtang2 /the_stack_v2_python_repos_pretraining_dataset_imported_context-datasettext1M<n<10M0 likes2.2k downloads1y agoHugging Face20nvidia /Nemotron-Pretraining-Specialized-v1.1 Nemotron-Pretraining-Specialized-v1.1 Dataset Description: The Nemotron-Pretraining-Specialized-v1.1 dataset is part of the Nemotron Pretraining Data collection of pretraining datasets. Designed for the NVIDIA Nemotron 3 family of LLMs, this dataset contains a collection of synthetic datasets aimed to improve LLM capabilities in code concepts and algorithms, formal logic, economics, and multiple choice questions. The code concepts dataset is an instance of a general… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Specialized-v1.1.texttext-generation10M<n<100M46 likes2.2k downloads7mo agoHugging Face21nvidia /Nemotron-Pretraining-SFT-v1gated Nemotron-Pre-Training-Dataset-v1 Release Data Overview This pretraining dataset, for generative AI model training, preserves high-value math and code while enriching it with diverse multilingual Q&A, fueling the next generation of intelligent, globally-capable models. This dataset supports NVIDIA Nemotron Nano 2, a family of large language models (LLMs) that consists of the NVIDIA-Nemotron-Nano-9B-v2, NVIDIA-Nemotron-Nano-9B-v2-Base, and NVIDIA-Nemotron-Nano-12B-v2-Base… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-SFT-v1.texttext-generation100M<n<1B74 likes2k downloads10mo agoHugging Face22orionweller /mmBERT-pretraining-data-chunk2 mmBERT Training Data (Ready-to-Use) Complete Training Dataset: Pre-randomized and ready-to-use multilingual training data (3T tokens) for encoder model pre-training. This dataset is part of the complete, pre-shuffled training data used to train the mmBERT encoder models. Unlike the individual phase datasets, this version is ready for immediate use but the mixture cannot be modified easily. The data is provided in decompressed MDS format ready for use with ModernBERT's Composer… See the full description on the dataset page: https://huggingface.co/datasets/orionweller/mmBERT-pretraining-data-chunk2.fill-mask0 likes1.9k downloads1y agoHugging Face23aslawliet /math-pretraining-corpustext10M<n<100M4 likes1.9k downloads3y agoHugging Face24Ronaldo-GOAT /audio_pretraining_sono_synthetic FORGE3D SLAT training data (flat layout) Two directories only: tars/ — every tar, flat. docs/ — manifests, CSVs, notes (former nested paths flattened with / → __). docs/MANIFEST.csv — one row per object: id, tar, has_slat, has_white_cond, has_recgen_image, recgen_image_tar, recgen_views, is_recgen, ... Use it to select what to download. What is in tars/ tars content used for SLAT training shard_*.tar, inc_*.tar (not inc_slat2*/inc_rgslat*)… See the full description on the dataset page: https://huggingface.co/datasets/Ronaldo-GOAT/audio_pretraining_sono_synthetic.0 likes1.8k downloads6d agoHugging Face25rooty2020 /text_audio_pretraining_for_GAN bw_jh_dataset — Robot 3D Point-Track Dataset (droid / hrdexdb / robocasa) A unified multi-source robot manipulation dataset with 3D point tracks of the robot gripper and arm, calibrated multi-view RGB video, language instructions, and success labels. Three splits: Split Episodes Size Views Source droid 27,615 (+~1,050 in droid-11) ~280 GB 2 exterior ZED views DROID (real, Franka) hrdexdb 1,601 ~998 GB 22 calibrated cameras HRDexDB (real, xArm6 + dexterous hands)… See the full description on the dataset page: https://huggingface.co/datasets/rooty2020/text_audio_pretraining_for_GAN.videorobotics1K<n<10K1 likes1.8k downloads14d agoHugging Face26VuSnow /Vn-OCR-Pretrainingimageimage-to-text10K<n<100K1 likes1.6k downloads1y agoHugging Face27simple-pretraining /wikipedia_chunked Dataset Card for "wikipedia_chunked" More Information needed text10M<n<100M2 likes1.4k downloads3y agoHugging Face28CrowdMind /nemotron-pretraining-specialized-collection Nemotron Pretraining Specialized Collection This repository is a convenience collection of the NVIDIA Nemotron Pretraining Specialized releases. It preserves each original subset as a separate Hugging Face configuration, so consumers can select a single domain-focused subset without visiting multiple source repositories. Contents Source release Configurations included nvidia/Nemotron-Pretraining-Specialized-v1 Wiki Rewrite, Math Textbooks, STEM SFT… See the full description on the dataset page: https://huggingface.co/datasets/CrowdMind/nemotron-pretraining-specialized-collection.texttext-generation100M<n<1B0 likes1.3k downloads11d agoHugging Face29JuaAI /ts-icl-pretraining-corpus TS-ICL Pretraining Corpus (community reconstruction) A unified, cleaned reconstruction of the univariate pretraining corpus described in Table 5 of TS-ICL: A Flexible Time-Indexed Foundation Model for Time Series via In-Context Learning (Le Naour, Nabil & Petralia, EDF R&D; arXiv:2606.05878). The TS-ICL authors did not release their pretraining data pipeline, so this corpus is rebuilt from the named upstream sources (LOTSA, Chronos, and the TempoPFN synthetic generators) and… See the full description on the dataset page: https://huggingface.co/datasets/JuaAI/ts-icl-pretraining-corpus.tabulartime-series-forecasting1M<n<10M0 likes1.3k downloads3mo agoHugging Face30nvidia /Nemotron-Pretraining-Code-v3 Nemotron-Pretraining-Code-v3 Dataset Description: The Nemotron-Pretraining-Code-v3 dataset is part of the Nemotron Pretraining Data collection of pretraining datasets. Designed for the NVIDIA Nemotron 3 family of LLMs, this dataset is intended to improve the coding capabilities of LLMs. The Nemotron-Pretraining-Code-v3 dataset contains the metadata corresponding to the raw source-code update to our Nemotron-Pretraining-Code-v2 and Nemotron-Pretraining-Code-v1… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Code-v3.texttext-generation100M<n<1B73 likes1.3k downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.