datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Emilia-Dataset
Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
This is the official repository 👑 for the Emilia dataset and the source code for the Emilia-Pipe speech data preprocessing pipeline.
News 🔥
2025/02/26: The Emilia-Large dataset, featuring over 200,000 hours of data, is now available!!! Emilia-Large combines the original 101k-hour Emilia dataset (licensed under CC BY-NC 4.0) with the brand-new 114k-hour Emilia-YODAS… See the full description on the dataset page: https://huggingface.co/datasets/amphion/Emilia-Dataset.Vchitect_T2V_DataVerse
Vchitect-T2V-Dataverse
Vchitect Team1
1Shanghai Artificial Intelligence Laboratory
Paper |
Project Page |
Data Overview
The Vchitect-T2V-Dataverse is the core dataset used to train our text-to-video diffusion model, Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models.
It comprises 14 million high-quality videos collected from the Internet, each paired with detailed textual… See the full description on the dataset page: https://huggingface.co/datasets/Vchitect/Vchitect_T2V_DataVerse.FTP-1-Dataset
FTP-1-Dataset
FTP-1-Dataset contains heterogeneous tactile manipulation data for FTP-1 pretraining.
This release currently includes 18 dataset archives:
FreeTacMan
MotionTrans
RDP
RDP_Bimanual
RH20TCfg5Franka
RH20TCfg6ATIAxia
RH20TCfg7Tactile
Unit
Unit_Bimanual
VLA_touch
ViTaMIn
VisuoTactile_D-WHEEL
VisuoTactile_QINGLOONG
exUMI
sharpa
Each dataset directory contains either a single <dataset>.tar file or split parts named <dataset>.tar.part-*. For split archives, concatenate… See the full description on the dataset page: https://huggingface.co/datasets/MJJJJ1064/FTP-1-Dataset.scientific-stuff-1robopoint-data
RoboPoint Dataset Card
Dataset details
This dataset contains 1432K image-QA instances used to fine-tune RoboPoint, a VLM for spatial affordance prediction. It consists of the following parts:
347K object reference instances from a synthetic data pipeline;
320K free space reference instances from a synthetic data pipeline;
100K object detection instaces from LVIS;
150K GPT-generated instruction-following instances from liuhaotian/LLaVA-Instruct-150K;
515K general-purpose… See the full description on the dataset page: https://huggingface.co/datasets/wentao-yuan/robopoint-data.medpmc-11m-dataset_jun24_baseline
MedPMC WebDataset
MedPMC is a large-scale medical image-text dataset curated from articles in the PubMed Central (PMC) collection. This release contains approximately 11 million image-text pairs collected from the June 2024 PMC baseline. MedPMC is an ongoing effort, and future releases will continue to expand the dataset with newly published literature, improved annotations, and additional resources.
This dataset is presented in the paper MedPMC: A Systematic Framework for… See the full description on the dataset page: https://huggingface.co/datasets/Yale-BIDS-Chen/medpmc-11m-dataset_jun24_baseline.describe-anything-dataset
Describe Anything: Detailed Localized Image and Video Captioning
NVIDIA, UC Berkeley, UCSF
Long Lian, Yifan Ding, Yunhao Ge, Sifei Liu, Hanzi Mao, Boyi Li, Marco Pavone, Ming-Yu Liu, Trevor Darrell, Adam Yala, Yin Cui
[Paper] | [Code] | [Project Page] | [Video] | [HuggingFace Demo] | [Model/Benchmark/Datasets] | [Citation]
Dataset Card for Describe Anything Datasets
Datasets used in the training of describe anything models (DAM).
The datasets are in tar files. These… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/describe-anything-dataset.olmoearth_pretrain_datasetThis is the pre-training dataset for training the OlmoEarth pre-trained remote sensing foundation models.
Documentation is on GitHub at https://github.com/allenai/olmoearth_pretrain/blob/main/docs/Pretraining-Dataset.md
The dataset is released under CC BY 4.0. It includes data from the following sources:
Sentinel-2 L2A imagery from the European Space Agency, available under the Copernicus Sentinel Data and Service Legal Notice
Sentinel-1 GRD IW vv+vh imagery from the European Space Agency… See the full description on the dataset page: https://huggingface.co/datasets/allenai/olmoearth_pretrain_dataset.ReCo-Data
ReCo-Data Dataset Card
Introduction
ReCo-Data is a large-scale, high-quality video editing dataset comprising 500K+ instruction-video pairs. This card provides its statistics, collection pipeline, and dataset format.
1. Dataset Statistics
Statistics
Figure Caption:
(a) Overview of scale
(b) Task distribution showing balanced quantities: Replace (156.6K), Style (130.6K), Remove (121.6K), and Add (115.6K). Human evaluation on 200 randomly… See the full description on the dataset page: https://huggingface.co/datasets/HiDream-ai/ReCo-Data.raw_primitive_datasets
Datacard
This is the official fine-tuning dataset provided by VLABench (raw data), with 500 episodes each task. The current version includes 10 primitive tasks.
Source
Project Page: https://vlabench.github.io/
Arxiv Paper: https://arxiv.org/abs/2412.18194
Code: https://github.com/OpenMOSS/VLABench
Uses
Download all archive files and use the following command to extract:
cat vlabench_primitive.tar.gz.* | tar -xzvf -
In the resulting… See the full description on the dataset page: https://huggingface.co/datasets/VLABench/raw_primitive_datasets.BridgeVLA_COLOSSEUM_EVAL_DATAarxiv: https://arxiv.org/abs/2506.07961
voice-data
Voice-Data: a curated multi-corpus voice dataset for voice–text contrastive training
voice-data is a single, globally-shuffled WebDataset that bundles several voice/speech corpora into one ready-to-train mixture for voice–text contrastive (CLAP-style) models such as VoiceCLAP. Each clip pairs 48 kHz mono FLAC audio with a natural-language text caption describing the voice — its emotion, prosody, timbre, speaking style, recording context, and speaker traits.
The distinguishing… See the full description on the dataset page: https://huggingface.co/datasets/gijs/voice-data.majestrino-dataX2Edit-Dataset
X2Edit
Introduction
X2Edit Dataset is a comprehensive image editing dataset that covers 14 diverse editing tasks and exhibits substantial advantages over existing open-source datasets including AnyEdit, HQ-Edit, UltraEdit, SEED-Data-Edit, ImgEdit and OmniEdit.
For the relevant data construction scripts, model training and inference scripts, please refer to X2Edit.
News
2025/09/16: We are about to release a dataset constructed by Qwen-Image and… See the full description on the dataset page: https://huggingface.co/datasets/OPPOer/X2Edit-Dataset.Ola-DataThis repository contains the data presented in Ola: Pushing the Frontiers of Omni-Modal Language Model with Progressive Modality Alignment.
Code: https://github.com/Ola-Omni/Ola
DataCompDR-12M-bf16
Dataset Card for DataCompDR-12M-BFloat16
This dataset contains synthetic captions, embeddings, and metadata for DataCompDR-12M.
The metadata has been generated using pretrained image-text models on a 12M subset of DataComp-1B.
For details on how to use the metadata, please visit our github repository.
The dataset with the original captions is now available at mlfoundations/DataComp-12M.
The UIDs per shards match between mlfoundations/DataComp-12M and apple/DataCompDR-12M-bf16.… See the full description on the dataset page: https://huggingface.co/datasets/apple/DataCompDR-12M-bf16.voiceclap-data
VoiceCLAP Data
The audio + dense-caption mixture used to train
laion/voiceclap-small and
laion/voiceclap-large.
Each tar shard is a WebDataset of
paired <key>.flac (48 kHz mono audio) + <key>.json (caption + metadata)
samples. Captions and structured attribute annotations are produced
automatically by a pipeline of audio-aware LLMs — Qwen-Audio, Gemini Flash 2.5,
and a thinking-mode reasoning model that scores emotion under the EmoNet
taxonomy plus per-clip vocal-burst, timbre… See the full description on the dataset page: https://huggingface.co/datasets/laion/voiceclap-data.Harmonizer-Dataset
HARMONIZER DATASET
Dataset Description
Training dataset for DiffusionHarmonizer: a generative AI model for image and video enhancement bridging neural reconstruction and photorealistic simulation .
Model checkpoints: https://huggingface.co/nvidia/Harmonizer/Training code: https://github.com/NVIDIA/harmonizer/
The dataset was curated to support the following functions of the model:
3D reconstruction artifact removal
Harmonization of inserted objects to blend… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Harmonizer-Dataset.X-Voice-Dataset-Train
X-Voice Training Dataset
Overview
The X-Voice training dataset is a large-scale multilingual speech corpus curated for high-performance speech models. It provides a robust foundation for cross-lingual phonetic and prosodic modeling.
Also the train set of X-Voice Model.
Core Statistics
Total Speech Duration: 420K hours
30 languages
European: bg (Bulgarian), cs (Czech), da (Danish), de (German), el (Greek), en (English), es (Spanish), et (Estonian), fi… See the full description on the dataset page: https://huggingface.co/datasets/XRXRX/X-Voice-Dataset-Train.Uni-Edit-Train-Data
Uni-Edit Training Data: Uni-Edit-148k
Project Page | GitHub Repository | Paper
👀 Intro
We introduce Uni-Edit, an intelligent image editing task that serves as the first general task for Unified Multimodal Model (UMM) tuning. Unlike conventional mixed multi-task training that suffers from inherent task conflicts and requires complex multi-stage pipelines, Uni-Edit breaks this paradigm. It achieves true mutual reinforcement by improving image… See the full description on the dataset page: https://huggingface.co/datasets/Uni-Edit/Uni-Edit-Train-Data.hf_dataplotqa-dataset
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Dodon/plotqa-dataset.DF_DiFF_FAS_dataset_in_FSFM_FSVFM
FSFM / FS-VFM Downstream Datasets
Processed downstream fine-tuning datasets for cross-dataset deepfake detection, cross-domain face anti-spoofing, and unseen diffusion-generated face detection used with FSFM and FS-VFM.
This archive is intended for research use with the FSFM/FS-VFM release scripts.
RaCig-DataHO-Cap-DatasetSTRIDE-QA-Dataset
STRIDE-QA Dataset
📦 Dataset
STRIDE-QA is a large-scale visual question answering (VQA) dataset for physically grounded spatiotemporal reasoning in autonomous driving. Constructed from 100 hours of multi-sensor driving data in Tokyo, it offers 16 M QA pairs over 270 K frames with dense annotations including 3D bounding boxes, segmentation masks, and multi-object tracks.
Category
Description
Object-centric Spatial QA
Spatial relations between two… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/STRIDE-QA-Dataset.Taiwan-Tongues-ASR-CE-dataset-hokkien
Taiwan-Tongues-ASR-CE-dataset-hokkien
本資料集為 Taiwan-Tongues-ASR-CE 專案所使用的預訓練資料,透過 WebDataset 格式打包,並上傳至 Hugging Face 以便研究人員與開發者自由取用。
📂 Dataset 結構
本資料集分為 Training 與 Test 兩個子集,均以 WebDataset tar 檔案形式存放:
Training set (WebDataset format)
train/train-000000.tar
train/train-000001.tar
...
Test set (WebDataset format)
test/test-000000.tar
...
tsv set
train.tsv
test.tsv
...
每個 tar 內部均包含對應的音檔與標註,方便直接搭配 WebDataset 與 PyTorch / Hugging Face datasets 進行訓練與測試。… See the full description on the dataset page: https://huggingface.co/datasets/adi-gov-tw/Taiwan-Tongues-ASR-CE-dataset-hokkien.scientific-stuff-2IXI-DatasetsPhysics-aware-videos中文 README
🧱 Physics-aware Video Dataset
This dataset is a high-quality real-world video dataset focused on physical phenomena, designed for learning and evaluating physical laws from videos. It primarily covers the following classic physical processes:
🧱 Rigid-body motion / 🌊 Fluid dynamics
💫 Collision and rebound
💥 Explosion and burst phenomena
🌫️ Smoke, dust, and particle scattering
🌍 Gravity and inertia (falling, rolling, acceleration, etc.)
The dataset is constructed… See the full description on the dataset page: https://huggingface.co/datasets/BAAI-DataCube/Physics-aware-videos.
