datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
us-stock-dataleaderboard-dataset
Arena Leaderboard Dataset
Historical snapshots of the Arena leaderboard.
Usage
from datasets import load_dataset
# Load all historical text style control data
ds = load_dataset("lmarena-ai/leaderboard-dataset", "text_style_control", split="full")
# Load the current text style control leaderboard
ds = load_dataset("lmarena-ai/leaderboard-dataset", "text_style_control", split="latest")
# Filter to overall category
ds =… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset.fev_datasets
Forecast evaluation datasets
This repository contains time series datasets that can be used for evaluation of univariate & multivariate forecasting models.
The main focus of this repository is on datasets that reflect real-world forecasting scenarios, such as those involving covariates, missing values, and other practical complexities.
The datasets follow a format that is compatible with the fev package.
Data format and usage
Each dataset satisfies the following… See the full description on the dataset page: https://huggingface.co/datasets/autogluon/fev_datasets.Hy-Embodied-0.5-VLA-Data
Hy-Embodied-0.5-VLA
From Vision-Language-Action Models to a Real-World Robot Learning Stack
Tencent Robotics X × Tencent Hy Team
📖 Abstract
We introduce Hy-Embodied-0.5-VLA (Hy-VLA) — an end-to-end Vision-Language-Action system that spans the full robot learning stack: data collection, model design, pre-training, supervised fine-tuning, RL post-training, and real-world deployment. Built on the Hy-Embodied-0.5 MoT backbone, Hy-VLA integrates a flow-matching… See the full description on the dataset page: https://huggingface.co/datasets/tencent/Hy-Embodied-0.5-VLA-Data.molmobot-data
MolmoBot-data
Training episode data (actions, visual inputs, and other sensor data) for 8 tasks on 2 robotic platforms:
DoorOpeningDataGenConfig
RBY1OpenDataGenConfig
RBY1PickDataGenConfig
FrankaPickOmniCamConfig
RBY1PickAndPlaceDataGenConfig
FrankaPickAndPlaceOmniCamConfig
FrankaPickAndPlaceColorOmniCamConfig
FrankaPickAndPlaceNextToOmniCamConfig
Please note that every package indexed by the parquet files can contain several instances of episode data.
We also provide an… See the full description on the dataset page: https://huggingface.co/datasets/allenai/molmobot-data.MolmoAct2-BimanualYAM-DatasetThis dataset was created using LeRobot.
MolmoAct2-BimanualYAM Dataset
This repository is the merged ckpt / merged LeRobot dataset artifact for the MolmoAct2-BimanualYAM Dataset, a large-scale collection of bimanual robot manipulation demonstrations collected for MolmoAct2. Across the full collection, MolmoAct2-BimanualYAM contains more than 720 hours of training demonstrations spanning diverse tabletop manipulation tasks.
Language Annotations
This dataset includes… See the full description on the dataset page: https://huggingface.co/datasets/allenai/MolmoAct2-BimanualYAM-Dataset.datacomp_xlarge
DataComp XLarge Pool
This repository contains metadata files for the xlarge pool of DataComp. For details on how to use the metadata, please visit our website and our github repository.
We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights.
Terms and Conditions
We have terms of service that are similar to those adopted by HuggingFace… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_xlarge.review-data-b
Review Dataset B
Anonymous review release of synthetic driving data collected in CARLA 0.9.15
with the PDM-Lite expert.
5,246 quality-filtered recordings, approximately 221 GB.
39 scenario types across Towns 01–05, Town10HD, Town12, and Town13.
768 × 768 DVS events, represented in 50 ms windows.
Includes ego-motion metadata and BEV occupancy labels at 0, +0.5, and +1.0 seconds.
Split
Recordings
Train
4,132
Validation
538
Test
576
Files are organized under… See the full description on the dataset page: https://huggingface.co/datasets/anon-data-7k4m/review-data-b.trex_dataset
T-Rex Dataset
A large-scale, tactile-reactive bimanual manipulation dataset, collected via teleoperation on a
Dexmate Vega-1 robot with two Sharpa Wave dexterous hands. Stored as a
LeRobotDataset v3.0.
🌐 Project Page · ✍️ Paper (arXiv) · 💻 Code (T-Rex) · 🚀 Dataset Quickstart · 📓 Colab notebook
One episode from each of 20 motor primitives (head-camera view, cropped to the workspace), each with a different object.
Teleoperation setup: Manus gloves + VIVE… See the full description on the dataset page: https://huggingface.co/datasets/zekaiwang/trex_dataset.zendo-synthetic-data
Zendo Synthetic Visual Reasoning Dataset
Synthetic Zendo-style scenes with associated rules and per-scene tensor
representations. Each scene either follows ("positive", label=1) or violates
("negative", label=0) a rule that is given in natural language and as a Prolog
query.
Splits
split
scenes
train
56475
test
3344
rules total
3439
Layout
images/<split>/<batch>/<rule_id>/<scene_id>.png — rendered scene… See the full description on the dataset page: https://huggingface.co/datasets/sophia1ch/zendo-synthetic-data.PPTAgent-parsed_datachronos_datasets
Chronos datasets
Time series datasets used for training and evaluation of the Chronos forecasting models.
Note that some Chronos datasets (ETTh, ETTm, brazilian_cities_temperature and spanish_energy_and_weather) that rely on a custom builder script are available in the companion repo autogluon/chronos_datasets_extra.
See the paper for more information.
Data format and usage
The recommended way to use these datasets is via https://github.com/autogluon/fev.
All datasets… See the full description on the dataset page: https://huggingface.co/datasets/autogluon/chronos_datasets.Polymarket_data
Polymarket Data
Complete Data Infrastructure for Polymarket — Fetch, Process, Analyze
A comprehensive dataset of 6.6 billion on-chain trading records from Polymarket, processed into multiple analysis-ready formats. Features cleaned data, unified token perspectives, and user-level transformations — ready for market research, behavioral studies, and quantitative analysis.
Zhengjie Wang1,2, Leiyu Chao1,3, Yu Bao1,4, Lian Cheng1,3, Jianhan Liao1,5, Yikang Li1,†
1Shanghai Innovation… See the full description on the dataset page: https://huggingface.co/datasets/SII-WANGZJ/Polymarket_data.product-database
Open Food Facts Database
What is 🍊 Open Food Facts?
A food products database
Open Food Facts is a database of food products with ingredients, allergens, nutrition facts and all the tidbits of information we can find on product labels.
Made by everyone
Open Food Facts is a non-profit association of volunteers. 25.000+ contributors like you have added 1.7 million + products from 150 countries using our Android or iPhone app or their… See the full description on the dataset page: https://huggingface.co/datasets/openfoodfacts/product-database.show3d-dataset
SHOW3D: Capturing Scenes of 3D Hands and Objects in the Wild (CVPR 2026)
📢 News
October 3, 2026: Hand pose v3 release: Significantly reduced joint angle limit violations, and increased valid frames. Now available in UmeTrack, MANO, and MHR. See hand_pose/README.md.
September 26, 2026: Updated 54 scenes in hand_pose/v2: re-tracked three with the correct hand profile, and in the other 51, kept hands that align with the images on frames with a repaired headset… See the full description on the dataset page: https://huggingface.co/datasets/facebook/show3d-dataset.opencs2_dataset
OpenCS2 - POV Renders
Browse with the OpenCS2 Viewer - every match, map and round, with all 10 player POVs synced on one timeline.
Tick-aligned Counter-Strike 2 POV training clips, rendered from
blanchon/cs2_dataset_demo. Each row
in the main table is one player's perspective for one round; ten POVs per round share the same tick
clock.
Per POV round:
Video - 1280x720 @ 32 fps, near-lossless H.264, faststart, muxed with audio.
Audio - per-player stereo, mixed from that… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/opencs2_dataset.drlc-leaderboard-datasoc-ratchakitcha
Royal Gazette Thailand (Ratchakitcha) Dataset
ชุดข้อมูลราชกิจจานุเบกษา (แบบ Machine Readable)
โครงการ Open Law Data Thailand ร่วมกับคณะกรรมาธิการการพาณิชย์และการอุตสาหกรรม วุฒิสภา ได้รับความอนุเคราะห์ข้อมูลจาก สำนักเลขาธิการคณะรัฐมนตรี (สลค.) เพื่อเผยแพร่ข้อมูลกฎหมายไทยสู่สาธารณะในรูปแบบที่ประมวลผลได้ด้วยคอมพิวเตอร์ (Machine Readable) เพื่อส่งเสริมนวัตกรรม Legal Tech และ AI ของประเทศไทย
Dataset Description
ชุดข้อมูลนี้รวบรวมรายการประกาศในราชกิจจานุเบกษา… See the full description on the dataset page: https://huggingface.co/datasets/open-law-data-thailand/soc-ratchakitcha.datacomp200m
Datacomp200m
This is a smaller version of the datacomp_1b dataset.
Filtering was done by taking all rows that had self similarity (inner product) above 0.32. This resulted in 213009083 (213 million) rows.
The results of the datacomp paper suggest that filtering by CLIP score is better than random sampling.
Included in this repo are search indices created using autofaiss, over the text and image embeddings. There are two ways to access metadata, either in .parquet files in the… See the full description on the dataset page: https://huggingface.co/datasets/adams-story/datacomp200m.po_qwen14b_tabular_data
BoLT Prompt Optimization — Tabular Dataset
For prompt optimization tasks in BoLT, an accessible benchmark for black-box optimization on LLM tasks.
Dataset Description
The dataset covers 5,014 evaluated instructions. Each row is a candidate system-prompt instruction paired with its empirically measured MATH-500 (4-shot, non-thinking mode) scores.
Evaluation details:
Model: Qwen/Qwen3-14B
Task: minerva_math500 (4-shot) (from lm-eval library)
System prompt:… See the full description on the dataset page: https://huggingface.co/datasets/chewwt/po_qwen14b_tabular_data.kalshi-15m-metals-datallm-network-study-data
LLM-Network-Study-Data
Per-request network captures (.pcapng) collected by the
LLM-Network-Study benchmark harness (benchmark.py and the
per-workload test scripts). Each directory holds one capture file per request,
named request_<id>_run<n>_<timestamp>.pcapng.
A directory name encodes four dimensions:
<capture-env>_<provider/model>_<workload>[_<dataset/variant>]_results
Dimension legend
Dimension
Values
Meaning
Capture env
ethernet
Wired connection to… See the full description on the dataset page: https://huggingface.co/datasets/wayslab/llm-network-study-data.go_emotions
Dataset Card for GoEmotions
Dataset Summary
The GoEmotions dataset contains 58k carefully curated Reddit comments labeled for 27 emotion categories or Neutral.
The raw data is included as well as the smaller, simplified version of the dataset with predefined train/val/test
splits.
Supported Tasks and Leaderboards
This dataset is intended for multi-class, multi-label emotion classification.
Languages
The data is in English.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/go_emotions.datacomp_1b
DataComp-1B
This repository contains metadata files for DataComp-1B. For details on how to use the metadata, please visit our website and our github repository.
We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights.
Terms and Conditions
We have terms of service that are similar to those adopted by HuggingFace (https://huggingface.co/terms-of-service), which covers… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_1b.CADS-dataset
CADS: A Comprehensive Anatomical Dataset and Segmentation for Whole-Body Anatomy in Computed Tomography
Overview
CADS is a robust, fully automated framework for segmenting 167 anatomical structures in Computed Tomography (CT), spanning from head to knee regions across diverse anatomical systems.
The framework consists of two main components:
CADS-dataset:
22,022 CT volumes with complete annotations for 167 anatomical structures.
Most extensive whole-body CT dataset… See the full description on the dataset page: https://huggingface.co/datasets/huggingface/CADS-dataset.Pumpfun_v2_datasetdeutsche-bahn-data
Deutsche Bahn Train Data
This dataset contains public historical data from Deutsche Bahn, the largest German train company. It includes train schedules, delays, and cancellations from stations across Germany.
For more info visit the project page at GitHub: https://github.com/piebro/deutsche-bahn-data
Dataset Structure
Monthly Processed Data
The monthly processed data is located in monthly_processed_data/ and contains files named… See the full description on the dataset page: https://huggingface.co/datasets/piebro/deutsche-bahn-data.toxigen-data
Dataset Card for ToxiGen
Sign up for Data Access
To access ToxiGen, first fill out this form.
Dataset Summary
This dataset is for implicit hate speech detection. All instances were generated using GPT-3 and the methods described in our paper.
Languages
All text is written in English.
Dataset Structure
Data Fields
We release TOXIGEN as a dataframe with the following fields:
prompt is the prompt used for generation.
generation is… See the full description on the dataset page: https://huggingface.co/datasets/toxigen/toxigen-data.icl-dataset
ICL Dataset
Bimanual robot manipulation data collected by human teleoperation on the YOR robot,
in LeRobot v2.1 format.
3,149 episodes · 2,385,871 frames · 22.1 hours · 36 tasks · 3 camera views
Each episode is a single teleoperated attempt at a manipulation task, recorded at 30 fps
with synchronised proprioception, commanded actions and three camera streams.
⚠️ Read this before training on it
This dataset contains failures, trimmed successes and unusable… See the full description on the dataset page: https://huggingface.co/datasets/adityx23/icl-dataset.Company-document-dataset-v2
Company Documents v2
Generation complete: all 13 document types have completed export and upload checkpoints.
Synthetic, born-digital business documents rendered from four open sample databases, with exact gold
labels: 353,580 PDFs (404,514 pages) of 13 document types in
English and French, issued by 60 synthetic companies,
each with its own letterhead, numbering and wording. Successor of
CompanyDocuments (2,677 PDFs, 4 types).
Dataset overview
property… See the full description on the dataset page: https://huggingface.co/datasets/AyoubChLin/Company-document-dataset-v2.
