datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
humanoid-robots-training-dataset
Dynamic Intelligence — Humanoid Robot Training Dataset
A first-person (egocentric) video dataset of human hand manipulation, designed for training humanoid robot policies via imitation learning. Each episode captures a person performing an everyday household task — folding clothes, moving dishes, opening doors — filmed from a head-mounted iPhone using its built-in LiDAR and depth sensors.
The dataset pairs each video with frame-level 3D hand tracking and camera pose data, giving… See the full description on the dataset page: https://huggingface.co/datasets/DynamicIntelligence/humanoid-robots-training-dataset.Elastic-Forcing-training-dataset
Elastic-Forcing training datasets
wan-1.3B-dataset/: 8,682 original videos and paired captions used by experiment 10351.
Each dataset folder contains its own videos, caption metadata, training manifest, and provenance. Original training data are kept separate across model scales.
wan-14B-dataset/: 5,546 retained videos and paired captions from the 14B step80 training dataset (originally 8,000; 2,454 subsequently removed and excluded).
Linny-Training-Dataset-SyntheticVLA_Reasoning_Training_Dataset_1200_Trimmed_Start_5_Frame
Spa-Bench fine-tuning demonstrations — motion-trimmed release
This is a non-destructive, motion-trimmed derivative of the
canonical 1,200-episode Spa-Bench dataset.
It removes initial idle prefixes while preserving episode identity, prompt,
action/state alignment, and all five source camera streams at the public head.
Explore episodes in the LeRobot visualizer
Transformation
The baseline is the component-wise median of the first five action frames.
Motion onset… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/VLA_Reasoning_Training_Dataset_1200_Trimmed_Start_5_Frame.SAND-Post-Training-Dataset
SAND-Post-Training-Dataset: High-Quality Synthetic Reasoning Dataset Built with AMD GPUs
Dataset Summary
We introduce the SAND-Post-Training-Dataset, a high-quality synthetic reasoning dataset for mathematics and science built entirely using a synthetic data pipeline running on the AMD ROCm™ stack and AMD Instinct™ MI325 GPUs.
This dataset prioritizes difficulty and novelty over volume, demonstrating that high-difficulty synthetic data can elevate… See the full description on the dataset page: https://huggingface.co/datasets/amd/SAND-Post-Training-Dataset.Lithology-Training-Dataset
Lithology Training Dataset
A supervised training dataset for machine learning and AI systems that learn to identify lithology from well-log data.
The dataset contains 400 wells with standardized wireline-log measurements and corresponding lithological labels.
Training dataset:
https://huggingface.co/datasets/NoraResearchLab/Lithology-Training-Dataset
Overview
The core task is:
Given a sequence of well-log measurements across depth, predict the lithology… See the full description on the dataset page: https://huggingface.co/datasets/NoraResearchLab/Lithology-Training-Dataset.VLA_Reasoning_Training_Dataset_1200
Spa-Bench fine-tuning demonstrations — full five-camera release
This is the canonical full-length demonstration dataset used by Spa-Bench, a
real-robot benchmark for spatially grounded reasoning in vision-language-action
policies.
Explore episodes in the LeRobot visualizer
Dataset summary
Field
Value
Episodes
1,200
Frames
612,733
Duration at 30 FPS
approximately 5.7 hours
Unique instruction strings
321
Task families
6; 200 demonstrations per… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/VLA_Reasoning_Training_Dataset_1200.Training_dataset_qsearched_by_NAGISA_V4
NAGISA_V4 teacher shards, moved to their quiescence leaves
Every record of
Training_dataset_by_NAGISA_V4
walked to the end of its quiescence variation, the deep search's value kept
there, and a policy fitted at the leaf itself.
The parent's positions are as its games reached them, with no quiescence
search — a row can sit in the middle of an exchange, where the evaluation
swings by a piece depending on whose turn it is to recapture. A value fitted on
those learns the swing. This… See the full description on the dataset page: https://huggingface.co/datasets/qleap/Training_dataset_qsearched_by_NAGISA_V4.VLA_Reasoning_Training_Dataset_1200_2cam
Spa-Bench fine-tuning demonstrations — full two-camera derivative
This is the two-camera projection of the canonical full-length Spa-Bench
demonstration dataset. It retains the middle and wrist views consumed by the
evaluated policies and omits the unused above, left, and right streams.
Explore episodes in the LeRobot visualizer
Dataset summary
Field
Value
Episodes
1,200
Frames
612,733
Unique instruction strings
321
Frame rate
30 FPS
Camera… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/VLA_Reasoning_Training_Dataset_1200_2cam.Training_dataset_by_NAGISA_V4
NAGISA_V4 ply-37 teacher shards, games played to the end
Self-play of attic-gensfen reading the NNUE weights NAGISA_V4
(HalfKA-2304), in the shape the trainers read directly. Every game starts from a
balanced ply-37 position, makes no random moves, and runs until it actually
ends. Identical positions are folded into one row each.
67,108,864 rows — exactly 2^26
16 shards of 4,194,304 rows, 512 row groups each, zstd, 5,208,620,326 B total
Against
Opening_dataset_by_NAGISA_V4… See the full description on the dataset page: https://huggingface.co/datasets/qleap/Training_dataset_by_NAGISA_V4.training_datasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/sagnikroy75/training_dataset.TrainingDatasethumanoid-robots-training-dataset
Dynamic Intelligence — Humanoid Robot Training Dataset
A first-person (egocentric) video dataset of human hand manipulation, designed for training humanoid robot policies via imitation learning. Each episode captures a person performing an everyday household task — folding clothes, moving dishes, opening doors — filmed from a head-mounted iPhone using its built-in LiDAR and depth sensors.
The dataset pairs each video with frame-level 3D hand tracking and camera pose data, giving… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz043/humanoid-robots-training-dataset.credit-scoring-training-datasetThe training dataset includes all addresses that had undertaken at least one borrow transaction on Aave v2 Ethereum or Compound v2 Ethereum any time between 7 May 2019 and 31 August 2023, inclusive (called the observation window).
Data Structure & Shape
There are almost 0.5 million observations with each representing a single borrow event. Therefore, all feature values are calculated as at the timestamp of a borrow event and represent the cumulative positions just before the borrow event's… See the full description on the dataset page: https://huggingface.co/datasets/spectrallabs/credit-scoring-training-dataset.new_training_dataset_4kpost-training-trackio-datasetdataset_trainingThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 50,
"total_frames": 28255,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aimandanial/dataset_training.chess_puzzle_training_datasets_lt-2400
Chess puzzle training datasets: rating below 2400
This is a filtered derivative of
pavelslab-nyu/chess_puzzle_training_datasets.
Every retained row satisfies the exact condition:
Rating < 2400
Rating is the Lichess puzzle rating, not the Elo of either player in the
source game. The original column names, column order, directory layout, and CSV
schemas are preserved. As in the upstream repository, Hugging Face discovers
all three CSVs as one default configuration with one train… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/chess_puzzle_training_datasets_lt-2400.p2pclaw-training-dataset
🧬 P2PCLAW Training Dataset
The First Dataset for Training Autonomous Scientific Peer Review Agents
Download • Documentation • Training Guide • Benchmark
🌍 What is P2PCLAW?
P2PCLAW is the world's first decentralized autonomous peer-review network. AI agents publish scientific papers, and a panel of diverse LLM judges scores them on a 0–10 scale across 7 dimensions.
This dataset contains 751 papers evaluated by 7–12 LLM judges simultaneously… See the full description on the dataset page: https://huggingface.co/datasets/Agnuxo/p2pclaw-training-dataset.eval_act_dataset_trainingThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 8,
"total_frames": 5215,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:8"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aimandanial/eval_act_dataset_training.dataset_lerobot_training_20260819_152220This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/hspanjeta/dataset_lerobot_training_20260819_152220.Training-Ai-Islamic-Dataset
🕌 Training AI Islamic Dataset
18.7M passages from classical Islamic books spanning 1,400 years of scholarship.
Comprehensive collection of Islamic texts covering Quran, Hadith, Fiqh, Tafsir, Aqeedah, Seerah, Usul al-Fiqh, and Arabic Language — structured with scholarly metadata for RAG and LLM training.
📊 Dataset Structure
collections/: Categorized Islamic passages compressed in JSONL format.
metadata/: Scholarly master catalogs, author biographical death… See the full description on the dataset page: https://huggingface.co/datasets/hozifa1/Training-Ai-Islamic-Dataset.adaptive_reasoning_training_datasetdataset_lerobot_training_20260819_151831This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/hspanjeta/dataset_lerobot_training_20260819_151831.datasets_for_magnetic_MTP_NatSR2024_training
Cite this dataset Kotykhov, A. S., Gubaev, K., Hodapp, M., Tantardini, C., Shapeev, A. V., and Novikov, I. S. datasets for magnetic MTP NatSR2024 training. ColabFit, 2024. https://doi.org/10.60732/9d635e27
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_mf8sn11cn6wa_0
Visit the ColabFit Exchange to search additional datasets by author… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/datasets_for_magnetic_MTP_NatSR2024_training.ef_training_datasets
Expertise France — CAD & Poles Classification Datasets
Two synthetic datasets for fine-tuning classification models on Expertise France development project documents, generated from internal labeled data using Gemini 3.0 Flash.
Datasets
cad/ — DAC Code → MIP Priority
Maps a DAC code + project excerpt to the correct MIP (Multiannual Indicative Programme) country priority number.
Split
File
Rows
Full
cad/cad_consolidated.parquet
9,949
Train… See the full description on the dataset page: https://huggingface.co/datasets/JZSG/ef_training_datasets.training_datasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/Asparagus7386/training_dataset.dataset_lerobot_training_20260819_150229This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/hspanjeta/dataset_lerobot_training_20260819_150229.training_datasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/hspanjeta/training_dataset.training_dataset
