datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aloha_sim_insertion_humanThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "aloha",
"total_episodes": 50,
"total_frames": 25000,
"total_tasks": 1,
"total_videos": 50,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/aloha_sim_insertion_human.aloha_sim_transfer_cube_humanThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "aloha",
"total_episodes": 50,
"total_frames": 20000,
"total_tasks": 1,
"total_videos": 50,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/aloha_sim_transfer_cube_human.Humanoid-Everyday-H1arena-human-preference-100k
Overview
This dataset contains leaderboard conversation data collected between June 2024 and August 2024.
It includes English human preference evaluations used to develop Arena Explorer.
Additionally, we provide an embedding file, which contains precomputed embeddings for the English conversations.
These embeddings are used in the topic modeling pipeline to categorize and analyze these conversations.
For a detailed exploration of the dataset and analysis methods, refer to the… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/arena-human-preference-100k.esg_reports_human_labeled_v2
Vidore Benchmark 2 - ESG Human Labeled
This dataset is part of the "Vidore Benchmark 2" collection, designed for evaluating visual retrieval applications. It focuses on the theme of ESG reports from the fast food industry.
Dataset Summary
Each query is in english.
This dataset provides a focused benchmark for visual retrieval tasks related to ESG reports for the fast food industry. It includes a curated set of documents, queries, relevance judgments (qrels), and page… See the full description on the dataset page: https://huggingface.co/datasets/vidore/esg_reports_human_labeled_v2.humanoid-robots-training-dataset
Dynamic Intelligence — Humanoid Robot Training Dataset
A first-person (egocentric) video dataset of human hand manipulation, designed for training humanoid robot policies via imitation learning. Each episode captures a person performing an everyday household task — folding clothes, moving dishes, opening doors — filmed from a head-mounted iPhone using its built-in LiDAR and depth sensors.
The dataset pairs each video with frame-level 3D hand tracking and camera pose data, giving… See the full description on the dataset page: https://huggingface.co/datasets/DynamicIntelligence/humanoid-robots-training-dataset.Humanoid-Everyday-G1ai-humanizer-benchmark
AI Humanizer Benchmark: AI humanizers tested against 7 AI detectors (October 2026)
AI Humanizer Benchmark measures how well AI humanizers rewrite AI-generated text so that AI detectors classify it as human-written, and how much meaning and readability the rewrite loses. In each monthly cycle, 11 AI humanizers rewrite the same 33 source texts across 7 writing categories, using each tool's default settings. Each output gets three kinds of score: 7 AI detectors (GPTZero… See the full description on the dataset page: https://huggingface.co/datasets/ai-humanizer-benchmark/ai-humanizer-benchmark.mt_bench_human_judgments
Content
This dataset contains 3.3K expert-level pairwise human preferences for model responses generated by 6 models in response to 80 MT-bench questions.
The 6 models are GPT-4, GPT-3.5, Claud-v1, Vicuna-13B, Alpaca-13B, and LLaMA-13B. The annotators are mostly graduate students with expertise in the topic areas of each of the questions. The details of data collection can be found in our paper.
Agreement Calculation
This Colab notebook shows how to compute the… See the full description on the dataset page: https://huggingface.co/datasets/lmsys/mt_bench_human_judgments.arena-human-preference-55kDataset for Kaggle competition on predicting human preference on Chatbot Arena battles.
The training dataset includes over 55,000 real-world user and LLM conversations and user preferences across over 70 state-of-the-art LLMs, such as GPT-4, Claude 2, Llama 2, Gemini, and Mistral models.
Each sample represents a battle consisting of 2 LLMs which answer the same question, with a user label of either prefer model A, prefer model B, tie, or tie (both bad).
Citation
Please cite the… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/arena-human-preference-55k.world-model-physics-human-preference-283k
Rapidata Physics Benchmark
Built by Rapidata.
Do video and world models understand physics? We gave 25 video- and world models the same
real-world starting frame and scene description from Physics-IQ and asked
them to predict what happens next. ~283,000 human votes, collected with the
Rapidata Python SDK, decided which continuation is more realistic — with the
real recording competing as a hidden 26th participant.
Each row is a head-to-head matchup between two participants on… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/world-model-physics-human-preference-283k.HumanTracker
Dataset Card for HumanTracker
Project page · Paper · Code
HumanTracker is a humanoid motion-tracking benchmark. This release contains two complementary subsets:
motions/ — the evaluation test split: retargeted 29-DoF reference trajectories, grouped into four motion families.
preference_pair/ — 6,000 human preference pairs, each stored with the two tracker rollouts that were compared and the source-motion clip they track.
The evaluation harness and HumanScore reward model live… See the full description on the dataset page: https://huggingface.co/datasets/GalaxyGeneralRobotics/HumanTracker.humanoid-everyday-stepit
Humanoid-Everyday · stepit action(G1 子集)
一个自包含的标准 LeRobot v2.1 数据集,由原始 Humanoid-Everyday 数据集的 Unitree G1 子集
重新表达而来。与原始数据唯一的区别是 action 列被替换成了一个 41 维、包含完整 base 状态
(位置 + 姿态 + 线速度)的动作向量;其余所有列均逐字节保持不变,可直接用标准 LeRobot 加载器读取。
动作空间:41 维 = base(10) + G1 身体 29 关节 + 双手开合 2;抽取相应切片即可喂给
stepit(Unitree G1 29-DOF)控制器(见下文 stepit qpos)。
规模:4068 episodes / 1,781,092 frames / 246 tasks / 9 chunks,fps = 30,robot_type = g1。
体积:≈ 394 GB(parquet ≈ 391 GB + 视频 ≈ 3 GB + meta ≈ 2 MB)。
原始数据是 G1/H1 混合的(共… See the full description on the dataset page: https://huggingface.co/datasets/UsanoCoCr/humanoid-everyday-stepit.robocasa_target_human_unifiedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "robocasa",
"total_episodes": 25307,
"total_frames": 14957899,
"total_tasks": 50,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:25307"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/robocasa_target_human_unified.text-2-video-human-preferences
Rapidata Video Generation Preference Dataset
This dataset was collected in ~12 hours using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
The data collected in this dataset informs our text-2-video model benchmark. We just started so currently only two models are represented in this set:
Sora
Hunyouan
Pika 2.0
Runway ML Alpha
Luma Ray 2
Explore our latest model rankings on our website.
If you get value from this dataset and would… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences.gspc-human-labour-index
GSPC — labour components facts (Eurostat)
In one line: Two cited Eurostat labour series, read as deterministic facts, behind the board's labour-components axis. Not an index and no composite score. For economists and policy analysts who want source-checkable numbers.
Use it
from datasets import load_dataset
ds = load_dataset("csoai/gspc-human-labour-index", split="train")
print(ds[0])
Verify a signed card in your browser, free, no account:… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-human-labour-index.camera-movement-human-preference-324k
Rapidata Camera Movement Benchmark
Built by Rapidata.
This dataset contains 324,044 human responses, collected with the
Rapidata Python SDK, comparing how well 15 image-to-video models and world models
execute a described camera movement from a single still image. Each row is a head-to-head comparison between
two models' clips generated from the same still and the same instruction, judged by human annotators who
watched a reference animation of the requested movement.
The task… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/camera-movement-human-preference-324k.text-2-video-human-preferences-seedance-1-pro
Rapidata Video Generation Seedance 1 Pro Human Preference
In this dataset, ~60k human responses from ~20k human annotators were collected to evaluate Seedance 1 Pro video generation model on our benchmark. This dataset was collected in roughtly 30 min using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-seedance-1-pro.gspc-humanoid-labour-index
GSPC — humanoid labour index facts (Disclosure)
In one line: Does a named humanoid-robot vendor publish a dated deployment count on a stable URL? Yes/no facts from 8 frozen vendor URLs, behind the board's humanoid-labour-index axis. Empty cells stay empty. For robotics and labour analysts.
Use it
from datasets import load_dataset
ds = load_dataset("csoai/gspc-humanoid-labour-index", split="train")
print(ds[0])
Verify a signed card in your browser, free, no… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-humanoid-labour-index.egopi_latal_humantext-2-video-human-preferences-veo3
Rapidata Video Generation Veo 3 Human Preference
In this dataset, ~46k human responses from ~20k human annotators were collected to evaluate Veo3 video generation model on our benchmark. This dataset was collected in roughly 35 minutes using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please consider liking… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-veo3.PPE-Human-Preference-V1
Overview
This contains the human preference evaluation set for Preference Proxy Evaluations.
This dataset is meant for benchmarking and evaluation, not for training.
Paper
Code
License
User prompts are licensed under CC-BY-4.0, and model outputs are governed by the terms of use set by the respective model providers.
Citation
@misc{frick2024evaluaterewardmodelsrlhf,
title={How to Evaluate Reward Models for RLHF},
author={Evan Frick and Tianle Li and… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/PPE-Human-Preference-V1.human_ref_dna
Dataset Card for "human_ref_dna"
More Information needed
humanoidtoolbench-teleop
ToolBook: HumanoidToolBench demonstrations
Paper: arXiv:2610.02089
Raw Meta Quest teleoperation demonstrations for
HumanoidToolBench, recorded on
the Unitree G1 in MuJoCo through its whole-body controller.
Each recording uses LeRobot v2.1 files: 50 Hz, an ego-view video and wrist-view
videos when recorded, and the recording contract in meta/simple_contract.json.
Camera resolutions are
declared in meta/info.json. Code versions, WBC settings, and the render
backend are in… See the full description on the dataset page: https://huggingface.co/datasets/snupilab/humanoidtoolbench-teleop.QIT
QIT Humanize-Physic Formalizations and Proofs
QIT (Quantum Information Theory) is a blind benchmark for formalizing theorems in quantum information. It evaluates whether an AI agent can faithfully translate natural-language and TeX problem statements into Lean 4 theorems and then construct formal proofs checked by the Lean kernel. Its 40 tasks cover quantum channels and Choi representations, entropy and coding, mixed-unitary obstructions and symmetry, norm and fidelity tools… See the full description on the dataset page: https://huggingface.co/datasets/humanfia-lab/QIT.humans-top
humans.top — LIVE Global ranking of influential people (open dataset)
This dataset ranks real, named living people by global influence — e.g. #1
Donald Trump, #2 Xi Jinping, #3 Vladimir Putin, alongside figures like Elon Musk,
Narendra Modi and Lionel Messi. Every row is a person: their live influence
rank, a concise biography in 15 languages, and Wikidata / Wikipedia links.
Published from the website humans.top (.top is the
domain name).
Available on (identical CC0… See the full description on the dataset page: https://huggingface.co/datasets/dsfox/humans-top.fastdetector-test-stat-test-unedited-human
Auto-generated FastDetector dataset
Fastdetector Test Stat Test Unedited Human
Best detectorGiga EditLens Llama-3.2-3B Score0.6438 TPR @ 1% FPRHardest prompt subsetrewrite0.2581 max detector TPR @ 1% FPRHardest generator configclaude-opus-5 (Temp: Unknown)0.4209 max detector TPR @ 1% FPR
15,695rows10generator configs4prompt subsets4detectors
01Leaderboard02Model analytics03Distances04Appendix
01Detector leaderboardScore-based detectors ranked by overall AUROC. Thresholds are placed exactly… See the full description on the dataset page: https://huggingface.co/datasets/G-reen/fastdetector-test-stat-test-unedited-human.PLM-Video-Human
Dataset Card for PLM-Video Human
PLM-Video-Human is a collection of human-annotated resources for training Vision Language Models,
focused on detailed video understanding. Training tasks include: fine-grained open-ended question answering (FGQA), Region-based Video Captioning (RCap),
Region-based Dense Video Captioning (RDCap) and Region-based Temporal Localization (RTLoc).
[📃 Tech Report]
[📂 Github]
Dataset Structure
Fine-Grained Question Answering (FGQA)… See the full description on the dataset page: https://huggingface.co/datasets/facebook/PLM-Video-Human.voxceleb2_dev
VoxCeleb2 Dev Dataset
数据集描述
VoxCeleb2 Dev 数据集是 VoxCeleb2 数据集的开发集(训练集),用于说话人识别和音频检索任务。VoxCeleb2 是 VoxCeleb1 的扩展版本,包含更多的说话人和音频样本。该数据集包含标准化的音频文件和对应的元数据。
数据集结构
voxceleb2_dev/
├── data.parquet # 数据集元数据文件(Parquet格式)
├── audio_0/ # 音频文件目录 0
├── audio_1/ # 音频文件目录 1
├── audio_2/ # 音频文件目录 2
├── audio_3/ # 音频文件目录 3
├── audio_4/ # 音频文件目录 4
├── audio_5/ # 音频文件目录 5
└── audio_6/ #… See the full description on the dataset page: https://huggingface.co/datasets/humanify/voxceleb2_dev.MM-Food-100K
Overview
This project aims to introduce and release a comprehensive food image dataset designed specifically for computer vision tasks, particularly food recognition, classification, and nutritional analysis. We hope this dataset will provide a reliable resource for researchers and developers to advance the field of food AI. By publishing on Hugging Face, we expect to foster community collaboration and accelerate innovation in applications such as smart recipe recommendations… See the full description on the dataset page: https://huggingface.co/datasets/Humanbased-AI/MM-Food-100K.
