datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DECO-50
DECO: Decoupled Multimodal Diffusion Transformer for Bimanual Dexterous Manipulation with a Plugin Tactile Adapter
DECO-50 is a bimanual dexterous manipulation dataset with tactile sensing, comprising 50 hours of teleoperated data across 4 scenarios and 28 subtasks, totaling over 5 million frames collected on real dual-arm robots.
Dataset Structure
DECO-50/
├── task1/
│ ├── sub_task_1/
│ │ ├── episode_000000/
│ │ │ ├──… See the full description on the dataset page: https://huggingface.co/datasets/BAAI-Humanoid/DECO-50.aloha_sim_insertion_human_imageThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "aloha",
"total_episodes": 50,
"total_frames": 25000,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null,
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/aloha_sim_insertion_human_image.aloha_sim_transfer_cube_human_imageThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "aloha",
"total_episodes": 50,
"total_frames": 20000,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null,
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/aloha_sim_transfer_cube_human_image.MPII_Human_Pose_Dataset
Dataset Card for MPII Human Pose
MPII Human Pose dataset is a state of the art benchmark for evaluation of articulated human pose estimation.
The dataset includes around 25K images containing over 40K people with annotated body joints.
The images were systematically collected using an established taxonomy of every day human activities.
Overall the dataset covers 410 human activities and each image is provided with an activity label.
Each image was extracted from a YouTube… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/MPII_Human_Pose_Dataset.esg_reports_human_labeled_v2
Vidore Benchmark 2 - ESG Human Labeled
This dataset is part of the "Vidore Benchmark 2" collection, designed for evaluating visual retrieval applications. It focuses on the theme of ESG reports from the fast food industry.
Dataset Summary
Each query is in english.
This dataset provides a focused benchmark for visual retrieval tasks related to ESG reports for the fast food industry. It includes a curated set of documents, queries, relevance judgments (qrels), and page… See the full description on the dataset page: https://huggingface.co/datasets/vidore/esg_reports_human_labeled_v2.human-motion-tracking-deeplabcutThis dataset is used to adapt DeepLabCut for Human motion tracking.
Structure of the dataset
videos contains 100+ videos of 4 candidates recorded during a game of darts.
labeled-data contains labels on the corresponding frames of the videos. These labels are used to adapt DeepLabCut for human motion tracking. Under labeled-data there are 2 folders for every video.
video_name has all the relevant frames extracted from the video, xy coordinates of the labels in the csv file and the… See the full description on the dataset page: https://huggingface.co/datasets/pratikshapai/human-motion-tracking-deeplabcut.humanitys-second-last-exam
Humanity's Second Last Exam
Benchmark design, curation and release maintenance: Shashank Agnihotri.
Original questions retain their recorded authorship and source attribution.
This owner-reviewed retained release contains 365 target questions, 730
context examples, and 365 ordered target/A/B links: 1,095 question rows.
The owner review concluded on 16 September 2026. This is an owner-reviewed
release after suspected-AI-content exclusions, not a software-certified
guarantee of… See the full description on the dataset page: https://huggingface.co/datasets/shashankskagnihotri/humanitys-second-last-exam.gimbaled-uav-tracking-dataset
Gimbaled UAV Tracking Dataset / 云台无人机追踪数据集
A real-world, multi-sensor dataset for active localization of a non-cooperative
UAV using a two-axis gimbaled LiDAR–camera fusion system. It contains 16 flight
sequences acquired with a ground vehicle platform, together with the associated
tracking outputs and an RTK-based position reference.
面向非合作无人机主动定位的实测多传感数据集,采集自一套基于两轴云台
LiDAR–相机融合的地面车辆平台。包含 16 个飞行序列,并附带相应的跟踪
输出与基于 RTK 的位置参考。
The sequences span two acquisition days, sunny and… See the full description on the dataset page: https://huggingface.co/datasets/humanoidro/gimbaled-uav-tracking-dataset.craft-gc-human-eval-results
CRAFT-GC Human Evaluation Results
Study version: v4-yesno-30x25grid (reset 20260628-171828 UTC)
Format
30 prompts randomly sampled from GCFairBench-100
5 diffusion seeds per prompt (150 images total)
Yes/No questions per image (realism; cultural appropriateness)
Three evaluators (E1, E2, E3) — scores summed as yes-vote counts
Files
ratings.jsonl — one JSON object per image answer (current round)
submissions/ — per-evaluator submission snapshots… See the full description on the dataset page: https://huggingface.co/datasets/nati1221/craft-gc-human-eval-results.human-motion-tracking-deeplabcutThis dataset is used to adapt DeepLabCut for Human motion tracking.
Structure of the dataset
videos contains 100+ videos of 4 candidates recorded during a game of darts.
labeled-data contains labels on the corresponding frames of the videos. These labels are used to adapt DeepLabCut for human motion tracking. Under labeled-data there are 2 folders for every video.
video_name has all the relevant frames extracted from the video, xy coordinates of the labels in the csv file and the… See the full description on the dataset page: https://huggingface.co/datasets/GT-Neuronext/human-motion-tracking-deeplabcut.text-2-video-human-preferences
Rapidata Video Generation Preference Dataset
This dataset was collected in ~12 hours using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
The data collected in this dataset informs our text-2-video model benchmark. We just started so currently only two models are represented in this set:
Sora
Hunyouan
Pika 2.0
Runway ML Alpha
Luma Ray 2
Explore our latest model rankings on our website.
If you get value from this dataset and would… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences.HumanEdit
Dataset Card for HumanEdit
Paper (CVPR 2025 AI for Content Creation (AI4CC) Workshop)
Usage
from datasets import load_dataset
from PIL import Image
# Load the dataset
ds = load_dataset("BryanW/HumanEdit")
# Print the total number of samples and show the first sample
print(f"Total number of samples: {len(ds['train'])}")
print("First sample in the dataset:", ds['train'][0])
# Retrieve the first sample's data
data_dict = ds['train'][0]
# Save the input image (INPUT_IMG)… See the full description on the dataset page: https://huggingface.co/datasets/BryanW/HumanEdit.camera-movement-human-preference-324k
Rapidata Camera Movement Benchmark
Built by Rapidata.
This dataset contains 324,044 human responses, collected with the
Rapidata Python SDK, comparing how well 15 image-to-video models and world models
execute a described camera movement from a single still image. Each row is a head-to-head comparison between
two models' clips generated from the same still and the same instruction, judged by human annotators who
watched a reference animation of the requested movement.
The task… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/camera-movement-human-preference-324k.hit-uav-thermal-human-detectiontext-2-video-human-preferences-seedance-1-pro
Rapidata Video Generation Seedance 1 Pro Human Preference
In this dataset, ~60k human responses from ~20k human annotators were collected to evaluate Seedance 1 Pro video generation model on our benchmark. This dataset was collected in roughtly 30 min using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-seedance-1-pro.Fashion-1K
Fashion 1K
Dataset Summary
Fashion 1K is a curated collection of 1,000 high-quality fashion images, focusing on apparel and outfit compositions without human models.
Unlike typical street-style datasets (like DeepFashion) that include human poses and complex backgrounds, this dataset provides clean, human-free images. The images primarily feature Flat Lay (clothing arranged on a flat surface) or Ghost Mannequin styles, making them ideal for tasks that require a… See the full description on the dataset page: https://huggingface.co/datasets/Humanbased-AI/Fashion-1K.human-coherence-preferences-images
Rapidata Image Generation Coherence Dataset
This dataset was collected in ~4 Days using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please consider liking it.
Overview
One of the largest human annotated coherence datasets for text-to-image models, this release contains over 1,200,000 human… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/human-coherence-preferences-images.text-2-video-human-preferences-veo3
Rapidata Video Generation Veo 3 Human Preference
In this dataset, ~46k human responses from ~20k human annotators were collected to evaluate Veo3 video generation model on our benchmark. This dataset was collected in roughly 35 minutes using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please consider liking… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-veo3.700k_Human_Preference_Dataset_FLUX_SD3_MJ_DALLE3
NOTE: A newer version of this dataset is available Imagen3_Flux1.1_Flux1_SD3_MJ_Dalle_Human_Preference_Dataset
Rapidata Image Generation Preference Dataset
This Dataset is a 1/3 of a 2M+ human annotation dataset that was split into three modalities: Preference, Coherence, Text-to-Image Alignment.
Link to the Coherence dataset: https://huggingface.co/datasets/Rapidata/Flux_SD3_MJ_Dalle_Human_Coherence_Dataset
Link to the Text-2-Image Alignment dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/700k_Human_Preference_Dataset_FLUX_SD3_MJ_DALLE3.human-style-preferences-images
Rapidata Image Generation Preference Dataset
This dataset was collected in ~4 Days using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please consider liking it.
Overview
One of the largest human preference datasets for text-to-image models, this release contains over 1,200,000 human preference… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/human-style-preferences-images.Cabin-Human-Behavior-Dataset
全球最大的智能座舱多模态开源高质量数据集来啦!
一. 数据集摘要 (Dataset Summary)
「CyberData塞塔」智能座舱用户行为数据集是一个专为加速智能座舱感知算法开发而设计的高质量、程序化生成的图像数据集。随着 C-NCAP、EU GSR 等全球汽车安全法规对驾驶员监控系统 (DMS) 和乘客监控系统 (OMS) 提出更高要求,安全、合规、多样化的训练数据变得至关重要。本数据集通过合成方式,旨在解决真实世界数据采集面临的隐私风险、高昂成本和长尾场景覆盖不足等核心挑战。
该数据集包含 5,000 张 由 XAI Lab 自主研发的数据集生成引擎合成的高保真座舱内用户行为图像,每张图像都附带丰富的、100% 精确的标注信息。
核心特点:
丰富的场景多样性: 涵盖不同年龄、性别、种族和衣着风格的虚拟人模型,以及多种驾驶与乘坐行为(如使用手机、喝水、疲劳、手势)和面部表情。
专为座舱感知优化: 数据集可直接用于智能座舱端侧视觉模型,尤其是 DMS/OMS 算法的训练、微调与验证,帮助模型精准理解座舱内复杂的交互与状态。… See the full description on the dataset page: https://huggingface.co/datasets/OpenSparX/Cabin-Human-Behavior-Dataset.Seedream-3_t2i_human_preference
Rapidata Seedream 3 Preference
This T2I dataset contains over ~400'000 human responses from over ~30'000 individual annotators, collected in less than 7h using the Rapidata Python API, accessible to anyone and ideal for large scale evaluation.
Evaluating OpenAI 4o (version from 26.3.2025) across three categories: preference, coherence, and alignment.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/Seedream-3_t2i_human_preference.MolmoWeb-HumanTrajs
MolmoWeb-HumanTrajs
A dataset of human collected web trajectories. Each example pairs an instruction with a sequence of webpage screenshots and the corresponding agent actions (clicks, typing, scrolling, etc.).
Dataset Usage
from datasets import load_dataset
# load a single subset
ds = load_dataset("allenai/MolmoWeb-HumanTrajs")
Working with images and trajectories
Each row has an images field (list of raw image bytes) and a corresponding… See the full description on the dataset page: https://huggingface.co/datasets/allenai/MolmoWeb-HumanTrajs.humanoidtoolbench-teleop
ToolBook: HumanoidToolBench demonstrations
Paper: arXiv:2610.02089
Raw Meta Quest teleoperation demonstrations for
HumanoidToolBench, recorded on
the Unitree G1 in MuJoCo through its whole-body controller.
Each recording uses LeRobot v2.1 files: 50 Hz, an ego-view video and wrist-view
videos when recorded, and the recording contract in meta/simple_contract.json.
Camera resolutions are
declared in meta/info.json. Code versions, WBC settings, and the render
backend are in… See the full description on the dataset page: https://huggingface.co/datasets/snupilab/humanoidtoolbench-teleop.human-alignment-preferences-images
Rapidata Image Generation Alignment Dataset
This dataset was collected in ~4 Days using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please consider liking it.
Overview
One of the largest human annotated alignment datasets for text-to-image models, this release contains over 1,200,000 human… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/human-alignment-preferences-images.Flux-2-pro_t2i_human_preference
Rapidata Flux 2 Pro Preference
This T2I dataset contains over ~400'000 human responses from over ~50'000 individual annotators, collected in less than 7h using the Rapidata Python API, accessible to anyone and ideal for large scale evaluation.
Evaluating Flux 2 Pro (version from 25.11.25) across three categories: preference, coherence, and alignment.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/Flux-2-pro_t2i_human_preference.text-2-image-Rich-Human-Feedback
Building upon Google's research Rich Human Feedback for Text-to-Image Generation we have collected over 1.5 million responses from 152'684 individual humans using Rapidata via the Python API. Collection took roughly 5 days.
If you get value from this dataset and would like to see more in the future, please consider liking it.
Overview
We asked humans to evaluate AI-generated images in style, coherence and prompt alignment. For images that contained flaws, participants were… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-image-Rich-Human-Feedback.human_parsing_dataset
Dataset Card for Human parsing data (ATR)
Dataset Summary
This dataset has 17,706 images and mask pairs. It is just a copy of
Deep Human Parsing ATR dataset. The mask labels are:
"0": "Background",
"1": "Hat",
"2": "Hair",
"3": "Sunglasses",
"4": "Upper-clothes",
"5": "Skirt",
"6": "Pants",
"7": "Dress",
"8": "Belt",
"9": "Left-shoe",
"10": "Right-shoe",
"11": "Face",
"12": "Left-leg",
"13": "Right-leg",
"14":… See the full description on the dataset page: https://huggingface.co/datasets/mattmdjaga/human_parsing_dataset.OpenAI-4o_t2i_human_preference
Rapidata OpenAI 4o Preference
This T2I dataset contains over 200'000 human responses from over ~45,000 individual annotators, collected in less than half a day using the Rapidata Python API, accessible to anyone and ideal for large scale evaluation.
Evaluating OpenAI 4o (version from 26.3.2025) across three categories: preference, coherence, and alignment.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/OpenAI-4o_t2i_human_preference.HumanRef-CoT-45k
🦖🧠 Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning 🦖🧠
We propose Rex-Thinker, a Chain-of-Thought (CoT) reasoning model for object referring that addresses two key challenges: lack of interpretability and inability to reject unmatched expressions. Instead of directly predicting bounding boxes, Rex-Thinker reasons step-by-step over candidate objects to determine which, if any, match a given expression.… See the full description on the dataset page: https://huggingface.co/datasets/IDEA-Research/HumanRef-CoT-45k.
