datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rlbenchfail_test_dataset
Guardian: RLBench-Fail Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data generated in the RLBench simulator for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Failures are produced by an automated pipeline that procedurally perturbs successful scripted trajectories in simulation, generating diverse planning… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/rlbenchfail_test_dataset.cbam-test-data
CBAM Test Data
Rows for CBAM reporting tests: declaration id, CN code, quantity, unit, embedded emissions, and country of origin. Synthetic values.
Free to use under CC0-1.0 — public domain dedication, no attribution required.
Files
cbam-small.json / cbam-small.csv — 100 rows (documents: 25)
cbam-medium.json / cbam-medium.csv — 2,000 rows (documents: 250)
cbam-large.json / cbam-large.csv — 20,000 rows (documents: 2,500)
Deterministic: the same size always… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/cbam-test-data.test-big-dataset
Dataset Card for Danish WIT
Dataset Summary
Google presented the Wikipedia Image Text (WIT) dataset in July
2021, a dataset which contains
scraped images from Wikipedia along with their descriptions. WikiMedia released
WIT-Base in September
2021,
being a modified version of WIT where they have removed the images with empty
"reference descriptions", as well as removing images where a person's face covers more
than 10% of the image surface, along with inappropriate images… See the full description on the dataset page: https://huggingface.co/datasets/huggingface/test-big-dataset.2025_Virtual_Cell_Challenge_Test_Datatest_shards_datasettest-dataset-v1iata-awb-test-data
IATA Air Waybill Test Data
Air waybill numbers combining a 3-digit airline prefix, a 7-digit serial, and a MOD-7 check digit (IATA Resolution 600a).
Free to use under CC0-1.0 — public domain dedication, no attribution required.
Files
iata-awb-small.json / iata-awb-small.csv — 100 rows (documents: 25)
iata-awb-medium.json / iata-awb-medium.csv — 2,000 rows (documents: 250)
iata-awb-large.json / iata-awb-large.csv — 20,000 rows (documents: 2,500)
Deterministic:… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/iata-awb-test-data.test_dataset
test_dataset
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
container-test-data
Container Number Test Data
Shipping container numbers with a correctly computed ISO 6346 check digit. Owner codes are synthetic.
Free to use under CC0-1.0 — public domain dedication, no attribution required.
Files
container-small.json / container-small.csv — 100 rows (documents: 25)
container-medium.json / container-medium.csv — 2,000 rows (documents: 250)
container-large.json / container-large.csv — 20,000 rows (documents: 2,500)
Deterministic: the same size… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/container-test-data.so100_test_FIRST_DATAThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 50,
"total_frames": 31309,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/BobBobbson/so100_test_FIRST_DATA.app_test_dataThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"j1",
"j2",
"j3",
"j4",
"j5",
"j6",
"gripper_state"
],
"shape": [
7
]
}… See the full description on the dataset page: https://huggingface.co/datasets/Bekhzod/app_test_data.test-dataso101_testdata_0820This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101",
"total_episodes": 30,
"total_frames": 9662,
"total_tasks": 1,
"total_videos": 60,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Tank-123/so101_testdata_0820.ur5fail_test_dataset
Guardian Failure Detection Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Guardian introduces an automated failure generation approach that procedurally perturbs successful robot trajectories to produce diverse planning failures and execution failures, each… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/ur5fail_test_dataset.lula-data-public-train-test
LULA Data Public Train/Test
This dataset contains 6,819,817 protein-small-molecule
pairs with binary bioactivity labels and fixed train/test splits.
Fields
pair_id: stable identifier for the protein-molecule pair.
canonical_smiles: canonical molecule SMILES.
protein_sequence: amino-acid sequence.
uniprot_id: normalized UniProt identifier.
label: 1 for binder/active and 0 for non-binder/inactive.
split: train or test.
sources: public databases contributing… See the full description on the dataset page: https://huggingface.co/datasets/omtx/lula-data-public-train-test.cucumber-dual-dataset15-testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 1,
"total_frames": 449,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/choiwoong/cucumber-dual-dataset15-test.lx7r_knockover_test_datasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "multi_robot",
"total_episodes": 16,
"total_frames": 5119,
"total_tasks": 1,
"total_videos": 32,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:16"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/arclabmit/lx7r_knockover_test_dataset.lerobot-testdata
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "agilex",
"total_episodes": 503,
"total_frames": 462517,
"total_tasks": 1,
"total_videos": 1509,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:503"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/yelanye/lerobot-testdata.arm-test-data_20260906_154504This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/zarianw/arm-test-data_20260906_154504.app_test_data_v3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"j1",
"j2",
"j3",
"j4",
"j5",
"j6",
"gripper_state"
],
"shape": [
7
]
}… See the full description on the dataset page: https://huggingface.co/datasets/Bekhzod/app_test_data_v3.arm-test-data_20260906_031327This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/zarianw/arm-test-data_20260906_031327.my_openarm_dataset_test_20260923_114306This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"joint_1.pos",
"joint_2.pos",
"joint_3.pos",
"joint_4.pos",
"joint_5.pos",
"joint_6.pos",
"joint_7.pos",
"gripper.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/LVelickovic/my_openarm_dataset_test_20260923_114306.data-studio-merge-test-002
Merged LeRobot Dataset
This dataset was created by merging multiple LeRobot datasets using the LeRobot merge tool.
Source Datasets
This merged dataset combines the following 25 datasets:
jackvial/koch_screwdriver_attach_orange_panel_e125
jackvial/koch_screwdriver_attach_orange_panel_1_e5
jackvial/koch_screwdriver_attach_orange_panel_3_e5
jackvial/koch_screwdriver_attach_orange_panel_5_e5
jackvial/koch_screwdriver_attach_orange_panel_6_e5… See the full description on the dataset page: https://huggingface.co/datasets/jackvial/data-studio-merge-test-002.bimanual-piper-dataset-threecam-test-11This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "bi_piper",
"total_episodes": 3,
"total_frames": 10691,
"total_tasks": 1,
"total_videos": 9,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:3"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/HCHoongChing/bimanual-piper-dataset-threecam-test-11.so101_test_dataset_2cams_20260926_012912This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/auntowen/so101_test_dataset_2cams_20260926_012912.Test_Audio_Generate_Dataset
Hinglish Audio Dataset
Generated by Sarvam AI.
marin-test-data-fixturesdataset_TESTThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 0,
"total_frames": 0,
"total_tasks": 0,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aicps/dataset_TEST.three-cam-test-datasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 1,
"total_frames": 1240,
"total_tasks": 1,
"total_videos": 1,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aaron-ser/three-cam-test-dataset.chess-test-data
