datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OpenThoughts-114k-math-correct-qwen3-14b-math-prepared-topk128-normalizedXDof-TshirtFolding-20hours-normalizedDepth-Normal-Videos-42K
Depth and Normal Videos Dataset
42,498 videos with depth and surface normals.
Usage
from huggingface_hub import hf_hub_download
video = hf_hub_download(
repo_id="Yanbin99/Depth-Normal-Videos-42K",
filename="Depth_and_Normal_42K/group_0000/videos/00000000.mp4",
repo_type="dataset"
)
OpenThoughts-114k-math-correct-qwen3-14b-math-prepared-normalizedmeld-open-normalized
MELD Open (Normalized)
MELD is a multilingual and multi-domain dataset for Named Entity Recognition (NER) constructed from 60 existing datasets. It includes gold-standard annotations across 60 languages and 14 domains. This dataset is a subset of 43 datasets for which licenses permit the redistribution of data in a new format. See the MELD GitHub repository for more details.
Note: This version of MELD Open uses normalized labels. For original labels from each source dataset, use… See the full description on the dataset page: https://huggingface.co/datasets/kgnlp/meld-open-normalized.lichess-stockfish-normalized
Lichess Chess Positions: ML-Ready Deduplicated Evaluations
Dataset Description
A curated dataset of 316,072,343 unique chess positions with Stockfish evaluations, optimized for training neural networks. This is a deduplicated, ML-ready version of the Lichess evaluation database.
Why This Dataset?
While Lichess provides deduplicated evaluations in JSONL.zst format, and HuggingFace hosts the full (non-deduplicated) version, this dataset offers:
Unique advantages:… See the full description on the dataset page: https://huggingface.co/datasets/mateuszgrzyb/lichess-stockfish-normalized.normal_ball_fullThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 50,
"total_frames": 44089,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/satvikahuja/normal_ball_full.CartonPickNPlace2Target-normalizedE_normal_over70_add
Dataset Card for "E_normal_over70_add"
More Information needed
grabette-tactile-normal-3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "grabette",
"total_episodes": 17,
"total_frames": 6800,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 50,
"splits": {
"train": "0:17"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/CarolinePascal/grabette-tactile-normal-3.normal_ball_part2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 50,
"total_frames": 44089,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/satvikahuja/normal_ball_part2.grabette-tactile-normal-4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "grabette",
"total_episodes": 17,
"total_frames": 6800,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 50,
"splits": {
"train": "0:17"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/CarolinePascal/grabette-tactile-normal-4.VFD_normalize_9_v1eval_100k_groot_normalThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "piper_follower",
"total_episodes": 15,
"total_frames": 21693,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:15"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/kaiseong/eval_100k_groot_normal.business-entity-resolution-normalized
Business Entity Resolution: normalised records
Normalised copies of the six source files of the ML Challenge 2026 Business Entity Resolution task
(business records from three sources, US / India in train, plus France in test).
The goal of the task is to find, for every Source 1 record, the Source 2 / Source 3 records that describe the same business.
File
Rows
train_s1.parquet
2,206,821
train_s2.parquet
5,034,616
train_s3.parquet
5,285,603
test_s1.parquet
1,732… See the full description on the dataset page: https://huggingface.co/datasets/vc940/business-entity-resolution-normalized.grabette-tactile-normal-2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "grabette",
"total_episodes": 17,
"total_frames": 6800,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 50,
"splits": {
"train": "0:17"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/CarolinePascal/grabette-tactile-normal-2.E_normal_over70
Dataset Card for "E_normal_over70"
More Information needed
Stuffed_Animal_V4.1_3cam_Normal_bboxes
Stuffed_Animal_V4.1_3cam_Normal
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
normalcamera_HKWSThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "so100",
"total_episodes": 50,
"total_frames": 16843,
"total_tasks": 1,
"total_videos": 50,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/zhaoraning/normalcamera_HKWS.grabette-tactile-normalThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "grabette",
"total_episodes": 17,
"total_frames": 6800,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 50,
"splits": {
"train": "0:17"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/CarolinePascal/grabette-tactile-normal.galaxea-r1-shelf-full-normalizedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": null,
"total_episodes": 99,
"total_frames": 48085,
"total_tasks": 1,
"total_videos": 297,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:99"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Yiheyihe/galaxea-r1-shelf-full-normalized.test_community_normalizedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/CarolinePascal/test_community_normalized.Y_normal
Dataset Card for "Y_normal"
More Information needed
70B_normal_llama_33_70b_instruct__swe_bench_verified_mini
Inspect Dataset: 70B_normal_llama_33_70b_instruct__swe_bench_verified_mini
Dataset Information
This dataset was created using the create_inspect_dataset function from the deception_sprint package on 2025-06-17.
Model Information
Model: vllm/meta-llama/Llama-3.3-70B-Instruct
Model args: {'max_model_len': 32768, 'gpu_memory_utilization': 0.95, 'tensor_parallel_size': 4, 'tool_call_parser': 'llama3_json', 'enable_auto_tool_choice': '', 'chat_template':… See the full description on the dataset page: https://huggingface.co/datasets/jordan-taylor-aisi/70B_normal_llama_33_70b_instruct__swe_bench_verified_mini.reward-bench-hacking-rewards-harmless-train-normalOpenThoughts-114k-math-correct-qwen3-14b-math-prepared-topk256-normalizedmy_lab_pick_load_50eps_normalThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 50,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
31
],
"names": [
"waist_yaw_joint",
"waist_roll_joint",
"waist_pitch_joint",
"left_shoulder_pitch_joint"… See the full description on the dataset page: https://huggingface.co/datasets/dtakehara/my_lab_pick_load_50eps_normal.test_community_normalized_1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/CarolinePascal/test_community_normalized_1.hoike_normal_expression_GTEx_Analysis_v10_log2tpmplus1
Hōʻike - Normal GTEx Gene Expression Data in Log2(TPM+1) Format
These data are for use in the Hoike gene expression data generation models as the normal_profile input.
These were obtained from the GTEx_Analysis_v10_RNASeQCv2.4.2_gene_tpm.gct.gz file from the GTEx Portal on 6/10/2026.
Mathlib-Normalized-Sexpr
Mathlib Normalized S-Expressions
Lean 4 proof states from Mathlib, paired with the tactic applied at each
step, in three representations extracted directly from the Lean kernel:
Source-faithful S-expressions of the goal and every hypothesis, as
Lean elaborated them.
Normalized S-expressions of the same state, with stable local-context
indices suitable for model input.
Annotated tactic syntax -- the original tactic's syntax tree with
identifier leaves resolved to the constants… See the full description on the dataset page: https://huggingface.co/datasets/jajostrains/Mathlib-Normalized-Sexpr.
