datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dataset_merged_preprocesssed_v2
Dataset Card for "dataset_merged_preprocesssed_v2"
More Information needed
srtm30m-mergedmisc-merged-claude-code-traces-v1
MISC Unification of Public Claude Code Traces
A unified dataset of 32,133 deduplicated Claude API conversation traces focused on software engineering and code generation tasks. This dataset merges and normalizes traces from 10 different source datasets into a single, consistent format.
Dataset Description
This dataset contains real Claude API interaction traces capturing software engineering workflows including:
Code generation and modification
Bug fixing and debugging… See the full description on the dataset page: https://huggingface.co/datasets/nlile/misc-merged-claude-code-traces-v1.CC-100-zh-Hant-merged
CC-100 zh-Hant (Traditional Chinese)
From https://data.statmt.org/cc-100/, only zh-Hant - Chinese (Traditional). Broken into paragraphs, with each paragraphs as a row.
Estimated to have around 4B tokens when tokenized with the bigscience/bloom tokenizer.
There's another version that the text is split by lines instead of paragraphs: zetavg/CC-100-zh-Hant.
References
Please cite the following if you found the resources in the CC-100 corpus useful.
Unsupervised… See the full description on the dataset page: https://huggingface.co/datasets/zetavg/CC-100-zh-Hant-merged.robotwin_merged
LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies
This repository contains the preprocessed dataset used in the paper LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies.
Project Page: https://rlinf.github.io/LaWAM/
Repository: https://github.com/RLinf/LaWAM
The dataset is formatted in LeRobot format and is designed for training and evaluating dynamics-aware robot policies.
Citation
@misc{chen2026lawam… See the full description on the dataset page: https://huggingface.co/datasets/jialei02/robotwin_merged.details_grimjim__Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge
Dataset Card for Evaluation run of grimjim/Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge
Dataset automatically created during the evaluation run of model grimjim/Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_grimjim__Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge.transformers-merge-experimentsGR1-Tabletop-Merged-1000x24
GR1 Tabletop Merged LeRobot Datasets
Merged and subsampled versions of the GR1 tabletop manipulation datasets from the NVIDIA PhysicalAI-Robotics-GR00T-X-Embodiment-Sim collection, formatted in LeRobot v2.0 format.
Dataset Variants
Variant
Demos/Task
Tasks
Total Episodes
Total Frames
Approx Size
1000x24/
1000
24 folders, 186 unique tasks
24,000
6,020,058
~40 GB
300x24/
300
24 folders, 186 unique tasks
7,200
1,803,236
~12 GB
100x24/
100
24 folders… See the full description on the dataset page: https://huggingface.co/datasets/Joocjun/GR1-Tabletop-Merged-1000x24.merged-swe-bench-lite-extra-verified
Merged SWE-bench Dataset
This dataset merges these source datasets:
princeton-nlp/SWE-bench_Lite
nebius/SWE-bench-extra
princeton-nlp/SWE-bench_Verified
Row counts
princeton-nlp/SWE-bench_Lite: 323 rows
nebius/SWE-bench-extra: 6,376 rows
princeton-nlp/SWE-bench_Verified: 500 rows
Splits used
princeton-nlp/SWE-bench_Lite splits: dev: 23, test: 300
nebius/SWE-bench-extra splits: train: 6,376
princeton-nlp/SWE-bench_Verified splits: test: 500… See the full description on the dataset page: https://huggingface.co/datasets/antontuzovAI/merged-swe-bench-lite-extra-verified.vqa_merged2Patent_FR_US_Merge_Radix_65536korean_hate_speech_merge11kasim-merged-filteredindic-align-merged-cleanedgit-commits-merged
Themis-Git-Commits-Merged
Overview
Themis-Git-Commits-Merged is a large-scale dataset of ~3.98M single-file code commits from permissively licensed GitHub repositories that have been cross-referenced with GHTorrent pull request data to retain only commits that are part of successfully merged, non-reverted pull requests. This provides implicit human validation of each code change — a merge decision by project maintainers confirms the intent and quality of… See the full description on the dataset page: https://huggingface.co/datasets/project-themis/git-commits-merged.openarm_teleop_udp_d435_merged_20260725_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
16
],
"names": [
"left_joint_1.pos",
"left_joint_2.pos",
"left_joint_3.pos",
"left_joint_4.pos",
"left_joint_5.pos"… See the full description on the dataset page: https://huggingface.co/datasets/nkmurst/openarm_teleop_udp_d435_merged_20260725_lerobot.merged_speech_datasetIndic-total-New-TTS-Merge
Indic Total TTS Merge
Merged TTS dataset with 13 Indic languages. All audio clips are >= 3.0 seconds duration.
Languages
assamese, bengali, english, gujarati, hindi, kannada, malayalam, marathi, nepali, odia, punjabi, tamil, telugu
Columns
audio: Audio data
text: Transcript text
duration: Duration in seconds (all >= 3.0s)
language: Language name
optim_policy_pretrain-pythia-160m_lr0.0001_bs24_wp1_wd0.01_ep0_cp35k-mergeddk1-merge-2026-03
DK-1 Merged Dataset
This dataset was created using LeRobot.
Dataset Description
Merged and deduplicated DK-1 bimanual robot dataset. All source videos are re-encoded to 640x360 h264 with 14D joint-space actions. Per-episode FPS is preserved from the source datasets.
Generated: 2026-03-27
Total episodes: 1,823
Total frames: 2,230,018
Total tasks: 17
Resolution: 640x360
Codec: h264
FPS: 30, 50 (per-episode, stored in episodes.jsonl)
Cameras: head, left_wrist… See the full description on the dataset page: https://huggingface.co/datasets/andreaskoepf/dk1-merge-2026-03.origami-v3-merged
Robotic Origami Challenge — Unified LeRobot v3.0
A single, ready-to-train LeRobot v3.0 dataset of real-world bimanual dexterous
paper-airplane folding (robot type north_ces, Sharpa Hands), consolidated from
the per-season releases of the Robotic Origami Challenge.
Provenance. The upstream release ships as 46 separate per-season datasets,
each a self-contained v3.0 tree whose episode/frame/file indices restart at 0, plus a
redundant v2.1 copy. This repo merges all lerobot3.0… See the full description on the dataset page: https://huggingface.co/datasets/suveen013/origami-v3-merged.dolma3_mix-150B-1025-merged-olmo3
dolma3_mix-150B-1025-merged (OLMo3)
Tokenized copy of the Dolma3 150B mix (dolma3_mix-150B-1025-merged) using the OLMo3 / Qwen3.5-base (q35base) tokenizer.
Format
Megatron-LM indexed binaries: paired *_text_document.bin and *_text_document.idx files (448 files total, ~613 GiB).
These are not Hugging Face datasets Arrow/Parquet shards. Load them with Megatron / Megatron-LM indexed dataset readers.
Source name
Local / blob name:… See the full description on the dataset page: https://huggingface.co/datasets/yangwang92/dolma3_mix-150B-1025-merged-olmo3.instruction_merge_set
Dataset Card for "instruction_merge_set"
本数据集由以下数据集构成:
数据(id in the merged set)
Hugging face 地址
notes
OIG (unified-任务名称) 15k
https://huggingface.co/datasets/laion/OIG
Open Instruction Generalist Dataset
Dolly databricks-dolly-15k
https://huggingface.co/datasets/databricks/databricks-dolly-15k
an open-source dataset of instruction-following records generated by thousands of Databricks employees in several of the behavioral categories
UltraChat… See the full description on the dataset page: https://huggingface.co/datasets/LinkSoul/instruction_merge_set.merged_valid
Dataset Card for 1g0rrr/merged_valid
Dataset Structure
This dataset follows the LeRobot v2.1 format with the following structure:
meta/info.json: Dataset metadata and configuration
meta/episodes.jsonl: Episode information including tasks and lengths
meta/tasks.jsonl: Task descriptions and indices
meta/episodes_stats.jsonl: Per-episode statistics
data/chunk_*/episode_*.parquet: Episode data files
videos/chunk_*/video_key_*/episode_*.mp4: Video files (if present)… See the full description on the dataset page: https://huggingface.co/datasets/1g0rrr/merged_valid.jitteredwebsites-merged-224-paraphrasedjetson1-061026-subtask-tilt-doris-mergedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos",
"tilt.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/jetson1-061026-subtask-tilt-doris-merged.merged_chaosbenchtibetan_monolingual_A_merged_123_linesin1k_clip_qwen25vl_3b_448res_256tokens_new_merged_ptNuscenes-QA-merge-front-imageUSAGE in Python
load train and valid dataset
add base_folder
