datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
openenv-python-repair
Python Repair Lab
An original OpenEnv curriculum of 1,200 deterministic Python function-repair episodes: 12 problem families, four distinct bug patterns per family, and 25 seeded case sets per pattern. There are 151,780 executable checks across the episodes. These are 48 repair patterns with data variants, not 1,200 unrelated algorithms. Tasks cover interval algorithms, rolling calculations, weighted statistics, stable deduplication, Unicode run-length encoding, Luhn checksums… See the full description on the dataset page: https://huggingface.co/datasets/Louistiti/openenv-python-repair.self_repair_gripper_screwdriver_bc
Flex-pi screwdriver pickup and tightening crops
Local LeRobot v2.1 dataset containing 800 episodes, 227,481 frames,
126.38 minutes at 30 Hz. Open review.html to browse every clip
and switch between overhead, left-wrist, and right-wrist cameras.
Task: Pick up the screwdriver and tighten the screw securing the gripper in its holder.
Contents
Original 32-dimensional observation.state and action, preserved bit-exact for retained rows. Dimension names and units follow… See the full description on the dataset page: https://huggingface.co/datasets/AdityaShah/self_repair_gripper_screwdriver_bc.retro-weave-agent-editor-repair-diffs-v0.1
RetroInstruct Weave Agent Editor Repair Diffs
This component of RetroInstruct trains weave-agent to use the WeaveEditor to fix synthetic corruptions in the vein of
the Easy Prose Repair Diffs component.
Each row in the dataset provides the pieces you need to make a synthetic episode
demonstrating the agent:
Singling out one of three files as corrupted and in need of repair
Writing out a patch to the file as either a series of WeaveEditor edit() commands or a unidiff
Observing the… See the full description on the dataset page: https://huggingface.co/datasets/jdpressman/retro-weave-agent-editor-repair-diffs-v0.1.SysMLv2_Repair_with_SLMs
SysMLv2 Repair with SLMs
Dataset used in "Automated Semantic Fault Localization in SysML v2: A Human-in-the-Loop Framework Using Knowledge-Graph Augmented LLMs", presented at INCOSE International Symposium 2026.
Dataset Structure
This dataset provides two configurations:
default: Contains train/validation/test splits used for fine-tuning small models. Samples exceeding 2048 tokens have been removed.
full: Contains complete dataset
Task
Given SysML v2 code… See the full description on the dataset page: https://huggingface.co/datasets/rohhaiil/SysMLv2_Repair_with_SLMs.bibletts-asante-twi-repaired
BibleTTS Asante Twi — Repaired Transcripts
The Asante Twi transcripts released with BibleTTS have had the
characters ɛ (U+025B) and ɔ (U+0254) stripped out. This dataset restores them.
Audio is not included. This is a drop-in replacement for the .txt files that ship with the
BibleTTS Asante Twi package, matched by clip ID.
The problem
Both are Twi vowels, and both are required by the orthography. Measured across the released
Asante Twi transcripts:
Character… See the full description on the dataset page: https://huggingface.co/datasets/danieldzikunuofmarvel/bibletts-asante-twi-repaired.tb21-eval-qwen35-action-only-20k-infra-repaired-c164-max32k-timeout2x
qwen35-action-only-20k — Terminal-Bench 2.1
Noncanonical Terminal-Bench 2.1 evaluation of violetxi/qwen35-4b-offline-echo-action-only-20k-tacc through the served
model ID qwen35-action-only-20k with Terminus-2.
Noncanonical run: timeout_multiplier=2 instead of 1.0; repair concurrency=164 exceeds 30. Do not compare this score directly with canonical TB2.1 leaderboard runs.
Result
Recorded trials: 445
Tasks / attempts: 89 × 5
Errored trials scored as zero: 250… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/tb21-eval-qwen35-action-only-20k-infra-repaired-c164-max32k-timeout2x.pixeldit-b-native-repa-w4mn-200k-20260925multi-bug-repair
CodeWalk — Multi-Bug Repair
Agentic co-located multi-bug software repair. A level-N task presents N coupled
bugs simultaneously at one repository snapshot; the agent must fix all of them so that the
union of their FAIL_TO_PASS tests passes. Part of the CodeWalk benchmark suite
(CodeWalk: Generating Coding Benchmarks by Walking a Problem Graph).
568 tasks across levels L1–L3 (1,104 bugs, 280 distinct repositories)
Every task is gold-verified: all bugs fail at the base commit… See the full description on the dataset page: https://huggingface.co/datasets/CodeWalk/multi-bug-repair.context-repair-benchmark
ThoughtDAG Context Repair Benchmark
What happens after one wrong assumption enters a long LLM conversation?
This dataset turns context editing into a measurable intervention. Each synthetic case starts with a clean fact, introduces a false update, lets the error propagate through one to three downstream turns, and then asks the same final question under five graph conditions:
clean
polluted
source_prune
subgraph_prune
recompute_descendants
The central question is not only… See the full description on the dataset page: https://huggingface.co/datasets/thoughtdag/context-repair-benchmark.D4J-Repair
Dataset Summary
D4J-Repair is a curated subset of Defects4J, containing 371 single-function Java bugs from real-world projects. Each example includes a buggy implementation, its corresponding fixed version, and unit tests for verification.
Supported Tasks
Program Repair: Fixing bugs in Java functions
Code Generation: Generating correct implementations from buggy code
Dataset Structure
Each row contains:
task_id: Unique identifier for the task (in format:… See the full description on the dataset page: https://huggingface.co/datasets/barty/D4J-Repair.gamma-g1-73-74-core140-multilingual-repair-20260617gamma-g1-72-trainable-multilingual-repair-head-20260617gamma-g1-317-am-th-semantic-unicode-repair-data-20260623
