datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
openenv-python-repair
Python Repair Lab
An original OpenEnv curriculum of 1,200 deterministic Python function-repair episodes: 12 problem families, four distinct bug patterns per family, and 25 seeded case sets per pattern. There are 151,780 executable checks across the episodes. These are 48 repair patterns with data variants, not 1,200 unrelated algorithms. Tasks cover interval algorithms, rolling calculations, weighted statistics, stable deduplication, Unicode run-length encoding, Luhn checksums… See the full description on the dataset page: https://huggingface.co/datasets/Louistiti/openenv-python-repair.polaris-53k-repaired
POLARIS-53K, label-repaired
49,289 of the 53,291 rows in
POLARIS-Project/Polaris-Dataset-53K,
with 4,580 stored answers corrected and 4,002 rows removed as unrepairable.
Measurements on the source set put its bad-label rate at roughly 15.9%
[14.3, 17.6] (two independent detectors agreeing on a 2,000-row sample).
Mislabelled rows are not uniformly distributed: they concentrate in the problems
models fail, which is exactly where a difficulty-calibration pipeline looks.… See the full description on the dataset page: https://huggingface.co/datasets/joanvelja/polaris-53k-repaired.multi-bug-repair
CodeWalk — Multi-Bug Repair
Agentic co-located multi-bug software repair. A level-N task presents N coupled
bugs simultaneously at one repository snapshot; the agent must fix all of them so that the
union of their FAIL_TO_PASS tests passes. Part of the CodeWalk benchmark suite
(CodeWalk: Generating Coding Benchmarks by Walking a Problem Graph).
568 tasks across levels L1–L3 (1,104 bugs, 280 distinct repositories)
Every task is gold-verified: all bugs fail at the base commit… See the full description on the dataset page: https://huggingface.co/datasets/CodeWalk/multi-bug-repair.
