datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lca-ci-builds-repair
🏟️ Long Code Arena (CI builds repair)
This is the benchmark for CI builds repair task as part of the
🏟️ Long Code Arena benchmark.
🛠️ Task. Given the logs of a failed GitHub Actions workflow and the corresponding repository snapshot,
repair the repository contents in order to make the workflow pass.
All the data is collected from repositories published under permissive licenses (MIT, Apache-2.0, BSD-3-Clause, and BSD-2-Clause). The datapoints can be removed upon request.
To… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/lca-ci-builds-repair.Repaired_videossemantic-repair-routing
semantic-repair-routing
The supervised pairs that train
SemanticRepair-270M:
a message somebody actually wrote, and the requests inside it restated
plainly, one per line. 84,819 pairs in five languages, plus 2,515 in
Italian and English aimed at what the model used to refuse.
It teaches one narrow thing. An embedding router compares a question with
the description of every capability it can reach. People do not write the
way capabilities are described — they hedge, they… See the full description on the dataset page: https://huggingface.co/datasets/Gramscii-IT/semantic-repair-routing.ci-repair-bench
CI-REPAIR-BENCH
Overview
CI-REPAIR-BENCH is a benchmark dataset for research on Continuous Integration (CI) failures and automated repair in Python repositories.
The dataset contains 567 CI failure instances collected from 105 real-world GitHub repositories, all written in Python.Each instance captures a CI workflow failure, its logs, the corresponding code diff, and repository-level metadata.
Dataset Statistics
Programming language: Python
Number… See the full description on the dataset page: https://huggingface.co/datasets/ci-benchmark-user/ci-repair-bench.ex-repairopenenv-python-repair
Python Repair Lab
An original OpenEnv curriculum of 1,200 deterministic Python function-repair episodes: 12 problem families, four distinct bug patterns per family, and 25 seeded case sets per pattern. There are 151,780 executable checks across the episodes. These are 48 repair patterns with data variants, not 1,200 unrelated algorithms. Tasks cover interval algorithms, rolling calculations, weighted statistics, stable deduplication, Unicode run-length encoding, Luhn checksums… See the full description on the dataset page: https://huggingface.co/datasets/Louistiti/openenv-python-repair.appliancedb-error-codes-repair-database
ApplianceDB: Home Appliance Error Codes & Ranked Repairs
Full dataset: appliancedb.dataengineered.io · $99 one-time (Repair Intelligence Snapshot: commercial licence + SQLite and Parquet builds; the same rows as this sample) → Buy on Stripe · the same sample on Kaggle
Relational database mapping 438 home-appliance error codes across 13 brands and 26 (brand, appliance-type) pairs to 288 ranked repair procedures with DIY difficulty tiers. Every code is identified by its… See the full description on the dataset page: https://huggingface.co/datasets/Ichlibitiche/appliancedb-error-codes-repair-database.llama-repair-v2data-pipeline-repair-trajectories
Data Pipeline Repair Trajectories
Rights & intended use: legacy public research corpus / portfolio
artifact. Hosted frontier-model outputs are research-only inputs under
project policy (synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json.
Release status: The raw, uncurated payload is now published under
data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/data-pipeline-repair-trajectories.uncgpt-conversations-semantic-approved-1p25-repaired-paperclip
UncGPT — Semantic-Approved 1.25σ Conversations (Leak-Repaired)
The 1.25σ semantic-gate cohort with uncle-diary leakage repaired and normalized diary fields. The auditable replacement for the older paperclip_all_1803 source that an earlier audit flagged for visible diary leakage.
Part of the UncGPT NeurIPS 2026 Competition collection.
Config
approved_manifest (default): one row per approved conversation, with metadata + path back to the full-schema JSON.… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/uncgpt-conversations-semantic-approved-1p25-repaired-paperclip.Minecraft-GLB2Schem-RepairPairs-v1
unfundedResearcher/Minecraft-GLB2Schem-RepairPairs-v1
Paired (generated input, ground-truth target) Minecraft schematics for training a
model that turns an approximate voxelisation into a real build.
What a sample is
Each sample is three files inside a WebDataset TAR shard:
File
Meaning
<id>.input.schem
GENERATED. Produced by voxelising the source .glb. Approximate and noisy.
<id>.target.schem
GROUND TRUTH. The original schematic, copied byte-for-byte… See the full description on the dataset page: https://huggingface.co/datasets/unfundedResearcher/Minecraft-GLB2Schem-RepairPairs-v1.OpenCOCO-I2T-Repack-Tiny
OpenCOCO-I2T-Repack-Tiny
OpenCOCO-I2T-Repack-Tiny is a compact image-to-text / image-text-to-text captioning dataset containing 158,958 image samples sourced from the COCO dataset and repackaged into a lightweight format suitable for vision-language model (VLM) fine-tuning. The dataset contains synthesized responses generated using a custom Qwen3.5 multimodal captioning pipeline. The input images undergo lossless image compression to significantly reduce the overall storage… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/OpenCOCO-I2T-Repack-Tiny.self_repair_gripper_screwdriver_bc
Flex-pi screwdriver pickup and tightening crops
Local LeRobot v2.1 dataset containing 800 episodes, 227,481 frames,
126.38 minutes at 30 Hz. Open review.html to browse every clip
and switch between overhead, left-wrist, and right-wrist cameras.
Task: Pick up the screwdriver and tighten the screw securing the gripper in its holder.
Contents
Original 32-dimensional observation.state and action, preserved bit-exact for retained rows. Dimension names and units follow… See the full description on the dataset page: https://huggingface.co/datasets/AdityaShah/self_repair_gripper_screwdriver_bc.SWESwiss-Repair-RL-SWEGym-SWESmith-12K
Overview
RL dataset for training SWE-Swiss models on the repair task. The prompts are based on issues from SWE-Gym and SWE-smith. To create a challenging task, the code content in each prompt consists of two components: "oracle" files, which are the ground-truth files requiring a patch, and "distractor" files, which are plausible but incorrect files predicted by an LLM.
Citation
@misc{SWESwiss2025,
title = {SWE-Swiss: A Multi-Task Fine-Tuning and RL Recipe for… See the full description on the dataset page: https://huggingface.co/datasets/SWE-Swiss/SWESwiss-Repair-RL-SWEGym-SWESmith-12K.swerl-tmax-15k-repairs-gpt-5-6-sol
swerl-tmax-15k task repairs (gpt-5-6-sol)
Proposed repairs for defective tasks in hamishivi/swerl-tmax-15k, generated from
the audit labels in
wAI-org/swerl-tmax-15k-rubric-gpt-5-6-sol.
These repairs are UNVALIDATED
No repair here has been executed, and none has been shown to be solvable.
Every repair was checked mechanically — valid bash, still writes a reward,
does not delete a path the instruction needs. None was checked empirically.
A hardened verifier can be… See the full description on the dataset page: https://huggingface.co/datasets/wAI-org/swerl-tmax-15k-repairs-gpt-5-6-sol.dfm11-toolace-native-tool-use-repaired
dfm11-toolace-native-tool-use-repaired
ToolACE conversations with declared-name parsing and complete parallel result binding.
This is a DFM11 replacement for schneiderkamplab/dfm10-toolace-native-tool-use. All rows pass exhaustive structural validation. See metadata/manifest.json.
Palace-Config-Repair
Palace Configuration Repair Benchmark
Execution-graded repair tasks for configuration files of
Palace, an open-source finite-element solver for
computational electromagnetics. Each task gives a model a perturbed Palace JSON configuration and
asks it to return a corrected one. A repair counts as correct only if it conforms to the schema,
is accepted by the solver, and reproduces the reference outputs of the original case when
Palace v0.14.0 runs it. A configuration that is valid… See the full description on the dataset page: https://huggingface.co/datasets/empirischtech/Palace-Config-Repair.mechanicdb-obd2-repair-sample
🔧 MechanicDB — OBD-II Diagnostic & Repair Database (Free Sample)
Full dataset: mechanicdb.dataengineered.io · $49 Standard (SAE) · $149 OEM Complete, one-time → Buy Standard · Buy OEM Complete · the same sample on Kaggle
The free developer sample of MechanicDB: an automotive dataset mapping OBD-II
Diagnostic Trouble Codes (DTCs) to authored repair procedures in consideration order with DIY
difficulty ratings, aftermarket parts-cost ranges (USD), labor-hour estimates,
and… See the full description on the dataset page: https://huggingface.co/datasets/Ichlibitiche/mechanicdb-obd2-repair-sample.SWEUniverse-Repaired-Indist-full-not-SWE-bench-pro-matched
VmaxRL/SWEUniverse-Repaired-Indist-full-not-SWE-bench-pro-matched
This dataset contains a 350-row subset selected from the Indist SWEUniverse training rows.
Selection policy: three-way repo overlap with Bugpilot and LM-Modify, deduped by repo plus introduction patch, then balanced round-robin across overlapping repos.
Rows: 350
Selected repos: 19
Deduped overlap capacity: 468
Source dataset: VmaxRL/SWEUniverse-Repaired-Indist-full-not-SWE-bench-pro-matched
repairllama-datasets
RepairLLaMA - Datasets
Contains the processed fine-tuning datasets for RepairLLaMA.
Instructions to explore the dataset
To load the dataset, you must define which revision (i.e., which input/output representation pair) you want to load.
from datasets import load_dataset
# Load ir1xor1
dataset = load_dataset("ASSERT-KTH/repairllama-datasets", "ir1xor1")
# Load irXxorY
dataset = load_dataset("ASSERT-KTH/repairllama-datasets", "irXxorY")
Citation
If you use… See the full description on the dataset page: https://huggingface.co/datasets/ASSERT-KTH/repairllama-datasets.python-function-repair-20261006
Python Function Repair Records (2026-10-06, deterministic)
Rights & intended use: public research snapshot, not training data.
Deterministic, project-owned generation over real MIT-licensed upstream
programs (TheAlgorithms/Python @ 2067ce6dfb3, MIT — see
provenance.json): intended_use: training_candidate,
project_training_policy: allowed per the factory registry — but this raw
snapshot is not training-ready (release-status.json:
release_stage: raw_uncurated_public… See the full description on the dataset page: https://huggingface.co/datasets/rmems/python-function-repair-20261006.emgena_rust_memory_leak_arc_cyclic_repair_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/emgena_rust_memory_leak_arc_cyclic_repair_teaser.repa
Dataset Card for REPA
Image Source
Dataset Description
REPA is a Russian language dataset which consists of 1k user queries categorized into nine types, along with responses from six open-source instruction-finetuned Russian LLMs. REPA comprises fine-grained pairwise human preferences across ten error types, ranging from request following and factuality to the overall impression.
Each data instance consists of a query and two LLM responses manually annotated to… See the full description on the dataset page: https://huggingface.co/datasets/RussianNLP/repa.db-migration-repair-trajectories
Db Migration Repair Trajectories
Rights & intended use: legacy public research corpus / portfolio
artifact. Hosted frontier-model outputs are research-only inputs under
project policy (synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json.
Release status: The raw, uncurated payload is now published under
data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/db-migration-repair-trajectories.douvras-lean-proof-repair
Douvras Lean Proof Repair Corpus
Exemplos sintéticos de erros comuns de reparo em Lean: importação ausente, incompatibilidade de
tipos, falha de tática, meta não resolvida, reescrita inválida e prova reflexiva. Os snippets não
foram executados no compilador (proof_status: NOT_EXECUTED); portanto o corpus não prova nenhum
teorema e não substitui validação com uma versão específica do Mathlib.
state-right-to-repair-laws
State Right-to-Repair Laws: Coverage, Requirements, and Effective Dates
Canonical, always-current version: https://referencesource.org/state-right-to-repair-laws/
Machine-readable: https://referencesource.org/state-right-to-repair-laws/data.json — this mirror is a point-in-time copy.
Last verified: 2026-10-06
Stale after: 2026-11-13 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 6
Which US states have enacted… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/state-right-to-repair-laws.TobiaSVG-repair
TobiaSVG Repair
TobiaSVG Repair contains 44,802 synthetic SVG repair pairs. Each row pairs a
corrupted SVG with its clean target. Raster images are rendered when examples
are loaded and are not stored.
Sources And Splits
Subset
Source
Rows
License
vfig_diagrams
VFIG-Data
33,624
ODC-BY 1.0
vfig_shapes
VFIG-Data
8,884
ODC-BY 1.0
animal_illustrations
SVG Animal Illustrations
2,294
CC0 1.0
Splits contain 35,833 training, 4,423 test, and 4,546… See the full description on the dataset page: https://huggingface.co/datasets/shravandoda/TobiaSVG-repair.omnimcp_react_hydration_repair_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_react_hydration_repair_teaser.mlqa_repairedThis is a repaired version of https://huggingface.co/datasets/facebook/mlqa made compatible with datasets>=4.X (no arbitrary code execution).
omnimcp_pytest_traceback_repair_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_pytest_traceback_repair_teaser.
