Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01JetBrains-Research /lca-ci-builds-repair 🏟️ Long Code Arena (CI builds repair) This is the benchmark for CI builds repair task as part of the 🏟️ Long Code Arena benchmark. 🛠️ Task. Given the logs of a failed GitHub Actions workflow and the corresponding repository snapshot, repair the repository contents in order to make the workflow pass. All the data is collected from repositories published under permissive licenses (MIT, Apache-2.0, BSD-3-Clause, and BSD-2-Clause). The datapoints can be removed upon request. To… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/lca-ci-builds-repair.tabularn<1K3 likes635 downloads2y agoHugging Face02Tri1 /Repaired_videosimage100K<n<1M0 likes473 downloads3mo agoHugging Face03Gramscii-IT /semantic-repair-routing semantic-repair-routing The supervised pairs that train SemanticRepair-270M: a message somebody actually wrote, and the requests inside it restated plainly, one per line. 84,819 pairs in five languages, plus 2,515 in Italian and English aimed at what the model used to refuse. It teaches one narrow thing. An embedding router compares a question with the description of every capability it can reach. People do not write the way capabilities are described — they hedge, they… See the full description on the dataset page: https://huggingface.co/datasets/Gramscii-IT/semantic-repair-routing.texttext-generation10K<n<100K0 likes357 downloads5d agoHugging Face04ci-benchmark-user /ci-repair-bench CI-REPAIR-BENCH Overview CI-REPAIR-BENCH is a benchmark dataset for research on Continuous Integration (CI) failures and automated repair in Python repositories. The dataset contains 567 CI failure instances collected from 105 real-world GitHub repositories, all written in Python.Each instance captures a CI workflow failure, its logs, the corresponding code diff, and repository-level metadata. Dataset Statistics Programming language: Python Number… See the full description on the dataset page: https://huggingface.co/datasets/ci-benchmark-user/ci-repair-bench.tabularn<1K2 likes306 downloads19d agoHugging Face05nus-yam /ex-repairtabular1M<n<10M3 likes252 downloads3y agoHugging Face06Louistiti /openenv-python-repair Python Repair Lab An original OpenEnv curriculum of 1,200 deterministic Python function-repair episodes: 12 problem families, four distinct bug patterns per family, and 25 seeded case sets per pattern. There are 151,780 executable checks across the episodes. These are 48 repair patterns with data variants, not 1,200 unrelated algorithms. Tasks cover interval algorithms, rolling calculations, weighted statistics, stable deduplication, Unicode run-length encoding, Luhn checksums… See the full description on the dataset page: https://huggingface.co/datasets/Louistiti/openenv-python-repair.tabulartext-generation1K<n<10K0 likes204 downloads1d agoHugging Face07Ichlibitiche /appliancedb-error-codes-repair-database ApplianceDB: Home Appliance Error Codes & Ranked Repairs Full dataset: appliancedb.dataengineered.io · $99 one-time (Repair Intelligence Snapshot: commercial licence + SQLite and Parquet builds; the same rows as this sample) → Buy on Stripe · the same sample on Kaggle Relational database mapping 438 home-appliance error codes across 13 brands and 26 (brand, appliance-type) pairs to 288 ranked repair procedures with DIY difficulty tiers. Every code is identified by its… See the full description on the dataset page: https://huggingface.co/datasets/Ichlibitiche/appliancedb-error-codes-repair-database.tabular1K<n<10K0 likes189 downloads7d agoHugging Face08bcckfdn /llama-repair-v2text1K<n<10K0 likes176 downloads2mo agoHugging Face09rmems /data-pipeline-repair-trajectories Data Pipeline Repair Trajectories Rights & intended use: legacy public research corpus / portfolio artifact. Hosted frontier-model outputs are research-only inputs under project policy (synthetic-factory#161): intended_use: research_only, project_training_policy: blocked. Not training data for any model-weight update. Machine-readable record: rights.json. Release status: The raw, uncurated payload is now published under data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/data-pipeline-repair-trajectories.text1K<n<10K0 likes169 downloads23d agoHugging Face10Reza2kn /uncgpt-conversations-semantic-approved-1p25-repaired-paperclip UncGPT — Semantic-Approved 1.25σ Conversations (Leak-Repaired) The 1.25σ semantic-gate cohort with uncle-diary leakage repaired and normalized diary fields. The auditable replacement for the older paperclip_all_1803 source that an earlier audit flagged for visible diary leakage. Part of the UncGPT NeurIPS 2026 Competition collection. Config approved_manifest (default): one row per approved conversation, with metadata + path back to the full-schema JSON.… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/uncgpt-conversations-semantic-approved-1p25-repaired-paperclip.texttext-generation1K<n<10K0 likes168 downloads5mo agoHugging Face11unfundedResearcher /Minecraft-GLB2Schem-RepairPairs-v1 unfundedResearcher/Minecraft-GLB2Schem-RepairPairs-v1 Paired (generated input, ground-truth target) Minecraft schematics for training a model that turns an approximate voxelisation into a real build. What a sample is Each sample is three files inside a WebDataset TAR shard: File Meaning <id>.input.schem GENERATED. Produced by voxelising the source .glb. Approximate and noisy. <id>.target.schem GROUND TRUTH. The original schematic, copied byte-for-byte… See the full description on the dataset page: https://huggingface.co/datasets/unfundedResearcher/Minecraft-GLB2Schem-RepairPairs-v1.textothern<1K0 likes165 downloads1mo agoHugging Face12prithivMLmods /OpenCOCO-I2T-Repack-Tiny OpenCOCO-I2T-Repack-Tiny OpenCOCO-I2T-Repack-Tiny is a compact image-to-text / image-text-to-text captioning dataset containing 158,958 image samples sourced from the COCO dataset and repackaged into a lightweight format suitable for vision-language model (VLM) fine-tuning. The dataset contains synthesized responses generated using a custom Qwen3.5 multimodal captioning pipeline. The input images undergo lossless image compression to significantly reduce the overall storage… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/OpenCOCO-I2T-Repack-Tiny.imageimage-text-to-text100K<n<1M3 likes157 downloads2mo agoHugging Face13AdityaShah /self_repair_gripper_screwdriver_bc Flex-pi screwdriver pickup and tightening crops Local LeRobot v2.1 dataset containing 800 episodes, 227,481 frames, 126.38 minutes at 30 Hz. Open review.html to browse every clip and switch between overhead, left-wrist, and right-wrist cameras. Task: Pick up the screwdriver and tighten the screw securing the gripper in its holder. Contents Original 32-dimensional observation.state and action, preserved bit-exact for retained rows. Dimension names and units follow… See the full description on the dataset page: https://huggingface.co/datasets/AdityaShah/self_repair_gripper_screwdriver_bc.tabularroboticsn<1K0 likes155 downloads7d agoHugging Face14SWE-Swiss /SWESwiss-Repair-RL-SWEGym-SWESmith-12K Overview RL dataset for training SWE-Swiss models on the repair task. The prompts are based on issues from SWE-Gym and SWE-smith. To create a challenging task, the code content in each prompt consists of two components: "oracle" files, which are the ground-truth files requiring a patch, and "distractor" files, which are plausible but incorrect files predicted by an LLM. Citation @misc{SWESwiss2025, title = {SWE-Swiss: A Multi-Task Fine-Tuning and RL Recipe for… See the full description on the dataset page: https://huggingface.co/datasets/SWE-Swiss/SWESwiss-Repair-RL-SWEGym-SWESmith-12K.text10K<n<100K4 likes154 downloads1y agoHugging Face15wAI-org /swerl-tmax-15k-repairs-gpt-5-6-sol swerl-tmax-15k task repairs (gpt-5-6-sol) Proposed repairs for defective tasks in hamishivi/swerl-tmax-15k, generated from the audit labels in wAI-org/swerl-tmax-15k-rubric-gpt-5-6-sol. These repairs are UNVALIDATED No repair here has been executed, and none has been shown to be solvable. Every repair was checked mechanically — valid bash, still writes a reward, does not delete a path the instruction needs. None was checked empirically. A hardened verifier can be… See the full description on the dataset page: https://huggingface.co/datasets/wAI-org/swerl-tmax-15k-repairs-gpt-5-6-sol.texttext-generation10K<n<100K0 likes146 downloads29d agoHugging Face16schneiderkamplab /dfm11-toolace-native-tool-use-repaired dfm11-toolace-native-tool-use-repaired ToolACE conversations with declared-name parsing and complete parallel result binding. This is a DFM11 replacement for schneiderkamplab/dfm10-toolace-native-tool-use. All rows pass exhaustive structural validation. See metadata/manifest.json. text10K<n<100K0 likes145 downloads1mo agoHugging Face17empirischtech /Palace-Config-Repair Palace Configuration Repair Benchmark Execution-graded repair tasks for configuration files of Palace, an open-source finite-element solver for computational electromagnetics. Each task gives a model a perturbed Palace JSON configuration and asks it to return a corrected one. A repair counts as correct only if it conforms to the schema, is accepted by the solver, and reproduces the reference outputs of the original case when Palace v0.14.0 runs it. A configuration that is valid… See the full description on the dataset page: https://huggingface.co/datasets/empirischtech/Palace-Config-Repair.texttext-generationn<1K0 likes138 downloads9d agoHugging Face18Ichlibitiche /mechanicdb-obd2-repair-sample 🔧 MechanicDB — OBD-II Diagnostic & Repair Database (Free Sample) Full dataset: mechanicdb.dataengineered.io · $49 Standard (SAE) · $149 OEM Complete, one-time → Buy Standard · Buy OEM Complete · the same sample on Kaggle The free developer sample of MechanicDB: an automotive dataset mapping OBD-II Diagnostic Trouble Codes (DTCs) to authored repair procedures in consideration order with DIY difficulty ratings, aftermarket parts-cost ranges (USD), labor-hour estimates, and… See the full description on the dataset page: https://huggingface.co/datasets/Ichlibitiche/mechanicdb-obd2-repair-sample.tabular1K<n<10K0 likes121 downloads2d agoHugging Face19VmaxRL /SWEUniverse-Repaired-Indist-full-not-SWE-bench-pro-matched VmaxRL/SWEUniverse-Repaired-Indist-full-not-SWE-bench-pro-matched This dataset contains a 350-row subset selected from the Indist SWEUniverse training rows. Selection policy: three-way repo overlap with Bugpilot and LM-Modify, deduped by repo plus introduction patch, then balanced round-robin across overlapping repos. Rows: 350 Selected repos: 19 Deduped overlap capacity: 468 Source dataset: VmaxRL/SWEUniverse-Repaired-Indist-full-not-SWE-bench-pro-matched texttext-generationn<1K0 likes120 downloads5mo agoHugging Face20ASSERT-KTH /repairllama-datasets RepairLLaMA - Datasets Contains the processed fine-tuning datasets for RepairLLaMA. Instructions to explore the dataset To load the dataset, you must define which revision (i.e., which input/output representation pair) you want to load. from datasets import load_dataset # Load ir1xor1 dataset = load_dataset("ASSERT-KTH/repairllama-datasets", "ir1xor1") # Load irXxorY dataset = load_dataset("ASSERT-KTH/repairllama-datasets", "irXxorY") Citation If you use… See the full description on the dataset page: https://huggingface.co/datasets/ASSERT-KTH/repairllama-datasets.texttext-generation100K<n<1M3 likes118 downloads2y agoHugging Face21rmems /python-function-repair-20261006 Python Function Repair Records (2026-10-06, deterministic) Rights & intended use: public research snapshot, not training data. Deterministic, project-owned generation over real MIT-licensed upstream programs (TheAlgorithms/Python @ 2067ce6dfb3, MIT — see provenance.json): intended_use: training_candidate, project_training_policy: allowed per the factory registry — but this raw snapshot is not training-ready (release-status.json: release_stage: raw_uncurated_public… See the full description on the dataset page: https://huggingface.co/datasets/rmems/python-function-repair-20261006.textn<1K0 likes110 downloads3d agoHugging Face22emgena /emgena_rust_memory_leak_arc_cyclic_repair_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/emgena_rust_memory_leak_arc_cyclic_repair_teaser.textn<1K1 likes107 downloads23d agoHugging Face23RussianNLP /repa Dataset Card for REPA Image Source Dataset Description REPA is a Russian language dataset which consists of 1k user queries categorized into nine types, along with responses from six open-source instruction-finetuned Russian LLMs. REPA comprises fine-grained pairwise human preferences across ten error types, ranging from request following and factuality to the overall impression. Each data instance consists of a query and two LLM responses manually annotated to… See the full description on the dataset page: https://huggingface.co/datasets/RussianNLP/repa.text1K<n<10K4 likes94 downloads2y agoHugging Face24rmems /db-migration-repair-trajectories Db Migration Repair Trajectories Rights & intended use: legacy public research corpus / portfolio artifact. Hosted frontier-model outputs are research-only inputs under project policy (synthetic-factory#161): intended_use: research_only, project_training_policy: blocked. Not training data for any model-weight update. Machine-readable record: rights.json. Release status: The raw, uncurated payload is now published under data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/db-migration-repair-trajectories.text1K<n<10K0 likes87 downloads1mo agoHugging Face25dougdotcon /douvras-lean-proof-repair Douvras Lean Proof Repair Corpus Exemplos sintéticos de erros comuns de reparo em Lean: importação ausente, incompatibilidade de tipos, falha de tática, meta não resolvida, reescrita inválida e prova reflexiva. Os snippets não foram executados no compilador (proof_status: NOT_EXECUTED); portanto o corpus não prova nenhum teorema e não substitui validação com uma versão específica do Mathlib. texttext-classificationn<1K0 likes84 downloads27d agoHugging Face26referencesource /state-right-to-repair-laws State Right-to-Repair Laws: Coverage, Requirements, and Effective Dates Canonical, always-current version: https://referencesource.org/state-right-to-repair-laws/ Machine-readable: https://referencesource.org/state-right-to-repair-laws/data.json — this mirror is a point-in-time copy. Last verified: 2026-10-06 Stale after: 2026-11-13 (past this date, prefer the canonical copy — it re-verifies on a cadence this snapshot does not) Records: 6 Which US states have enacted… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/state-right-to-repair-laws.textn<1K0 likes82 downloads4d agoHugging Face27shravandoda /TobiaSVG-repair TobiaSVG Repair TobiaSVG Repair contains 44,802 synthetic SVG repair pairs. Each row pairs a corrupted SVG with its clean target. Raster images are rendered when examples are loaded and are not stored. Sources And Splits Subset Source Rows License vfig_diagrams VFIG-Data 33,624 ODC-BY 1.0 vfig_shapes VFIG-Data 8,884 ODC-BY 1.0 animal_illustrations SVG Animal Illustrations 2,294 CC0 1.0 Splits contain 35,833 training, 4,423 test, and 4,546… See the full description on the dataset page: https://huggingface.co/datasets/shravandoda/TobiaSVG-repair.texttext-generation10K<n<100K0 likes79 downloads3mo agoHugging Face28emgena /omnimcp_react_hydration_repair_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_react_hydration_repair_teaser.texttext-generationn<1K0 likes78 downloads23d agoHugging Face29AdaMLLab /mlqa_repairedThis is a repaired version of https://huggingface.co/datasets/facebook/mlqa made compatible with datasets>=4.X (no arbitrary code execution). textquestion-answering100K<n<1M0 likes74 downloads10mo agoHugging Face30emgena /omnimcp_pytest_traceback_repair_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_pytest_traceback_repair_teaser.texttext-generationn<1K1 likes72 downloads23d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.