Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Gramscii-IT /semantic-repair-routing semantic-repair-routing The supervised pairs that train SemanticRepair-270M: a message somebody actually wrote, and the requests inside it restated plainly, one per line. 84,819 pairs in five languages, plus 2,515 in Italian and English aimed at what the model used to refuse. It teaches one narrow thing. An embedding router compares a question with the description of every capability it can reach. People do not write the way capabilities are described — they hedge, they… See the full description on the dataset page: https://huggingface.co/datasets/Gramscii-IT/semantic-repair-routing.texttext-generation10K<n<100K0 likes357 downloads5d agoHugging Face02Louistiti /openenv-python-repair Python Repair Lab An original OpenEnv curriculum of 1,200 deterministic Python function-repair episodes: 12 problem families, four distinct bug patterns per family, and 25 seeded case sets per pattern. There are 151,780 executable checks across the episodes. These are 48 repair patterns with data variants, not 1,200 unrelated algorithms. Tasks cover interval algorithms, rolling calculations, weighted statistics, stable deduplication, Unicode run-length encoding, Luhn checksums… See the full description on the dataset page: https://huggingface.co/datasets/Louistiti/openenv-python-repair.tabulartext-generation1K<n<10K0 likes204 downloads23h agoHugging Face03rmems /data-pipeline-repair-trajectories Data Pipeline Repair Trajectories Rights & intended use: legacy public research corpus / portfolio artifact. Hosted frontier-model outputs are research-only inputs under project policy (synthetic-factory#161): intended_use: research_only, project_training_policy: blocked. Not training data for any model-weight update. Machine-readable record: rights.json. Release status: The raw, uncurated payload is now published under data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/data-pipeline-repair-trajectories.text1K<n<10K0 likes169 downloads23d agoHugging Face04Reza2kn /uncgpt-conversations-semantic-approved-1p25-repaired-paperclip UncGPT — Semantic-Approved 1.25σ Conversations (Leak-Repaired) The 1.25σ semantic-gate cohort with uncle-diary leakage repaired and normalized diary fields. The auditable replacement for the older paperclip_all_1803 source that an earlier audit flagged for visible diary leakage. Part of the UncGPT NeurIPS 2026 Competition collection. Config approved_manifest (default): one row per approved conversation, with metadata + path back to the full-schema JSON.… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/uncgpt-conversations-semantic-approved-1p25-repaired-paperclip.texttext-generation1K<n<10K0 likes168 downloads5mo agoHugging Face05unfundedResearcher /Minecraft-GLB2Schem-RepairPairs-v1 unfundedResearcher/Minecraft-GLB2Schem-RepairPairs-v1 Paired (generated input, ground-truth target) Minecraft schematics for training a model that turns an approximate voxelisation into a real build. What a sample is Each sample is three files inside a WebDataset TAR shard: File Meaning <id>.input.schem GENERATED. Produced by voxelising the source .glb. Approximate and noisy. <id>.target.schem GROUND TRUTH. The original schematic, copied byte-for-byte… See the full description on the dataset page: https://huggingface.co/datasets/unfundedResearcher/Minecraft-GLB2Schem-RepairPairs-v1.textothern<1K0 likes165 downloads1mo agoHugging Face06AdityaShah /self_repair_gripper_screwdriver_bc Flex-pi screwdriver pickup and tightening crops Local LeRobot v2.1 dataset containing 800 episodes, 227,481 frames, 126.38 minutes at 30 Hz. Open review.html to browse every clip and switch between overhead, left-wrist, and right-wrist cameras. Task: Pick up the screwdriver and tighten the screw securing the gripper in its holder. Contents Original 32-dimensional observation.state and action, preserved bit-exact for retained rows. Dimension names and units follow… See the full description on the dataset page: https://huggingface.co/datasets/AdityaShah/self_repair_gripper_screwdriver_bc.tabularroboticsn<1K0 likes155 downloads7d agoHugging Face07schneiderkamplab /dfm11-toolace-native-tool-use-repaired dfm11-toolace-native-tool-use-repaired ToolACE conversations with declared-name parsing and complete parallel result binding. This is a DFM11 replacement for schneiderkamplab/dfm10-toolace-native-tool-use. All rows pass exhaustive structural validation. See metadata/manifest.json. text10K<n<100K0 likes145 downloads1mo agoHugging Face08empirischtech /Palace-Config-Repair Palace Configuration Repair Benchmark Execution-graded repair tasks for configuration files of Palace, an open-source finite-element solver for computational electromagnetics. Each task gives a model a perturbed Palace JSON configuration and asks it to return a corrected one. A repair counts as correct only if it conforms to the schema, is accepted by the solver, and reproduces the reference outputs of the original case when Palace v0.14.0 runs it. A configuration that is valid… See the full description on the dataset page: https://huggingface.co/datasets/empirischtech/Palace-Config-Repair.texttext-generationn<1K0 likes138 downloads9d agoHugging Face09VmaxRL /SWEUniverse-Repaired-Indist-full-not-SWE-bench-pro-matched VmaxRL/SWEUniverse-Repaired-Indist-full-not-SWE-bench-pro-matched This dataset contains a 350-row subset selected from the Indist SWEUniverse training rows. Selection policy: three-way repo overlap with Bugpilot and LM-Modify, deduped by repo plus introduction patch, then balanced round-robin across overlapping repos. Rows: 350 Selected repos: 19 Deduped overlap capacity: 468 Source dataset: VmaxRL/SWEUniverse-Repaired-Indist-full-not-SWE-bench-pro-matched texttext-generationn<1K0 likes120 downloads5mo agoHugging Face10ASSERT-KTH /repairllama-datasets RepairLLaMA - Datasets Contains the processed fine-tuning datasets for RepairLLaMA. Instructions to explore the dataset To load the dataset, you must define which revision (i.e., which input/output representation pair) you want to load. from datasets import load_dataset # Load ir1xor1 dataset = load_dataset("ASSERT-KTH/repairllama-datasets", "ir1xor1") # Load irXxorY dataset = load_dataset("ASSERT-KTH/repairllama-datasets", "irXxorY") Citation If you use… See the full description on the dataset page: https://huggingface.co/datasets/ASSERT-KTH/repairllama-datasets.texttext-generation100K<n<1M3 likes118 downloads2y agoHugging Face11rmems /db-migration-repair-trajectories Db Migration Repair Trajectories Rights & intended use: legacy public research corpus / portfolio artifact. Hosted frontier-model outputs are research-only inputs under project policy (synthetic-factory#161): intended_use: research_only, project_training_policy: blocked. Not training data for any model-weight update. Machine-readable record: rights.json. Release status: The raw, uncurated payload is now published under data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/db-migration-repair-trajectories.text1K<n<10K0 likes87 downloads1mo agoHugging Face12dougdotcon /douvras-lean-proof-repair Douvras Lean Proof Repair Corpus Exemplos sintéticos de erros comuns de reparo em Lean: importação ausente, incompatibilidade de tipos, falha de tática, meta não resolvida, reescrita inválida e prova reflexiva. Os snippets não foram executados no compilador (proof_status: NOT_EXECUTED); portanto o corpus não prova nenhum teorema e não substitui validação com uma versão específica do Mathlib. texttext-classificationn<1K0 likes84 downloads27d agoHugging Face13referencesource /state-right-to-repair-laws State Right-to-Repair Laws: Coverage, Requirements, and Effective Dates Canonical, always-current version: https://referencesource.org/state-right-to-repair-laws/ Machine-readable: https://referencesource.org/state-right-to-repair-laws/data.json — this mirror is a point-in-time copy. Last verified: 2026-10-06 Stale after: 2026-11-13 (past this date, prefer the canonical copy — it re-verifies on a cadence this snapshot does not) Records: 6 Which US states have enacted… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/state-right-to-repair-laws.textn<1K0 likes82 downloads4d agoHugging Face140xkamal7 /code-contract-repairtext1K<n<10K0 likes67 downloads2mo agoHugging Face15NadavSalem /Titanius-1.2-sft-repair Titanius-1.2-sft-repair Synthetic instructions used in the two mixed-data continuations of Titanius-1.2-128m-sft-fp16. They target short answers, arithmetic, copying, JSON, facts, concise definitions and multi-turn recall. Config Training conversations Used by phase1 5,917 First continuation, selected at step 5,500 phase2 16,527 Second continuation, selected at step 6,000 Each config is a complete synthetic pool for its phase. They overlap; do not concatenate… See the full description on the dataset page: https://huggingface.co/datasets/NadavSalem/Titanius-1.2-sft-repair.texttext-generation10K<n<100K0 likes66 downloads5d agoHugging Face16dipenbhuva /home-diy-repair-qa Home DIY Repair Q&A A synthetic dataset of 5,000 Q&A pairs covering common home DIY repair scenarios. Each example includes a detailed step-by-step answer, required tools, safety warnings, and practical tips. Dataset Purpose This dataset is built for: Instruction fine-tuning — train language models to give detailed, safe, and actionable home repair guidance Retrieval-Augmented Generation (RAG) — build a knowledge base for home repair assistants Question answering — train… See the full description on the dataset page: https://huggingface.co/datasets/dipenbhuva/home-diy-repair-qa.textquestion-answering1K<n<10K1 likes56 downloads7mo agoHugging Face17jdpressman /retro-weave-agent-editor-repair-diffs-v0.1 RetroInstruct Weave Agent Editor Repair Diffs This component of RetroInstruct trains weave-agent to use the WeaveEditor to fix synthetic corruptions in the vein of the Easy Prose Repair Diffs component. Each row in the dataset provides the pieces you need to make a synthetic episode demonstrating the agent: Singling out one of three files as corrupted and in need of repair Writing out a patch to the file as either a series of WeaveEditor edit() commands or a unidiff Observing the… See the full description on the dataset page: https://huggingface.co/datasets/jdpressman/retro-weave-agent-editor-repair-diffs-v0.1.tabularn<1K0 likes55 downloads2y agoHugging Face18rohhaiil /SysMLv2_Repair_with_SLMs SysMLv2 Repair with SLMs Dataset used in "Automated Semantic Fault Localization in SysML v2: A Human-in-the-Loop Framework Using Knowledge-Graph Augmented LLMs", presented at INCOSE International Symposium 2026. Dataset Structure This dataset provides two configurations: default: Contains train/validation/test splits used for fine-tuning small models. Samples exceeding 2048 tokens have been removed. full: Contains complete dataset Task Given SysML v2 code… See the full description on the dataset page: https://huggingface.co/datasets/rohhaiil/SysMLv2_Repair_with_SLMs.tabular10K<n<100K0 likes53 downloads6mo agoHugging Face19skonml /code-contract-repair APIContractRepair APIContractRepair is a provenance-tracked instruction-tuning dataset for software engineers and code-model researchers who need contract-faithful, minimal repairs with tests that distinguish a broken implementation from its fix. Magicoder-OSS-Instruct-75K supplies real function-identifier seeds, but it does not provide these documented contracts, deliberately buggy implementations, minimal corrected implementations, or paired regression tests. This release… See the full description on the dataset page: https://huggingface.co/datasets/skonml/code-contract-repair.texttext-generation1K<n<10K0 likes53 downloads2mo agoHugging Face20Benitoow /OfficeSmith-PPTX-Repair OfficeSmith PPTX Repair Deterministically degraded PPTX IR objects paired with validated repairs. Dataset summary This dataset is part of the OfficeSmith collection for training models to plan, build, clarify, critique, and repair editable business presentations. It contains observable outputs only: no hidden chain of thought, secret benchmark prompt, personal data, or API credential is included. Train rows: 160 Validation rows: 0 Test rows: 0 Languages: French… See the full description on the dataset page: https://huggingface.co/datasets/Benitoow/OfficeSmith-PPTX-Repair.texttext-generationn<1K0 likes47 downloads2mo agoHugging Face21schneiderkamplab /dfm11-synthetic-native-tool-calling-repaired dfm11-synthetic-native-tool-calling-repaired DFM8 synthetic tool trajectories with compatibility normalization materialized in source data. This is a DFM11 replacement for schneiderkamplab/dfm8-synthetic-native-tool-calling. All rows pass exhaustive structural validation. See metadata/manifest.json. text100K<n<1M0 likes47 downloads1mo agoHugging Face22danieldzikunuofmarvel /bibletts-asante-twi-repaired BibleTTS Asante Twi — Repaired Transcripts The Asante Twi transcripts released with BibleTTS have had the characters ɛ (U+025B) and ɔ (U+0254) stripped out. This dataset restores them. Audio is not included. This is a drop-in replacement for the .txt files that ship with the BibleTTS Asante Twi package, matched by clip ID. The problem Both are Twi vowels, and both are required by the orthography. Measured across the released Asante Twi transcripts: Character… See the full description on the dataset page: https://huggingface.co/datasets/danieldzikunuofmarvel/bibletts-asante-twi-repaired.tabularautomatic-speech-recognition10K<n<100K0 likes45 downloads2mo agoHugging Face23Graunt /json-repair-eval-sample JSON repair eval (sample) 30 cases of broken JSON. Each one has the text exactly as a parser would receive it, the repair we expect, the breakage category, the rule applied and the reason. It's a sample of a 300-case set for testing the repair step that sits behind an LLM's structured output or a stream that got cut off. There are ten categories: truncation, trailing commas, single quotes, unescaped control characters, NaN and Infinity, comments, concatenated objects, unquoted… See the full description on the dataset page: https://huggingface.co/datasets/Graunt/json-repair-eval-sample.texttext-generationn<1K0 likes45 downloads14d agoHugging Face24MichaelAnthony /hedgehog-stopping-repair-r5 hedgehog-stopping-repair-r5 Hedgehog — stopping-repair round 5. Contents train.jsonl (1848 rows) validation.jsonl (438 rows) Format JSON Lines (.jsonl), one example per line. Provenance Original content for the Hedgehog extraction model (Michael Anthony Falabella). textquestion-answering1K<n<10K0 likes40 downloads2mo agoHugging Face25MichaelAnthony /hedgehog-precision-repair hedgehog-precision-repair Hedgehog — precision-repair round (complete merchant extraction). Contents train.jsonl (1180 rows) validation.jsonl (116 rows) Format JSON Lines (.jsonl), one example per line. Provenance Original content for the Hedgehog extraction model (Michael Anthony Falabella). textquestion-answering1K<n<10K0 likes36 downloads2mo agoHugging Face26smainye /mpesa-loan-repayment-profiles This dataset is a remastered version prepared using Adaption's Adaptive Data platform. mpesa_loan_repayment_profiles This dataset contains behavioral profiles of M-Pesa users in Kenya, detailing transaction metrics such as frequency, amounts, and cash flow stability across various user types like farmers and salaried workers. Each sample includes derived features like the coefficient of variation for income stability and the proportion of Paybill or airtime transactions. The… See the full description on the dataset page: https://huggingface.co/datasets/smainye/mpesa-loan-repayment-profiles.texttabular-classification1K<n<10K0 likes32 downloads27d agoHugging Face27ProjectScugnizz /scugnizz-agentic-repair-v6-pro-300ktext100K<n<1M0 likes32 downloads3mo agoHugging Face28violetxi /tb21-eval-qwen35-action-only-20k-infra-repaired-c164-max32k-timeout2x qwen35-action-only-20k — Terminal-Bench 2.1 Noncanonical Terminal-Bench 2.1 evaluation of violetxi/qwen35-4b-offline-echo-action-only-20k-tacc through the served model ID qwen35-action-only-20k with Terminus-2. Noncanonical run: timeout_multiplier=2 instead of 1.0; repair concurrency=164 exceeds 30. Do not compare this score directly with canonical TB2.1 leaderboard runs. Result Recorded trials: 445 Tasks / attempts: 89 × 5 Errored trials scored as zero: 250… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/tb21-eval-qwen35-action-only-20k-infra-repaired-c164-max32k-timeout2x.tabularreinforcement-learningn<1K0 likes32 downloads2mo agoHugging Face29ziqiaow /pixeldit-b-native-repa-w4mn-200k-20260925tabularn<1K0 likes32 downloads16d agoHugging Face30MichaelAnthony /hedgehog-complex-repair hedgehog-complex-repair Hedgehog — complex-extraction repair round. Contents train.jsonl (3840 rows) validation.jsonl (304 rows) Format JSON Lines (.jsonl), one example per line. Provenance Original content for the Hedgehog extraction model (Michael Anthony Falabella). textquestion-answering1K<n<10K0 likes31 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.