Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Gramscii-IT /semantic-repair-routing semantic-repair-routing The supervised pairs that train SemanticRepair-270M: a message somebody actually wrote, and the requests inside it restated plainly, one per line. 84,819 pairs in five languages, plus 2,515 in Italian and English aimed at what the model used to refuse. It teaches one narrow thing. An embedding router compares a question with the description of every capability it can reach. People do not write the way capabilities are described — they hedge, they… See the full description on the dataset page: https://huggingface.co/datasets/Gramscii-IT/semantic-repair-routing.texttext-generation10K<n<100K0 likes357 downloads5d agoHugging Face02Emulated-Inc /python-program-repair-training-pool Python program-repair training pool A pool of public data for training a model to repair broken Python. Every row is a program that does the wrong thing and the program that replaces it. It is a straight collection of open datasets plus a rule-generated layer built from open functions, not a new corpus: every row comes from one of the sources below, at the revision named, and every row was put through an overlap filter against held-out material this pool is kept separate from.… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/python-program-repair-training-pool.text-generation100K<n<1M0 likes272 downloads29d agoHugging Face03Louistiti /openenv-python-repair Python Repair Lab An original OpenEnv curriculum of 1,200 deterministic Python function-repair episodes: 12 problem families, four distinct bug patterns per family, and 25 seeded case sets per pattern. There are 151,780 executable checks across the episodes. These are 48 repair patterns with data variants, not 1,200 unrelated algorithms. Tasks cover interval algorithms, rolling calculations, weighted statistics, stable deduplication, Unicode run-length encoding, Luhn checksums… See the full description on the dataset page: https://huggingface.co/datasets/Louistiti/openenv-python-repair.tabulartext-generation1K<n<10K0 likes204 downloads1d agoHugging Face04VmaxRL /SWE-universe-repaired-bug-pilot-trajectories SWE-universe repaired BugPilot trajectories Combined trajectory artifacts for the Qwen3.6 + mini-swe-agent evaluation of VmaxRL/SWEUniverse-Repaired-Bugpilot. This dataset contains one row per evaluated task in metadata.jsonl, plus per-task files under trajectories//. The combined set uses the main full eval and replaces the two original infra-failure rows with the clean infra rerun trajectories. Summary: rows: 804 effective attempts: 804 passes: 629 pass rate: 0.782338 infra… See the full description on the dataset page: https://huggingface.co/datasets/VmaxRL/SWE-universe-repaired-bug-pilot-trajectories.text-generation0 likes175 downloads5mo agoHugging Face05Reza2kn /uncgpt-conversations-semantic-approved-1p25-repaired-paperclip UncGPT — Semantic-Approved 1.25σ Conversations (Leak-Repaired) The 1.25σ semantic-gate cohort with uncle-diary leakage repaired and normalized diary fields. The auditable replacement for the older paperclip_all_1803 source that an earlier audit flagged for visible diary leakage. Part of the UncGPT NeurIPS 2026 Competition collection. Config approved_manifest (default): one row per approved conversation, with metadata + path back to the full-schema JSON.… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/uncgpt-conversations-semantic-approved-1p25-repaired-paperclip.texttext-generation1K<n<10K0 likes168 downloads5mo agoHugging Face06wAI-org /swerl-tmax-15k-repairs-gpt-5-6-sol swerl-tmax-15k task repairs (gpt-5-6-sol) Proposed repairs for defective tasks in hamishivi/swerl-tmax-15k, generated from the audit labels in wAI-org/swerl-tmax-15k-rubric-gpt-5-6-sol. These repairs are UNVALIDATED No repair here has been executed, and none has been shown to be solvable. Every repair was checked mechanically — valid bash, still writes a reward, does not delete a path the instruction needs. None was checked empirically. A hardened verifier can be… See the full description on the dataset page: https://huggingface.co/datasets/wAI-org/swerl-tmax-15k-repairs-gpt-5-6-sol.texttext-generation10K<n<100K0 likes146 downloads29d agoHugging Face07empirischtech /Palace-Config-Repair Palace Configuration Repair Benchmark Execution-graded repair tasks for configuration files of Palace, an open-source finite-element solver for computational electromagnetics. Each task gives a model a perturbed Palace JSON configuration and asks it to return a corrected one. A repair counts as correct only if it conforms to the schema, is accepted by the solver, and reproduces the reference outputs of the original case when Palace v0.14.0 runs it. A configuration that is valid… See the full description on the dataset page: https://huggingface.co/datasets/empirischtech/Palace-Config-Repair.texttext-generationn<1K0 likes138 downloads9d agoHugging Face08VmaxRL /SWEUniverse-Repaired-Indist-full-not-SWE-bench-pro-matched VmaxRL/SWEUniverse-Repaired-Indist-full-not-SWE-bench-pro-matched This dataset contains a 350-row subset selected from the Indist SWEUniverse training rows. Selection policy: three-way repo overlap with Bugpilot and LM-Modify, deduped by repo plus introduction patch, then balanced round-robin across overlapping repos. Rows: 350 Selected repos: 19 Deduped overlap capacity: 468 Source dataset: VmaxRL/SWEUniverse-Repaired-Indist-full-not-SWE-bench-pro-matched texttext-generationn<1K0 likes120 downloads5mo agoHugging Face09ASSERT-KTH /repairllama-datasets RepairLLaMA - Datasets Contains the processed fine-tuning datasets for RepairLLaMA. Instructions to explore the dataset To load the dataset, you must define which revision (i.e., which input/output representation pair) you want to load. from datasets import load_dataset # Load ir1xor1 dataset = load_dataset("ASSERT-KTH/repairllama-datasets", "ir1xor1") # Load irXxorY dataset = load_dataset("ASSERT-KTH/repairllama-datasets", "irXxorY") Citation If you use… See the full description on the dataset page: https://huggingface.co/datasets/ASSERT-KTH/repairllama-datasets.texttext-generation100K<n<1M3 likes118 downloads2y agoHugging Face10shravandoda /TobiaSVG-repair TobiaSVG Repair TobiaSVG Repair contains 44,802 synthetic SVG repair pairs. Each row pairs a corrupted SVG with its clean target. Raster images are rendered when examples are loaded and are not stored. Sources And Splits Subset Source Rows License vfig_diagrams VFIG-Data 33,624 ODC-BY 1.0 vfig_shapes VFIG-Data 8,884 ODC-BY 1.0 animal_illustrations SVG Animal Illustrations 2,294 CC0 1.0 Splits contain 35,833 training, 4,423 test, and 4,546… See the full description on the dataset page: https://huggingface.co/datasets/shravandoda/TobiaSVG-repair.texttext-generation10K<n<100K0 likes79 downloads3mo agoHugging Face11ChanMeng666 /archlang-repair-trajectories ArchLang Repair Trajectories A fully synthetic, procedurally generated dataset of floor-plan program-repair and authoring examples for ArchLang — a small declarative language that compiles .arch floor-plan source to professional SVG. Every row is self-verifying through the deterministic ArchLang compiler, with zero model or API involvement in its construction. Generator + seed: open source in the main repository (dataset/, npm run dataset:gen), so the corpus is reproducible… See the full description on the dataset page: https://huggingface.co/datasets/ChanMeng666/archlang-repair-trajectories.text-generation1K<n<10K0 likes78 downloads3mo agoHugging Face12emgena /omnimcp_react_hydration_repair_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_react_hydration_repair_teaser.texttext-generationn<1K0 likes78 downloads24d agoHugging Face13emgena /omnimcp_pytest_traceback_repair_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_pytest_traceback_repair_teaser.texttext-generationn<1K1 likes72 downloads24d agoHugging Face14NadavSalem /Titanius-1.2-sft-repair Titanius-1.2-sft-repair Synthetic instructions used in the two mixed-data continuations of Titanius-1.2-128m-sft-fp16. They target short answers, arithmetic, copying, JSON, facts, concise definitions and multi-turn recall. Config Training conversations Used by phase1 5,917 First continuation, selected at step 5,500 phase2 16,527 Second continuation, selected at step 6,000 Each config is a complete synthetic pool for its phase. They overlap; do not concatenate… See the full description on the dataset page: https://huggingface.co/datasets/NadavSalem/Titanius-1.2-sft-repair.texttext-generation10K<n<100K0 likes66 downloads5d agoHugging Face15darrowoflykos /nemotron-sft-safety-v2-repaired Nemotron-SFT-Safety-v2 (repaired) A cleaned, schema-normalized copy of nvidia/Nemotron-SFT-Safety-v2. Why this exists The upstream data/train.jsonl is not directly loadable by datasets' JSON loader: The metadata object is heterogeneous -- different records carry different sets of keys (6 observed variants, 6 to 20 fields). metadata.source_id is type-inconsistent -- a string, an int, or null across records. This dataset normalizes every record to a fixed union… See the full description on the dataset page: https://huggingface.co/datasets/darrowoflykos/nemotron-sft-safety-v2-repaired.texttext-generation100K<n<1M1 likes62 downloads24d agoHugging Face16dipenbhuva /home-diy-repair-qa Home DIY Repair Q&A A synthetic dataset of 5,000 Q&A pairs covering common home DIY repair scenarios. Each example includes a detailed step-by-step answer, required tools, safety warnings, and practical tips. Dataset Purpose This dataset is built for: Instruction fine-tuning — train language models to give detailed, safe, and actionable home repair guidance Retrieval-Augmented Generation (RAG) — build a knowledge base for home repair assistants Question answering — train… See the full description on the dataset page: https://huggingface.co/datasets/dipenbhuva/home-diy-repair-qa.textquestion-answering1K<n<10K1 likes56 downloads7mo agoHugging Face17skonml /code-contract-repair APIContractRepair APIContractRepair is a provenance-tracked instruction-tuning dataset for software engineers and code-model researchers who need contract-faithful, minimal repairs with tests that distinguish a broken implementation from its fix. Magicoder-OSS-Instruct-75K supplies real function-identifier seeds, but it does not provide these documented contracts, deliberately buggy implementations, minimal corrected implementations, or paired regression tests. This release… See the full description on the dataset page: https://huggingface.co/datasets/skonml/code-contract-repair.texttext-generation1K<n<10K0 likes53 downloads2mo agoHugging Face18joanvelja /polaris-53k-repaired POLARIS-53K, label-repaired 49,289 of the 53,291 rows in POLARIS-Project/Polaris-Dataset-53K, with 4,580 stored answers corrected and 4,002 rows removed as unrepairable. Measurements on the source set put its bad-label rate at roughly 15.9% [14.3, 17.6] (two independent detectors agreeing on a 2,000-row sample). Mislabelled rows are not uniformly distributed: they concentrate in the problems models fail, which is exactly where a difficulty-calibration pipeline looks.… See the full description on the dataset page: https://huggingface.co/datasets/joanvelja/polaris-53k-repaired.tabulartext-generation10K<n<100K0 likes49 downloads1mo agoHugging Face19Benitoow /OfficeSmith-PPTX-Repair OfficeSmith PPTX Repair Deterministically degraded PPTX IR objects paired with validated repairs. Dataset summary This dataset is part of the OfficeSmith collection for training models to plan, build, clarify, critique, and repair editable business presentations. It contains observable outputs only: no hidden chain of thought, secret benchmark prompt, personal data, or API credential is included. Train rows: 160 Validation rows: 0 Test rows: 0 Languages: French… See the full description on the dataset page: https://huggingface.co/datasets/Benitoow/OfficeSmith-PPTX-Repair.texttext-generationn<1K0 likes47 downloads2mo agoHugging Face20Graunt /json-repair-eval-sample JSON repair eval (sample) 30 cases of broken JSON. Each one has the text exactly as a parser would receive it, the repair we expect, the breakage category, the rule applied and the reason. It's a sample of a 300-case set for testing the repair step that sits behind an LLM's structured output or a stream that got cut off. There are ten categories: truncation, trailing commas, single quotes, unescaped control characters, NaN and Infinity, comments, concatenated objects, unquoted… See the full description on the dataset page: https://huggingface.co/datasets/Graunt/json-repair-eval-sample.texttext-generationn<1K0 likes45 downloads14d agoHugging Face21darkone01 /synthetic-everyday-text-repair-corpus Synthetic Everyday Text Repair Corpus Description This dataset contains 3,080 original AI-generated synthetic English sentences about retail operations, delivery, maintenance, training, inventory, and workplace communication. It provides clean reference text for the challenge Noisy Text Repair: Meaning-Preserving Text Correction. A separate preparation script creates noisy inputs and splits the data by template family. Data File clean.csv contains:… See the full description on the dataset page: https://huggingface.co/datasets/darkone01/synthetic-everyday-text-repair-corpus.texttext-generation1K<n<10K0 likes45 downloads3d agoHugging Face22Eyerf /agentblackbox-rag-repair-outcomes AgentBlackBox RAG Repair Outcome Dataset This dataset contains replay-labeled repair outcome data for AgentBlackBox, a counterfactual debugging framework for language agents. The data is built around failed RAG/document-recall agent traces, candidate repairs, counterfactual replay labels, and repair-ranking evaluation outputs. Contents datasets/ world_model_ranker_dataset_v2_train10k/ pointwise/ listwise/ stats.json… See the full description on the dataset page: https://huggingface.co/datasets/Eyerf/agentblackbox-rag-repair-outcomes.text-generation0 likes43 downloads3mo agoHugging Face23KorolOrol /gpt2-steering-repair-results GPT-2 Steering Repair Results Итоговые machine-readable результаты исследования gpt2-stearing-repair. Опубликованный checkpoint: gpt2-steering-denoiser. Датасет содержит только метрики, без текстов prompts и сгенерированных продолжений. Файлы Файл Строки Назначение confirm_neural_v2.csv 80 000 Итоговая common-RNG оценка пяти методов confirm_isotropic_v2_seed1.csv 16 000 Независимое повторение isotropic checkpoint pareto_neural_v2.csv 50 Агрегаты по… See the full description on the dataset page: https://huggingface.co/datasets/KorolOrol/gpt2-steering-repair-results.text-generation0 likes38 downloads2mo agoHugging Face24CodeWalk /multi-bug-repair CodeWalk — Multi-Bug Repair Agentic co-located multi-bug software repair. A level-N task presents N coupled bugs simultaneously at one repository snapshot; the agent must fix all of them so that the union of their FAIL_TO_PASS tests passes. Part of the CodeWalk benchmark suite (CodeWalk: Generating Coding Benchmarks by Walking a Problem Graph). 568 tasks across levels L1–L3 (1,104 bugs, 280 distinct repositories) Every task is gold-verified: all bugs fail at the base commit… See the full description on the dataset page: https://huggingface.co/datasets/CodeWalk/multi-bug-repair.tabulartext-generationn<1K0 likes30 downloads2d agoHugging Face25Occupying-Mars /glm-base-ood-repair-mix-10k GLM base OOD repair mix 10k Bucket-targeted BFCL-style tool-calling repair dataset for GLM native tool-call finetuning. Built from public OOD tool-call datasets and filtered against BFCL single-call eval prompts. Primary file: train.jsonl Rows: 9788 after dropping exact BFCL eval prompt overlaps. Format: messages, tools, target_call. Training should use GLM native target formatting from target_call, not the legacy target_text_cohere field. Audit files included:… See the full description on the dataset page: https://huggingface.co/datasets/Occupying-Mars/glm-base-ood-repair-mix-10k.texttext-generation1K<n<10K0 likes24 downloads4mo agoHugging Face26SerFabio89 /italian-logic-repair-sft-dataset Italian Logic Repair SFT Dataset Teacher-backed synthetic Italian-first dataset designed for supervised fine-tuning repair. It targets exact arithmetic, concise direct QA, executable Python functions, JSON-only output, constraint following, stop behavior, and reasoning final-answer-marker behavior. Teacher outputs are used as candidates, then validated, corrected, or rejected by deterministic checks. Dataset Details Field Value Repository… See the full description on the dataset page: https://huggingface.co/datasets/SerFabio89/italian-logic-repair-sft-dataset.texttext-generation100K<n<1M0 likes23 downloads5mo agoHugging Face27ProjectScugnizz /scugnizz-agentic-repair-50k-v2 Scugnizz Agentic Repair 50k Dataset sintetico bilanciato per correggere renderer, copia esatta e tool calling. Train: 49500 Validation: 500 Categorie: { "renderer_weather": 3750, "renderer_finance": 3750, "renderer_spotify": 8, "renderer_mail": 3750, "renderer_calendar": 448, "renderer_dns": 36, "renderer_whois": 3750, "renderer_json_complex": 3750, "exact_hash": 64, "exact_network": 180, "exact_url_domain": 2424, "tool_weather": 48, "tool_finance": 36… See the full description on the dataset page: https://huggingface.co/datasets/ProjectScugnizz/scugnizz-agentic-repair-50k-v2.texttext-generation10K<n<100K0 likes23 downloads3mo agoHugging Face28frank-rg /bacardi-breaking-update-repair Bacardi Breaking-Update Repair Results from evaluating five self-hosted, open-weight LLMs on the Bacardi benchmark: automatically repairing Java projects broken by upstream dependency updates. Each model is run against the same 103-case breaking-dependency-update benchmark, across all 8 Bacardi prompt pipelines, served locally via vLLM on the Berzelius (NSC) HPC cluster. Benchmark: 103 real-world Java "breaking update" commits (from the chains-project/breaking-updates corpus)… See the full description on the dataset page: https://huggingface.co/datasets/frank-rg/bacardi-breaking-update-repair.text-generation1 likes20 downloads1mo agoHugging Face29jmp1987 /simson-repair-manual 🔧 Simson Repair Manual & Technical Data Strukturierte technische Daten aus DDR-Werkstatthandbüchern für klassische Simson-Mopeds. Inhalt (14 Records in 6 Sektionen) Sektion Inhalt Modelle Technische_Daten Vollständige Technische Daten je Modell S50, S51, S70, KR51/2, SR50 Anzugsmomente Drehmoment-Tabellen für alle Schrauben S51/S50/S70, KR51 Einstellwerte Zündung, Vergaser, Kupplung, Reifen S51/S50/S70, KR51 Wartungsintervalle 500/2500/5000km +… See the full description on the dataset page: https://huggingface.co/datasets/jmp1987/simson-repair-manual.textquestion-answeringn<1K1 likes16 downloads4mo agoHugging Face30ProjectScugnizz /scugnizz-agentic-repair-50k-v4 Scugnizz Agentic Repair 50k v3 { "renderer_weather": 3750, "renderer_finance": 3750, "renderer_spotify": 3750, "renderer_mail": 3750, "renderer_calendar": 3750, "renderer_dns": 3750, "renderer_whois": 3750, "renderer_json_complex": 3750, "exact_hash": 3334, "exact_network": 3333, "exact_url_domain": 3333, "tool_weather": 1667, "tool_finance": 1667, "tool_dns": 1667, "tool_spotify": 1667, "tool_mail": 1666, "tool_calendar": 1666 } texttext-generation10K<n<100K0 likes12 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.