Team Ai
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01greghavens /fable-5-coding-and-debugging-traces-synthetic-corrections Model Synthetic Corrections 1 TRAJECTORIES · 2 TRAINING ROWS · 16 kB Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Synthetic Corrections companion dataset. The original dataset is greghavens/fable-5-coding-and-debugging-traces. These are narrowly, synthetically corrected, independently re-judged traces that never passed in the original dataset. Behavior-preserving instruction-following… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/fable-5-coding-and-debugging-traces-synthetic-corrections.tabulartext-generationn<1K0 likes169 downloads2mo agoHugging Face02stindardlogic /code-debugging-sft-50k Code Debugging SFT (50K) 50,000 ShareGPT-format conversations where the user presents buggy code and the assistant provides root-cause analysis and a corrected solution. Covers Python, JavaScript, Go, TypeScript, and SQL across 14 bug categories. Motivation Debugging is one of the most frequent developer tasks — and one of the hardest to train. Most coding datasets focus on writing code from scratch. This dataset trains models to: Identify the precise root cause… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/code-debugging-sft-50k.texttext-generation10K<n<100K0 likes165 downloads3mo agoHugging Face0311-47 /fable-5-coding-and-debugging-traces-synthetic Model Synthetic Corrections 1 TRAJECTORIES · 2 TRAINING ROWS · 16 kB Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Synthetic Corrections companion dataset. The original dataset is greghavens/fable-5-coding-and-debugging-traces. These are narrowly, synthetically corrected, independently re-judged traces that never passed in the original dataset. Behavior-preserving instruction-following… See the full description on the dataset page: https://huggingface.co/datasets/11-47/fable-5-coding-and-debugging-traces-synthetic.tabulartext-generationn<1K0 likes122 downloads22d agoHugging Face04gbeck /kimi-k3-coding-and-debugging-traces Kimi K3 Coding & Debugging Agent Traces Generated by moonshiner — an open harness for distilling verified, model-attested agentic coding traces. Real, end-to-end agentic coding trajectories produced by moonshotai/kimi-k3 driving the pi coding-agent runtime over openrouter, at max reasoning. Each trajectory solves a concrete repair or build task in a real repository — reading, editing, and running code with tools — and is published only after its work verifiably passes —… See the full description on the dataset page: https://huggingface.co/datasets/gbeck/kimi-k3-coding-and-debugging-traces.texttext-generationn<1K3 likes112 downloads3mo agoHugging Face05creeperdatasets /python_debugging Python Debugging A synthetic instruction-tuning dataset for training AI models to identify and fix bugs in Python code. Dataset Summary Field Value Entries 75 Format input / output pairs Language English Topic Finding and fixing bugs in Python code Synthetic Yes, generated with DeepSeek License MIT Dataset Description Each entry presents a snippet of Python code containing a deliberate bug, along with a corrected version… See the full description on the dataset page: https://huggingface.co/datasets/creeperdatasets/python_debugging.texttext-generationn<1K0 likes94 downloads1mo agoHugging Face06failuremap /failuremap-debugging Failure Map: 20,168 open debugging tasks with executable checks Explore the archive · Try a case · Methodology Failure Map is a corpus of compact Python debugging tasks for coding-model training experiments, repair evaluation, and reinforcement learning with execution feedback. Each open task includes an explicit contract, a broken implementation, a plausible repair that still fails, and embedded executable boundary checks. The implementations use only the Python standard… See the full description on the dataset page: https://huggingface.co/datasets/failuremap/failuremap-debugging.texttext-generation10K<n<100K1 likes49 downloads6d agoHugging Face07dmeldrum6 /Code_Debugging_QA Code Debugging Q&A Dataset By dmeldrum6 A curated dataset of 1,073 question-and-answer pairs covering common debugging scenarios across Python, JavaScript, SQL, and Bash. Designed for fine-tuning and instruction-tuning language models on code debugging tasks. Dataset Summary Each pair presents a realistic bug symptom as a question and a structured answer containing: A buggy code block demonstrating the problem A corrected code block showing the fix A plain-language… See the full description on the dataset page: https://huggingface.co/datasets/dmeldrum6/Code_Debugging_QA.text1K<n<10K1 likes39 downloads6mo agoHugging Face08kalaiarasan27 /Code_Debugging_QA Code Debugging Q&A Dataset By dmeldrum6 A curated dataset of 1,073 question-and-answer pairs covering common debugging scenarios across Python, JavaScript, SQL, and Bash. Designed for fine-tuning and instruction-tuning language models on code debugging tasks. Dataset Summary Each pair presents a realistic bug symptom as a question and a structured answer containing: A buggy code block demonstrating the problem A corrected code block showing the fix A… See the full description on the dataset page: https://huggingface.co/datasets/kalaiarasan27/Code_Debugging_QA.text1K<n<10K0 likes22 downloads1mo agoHugging Face0911-47 /self_evolving_self_debugging_250_implementations-2textn<1K1 likes19 downloads10mo agoHugging Face10schneiderkamplab /dfm8-synthetic-code-debugging Code Generation and Debugging Synthetic DFM8 training data generated with Gemma 4 31B and filtered by deterministic checks plus a Gemma 4 31B judge. Schema Rows are JSONL chat records: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]} Tool-calling rows may also include a top-level tools list and assistant tool_calls. Counts accepted rows: 340711 generated rows seen: 4800000 audit rows seen: 4580233… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm8-synthetic-code-debugging.text100K<n<1M0 likes12 downloads3mo agoHugging Face1111-47 /self_evolving_self_debugging_250_implementationstextn<1K1 likes10 downloads10mo agoHugging Face12Maitreyajayaraj /telugu_compiler_debugging_v7textn<1K0 likes7 downloads6mo agoHugging Face13alucent /mirror-Code_Debugging_QAgated Code Debugging Q&A Dataset By dmeldrum6 A curated dataset of 1,073 question-and-answer pairs covering common debugging scenarios across Python, JavaScript, SQL, and Bash. Designed for fine-tuning and instruction-tuning language models on code debugging tasks. Dataset Summary Each pair presents a realistic bug symptom as a question and a structured answer containing: A buggy code block demonstrating the problem A corrected code block showing the fix A… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-Code_Debugging_QA.text1K<n<10K0 likes5 downloads3mo agoHugging Face14Maitreyajayaraj /debugging-failure-dilemma-v10textn<1K0 likes4 downloads6mo agoHugging Face15Maitreyajayaraj /telugu_debugging_failures_v7textn<1K0 likes4 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.