datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Agentic-Chain-of-Thought-Coding-SFT-Dataset
🤖 Agentic Coding CoT Dataset
A high-quality supervised fine-tuning (SFT) dataset for training agentic coding assistants with Chain-of-Thought reasoning capabilities.
📋 Dataset Description
This dataset was created by processing and distilling ~20GB of GitHub crawl data using Minimax-M2 to generate structured, reasoning-rich coding examples. Each sample demonstrates systematic problem-solving with explicit tool usage patterns.
🏗️ Assistant Data Structure… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Agentic-Chain-of-Thought-Coding-SFT-Dataset.Roblox-luau-coding_L1
8BitStudio/Roblox-luau-coding_L1
A dataset for training and fine-tuning AI models on Roblox Luau scripting.
Covers a wide range of scripting topics from beginner to advanced.
Summary
This dataset contains 12,306 Luau code examples designed to teach AI models
how to write scripts for Roblox. Topics range from basic part manipulation
to complex datastore systems.
Dataset Structure
Data Format
Each example is a tab-separated pair of a… See the full description on the dataset page: https://huggingface.co/datasets/8BitStudio/Roblox-luau-coding_L1.Paragon-coding
NOTICE
This was done by me, someone with a learning Disability. So please do bare with me when updating this with more working data.
Multi-Language Programming Code Dataset
A curated dataset of original, non-scraped code examples across 7 programming
environments: Python, JavaScript, Node.js, Java, C, C++, and Rust.
The dataset ships in two parts that can be used separately or combined:
File
Rows
Description
code_dataset.jsonl / .csv
105
Hand-written… See the full description on the dataset page: https://huggingface.co/datasets/TGPRO32/Paragon-coding.fable5-agentic-coding-sft
FABLE.5 Agentic Coding SFT (curated)
~159,972 supervised fine-tuning examples for agentic coding — multi-turn conversations where the
assistant drives a tool-call loop (shell, file edits, tests) and commits to complete solutions. Used to train
VibeThinker-Fable-Nano-Agentic-3B.
Provenance & license
Curated/distilled from the Complete-FABLE.5-traces-2M trace set:
Original source: Glint-Research/Complete-FABLE.5-traces-2M (currently gated).
Pulled from:… See the full description on the dataset page: https://huggingface.co/datasets/Nexlab/fable5-agentic-coding-sft.swe-coding-instruction-following
SWE Coding Instruction-Following
A curated collection of real-world software engineering tasks in the SWE-bench format, designed for evaluating instruction-following capabilities of coding agents. Each task represents a genuine GitHub issue with a reproducible environment, test suite, and reference solution — the agent must precisely follow the issue instructions to produce a correct fix.
Overview
Item
Details
Total Tasks
50
Repositories
2 (pallets/click… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/swe-coding-instruction-following.Kimi-K2.7-CodingTraces-9000x
Kimi K2.7 Coding Traces 9000x
A validated 9,014-row coding and software-engineering reasoning
dataset generated with moonshotai/Kimi-K2.7-Code.
Every row contains a coding-focused prompt, a separated reasoning trace, and a
final answer. The release was built from a durable Google Drive generation
pipeline and underwent a complete two-pass schema and delimiter audit before
publication.
Generation configuration
Setting
Value
Teacher… See the full description on the dataset page: https://huggingface.co/datasets/trjxter/Kimi-K2.7-CodingTraces-9000x.fable-5-coding-and-debugging-traces-synthetic-corrections
Model Synthetic Corrections
1 TRAJECTORIES · 2 TRAINING ROWS · 16 kB
Generated by moonshiner — an open harness for
distilling verified instruction-following, tool-use, and agentic coding traces.
Synthetic Corrections companion dataset. The original dataset is greghavens/fable-5-coding-and-debugging-traces. These are narrowly, synthetically corrected, independently re-judged traces that never passed in the original dataset.
Behavior-preserving instruction-following… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/fable-5-coding-and-debugging-traces-synthetic-corrections.Vibe-Coding-Instructfable-5-coding-and-debugging-traces-synthetic
Model Synthetic Corrections
1 TRAJECTORIES · 2 TRAINING ROWS · 16 kB
Generated by moonshiner — an open harness for
distilling verified instruction-following, tool-use, and agentic coding traces.
Synthetic Corrections companion dataset. The original dataset is greghavens/fable-5-coding-and-debugging-traces. These are narrowly, synthetically corrected, independently re-judged traces that never passed in the original dataset.
Behavior-preserving instruction-following… See the full description on the dataset page: https://huggingface.co/datasets/11-47/fable-5-coding-and-debugging-traces-synthetic.kimi-k3-coding-and-debugging-traces
Kimi K3 Coding & Debugging Agent Traces
Generated by moonshiner — an open harness for
distilling verified, model-attested agentic coding traces.
Real, end-to-end agentic coding trajectories produced by
moonshotai/kimi-k3 driving the pi coding-agent runtime over
openrouter, at max reasoning. Each trajectory solves a concrete
repair or build task in a real repository — reading, editing, and running code
with tools — and is published only after its work verifiably passes —… See the full description on the dataset page: https://huggingface.co/datasets/gbeck/kimi-k3-coding-and-debugging-traces.Agentic-Chain-of-Thought-Coding-SFT-Dataset
🤖 Agentic Coding CoT Dataset
A high-quality supervised fine-tuning (SFT) dataset for training agentic coding assistants with Chain-of-Thought reasoning capabilities.
📋 Dataset Description
This dataset was created by processing and distilling ~20GB of GitHub crawl data using Minimax-M2 to generate structured, reasoning-rich coding examples. Each sample demonstrates systematic problem-solving with explicit tool usage patterns.
🏗️ Assistant Data… See the full description on the dataset page: https://huggingface.co/datasets/ArkhAngelLifeJiggy/Agentic-Chain-of-Thought-Coding-SFT-Dataset.local-agentic-coding-bench-8gb-vram-2026-05
agentic coding benchmark: local LLMs on 8GB VRAM
can local LLMs do agentic coding (multi-turn tool calling, file creation, debugging) on consumer hardware? this dataset captures real test results.
hardware
GPU: NVIDIA RTX 4060 Ti 8GB
CPU: Intel i7-14700F
RAM: 32 GB DDR5
OS: Windows 11 + WSL2 (Ubuntu)
inference: llama-server (turboquant fork of llama.cpp)
what was tested
two agent frameworks:
Hermes Agent (NousResearch): structured tool calling with… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/local-agentic-coding-bench-8gb-vram-2026-05.mimo-coding-synthetic-5k
MiMo Coding Synthetic 5.4K
MiMo Coding Synthetic 5.4K is a purely synthetic coding instruction dataset generated with Xiaomi MiMo mimo-v2.5-pro.
It contains 5,411 validated examples across programming languages, coding task types, and difficulty levels. The dataset is provided in two formats:
A canonical rich JSONL format with metadata and labels.
An OpenAI chat messages JSONL format for supervised fine-tuning pipelines.
The generation run used 20 parallel workers for roughly… See the full description on the dataset page: https://huggingface.co/datasets/himanshunakrani9/mimo-coding-synthetic-5k.Agentic-Chain-of-Thought-Coding-SFT-Dataset-v1.1
🤖 Agentic Coding CoT Dataset v1.1
A high-quality supervised fine-tuning (SFT) dataset for training agentic coding assistants with Chain-of-Thought reasoning capabilities.
📋 Dataset Description
This dataset was created by processing and distilling ~20GB of GitHub crawl data using Minimax-M2 & MiniMax M2.1 to generate structured, reasoning-rich coding examples. Each sample demonstrates systematic problem-solving with explicit tool usage patterns.
🏗️ Assistant… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Agentic-Chain-of-Thought-Coding-SFT-Dataset-v1.1.coding-master-dataset
Coding Master Dataset
Overview
A large-scale coding instruction-tuning dataset in ShareGPT conversational format, assembled from multiple open sources and deduplicated.
Records: 766,987
Format: JSONL / ShareGPT
License: Apache 2.0
Sources
CodeX-2M-Thinking (430,542 records)
python-code-dataset-500k (559,515 records)
StackPulse high-quality subset (20,205 records)
CodeFeedback-Filtered-Instruction (156,525 records)
secure_programming_dpo (4,656… See the full description on the dataset page: https://huggingface.co/datasets/Voidreaper2026/coding-master-dataset.coding-interview-sft-100k
Coding Interview SFT (100K)
100,000 ShareGPT conversations demonstrating expert-level coding interview preparation across algorithms, data structures, system design, and behavioral questions. Each example provides a complete solution with detailed explanation of the approach, step-by-step reasoning, time/space complexity analysis, and edge case handling.
Motivation
Coding interview preparation is one of the highest-demand AI assistant use cases. Models commonly… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/coding-interview-sft-100k.Coding_GPT4_Data
Dataset Info
** This dataset is generated by the GPT-4 based model.
** The whole dataset is about coding.
Dataset Structure
[
{
"user": "How can I implement a Python function to check if a given string is a palindrome or not? \n\nPlease generate the code for this task in Python.",
"assistant": "Sure! Here's a Python function that checks whether a given string is a palindrome or not:\n\n```python\ndef is_palindrome(input_string):\n # Convert the string to… See the full description on the dataset page: https://huggingface.co/datasets/MAsad789565/Coding_GPT4_Data.dense-reasoning-coding-1k
Dense-Reasoning-Coding-1K
Dataset Description
This dataset is an optimized, highly dense Supervised Fine-Tuning (SFT) subset designed to teach smaller language models (e.g., 1B to 8B architectures) how to reason about complex coding problems without overwhelming their context windows.
It is derived from the verified_90k split of IIGroup/X-Coder-SFT-376k, which features advanced programming tasks and solutions.
About the Creator & Origin
This… See the full description on the dataset page: https://huggingface.co/datasets/Phips/dense-reasoning-coding-1k.Rhea-Coding
Rhea Multi-Pass Coding Dataset
A curated dataset for fine-tuning coding AI models with 3-pass reasoning capabilities.
Dataset Description
This dataset contains Python programming examples with structured multi-pass reasoning:
Pass 1: Quick first implementation
Pass 2: Self-review with structured checklist
Pass 3: Final optimized version
Languages
Python (primary)
Dataset Structure
Data Instances
Each example follows… See the full description on the dataset page: https://huggingface.co/datasets/roskosmos19/Rhea-Coding.Synthetic-JP-Coding-Dataset-Magpie-Nemotron-4-10k
Synthetic-JP-Coding-Dataset-Magpie-Nemotron-4-10k
Magpieの手法をnvidia/Nemotron-4-340B-Instructに対して適用し作成した、約10000件の日本語のコーディング用対話データセットです。
データセットの作成にはDeepInfraを利用しました。
また、このリポジトリでデータセット作成に用いたコードを公開しています。これをベースに、システムプロンプトとstopを一部変更することで生成しています。
特に事後的なフィルタ処理は加えていないため、クオリティの低いレコードが含まれている可能性があります。ご注意ください。
sota-codingvcl-ai-coding-prompts
VCL AI Coding Power Prompts
50 battle-tested prompts for Claude Code, Codex, Gemini CLI, and Cursor — by Vibe Coder's Life.
Free catalog for vibe coders. Replace {{PLACEHOLDERS}} with your facts. Not the paid Apify Playbook prompt pack (those stay private).
Load
from datasets import load_dataset
ds = load_dataset("kondasviktor/vcl-ai-coding-prompts", "prompts")
print(ds["train"][0]["title"])
Columns
Column
Description
id
Stable id… See the full description on the dataset page: https://huggingface.co/datasets/kondasviktor/vcl-ai-coding-prompts.agentic-cot-coding-sft
🤖 Agentic Coding CoT Dataset v1.1
A high-quality supervised fine-tuning (SFT) dataset for training agentic coding assistants with Chain-of-Thought reasoning capabilities.
📋 Dataset Description
This dataset was created by processing and distilling ~20GB of GitHub crawl data using Minimax-M2 & MiniMax M2.1 to generate structured, reasoning-rich coding examples. Each sample demonstrates systematic problem-solving with explicit tool usage patterns.
🏗️… See the full description on the dataset page: https://huggingface.co/datasets/ArkhAngelLifeJiggy/agentic-cot-coding-sft.agentic-coding-traces
Agentic Coding Mooncake Traces
Synthetic agentic coding benchmark datasets in Mooncake trace (JSONL) format,
generated with AIPerf 0.9.0 for LLM inference benchmarking.
Designed for use with InferenceX via the
agentic-replay scenario-type and aiperf_adapter.py.
Files
File
Sessions
Turns
max_prompt_tokens
Seed
64k/dataset.jsonl
1,000
18,595
65,536
42
128k/dataset.jsonl
1,000
16,957
131,072
42
Format
Each line is a Mooncake trace… See the full description on the dataset page: https://huggingface.co/datasets/thangquang09/agentic-coding-traces.Synthetic-JP-EN-Coding-Dataset-Magpie-69k
Synthetic-JP-EN-Coding-Dataset-Magpie-69k
Magpieの手法を様々なモデルに対して適用し作成した、約69000件の日本語・英語のコーディング対話データセットです。
作成に利用したモデルは以下の通りです。modelキーに該当レコードの作成に利用したモデル情報があります。
nvidia/Nemotron-4-340B-Instruct
microsoft/Phi-3-medium-4k-instruct
mistralai/Mixtral-8x22B-Instruct-v0.1
cyberagent/calm3-22b-chat
データセットの作成にはDeepInfraを利用しました。
また、このリポジトリでデータセット作成に用いたコードを公開しています。これをベースに、プロンプトテンプレートやシステムプロンプト等を一部変更することで生成しています。特に事後的なフィルタ処理は加えていないため、クオリティの低いレコードが含まれている可能性があります。ご注意ください。
preference-pairs-coding-50k
Code Quality Preference Pairs (50K)
50,000 DPO preference pairs for training LLMs to write high-quality, secure, and idiomatic code.
Motivation
Code generation models frequently produce code that "works" but has critical issues: SQL injection vulnerabilities, O(n²) algorithms where O(n) is trivial, swallowed exceptions, thread safety bugs, and non-idiomatic patterns. This dataset trains models to produce code that a senior engineer would actually approve.… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/preference-pairs-coding-50k.Agentic-Chain-of-Thought-Coding-SFT-Dataset-v1.1
🤖 Agentic Coding CoT Dataset v1.1
A high-quality supervised fine-tuning (SFT) dataset for training agentic coding assistants with Chain-of-Thought reasoning capabilities.
📋 Dataset Description
This dataset was created by processing and distilling ~20GB of GitHub crawl data using Minimax-M2 & MiniMax M2.1 to generate structured, reasoning-rich coding examples. Each sample demonstrates systematic problem-solving with explicit tool usage patterns.
🏗️ Assistant… See the full description on the dataset page: https://huggingface.co/datasets/mepartha/Agentic-Chain-of-Thought-Coding-SFT-Dataset-v1.1.DSA-Coding-Problems-and-Solutions-Dataset
Dataset Description
This dataset is a large-scale collection of Data Structures and Algorithms (DSA) code, containing 12,385 code files with 3.86 million lines of code and 25.01 million lexical tokens, designed to support the development of advanced code generation models, programming assistants, software engineering AI systems, and code intelligence applications.
It consists of real-world DSA implementations covering a wide range of algorithms, data structures, problem-solving… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/DSA-Coding-Problems-and-Solutions-Dataset.Vibe-Coding-InstructBashCoder
🌀 Claude-opus-4.7-TraceInversion-5000x
v1.0 Release
A High-Fidelity Reconstructed CoT Dataset Saturated with the 'Opus Deep Logic Style' via Trace Inversion
📊 5,000 Samples
🧬 Trace Inversion & Negentropy
🛠 SFT & DPO Ready
🔥 Claude 4.7-Max Distillation
🌐 English & Multilingual
💡 What is Trace Inversion?
In Large Language Model (LLM) reasoning distillation, proprietary API models (such as GPT-4/5 and Claude)… See the full description on the dataset page: https://huggingface.co/datasets/Coding-With-Bashir/BashCoder.
