datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
coding
Dataset Card for "livebench/coding"
LiveBench is a benchmark for LLMs designed with test set contamination and objective evaluation in mind. It has the following properties:
LiveBench is designed to limit potential contamination by releasing new questions monthly, as well as having questions based on recently-released datasets, arXiv papers, news articles, and IMDb movie synopses.
Each question has verifiable, objective ground-truth answers, allowing hard questions to be scored… See the full description on the dataset page: https://huggingface.co/datasets/livebench/coding.Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection
🧬 Omni-Frontier Collection
Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package
A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible.
📖 Jump to
What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/SHSLab/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.verifiable-coding-problems
SYNTHETIC-1
This is a subset of the task data used to construct SYNTHETIC-1. You can find the full collection here
verifiable-coding-problems-python
Dataset Card for Verifiable Coding Problems Python 10k
This dataset contains all Python problems from PrimeIntellect's verifiable-coding-problems dataset. We have formatted the verification_info and metadata columns to be proper dictionaries, but otherwise the data is the same. Please see their dataset for more details.
S10-Citadel-Core
Run and deploy your AI Studio app
This contains everything you need to run your app locally.
Run Locally
Prerequisites: Node.js
Install dependencies:
npm install
Set the GEMINI_API_KEY in .env.local to your Gemini API key
Run the app:
npm run dev
llama-coding-agent-evalsreal-pi-coding-agent-traces-sessions
Real Pi Coding Agent Traces Sessions
An aggregated dataset of real human–AI coding agent sessions, collected from 21 independently published Hugging Face datasets and hand-filtered to exclude synthetic or AI-generated content.
Every session is an unedited (but redacted) trace of a real person using pi — an open-source AI coding agent harness — to build, debug, and ship real open-source software. Real prompts, real tool calls, real errors, real backtracking.
Why this… See the full description on the dataset page: https://huggingface.co/datasets/MaxDevv/real-pi-coding-agent-traces-sessions.Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection
🌌 Omni-Frontier Distillation SFT
The Definitive Evolution of Open-Source Distillation & Human-Crafted Expertise
Repository: Manusagents/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection
"The most comprehensive multi‑domain SFT corpus ever assembled — fusing 6.86 million cleaned distillation samples with 9.14 million human‑crafted expert examples across medical, cybersecurity, chemical, robotics, humanities, and more. 16 million… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.kimi-k3-coding-and-debugging-traces
Kimi K3 Coding, Tool Use & Instruction Following Traces
582 TRAJECTORIES · 3,956 TRAINING ROWS · 3 MB PARQUET · 72 MB JSONL
Generated by moonshiner — an open harness for
distilling verified instruction-following, tool-use, and agentic coding traces.
Behavior-preserving instruction-following, tool-use, and agent trajectories
from Kimi K3 (moonshotai/kimi-k3). The category and row-share tables
below describe the actual mix seen during training rather than assuming a… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/kimi-k3-coding-and-debugging-traces.fable-5-coding-and-debugging-traces
Claude Fable 5 Agent Traces
2,380 TRAJECTORIES · 12,490 TRAINING ROWS · 14 MB PARQUET · 663 MB JSONL
Generated by moonshiner — an open harness for
distilling verified instruction-following, tool-use, and agentic coding traces.
Behavior-preserving instruction-following, tool-use, and agent trajectories
from Claude Fable 5 (anthropic/claude-fable-5). The category and row-share tables
below describe the actual mix seen during training rather than assuming a
particular task… See the full description on the dataset page: https://huggingface.co/datasets/DSFFGFG456/fable-5-coding-and-debugging-traces.Nemotron-RL-coding-competitive_coding
Dataset Description:
The Nemotron-RL-coding-competitive_coding dataset is a python-only, reasoning-based, synthetic dataset. It contains competitive coding style problems and their unit test cases. These questions and test cases are collected from CodeContests (deepmind/code_contests), and Open-R1 (open-r1/codeforces)
.
This dataset is released as part of NVIDIA NeMo Gym, a framework for building reinforcement learning environments to train large language models. NeMo Gym… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-coding-competitive_coding.starcoder-python-instruct
StarCoder-Python-Qwen-Instruct
Dataset Description
This dataset contains Python code samples paired with synthetically generated natural language instructions. It is designed for supervised fine-tuning of language models for code generation tasks. The dataset is derived from the Python subset of the bigcode/starcoderdata corpus, and the instructional text for each code sample was generated using the Qwen/Qwen3-Coder-30B-A3B-Instruct-FP8 model.
Creation… See the full description on the dataset page: https://huggingface.co/datasets/OLMo-Coding/starcoder-python-instruct.gpt-5.6-sol-coding-and-debugging-traces
GPT-5.6 Sol Coding & Debugging Traces
Verified software-engineering, independent model-judging, seed-authoring,
defensive-security, and training-harness trajectories from
GPT-5.6 Sol (gpt-5.6-sol) running through the Codex CLI as an
autonomous coding agent. Sessions show the observable development loop:
inspecting repositories, reproducing failures, explaining evidence, editing
files, running compilers and test suites, correcting mistakes, and verifying
the completed result.… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/gpt-5.6-sol-coding-and-debugging-traces.web-research-coding-5m
Web Research + GitHub + Website Coding Dataset
Version: 1.1.0
Total examples: 5,000,000
Splits
train: 4,750,000
validation: 125,000
test: 125,000
Core capabilities
Web search
Web research
Evidence extraction
Fact verification
Multi-hop research
Multi-layer technical analysis
Architecture analysis
Root-cause analysis
Security analysis
Performance analysis
UX analysis
Design analysis
Refactoring
Code review
Debugging
Website coding
Design systems… See the full description on the dataset page: https://huggingface.co/datasets/Lelonthecodeur/web-research-coding-5m.coding_agent_tracesWe release coding agent traces using Claude Code for
Opus
ISL, OSL, ISL_new counts
GPT-oss-120B
ISL, OSL, ISL_new counts and their raw texts
For Opus, only the locally saved files from the harness were used for analysis.
Coding agents take multiple turns to carry out a task from the input prompt. To analyze the token distribution, two models were selected: Anthropic's Opus and OpenAI's gpt-oss-120B. The input sequence length (ISL), output sequence length (OSL) and the uncached, new input… See the full description on the dataset page: https://huggingface.co/datasets/netpreme/coding_agent_traces.Nemotron-RL-coding-quality-filtered
Nemotron coding — quality pool, revision 2
8,201 retained tasks from 16,083 upstream train rows (51.0%).
This is a static quality screen for Python standard-input/standard-output coding RL. It is not a reference-verified gold dataset.
All difficulty levels are eligible. There is no model-accuracy filter, rollout generation, rating cutoff, random subsampling, or 3,200-row cap.
Source: NVIDIA Nemotron-RL-coding-competitive_coding, revision 755d5910fc8646b385e3926eec08c152051cdc07.… See the full description on the dataset page: https://huggingface.co/datasets/hi-todayis-jh/Nemotron-RL-coding-quality-filtered.verifiable-coding-problems-python_decontaminated-testedCoding-Conversational-Dataset-Indicglm-5.2-coding-and-debugging-traces
GLM 5.2 Agent Traces
207 TRAJECTORIES · 1,821 TRAINING ROWS · 1 MB PARQUET · 35 MB JSONL
Generated by moonshiner — an open harness for
distilling verified instruction-following, tool-use, and agentic coding traces.
Behavior-preserving instruction-following, tool-use, and agent trajectories
from GLM 5.2 (glm-5.2). The category and row-share tables
below describe the actual mix seen during training rather than assuming a
particular task domain.
This is an actively growing… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/glm-5.2-coding-and-debugging-traces.wizardlm8x22b-logical-math-coding-sft
自動生成したテキスト
WizardLM 8x22bで生成した論理・数学・コード系のデータです。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。
german-wikipedia-articlesverifiable-coding-problems-python_decontaminated-tested-shuffledhf-coding-tools-traces_april26
HuggingFace AI Coding Tools — Agent Traces
This dataset rehydrates the benchmark results from
davidkling/hf-coding-tools-dashboard
into the JSONL session format consumed by the
Hugging Face Agent Trace Viewer.
What's inside
32 sessions, one per (tool, model, effort, thinking) configuration
9,130 query → response turns total (≈18,260 events)
Tools covered: claude_code, codex, copilot, cursor
Models: claude-opus-4-6, claude-sonnet-4-6, claude-sonnet-4.6, composer-2… See the full description on the dataset page: https://huggingface.co/datasets/clem/hf-coding-tools-traces_april26.hf-coding-tools-traces-run-april12
HuggingFace AI Coding Tools — Agent Traces
This dataset rehydrates the benchmark results from
davidkling/hf-coding-tools-dashboard
into the JSONL session format consumed by the
Hugging Face Agent Trace Viewer.
What's inside
31 sessions, one per (tool, model, effort, thinking) configuration
8,875 query → response turns total (≈17,750 events)
Tools covered: claude_code, codex, copilot, cursor
Models: claude-opus-4-6, claude-sonnet-4-6, claude-sonnet-4.6, composer-2… See the full description on the dataset page: https://huggingface.co/datasets/davidkling/hf-coding-tools-traces-run-april12.Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2
🧬 Omni-Frontier Collection
Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package
A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible.
📖 Jump to
What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2.Nemotron-LeetCode-coding-clean-3.2k
Nemotron + LeetCode Coding Clean 3.2k
3,200 distinct training problems, seed 42, intended for Python coding reinforcement learning. This is a training mix, not a held-out benchmark. It combines the pinned default train Parquet split of Nemotron-RL-coding-competitive_coding with the train JSONL of LeetCodeDataset.
Composition
Source
Questions
Selection
Nemotron / Codeforces
1,856
1000–1600, inclusive
Nemotron / AtCoder
366
300–2000 display difficulty… See the full description on the dataset page: https://huggingface.co/datasets/hi-todayis-jh/Nemotron-LeetCode-coding-clean-3.2k.hf-coding-tools-traces
HuggingFace AI Coding Tools — Agent Traces
This dataset rehydrates the benchmark results from
davidkling/hf-coding-tools-dashboard
into the JSONL session format consumed by the
Hugging Face Agent Trace Viewer.
What's inside
32 sessions, one per (tool, model, effort, thinking) configuration
9,130 query → response turns total (≈18,260 events)
Tools covered: claude_code, codex, copilot, cursor
Models: claude-opus-4-6, claude-sonnet-4-6, claude-sonnet-4.6, composer-2… See the full description on the dataset page: https://huggingface.co/datasets/davidkling/hf-coding-tools-traces.niftyoracle-coding-datawizardlm8x22b-logical-math-coding-sft_additional
自動生成したテキスト
WizardLM 8x22bで生成した論理・数学・コード系のデータです。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。
pi-coding-sessions
