datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Code-Reasoning
Code-Reasoning
Dataset Description
Code problem-solving data with reasoning, direct-answer, and task-synthesis subsets. This repository is part of the K2 Horizon collection.
The repository is organized into multiple subsets. Every subset has a train split backed by Parquet shards, which supports Dataset Viewer inspection and streaming access.
K2 Horizon Dataset Series
Dataset repository
Focus
Subsets
IFM/TxT360-v2
Web and… See the full description on the dataset page: https://huggingface.co/datasets/IFM/Code-Reasoning.rStar-Coder
rStar-Coder Dataset
Project GitHub | Paper
Dataset Description
rStar-Coder is a large-scale competitive code problem dataset containing 418K programming problems, 580K long-reasoning solutions, and rich test cases of varying difficulty levels. This dataset aims to enhance code reasoning capabilities in large language models, particularly in handling competitive code problems.
Experiments on Qwen models (1.5B-14B) across various code reasoning benchmarks demonstrate… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/rStar-Coder.CoderForge-Preview
CoderForge-Preview: SOTA Open Dataset for Training Efficient Agents
CoderForge-Preview is the largest open test-verified coding agent dataset.
Fine-tuning Qwen-3 32B on it, we boost SWE-Bench Verified performance 23.0% → 59.4% pass@1 and rank #1 among open-data and #2 among open-weight models ≤32B parameters.
Limitations
Adaptability to different scaffolds: We generated all trajectories using a single scaffold and fixed tool set (no permutations). Models trained via… See the full description on the dataset page: https://huggingface.co/datasets/togethercomputer/CoderForge-Preview.github-repos-pythonThe Github repository retrieval source for [code-rag-bench], containing all Python files from the entire GitHub dump (in github-repos)
LiteCoder-Terminal-RL-preview
LiteCoder-Terminal-RL-preview
Paper | Code | Blog Post
This dataset contains 602 standardized Harbor terminal environments and was released as part of the paper LiteCoder-Terminal: Scaling Long-Horizon Terminal Environments for Learning Language Agents.
Unlike static text-only instructions, these environments are fully executable and are designed to support the training of terminal-based agents.
Environment Generation Pipeline
The lack of high-quality, executable… See the full description on the dataset page: https://huggingface.co/datasets/Lite-Coder/LiteCoder-Terminal-RL-preview.Code-Reasoning
Code-Reasoning
Dataset Description
Code problem-solving data with reasoning, direct-answer, and task-synthesis subsets. This repository is part of the K2 Horizon collection.
The repository is organized into multiple subsets. Every subset has a train split backed by Parquet shards, which supports Dataset Viewer inspection and streaming access.
K2 Horizon Dataset Series
Dataset repository
Focus
Subsets
IFM/TxT360-v2
Web and… See the full description on the dataset page: https://huggingface.co/datasets/LexyJawa/Code-Reasoning.Ling-Coder-SFT
🤗 Hugging Face
🤖 ModelScope
🖥️ GitHub
Ling-Coder Dataset
The Ling-Coder Dataset comprises the following components:
Ling-Coder-SFT: A subset of SFT data used for training Ling-Coder Lite, containing more than 5 million samples.
Ling-Coder-DPO: A subset of DPO data used for training Ling-Coder Lite, containing 250k samples.
Ling-Coder-SyntheticQA: A subset of synthetic data used for annealing training of Ling-Coder Lite, containing more… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/Ling-Coder-SFT.github-codereview
Code Review Dataset
A large-scale dataset of the best human-written code reviews from top GitHub repositories.
Each row captures a moment where a human code reviewer left an inline comment on a pull request, and the author subsequently modified the code in response.
The dataset also includes negative examples — code from the same PRs that passed review without comments — to help models learn when code is acceptable.
This provides a natural signal for training models to:
Generate… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/github-codereview.dolphin-coder
dolphin-coder
This dataset is transformed from https://www.kaggle.com/datasets/erichartford/leetcode-rosetta
it is used to train dolphin-coder model
dream-coder
Program Synthesis Data
Generated program synthesis datasets used to train dreamcoder.
Currently just supports text & list data.
High-Coder-Reasoning-Multi-Turn
High-Coder-Reasoning-Multi-Turn
Dataset Description
This dataset contains high-quality, multi-turn coding conversations focused on code critique, transformation (fixing, translating, and repurposing), and architectural analysis. It was generated using a proprietary pipeline targeting the openrouter/hunter-alpha model to simulate expert-level software engineering workflows.
Pipeline Details:
Each sample consists of three turns:
Critique: A detailed… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/High-Coder-Reasoning-Multi-Turn.mini-coder-trajs-400kGenerated using Qwen 3 Coder 30B A3B, mini-swe-agent, and SWE-smith.
Used to train the mini-coder models
Citation
@article{olmedo2026computational,
title={Computational Arbitrage in AI Model Markets},
author={Olmedo, Ricardo and Sch{\"o}lkopf, Bernhard and Hardt, Moritz},
journal={The International Conference on Machine Learning},
year={2026}
}
X-Coder-SFT-376k
X-Coder: Advancing Competitive Programming with Fully Synthetic Tasks, Solutions, and Tests
Dataset Overview
X-Coder-SFT-376k is a large-scale, fully synthetic dataset for advancing competitive programming.
The dataset comprises 4 subsets with a total of 887,321 synthetic records across 423,883 unique queries.
It is designed for supervised fine-tuning and suitbale for cold start to train code reasoning foundations.
X-Coder-SFT-376k is curated by sota reasoning models.… See the full description on the dataset page: https://huggingface.co/datasets/IIGroup/X-Coder-SFT-376k.MMSD2.0
MMSD2.0: Towards a Reliable Multi-modal Sarcasm Detection System
This is a copy of the dataset uploaded on Hugging Face for easy access. The original data comes from this work, which is an improvement upon a previous study.
Usage
from typing import TypedDict, cast
import pytorch_lightning as pl
from datasets import Dataset, load_dataset
from torch import Tensor
from torch.utils.data import DataLoader
from transformers import CLIPProcessor
class… See the full description on the dataset page: https://huggingface.co/datasets/coderchen01/MMSD2.0.CodeR-Pile
Towards A Generalist Code Embedding Model Based On Massive Data Synthesis
Introduction
This repository contains the synthetic training data introduced in the paper Towards A Generalist Code Embedding Model Based On Massive Data Synthesis. The dataset is designed to enhance text embeddings for code retrieval tasks.
For more details, please refer to our Github repo: CodeR.
Load Dataset
Simple Example
An example to load the dataset:… See the full description on the dataset page: https://huggingface.co/datasets/nebula2025/CodeR-Pile.ClankerDatasetgithub-reposThe entire dump of GitHub repositories.
code-review-bench
Code Review Bench
A paired online-offline benchmark for AI code review.
Splits
online — Stratified sample of 1,135 bot-reviewed PRs, scraped from open-source Github repositories and scored by the online benchmark (15 tools, Feb–Apr 2026).
offline — 136 expert-curated golden issues across 50 PRs (5 repositories).
Provenance
The offline golden issues extend the 50-PR benchmark originally created by Greptile (2025) and refined by Augment (2025). Our… See the full description on the dataset page: https://huggingface.co/datasets/code-review-bench/code-review-bench.CAD-Coder
CAD-Coder Dataset
This is the official dataset for the paper "CAD-Coder: Text-to-CAD Generation with Chain-of-Thought and Geometric Reward".
Accepted at NeurIPS 2025 (Poster)
Dataset Description
CAD-Coder Dataset is a large-scale Text-to-CadQuery dataset containing natural language descriptions of 3D CAD models paired with executable CadQuery Python code. The dataset enables training and evaluating language models to generate parametric CAD code from textual descriptions.… See the full description on the dataset page: https://huggingface.co/datasets/gudo7208/CAD-Coder.Ling-Coder-SFT-English-Clean
Ling-Coder-SFT-English-Clean
A cleaned, English-only version of inclusionAI/Ling-Coder-SFT — one of the largest open-source coding instruction datasets (~5.1M samples). Split by programming language for easy access.
Curated by MnemicAI
Origin Story
While building our Mnemic COCM-COT training pipeline — a multi-language coding instruction dataset with stratified topic sampling — we discovered that 11.44% of Ling-Coder-SFT contains Chinese/CJK characters mixed into what… See the full description on the dataset page: https://huggingface.co/datasets/MnemicAI/Ling-Coder-SFT-English-Clean.swe-mt-combined-coderforge-hero-lego-nex-swezero
fan-shu/swe-mt-combined-coderforge-hero-lego-nex-swezero
Concatenated mid-train dataset for Qwen3 Thinking SFT. Each source subset is loaded
in order and concatenated into a single config so one training epoch visits every
trajectory exactly once (no interleave / no oversampling).
Built from fan-shu/swe-instruct-trajectories-empty-think-inserted.
Source subsets (7)
togethercomputer__CoderForge-Preview
nvidia__SWE-Zero-openhands-trajectories
nex-agi__agent-sft… See the full description on the dataset page: https://huggingface.co/datasets/fan-shu/swe-mt-combined-coderforge-hero-lego-nex-swezero.LiteCoder-Terminal-SFT
LiteCoder-SFT-Terminal
Paper | Code | Blog Post
LiteCoder-SFT-Terminal is a dataset of 11,255 agent trajectories in terminal environments, introduced in the paper LiteCoder-Terminal: Scaling Long-Horizon Terminal Environments for Learning Language Agents.
Fine-tuned on this data, the LiteCoder-Terminal-30b-a3b-sft model achieves 31.5% Pass@1 on Terminal Bench Pro, while the LiteCoder-Terminal-4b-sft model shows distinct gains over its baseline.
Released Artifacts… See the full description on the dataset page: https://huggingface.co/datasets/Lite-Coder/LiteCoder-Terminal-SFT.Ling-Coder-SyntheticQA
🤗 Hugging Face
🤖 ModelScope
🖥️ GitHub
Ling-Coder Dataset
The Ling-Coder Dataset comprises the following components:
Ling-Coder-SFT: A subset of SFT data used for training Ling-Coder Lite, containing more than 5 million samples.
Ling-Coder-DPO: A subset of DPO data used for training Ling-Coder Lite, containing 250k samples.
Ling-Coder-SyntheticQA: A subset of synthetic data used for annealing training of Ling-Coder Lite, containing more… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/Ling-Coder-SyntheticQA.golang-coderQ&A style combined, deduplicated dataset including portions of:
Golang best practices and coding guides (general Q&A) https://huggingface.co/datasets/smcleod/golang-programming-style-best-practices (MIT)
Golang questions (general Q&A) https://huggingface.co/datasets/ExAi/Code-Golang-QA-2k (Apache2)
Golang functions (code & description) https://huggingface.co/datasets/google/code_x_glue_ct_code_to_text (c-uda)
Golang snippets (code & description)… See the full description on the dataset page: https://huggingface.co/datasets/smcleod/golang-coder.X-Coder-RL-40k
X-Coder-RL-40k
X-Coder-RL-40k is a fully synthetic reinforcement learning dataset for competitive programming, containing 40k high-quality tasks with verified test cases.
Dataset Structure
The dataset is organized by difficulty level:
File
Difficulty
part_0000.parquet
Easiest
part_0001.parquet
Easy
part_0002.parquet
Medium
part_0003.parquet
Hard
part_0004.parquet
Hardest
Task Difficulty Distribution
Table: Distribution of Proprietary… See the full description on the dataset page: https://huggingface.co/datasets/IIGroup/X-Coder-RL-40k.code-retrieval-training-datasetds-coder-instruct-v1
Dataset Card for DS Coder Instruct Dataset
DS Coder is a dataset for instruction fine tuning of language models. It is a specialized dataset focusing only on
data science (eg. plotting, data wrangling, machine learnig models, deep learning, and numerical computations). The dataset contains code examples both in R and Python.
The goal of this dataset is to enable creation of small-scale, specialized language model assistants for data science projects.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/ed001/ds-coder-instruct-v1.Rust-Coder
Rust-Coder
Rust-Coder is a comprehensive text dataset designed for Rust programming language learning. It contains 12,000 unique samples focusing on distinct Rust concepts, code snippets, and explanations.
Dataset Structure
Each sample consists of:
id: A unique UUID.
instruction: A prompt or question about a Rust concept.
code: An idiomatic Rust code snippet.
explanation: A detailed explanation of the concept and code.
category: The high-level Rust category (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/Convence/Rust-Coder.Luau-Coder-1.0-Preview-SFT
Luau Coder 1.0 Preview SFT 🦭
This dataset is exceptionally high-quality supervised fine-tuning conversations for a highly capable coding model in Roblox Luau domain.
It prioritize technical correctness, useful engineering judgment, realistic interaction, and efficient explanations over output volume.
This dataset includes & covering:
Multi-turns (4-10 turns)
Dynamic CoT (length)
Dynamic Interleaved Reasoning
Long Context Session
Q/A
Review
Debugging
Bug Fix… See the full description on the dataset page: https://huggingface.co/datasets/khtsly/Luau-Coder-1.0-Preview-SFT.CodeRouterBench
CodeRouterBench
CodeRouterBench is the benchmark data released with Agent-as-a-Router. The
core unit is a complete task-by-model result matrix: every benchmark task has
one recorded result for each of the eight canonical backend models.
Repository: https://github.com/LanceZPF/agent-as-a-router
Optional trained router adapter: Lance1573/acrouter-qwen35-08b-router-lora
Associated Paper
Hugging Face Daily Papers: Agent-as-a-Router: Agentic Model Routing for Coding… See the full description on the dataset page: https://huggingface.co/datasets/Lance1573/CodeRouterBench.
