datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LiteCoder-Terminal-RL-preview
LiteCoder-Terminal-RL-preview
Paper | Code | Blog Post
This dataset contains 602 standardized Harbor terminal environments and was released as part of the paper LiteCoder-Terminal: Scaling Long-Horizon Terminal Environments for Learning Language Agents.
Unlike static text-only instructions, these environments are fully executable and are designed to support the training of terminal-based agents.
Environment Generation Pipeline
The lack of high-quality, executable… See the full description on the dataset page: https://huggingface.co/datasets/Lite-Coder/LiteCoder-Terminal-RL-preview.github-codereview
Code Review Dataset
A large-scale dataset of the best human-written code reviews from top GitHub repositories.
Each row captures a moment where a human code reviewer left an inline comment on a pull request, and the author subsequently modified the code in response.
The dataset also includes negative examples — code from the same PRs that passed review without comments — to help models learn when code is acceptable.
This provides a natural signal for training models to:
Generate… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/github-codereview.code-review-bench
Code Review Bench
A paired online-offline benchmark for AI code review.
Splits
online — Stratified sample of 1,135 bot-reviewed PRs, scraped from open-source Github repositories and scored by the online benchmark (15 tools, Feb–Apr 2026).
offline — 136 expert-curated golden issues across 50 PRs (5 repositories).
Provenance
The offline golden issues extend the 50-PR benchmark originally created by Greptile (2025) and refined by Augment (2025). Our… See the full description on the dataset page: https://huggingface.co/datasets/code-review-bench/code-review-bench.CodeRouterBench
CodeRouterBench
CodeRouterBench is the benchmark data released with Agent-as-a-Router. The
core unit is a complete task-by-model result matrix: every benchmark task has
one recorded result for each of the eight canonical backend models.
Repository: https://github.com/LanceZPF/agent-as-a-router
Optional trained router adapter: Lance1573/acrouter-qwen35-08b-router-lora
Associated Paper
Hugging Face Daily Papers: Agent-as-a-Router: Agentic Model Routing for Coding… See the full description on the dataset page: https://huggingface.co/datasets/Lance1573/CodeRouterBench.qwen3-coder-gb10-vs-rtx5090-benchmark
NVIDIA GB10 vs. GeForce RTX 5090 - Local LLM Inference Benchmark
Model: Qwen3-Coder-30B-A3B-InstructFormat: GGUF, Q4_K_M, 18.63 GBRuntime: LM Studio / llama.cppAuthor: Efehan A.Benchmark date: 5 August 2026
This repository contains a decode-focused local inference benchmark comparing an NVIDIA GB10 system with a Windows workstation containing two GeForce RTX 5090 GPUs. Telemetry shows that the inference workload was carried primarily by a single RTX 5090 (GPU 0), while GPU 1… See the full description on the dataset page: https://huggingface.co/datasets/mreltera/qwen3-coder-gb10-vs-rtx5090-benchmark.CoderForge-Preview-32B-SWE-Bench-Verified-Evaluation-trajectoriescodereviewerds-coder-instruct-v2
Dataset Card for DS Coder Instruct v2 Dataset
Changes from v1:
Added WizardLM evol data science samples
Removed R samples from v2
DS Coder is a dataset for instruction fine tuning of language models. It is a specialized dataset focusing only on
data science (eg. plotting, data wrangling, machine learnig models, deep learning, and numerical computations). The dataset contains code examples both in Python (R samples were removed in v2).
The goal of this dataset is to enable… See the full description on the dataset page: https://huggingface.co/datasets/ed001/ds-coder-instruct-v2.Magpie-Qwen2.5-Coder-Pro-300K-v0.1
Project Web: https://magpie-align.github.io/
Arxiv Technical Report: https://arxiv.org/abs/2406.08464
Codes: https://github.com/magpie-align/magpie
Abstract
Click Here
High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Qwen2.5-Coder-Pro-300K-v0.1.Fake_News_GossipCopDataset Source: Ahren09/MMSoc_GossipCop
This is a copied and reformatted version of the Ahren09/MMSoc_GossipCop
text: text of the article (str)
bert_embeddings: (768, )
roberta_embeddings: (768, )
label: (int)
0: real
1: fake
Datasets Distribution:
Train: 9988 (real: 7955, fake: 2033)
Test: 2672 (real: 2169, 503)
Qwen3-Coder-Next-OpenCode-Preference
Dataset Card — OpenCode Rejection Sampling (Preference)
Overview
This dataset contains 10,920 preference pairs for preference-based training (DPO, KTO, SimPO, ORPO, etc.) on competitive programming tasks. Each pair consists of:
Chosen: a candidate solution that passes 100% of test cases
Rejected: a candidate solution that fails, with a fine-grained rejection type label
Pairs are produced via rejection sampling with Qwen3-Coder-Next: 8 candidate solutions are… See the full description on the dataset page: https://huggingface.co/datasets/zake7749/Qwen3-Coder-Next-OpenCode-Preference.spotify-top-10k-songsthis list has been extracted from anna's archive : https://annas-archive.li/blog/spotify/spotify-top-10k-songs-table.html
the script used to scrape can be found here : https://gist.github.com/the-code-rider/96838f5d6ff538377776b6ddbb1c633d
Qwen3-Coder-Next-Open-Code-SFT
Dataset Card — OpenCode Rejection Sampling
Overview
This dataset contains high-quality code reasoning data for training language models on competitive programming tasks. It is produced via rejection sampling with Qwen3-Coder-Next, which would generate multiple candidate solutions per problem, each candidate is executed against test cases in a sandboxed environment, and the results are used to build two complementary training datasets:
SFT dataset (49,374 examples)… See the full description on the dataset page: https://huggingface.co/datasets/zake7749/Qwen3-Coder-Next-Open-Code-SFT.strudel-coder
Claude Code session traces for JohnBeanerson/strudel-coder
This dataset contains redacted Claude Code session traces collected while working on https://github.com/ultralazr/strudel-coder.git. The traces were exported with cc-share-hf and filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each file in the repo root is a redacted Claude Code session in its native JSONL format (one entry per line). HuggingFace… See the full description on the dataset page: https://huggingface.co/datasets/JohnBeanerson/strudel-coder.code_contests_qwen_coder
Dataset Card for code_contests_qwen_coder
This dataset has been created with distilabel.
The pipeline script was uploaded to easily reproduce the dataset:
pipeline.py.
It can be run directly using the CLI:
distilabel pipeline run --script "https://huggingface.co/datasets/argilla/code_contests_qwen_coder/raw/main/pipeline.py"
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel… See the full description on the dataset page: https://huggingface.co/datasets/argilla/code_contests_qwen_coder.details_Qwen__Qwen2.5-Coder-14B-Instruct
Dataset Card for Evaluation run of Qwen/Qwen2.5-Coder-14B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-Coder-14B-Instruct.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Qwen__Qwen2.5-Coder-14B-Instruct.CodeJudge-Eval
CodeJudge-Eval: Can Large Language Models be Good Judges in Code Understanding?
If our project helps you, please give us a star ⭐ on GitHub to support us. 🙏🙏
Introduction
Recent advancements in large language models (LLMs) have showcased impressive code generation capabilities, primarily evaluated through language-to-code benchmarks. However, these benchmarks may not fully capture a model's code understanding abilities. We introduce CodeJudge-Eval (CJ-Eval), a novel… See the full description on the dataset page: https://huggingface.co/datasets/CodeResearch/CodeJudge-Eval.Coder-Stat
Coder-Stat Dataset
Overview
The Coder-Stat dataset is a collection of programming-related data, including problem IDs, programming languages, original statuses, and source code snippets. This dataset is designed to assist in the analysis of coding patterns, error types, and performance metrics.
Dataset Details
Modalities
Tabular: The dataset is structured in a tabular format.
Text: Contains text data, including source code snippets.
Formats… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Coder-Stat.generations-qwen3-coder-next-pre_valrebus-dataset
|🔄 🚍| Re-Bus: A Large and Diverse Multimodal Benchmark for evaluating the ability of Vision-Language Models to understand Rebus Puzzles
Understanding Rebus Puzzles requires a variety of skills such as image recognition, cognitive skills, commonsense reasoning, and multi-step reasoning, making this a challenging task for current Vision-Language Models. In this paper, we present Re-Bus, a large and diverse benchmark of 1,333 English Rebus Puzzles containing different artistic… See the full description on the dataset page: https://huggingface.co/datasets/codergautam/rebus-dataset.Qwen2.5-Coder-0.5B-Flutter-steps-eval
Qwen2.5-Coder-0.5B Flutter — Steps Mode — Validation Results
Dataset Summary
Held-out evaluation results for bbidpa/Qwen2.5-Coder-0.5B-Flutter-steps,
a fine-tune of Qwen2.5-Coder-0.5B for editing Flutter/Dart source files. In steps
mode, the model is given an existing file and an edit instruction and generates a
sequence of localized search/replace edit actions, each mechanically applied to the
current file state before the next action is generated, until the… See the full description on the dataset page: https://huggingface.co/datasets/bbidpa/Qwen2.5-Coder-0.5B-Flutter-steps-eval.Qwen2.5-Coder-0.5B-Flutter-direct-eval
Qwen2.5-Coder-0.5B Flutter — Direct Mode — Validation Results
Dataset Summary
Held-out evaluation results for bbidpa/Qwen2.5-Coder-0.5B-Flutter-direct,
a fine-tune of Qwen2.5-Coder-0.5B for editing Flutter/Dart source files. In direct
mode, the model is given an existing file and an edit instruction and generates the
complete modified file in a single forward pass (as opposed to the steps /
iterative diff-based mode — see the sibling dataset… See the full description on the dataset page: https://huggingface.co/datasets/bbidpa/Qwen2.5-Coder-0.5B-Flutter-direct-eval.github-codereview-dataset
Github-Codereview-Dataset
Made with ❤️ using 🦥 Unsloth Studio
github-codereview-dataset was generated with Unsloth Recipe Studio. It contains 10,000 generated records.
🚀 Quick Start
from datasets import load_dataset
# Load the main dataset
dataset = load_dataset("manishsaini1/github-codereview-dataset", "data", split="train")
df = dataset.to_pandas()
📊 Dataset Summary
📈 Records: 10,000
📋 Columns: 23
📋 Schema & Statistics… See the full description on the dataset page: https://huggingface.co/datasets/manishsaini1/github-codereview-dataset.data_CodeRLPLUSdetails_Qwen__Qwen2.5-Coder-7B-Instruct
Dataset Card for Evaluation run of Qwen/Qwen2.5-Coder-7B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-Coder-7B-Instruct.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Qwen__Qwen2.5-Coder-7B-Instruct.Persian-Wikipedia-Corpus
Overview
This dataset is derived from the Persian Wikipedia Corpus project, which contains parsed articles from the Persian Wikipedia. The original data has been converted into a more accessible format and made available through the HuggingFace datasets library.
Usage
from datasets import load_dataset
dataset = load_dataset("codersan/Persian-Wikipedia-Corpus")
Persian-Wikipedia-Corpus
A complete copy of Persian Wikimedia pages, The dataset contains articles… See the full description on the dataset page: https://huggingface.co/datasets/codersan/Persian-Wikipedia-Corpus.code-route
Code de la route, non-instruct (2025-07-11)
The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects.
Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source language models… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-route.Qwen__Qwen2.5-Coder-7B-Instruct-details
Dataset Card for Evaluation run of Qwen/Qwen2.5-Coder-7B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-Coder-7B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2.5-Coder-7B-Instruct-details.cleand_microsoft_rStar-Coder元データ: https://huggingface.co/datasets/microsoft/rStar-Coder
データ件数: 269,863
平均トークン数: 11674
最大トークン数: 31,184
合計トークン数: 3,150,447,484
ファイル形式: JSONL
ファイルサイズ: 不明
加工内容
synthetic_sftを使用
トークン処理が重たいので、文字数でフィルター
seed_question < 6000
generation < 80000
thinkタグ除去 が中途半端なものを除外
トークナイズ処理(速度向上アップデート
繰り返し除去
Code-Evol-Instruct-OSS
Code-Evol-Instruct-OSS
Summary
Code-Evol-Instruct-OSS is a dataset that was generated with Code Evol-Instruct by prompting open-souce LLMs, WizardLM-13B-v1.2 and WizardCoder-34B-Python.
The underlying process is explained in the paper code-evol-instruct. This algorithm gave birth to famous open-souce code LLMs, WizardCoder-Family.
Our approach
We did not use any closed-source LLMs.
Our seed dataset is sourced from self-instruct-starcoder.
We leverage the… See the full description on the dataset page: https://huggingface.co/datasets/CodeResearch/Code-Evol-Instruct-OSS.
