datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Code-Reasoning
Code-Reasoning
Dataset Description
Code problem-solving data with reasoning, direct-answer, and task-synthesis subsets. This repository is part of the K2 Horizon collection.
The repository is organized into multiple subsets. Every subset has a train split backed by Parquet shards, which supports Dataset Viewer inspection and streaming access.
K2 Horizon Dataset Series
Dataset repository
Focus
Subsets
IFM/TxT360-v2
Web and… See the full description on the dataset page: https://huggingface.co/datasets/IFM/Code-Reasoning.verifiable-code-reasoning
Verifiable Code Reasoning
Execution-verified Python problems with chain-of-thought
Sandbox-checked solutions · Multi-test unit checks · Deduplicated instances · Training-ready sft_text
Overview
Verifiable Code Reasoning is a large-scale dataset of Python coding problems where every kept solution has passed sandboxed unit tests.
Unlike scraped contest dumps or unverified LLM traces, an example enters this release only if:
a reference… See the full description on the dataset page: https://huggingface.co/datasets/smshahbaj/verifiable-code-reasoning.Code-Reasoning
Code-Reasoning
Dataset Description
Code problem-solving data with reasoning, direct-answer, and task-synthesis subsets. This repository is part of the K2 Horizon collection.
The repository is organized into multiple subsets. Every subset has a train split backed by Parquet shards, which supports Dataset Viewer inspection and streaming access.
K2 Horizon Dataset Series
Dataset repository
Focus
Subsets
IFM/TxT360-v2
Web and… See the full description on the dataset page: https://huggingface.co/datasets/LexyJawa/Code-Reasoning.Z1-Code-Reasoning-107K
Z1: Efficient Test-time Scaling with Code
Train Large Language Model to Reason with Shifted Thinking
[📜 Paper] •
[🤗 HF Models] •
[🐱 GitHub]
Details
Please refer to https://github.com/efficientscaling/Z1.
Usage
from datasets import load_dataset
ds = load_dataset("efficientscaling/Z1-Code-Reasoning-107K")["train"]
ds[0]
Citation
@misc{yu2025efficientscaling,
title={Z1: Efficient Test-time Scaling with Code}… See the full description on the dataset page: https://huggingface.co/datasets/efficientscaling/Z1-Code-Reasoning-107K.code-reasoning-phi4-templateSAGE-Code-ReasoningNOESIS-1M-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54
NOESIS DORA SFT Dataset
Multilingual supervised fine-tuning dataset built for the NOESIS QwQ+DeepSeek-R1 MoE pipeline.
Released as part of the NOESIS Professional Multilingual Dubbing Automation Platform(framework: DHCF-FNO — Deterministic Hybrid Control Framework for Frozen Neural Operators).
Founder: Ilia Bolotnikov
Organization: AMAImedia.com
X (Twitter): @AMAImediacom
LinkedIn: Ilia Bolotnikov
Telegram: @djbionicl
NOESIS version: v14.8-NT89
Build date: 2026-04… See the full description on the dataset page: https://huggingface.co/datasets/SMH-DEV-AI/NOESIS-1M-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54.NOESIS-50K-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54
⚡ Each donation funds the next large quant.
I host free GGUF or MoE quants as independent research.
Local hardware: Mechrevo Kuangshi GM7AG0M — RTX 3060 Laptop 6GB GDDR6, 64GB DDR5, i7-12700H (14C/20T, 4.7GHz), Windows 11, Samsung 990 Pro.
Good for imatrix and 0.6–35B-class work in RAM. 9B+ and searches need rented H200/Blackwell, typically $100 per quant.
🎉 Boosty🦄 |
☕ Buy Me a Coffee🦄 |
⭐ DonationAlerts🦄
💚 Thanks to Hugging Face for extra storage.🦄… See the full description on the dataset page: https://huggingface.co/datasets/AMAImedia/NOESIS-50K-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54.NOESIS-1M-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54
⚡ Each donation funds the next large quant.
I host free GGUF or MoE quants as independent research.
Local hardware: Mechrevo Kuangshi GM7AG0M — RTX 3060 Laptop 6GB GDDR6, 64GB DDR5, i7-12700H (14C/20T, 4.7GHz), Windows 11, Samsung 990 Pro.
Good for imatrix and 0.6–35B-class work in RAM. 9B+ and searches need rented H200/Blackwell, typically $100 per quant.
🎉 Boosty🦄 |
☕ Buy Me a Coffee🦄 |
⭐ DonationAlerts🦄
💚 Thanks to Hugging Face for extra storage.🦄… See the full description on the dataset page: https://huggingface.co/datasets/AMAImedia/NOESIS-1M-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54.code-meta-reasoning-cleaned-final-string-idCode-Reasoning
Code-Reasoning: Quality Filtered Dataset
A high-quality, curated subset of the OpenCodeReasoning-2 dataset for training competitive programming models. This dataset has been processed with rigorous quality filtering and question reconstruction to ensure optimal training data.
📊 Dataset Overview
This dataset contains competitive programming problems with high-quality solutions, specifically filtered and processed from the original nvidia/OpenCodeReasoning-2 dataset.… See the full description on the dataset page: https://huggingface.co/datasets/GetSoloTech/Code-Reasoning.Deepseek-V4-Reasoning-Code-2500
DeepSeek Reasoning and Code Distillation Dataset
This dataset contains synthetic instruction-response examples generated from coding, reasoning, and math prompts. It was generated with enforce_distillable_text enabled using DeepSeek V4 Pro and DeepSeek V4 Flash through OpenRouter. It is intended for experimentation with supervised fine-tuning, response-style distillation, reasoning-format analysis, and code-assistant behavior research.
The dataset file is:
train.csv
It contains 2… See the full description on the dataset page: https://huggingface.co/datasets/Banaxi-Tech/Deepseek-V4-Reasoning-Code-2500.open-code-reasoning-sftcode-meta-reasoning-filteredCode-Reasoning
Code-Reasoning: Quality Filtered Dataset
A high-quality, curated subset of the OpenCodeReasoning-2 dataset for training competitive programming models. This dataset has been processed with rigorous quality filtering and question reconstruction to ensure optimal training data.
📊 Dataset Overview
This dataset contains competitive programming problems with high-quality solutions, specifically filtered and processed from the original nvidia/OpenCodeReasoning-2 dataset.… See the full description on the dataset page: https://huggingface.co/datasets/cublya/Code-Reasoning.code-reasoning-thinking-3k-5kLinny-Code-Reasoning-Shrunkenmala-code-reasoning-v3open-code-reasoning-sft-n-32reasoning_code_advanced_1m
💻 Reasoning Code Advanced 1M
📖 Dataset Summary
Reasoning Code Advanced 1M is a massive-scale, synthetic dataset specifically engineered to improve the algorithmic reasoning and problem-solving capabilities of Large Language Models (LLMs). Featuring 1,000,000 unique coding samples, this dataset spans multiple programming languages (Python, JS, C++, etc.) and focuses on logic-heavy development tasks.
A key feature of this dataset is its Adaptive Reasoning Architecture.… See the full description on the dataset page: https://huggingface.co/datasets/naimulislam/reasoning_code_advanced_1m.open-code-reasoning-rlvrZ1-Code-Reasoning-Shortest-90Kmala-code-reasoning-v2
MaLA Corpus: Massive Language Adaptation Corpus
This MaLA code and reasoning dataset (V2) is used for training EMMA-500 Llama 3(.1) Mono/Bi model series.
🤗MaLA-LM/emma-500-llama3-8b-mono: CPT model trained on monolingual data mix in 500+ languages
🤗MaLA-LM/emma-500-llama3-8b-bi: CPT model trained on monolingual data mix in 500+ languages + bilingual translation data in 2,500+ language pairs
🤗MaLA-LM/emma-500-llama3.1-8b-mono: CPT model trained on monolingual data mix in… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/mala-code-reasoning-v2.Z1-Code-Reasoning-Longest-33Knvidia-code-reasoning-cleanCodeReasoningPro
CodeReasoningPro
Dataset Summary
CodeReasoningPro is a large-scale synthetic dataset comprising 1,785,725 competitive programming problems in Python, created by XythicK, an MLOps Engineer. Designed for supervised fine-tuning (SFT) of machine learning models for coding tasks, it draws inspiration from datasets like OpenCodeReasoning. The dataset includes problem statements, Python solutions, and reasoning explanations, covering algorithmic topics such as arrays, subarrays… See the full description on the dataset page: https://huggingface.co/datasets/XythicK/CodeReasoningPro.Deepseek-V4-Reasoning-Code-2500
DeepSeek Reasoning and Code Distillation Dataset
This dataset contains synthetic instruction-response examples generated from coding, reasoning, and math prompts. It was generated with enforce_distillable_text enabled using DeepSeek V4 Pro and DeepSeek V4 Flash through OpenRouter. It is intended for experimentation with supervised fine-tuning, response-style distillation, reasoning-format analysis, and code-assistant behavior research.
The dataset file is:
train.csv
It contains… See the full description on the dataset page: https://huggingface.co/datasets/lucsaint/Deepseek-V4-Reasoning-Code-2500.deepseek-v4-reasoning-code-2500
Mirror: lucsaint/Deepseek-V4-Reasoning-Code-2500
Pinned snapshot / mirror of lucsaint/Deepseek-V4-Reasoning-Code-2500, re-hosted for PROTISEC
research reproducibility. Redistributed under the upstream license (apache-2.0)
with attribution — all credit to the original author.
Original author: lucsaint
Source dataset: lucsaint/Deepseek-V4-Reasoning-Code-2500
License: apache-2.0
Family: coding_traces
Mode: full
Rows cached: 2556
Changes vs upstream: cached snapshot, possibly… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/deepseek-v4-reasoning-code-2500.math_reasoning_automated_problem_solving_with_code_track_3mala-code-reasoning
MaLA Corpus: Massive Language Adaptation Corpus
This MaLA code and reasoning dataset is used for training 🤗MaLA-LM/emma-500-llama2-7b.
Dataset Summary
The MaLA Corpus (Massive Language Adaptation) is a comprehensive, multilingual dataset designed to support the continual pre-training of large language models.
This subset contains code, reasoning data, and scientific papers.
Project page: https://mala-lm.github.io
Paper: https://arxiv.org/abs/2409.17892… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/mala-code-reasoning.
