datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Code-Reasoning
Code-Reasoning
Dataset Description
Code problem-solving data with reasoning, direct-answer, and task-synthesis subsets. This repository is part of the K2 Horizon collection.
The repository is organized into multiple subsets. Every subset has a train split backed by Parquet shards, which supports Dataset Viewer inspection and streaming access.
K2 Horizon Dataset Series
Dataset repository
Focus
Subsets
IFM/TxT360-v2
Web and… See the full description on the dataset page: https://huggingface.co/datasets/IFM/Code-Reasoning.verifiable-code-reasoning
Verifiable Code Reasoning
Execution-verified Python problems with chain-of-thought
Sandbox-checked solutions · Multi-test unit checks · Deduplicated instances · Training-ready sft_text
Overview
Verifiable Code Reasoning is a large-scale dataset of Python coding problems where every kept solution has passed sandboxed unit tests.
Unlike scraped contest dumps or unverified LLM traces, an example enters this release only if:
a reference… See the full description on the dataset page: https://huggingface.co/datasets/smshahbaj/verifiable-code-reasoning.Code-Reasoning
Code-Reasoning
Dataset Description
Code problem-solving data with reasoning, direct-answer, and task-synthesis subsets. This repository is part of the K2 Horizon collection.
The repository is organized into multiple subsets. Every subset has a train split backed by Parquet shards, which supports Dataset Viewer inspection and streaming access.
K2 Horizon Dataset Series
Dataset repository
Focus
Subsets
IFM/TxT360-v2
Web and… See the full description on the dataset page: https://huggingface.co/datasets/LexyJawa/Code-Reasoning.Z1-Code-Reasoning-107K
Z1: Efficient Test-time Scaling with Code
Train Large Language Model to Reason with Shifted Thinking
[📜 Paper] •
[🤗 HF Models] •
[🐱 GitHub]
Details
Please refer to https://github.com/efficientscaling/Z1.
Usage
from datasets import load_dataset
ds = load_dataset("efficientscaling/Z1-Code-Reasoning-107K")["train"]
ds[0]
Citation
@misc{yu2025efficientscaling,
title={Z1: Efficient Test-time Scaling with Code}… See the full description on the dataset page: https://huggingface.co/datasets/efficientscaling/Z1-Code-Reasoning-107K.code-reasoning-phi4-templateSAGE-Code-Reasoningcode-meta-reasoning-cleaned-final-string-idopen-code-reasoning-sftCode-Reasoning
Code-Reasoning: Quality Filtered Dataset
A high-quality, curated subset of the OpenCodeReasoning-2 dataset for training competitive programming models. This dataset has been processed with rigorous quality filtering and question reconstruction to ensure optimal training data.
📊 Dataset Overview
This dataset contains competitive programming problems with high-quality solutions, specifically filtered and processed from the original nvidia/OpenCodeReasoning-2 dataset.… See the full description on the dataset page: https://huggingface.co/datasets/GetSoloTech/Code-Reasoning.code-meta-reasoning-filteredCode-Reasoning
Code-Reasoning: Quality Filtered Dataset
A high-quality, curated subset of the OpenCodeReasoning-2 dataset for training competitive programming models. This dataset has been processed with rigorous quality filtering and question reconstruction to ensure optimal training data.
📊 Dataset Overview
This dataset contains competitive programming problems with high-quality solutions, specifically filtered and processed from the original nvidia/OpenCodeReasoning-2 dataset.… See the full description on the dataset page: https://huggingface.co/datasets/cublya/Code-Reasoning.code-reasoning-thinking-3k-5kopen-code-reasoning-sft-n-32Linny-Code-Reasoning-ShrunkenZ1-Code-Reasoning-Shortest-90KZ1-Code-Reasoning-Longest-33Kopen-code-reasoning-rlvrmath_reasoning_automated_problem_solving_with_code_track_3open-code-reasoning-rlvr-stdioopen-code-reasoning-rlvr-sft-stdioolmo-3-preference-mix-deltas_reasoning-yolo_victoria_hates_code-DECONCode-Reasoning-4k
Code-Reasoning-4k
A curated dataset for code-centric reasoning tasks, combining programming problems, mathematical reasoning, and general instruction-following samples. Despite the name, the dataset currently contains 38,140 samples, reflecting a significant expansion beyond the original “4k” scale.
Overview
This dataset is designed to support training and evaluation of models on:
Code generation and debugging
Algorithmic reasoning
Mathematical problem solving
Light… See the full description on the dataset page: https://huggingface.co/datasets/nphearum/Code-Reasoning-4k.filtered_sky_code_8k_math_10k_rubric_reasoningcrim-code-reasoning-flash-2.5-5k_chunksopen-code-reasoning-sft-n-1olmo-3-preference-mix-deltas_reasoning_nothink-yolo_victoria_hates_codeolmo-3-preference-mix-deltas_reasoning_nothink-yolo_victoria_hates_code-DECONodin-c1-code-reasoning-mix
Entity-27th/odin-c1-code-reasoning-mix
Generated by ML Intern
This dataset repository was generated by ML Intern, an agent for machine learning research and development on the Hugging Face Hub.
Try ML Intern: https://smolagents-ml-intern.hf.space
Source code: https://github.com/huggingface/ml-intern
Usage
from datasets import load_dataset
dataset = load_dataset('Entity-27th/odin-c1-code-reasoning-mix')
open-code-reasoning-sft-n-16Code_obligation_contrats_Tunisie_ReasoningFromat
