datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Code-Reasoning
Code-Reasoning
Dataset Description
Code problem-solving data with reasoning, direct-answer, and task-synthesis subsets. This repository is part of the K2 Horizon collection.
The repository is organized into multiple subsets. Every subset has a train split backed by Parquet shards, which supports Dataset Viewer inspection and streaming access.
K2 Horizon Dataset Series
Dataset repository
Focus
Subsets
IFM/TxT360-v2
Web and… See the full description on the dataset page: https://huggingface.co/datasets/IFM/Code-Reasoning.Math-Reasoning
Math-Reasoning
Dataset Description
Mathematical problem-solving, rewriting, and dialogue data for reasoning-oriented language-model training. This repository is part of the K2 Horizon collection.
The repository is organized into multiple subsets. Every subset has a train split backed by Parquet shards, which supports Dataset Viewer inspection and streaming access.
K2 Horizon Dataset Series
Dataset repository
Focus
Subsets
IFM/TxT360-v2… See the full description on the dataset page: https://huggingface.co/datasets/IFM/Math-Reasoning.medical-o1-reasoning-SFT
News
[2025/04/22] We split the data and kept only the medical SFT dataset (medical_o1_sft.json). The file medical_o1_sft_mix.json contains a mix of medical and general instruction data.
[2025/02/22] We released the distilled dataset from Deepseek-R1 based on medical verifiable problems. You can use it to initialize your models with the reasoning chain from Deepseek-R1.
[2024/12/25] We open-sourced the medical reasoning dataset for SFT, built on medical verifiable problems and an… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/medical-o1-reasoning-SFT.FinMMDocRZebra-CoT
Zebra‑CoT
A diverse large-scale dataset for interleaved vision‑language reasoning traces.
Dataset Description
Zebra‑CoT is a diverse large‑scale dataset with 182,384 samples containing logically coherent interleaved text‑image reasoning traces across four major categories: scientific reasoning, 2D visual reasoning, 3D visual reasoning, and visual logic & strategic games.
Dataset Structure
Each example in Zebra‑CoT consists of:
Problem statement:… See the full description on the dataset page: https://huggingface.co/datasets/multimodal-reasoning-lab/Zebra-CoT.reasoning
Dataset Card for "livebench/reasoning"
LiveBench is a benchmark for LLMs designed with test set contamination and objective evaluation in mind. It has the following properties:
LiveBench is designed to limit potential contamination by releasing new questions monthly, as well as having questions based on recently-released datasets, arXiv papers, news articles, and IMDb movie synopses.
Each question has verifiable, objective ground-truth answers, allowing hard questions to be scored… See the full description on the dataset page: https://huggingface.co/datasets/livebench/reasoning.3d-spatial-reasoning-13d-spatial-reasoningchankhavu-imo-reasoning-traces3d-spatial-reasoning-2stagingRaw generator output behind procedural-pile, one configuration per build, before deduplication and the train/test split.
load_dataset("reasoning-core/staging", "rc14", split="train", streaming=True)
SFT-Reasoning
SFT-Reasoning
Dataset Description
Instruction-following and reasoning data prepared for supervised fine-tuning. This repository is part of the K2 Horizon collection.
The repository is organized into multiple subsets. Every subset has a train split backed by Parquet shards, which supports Dataset Viewer inspection and streaming access.
K2 Horizon Dataset Series
Dataset repository
Focus
Subsets
IFM/TxT360-v2
Web and question-answering text… See the full description on the dataset page: https://huggingface.co/datasets/IFM/SFT-Reasoning.SWE-Bench-Verified-O1-reasoning-high-results
SWE-Bench Verified O1 Dataset
Executive Summary
This repository contains verified reasoning traces from the O1 model evaluating software engineering tasks. Using OpenHands + CodeAct v2.2, we tested O1's bug-fixing capabilities on the SWE-Bench Verified dataset, achieving a 28.8% success rate across 500 test instances.
Overview
This dataset was generated using the CodeAct framework, which aims to improve code generation through enhanced action-based reasoning.… See the full description on the dataset page: https://huggingface.co/datasets/AlexCuadron/SWE-Bench-Verified-O1-reasoning-high-results.Edge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Edge-Agent-Reasoning-WebSearch-260K.Fable-5.1-Max-Reasoning-Filtered-10000x
Dataset Description
This dataset contains 10,000 agentic coding and reasoning multi-turn high-quality traces generated by the new Fable 5.1 model using max reasoning effort.
It holds almost 500,000,000 tokens of step-by-step chain-of-thought programming across multiple complex domains.
It has also been deduplicated and heavily filtered to remove low-quality traces, keeping only high-quality traces.
Dataset Statistics
Metric
Value
Total Examples
10,000… See the full description on the dataset page: https://huggingface.co/datasets/MoreThought/Fable-5.1-Max-Reasoning-Filtered-10000x.gsm-hard
Dataset Summary
This is the harder version of gsm8k math reasoning dataset (https://huggingface.co/datasets/gsm8k).
We construct this dataset by replacing the numbers in the questions of GSM8K with larger numbers that are less common.
Supported Tasks and Leaderboards
This dataset is used to evaluate math reasoning
Languages
English - Numbers
Dataset Structure
dataset = load_dataset("reasoning-machines/gsm-hard")
DatasetDict({
train: Dataset({… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-machines/gsm-hard.II-Medical-Reasoning-SFT
II-Medical-Reasoning-SFT
II-Medical SFT is a curated dataset designed to support the supervised fine-tuning of large language models (LLMs) for medical reasoning tasks. It comprises multi-turn dialogues, clinical case scenarios, and question-answer pairs that reflect the complex reasoning processes encountered in real-world clinical practice.
The dataset is intended to help models develop key competencies such as differential diagnosis, evidence-based decision-making, patient… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Internet/II-Medical-Reasoning-SFT.reasoning-v1-20m
We are excited to release a synthetic reasoning dataset containing 22mil+ general reasoning questions and responses generated using deepseek-ai/DeepSeek-R1-Distill-Llama-70B. While there have been multiple efforts to build open reasoning datasets for math and code tasks, we noticed a lack of large datasets containing reasoning traces for diverse non code/math topics like social and natural sciences, education, creative writing and general conversations, which is why we decided to release this… See the full description on the dataset page: https://huggingface.co/datasets/glaiveai/reasoning-v1-20m.AIME_1983_2024-Reasoning-Paths
News
🌟🌟🌟 Try this dataset in our HuggingFace Space!
🥳🥳🥳 Thrilled to share that this NeurIPS paper was selected as 🏆 #1 Paper of the Day on Oct. 20th!
Sampled Reasoning Paths for the AIME dataset (from 1983 to 2024)
This dataset contains sampled reasoning paths for the AIME_1983_2024 dataset, released as part of the NeurIPS 2025 paper: "A Theoretical Study on Bridging Internal Probability and Self-Consistency for LLM Reasoning" (Arxiv).… See the full description on the dataset page: https://huggingface.co/datasets/WNJXYK/AIME_1983_2024-Reasoning-Paths.Medical-Reasoning-SFT-Baichuan-M3-235B
Medical-Reasoning-SFT-Baichuan-M3-235B
A large-scale medical reasoning dataset generated using baichuan-inc/Baichuan-M3-235B, containing over 124,000 samples with detailed chain-of-thought reasoning for medical and healthcare questions.
Baichuan-M3-235B is ranked #1 on HealthBench Total leaderboard and achieves state-of-the-art performance on medical reasoning benchmarks.
Dataset Overview
Metric
Value
Model
baichuan-inc/Baichuan-M3-235B
Total Samples
124… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-Baichuan-M3-235B.hermes_reasoning_tool_use
TL;DR
51 004 ShareGPT conversations that teach LLMs when, how and whether to call tools.Built with the Nous Research Atropos RL stack in Atropos using a custom MultiTurnToolCallingEnv, and aligned with BFCL v3 evaluation scenarios.Released by @interstellarninja under Apache-2.0.
1 Dataset Highlights
Count
Split
Scenarios covered
Size
51 004
train
single-turn · multi-turn · multi-step · relevance
392 MB
Each row: OpenAI-style conversations… See the full description on the dataset page: https://huggingface.co/datasets/interstellarninja/hermes_reasoning_tool_use.ScholScan
ScholScan
copying_reasoning_task_improved
copying_reasoning_task_improved
Dataset Description
The Enhanced Copying Reasoning Task Dataset is designed to provide a rich resource for analyzing promotional texts and their key elements. This dataset includes a variety of question-and-answer formats, focusing on whether specific phrases are mentioned within the text. Its purpose is to assist in the training of models for natural language understanding tasks, particularly in identifying relevant information in… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/copying_reasoning_task_improved.Fino1_Reasoning_Path_FinQAFino1 is a financial reasoning dataset based on FinQA, with GPT-4o-generated reasoning paths to enhance structured financial question answering.
For more details, please check our paper arxiv.org/abs/2502.08127.
Source Data
Initial Data Collection and Normalization
The dataset originates from FinQA dataset.
Annotations
Annotation Process
We add a prompt and create a reasoning process using GPT-4o for each question-answer pair.
💡 Citation… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/Fino1_Reasoning_Path_FinQA.natural_reasoningNaturalReasoning is a large-scale dataset for general reasoning tasks. It consists of high-quality challenging reasoning questions backtranslated from pretraining corpora DCLM and FineMath. The questions have been deduplicated and decontaminated from popular reasoning benchmarks including MATH, GPQA, MMLU-Pro, MMLU-STEM. For each question, we extract the reference final answer from the original document from the pretraining corpora if possible. We also provide a model-generated response from… See the full description on the dataset page: https://huggingface.co/datasets/facebook/natural_reasoning.reasoningMedical-Reasoning-SFT-Mega
Medical-Reasoning-SFT-Mega
The ultimate medical reasoning dataset - combining 7 state-of-the-art AI models with fair distribution deduplication. 1.79 million unique samples with 3.78 billion tokens of medical chain-of-thought reasoning.
Dataset Overview
Metric
Value
Total Samples
1,789,998 (after deduplication)
Total Tokens
~3.78 Billion
Content Tokens
~2.22 Billion
Reasoning Tokens
~1.56 Billion
Samples with Reasoning
1,789,764 (100.0%)
Unique… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-Mega.kimi-cyber-reasoning
Kimi Cyber Reasoning
997 chain-of-thought records covering 13 cybersecurity disciplines and 4 systems engineering domains, distilled from the Kimi K3 reasoning model via API. Every record provides an explicit step-by-step <think> reasoning trace followed by a technical resolution, unified code diff fix, or structured tool invocation.
The dataset was curated as an anchor set for training, healing, and specializing compact reasoning models on systems security and tool calling… See the full description on the dataset page: https://huggingface.co/datasets/echel0nn1881/kimi-cyber-reasoning.ww2-temporal-reasoning
WWII Temporal Reasoning
A very large synthetic question-answering dataset of calendar arithmetic over World War II events: how many days or years separate two events, what weekday a date fell on, which of several events came first, how long a campaign ran, whether a claimed date or ordering is correct, and dozens of related question shapes -- plus the reverse lookup, what happened on a given date.
Every answer is computed by code, not written freehand. The event names and their… See the full description on the dataset page: https://huggingface.co/datasets/wayneworkman2012/ww2-temporal-reasoning.medical-clinical-reasoning-sft-100k
Medical Clinical Reasoning SFT 100K
A synthetic supervised fine-tuning dataset of 100,000 high-quality medical and clinical reasoning conversations designed to train AI assistants capable of supporting clinical decision-making, documentation, and medical education.
Dataset Description
This dataset covers a broad spectrum of clinical practice scenarios across 10 medical specialty categories. Each record follows the ShareGPT conversation format with a detailed human… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/medical-clinical-reasoning-sft-100k.
