datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
medical-o1-reasoning-SFT
News
[2025/04/22] We split the data and kept only the medical SFT dataset (medical_o1_sft.json). The file medical_o1_sft_mix.json contains a mix of medical and general instruction data.
[2025/02/22] We released the distilled dataset from Deepseek-R1 based on medical verifiable problems. You can use it to initialize your models with the reasoning chain from Deepseek-R1.
[2024/12/25] We open-sourced the medical reasoning dataset for SFT, built on medical verifiable problems and an… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/medical-o1-reasoning-SFT.Fable-5.1-Max-Reasoning-Filtered-10000x
Dataset Description
This dataset contains 10,000 agentic coding and reasoning multi-turn high-quality traces generated by the new Fable 5.1 model using max reasoning effort.
It holds almost 500,000,000 tokens of step-by-step chain-of-thought programming across multiple complex domains.
It has also been deduplicated and heavily filtered to remove low-quality traces, keeping only high-quality traces.
Dataset Statistics
Metric
Value
Total Examples
10,000… See the full description on the dataset page: https://huggingface.co/datasets/MoreThought/Fable-5.1-Max-Reasoning-Filtered-10000x.gsm-hard
Dataset Summary
This is the harder version of gsm8k math reasoning dataset (https://huggingface.co/datasets/gsm8k).
We construct this dataset by replacing the numbers in the questions of GSM8K with larger numbers that are less common.
Supported Tasks and Leaderboards
This dataset is used to evaluate math reasoning
Languages
English - Numbers
Dataset Structure
dataset = load_dataset("reasoning-machines/gsm-hard")
DatasetDict({
train: Dataset({… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-machines/gsm-hard.natural_reasoningNaturalReasoning is a large-scale dataset for general reasoning tasks. It consists of high-quality challenging reasoning questions backtranslated from pretraining corpora DCLM and FineMath. The questions have been deduplicated and decontaminated from popular reasoning benchmarks including MATH, GPQA, MMLU-Pro, MMLU-STEM. For each question, we extract the reference final answer from the original document from the pretraining corpora if possible. We also provide a model-generated response from… See the full description on the dataset page: https://huggingface.co/datasets/facebook/natural_reasoning.reasoningmedical-clinical-reasoning-sft-100k
Medical Clinical Reasoning SFT 100K
A synthetic supervised fine-tuning dataset of 100,000 high-quality medical and clinical reasoning conversations designed to train AI assistants capable of supporting clinical decision-making, documentation, and medical education.
Dataset Description
This dataset covers a broad spectrum of clinical practice scenarios across 10 medical specialty categories. Each record follows the ShareGPT conversation format with a detailed human… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/medical-clinical-reasoning-sft-100k.Reasoning-Table
Reasoning-Table: Exploring Reinforcement Learning for Table Reasoning
The Reasoning-Table dataset is a high-quality, reasoning dataset designed for table reasoning tasks.
📁 Directory Structure
This repository is organized by task. Each subfolder contains task-specific reasoning data, including raw and filtered versions. Here is an overview:
├── fetaqa/
├── feverous/
├── finqa/
├── gsm8k/
├── hitab/
├── hybridqa/
├── multihierttt/
├── ottqa/
├── tabfact/
├── tatqa/
├──… See the full description on the dataset page: https://huggingface.co/datasets/TableQAKit/Reasoning-Table.kimi-cyber-reasoning
Kimi Cyber Reasoning
997 chain-of-thought records covering 13 cybersecurity disciplines and 4 systems engineering domains, distilled from the Kimi K3 reasoning model via API. Every record provides an explicit step-by-step <think> reasoning trace followed by a technical resolution, unified code diff fix, or structured tool invocation.
The dataset was curated as an anchor set for training, healing, and specializing compact reasoning models on systems security and tool calling… See the full description on the dataset page: https://huggingface.co/datasets/echel0nn1881/kimi-cyber-reasoning.Cosmos-Reason1-SFT-Dataset
Dataset Description:
The data format is a pair of video and text annotations. We summarize the data and annotations in Table 4 (SFT), Table 5 (RL), and Table 6 (Benchmark) of the Cosmos-Reason1 paper. We release the annotations for embodied reasoning tasks for BridgeDatav2, RoboVQA, Agibot, HoloAssist, AV, and the videos for the RoboVQA and AV datasets. We additionally release the annotations and videos for the RoboFail dataset for benchmarks. By releasing the dataset, NVIDIA… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Cosmos-Reason1-SFT-Dataset.GMAI-Reasoning10K
GMAI-Reasoning10K
Medical Reasoning dataset used in GMAI-VL-R1
Data description
GMAI-Reasoning10K is a high-quality medical image reasoning dataset containing 10,000 carefully selected samples. The data was collected from 95 medical datasets from reliable sources such as Kaggle, GrandChallenge, and Open-Release, covering 12 imaging modalities including X-ray, CT, and MRI.
Data preprocessing followed the standardization methods from SAMed-20M: 3D data (CT/MRI) had… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/GMAI-Reasoning10K.Superior-Reasoning-SFT-gpt-oss-120b
Superior-Reasoning-SFT-gpt-oss-120b
📣 News
Our dataset ranked #1 on the Hugging Face Datasets Trending leaderboard from January 20 to January 30.
🚀 Overview
The Superior-Reasoning-SFT-gpt-oss-120b dataset is a high-quality, open-source collection containing 435K samples designed to democratize the training of high-performance Long Chain-of-Thought (Long-CoT) models. Unlike standard distilled datasets that rely on random sampling or… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-Apsara/Superior-Reasoning-SFT-gpt-oss-120b.python_functions_reasoningThis is the Python (functions) coding reasoning dataset used to train
Notbad v1.0 Mistral 24B reasoning model.
The reasoning data were sampled from an RL-based self-improved
Mistral-Small-24B-Instruct-2501 model.
The Python functions and instructions were sourced from OpenCoder Dataset Stage1
and from open source projects on Github.
You can try Notbad v1.0 Mistral 24B on chat.labml.ai.
natural_reasoning_rubricsreason-tool-use-demo-1500
Dataset info
The dataset is a selection of reasoning toolcalls data from https://huggingface.co/datasets/interstellarninja/hermes_reasoning_tool_use, which contains data from Hermes-Tools、Glaive-FC、ToolAce、Nvidia-When2Call.
The format has been transformed to adapt llama-factory v1 training pipeline.
virl39k_reasoningclaude-opus-4.6-4.7-reasoning-8.7k
Background
Ended up with some tokens to burn on a Claude Max plan. Assembly began during 4.6 and moved to 4.7. Model is tagged. The development evolved as it went along. The dataset has not been manually reviewed. It's entirely Claude developed.
Clarification on Reasoning
The reasoning is not Claude's actual chain-of-thought (cot) and is not summarized cot. It's a fully synthetic cot created as part of the Assistant response to mimic the type of "thinking"… See the full description on the dataset page: https://huggingface.co/datasets/angrygiraffe/claude-opus-4.6-4.7-reasoning-8.7k.medical-reasoningreasoning-corpus-4K-5M-v1 Reasoning Corpus 5M · Within 5k sequence length
About Dataset
This dataset contains reasoning chains from major AI models, such as: DeepSeek-v4 (both Pro and Flash), DeepSeek-r1 (DS-r1, Llama-DS, Qwen-DS), Qwen3, Qwen3.5/3.6 (both OpenSource and API models), Gemma4-31B derived from many other repositories, and properly filtered to train SLMs.
The dataset has these columns for users to filter out:
repo_id
tok_len
user
thought_trace
assistant
ChatML
Repositories… See the full description on the dataset page: https://huggingface.co/datasets/danie1111/reasoning-corpus-4K-5M-v1.Reason-RFT-CoT-Dataset
🤗 Reason-RFT CoT Dateset
The full dataset used in our project "Reason-RFT: Reinforcement Fine-Tuning for Visual Reasoning".
⭐️ Project │ 🌎 Github │ 🔥 Models │ 📑 ArXiv │ 💬 WeChat
🤖 RoboBrain: Aim to Explore ReasonRFT Paradigm to Enhance RoboBrain's Embodied Reasoning Capabilities.
♣️ Quick Start
Please refer to Dataset Preparation
🔥 Overview
Visual reasoning abilities play a crucial role in understanding complex multimodal… See the full description on the dataset page: https://huggingface.co/datasets/tanhuajie2001/Reason-RFT-CoT-Dataset.rollouts-olmo7b-cue-search
rollouts-olmo7b-cue-search
Model: allenai/Olmo-3-1025-7B (snapshot a81bae42).
Tokenizer: allenai/Olmo-3-1025-7B (snapshot a81bae42).
Protocol: RL-Zero prompt, MATH-500 x 4 rollouts, budget 31,744, T 0.6, top-p 0.95, seed 20260819 (depth-2 exhaustive and n-gram chain: seed 20260821); the top-20 beam nominee screen, ten random-opener arms, every depth-2 opener (84 shards, arm names unique across shards) and the n-gram chain arms.
Rollouts generated on the CSAIL cluster for the… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-cues/rollouts-olmo7b-cue-search.Cybersecurity_Reasoning_Dataset
Cybersecurity Reasoning Dataset (Model-Agnostic)
A model-agnostic re-architecture of the Cybersecurity Reasoning Dataset. The original
corpus was format-bound to the Mistral/Llama ### Instruction: / ### Response: template;
this dataset losslessly separates reasoning content from format, providing one
neutral canonical corpus plus four per-family rendered training variants
(Mistral/Llama, DeepSeek, ChatML, Gemma).
Why this exists. Identical content scored 88.1 on a… See the full description on the dataset page: https://huggingface.co/datasets/dpevzner/Cybersecurity_Reasoning_Dataset.GLM-5.1-Reasoning-1M-Cleaned
GLM-5.1-Reasoning-1M-Cleaned
GLM-5.1-Reasoning-1M-Cleaned is a cleaned and reformatted derivative of Kassadin88/GLM-5.1-1000000x. It preserves the original four-subset layout (main, PHD-Science, Multilingual-STEM, Math) while converting every example into a unified SFT-ready schema with explicit conversations, input, output, domain, and meta fields.
This release was prepared from the original dataset published by Kassadin88.
Summary
Teacher model in the data: GLM-5.1… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/GLM-5.1-Reasoning-1M-Cleaned.High-Coder-Reasoning-Multi-Turn
High-Coder-Reasoning-Multi-Turn
Dataset Description
This dataset contains high-quality, multi-turn coding conversations focused on code critique, transformation (fixing, translating, and repurposing), and architectural analysis. It was generated using a proprietary pipeline targeting the openrouter/hunter-alpha model to simulate expert-level software engineering workflows.
Pipeline Details:
Each sample consists of three turns:
Critique: A detailed… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/High-Coder-Reasoning-Multi-Turn.adaption-financial-reasoning-steps
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-financial_reasoning_steps
This dataset contains pairs of financial analysis questions and their corresponding step-by-step reasoning processes to derive numerical answers. Each entry includes a specific query about corporate metrics like growth rates, percentages, or net changes, followed by explicit arithmetic operations and a final calculated value. The content is structured to… See the full description on the dataset page: https://huggingface.co/datasets/asadullahdogarr/adaption-financial-reasoning-steps.embodied-spatial-reasoning
Embodied Spatial Reasoning Tasks
Dataset Description
This dataset is part of the embodied-spatial-reasoning project, where the agent has to actively explore the environment to determine if certain spatial relationships hold true. The tasks involve spatial reasoning with various objects and scenes. Each task includes a query about the spatial relationships between objects within a scene, which the agent must verify through exploration.
Dataset Structure
The… See the full description on the dataset page: https://huggingface.co/datasets/thanhqt2002/embodied-spatial-reasoning.LongVideo-Reason-4k-Video-Crop-Handoff-20260911
LongVideo-Reason 4k · Video Crop 合成移交包
公开仓库,文件访问需要人工审批。 只有仓库根目录出现 READY.json 且 complete=true 时,才表示所有 QA、视频、pipeline 和校验信息已齐备;此前为准备/上传阶段。
本包用于将原视频和原始 QA 重新合成为视频工具轨迹。它不是已经审核通过的 SFT 数据,也不把原论文 reasoning 当作工具轨迹监督。
内容
文件
用途
data/qa.jsonl
4,000 条原始 LongVideo-Reason train QA、原选项、原答案和来源
videos/*.mp4
配套原视频;与 QA 的 video_path 对应
data/video_manifest.jsonl
每个视频的 SHA-256、CRC、ffprobe 时长、尺寸和镜像来源
data/selection_report.json
最终数量、时长分布、去重和筛选范围… See the full description on the dataset page: https://huggingface.co/datasets/b1intern/LongVideo-Reason-4k-Video-Crop-Handoff-20260911.gsm8k-multilingual-reasoning
gsm8k-multilingual-reasoning
GSM8K with reasoning translated to multiple languages
Schema
{"prompt": "...", "answer": "...", "reasoning": "...", "metadata": {...}}
Usage
from datasets importload_dataset
ds = load_dataset("eddie-OB/gsm8k-multilingual-reasoning")
print(ds["train"][0])
Source
Derived from OpenAI GSM8K.
grounded-visual-spatial-reasoning
Grounded Visual Spatial Reasoning
Code for generating the annotations can be found here: github.com
Dataset Summary
This dataset extends the Visual Spatial Reasoning (VSR) dataset with visual grounding annotations: each caption is annotated with COCO-category object mentions, their positions , and corresponding bounding boxes in the image.
Data instance
Each sample instance has the following structure:
Field
Type
Description
image_file
string… See the full description on the dataset page: https://huggingface.co/datasets/tomhodemon/grounded-visual-spatial-reasoning.Kimi-K2.5-Reasoning-1M-Cleaned
🪐 Kimi-K2.5-Reasoning-1M-Cleaned
Kimi-K2.5-Reasoning-1M-Cleaned is a cleaned derivative of ianncity/KIMI-K2.5-1000000x. It preserves the original four-config layout from the source dataset and rewrites each record into a unified reasoning-SFT schema with id, conversations, input, output, domain, and meta.
Summary
Source dataset: ianncity/KIMI-K2.5-1000000x
Source author: ianncity
Teacher model recorded in meta.teacher_model: KIMI-K2.5
Token lengths computed with… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Kimi-K2.5-Reasoning-1M-Cleaned.Bode-reasoning
Bode-Reasoning
Bode-Reasoning is a comprehensive Portuguese-language dataset specifically designed to enhance reasoning capabilities in Large Language Models (LLMs). This dataset comprises 11,715 instances featuring reasoning traces across multiple-choice and open-ended questions from Brazilian standardized examinations, mathematical problems, and diverse general knowledge topics.
Dataset Details
Dataset Description
This dataset was created to address the… See the full description on the dataset page: https://huggingface.co/datasets/recogna-nlp/Bode-reasoning.
