datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nemotron-student-fail-v41-clean-thinking
DeepSeek-V4.1 clean and action-only trajectories with Nemotron outcomes
DeepSeek-V4.1 reward-1 trajectories rebuilt from the complete teacher audit
under v57-test-path-component-boundary+v57-target-source-recheck. The V4.1 reward and trajectory tier do not by themselves prove
that Nemotron failed. Student outcomes are joined from
nemotron-prolike-coverage-audit-20261001.json. A student failure requires either complete
required-test results with reward 0, or an individually… See the full description on the dataset page: https://huggingface.co/datasets/zhiyuanhucs/nemotron-student-fail-v41-clean-thinking.Multilingual-Thinking
Dataset summary
Multilingual-Thinking is a reasoning dataset where the chain-of-thought has been translated from English into one of 4 languages: Spanish, French, Italian, and German. The dataset was created by sampling 1k training samples from the SystemChat subset of SmolTalk2 and translating the reasoning traces with another language model.
This dataset was used in the OpenAI Cookbook to fine-tune the OpenAI gpt-oss models.
You can load the dataset using:
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceH4/Multilingual-Thinking.thinking_droid_lerobot_output_qwen3vlllm-jp-4.1-thinking-sft-data
llm-jp-4.1-thinking-sft-data
Overview
This dataset is a supervised fine-tuning (SFT) dataset used to train llm-jp-4.1-*-thinking models.
This dataset is constructed from prompts and conversations collected from multiple data sources. For most subsets, reasoning processes and final responses used for LLM-jp-4.1 SFT were generated or augmented using gpt-oss-120b.
For the tool-calling and agentic data derived from NVIDIA Nemotron datasets, the original conversations… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/llm-jp-4.1-thinking-sft-data.MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking
MMFineReason-Full-2.3M
The Complete Pre-Selection Dataset — Before Quality Filtering
📖 Overview
MMFineReason-Full-2.3M is the complete pre-selection dataset containing 2.3M samples and 8.8B solution tokens, generated through our reasoning distillation pipeline before the data selection stage. This dataset includes all samples that passed basic template and length validation, but have not undergone correctness verification filtering.
🎯 Key Characteristics… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking.llm-jp-4-thinking-sft-data
llm-jp-4-thinking-sft-data
Overview
This dataset is a supervised fine-tuning (SFT) dataset used to train llm-jp-4-*-thinking models.
This dataset is constructed by extracting prompts from multiple data sources and generating reasoning processes and final responses using gpt-oss-120b.
The splits reasoning_low, reasoning_medium, and reasoning_high correspond to different reasoning effort settings used during generation with gpt-oss-120b.
To support the continued development… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/llm-jp-4-thinking-sft-data.thinking_furniture_bench_dataset_lerobot_output_qwen3vlFineReason-1.8M-Qwen3-VL-235B-Thinking
MMFineReason
Closing the Multimodal Reasoning Gap via Open Data-Centric Methods
Average score across mathematical reasoning and multimodal understanding benchmarks.
📖 Overview
MMFineReason is a large-scale, high-quality multimodal reasoning dataset comprising 1.8M samples and 5.1B solution tokens, featuring detailed reasoning annotations distilled from Qwen3-VL-235B-A22B-Thinking.
🎯 Key Highlights
1.8M High-Quality Samples with 5.1B Solution Tokens… See the full description on the dataset page: https://huggingface.co/datasets/NarsAI/FineReason-1.8M-Qwen3-VL-235B-Thinking.MMFineReason-1.8M-Qwen3-VL-235B-Thinking
MMFineReason
Closing the Multimodal Reasoning Gap via Open Data-Centric Methods
Average score across mathematical reasoning and multimodal understanding benchmarks.
📖 Overview
MMFineReason is a large-scale, high-quality multimodal reasoning dataset comprising 1.8M samples and 5.1B solution tokens, featuring detailed reasoning annotations distilled from Qwen3-VL-235B-A22B-Thinking.
🎯 Key Highlights
1.8M High-Quality Samples with 5.1B Solution Tokens… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-1.8M-Qwen3-VL-235B-Thinking.ioi-eval-openrouter_anthropic_claude-3_7-sonnet_thinking-prompt-mem-limitioi-eval-openrouter_google_gemini-2_0-flash-thinking-exp-prompt-mem-limitMMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking
MMFineReason-Full-2.3M
The Complete Pre-Selection Dataset — Before Quality Filtering
📖 Overview
MMFineReason-Full-2.3M is the complete pre-selection dataset containing 2.3M samples and 8.8B solution tokens, generated through our reasoning distillation pipeline before the data selection stage. This dataset includes all samples that passed basic template and length validation, but have not undergone correctness verification filtering.
🎯 Key Characteristics… See the full description on the dataset page: https://huggingface.co/datasets/ericktwo/MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking.Qwen3-235B-A22B-Thinking-2507_Qwen3-1.7B_AIME_1983_2024llm-jp-4.1-33b-thinking-dpo-data
llm-jp-4.1-33b-thinking-dpo-data
Overview
This dataset is a Direct Preference Optimization (DPO) dataset used to train llm-jp-4.1-33b-thinking.
It is constructed by pairing multiple candidate responses for a given prompt and selecting preferred (chosen) and non-preferred (rejected) responses. The splits reasoning_low, reasoning_medium, and reasoning_high correspond to different reasoning effort settings used during response generation.
The fields chosen_analysis… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/llm-jp-4.1-33b-thinking-dpo-data.llm-jp-4.1-32b-a3b-thinking-dpo-data
llm-jp-4.1-32b-a3b-thinking-dpo-data
Overview
This dataset is a Direct Preference Optimization (DPO) dataset used to train llm-jp-4.1-32b-a3b-thinking.
It is constructed by pairing multiple candidate responses for a given prompt and selecting preferred (chosen) and non-preferred (rejected) responses. The splits reasoning_low, reasoning_medium, and reasoning_high correspond to different reasoning effort settings used during response generation.
The fields… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/llm-jp-4.1-32b-a3b-thinking-dpo-data.epic-thinking
Source
Rows
glaiveai/reasoning-v1-20m
1 999 793
BAAI/OpenSeek-Synthetic-Reasoning-Data-Examples CC
1 267 534
PrimeIntellect/INTELLECT-3-SFT openreasoning_science
1 000 000
PrimeIntellect/INTELLECT-3-SFT am_chat
852 816
nvidia/Nemotron-Cascade-SFT-Stage-1 general
583 612
open-thoughts/OpenThoughts2-1M
541 898
PrimeIntellect/SYNTHETIC-1-SFT-Data
474 810
allenai/Dolci-Think-SFT-7B
334 908
allenai/Dolci-Think-SFT-32B
327 491
GeneralReasoning/GeneralThought-430K
291 946… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/epic-thinking.ablation_nemotron_thinking_32k_with_reasoning_effort
Dataset: ablation_nemotron_thinking_32k_with_reasoning_effort
This dataset was uploaded from /mnt/yulan_pretrain/mount/data_final_train/ablation_nemotron_thinking_32k_with_reasoning_effort/stage_1/tmp/.
explore-thinking-models-internalablation_openmathreasoning_thinking_with_reasoning_effort
Dataset: ablation_openmathreasoning_thinking_with_reasoning_effort
This dataset was uploaded from /mnt/yulan_pretrain/mount/data_final_train/ablation_openmathreasoning_thinking_with_reasoning_effort/stage_1/tmp/.
MMFineReason-1.8M-Qwen3-VL-235B-Thinking
MMFineReason
Closing the Multimodal Reasoning Gap via Open Data-Centric Methods
Average score across mathematical reasoning and multimodal understanding benchmarks.
📖 Overview
MMFineReason is a large-scale, high-quality multimodal reasoning dataset comprising 1.8M samples and 5.1B solution tokens, featuring detailed reasoning annotations distilled from Qwen3-VL-235B-A22B-Thinking.
🎯 Key Highlights
1.8M High-Quality Samples with 5.1B Solution Tokens… See the full description on the dataset page: https://huggingface.co/datasets/Sandeepthakur/MMFineReason-1.8M-Qwen3-VL-235B-Thinking.llm-jp-4.1-8b-thinking-dpo-data
llm-jp-4.1-8b-thinking-dpo-data
Overview
This dataset is a Direct Preference Optimization (DPO) dataset used to train llm-jp-4.1-8b-thinking.
It is constructed by pairing multiple candidate responses for a given prompt and selecting preferred (chosen) and non-preferred (rejected) responses. The splits reasoning_low, reasoning_medium, and reasoning_high correspond to different reasoning effort settings used during response generation.
The fields chosen_analysis… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/llm-jp-4.1-8b-thinking-dpo-data.CodeX-2M-Thinking
Modotte
Note: This dataset is part of the lineup CodeX by Modotte. You can get lots of datasets in this same lineup, with the main focus on providing very high-quality datasets for model training and fine-tuning.
This dataset is fully synthetic, curated from high-quality public sources and enhanced with synthetic data generated using both closed and open-source models. It serves as a strong foundation for instruction-based model tuning and fine-tuning, offering one of the… See the full description on the dataset page: https://huggingface.co/datasets/Modotte/CodeX-2M-Thinking.VLAA-Thinking
SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
🌐 Project Page
• 📄 Arxiv
• 💻 Code
🤗 VLAA-Thinker Family
• 🤔 VLAA-Thinking Dataset
🤗 VLAA-Thinker-Qwen2.5-3B
• 🤗 VLAA-Thinker-Qwen2.5-7B
Both VLAA-Thinker-Qwen2.5-3B and VLAA-Thinker-Qwen2.5-7Bachieve SOTA performance on OpenCompass Multimodal Reasoning Leaderboard as of April 7th, 2025.
Contents
Quick Start 🚀… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/VLAA-Thinking.hermes-function-calling-thinking-V1thinking_fmb_dataset_lerobot_output_qwen3vlthinking-rollouts
thinking-rollouts
Unconstrained rollouts from thinking (chain-of-thought) models on DS-1000 and LiveCodeBench, CoT
saved verbatim alongside the final answer. Format per genlm/rollouts issue #5; schema is a superset
of temperature-sweep-data.
Hive-partitioned Parquet, thinking_mode folded into the model tag:
rollouts/domain=<dataset>/model=<tag>/temp=<temp>/data.parquet (tags like qwen3-8b-think,
qwen3-1.7b-nothink). 100 samples/instance.
Columns: model, thinking_mode, temp… See the full description on the dataset page: https://huggingface.co/datasets/samuki-hf/thinking-rollouts.thinking-benchmark-90
Thinking Benchmark
A calibration pool of 90 competition-mathematics problems assembled to study how output / reasoning-trace length varies with problem difficulty across frontier language models. Part of the Cost of Overthinking research project.
Dataset at a glance
Source
n
Difficulty
Contamination risk
AIME 2026
29
3–5
low
OlymMATH
41
4–6
medium
HMMT February 2026
12
4–5
low
MATH-500
5
2–3
high
FrontierMath-style
3
6
medium
Difficulty is… See the full description on the dataset page: https://huggingface.co/datasets/tyrtleli/thinking-benchmark-90.ablation_openmathreasoning_thinking
Dataset: ablation_openmathreasoning_thinking
This dataset was uploaded from /mnt/yulan_pretrain/mount/data_final_train/ablation_openmathreasoning_thinking/stage_1/tmp/.
openthoughts4-code-9168-prompts-qwen3-30b-a3b-thinking-2507-n16-flattened-logprobs-k16
OpenThoughts-4 Code SDG: Qwen3-30B-A3B-Thinking-2507 (n=16, top-16 logprobs)
Synthetic generations from
Qwen/Qwen3-30B-A3B-Thinking-2507
on the Marin OpenThoughts-4 code SDG prompt
set.
Each prompt is sampled n=16 times, and for every generated token the dataset
stores the chosen-token log probability plus the top-16 log probabilities
over the vocabulary, enabling distillation, KL-style fine-tuning,
reranking, and uncertainty analysis.
Generation setup
Field… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-code-9168-prompts-qwen3-30b-a3b-thinking-2507-n16-flattened-logprobs-k16.mhlc-training-qwen3vl-qwen3_vl_2b_thinking_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Qwen3-VL 2B Thinking hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3vl-qwen3_vl_2b_thinking_hard_mixed_sources_120k.
