datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
controlling-reasoning-models-privacy-outputs
Model Outputs Dataset Card
Go to **Files and versions** tab to access the data.
Dataset Description
This dataset contains the raw model generations (reasoning traces and final answers) produced in the experiments described in our paper Controllable Reasoning Models are Private Thinkers. It aggregates outputs for:
two model families: Qwen 3 and Phi 4,
multiple model sizes (1.7B–14B),
five variants per model (baseline, RT-IF–optimized, FA-IF–optimized… See the full description on the dataset page: https://huggingface.co/datasets/haritzpuerto/controlling-reasoning-models-privacy-outputs.reasoning-models-noncritical-artifactsreasoning_world_modelWhy_Reasoning_Models_Collapse_Themselves_in_Reasoning
Why Reasoning Models Collapse Themselves in Reasoning
4 Algorithmic Atoms Reveal the Geometric Truth
Abstract
This paper presents LeftAndRight, a diagnostic framework using four algorithmic primitives (>>, <<, 1, 0) to reveal a fundamental property of transformer representations: they geometrically collapse backward operations, regardless of attention architecture.
Key Discovery
The counterintuitive finding: We initially hypothesized that causal attention masks… See the full description on the dataset page: https://huggingface.co/datasets/OzTianlu/Why_Reasoning_Models_Collapse_Themselves_in_Reasoning.where-larger-models-excel-reasoning-tracesA_Reasoning_Critique_of_Diffusion_Models
A Reasoning Critique of Diffusion Models
Author: Zixi "Oz" Li (李籽溪)
Date: December 12, 2025
Type: Theoretical AI Research (Geometry, Reasoning Theory)
Citation
@misc{oz_lee_2025,
author = { Oz Lee },
title = { A_Reasoning_Critique_of_Diffusion_Models (Revision 267326d) },
year = 2025,
url = { https://huggingface.co/datasets/OzTianlu/A_Reasoning_Critique_of_Diffusion_Models },
doi = { 10.57967/hf/7243 }… See the full description on the dataset page: https://huggingface.co/datasets/OzTianlu/A_Reasoning_Critique_of_Diffusion_Models.LFM2.5-KO-SFT-Stage2-Diverse-KoSWE-Reasoning-LFMChat-Raw
LFM2.5-KO-SFT-Stage2-Diverse-KoSWE-Reasoning-LFMChat-Raw
Stage2 raw LFM chat JSONL shards: Korean domain, behavior, SWE/coding, reasoning, finance, legal, Text2SQL.
This dataset is part of the LFM2.5-8B-A1B-KO-SFT / Agentic SFT workflow.
Main SFT model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-SFT
CPT base model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-CPT-FULL
Agentic follow-up model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-Agentic-SFT… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/LFM2.5-KO-SFT-Stage2-Diverse-KoSWE-Reasoning-LFMChat-Raw.MMMU-Reasoning-Distill-Validation中文版本
Description
MMMU-Reasoning-Distill-Validation is a Multi-Modal reasoning dataset that contains 839 image descriptions and natural language inference data samples. This dataset is built upon the validation set of MMMU. The construction process begins with using Qwen2.5-VL-72B-Instruct for image understanding and generating detailed image descriptions, followed by generating reasoning conversations using the DeepSeek-R1 model. Its main features are as follows:
Use the… See the full description on the dataset page: https://huggingface.co/datasets/modelscope/MMMU-Reasoning-Distill-Validation.Semigroup_Reasoning_Model_A_Scalpel
Semigroup Reasoning Model: A Scalpel
Formalizing Sparse Neural Circuits as Reasoning Dynamics
🎯 Central Question
How do we formalize the interpretability of reasoning processes?
This work establishes reasoning as a semigroup dynamical system, providing the first formal equivalence between sparse neural circuits and algebraic reasoning dynamics. We prove that:
Reasoning is a semigroup orbit problem, not a vector space embedding task.
🔬 Key Contributions… See the full description on the dataset page: https://huggingface.co/datasets/OzTianlu/Semigroup_Reasoning_Model_A_Scalpel.koen-reasoning-calibration-v2ida-reasoning-model
IDA Reasoning Model
This model was trained using Imitation, Distillation, and Amplification (IDA) on multiple reasoning datasets.
Training Details
Teacher Model: deepseek-ai/DeepSeek-R1-Distill-Qwen-7B
Student Model: Qwen/Qwen3-1.7B
Datasets: 4 reasoning datasets
Total Samples: 600
Training Method: IDA (Iterative Distillation and Amplification)
Datasets Used
gsm8k
HuggingFaceH4/MATH-500
MuskumPillerum/General-Knowledge
SAGI-1/reasoningData_200k… See the full description on the dataset page: https://huggingface.co/datasets/ziadrone/ida-reasoning-model.reasoning-models-interpretability-artifacts
Reasoning Models Interpretability Artifacts
This dataset contains intermediate artifacts for studying reasoning traces in open-weight language models. It includes annotated-trace hidden representations and spectral metrics computed over reasoning-step categories.
The artifacts are intended for analysis and sharing, not for direct datasets.load_dataset(...) loading as a tabular dataset.
Contents
annotated_traces_reprs/
<model>/
config.json
index.json… See the full description on the dataset page: https://huggingface.co/datasets/jaygala24/reasoning-models-interpretability-artifacts.ida-reasoning-model1jailbreak-classification-reasoning-modelsnvidia-nemotron-model-reasoning-dataset-turkish
Nemotron Reasoning Challenge - Turkish
Turkish translation of the training data from NVIDIA's Nemotron Model Reasoning Challenge
Each row is a reasoning puzzle framed in an "Alice's Wonderland" setting. Given a few input/output examples, the model needs to figure out the hidden rule and apply it to a new input.
Category
Rows
Description
bit
1602
Hidden bit manipulation rule on 8-bit binary numbers
grav
1597
Falling distance with a modified gravitational constant… See the full description on the dataset page: https://huggingface.co/datasets/mramazan/nvidia-nemotron-model-reasoning-dataset-turkish.LFM2.5-KO-SFT-Stage2-Diverse-KoSWE-Reasoning-LFMChat-4K
LFM2.5-KO-SFT-Stage2-Diverse-KoSWE-Reasoning-LFMChat-4K
Stage2 diverse Korean/SWE/reasoning prepared SFT arrays, LFM tokenizer.
This dataset is part of the LFM2.5-8B-A1B-KO-SFT / Agentic SFT workflow.
Main SFT model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-SFT
CPT base model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-CPT-FULL
Agentic follow-up model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-Agentic-SFT
SFT GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/LFM2.5-KO-SFT-Stage2-Diverse-KoSWE-Reasoning-LFMChat-4K.NVIDIA-Nemotron-Model-Reasoning-Challengehumanoid-world-model-reasoning
Humanoid World Model Reasoning
Dataset for training internal world models and reasoning loops in humanoid AI.
test-reasoning-model-datasetMedical-Knowledge-Benchmark-for-Reasoning-AI-ModelsLink to the original dataset on HuggingFace:
https://huggingface.co/datasets/FreedomIntelligence/medical-o1-reasoning-SFT
We are going to test on 4 free Reasoning AI models:
Kimi K1.5-extended-thinking
Deepseek R1
Qwen3-235B-A22B
Gemini 2.5 Pro
Settings:
Web Search: Disabled❌
Reasoning: Enabled✅
Temperature: Default
korean-reasoning-mixture-20250203-previewCore-Model-SFT-Adv-Reasoningreasoning-enhanced-image-gen-modelquantitative_modeling_reasoning_v2koen-reasoning-calibration-v1scientific_modeling_reasoningadvanced_mathematical_modelling_reasoning_v1reasoning-model-kbkorean-reasoning-mixture-20250203-preview-calibration
