datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Reason_Tuning
🎨 UniReason • Unified Reasoning Framework for World Knowledge–Aligned Image Generation and Editing
UniReason is a unified framework that harmonizes text-to-image generation and image editing through a dual reasoning paradigm. We formulate generation as world knowledge-enhanced planning to inject implicit constraints, and leverage editing capabilities for fine-grained visual refinement to further correct visual errors via self-reflection. This approach… See the full description on the dataset page: https://huggingface.co/datasets/Alex11556666/Reason_Tuning.judged_science_completionsjetson-non-reasoning-benchmark-ollama-15w
Tiny LLM Benchmark — Jetson Orin Nano Super 8GB
Date: 2026-06-07 02:35Backends: ollamaSweep: prompt ∈ {128,512,1024,2048} tok × gen ∈ {64,128,256} tokArtifacts: /home/yuvrajsingh/Desktop/benchmark/smolbenchmark/non-reasoning-models/artifacts/blog-all-20260606-0139-15w
Full Results — ollama
Power = VDD_CPU_GPU_CV avg over aiperf window.
Model
Quant
ISL
OSL
OSL mis%
TTFT avg
p50
p90
p99
T2T avg
p50
p90
p99
ITL avg
p50
p90
p99
Tok/s
Req/s
E2E avg
p50
p90
p99… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/jetson-non-reasoning-benchmark-ollama-15w.jetson-non-reasoning-benchmark-ollama-7w
Tiny LLM Benchmark — Jetson Orin Nano Super 8GB
Date: 2026-06-09 02:38Backends: ollamaSweep: prompt ∈ {128,512,1024,2048} tok × gen ∈ {64,128,256} tokArtifacts: /home/yuvrajsingh/Desktop/benchmark/smolbenchmark/non-reasoning-models/artifacts/blog-all-20260607-0403-7w
Full Results — ollama
Power = VDD_CPU_GPU_CV avg over aiperf window.
Model
Quant
ISL
OSL
OSL mis%
TTFT avg
p50
p90
p99
T2T avg
p50
p90
p99
ITL avg
p50
p90
p99
Tok/s
Req/s
E2E avg
p50
p90
p99… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/jetson-non-reasoning-benchmark-ollama-7w.lsat_logic_games-analytical_reasoningNovel annotated evaluation dataset of LSAT logic games associated with paper:
Lost in the Logic: An Evaluation of Large Language Models’ Reasoning Capabilities on LSAT Logic Games
Arxiv: http://arxiv.org/pdf/2409.19012
If you find this dataset useful, please cite the paper!
@misc{malik2024lostlogicevaluationlarge,
title={Lost in the Logic: An Evaluation of Large Language Models' Reasoning Capabilities on LSAT Logic Games},
author={Saumya Malik},
year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saumyamalik/lsat_logic_games-analytical_reasoning.jetson-non-reasoning-benchmark-ollama-25w
Tiny LLM Benchmark — Jetson Orin Nano Super 8GB
Date: 2026-06-23 06:04Backends: ollamaSweep: prompt ∈ {128,512,1024,2048} tok × gen ∈ {64,128,256} tokArtifacts: /home/yuvrajsingh/Desktop/benchmark/smolbenchmark/benchmark-jetson-nano-orin-super/non-reasoning-models/artifacts/blog-all-20260622-0159-25w
Full Results — ollama
Power = VDD_CPU_GPU_CV avg over aiperf window.
Model
Quant
ISL
OSL
OSL mis%
TTFT avg
p50
p90
p99
T2T avg
p50
p90
p99
ITL avg
p50
p90
p99… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/jetson-non-reasoning-benchmark-ollama-25w.jetson-non-reasoning-benchmark-7w
Tiny LLM Benchmark — Jetson Orin Nano Super 8GB (7W / nvpmodel -m 3)
Date: 2026-05-28 18:03Backend: llama.cpp CUDA (-ngl 99) Concurrency: 1Sweep: prompt in {128,512,1024,2048} tok × gen in {64,128,256} tok
Artifacts: https://huggingface.co/datasets/YuvrajSingh9886/jetson-non-reasoning-benchmark-7w
Note: tok/J computed from per-run start_time/end_time in each aiperf JSON
Full Results
Power = VDD_CPU_GPU_CV average over each aiperf run window (per-run timestamps… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/jetson-non-reasoning-benchmark-7w.jetson-non-reasoning-benchmark-maxn
Tiny LLM Benchmark — Jetson Orin Nano Super 8GB
Date: 2026-05-26 18:18Backend: llama.cpp CUDA (-ngl 99) Concurrency: 1Sweep: prompt in {128,512,1024,2048} tok × gen in {64,128,256} tok
Artifacts: https://huggingface.co/datasets/YuvrajSingh9886/jetson-non-reasoning-benchmark-maxn
Skipped / Failed Models
gemma3-4b (OOM — server failed to start)
Full Results
Cells marked — = OOM (server crashed or skipped). Power = VDD_CPU_GPU_CV average over… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/jetson-non-reasoning-benchmark-maxn.jetson-non-reasoning-benchmark-25w
Tiny LLM Benchmark — Jetson Orin Nano Super 8GB
Date: 2026-05-27 21:12Backend: llama.cpp CUDA (-ngl 99) Concurrency: 1Sweep: prompt in {128,512,1024,2048} tok × gen in {64,128,256} tok
Artifacts: https://huggingface.co/datasets/YuvrajSingh9886/jetson-non-reasoning-benchmark-25w
Skipped / Failed Models
gemma3-4b (server failed to start)
Full Results
Cells marked — = OOM (server crashed or skipped). Power = VDD_CPU_GPU_CV average over aiperf run… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/jetson-non-reasoning-benchmark-25w.jetson-non-reasoning-benchmark-ollama-maxn
Tiny LLM Benchmark — Jetson Orin Nano Super 8GB
Date: 2026-06-22 01:58Backends: ollamaSweep: prompt ∈ {128,512,1024,2048} tok × gen ∈ {64,128,256} tokArtifacts: /home/yuvrajsingh/Desktop/benchmark/smolbenchmark/benchmark-jetson-nano-orin-super/non-reasoning-models/artifacts/blog-all-20260621-1401-maxn
Full Results — ollama
Power = VDD_CPU_GPU_CV avg over aiperf window.
Model
Quant
ISL
OSL
OSL mis%
TTFT avg
p50
p90
p99
T2T avg
p50
p90
p99
ITL avg
p50
p90… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/jetson-non-reasoning-benchmark-ollama-maxn.TranNhiem-Vietnamese-ImageText-Reasoning
TranNhiem Vietnamese Image-Text Reasoning (V-LAION)
Large-scale Vietnamese multimodal reasoning: multi-turn visual question–answering grounded on
natural images, where every answer ships with an explicit chain-of-thought. Reasoning traces
and Answer were synthesized by Qwen3.5-397B-A17B over images from the LAION-derived Vi-Laion-gemini-VQA set.
Curated by: Trần Nhiệm Mình rất welcome cho các hợp tác liên quan tới building Data Engine và Model Training at Scale. Contact… See the full description on the dataset page: https://huggingface.co/datasets/minhnguyent546/TranNhiem-Vietnamese-ImageText-Reasoning.jetson-non-reasoning-benchmark-15w
Tiny LLM Benchmark — Jetson Orin Nano Super 8GB
Date: 2026-05-26 19:00Backend: llama.cpp CUDA (-ngl 99) Concurrency: 1Sweep: prompt in {128,512,1024,2048} tok × gen in {64,128,256} tok
Artifacts: https://huggingface.co/datasets/YuvrajSingh9886/jetson-non-reasoning-benchmark-15w
Skipped / Failed Models
gemma3-4b (OOM — server failed to start)
Full Results
Cells marked — = OOM (server crashed or skipped). Power = VDD_CPU_GPU_CV average over… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/jetson-non-reasoning-benchmark-15w.severity_ablation_scienceseverity_ablation_logicrollouts-olmo7b-cue-search
rollouts-olmo7b-cue-search
Model: allenai/Olmo-3-1025-7B (snapshot a81bae42).
Tokenizer: allenai/Olmo-3-1025-7B (snapshot a81bae42).
Protocol: RL-Zero prompt, MATH-500 x 4 rollouts, budget 31,744, T 0.6, top-p 0.95, seed 20260819 (depth-2 exhaustive and n-gram chain: seed 20260821); the top-20 beam nominee screen, ten random-opener arms, every depth-2 opener (84 shards, arm names unique across shards) and the n-gram chain arms.
Rollouts generated on the CSAIL cluster for the… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-cues/rollouts-olmo7b-cue-search.reason-to-play
VGDL-fMRI: Reason to Play
Human video-game learning, fMRI recordings, model gameplay and representations for
Reason to Play: Behavioral and Brain Alignment Between Frontier LRMs and Human Game Learners,
accepted at NeurIPS 2026.
Research code ·
Interactive results ·
Original human dataset
TL;DR: Explore the replays on the website, or download the human recordings,
model features and processed fMRI inputs for analysis. The complete release is
33,557 files, 4.80 TB (4,804,604,182… See the full description on the dataset page: https://huggingface.co/datasets/csbotos/reason-to-play.ReasonXL-SFT
ReasonXL: A Multilingual Cross-Domain Reasoning Corpus
ReasonXL is a large-scale multilingual reasoning corpus spanning five languages, with 2,538,450 positionally aligned examples per language (12,692,250 rows total). It is designed to support supervised fine-tuning of reasoning models with in-language chain-of-thought traces across diverse technical domains.
Data Generation
English source samples were drawn from 10 existing reasoning datasets, filtered and… See the full description on the dataset page: https://huggingface.co/datasets/toroe/ReasonXL-SFT.Natural-Reasoning-STEM-25KLongVideo-Reason-4k-Video-Crop-Handoff-20260911
LongVideo-Reason 4k · Video Crop 合成移交包
公开仓库,文件访问需要人工审批。 只有仓库根目录出现 READY.json 且 complete=true 时,才表示所有 QA、视频、pipeline 和校验信息已齐备;此前为准备/上传阶段。
本包用于将原视频和原始 QA 重新合成为视频工具轨迹。它不是已经审核通过的 SFT 数据,也不把原论文 reasoning 当作工具轨迹监督。
内容
文件
用途
data/qa.jsonl
4,000 条原始 LongVideo-Reason train QA、原选项、原答案和来源
videos/*.mp4
配套原视频;与 QA 的 video_path 对应
data/video_manifest.jsonl
每个视频的 SHA-256、CRC、ffprobe 时长、尺寸和镜像来源
data/selection_report.json
最终数量、时长分布、去重和筛选范围… See the full description on the dataset page: https://huggingface.co/datasets/b1intern/LongVideo-Reason-4k-Video-Crop-Handoff-20260911.grounded-visual-spatial-reasoning
Grounded Visual Spatial Reasoning
Code for generating the annotations can be found here: github.com
Dataset Summary
This dataset extends the Visual Spatial Reasoning (VSR) dataset with visual grounding annotations: each caption is annotated with COCO-category object mentions, their positions , and corresponding bounding boxes in the image.
Data instance
Each sample instance has the following structure:
Field
Type
Description
image_file
string… See the full description on the dataset page: https://huggingface.co/datasets/tomhodemon/grounded-visual-spatial-reasoning.multilingual-medical-reasoning-tracesThis datasets containes the traces generated to answer multiple-choice medical questions in Italian, Englihs, and Spanish.
The dataset is structured in 3 parts, one per language. Each part is composed by 2 splits, one containing the examples generated from medqa, one from medmcqa.
The columns are:
id, representing an unique identifier
full_question, representing the medical question
options, a dictionary of options to answer the question and their identifiers
list_of_options, a list of the… See the full description on the dataset page: https://huggingface.co/datasets/NLP-FBK/multilingual-medical-reasoning-traces.32B-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids
32B-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids
Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by Qwen3-32B
with thinking mode off. Each thought is a few dense sentences of reasoning about the next
8 tokens after a cut, written from the document prefix alone — the generator never sees the
continuation. Stored thought_text includes the <thought>/</thought> wrapper.
The 4B parity counterpart is… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/32B-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.context
On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models
Charlie Zhang, Graham Neubig,
Xiang Yue
Carnegie Mellon University, Language Technologies Institute
Does Reinforcement Learning Truly Extend Reasoning?
This work explores the discrepancy in views on RL's effectiveness in extending language models' reasoning abilities. Some characterize RL as a capability refiner, while others see it as inducing new compositional skills. This challenge… See the full description on the dataset page: https://huggingface.co/datasets/Interplay-LM-Reasoning/context.chart-reasoning-verified
chart-reasoning-verified
Chart reasoning examples generated from an explicit latent representation.
The data, the question and the answer are computed before the chart is
drawn, so the image is a rendering of known ground truth rather than the
source of it. No model was asked to label anything.
Each row carries both a rendered chart and a text serialisation of the same
chart, so the set is usable for vision-language training and for text-only
language model training without… See the full description on the dataset page: https://huggingface.co/datasets/vinod-anbalagan/chart-reasoning-verified.reason-embed-data
ReasonEmbed: Enhanced Text Embeddings for Reasoning-Intensive Document Retrieval
This repository contains the synthetic training data introduced in the paper ReasonEmbed: Enhanced Text Embeddings for Reasoning-Intensive Document Retrieval. The dataset is designed to enhance text embeddings for reasoning-intensive document retrieval tasks.
Dataset Overview
v0928
This version corresponds to the 81,659 training samples used in the paper. The… See the full description on the dataset page: https://huggingface.co/datasets/hanhainebula/reason-embed-data.4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids
4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids
Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by
Qwen3-4B-Instruct-2507. Each thought is a few dense sentences of reasoning about the next
8 tokens after a cut, written from the document prefix alone — the generator never sees the
continuation. Stored thought_text includes the <thought>/</thought> wrapper.
This is the small-generator parity counterpart of… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.eval-Qwen3-235B-A22B-reasoning
qwen-235b-a22-thinking Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.782
math_pass@1:64_samples
64
0.5%
aime25
0.718
math_pass@1:64_samples
64
0.1%
arenahard
0.939
eval/overall_winrate
500
0.0%
bbh_generative
0.884
extractive_match
1
0.0%
creative-writing-v3
0.775
creative_writing_score
96
0.0%
drop_generative_nous
0.903
drop_acc
1
0.0%
eqbench3
0.800
eqbench_score
135
0.0%
gpqa_diamond
0.697… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Qwen3-235B-A22B-reasoning.financial-economics-reasoning
Model Card
📌 Summary
financial-economics-reasoning dataset was constructed using advanced Inference Distillation techniques. We employed the qwen-3-235b-a22b-thinking-2507 model as the Teacher Model to process the open-source BAAI/IndustryInstruction_Finance-Economics dataset, which contains 122,378 bilingual (Chinese-English) entries in finance, economics, and business.
Unlike standard distillation datasets that only provide final answers, this dataset retains the… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/financial-economics-reasoning.Merged_Reasoning_TasksThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/Merged_Reasoning_Tasks.deepseek-hermes-reasoning-traces
DeepSeek V4 Pro Hermes Reasoning Traces
19,331 multi-turn ChatML + Hermes reasoning traces generated by DeepSeek V4 Pro. Designed for LoRA fine-tuning local models to operate as Hermes Agent instances.
Quick Start
\
Splits
Split
Traces
train
16,431
valid
1,933
test
967
Variants (VRAM-Tiered)
Variant
Max Tokens
Traces
GPU
nano
2,048
15,948
Dev / 7B
budget
4,096
2,149
48GB
standard
8,192
990
64GB
spark
16,384
244… See the full description on the dataset page: https://huggingface.co/datasets/r0b0tlab/deepseek-hermes-reasoning-traces.
