datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
l1-max-100-6000-math-100plus-topup160k-qwen3-1.7b-verl091-dual-20261005
L1-Max 100–6000: at least 100 responses and 160k output tokens per question
112,758 complete, graded responses across all 1,119 questions in six math datasets. Every question has at least 163,840 cumulative full-output tokens, including thinking.
Model: hi-todayis-jh/l1-max-qwen3-1.7b-compression-100-6k-from80-mixed50-alpha3e-4-bs32-n8-32k-t1-146102-final at 8dee61ed8fbd53f2783a32c9ba6283b4a3d53563. Two compute nodes, 16 independent TP=1 GPU workers.
Sampling: temperature 0.6… See the full description on the dataset page: https://huggingface.co/datasets/hi-todayis-jh/l1-max-100-6000-math-100plus-topup160k-qwen3-1.7b-verl091-dual-20261005.l1-exact-100-6000-math-100plus-topup160k-qwen3-1.7b-verl091-dual-20261005
L1-Exact 100–6000: at least 100 responses and 160k output tokens per question
112,516 complete, graded responses across all 1,119 questions in six math datasets. Every question has at least 163,840 cumulative full-output tokens, including thinking.
Model: hi-todayis-jh/l1-exact-qwen3-1.7b-compression-100-6k-alpha3e-4-bs32-n8-32k-t1-146102-final at ddd566858009b5f97b8ce6d55dcd5060977ed7b7. Two compute nodes, 16 independent TP=1 GPU workers.
Sampling: temperature 0.6, top-p 0.95… See the full description on the dataset page: https://huggingface.co/datasets/hi-todayis-jh/l1-exact-100-6000-math-100plus-topup160k-qwen3-1.7b-verl091-dual-20261005.distilabel-math-preference-dpo
Dataset Card for "distilabel-math-preference-dpo"
More Information needed
brokenarxiv-0826
BrokenArXiv August 2026
Homepage and repository
Homepage: MathArena
Repository: MathArena evaluation code
Benchmark description and prompts: August benchmark update
Dataset summary
This dataset contains 56 plausible but false mathematical statements drawn from arXiv papers submitted in August 2026. Models are asked to prove the statements, and their responses are evaluated for whether they recognize the falsity, acknowledge their inability to… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/brokenarxiv-0826.ODA-Math-460k
ODA-Math-460k
ODA-Math-460k is a large-scale math reasoning dataset curated from top-performing open mathematics corpora (selected via the OpenDataArena leaderboard) and refined through deduplication, benchmark decontamination, LLM-based filtering, and verifier-backed response distillation.It targets a “learnable but challenging” difficulty band: non-trivial for smaller models yet solvable by stronger reasoning models.
🧠 Dataset Summary
Domain: Mathematics… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/ODA-Math-460k.mathdial
Mathdial dataset
https://arxiv.org/abs/2305.14536
MathDial: A Dialogue Tutoring Dataset with Rich Pedagogical Properties Grounded in Math Reasoning Problems.
MathDial is grounded in math word problems as well as student confusions which provide a challenging testbed for creating faithful and equitable dialogue tutoring models able to reason over complex information. Current models achieve high accuracy in solving such problems but they fail in the task of teaching.
Data… See the full description on the dataset page: https://huggingface.co/datasets/eth-nlped/mathdial.MathLake
MathLake: A Large-Scale Mathematics Dataset
MathLake is a massive collection of 8.3 million mathematical problems aggregated from over 50 open-source datasets. Unlike datasets focused on filtering for the highest quality solutions immediately, MathLake prioritizes query comprehensiveness, serving as a universal "raw ore" for researchers to curate, distill, or annotate further. The dataset provides annotations for Difficulty, Format, and Subject, covering diverse mathematical fields… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MathLake.OpenR1-SFT-Math-20k
OpenR1 SFT Math 20k
This is the exact prepared SFT population used in the OpenR1 SFT to GRPO token-level study. It contains 20,144 distinct problems: 20,016 training examples and 128 held-out examples. The 128-example training probe is a subset of the training split. The same examples and split order are used for the four Qwen models and DeepSeek-R1-Distill-Qwen-7B in this experiment.
Splits and configurations
Split
Rows
Purpose
train
20,016
Full SFT… See the full description on the dataset page: https://huggingface.co/datasets/zbeeb/OpenR1-SFT-Math-20k.Crystal-Math-Preview
CrystalMath (Preview)
🏆 AI Mathematical Olympiad – Progress Prize 3: MathCorpus Prize Winner
CrystalMath was selected as the winning submission for the MathCorpus Prize.
A curated dataset of 2,129 competition-level mathematics problems designed for reinforcement learning with verifiable rewards (RLVR) training of frontier reasoning models. CrystalMath is distilled from over 800,000 candidates across 12 public sources through a multi-stage pipeline that enforces high… See the full description on the dataset page: https://huggingface.co/datasets/ycchen/Crystal-Math-Preview.compression-math-100plus-topup128k-qwen3-1.7b-verl091-146103-20261004
Compression math responses: every question has at least 160k output tokens
585,003 complete responses from five Qwen3-1.7B models on six datasets. Every model-question pair now has at least 163,840 total generated output tokens (160 × 1,024).
This update adds 21,973 complete responses (21,972,891 output tokens) to the previous 563,030-response release. The repository name records the original 128k release; the current data includes the completed 160k supplementation.
Metrics… See the full description on the dataset page: https://huggingface.co/datasets/hi-todayis-jh/compression-math-100plus-topup128k-qwen3-1.7b-verl091-146103-20261004.olympiad_style_integer_math_problems
Olympiad Math Corpus
Version: v2.1.1
Release date: 2026-05-03
59,486 synthetically generated olympiad-style math problems with verified integer answers and formal computation graphs.
Loading
from datasets import load_dataset
ds = load_dataset("mihailgribov/olympiad_style_integer_math_problems", split="train")
lemma_applicability is stored as list[{lemma, status}] rather than a sparse dict (required for Arrow-based consumers). To convert to a dict for local use:… See the full description on the dataset page: https://huggingface.co/datasets/mihailgribov/olympiad_style_integer_math_problems.ample-math
AMPLE-Math
5,319 mathematics problems, each with a verified final answer and six references to that same
answer. The references differ only in how much of the reasoning they show, which makes them useful
for studying what a teacher's reference content contributes during distillation.
Problems and original reasoning come from the metadata configuration of
OpenThoughts-114k, and keep its
Apache-2.0 attribution. A question was kept only if all six references exist, every generated… See the full description on the dataset page: https://huggingface.co/datasets/xiuyuz/ample-math.olympiad_style_integer_math_reasoning
Olympiad Math Reasoning Traces
Version: v1.0.2
Release date: 2026-04-19
64,763 full model reasoning traces for olympiad-style math problems with verified integer answers. This dataset contains only correct and non-truncated traces — every record contains a terminal \boxed{...} answer (within the last 500 characters of the response) that matches the expected integer exactly, and none of the responses hit the model's generation-token cap. Intended for distillation and supervised… See the full description on the dataset page: https://huggingface.co/datasets/mihailgribov/olympiad_style_integer_math_reasoning.linalg-bench-math-ai-neurips
LinAlg-Bench: A Benchmark Exposing Structural Failure Modes in LLM Linear Algebra — Where Models Stop Computing and Start Hallucinating
This dataset is the official release accompanying the MATH-AI 2026 NeurIPS workshop paper, "LinAlg-Bench: A Benchmark Exposing Structural Failure Modes in LLM Linear Algebra — Where Models Stop Computing and Start Hallucinating." The exact 660-problem core evaluated in that paper (9 tasks × 3 matrix sizes, 6,600 model outputs, 1,156… See the full description on the dataset page: https://huggingface.co/datasets/mst-ai/linalg-bench-math-ai-neurips.deepseek-v4-pro-math-cot-1k
DeepSeek V4 Pro Math CoT 1K
A small, high-signal supervised-fine-tuning (SFT) dataset of math reasoning traces. Problems were sampled from a Nemotron math problem set (originally sourced from StackExchange-Math and AoPS), answered by DeepSeek V4 Pro with thinking enabled at high reasoning effort, then independently reviewed by DeepSeek V4 Flash for correctness against the expected answer. Pathological reasoning traces (looping, run-away length, excessive Wait-style backtracking)… See the full description on the dataset page: https://huggingface.co/datasets/blythet/deepseek-v4-pro-math-cot-1k.pika-math-generations
PIKA MATH Generations Dataset
A comprehensive dataset of MATH problem solutions generated by different language models with various sampling parameters.
Paper (arXiv) | GitHub Repository
Dataset Description
This dataset contains code generation results from the MATH Dataset evaluated across multiple models. It was created to support the PIKA (Probe-Informed K-Aware Routing) project, which investigates how LLMs encode their own likelihood of success in their internal… See the full description on the dataset page: https://huggingface.co/datasets/CoffeeGitta/pika-math-generations.alpaca_jp_math
alpaca_jp_math
alpaca_jp_mathは、
Stanford Alpacaの手法
mistralai/Mixtral-8x22B-Instruct-v0.1
で作った合成データ(Synthetic data)です。モデルの利用にはDeepinfraを利用しています。
また、"_cleaned"がついたデータセットは以下の手法で精査されています。
pythonの計算結果がきちんと、テキストの計算結果が同等であるか確認
LLM(mistralai/Mixtral-8x22B-Instruct-v0.1)による確認(詳細は下記)
code_result, text_resultは小数第三位で四捨五入してあります。
Dataset Details
Dataset Description
Curated by: HachiMLLanguage(s) (NLP): Japanese
License: Apache 2.0
Github: Alpaca-jp… See the full description on the dataset page: https://huggingface.co/datasets/HachiML/alpaca_jp_math.math_benchmark_test_saturation
LLM Leaderboard Data for Hendrycks MATH Dataset (2022–2024)
This dataset aggregates yearly performance (2022–2024) of large language models (LLMs) on the Hendrycks MATH benchmark. It is specifically compiled to explore performance evolution, benchmark saturation, parameter scaling trends, and evaluation metrics of foundation models solving complex math word problems.
Original source data: Math Word Problem Solving on MATH (Papers with Code)
About Hendrycks' MATH… See the full description on the dataset page: https://huggingface.co/datasets/nlile/math_benchmark_test_saturation.dart-math-pool-gsm8k-query-info
[!NOTE]
This dataset is the synthesis information of queries from the GSM8K training set,
such as the numbers of raw/correct samples of each synthesis job.
Usually used with dart-math-pool-gsm8k.
🎯 DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving
📝 Paper@arXiv | 🤗 Datasets&Models@HF | 🐱 Code@GitHub
🐦 Thread@X(Twitter) | 🐶 中文博客@知乎 | 📊 Leaderboard@PapersWithCode | 📑 BibTeX
Datasets: DART-Math
DART-Math datasets are the… See the full description on the dataset page: https://huggingface.co/datasets/hkust-nlp/dart-math-pool-gsm8k-query-info.math-intuition-reasoning-traces
math-intuition reasoning traces
Full chain-of-thought traces from 7 reasoning models on the same 4,020 problems, graded
by each problem family's own verifier.
Questions come from
amphora/math-intuition-20260908-402-easy-10
— 402 arXiv-derived problem families x 10 seeds, easy preset. Every row here refers to an id
in that dataset, so prompts and the instance cache can be joined from it.
Generation settings
Identical for every model, so the traces are directly… See the full description on the dataset page: https://huggingface.co/datasets/amphora/math-intuition-reasoning-traces.adaption-math-worked-solutions-raw
NuminaMath Worked Solutions
Competition and school maths problems with step-by-step solutions.
Rows
3,000
Domain
mathematics
Format
data.parquet, one row per example
Licence
apache-2.0
Built for
supervised fine-tuning (SFT) experiments on Adaption AutoScientist
Columns
Column
Description
original_prompt
The prompt (user turn) as uploaded.
original_completion
The target response as uploaded.
enhanced_prompt
Empty in this… See the full description on the dataset page: https://huggingface.co/datasets/rodriguescarson/adaption-math-worked-solutions-raw.math500-cot-deepseek-r1-1.5b
MATH-500 CoT completions (DeepSeek-R1-Distill-Qwen-1.5B)
Successful chain-of-thought completions for HuggingFaceH4/MATH-500 test problems, generated with deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B via vLLM.
Files
File
Description
records.parquet
Main dataset: correct completions as token IDs
manifest.json
Schema, tokenizer, run ids, decoding config
problem_index.json
unique_id → problem_idx in MATH-500 test
subject_max_tokens.json
Per-subject completion… See the full description on the dataset page: https://huggingface.co/datasets/ChrisMcCormick/math500-cot-deepseek-r1-1.5b.qwen25-math-sft-long-15k-v1
Qwen2.5 Math SFT Long 15K v1
A reproducible Long CoT SFT dataset for Qwen2.5-Math-1.5B with max context 3,840 tokens (below 4,096 limit).
Source: open-r1/OpenR1-Math-220k
Total samples: 15,000
Decontaminated against GSM8K, SVAMP, MATH-500, AIME 2026.
speculative-reasoning-matheval-4b-sweep
MathEval sweep — speculative reasoning on Qwen3-4B
Per-sample generations and grading for four arms of a MathEval run, measuring what
speculative reasoning costs and saves against a base model that does not speculate.
Code and write-up: yurun-yuan/speculative-reasoning
— see docs/05-rl-4b.md.
The arms
All four answer the same 1,547 MathEval problems under a 20,000-token response budget.
split
model
runtime
base_plain
Qwen/Qwen3-4B
plain — no speculation… See the full description on the dataset page: https://huggingface.co/datasets/yyuan244/speculative-reasoning-matheval-4b-sweep.qwen25-math-sft-long-think7k-v1
Qwen2.5 Math Long-CoT thinking 7k
2,500 train and 250 validation rows. Long examples have 4,096–7,168 thinking tokens; complete ChatML samples including final answer and EOS fit in 8,192 tokens.
Long source: open-r1/OpenR1-Math-220k at e4e141ec9dea9f8326f4d347be56105859b2bd68 (only math-verified generations).
Short replay source: danghoang2005/qwen25-math-sft-long-15k-v1 at 32f0afc54a09c4963ae2478cf93871025abda901.
Tokenizer: Qwen/Qwen2.5-Math-1.5B at… See the full description on the dataset page: https://huggingface.co/datasets/danghoang2005/qwen25-math-sft-long-think7k-v1.dart-math-pool-math-query-info
[!NOTE]
This dataset is the synthesis information of queries from the MATH training set,
such as the numbers of raw/correct samples of each synthesis job.
Usually used with dart-math-pool-math.
🎯 DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving
📝 Paper@arXiv | 🤗 Datasets&Models@HF | 🐱 Code@GitHub
🐦 Thread@X(Twitter) | 🐶 中文博客@知乎 | 📊 Leaderboard@PapersWithCode | 📑 BibTeX
Datasets: DART-Math
DART-Math datasets are the state-of-the-art… See the full description on the dataset page: https://huggingface.co/datasets/hkust-nlp/dart-math-pool-math-query-info.dapo-math-17k-qwen3-1.7b-base-n8
DAPO-Math-17k sampled with Qwen3-1.7B-Base, n=8
17398 problems from the RL training set, each sampled 8 times and scored with
the reward function the RL runs themselves used.
The point of this dataset is to be comparable with what the RL runs actually saw, so
every sampling knob is taken from the live training config or from the default that
config falls through to. Two of them are not in the config file at all and would be
wrong if guessed: top_k = -1 and min_tokens = 1.… See the full description on the dataset page: https://huggingface.co/datasets/RyanYr/dapo-math-17k-qwen3-1.7b-base-n8.hintedselfteacher-nemotron-math-v2-AoPS
hintedselfteacher-nemotron-math-v2-AoPS
This dataset contains a training-ready hinted self-teacher split derived from the AoPS split of nvidia/Nemotron-Math-v2.
The source problems were filtered to the AoPS split with the medium/notool solve rate between 2 and 6. Hints were generated with GPT-5.5 medium using an h17_nt hint-generation prompt. This hint type was close to the best hint type found after doing hint mutations, based on qualitative analysis of token-level hinted… See the full description on the dataset page: https://huggingface.co/datasets/ar0cket1/hintedselfteacher-nemotron-math-v2-AoPS.adaption-math-worked-solutions-mixtral-raw
NuminaMath Worked Solutions
Competition and school maths problems with step-by-step solutions.
Rows
3,000
Domain
mathematics
Format
data.parquet, one row per example
Licence
apache-2.0
Built for
supervised fine-tuning (SFT) experiments on Adaption AutoScientist
Columns
Column
Description
original_prompt
The prompt (user turn) as uploaded.
original_completion
The target response as uploaded.
enhanced_prompt
Empty in this… See the full description on the dataset page: https://huggingface.co/datasets/rodriguescarson/adaption-math-worked-solutions-mixtral-raw.PUM-MATH
PUM Prefix Pair Dataset
This dataset contains pairwise prefix preference examples for gain-based evaluation of LLM reasoning. Each example compares two partial reasoning prefixes for the same math problem and records which prefix is preferred according to outcome-grounded prefix utility.
The dataset is associated with the paper:
From Correctness to Utility: Gain-Based Prefix Evaluation for LLM Reasoning
Dataset Description
Reasoning prefixes can strongly affect… See the full description on the dataset page: https://huggingface.co/datasets/zhiqix/PUM-MATH.
