datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MATH-500
Dataset Card for MATH-500
This dataset contains a subset of 500 problems from the MATH benchmark that OpenAI created in their Let's Verify Step by Step paper. See their GitHub repo for the source file: https://github.com/openai/prm800k/tree/main?tab=readme-ov-file#math-splits
MathInstruct
🦣 MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning
MathInstruct is a meticulously curated instruction tuning dataset that is lightweight yet generalizable. MathInstruct is compiled from 13 math rationale datasets, six of which are newly curated by this work. It uniquely focuses on the hybrid use of chain-of-thought (CoT) and program-of-thought (PoT) rationales, and ensures extensive coverage of diverse mathematical fields.
Project Page:… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/MathInstruct.swallow-math-v2
SwallowMath-v2
Resources
📑 arXiv: Read our paper for detailed methodology at arXiv:2505.02881.
🤗 Sister Dataset: Discover SwallowCode2, our companion dataset for code generation.
🧮 What is it?
SwallowMath-v2 is a large-scale mathematical dataset containing 32 billion tokens, developed as the successor to SwallowMath-v1.
Building on the success of v1, this release aims to construct a larger-scale and more permissively licensed corpus to support open and… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-math-v2.Nemotron-Math-Proofs-v3-SFT
Nemotron-Math-Proofs-v3-SFT
Dataset Description:
Nemotron-Math-Proofs-v3-SFT is a long-form mathematical reasoning dataset containing proof-generation, proof-refinement, verification, and meta-verification traces. The release contains 414,890 samples representing 15,818 unique problems after quality filtering.
The source pool contains 15,879 hard proof problems selected from the AoPS subset of nvidia/Nemotron-Math-Proofs-v1. Responses are generated using… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Math-Proofs-v3-SFT.formal-math-autoformalization
Formal Math Autoformalization Dataset
A growing, CC0 public-domain corpus of ⟨natural-language statement ↔ Lean 4 statement + proof⟩ pairs, contributed through the Agentic Commons network.
Why this is scarce data. Mathlib already contains millions of proven Lean theorems — but as bare Lean, with no paired natural language:
theorem add_comm (a b : ℕ) : a + b = b + a := ... -- no "addition on naturals is commutative" attached
The scarce, valuable artifact is the pairing of the… See the full description on the dataset page: https://huggingface.co/datasets/AgenticCommons/formal-math-autoformalization.open-math-courses
Open Mathematics Courses
Graduate and research-level mathematics lessons with complete proofs, worked examples and solved exercises, one row
per lesson. The lessons are the Markdown sources of the public site https://kokunoyumeto.github.io/open-math-courses-public/ (snapshot commit 4219d3885c68).
3,145 lessons in 140 courses, about 16,927 thousand words.
Fields: course_id, course_title, lesson_title, source_path, page_url, authorship, license, words,
text (Markdown with TeX… See the full description on the dataset page: https://huggingface.co/datasets/KokunoYumeto/open-math-courses.StackMathQA
StackMathQA
StackMathQA: A Curated Collection of 2 Million Mathematical Questions and Answers Sourced from Stack Exchange
StackMathQA is a meticulously curated collection of 2 million mathematical questions and answers, sourced from various Stack Exchange sites. This repository is designed to serve as a comprehensive resource for researchers, educators, and enthusiasts in the field of mathematics and AI research.
Configs
configs:
- config_name: stackmathqa1600k… See the full description on the dataset page: https://huggingface.co/datasets/math-ai/StackMathQA.mathmetics-dataset-intmax
Transformer Math Dataset (200,000,000 Samples Sharded)
High-precision synthetic mathematical expression dataset generated for training sequence-to-sequence math Transformers in JAX/Flax.
Dataset Structure
Total Samples: 200,000,000
Shard Format: JSONL sharded files (100,000 samples per shard)
Supported Operations: +, -, *, /, ^, sin, cos, tan, log, ln, exp, sqrt, abs
Expression Depth Range: Depth 4 to 6
Integer Operand Ratio: 80%
Data Fields
Each… See the full description on the dataset page: https://huggingface.co/datasets/saidurga001301/mathmetics-dataset-intmax.AceReason-Math
AceReason-Math Dataset
Overview
AceReason-Math is a high quality, verfiable, challenging and diverse math dataset for training math reasoning model using reinforcement leraning. This dataset contains
49K math problems and answer sourced from NuminaMath and DeepScaler-Preview
applying filtering rules to exclude unsuitable data (e.g., multiple sub-questions, multiple-choice, true/false, long and complex answers, proof, figure)
this dataset was used to train… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/AceReason-Math.swallow-math
SwallowMath
October 21, 2025: Newer versions are available: SwallowCode-v2 and SwallowMath-v2 have been released with improved rewriting pipelines.
Resources
🐙 GitHub: Explore the project repository, including pipeline code and prompts at rioyokotalab/swallow-code-math.
📑 arXiv: Read our paper for detailed methodology and results at arXiv:2505.02881.
🤗 Sister Dataset: Discover SwallowCode, our companion dataset for code generation.
What is it?… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-math.Nemotron-Math-Proofs-v1
Nemotron-Math-Proofs-v1
Paper: Nemotron-Math: Efficient Long-Context Distillation of Mathematical Reasoning from Multi-Mode SupervisionCode: https://github.com/NVIDIA/NeMo-SkillsDocumentation: Nemotron-MathProofs-v1 documentation
Dataset Description:
Nemotron-Math-Proofs-v1 is a large-scale mathematical reasoning dataset containing ~580k natural language proof problems, ~550k formalizations into theorem statements in Lean 4, and ~900k model-generated reasoning… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Math-Proofs-v1.MSVAMPMathOlympiadBenchThis repository contains the MathOlympiadBench dataset, which is introduced in the paper Goedel-Prover-V2: Scaling Formal Theorem Proving with Scaffolded Data Synthesis and Self-Correction.
Project Page: https://blog.goedel-prover.com
Code Repository: https://github.com/Goedel-LM/Goedel-Prover-V2
MathOlympiadBench (Math Olympiad) comprises human-verified formalizations of Olympiad-level mathematical competition problems, sourced from Compfiles and IMOSLLean4 repository. MathOlympiadBench… See the full description on the dataset page: https://huggingface.co/datasets/Goedel-LM/MathOlympiadBench.Nemotron-Math-Proofs-v2
Nemotron-Math-Proofs-v2
Dataset Description:
Nemotron-Math-Proofs-v2 is a mathematical proof-generation, verification, and meta-verification trace dataset. The problems are sourced from nvidia/Nemotron-Math-Proofs-v1 only taking the AoPS subset. The release contains 82,737 samples across 5,752 unique problems.
For this version, solutions are generated using DeepSeek-V4-Pro on Max inference mode. The generation pipeline produces proofs, verification traces, and… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Math-Proofs-v2.Nemotron-RL-Math-v2
Nemotron-RL-Math-v2
Dataset Description:
Nemotron-RL-Math-v2 is a small curated set of mathematical problems selected for reinforcement learning. The dataset is designed for RL training workflows where problems have verifiable answers or other validation signals suitable for Reinforcement Learning from Verifiable Rewards (RLVR).
Problems are sourced from AoPS, StackExchange-derived math data held out from the Nemotron-SFT-Math-v4 SFT set… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Math-v2.math500
Dataset Card for MATH-500
This dataset contains a subset of 500 problems from the MATH benchmark that OpenAI created in their Let's Verify Step by Step paper. See their GitHub repo for the source file: https://github.com/openai/prm800k/tree/main?tab=readme-ov-file#math-splits
mathdial
Mathdial dataset
https://arxiv.org/abs/2305.14536
MathDial: A Dialogue Tutoring Dataset with Rich Pedagogical Properties Grounded in Math Reasoning Problems.
MathDial is grounded in math word problems as well as student confusions which provide a challenging testbed for creating faithful and equitable dialogue tutoring models able to reason over complex information. Current models achieve high accuracy in solving such problems but they fail in the task of teaching.
Data… See the full description on the dataset page: https://huggingface.co/datasets/eth-nlped/mathdial.Nemotron-Math-Proofs-v3-RL
Nemotron-Math-Proofs-v3-RL
Dataset Description:
Nemotron-Math-Proofs-v3-RL is a long-form mathematical reasoning dataset for reinforcement learning. The release contains 9,597 proof-generation prompts.
The dataset uses NeMo Gym-compatible, single-turn user prompts derived from hard proof problems in the AoPS subset of nvidia/Nemotron-Math-Proofs-v1. The train split asks the policy to produce a rigorous solution and self-evaluation. Policy responses and realized… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Math-Proofs-v3-RL.mathmetics-dataset-custom
Transformer Math Dataset (54,000,000 Samples Sharded)
High-precision synthetic mathematical expression dataset generated for training sequence-to-sequence math Transformers in JAX/Flax.
Dataset Structure
Total Samples: 54,000,000
Shard Format: JSONL sharded files (100,000 samples per shard)
Supported Operations: +, -, *, /, ^, sin, cos, tan, log, ln, exp, sqrt, abs
Expression Depth Range: Depth 1 to 2
Integer Operand Ratio: 0%
Data Fields
Each… See the full description on the dataset page: https://huggingface.co/datasets/saidurga001301/mathmetics-dataset-custom.Maths-CollegeMaths-College
I am releasing a large Mathematics dataset in the instrution format.
This extensive dataset, comprising nearly one million instructions in JSON format, encapsulates a wide array of mathematical disciplines essential for a profound understanding of the subject.
This dataset is very useful to Researchers & Model developers.
Following Fields & sub Fields are covered:
Probability
Statistics
Liner Algebra
Algebra
Group Theory
Topology
Abstract Algebra
Graph Theory
Combinatorics… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/Maths-College.reasoning-math-advanced-1m
🧠 Reasoning Math Advanced 1M
📖 Dataset Summary
Reasoning Math Advanced 1M is a large-scale, synthetic dataset designed to enhance the reasoning capabilities of Large Language Models (LLMs). Comprising 1,000,000 unique samples, this dataset focuses on Math, Logic, and Common Sense reasoning tasks.
A unique feature of this dataset is its adaptive reasoning structure, where the presence of Chain-of-Thought (CoT) reasoning scales with difficulty. All reasoning traces are… See the full description on the dataset page: https://huggingface.co/datasets/naimulislam/reasoning-math-advanced-1m.dfm13-multilingual-math-code-hu
dfm13_wave4_synthetic_hu_math_code
6000 complete conversations; 6000 native assistant targets. All user/tool history and tool definitions are preserved. Gemma native student rendering, thinking disabled. Generated and separately model-reviewed by Gemma4 26B A4B; automated judgments are fallible, not human or native-speaker certification.
Includes unchanged original accepted conversations and narrowly recovered complete keep reviews rejected solely for an empty rationale.… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm13-multilingual-math-code-hu.dfm13-multilingual-math-code-sr
dfm13_wave4_synthetic_sr_math_code
6000 complete conversations; 6000 native assistant targets. All user/tool history and tool definitions are preserved. Gemma native student rendering, thinking disabled. Generated and separately model-reviewed by Gemma4 26B A4B; automated judgments are fallible, not human or native-speaker certification.
Includes unchanged original accepted conversations and narrowly recovered complete keep reviews rejected solely for an empty rationale.… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm13-multilingual-math-code-sr.Maths-Grade-SchoolMaths-Grade-School
I am releasing large Grade School level Mathematics datatset.
This extensive dataset, comprising nearly one million instructions in JSON format, encapsulates a diverse array of topics fundamental to building a strong mathematical foundation.
This dataset is in instruction format so that model developers, researchers etc. can easily use this dataset.
Following Fields & sub Fields are covered:
Calculus
Probability
Algebra
Liner Algebra
Trigonometry
Differential Equations… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/Maths-Grade-School.School-Math-R1-Distil-Chinese-220K从原数据集 BelleGroup/school_math_0.25M 提取指令,然后重新合成回复。
每条数据的格式如下:
{
"id": <<12位nanoid>>,
"prompt": <<提示词>>,
"reasoning": <<模型思考过程>>,
"response": <<模型最终回复>>
}
请注意:本数据集有如下已知缺陷
问题可解性无法保证:这是由于原数据集本身就是纯合成数据集,未经过校验。尽管本数据集已经尽力筛选过滤了一部分,但仍然无法保证余下数据的指令正确性和可解性。
答案未经过校验:所有回答均为合成,且未经过校验。
NOESIS-1M-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54
NOESIS DORA SFT Dataset
Multilingual supervised fine-tuning dataset built for the NOESIS QwQ+DeepSeek-R1 MoE pipeline.
Released as part of the NOESIS Professional Multilingual Dubbing Automation Platform(framework: DHCF-FNO — Deterministic Hybrid Control Framework for Frozen Neural Operators).
Founder: Ilia Bolotnikov
Organization: AMAImedia.com
X (Twitter): @AMAImediacom
LinkedIn: Ilia Bolotnikov
Telegram: @djbionicl
NOESIS version: v14.8-NT89
Build date: 2026-04… See the full description on the dataset page: https://huggingface.co/datasets/SMH-DEV-AI/NOESIS-1M-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54.dfm13-multilingual-math-code-fa
dfm13_wave4_synthetic_fa_math_code
6000 complete conversations; 6000 native assistant targets. All user/tool history and tool definitions are preserved. Gemma native student rendering, thinking disabled. Generated and separately model-reviewed by Gemma4 26B A4B; automated judgments are fallible, not human or native-speaker certification.
Includes unchanged original accepted conversations and narrowly recovered complete keep reviews rejected solely for an empty rationale.… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm13-multilingual-math-code-fa.dfm13-multilingual-math-code-sl
dfm13_wave4_synthetic_sl_math_code
6000 complete conversations; 6000 native assistant targets. All user/tool history and tool definitions are preserved. Gemma native student rendering, thinking disabled. Generated and separately model-reviewed by Gemma4 26B A4B; automated judgments are fallible, not human or native-speaker certification.
Includes unchanged original accepted conversations and narrowly recovered complete keep reviews rejected solely for an empty rationale.… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm13-multilingual-math-code-sl.dfm13-multilingual-math-code-lb
dfm13_wave4_synthetic_lb_math_code
24107 complete conversations; 24107 native assistant targets. All user/tool history and tool definitions are preserved. Gemma native student rendering, thinking disabled. Generated and separately model-reviewed by Gemma4 26B A4B; automated judgments are fallible, not human or native-speaker certification.
Includes unchanged original accepted conversations and narrowly recovered complete keep reviews rejected solely for an empty rationale.… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm13-multilingual-math-code-lb.dfm13-multilingual-math-code-sq
dfm13_wave4_synthetic_sq_math_code
6000 complete conversations; 6000 native assistant targets. All user/tool history and tool definitions are preserved. Gemma native student rendering, thinking disabled. Generated and separately model-reviewed by Gemma4 26B A4B; automated judgments are fallible, not human or native-speaker certification.
Includes unchanged original accepted conversations and narrowly recovered complete keep reviews rejected solely for an empty rationale.… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm13-multilingual-math-code-sq.
