datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
chatbot-arena-elo
LMSYS Chatbot Arena ELO Scores
This dataset is a datasets-friendly version of Chatbot Arena ELO scores,
updated daily from the leaderboard API at
https://huggingface.co/spaces/lmarena-ai/chatbot-arena-leaderboard.
Updated: 20250717
Loading Data
from datasets import load_dataset
dataset = load_dataset("mathewhe/chatbot-arena-elo", split="train")
The main branch of this dataset will always be updated to the latest ELO and
leaderboard version. If you need a fixed dataset… See the full description on the dataset page: https://huggingface.co/datasets/mathewhe/chatbot-arena-elo.math-graph
Math-Graph
Math-Graph is the dataset behind TheoremGraph, a unified, statement-level dependency
graph spanning both informal and formal mathematics. On the informal side it parses millions of
theorem-like environments from mathematics arXiv and recovers directed dependency edges within and
across papers; on the formal side it releases LeanGraph, an elaborator-level extraction of typed
declaration dependencies across 25 Lean 4 projects. The two graphs are bridged into one… See the full description on the dataset page: https://huggingface.co/datasets/uw-math-ai/math-graph.sat_multiple_choice_math_may_23This is the set of math SAT questions from the May 2023 SAT, taken from here: https://www.mcelroytutoring.com/lower.php?url=44-official-sat-pdfs-and-82-official-act-pdf-practice-tests-free.
Questions that included images were not included but all other math questions, including those that have tables were included.
Mathverse_VLMEvalKitlinalg-bench-math-ai-neurips
LinAlg-Bench: A Benchmark Exposing Structural Failure Modes in LLM Linear Algebra — Where Models Stop Computing and Start Hallucinating
This dataset is the official release accompanying the MATH-AI 2026 NeurIPS workshop paper, "LinAlg-Bench: A Benchmark Exposing Structural Failure Modes in LLM Linear Algebra — Where Models Stop Computing and Start Hallucinating." The exact 660-problem core evaluated in that paper (9 tasks × 3 matrix sizes, 6,600 model outputs, 1,156… See the full description on the dataset page: https://huggingface.co/datasets/mst-ai/linalg-bench-math-ai-neurips.LLM-Hallucination-Detection-complex-mathematics
AIME Hallucination Detection Dataset
This dataset is created for detecting hallucinations in Large Language Models (LLMs), particularly focusing on complex mathematical problems. It can be used for tasks like model evaluation, fine-tuning, and research.
Dataset Details
Name: AIME Hallucination Detection Dataset
Format: CSV
Size: (add size, e.g., 10MB)
Files Included:
AIME-hallucination-detection-dataset.csv: Contains the dataset.
Content Description… See the full description on the dataset page: https://huggingface.co/datasets/tourist800/LLM-Hallucination-Detection-complex-mathematics.Nemotron-RL-Math-v2-prompt-only
Nemotron-RL-Math-v2-prompt-only
Prompt-only extraction from nvidia/Nemotron-RL-Math-v2.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.
null_or_empty_rows.md: row indexes where prompt extraction produced a… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Math-v2-prompt-only.math-graph
Math-Graph
Math-Graph is the dataset behind TheoremGraph, a unified, statement-level dependency
graph spanning both informal and formal mathematics. On the informal side it parses millions of
theorem-like environments from mathematics arXiv and recovers directed dependency edges within and
across papers; on the formal side it releases LeanGraph, an elaborator-level extraction of typed
declaration dependencies across 25 Lean 4 projects. The two graphs are bridged into one… See the full description on the dataset page: https://huggingface.co/datasets/jessteru/math-graph.llm-math-evaluation-dataset
LLM Math Response Evaluation Dataset
Dataset Summary
A human-annotated dataset of 150 AI-generated math responses
evaluated across GPT-4o, Claude, and Gemini. Each response is
scored on Correctness, Reasoning, and Clarity using a structured
rubric, with written justification for every score.
Supported Tasks
LLM evaluation and benchmarking
Math reasoning quality assessment
Error type classification in AI responses
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/nbvbharath-1729/llm-math-evaluation-dataset.Par-Four-Fineweb-Edu-Fortified-Chemistry-Physics-Astronomy-Math-ReasonExtracted Chemistry Physics Asronomy Math and Logic portions from the original.
Script used for the extraction:
https://huggingface.co/datasets/Josephgflowers/Par-Four-Fineweb-Edu-Fortified-Chemistry-Physics-Astronomy-Math-Reason/resolve/main/find-science-fine.py
Bangla-Math
Overview
The Bangla-Math Dataset is a valuable resource that addresses the critical need for Bangla-language mathematical problem-solving datasets. Currently, there are no publicly available datasets for math problems in Bangla, making this dataset a unique and valuable contribution to the fields of natural language processing (NLP). This dataset bridges the gap by enabling AI models to understand and work effectively with Bengali math content, thereby supporting advancements in… See the full description on the dataset page: https://huggingface.co/datasets/kawchar85/Bangla-Math.Math_CoT_Arabic_English_Reasoning
Math CoT Arabic English Dataset
A high-quality, bilingual (English & Arabic) dataset for Chain-of-Thought (COT) reasoning in mathematics and related disciplines, developed by Miscovery AI.
Overview
Math-COT is a unique dataset designed to facilitate and benchmark the development of chain-of-thought reasoning capabilities in language models across mathematical domains. With meticulously crafted examples, explicit reasoning steps, and bilingual support, this dataset offers… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/Math_CoT_Arabic_English_Reasoning.orz_math_difficulty
Difficulty Estimation on Open Reasoner Zero
We annotate the entire Open Reasoner Zero dataset with a difficulty score based on the performance of the Qwen 2.5-MATH-7B model. This provides an adaptive signal for curriculum construction.
Open Reasoner Zero is a curated a dataset of 57,000 reasoning-intensive problems used to train and evaluate reinforcement learning-based methods for large language models.
Difficulty Scoring Method
Difficulty scores are estimated using… See the full description on the dataset page: https://huggingface.co/datasets/lime-nlp/orz_math_difficulty.mathoverflow_text_arxiv_labelsDownloaded from https://archive.org/download/stackexchange
Used TexSoup to replace all text in math environments with [UNK]. For instance the text:
"The integral $\int_a^b f(x) \textrm{ d}x$ is easy to evaluate if..."
was replaced with
"The integral [UNK] is easy to evaluate if..."
Note: There is still some "ascii math". For instance, people sometimes write things like f: X --> Y. This is retained.
Concatenated title and body.
Some of these are "answer" posts rather than "question" posts.… See the full description on the dataset page: https://huggingface.co/datasets/stevengubkin/mathoverflow_text_arxiv_labels.math-intuition-20260906-403-easy-30
math-intuition-20260906-403-easy-30
12,090 synthetic mathematics problems drawn from 403 problem families, each family
derived from a distinct arXiv paper. Every problem is generated answer-first, so the
answer is known by construction and is checked by the family's own verify() before
the row is written. No row in this file is ungraded.
This is the easy slice: 30 instances per family at each family's easiest difficulty
preset. It is not the hard benchmark — see Difficulty… See the full description on the dataset page: https://huggingface.co/datasets/amphora/math-intuition-20260906-403-easy-30.MATH_Difficulty
Difficulty Estimation on MATH
We annotate the entire MATH dataset with a difficulty score based on the performance of the Qwen 2.5-MATH-7B model. This provides an adaptive signal for curriculum construction and model evaluation.
The Mathematics Aptitude Test of Heuristics (MATH) dataset consists of problems from mathematics competitions, including the AMC 10, AMC 12, AIME, and more. Each problem in MATH has a full step-by-step solution, which can be used to teach models to generate… See the full description on the dataset page: https://huggingface.co/datasets/lime-nlp/MATH_Difficulty.Nemotron-Math-Proofs-v2-prompt-only
Nemotron-Math-Proofs-v2-prompt-only
Prompt-only extraction from nvidia/Nemotron-Math-Proofs-v2.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.
null_or_empty_rows.md: row indexes where prompt extraction… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-Math-Proofs-v2-prompt-only.math-intuition-20260906-403-demo-10
math-intuition-20260906-403-demo-10
3,936 mathematics problems drawn from 403 problem families, each derived from a
distinct arXiv paper. Every problem is generated answer-first, so the answer is known by
construction and is checked by the family's own verify() before the row is written.
No row in this file is ungraded.
This is the demo rung — read this before using it
Each family exposes a four-rung ladder: demo, easy, medium, hard. This file samples
demo, which… See the full description on the dataset page: https://huggingface.co/datasets/amphora/math-intuition-20260906-403-demo-10.DDCF-fullcorpus-mathmath_onetNemotron-Math-Proofs-v1-prompt-only
Nemotron-Math-Proofs-v1-prompt-only
Prompt-only extraction from nvidia/Nemotron-Math-Proofs-v1.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.
null_or_empty_rows.md: row indexes where prompt extraction… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-Math-Proofs-v1-prompt-only.MathVista-tsvng-jss1-math-lesson-plans
Nigerian JSS1 Mathematics Lesson Plans
Six mathematics lessons that a practising teacher planned and taught to a Junior Secondary School 1
(JSS1) class in Nigeria in the 2025/26 school year, following the national NERDC curriculum.
Each lesson is here as the teacher wrote it, as a cleaned transcription, as structured JSON, and
as a declared XML document valid against a RELAX NG grammar. The three lessons that ended in a
class exercise also carry the learners' de-identified… See the full description on the dataset page: https://huggingface.co/datasets/tosinamuda/ng-jss1-math-lesson-plans.ng-jss1-math-alignment-ratings
Nigerian JSS1 Mathematics Alignment Ratings
What the LPCG framework generated from six JSS1 mathematics lessons, and how four blinded raters and
a model judge rated it: the inputs as frozen, every run with its record, the documents the raters
received and returned, and the ratings.
This is one of four datasets released with the LPCG framework from the MSc study Design and Evaluation of a Lesson-Plan-Driven Framework for Curriculum-Constrained Generation and Personalisation of… See the full description on the dataset page: https://huggingface.co/datasets/tosinamuda/ng-jss1-math-alignment-ratings.class-zbmath-identifier
class-zbmath-identifier
This is a proxy dataset to model semantic similarity of short mathematical texts from zbMath.This proxy only contains zbMath.org identifiers (aka an) instead of full titles / abstracts.
Columns
an_a (string): zbMath.org identifier of work a
MSC_a (string): primary MSC5 of work a
MSC2_a (list(string)): secondary MSC5s of work a
an_b (string): zbMath.org identifier of work b
MSC_b (string): primary MSC5 of work b
MSC2_b (list(string)): secondary… See the full description on the dataset page: https://huggingface.co/datasets/math-similarity/class-zbmath-identifier.ktt-math-tutor-data
KTT Math Tutor — Data
Data artefacts for the AIMS KTT Hackathon Tier-3 submission
S2.T3.1 AI Math Tutor for Early Learners. Source code:
https://github.com/DrUkachi/ktt-math-tutor.
Contents
T3.1_Math_Tutor/
Core curriculum + seeds.
curriculum.json — 80 items × 5 sub-skills (counting, number
sense, addition, subtraction, word problem) with EN / FR / KIN
stems, difficulty 1–10, age bands 5–6 / 6–7 / 7–8 / 8–9, visual
asset keys, expected integer answer.… See the full description on the dataset page: https://huggingface.co/datasets/DrUkachi/ktt-math-tutor-data.dumy-zno-ukrainian-math-history-geo-r1-o1
DUMY («Думи»): Ukrainian Multidomain Reasoning Dataset (Part 1: ZNO/NMT tasks with DeepSeek R1 and OpenAI o1 answers)
DUMY is an open benchmark and dataset designed for training, distillation, and evaluation of language models focused on Ukrainian reasoning tasks.
The word “Dumy” comes from Taras Shevchenko’s famous poem and literally means “thoughts” in Ukrainian:
Думи мої, думи мої,
Лихо мені з вами!
Нащо стали на папері
Сумними рядами?..
Work in progress. Stay tuned.… See the full description on the dataset page: https://huggingface.co/datasets/NLPForUA/dumy-zno-ukrainian-math-history-geo-r1-o1.math_hard_problemsPar-Four-Fineweb-Edu-Fortified-MathFiltered for a math focus.
Script used to create the dataset:
https://huggingface.co/datasets/Josephgflowers/Par-Four-Fineweb-Edu-Fortified-Math/resolve/main/find-math-fine.py
MathVerse_TH
