datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
chatbot-arena-elo
LMSYS Chatbot Arena ELO Scores
This dataset is a datasets-friendly version of Chatbot Arena ELO scores,
updated daily from the leaderboard API at
https://huggingface.co/spaces/lmarena-ai/chatbot-arena-leaderboard.
Updated: 20250717
Loading Data
from datasets import load_dataset
dataset = load_dataset("mathewhe/chatbot-arena-elo", split="train")
The main branch of this dataset will always be updated to the latest ELO and
leaderboard version. If you need a fixed dataset… See the full description on the dataset page: https://huggingface.co/datasets/mathewhe/chatbot-arena-elo.clef2025-bioasq-task13Bmath-graph
Math-Graph
Math-Graph is the dataset behind TheoremGraph, a unified, statement-level dependency
graph spanning both informal and formal mathematics. On the informal side it parses millions of
theorem-like environments from mathematics arXiv and recovers directed dependency edges within and
across papers; on the formal side it releases LeanGraph, an elaborator-level extraction of typed
declaration dependencies across 25 Lean 4 projects. The two graphs are bridged into one… See the full description on the dataset page: https://huggingface.co/datasets/uw-math-ai/math-graph.sat_multiple_choice_math_may_23This is the set of math SAT questions from the May 2023 SAT, taken from here: https://www.mcelroytutoring.com/lower.php?url=44-official-sat-pdfs-and-82-official-act-pdf-practice-tests-free.
Questions that included images were not included but all other math questions, including those that have tables were included.
Maths_competition_questionslinalg-bench-math-ai-neurips
LinAlg-Bench: A Benchmark Exposing Structural Failure Modes in LLM Linear Algebra — Where Models Stop Computing and Start Hallucinating
This dataset is the official release accompanying the MATH-AI 2026 NeurIPS workshop paper, "LinAlg-Bench: A Benchmark Exposing Structural Failure Modes in LLM Linear Algebra — Where Models Stop Computing and Start Hallucinating." The exact 660-problem core evaluated in that paper (9 tasks × 3 matrix sizes, 6,600 model outputs, 1,156… See the full description on the dataset page: https://huggingface.co/datasets/mst-ai/linalg-bench-math-ai-neurips.theorem-matching
TheoremGraph Matching
Formal–informal theorem matches from the TheoremGraph paper. Each row pairs a
Lean declaration with the most similar natural-language statement from arXiv,
found by cosine similarity over slogan embeddings, and labeled by an LLM judge
as exact, inexact, or wrong (the first two count as a match).
The file contains every candidate pair at cosine similarity 0.80 and above:
100,831 pairs. Our primary judge, GPT-5.4, labels 47,952 of them as matches; a
second… See the full description on the dataset page: https://huggingface.co/datasets/uw-math-ai/theorem-matching.Mathverse_VLMEvalKitMathVista_V2gsm8k-translated
Multilingual GSM8K Translations
This dataset contains machine-translated versions of GSM8K in these languages:
French (fr)
German (de)
Hindi (hi)
Dataset Structure
For each language, we provide the original GSM8K train and test splits:
train: 7,473 samples
test: 1,319 samples
Each sample consists of a question and an answer.
The question describes a grade-school-level math word problem that requires multi-step mathematical reasoning. The answer contains a… See the full description on the dataset page: https://huggingface.co/datasets/math-across-languages/gsm8k-translated.Fast-Math-R1-SFTThis repository contains the First stage SFT dataset as presented in the paper A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning.
This dataset is used for the intensive Supervised Fine-Tuning (SFT) phase, crucial for pushing the model's mathematical accuracy.
Project GitHub Repository: https://github.com/RabotniKuma/Kaggle-AIMO-Progress-Prize-2-9th-Place-Solution
Dataset Construction
This dataset was… See the full description on the dataset page: https://huggingface.co/datasets/RabotniKuma/Fast-Math-R1-SFT.LLM-Hallucination-Detection-complex-mathematics
AIME Hallucination Detection Dataset
This dataset is created for detecting hallucinations in Large Language Models (LLMs), particularly focusing on complex mathematical problems. It can be used for tasks like model evaluation, fine-tuning, and research.
Dataset Details
Name: AIME Hallucination Detection Dataset
Format: CSV
Size: (add size, e.g., 10MB)
Files Included:
AIME-hallucination-detection-dataset.csv: Contains the dataset.
Content Description… See the full description on the dataset page: https://huggingface.co/datasets/tourist800/LLM-Hallucination-Detection-complex-mathematics.math-graph
Math-Graph
Math-Graph is the dataset behind TheoremGraph, a unified, statement-level dependency
graph spanning both informal and formal mathematics. On the informal side it parses millions of
theorem-like environments from mathematics arXiv and recovers directed dependency edges within and
across papers; on the formal side it releases LeanGraph, an elaborator-level extraction of typed
declaration dependencies across 25 Lean 4 projects. The two graphs are bridged into one… See the full description on the dataset page: https://huggingface.co/datasets/jessteru/math-graph.Nemotron-RL-Math-v2-prompt-only
Nemotron-RL-Math-v2-prompt-only
Prompt-only extraction from nvidia/Nemotron-RL-Math-v2.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.
null_or_empty_rows.md: row indexes where prompt extraction produced a… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Math-v2-prompt-only.Par-Four-Fineweb-Edu-Fortified-Chemistry-Physics-Astronomy-Math-ReasonExtracted Chemistry Physics Asronomy Math and Logic portions from the original.
Script used for the extraction:
https://huggingface.co/datasets/Josephgflowers/Par-Four-Fineweb-Edu-Fortified-Chemistry-Physics-Astronomy-Math-Reason/resolve/main/find-science-fine.py
llm-math-evaluation-dataset
LLM Math Response Evaluation Dataset
Dataset Summary
A human-annotated dataset of 150 AI-generated math responses
evaluated across GPT-4o, Claude, and Gemini. Each response is
scored on Correctness, Reasoning, and Clarity using a structured
rubric, with written justification for every score.
Supported Tasks
LLM evaluation and benchmarking
Math reasoning quality assessment
Error type classification in AI responses
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/nbvbharath-1729/llm-math-evaluation-dataset.easy_turkish_math_reasoning
Easy Turkish Math Reasoning
Dataset Summary
The Easy Turkish Math Reasoning dataset is the first phase of a multi-stage curriculum learning pipeline designed to enhance the reasoning abilities of compact language models. This dataset focuses on elementary-level arithmetic and logic problems in Turkish, serving as a warm-up stage for supervised fine-tuning (SFT).
Use Case
Primarily used for:
Bootstrapping reasoning ability in Turkish for compact LLMs.
Phase 1… See the full description on the dataset page: https://huggingface.co/datasets/erayalp/easy_turkish_math_reasoning.Fast-Math-R1-GRPOThis repository contains the second-stage GRPO dataset for the paper A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning.
This dataset is crucial for the second stage of the training recipe, aiming to improve token efficiency while preserving peak mathematical reasoning performance in Large Language Models (LLMs) through Reinforcement Learning from online inference (GRPO).
We extracted the answers from the 2nd stage SFT… See the full description on the dataset page: https://huggingface.co/datasets/RabotniKuma/Fast-Math-R1-GRPO.Math_CoT_Arabic_English_Reasoning
Math CoT Arabic English Dataset
A high-quality, bilingual (English & Arabic) dataset for Chain-of-Thought (COT) reasoning in mathematics and related disciplines, developed by Miscovery AI.
Overview
Math-COT is a unique dataset designed to facilitate and benchmark the development of chain-of-thought reasoning capabilities in language models across mathematical domains. With meticulously crafted examples, explicit reasoning steps, and bilingual support, this dataset offers… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/Math_CoT_Arabic_English_Reasoning.mathoverflow_text_arxiv_labelsDownloaded from https://archive.org/download/stackexchange
Used TexSoup to replace all text in math environments with [UNK]. For instance the text:
"The integral $\int_a^b f(x) \textrm{ d}x$ is easy to evaluate if..."
was replaced with
"The integral [UNK] is easy to evaluate if..."
Note: There is still some "ascii math". For instance, people sometimes write things like f: X --> Y. This is retained.
Concatenated title and body.
Some of these are "answer" posts rather than "question" posts.… See the full description on the dataset page: https://huggingface.co/datasets/stevengubkin/mathoverflow_text_arxiv_labels.bengali-math-cotBangla-Math
Overview
The Bangla-Math Dataset is a valuable resource that addresses the critical need for Bangla-language mathematical problem-solving datasets. Currently, there are no publicly available datasets for math problems in Bangla, making this dataset a unique and valuable contribution to the fields of natural language processing (NLP). This dataset bridges the gap by enabling AI models to understand and work effectively with Bengali math content, thereby supporting advancements in… See the full description on the dataset page: https://huggingface.co/datasets/kawchar85/Bangla-Math.orz_math_difficulty
Difficulty Estimation on Open Reasoner Zero
We annotate the entire Open Reasoner Zero dataset with a difficulty score based on the performance of the Qwen 2.5-MATH-7B model. This provides an adaptive signal for curriculum construction.
Open Reasoner Zero is a curated a dataset of 57,000 reasoning-intensive problems used to train and evaluate reinforcement learning-based methods for large language models.
Difficulty Scoring Method
Difficulty scores are estimated using… See the full description on the dataset page: https://huggingface.co/datasets/lime-nlp/orz_math_difficulty.medium_turkish_math_reasoning
Dataset Summary
The Medium Turkish Math Reasoning dataset is Phase 2 of a curriculum learning pipeline to teach compact models multi-step reasoning in Turkish. It includes moderately difficult math problems involving multiple reasoning steps, such as two-part arithmetic, comparisons, and logical reasoning.
Use Case
This dataset is ideal for:
Continuing SFT after foundational training with simpler problems.
Bridging the gap between basic arithmetic and complex GSM8K-style… See the full description on the dataset page: https://huggingface.co/datasets/erayalp/medium_turkish_math_reasoning.math_onetmath-intuition-20260906-403-easy-30
math-intuition-20260906-403-easy-30
12,090 synthetic mathematics problems drawn from 403 problem families, each family
derived from a distinct arXiv paper. Every problem is generated answer-first, so the
answer is known by construction and is checked by the family's own verify() before
the row is written. No row in this file is ungraded.
This is the easy slice: 30 instances per family at each family's easiest difficulty
preset. It is not the hard benchmark — see Difficulty… See the full description on the dataset page: https://huggingface.co/datasets/amphora/math-intuition-20260906-403-easy-30.MathWizard-mathword-problem-dataset-with-grade-section
Dataset Card for MathWizard-mathword-problem-dataset-with-grade-section
This dataset consists of approximately 4,000 Elementary Math Word Problems (MWPs) generated using Large Language Models (LLMs) and comprehensively annotated for errors by humans and LLM judges. It is designed to support the generation and evaluation of high-quality, grade-appropriate math problems.
Dataset Details
Dataset Description
Curated by: [Nimesh Ariyarathne, Harshani Bandara… See the full description on the dataset page: https://huggingface.co/datasets/MathWizards/MathWizard-mathword-problem-dataset-with-grade-section.dart_math_banglaThe dataset contains math problems in bangla. hkust-nlp/dart-math-uniform is translated using facebook/nllb-200-3.3B. To achive better performance english sentences are splitted and then fed into the translation model.
math-intuition-20260906-403-demo-10
math-intuition-20260906-403-demo-10
3,936 mathematics problems drawn from 403 problem families, each derived from a
distinct arXiv paper. Every problem is generated answer-first, so the answer is known by
construction and is checked by the family's own verify() before the row is written.
No row in this file is ungraded.
This is the demo rung — read this before using it
Each family exposes a four-rung ladder: demo, easy, medium, hard. This file samples
demo, which… See the full description on the dataset page: https://huggingface.co/datasets/amphora/math-intuition-20260906-403-demo-10.MathDisAM
MathDisAm: Towards Math Query Disambiguation
Welcome to the official repository for the paper MathDisAm: Towards Math Query Disambiguation. This repository contains the dataset, query generation pipelines, and baseline classifier models designed to distinguish between ambiguous and unambiguous mathematical queries.
Dataset
The data directory contains the core datasets used for training and evaluating the disambiguation models:
MathDisAM_train.tsv: The training… See the full description on the dataset page: https://huggingface.co/datasets/AIIRLab/MathDisAM.
