Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hi-todayis-jh /l1-max-100-6000-math-100plus-topup160k-qwen3-1.7b-verl091-dual-20261005 L1-Max 100–6000: at least 100 responses and 160k output tokens per question 112,758 complete, graded responses across all 1,119 questions in six math datasets. Every question has at least 163,840 cumulative full-output tokens, including thinking. Model: hi-todayis-jh/l1-max-qwen3-1.7b-compression-100-6k-from80-mixed50-alpha3e-4-bs32-n8-32k-t1-146102-final at 8dee61ed8fbd53f2783a32c9ba6283b4a3d53563. Two compute nodes, 16 independent TP=1 GPU workers. Sampling: temperature 0.6… See the full description on the dataset page: https://huggingface.co/datasets/hi-todayis-jh/l1-max-100-6000-math-100plus-topup160k-qwen3-1.7b-verl091-dual-20261005.tabulartext-generation100K<n<1M0 likes1k downloads4d agoHugging Face02hi-todayis-jh /l1-exact-100-6000-math-100plus-topup160k-qwen3-1.7b-verl091-dual-20261005 L1-Exact 100–6000: at least 100 responses and 160k output tokens per question 112,516 complete, graded responses across all 1,119 questions in six math datasets. Every question has at least 163,840 cumulative full-output tokens, including thinking. Model: hi-todayis-jh/l1-exact-qwen3-1.7b-compression-100-6k-alpha3e-4-bs32-n8-32k-t1-146102-final at ddd566858009b5f97b8ce6d55dcd5060977ed7b7. Two compute nodes, 16 independent TP=1 GPU workers. Sampling: temperature 0.6, top-p 0.95… See the full description on the dataset page: https://huggingface.co/datasets/hi-todayis-jh/l1-exact-100-6000-math-100plus-topup160k-qwen3-1.7b-verl091-dual-20261005.tabulartext-generation100K<n<1M0 likes913 downloads4d agoHugging Face03argilla /distilabel-math-preference-dpo Dataset Card for "distilabel-math-preference-dpo" More Information needed tabulartext-generation1K<n<10K87 likes696 downloads2y agoHugging Face04MathArena /brokenarxiv-0826 BrokenArXiv August 2026 Homepage and repository Homepage: MathArena Repository: MathArena evaluation code Benchmark description and prompts: August benchmark update Dataset summary This dataset contains 56 plausible but false mathematical statements drawn from arXiv papers submitted in August 2026. Models are asked to prove the statements, and their responses are evaluated for whether they recognize the falsity, acknowledge their inability to… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/brokenarxiv-0826.tabulartext-generationn<1K0 likes661 downloads27d agoHugging Face05OpenDataArena /ODA-Math-460k ODA-Math-460k ODA-Math-460k is a large-scale math reasoning dataset curated from top-performing open mathematics corpora (selected via the OpenDataArena leaderboard) and refined through deduplication, benchmark decontamination, LLM-based filtering, and verifier-backed response distillation.It targets a “learnable but challenging” difficulty band: non-trivial for smaller models yet solvable by stronger reasoning models. 🧠 Dataset Summary Domain: Mathematics… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/ODA-Math-460k.tabularquestion-answering100K<n<1M105 likes568 downloads9mo agoHugging Face06eth-nlped /mathdial Mathdial dataset https://arxiv.org/abs/2305.14536 MathDial: A Dialogue Tutoring Dataset with Rich Pedagogical Properties Grounded in Math Reasoning Problems. MathDial is grounded in math word problems as well as student confusions which provide a challenging testbed for creating faithful and equitable dialogue tutoring models able to reason over complex information. Current models achieve high accuracy in solving such problems but they fail in the task of teaching. Data… See the full description on the dataset page: https://huggingface.co/datasets/eth-nlped/mathdial.tabulartext-generation1K<n<10K18 likes541 downloads2y agoHugging Face07OpenDataArena /MathLake MathLake: A Large-Scale Mathematics Dataset MathLake is a massive collection of 8.3 million mathematical problems aggregated from over 50 open-source datasets. Unlike datasets focused on filtering for the highest quality solutions immediately, MathLake prioritizes query comprehensiveness, serving as a universal "raw ore" for researchers to curate, distill, or annotate further. The dataset provides annotations for Difficulty, Format, and Subject, covering diverse mathematical fields… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MathLake.tabularquestion-answering1M<n<10M21 likes538 downloads6mo agoHugging Face08zbeeb /OpenR1-SFT-Math-20k OpenR1 SFT Math 20k This is the exact prepared SFT population used in the OpenR1 SFT to GRPO token-level study. It contains 20,144 distinct problems: 20,016 training examples and 128 held-out examples. The 128-example training probe is a subset of the training split. The same examples and split order are used for the four Qwen models and DeepSeek-R1-Distill-Qwen-7B in this experiment. Splits and configurations Split Rows Purpose train 20,016 Full SFT… See the full description on the dataset page: https://huggingface.co/datasets/zbeeb/OpenR1-SFT-Math-20k.tabulartext-generation10K<n<100K0 likes437 downloads5d agoHugging Face09ycchen /Crystal-Math-Preview CrystalMath (Preview) 🏆 AI Mathematical Olympiad – Progress Prize 3: MathCorpus Prize Winner CrystalMath was selected as the winning submission for the MathCorpus Prize. A curated dataset of 2,129 competition-level mathematics problems designed for reinforcement learning with verifiable rewards (RLVR) training of frontier reasoning models. CrystalMath is distilled from over 800,000 candidates across 12 public sources through a multi-stage pipeline that enforces high… See the full description on the dataset page: https://huggingface.co/datasets/ycchen/Crystal-Math-Preview.documenttext-generation1K<n<10K14 likes419 downloads2mo agoHugging Face10hi-todayis-jh /compression-math-100plus-topup128k-qwen3-1.7b-verl091-146103-20261004 Compression math responses: every question has at least 160k output tokens 585,003 complete responses from five Qwen3-1.7B models on six datasets. Every model-question pair now has at least 163,840 total generated output tokens (160 × 1,024). This update adds 21,973 complete responses (21,972,891 output tokens) to the previous 563,030-response release. The repository name records the original 128k release; the current data includes the completed 160k supplementation. Metrics… See the full description on the dataset page: https://huggingface.co/datasets/hi-todayis-jh/compression-math-100plus-topup128k-qwen3-1.7b-verl091-146103-20261004.tabulartext-generation100K<n<1M0 likes392 downloads6d agoHugging Face11mihailgribov /olympiad_style_integer_math_problems Olympiad Math Corpus Version: v2.1.1 Release date: 2026-05-03 59,486 synthetically generated olympiad-style math problems with verified integer answers and formal computation graphs. Loading from datasets import load_dataset ds = load_dataset("mihailgribov/olympiad_style_integer_math_problems", split="train") lemma_applicability is stored as list[{lemma, status}] rather than a sparse dict (required for Arrow-based consumers). To convert to a dict for local use:… See the full description on the dataset page: https://huggingface.co/datasets/mihailgribov/olympiad_style_integer_math_problems.documenttext-generation10K<n<100K1 likes390 downloads5mo agoHugging Face12xiuyuz /ample-math AMPLE-Math 5,319 mathematics problems, each with a verified final answer and six references to that same answer. The references differ only in how much of the reasoning they show, which makes them useful for studying what a teacher's reference content contributes during distillation. Problems and original reasoning come from the metadata configuration of OpenThoughts-114k, and keep its Apache-2.0 attribution. A question was kept only if all six references exist, every generated… See the full description on the dataset page: https://huggingface.co/datasets/xiuyuz/ample-math.tabulartext-generation1K<n<10K1 likes282 downloads20d agoHugging Face13mihailgribov /olympiad_style_integer_math_reasoning Olympiad Math Reasoning Traces Version: v1.0.2 Release date: 2026-04-19 64,763 full model reasoning traces for olympiad-style math problems with verified integer answers. This dataset contains only correct and non-truncated traces — every record contains a terminal \boxed{...} answer (within the last 500 characters of the response) that matches the expected integer exactly, and none of the responses hit the model's generation-token cap. Intended for distillation and supervised… See the full description on the dataset page: https://huggingface.co/datasets/mihailgribov/olympiad_style_integer_math_reasoning.tabulartext-generation10K<n<100K0 likes180 downloads6mo agoHugging Face14mst-ai /linalg-bench-math-ai-neurips LinAlg-Bench: A Benchmark Exposing Structural Failure Modes in LLM Linear Algebra — Where Models Stop Computing and Start Hallucinating This dataset is the official release accompanying the MATH-AI 2026 NeurIPS workshop paper, "LinAlg-Bench: A Benchmark Exposing Structural Failure Modes in LLM Linear Algebra — Where Models Stop Computing and Start Hallucinating." The exact 660-problem core evaluated in that paper (9 tasks × 3 matrix sizes, 6,600 model outputs, 1,156… See the full description on the dataset page: https://huggingface.co/datasets/mst-ai/linalg-bench-math-ai-neurips.tabulartext-generation10K<n<100K0 likes157 downloads8d agoHugging Face15blythet /deepseek-v4-pro-math-cot-1k DeepSeek V4 Pro Math CoT 1K A small, high-signal supervised-fine-tuning (SFT) dataset of math reasoning traces. Problems were sampled from a Nemotron math problem set (originally sourced from StackExchange-Math and AoPS), answered by DeepSeek V4 Pro with thinking enabled at high reasoning effort, then independently reviewed by DeepSeek V4 Flash for correctness against the expected answer. Pathological reasoning traces (looping, run-away length, excessive Wait-style backtracking)… See the full description on the dataset page: https://huggingface.co/datasets/blythet/deepseek-v4-pro-math-cot-1k.tabulartext-generation1K<n<10K4 likes150 downloads5mo agoHugging Face16CoffeeGitta /pika-math-generations PIKA MATH Generations Dataset A comprehensive dataset of MATH problem solutions generated by different language models with various sampling parameters. Paper (arXiv) | GitHub Repository Dataset Description This dataset contains code generation results from the MATH Dataset evaluated across multiple models. It was created to support the PIKA (Probe-Informed K-Aware Routing) project, which investigates how LLMs encode their own likelihood of success in their internal… See the full description on the dataset page: https://huggingface.co/datasets/CoffeeGitta/pika-math-generations.tabulartext-generation10K<n<100K0 likes149 downloads7mo agoHugging Face17HachiML /alpaca_jp_math alpaca_jp_math alpaca_jp_mathは、 Stanford Alpacaの手法 mistralai/Mixtral-8x22B-Instruct-v0.1 で作った合成データ(Synthetic data)です。モデルの利用にはDeepinfraを利用しています。 また、"_cleaned"がついたデータセットは以下の手法で精査されています。 pythonの計算結果がきちんと、テキストの計算結果が同等であるか確認 LLM(mistralai/Mixtral-8x22B-Instruct-v0.1)による確認(詳細は下記) code_result, text_resultは小数第三位で四捨五入してあります。 Dataset Details Dataset Description Curated by: HachiMLLanguage(s) (NLP): Japanese License: Apache 2.0 Github: Alpaca-jp… See the full description on the dataset page: https://huggingface.co/datasets/HachiML/alpaca_jp_math.tabulartext-generation10K<n<100K6 likes135 downloads2y agoHugging Face18nlile /math_benchmark_test_saturation LLM Leaderboard Data for Hendrycks MATH Dataset (2022–2024) This dataset aggregates yearly performance (2022–2024) of large language models (LLMs) on the Hendrycks MATH benchmark. It is specifically compiled to explore performance evolution, benchmark saturation, parameter scaling trends, and evaluation metrics of foundation models solving complex math word problems. Original source data: Math Word Problem Solving on MATH (Papers with Code) About Hendrycks' MATH… See the full description on the dataset page: https://huggingface.co/datasets/nlile/math_benchmark_test_saturation.tabularquestion-answeringn<1K0 likes135 downloads2y agoHugging Face19hkust-nlp /dart-math-pool-gsm8k-query-info [!NOTE] This dataset is the synthesis information of queries from the GSM8K training set, such as the numbers of raw/correct samples of each synthesis job. Usually used with dart-math-pool-gsm8k. 🎯 DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving 📝 Paper@arXiv | 🤗 Datasets&Models@HF | 🐱 Code@GitHub 🐦 Thread@X(Twitter) | 🐶 中文博客@知乎 | 📊 Leaderboard@PapersWithCode | 📑 BibTeX Datasets: DART-Math DART-Math datasets are the… See the full description on the dataset page: https://huggingface.co/datasets/hkust-nlp/dart-math-pool-gsm8k-query-info.tabulartext-generation1K<n<10K2 likes124 downloads2y agoHugging Face20amphora /math-intuition-reasoning-traces math-intuition reasoning traces Full chain-of-thought traces from 7 reasoning models on the same 4,020 problems, graded by each problem family's own verifier. Questions come from amphora/math-intuition-20260908-402-easy-10 — 402 arXiv-derived problem families x 10 seeds, easy preset. Every row here refers to an id in that dataset, so prompts and the instance cache can be joined from it. Generation settings Identical for every model, so the traces are directly… See the full description on the dataset page: https://huggingface.co/datasets/amphora/math-intuition-reasoning-traces.tabulartext-generation10K<n<100K1 likes118 downloads29d agoHugging Face21rodriguescarson /adaption-math-worked-solutions-raw NuminaMath Worked Solutions Competition and school maths problems with step-by-step solutions. Rows 3,000 Domain mathematics Format data.parquet, one row per example Licence apache-2.0 Built for supervised fine-tuning (SFT) experiments on Adaption AutoScientist Columns Column Description original_prompt The prompt (user turn) as uploaded. original_completion The target response as uploaded. enhanced_prompt Empty in this… See the full description on the dataset page: https://huggingface.co/datasets/rodriguescarson/adaption-math-worked-solutions-raw.tabulartext-generation1K<n<10K0 likes106 downloads14d agoHugging Face22ChrisMcCormick /math500-cot-deepseek-r1-1.5b MATH-500 CoT completions (DeepSeek-R1-Distill-Qwen-1.5B) Successful chain-of-thought completions for HuggingFaceH4/MATH-500 test problems, generated with deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B via vLLM. Files File Description records.parquet Main dataset: correct completions as token IDs manifest.json Schema, tokenizer, run ids, decoding config problem_index.json unique_id → problem_idx in MATH-500 test subject_max_tokens.json Per-subject completion… See the full description on the dataset page: https://huggingface.co/datasets/ChrisMcCormick/math500-cot-deepseek-r1-1.5b.tabulartext-generation1K<n<10K0 likes99 downloads5mo agoHugging Face23danghoang2005 /qwen25-math-sft-long-15k-v1 Qwen2.5 Math SFT Long 15K v1 A reproducible Long CoT SFT dataset for Qwen2.5-Math-1.5B with max context 3,840 tokens (below 4,096 limit). Source: open-r1/OpenR1-Math-220k Total samples: 15,000 Decontaminated against GSM8K, SVAMP, MATH-500, AIME 2026. tabulartext-generation10K<n<100K0 likes94 downloads19d agoHugging Face24yyuan244 /speculative-reasoning-matheval-4b-sweep MathEval sweep — speculative reasoning on Qwen3-4B Per-sample generations and grading for four arms of a MathEval run, measuring what speculative reasoning costs and saves against a base model that does not speculate. Code and write-up: yurun-yuan/speculative-reasoning — see docs/05-rl-4b.md. The arms All four answer the same 1,547 MathEval problems under a 20,000-token response budget. split model runtime base_plain Qwen/Qwen3-4B plain — no speculation… See the full description on the dataset page: https://huggingface.co/datasets/yyuan244/speculative-reasoning-matheval-4b-sweep.tabulartext-generation10K<n<100K0 likes92 downloads27d agoHugging Face25danghoang2005 /qwen25-math-sft-long-think7k-v1 Qwen2.5 Math Long-CoT thinking 7k 2,500 train and 250 validation rows. Long examples have 4,096–7,168 thinking tokens; complete ChatML samples including final answer and EOS fit in 8,192 tokens. Long source: open-r1/OpenR1-Math-220k at e4e141ec9dea9f8326f4d347be56105859b2bd68 (only math-verified generations). Short replay source: danghoang2005/qwen25-math-sft-long-15k-v1 at 32f0afc54a09c4963ae2478cf93871025abda901. Tokenizer: Qwen/Qwen2.5-Math-1.5B at… See the full description on the dataset page: https://huggingface.co/datasets/danghoang2005/qwen25-math-sft-long-think7k-v1.tabulartext-generation1K<n<10K0 likes91 downloads16d agoHugging Face26hkust-nlp /dart-math-pool-math-query-info [!NOTE] This dataset is the synthesis information of queries from the MATH training set, such as the numbers of raw/correct samples of each synthesis job. Usually used with dart-math-pool-math. 🎯 DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving 📝 Paper@arXiv | 🤗 Datasets&Models@HF | 🐱 Code@GitHub 🐦 Thread@X(Twitter) | 🐶 中文博客@知乎 | 📊 Leaderboard@PapersWithCode | 📑 BibTeX Datasets: DART-Math DART-Math datasets are the state-of-the-art… See the full description on the dataset page: https://huggingface.co/datasets/hkust-nlp/dart-math-pool-math-query-info.tabulartext-generation1K<n<10K0 likes87 downloads2y agoHugging Face27RyanYr /dapo-math-17k-qwen3-1.7b-base-n8 DAPO-Math-17k sampled with Qwen3-1.7B-Base, n=8 17398 problems from the RL training set, each sampled 8 times and scored with the reward function the RL runs themselves used. The point of this dataset is to be comparable with what the RL runs actually saw, so every sampling knob is taken from the live training config or from the default that config falls through to. Two of them are not in the config file at all and would be wrong if guessed: top_k = -1 and min_tokens = 1.… See the full description on the dataset page: https://huggingface.co/datasets/RyanYr/dapo-math-17k-qwen3-1.7b-base-n8.tabulartext-generation10K<n<100K0 likes87 downloads25d agoHugging Face28ar0cket1 /hintedselfteacher-nemotron-math-v2-AoPS hintedselfteacher-nemotron-math-v2-AoPS This dataset contains a training-ready hinted self-teacher split derived from the AoPS split of nvidia/Nemotron-Math-v2. The source problems were filtered to the AoPS split with the medium/notool solve rate between 2 and 6. Hints were generated with GPT-5.5 medium using an h17_nt hint-generation prompt. This hint type was close to the best hint type found after doing hint mutations, based on qualitative analysis of token-level hinted… See the full description on the dataset page: https://huggingface.co/datasets/ar0cket1/hintedselfteacher-nemotron-math-v2-AoPS.tabulartext-generation10K<n<100K1 likes82 downloads4mo agoHugging Face29rodriguescarson /adaption-math-worked-solutions-mixtral-raw NuminaMath Worked Solutions Competition and school maths problems with step-by-step solutions. Rows 3,000 Domain mathematics Format data.parquet, one row per example Licence apache-2.0 Built for supervised fine-tuning (SFT) experiments on Adaption AutoScientist Columns Column Description original_prompt The prompt (user turn) as uploaded. original_completion The target response as uploaded. enhanced_prompt Empty in this… See the full description on the dataset page: https://huggingface.co/datasets/rodriguescarson/adaption-math-worked-solutions-mixtral-raw.tabulartext-generation1K<n<10K0 likes80 downloads14d agoHugging Face30zhiqix /PUM-MATH PUM Prefix Pair Dataset This dataset contains pairwise prefix preference examples for gain-based evaluation of LLM reasoning. Each example compares two partial reasoning prefixes for the same math problem and records which prefix is preferred according to outcome-grounded prefix utility. The dataset is associated with the paper: From Correctness to Utility: Gain-Based Prefix Evaluation for LLM Reasoning Dataset Description Reasoning prefixes can strongly affect… See the full description on the dataset page: https://huggingface.co/datasets/zhiqix/PUM-MATH.tabulartext-generation100K<n<1M12 likes75 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.