Team Ai
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01test-time-compute /aime_2025 AIME 2025 - Unified Test-Time Scaling Format This is the AIME (American Invitational Mathematics Examination) 2025 dataset in a unified format for test-time scaling experiments. Dataset Description Source: MathArena/aime_2025 Size: 30 competition-level mathematics problems Format: Unified TTS format (question, answer, metadata) Dataset Structure Fields question (string): The mathematical problem statement answer (string): The numerical answer… See the full description on the dataset page: https://huggingface.co/datasets/test-time-compute/aime_2025.textquestion-answeringn<1K0 likes366 downloads1y agoHugging Face02test-time-compute /test_MATHtextn<1K0 likes85 downloads11mo agoHugging Face03test-time-compute /test_olympiadbenchtextn<1K0 likes70 downloads9mo agoHugging Face04test-time-compute /game-of-24 Game of 24 Dataset Dataset Description The Game of 24 is a mathematical reasoning puzzle where players must use four numbers and basic arithmetic operations (+, -, *, /) to obtain the result 24. Each number must be used exactly once. This dataset contains 1,361 unique Game of 24 puzzles ranked by difficulty based on human performance from Amazon Mechanical Turk studies. Example Input: 4 5 6 10 Output: (5 * (10 - 4)) - 6 = 24 Step-by-step solution: 10 - 4 = 6… See the full description on the dataset page: https://huggingface.co/datasets/test-time-compute/game-of-24.tabularquestion-answering1K<n<10K3 likes49 downloads1y agoHugging Face05nkh /test-time-compute-for-tabular-foundation-models-results Test-Time Compute for Tabular Foundation Models: TabArena Results Per-cell TabArena test errors for the 12 methods ("arms") in the main comparison of Test-Time Compute for Tabular Foundation Models: Mechanisms, Gains, and Limits. Every arm covers all 51 TabArena datasets (816 cells), so these files are enough to recompute the paper's Elo ratings and paired comparisons. Code: GitHub · Paper: arXiv:2610.12005 Files file content data/results.parquet config… See the full description on the dataset page: https://huggingface.co/datasets/nkh/test-time-compute-for-tabular-foundation-models-results.tabulartabular-classification1K<n<10K0 likes33 downloads2d agoHugging Face06test-time-compute /test_gpqa_diamondtextn<1K0 likes31 downloads9mo agoHugging Face07test-time-compute /test_gsm8ktext1K<n<10K0 likes25 downloads1y agoHugging Face08test-time-compute /test_gaokao2023entextn<1K0 likes21 downloads9mo agoHugging Face09test-time-compute /cudabenchtextn<1K0 likes17 downloads10mo agoHugging Face10fineset-io /test-time-compute-papers Test-Time Compute & Reasoning Scaling Papers — FineSet A research-paper dataset on Test-Time Compute & Reasoning Scaling Papers, assembled, deduplicated, and quality-scored by FineSet from arXiv and Semantic Scholar. 📸 This is a dated snapshot — generated 2026-06-19. It is not auto-updated. Research on Test-Time Compute & Reasoning Scaling Papers moves fast — new papers land on arXiv every week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/test-time-compute-papers.tabulartext-classificationn<1K0 likes17 downloads4mo agoHugging Face11test-time-compute /test_Proofnettextn<1K0 likes15 downloads1y agoHugging Face12test-time-compute /test_minerva_mathtextn<1K0 likes15 downloads9mo agoHugging Face13test-time-compute /test_svamptext1K<n<10K0 likes9 downloads1y agoHugging Face14test-time-compute /test_amc23textn<1K0 likes9 downloads8mo agoHugging Face15test-time-compute /test_aime24textn<1K0 likes5 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.