test-time-compute
aime_2025
AIME 2025 - Unified Test-Time Scaling Format
This is the AIME (American Invitational Mathematics Examination) 2025 dataset in a unified format for test-time scaling experiments.
Dataset Description
Source: MathArena/aime_2025
Size: 30 competition-level mathematics problems
Format: Unified TTS format (question, answer, metadata)
Dataset Structure
Fields
question (string): The mathematical problem statement
answer (string): The numerical answer… See the full description on the dataset page: https://huggingface.co/datasets/test-time-compute/aime_2025.test_MATHtest_olympiadbenchgame-of-24
Game of 24 Dataset
Dataset Description
The Game of 24 is a mathematical reasoning puzzle where players must use four numbers and basic arithmetic operations (+, -, *, /) to obtain the result 24. Each number must be used exactly once.
This dataset contains 1,361 unique Game of 24 puzzles ranked by difficulty based on human performance from Amazon Mechanical Turk studies.
Example
Input: 4 5 6 10
Output: (5 * (10 - 4)) - 6 = 24
Step-by-step solution:
10 - 4 = 6… See the full description on the dataset page: https://huggingface.co/datasets/test-time-compute/game-of-24.test-time-compute-for-tabular-foundation-models-results
Test-Time Compute for Tabular Foundation Models: TabArena Results
Per-cell TabArena test errors for the 12 methods ("arms") in the main comparison of
Test-Time Compute for Tabular Foundation Models: Mechanisms, Gains, and Limits.
Every arm covers all 51 TabArena datasets (816 cells), so these files are enough to recompute
the paper's Elo ratings and paired comparisons.
Code: GitHub · Paper: arXiv:2610.12005
Files
file
content
data/results.parquet
config… See the full description on the dataset page: https://huggingface.co/datasets/nkh/test-time-compute-for-tabular-foundation-models-results.test_gpqa_diamond
