datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aimo-validation-aime
Dataset Card for AIMO Validation AIME
All 90 problems come from AIME 22, AIME 23, and AIME 24, and have been extracted directly from the AOPS wiki page https://artofproblemsolving.com/wiki/index.php/AIME_Problems_and_Solutions
This dataset serves as an internal validation set during our participation in the AIMO progress prize competition. Using data after 2021 is to avoid potential overlap with the MATH training set.
Here are the different columns in the dataset:
problem: the… See the full description on the dataset page: https://huggingface.co/datasets/AI-MO/aimo-validation-aime.esb-datasets-earnings22-validation-tiny-filteredA filtered (<=30s duration) slice (512 samples) of the Earnings22 dataset.
def add_duration(sample):
y, sr = sample['audio']["array"], sample['audio']["sampling_rate"]
sample['duration_ms']=librosa.get_duration(y=y, sr=sr) * 1000
return sample
tedlium = load_dataset("esb/datasets", "earnings22", split='validation', trust_remote_code=True)
# compute duration to filter
tedlium = tedlium.map(add_duration)
tedlium = tedlium.select(range(512))
# Whisper max supported duration
tedlium… See the full description on the dataset page: https://huggingface.co/datasets/D4nt3/esb-datasets-earnings22-validation-tiny-filtered.aimo-validation-amc
Dataset Card for AIMO Validation AMC
All 83 come from AMC12 2022, AMC12 2023, and have been extracted from the AOPS wiki page https://artofproblemsolving.com/wiki/index.php/AMC_12_Problems_and_Solutions
This dataset serves as an internal validation set during our participation in the AIMO progress prize competition. Using data after 2021 is to avoid potential overlap with the MATH training set.
Here are the different columns in the dataset:
problem: the modified problem statement… See the full description on the dataset page: https://huggingface.co/datasets/AI-MO/aimo-validation-amc.QuantiPhy-validation
QuantiPhy (Validation Set)
Dataset Summary
QuantiPhy is a benchmark for evaluating whether vision–language models (VLMs) can perform quantitative physical inference from visual evidence, rather than producing plausible but ungrounded numerical guesses.
This repository contains the official validation set of QuantiPhy, released to support model development, ablation studies, and preliminary evaluation.The validation set represents approximately 4% of the full… See the full description on the dataset page: https://huggingface.co/datasets/PaulineLi/QuantiPhy-validation.gaia_validationgaia-validation-sampled_50ii-agent_gaia-benchmark_validationCOCO_captions_validation
Dataset Card for "COCO_captions_validation"
More Information needed
VQAv2_validation
Dataset Card for "VQAv2_validation"
More Information needed
aimo-validation-math-level-5
Dataset Card for AIMO Validation MATH Level 5
A subset of level 5 problems from https://huggingface.co/datasets/lighteval/MATH
We have extracted the final answer from boxed, and only keep those with integer outputs.
spider-context-validation
Dataset Card for Spider Context Validation
Dataset Summary
Spider is a large-scale complex and cross-domain semantic parsing and text-to-SQL dataset annotated by 11 Yale students
The goal of the Spider challenge is to develop natural language interfaces to cross-domain databases.
This dataset was created to validate spider-fine-tuned LLMs with database context.
Yale Lily Spider Leaderboards
The leaderboard can be seen at https://yale-lily.github.io/spider… See the full description on the dataset page: https://huggingface.co/datasets/richardr1126/spider-context-validation.korea_speech_mfa_aligned_validationrcp-ndcg-external-validation
RCP-nDCG: external validation
This dataset holds the external validation of the paper "Rubric-Calibrated Preferences:
Cross-Query Calibration of LLM Judgments via Item Response Theory" (Schmidt, Crisostomi, Lassance, Reimers, 2026; arXiv:2609.35739): the blind human study and the blind comparisons by external LLM judges.
It contains:
every grade, tie-break, review and verdict of the human study;
every prompt and response of the three external LLM judges, reasoning traces… See the full description on the dataset page: https://huggingface.co/datasets/fabianschmidt-cohere/rcp-ndcg-external-validation.TextVQA_validation
Dataset Card for "TextVQA_validation"
More Information needed
VQAv2_sample_validation
Dataset Card for "VQAv2_sample_validation"
More Information needed
FineFineWeb-validation
FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus
arXiv: Coming Soon
Project Page: Coming Soon
Blog: Coming Soon
Data Statistics
Domain (#tokens/#samples)
Iteration 1 Tokens
Iteration 2 Tokens
Iteration 3 Tokens
Total Tokens
Iteration 1 Count
Iteration 2 Count
Iteration 3 Count
Total Count
aerospace
5.77B
261.63M
309.33M
6.34B
9100000
688505
611034
10399539
agronomy
13.08B
947.41M
229.04M
14.26B
15752828
2711790
649404
19114022
artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-validation.validationValidation dataset for General AI Assistant
music-validation-datasetCOVID-QA-unique-context-test-10-percent-validation-10-percent
Dataset Card for "COVID-QA-unique-context-test-10-percent-validation-10-percent"
More Information needed
pile-validationThis dataset has been created as an artefact of the paper Causal Estimation of Memorisation Profiles (Lesci et al., 2024).
More info about this dataset in the related collection Memorisation-Profiles.
The validation data used in our study. The Pythia suite does not have an official validation. However, we confirmed with the authors that the Pile validation split (this one) was not seen during training.
It is still a bit confusing whether the Pile data can be released freely. Thus, we will… See the full description on the dataset page: https://huggingface.co/datasets/pietrolesci/pile-validation.gol-rl-fixed-validation-37156495
GoL World Model — Genie at Tiny Scale
Current implementation status is tracked in STATUS_2026-05-19.md.
The repo currently has three executable tracks: the recursive world-model
demos/training path, agentic trajectory collection, and online GRPO training
via train_grpo.py.
A miniature implementation of the Genie world model
architecture using Conway's Game of Life as the substrate.
Goal: Show that resource-constrained researchers can experiment with world model ideas
using a… See the full description on the dataset page: https://huggingface.co/datasets/brysgo/gol-rl-fixed-validation-37156495.Imagenet1k_sample_validation
Dataset Card for "Imagenet1k_sample_validation"
More Information needed
sae-skeskinen-TinyStories-hf-validation-tokenizer-gpt2_playwikipedia-20220301.en-0.005-validationKorea-AIHub-middlesenior-dialect-speech-validation-part2brand-spectrometer-validation
Brand Spectrometer — Validation Study
Reproducible validation data for the Brand Spectrometer, an instrument that reads
cohort-resolved, eight-dimensional brand-perception specifications from public artifacts
via cross-operator LLM pipelines.
This dataset accompanies the Brand Spectrometer methods paper and holds the raw,
fully-reproducible outputs of its validation battery. The instrument is ground-truth
absent by design: it does not recover a "true" brand spec, and cohort… See the full description on the dataset page: https://huggingface.co/datasets/spectralbranding/brand-spectrometer-validation.validationValidation dataset for General AI Assistant
c4-en-validationnq_open-validation
Dataset Card for "nq_open-validation"
More Information needed
OMol25_validation
Cite this dataset Levine, D. S., Shuaibi, M., Spotte-Smith, E. W. C., Taylor, M. G., Hasyim, M. R., Michel, K., Batatia, I., Csányi, G., Dzamba, M., Eastman, P., Frey, N. C., Fu, X., Gharakhanyan, V., Krishnapriyan, A. S., Rackers, J. A., Raja, S., Rizvi, A., Rosen, A. S., Ulissi, Z., Vargas, S., Zitnick, C. L., Blau, S. M., and Wood, B. M. OMol25 validation. ColabFit, 2025. https://doi.org/10.60732/8baea040
This dataset has been curated and formatted for the ColabFit… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/OMol25_validation.
