datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aimo-validation-amc
Dataset Card for AIMO Validation AMC
All 83 come from AMC12 2022, AMC12 2023, and have been extracted from the AOPS wiki page https://artofproblemsolving.com/wiki/index.php/AMC_12_Problems_and_Solutions
This dataset serves as an internal validation set during our participation in the AIMO progress prize competition. Using data after 2021 is to avoid potential overlap with the MATH training set.
Here are the different columns in the dataset:
problem: the modified problem statement… See the full description on the dataset page: https://huggingface.co/datasets/AI-MO/aimo-validation-amc.QuantiPhy-validation
QuantiPhy (Validation Set)
Dataset Summary
QuantiPhy is a benchmark for evaluating whether vision–language models (VLMs) can perform quantitative physical inference from visual evidence, rather than producing plausible but ungrounded numerical guesses.
This repository contains the official validation set of QuantiPhy, released to support model development, ablation studies, and preliminary evaluation.The validation set represents approximately 4% of the full… See the full description on the dataset page: https://huggingface.co/datasets/PaulineLi/QuantiPhy-validation.real-routing-d2-c00-teleop-validation
real-routing-d2-c00-teleop-validation
Part of Mulligan. Browse the release at Policy Arena.
Property
Value
Task
real-routing-d2
Role
validation-view
Release grouping
mainline
Variant
validation
Model rounds
R0
Episodes
50
Frames
13768
Recording FPS
15
Cameras
observation.images.wrist_left, observation.images.wrist_right, observation.images.side_1, observation.images.side_2
Release tag
release-1
Parent
mulligan/real-routing-d2-c00-teleop-mixed… See the full description on the dataset page: https://huggingface.co/datasets/mulligan/real-routing-d2-c00-teleop-validation.real-square-d2-c00-teleop-validation
real-square-d2-c00-teleop-validation
Part of Mulligan. Browse the release at Policy Arena.
Property
Value
Task
real-square-d2
Role
validation-view
Release grouping
mainline
Variant
validation
Model rounds
R0
Episodes
50
Frames
7915
Recording FPS
15
Cameras
observation.images.wrist_left, observation.images.wrist_right, observation.images.side_1, observation.images.side_2
Release tag
release-1
Parent
mulligan/real-square-d2-c00-teleop-mixed… See the full description on the dataset page: https://huggingface.co/datasets/mulligan/real-square-d2-c00-teleop-validation.real-marker-d2-c00-teleop-validation
real-marker-d2-c00-teleop-validation
Part of Mulligan. Browse the release at Policy Arena.
Property
Value
Task
real-marker-d2
Role
validation-view
Release grouping
mainline
Variant
validation
Model rounds
R0
Episodes
50
Frames
9483
Recording FPS
15
Cameras
observation.images.wrist_left, observation.images.wrist_right, observation.images.side_1, observation.images.side_2
Release tag
release-1
Parent
mulligan/real-marker-d2-c00-teleop-mixed… See the full description on the dataset page: https://huggingface.co/datasets/mulligan/real-marker-d2-c00-teleop-validation.rcp-ndcg-external-validation
RCP-nDCG: external validation
This dataset holds the external validation of the paper "Rubric-Calibrated Preferences:
Cross-Query Calibration of LLM Judgments via Item Response Theory" (Schmidt, Crisostomi, Lassance, Reimers, 2026; arXiv:2609.35739): the blind human study and the blind comparisons by external LLM judges.
It contains:
every grade, tie-break, review and verdict of the human study;
every prompt and response of the three external LLM judges, reasoning traces… See the full description on the dataset page: https://huggingface.co/datasets/fabianschmidt-cohere/rcp-ndcg-external-validation.FineFineWeb-validation
FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus
arXiv: Coming Soon
Project Page: Coming Soon
Blog: Coming Soon
Data Statistics
Domain (#tokens/#samples)
Iteration 1 Tokens
Iteration 2 Tokens
Iteration 3 Tokens
Total Tokens
Iteration 1 Count
Iteration 2 Count
Iteration 3 Count
Total Count
aerospace
5.77B
261.63M
309.33M
6.34B
9100000
688505
611034
10399539
agronomy
13.08B
947.41M
229.04M
14.26B
15752828
2711790
649404
19114022
artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-validation.COVID-QA-unique-context-test-10-percent-validation-10-percent
Dataset Card for "COVID-QA-unique-context-test-10-percent-validation-10-percent"
More Information needed
gol-rl-fixed-validation-37156495
GoL World Model — Genie at Tiny Scale
Current implementation status is tracked in STATUS_2026-05-19.md.
The repo currently has three executable tracks: the recursive world-model
demos/training path, agentic trajectory collection, and online GRPO training
via train_grpo.py.
A miniature implementation of the Genie world model
architecture using Conway's Game of Life as the substrate.
Goal: Show that resource-constrained researchers can experiment with world model ideas
using a… See the full description on the dataset page: https://huggingface.co/datasets/brysgo/gol-rl-fixed-validation-37156495.brand-spectrometer-validation
Brand Spectrometer — Validation Study
Reproducible validation data for the Brand Spectrometer, an instrument that reads
cohort-resolved, eight-dimensional brand-perception specifications from public artifacts
via cross-operator LLM pipelines.
This dataset accompanies the Brand Spectrometer methods paper and holds the raw,
fully-reproducible outputs of its validation battery. The instrument is ground-truth
absent by design: it does not recover a "true" brand spec, and cohort… See the full description on the dataset page: https://huggingface.co/datasets/spectralbranding/brand-spectrometer-validation.OMol25_validation
Cite this dataset Levine, D. S., Shuaibi, M., Spotte-Smith, E. W. C., Taylor, M. G., Hasyim, M. R., Michel, K., Batatia, I., Csányi, G., Dzamba, M., Eastman, P., Frey, N. C., Fu, X., Gharakhanyan, V., Krishnapriyan, A. S., Rackers, J. A., Raja, S., Rizvi, A., Rosen, A. S., Ulissi, Z., Vargas, S., Zitnick, C. L., Blau, S. M., and Wood, B. M. OMol25 validation. ColabFit, 2025. https://doi.org/10.60732/8baea040
This dataset has been curated and formatted for the ColabFit… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/OMol25_validation.drawvla-prompt-validation-v3
DrawVLA — Sketch-Prompt Validation
Circle (which) + arrow (where) + caption (what) visual instructions overlaid on
LIBERO observations, each labelled with a binary
verdict for training a prompt validator or a self-checking VLA:
right — every channel is correct and exactly one reading survives; execute.
wrong — a channel is incorrect or the deictic prompt remains under-determined;
reject. Formerly ambiguous prompts are retained in this class.
All captions are name-free L2/L3… See the full description on the dataset page: https://huggingface.co/datasets/shibuina/drawvla-prompt-validation-v3.pen_pick_place_clean_200ep_v3_with_validationThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/ngustj/pen_pick_place_clean_200ep_v3_with_validation.drawvla-prompt-validation-clean
DrawVLA — Sketch-Prompt Validation
Circle (which) + arrow (where) + caption (what) visual instructions overlaid on
LIBERO observations, each labelled with a binary
verdict for training a prompt validator or a self-checking VLA:
right — every channel is correct and exactly one reading survives; execute.
wrong — a channel is incorrect or the deictic prompt remains under-determined;
reject. Formerly ambiguous prompts are retained in this class.
All captions are name-free L2/L3… See the full description on the dataset page: https://huggingface.co/datasets/shibuina/drawvla-prompt-validation-clean.sroiv2_strawberry_picking_lab_validationThis dataset was created using LeRobot.
SROI v2 — Strawberry Picking (Lab) — Validation Set
Held-out validation set for the SROI v2 strawberry-picking data (project page, Zhejiang University): 100 human strawberry-picking demonstrations recorded with the SROI V2 handheld data-acquisition device — a UMI-style gripper with an integrated Intel RealSense D405 stereo camera — on live plants in a laboratory setup. No robot arm is involved during collection: the 7-DoF… See the full description on the dataset page: https://huggingface.co/datasets/zfff/sroiv2_strawberry_picking_lab_validation.drawvla-prompt-validation
DrawVLA — Sketch-Prompt Validation
Circle (which) + arrow (where) + caption (what) visual instructions overlaid on
LIBERO observations, each labelled with a binary
verdict for training a prompt validator or a self-checking VLA:
right — every channel is correct and exactly one reading survives; execute.
wrong — a channel is incorrect or the deictic prompt remains under-determined;
reject. Formerly ambiguous prompts are retained in this class.
All captions are name-free L2/L3… See the full description on the dataset page: https://huggingface.co/datasets/shibuina/drawvla-prompt-validation.configreach-validation
ConfigReach Curated 50K Benchmark
The ConfigReach Curated 50K Benchmark is the committed controlled benchmark used to evaluate ConfigReach configuration-input detection across supported programming languages and configuration formats.
It contains 50,000 scenarios with deterministic ground-truth labels: 25,000 positive and 25,000 negative cases across 24 language/configuration groups. The benchmark is scored through the production configreach.engine.scan entry point.
Project hub:… See the full description on the dataset page: https://huggingface.co/datasets/sauravsingla08/configreach-validation.VQAv2_validation_no_image
Dataset Card for "VQAv2_validation_no_image"
More Information needed
pen_pick_place_clean_v2_validation_20260930_113146This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/ngustj/pen_pick_place_clean_v2_validation_20260930_113146.lalm-judge-validation-full-duplex
LALM Judge Validation on Full-Duplex Voice Agents
Companion dataset for the paper A Reliability Assessment of
LALM Audio Judges for Full-Duplex Voice Agents.
This repository contains the anonymised ratings, adversarial-defect
recall tables, JSON schemas, and analysis scripts used to produce
every headline number, table, and figure in that paper.
Summary
209 rated stereo sessions: 152 full-duplex agent-client
conversations across 13 accent-and-condition strata… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/lalm-judge-validation-full-duplex.Open_Molecular_Crystals_2025_OMC25_validation
Cite this dataset Gharakhanyan, V., Barroso-Luque, L., Yang, Y., Shuaibi, M., Michel, K., Levine, D. S., Dzamba, M., Fu, X., Gao, M., Liu, X., Ni, H., Noori, K., Wood, B. M., Uyttendaele, M., Boromand, A., Zitnick, C. L., Marom, N., Ulissi, Z. W., and Sriram, A. Open Molecular Crystals 2025 OMC25 validation. ColabFit, 2025. https://doi.org/None
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/Open_Molecular_Crystals_2025_OMC25_validation.OMol25_neutral_validation
Cite this dataset Levine, D. S., Shuaibi, M., Spotte-Smith, E. W. C., Taylor, M. G., Hasyim, M. R., Michel, K., Batatia, I., Csányi, G., Dzamba, M., Eastman, P., Frey, N. C., Fu, X., Gharakhanyan, V., Krishnapriyan, A. S., Rackers, J. A., Raja, S., Rizvi, A., Rosen, A. S., Ulissi, Z., Vargas, S., Zitnick, C. L., Blau, S. M., and Wood, B. M. OMol25 neutral validation. ColabFit, 2025. https://doi.org/10.60732/0d5818c5
This dataset has been curated and formatted for the… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/OMol25_neutral_validation.eskulap_validation_datasetsxxMD-CASSCF_validation
Cite this dataset Pengmei, Z., Shu, Y., and Liu, J. xxMD-CASSCF validation. ColabFit, 2023. https://doi.org/10.60732/cea2a8c1
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_e0nl35x6dl85_0
Visit the ColabFit Exchange to search additional datasets by author, description, element content and more.
https://materials.colabfit.org… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/xxMD-CASSCF_validation.OC22-IS2RE-Validation-out-of-domain
Cite this dataset Tran, R., Lan, J., Shuaibi, M., Wood, B. M., Goyal, S., Das, A., Heras-Domingo, J., Kolluru, A., Rizvi, A., Shoghi, N., Sriram, A., Therrien, F., Abed, J., Voznyy, O., Sargent, E. H., Ulissi, Z., and Zitnick, C. L. OC22-IS2RE-Validation-out-of-domain. ColabFit, 2023. https://doi.org/10.60732/15fa94f2
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/OC22-IS2RE-Validation-out-of-domain.so101_pick_cup1_validation_finalThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 5,
"total_frames": 760,
"total_tasks": 1,
"total_videos": 10,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/rowb1/so101_pick_cup1_validation_final.xxMD-DFT_validation
Cite this dataset Pengmei, Z., Shu, Y., and Liu, J. xxMD-DFT validation. ColabFit, 2023. https://doi.org/10.60732/bd646241
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_d1nmrzy4csx1_0
Visit the ColabFit Exchange to search additional datasets by author, description, element content and more.
https://materials.colabfit.org… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/xxMD-DFT_validation.OC22-IS2RE-Validation-in-domain
Cite this dataset Tran, R., Lan, J., Shuaibi, M., Wood, B. M., Goyal, S., Das, A., Heras-Domingo, J., Kolluru, A., Rizvi, A., Shoghi, N., Sriram, A., Therrien, F., Abed, J., Voznyy, O., Sargent, E. H., Ulissi, Z., and Zitnick, C. L. OC22-IS2RE-Validation-in-domain. ColabFit, 2023. https://doi.org/10.60732/ced227e5
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/OC22-IS2RE-Validation-in-domain.Training-Validation-Test-ILNThis Dataset is saved by Students of the Subject: Interação em Linguagem Natural 2023/24
From: FCUL (Faculdade de Ciências da Universidade de Lisboa)
Used for Sentiment Analysis with a pre-trained Roberta Model.
OMat24_validation_rattled_300
Cite this dataset Barroso-Luque, L., Shuaibi, M., Fu, X., Wood, B. M., Dzamba, M., Gao, M., Rizvi, A., Zitnick, C. L., and Ulissi, Z. W. OMat24 validation rattled 300. ColabFit, 2025. https://doi.org/10.60732/b3c0c67d
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_mm4npn96qxo1_0
Visit the ColabFit Exchange to search additional datasets… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/OMat24_validation_rattled_300.
