validation_5
flan-t5-large-vin-dob-ssn-validationbloom-560m-finetuned-aings-validation-data-1timl_varied_10_realistic_vision_v2.0_dreambooth_lora_500_steps_validation_ohwxbloom-560m-finetuned-aings-validation-data-3XGLM-564M-finetuned-aings-validation-data-2timl_images_stable-diffusion-v1-5_dreambooth_2500_steps_validation_ohwxCross_Project_5_Classic_without_validationneuromamba-v5-fineweb-validation
xsum_validation_t5aimo-validation-math-level-5
Dataset Card for AIMO Validation MATH Level 5
A subset of level 5 problems from https://huggingface.co/datasets/lighteval/MATH
We have extracted the final answer from boxed, and only keep those with integer outputs.
theo_impulsive-qwen_2_5-7b-14b-whitebox_probe_validation
Status: NOT the paper's results. White-box probe validation for 7B/14B (2026-09-13). The paper's canonical results are Misalignment-Empirics/theo_impulsive-qwen_2_5-7b-14b-32b-big_eval_results, run runs/impulsive-qwen_2_5-7b-14b-32b-20260925/.
MO_evals results
Raw per-sample results from MO_evals runs (private, CLAUDE.md §5). One directory per
upload; nothing here is aggregated — the Parquet and the .eval logs are the primary
evidence, the scorecard is a summary of them.… See the full description on the dataset page: https://huggingface.co/datasets/Misalignment-Empirics/theo_impulsive-qwen_2_5-7b-14b-whitebox_probe_validation.wmt14_de-en_validation_t5latent_vae_validation_archivedminicpm5-sft-swe-validation-200
MiniCPM5 SFT SWE Validation 200
A fixed, public task index for small-scale software-engineering evaluation. It contains two configurations, verified and pro, each with exactly 100 distinct tasks: 70 that the historical MiniCPM5-2B SFT baseline resolved and 30 that it did not resolve. Every row has benchmark, instance_id, task_id, sft_resolved, repo, and base_commit.
The task statements, repository contents, reference patches, and tests are not copied into this dataset. Join… See the full description on the dataset page: https://huggingface.co/datasets/eigentom/minicpm5-sft-swe-validation-200.
