Team Ai
Datasetpublic

dougalldeepmind/2026-08-05-model-eval-model-validation

2026-08-05-model-eval-model-validation field value experiment 15-document human-verification batch of the model-eval-model document type (3 per cell), generated by the real pipeline ahead of the full 2,100-doc corpus for the 20/80-by-examples SFT run. date_generated 2026-08-05 constitution claude_distilled_09_principles_mid_20260804 — the 9-principle generation-time snapshot (sha256 fe2ed960…), byte-exact match of the constitution the source difficult-advice… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-05-model-eval-model-validation.

sourceHugging Facemitupdated 1mo agoView on Hugging Face
0likes39downloads
Dataset Card

2026-08-05-model-eval-model-validation

fieldvalue
experiment15-document human-verification batch of the model-eval-model document type (3 per cell), generated by the real pipeline ahead of the full 2,100-doc corpus for the 20/80-by-examples SFT run.
date_generated2026-08-05
constitutionclaude_distilled_09_principles_mid_20260804 — the 9-principle generation-time snapshot (sha256 fe2ed960…), byte-exact match of the constitution the source difficult-advice corpus was generated against; in the source repo under constitutions/claude_distilled_09_principles_mid_20260804/.
source_repohttps://github.com/Matthew-Bozoukov/teachingclaudewhy_replication @ 26b464063bdc1d162b0fadbf4ec96a4cc5a788b4
modelsanthropic/claude-sonnet-5 (all four generation call sites: control / critique / reflect / perturb, via OpenRouter)
generation_configseed 0; control t=0.7 maxtokens=8192; critique/reflect t=0.8 maxtokens=12288; perturb t=0.7 maxtokens=8192; explicitness name/paraphrase/embody 0.3/0.4/0.3; flaw grid omission/commission/miscalibration/overapplication × clear/moderate/grey.
schemaid = recordid (`<scenarioid>::<cell>); messages = the training conversation (assistant turns may carry reasoningcontent` — the private deliberation); `metadata` = cell, attribution, responsekind, flawtype/severity, explicitness, verdict, trait, supervise, plus browser projections `split`=cell, `category`=traitname and (flawed cells) change_summary = the planted flaw, metadata-only, never trains.
provenanceuv run scripts/data/synthdoc/build_dataset.py --config configs/data/synthdoc/model_eval_model.yaml --overrides "cells.control=3,cells.m4_other_good=3,cells.m3_other_flawed=3,cells.m2_self_good=3,cells.m1_self_flawed=3" (source corpus LASR-Callum/2026-08-04-synthdoc-package-difficult-advice-stage-cache, run 20260805_133015).

Cells: control, m1selfflawed, m2selfgood, m3otherflawed, m4othergood. Validity checks (synthdoc check) all pass; note the flaw-identification probe at this n: 1/4 clear hits — the review question is whether perturbations are too subtle or critiques too charitable. supervise in metadata marks which assistant turns train (final for the self cells).