BrainAlign/brain-lm-alignment-ds003604
Brain-LM alignment: ds003604 Representational-similarity alignment between language-model hidden states and child fMRI RDMs for ds003604 (children ages 5/7/9, auditory). Tasks: Sem, Phon, Gram, Plaus Sessions: ses-5, ses-7, ses-9 Cells: 12 Models: 14 families (5 real + 9 PARC noise-seed baselines) Rows: 1848 (family x checkpoint x task x session) Generated: 2026-08-29 Headline: no model is distinguishable from a random seed Alignment is computed as Spearman… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds003604.
Brain-LM alignment: ds003604
Representational-similarity alignment between language-model hidden states and child fMRI RDMs for ds003604 (children ages 5/7/9, auditory).
- Tasks: Sem, Phon, Gram, Plaus Sessions: ses-5, ses-7, ses-9 Cells: 12
- Models: 14 families (5 real + 9 PARC noise-seed baselines)
- Rows: 1848 (family x checkpoint x task x session)
- Generated: 2026-08-29
Headline: no model is distinguishable from a random seed
Alignment is computed as Spearman correlation over the upper triangle of the model RDM (1 - corrcoef over mean-pooled final-layer states) against the brain RDM, reported raw (rsa) and as a fraction of the noise ceiling.
Every value is referenced against a PARC noise-seed null (9 randomly initialised seeds across 3 architectures, same cells, same pipeline) rather than against zero -- see null_referenced.csv for per-cell z-scores.
Across all three datasets and 130 (family x cell) combinations:
Real models beat the noise seeds less often than chance. The only reliable structure in these numbers is the stimulus set and the RDM, not the model.
Scale trend for ds003604: Spearman(params, mean RSA) = -0.10 (p = 0.87), i.e. a 16x parameter increase buys nothing.
What this does and does not license
The RDMs have demonstrated inter-subject reliability (noise ceilings 0.23-0.88), but the upstream pipeline's positive controls fail on all three datasets (0/108, 0/6 and 0/8 stimulus controls significant). The instrument has therefore not been shown to have power against alignment that does exist.
The defensible claim is "no LM alignment is detectable by this measurement" -- not "language models do not align with the developing brain." This is a result about the benchmark.
Coverage correction
The previously published grid did not pass --sessions, so it fell back to ds003604's ses-5/7/9. ds002236 matched only ses-9 (2 of 6 cells) and ds006239 matched nothing, producing zero alignment rows. Deriving sessions from the RDM tree takes coverage from 14 to 26 cells, so this release contains the first measurement of 12 previously unscored cells -- including ds006239/SemLocal, the only run x stimulus crossed cell, where the scanner-run confound cannot arise.
On SemLocal, untrained (step-0) alignment falls inside the random-seed band on both sessions (z = -0.84, -0.30). The "untrained models align better, training destroys it" reading is not supported: the decline is drift within noise, and there was no alignment to destroy.
Files
Precision: fp32 throughout. A bf16 spot check shifted RSA by up to 2.8e-3 at the final checkpoint (~1/3 of the across-seed noise sd), so precision is not mixed.
