Compactbot/slm-arch-scores
SLM Architecture → Score (controlled ablation panel) A small, controlled dataset of per-task zero-shot benchmark scores across different architectures, harvested from the model cards of the d0rj/tiny-llm-ablation family. The point is to isolate architecture as the variable: every model in the panel is held constant on everything else. Why this panel is controlled All models share: ~51M parameters, trained from scratch (not finetunes) Same data: FineWeb-Edu… See the full description on the dataset page: https://huggingface.co/datasets/Compactbot/slm-arch-scores.
SLM Architecture → Score (controlled ablation panel)
A small, controlled dataset of per-task zero-shot benchmark scores across different architectures, harvested from the model cards of the d0rj/tiny-llm-ablation family. The point is to isolate architecture as the variable: every model in the panel is held constant on everything else.
Why this panel is controlled
All models share:
- ~51M parameters, trained from scratch (not finetunes)
- Same data: FineWeb-Edu
sample-10BT, 3,932,160,000 source tokens - Same budget: 15,000 optimizer steps
- Same tokenizer: 32,768 tokens
- Same eval protocol: lm-eval 0.4.12, zero-shot, 8 tasks, full official splits, 95% Wilson confidence intervals, BF16
- Same width/heads/FFN: d=512, 8 query / 2 KV heads, SwiGLU ffn=1792, ctx 2048
The only thing that varies is the architecture family. That is what makes an architecture→score comparison meaningful — most "which arch is best" threads confound architecture with scale, data and tokenizer.
The models
d0rj/prefixlm-51M-baseis a duplicate ofq-prefixlm-51M-base(identical eval scores; the only difference is whether a 32-element RoPE buffer is counted in the parameter total: 50,866,720 vs 50,866,688). It is included for completeness and flaggedduplicate_of.
Headline result (AR-comparable models only)
- Looped (weight-shared depth) ≥ causal on 7 of 8 tasks (macro 39.74 vs 39.45); it wins most on ARC-Easy, PIQA, OBQA, BoolQ.
- Prefix-LM < causal on 7 of 8 tasks (macro 36.76 vs 39.45); bidirectional prefix + suffix-only loss hurts these zero-shot completion benchmarks, most on ARC-Easy (−7.1) and PIQA (−6.8). Its one win is LAMBADA (+2.5), where bidirectional context helps predict the final word.
Caveats (read before trusting this)
- n = 3 distinct AR-comparable architectures. This is a pairwise comparison panel, not a correlation. You cannot fit an architecture→score regression on three points; the honest claim is "in this controlled panel, looped ≥ causal and prefix-LM < causal", not "deeper/shared archs correlate with score".
- Single training seed. Differences of ~1–2 pts are within the 95% Wilson CIs on most tasks (e.g. HellaSwag causal CI [28.30, 30.07] overlaps both rivals). The directional pattern (looped up, prefix down, 7/8 tasks each) is more robust than any single-task gap.
- Diffusion row is a different metric.
diffusion-51M-baseis scored with continuation perplexity (its LAMBADA 42.21 is a PLL, not AR loglikelihood), so it is excluded from the AR macro and must not be mixed into the comparison. Its card says so explicitly. - Zero-shot, uncorrected for contamination. One seed, no multiple-comparison correction (as the source cards state).
Source & provenance
Scores are author-reported model-index / evaluation/results.json values from the four d0rj repos, harvested 2026-09-24. This dataset is a harvest + honest-analysis artifact: it does not re-run the evals, it re-states the source numbers with the controlled-design framing and the metric caveat made explicit. To reproduce the underlying evals, see each repo's evaluation/run_core.py.
Files
slm_arch_scores.jsonl— one row per model:archfeatures, per-task{metric, value, ci95, n},ar_macro(null for the diffusion row),metric_type,duplicate_of.
