Team Ai
Datasetpublic

Compactbot/slm-arch-scores

SLM Architecture → Score (controlled ablation panel) A small, controlled dataset of per-task zero-shot benchmark scores across different architectures, harvested from the model cards of the d0rj/tiny-llm-ablation family. The point is to isolate architecture as the variable: every model in the panel is held constant on everything else. Why this panel is controlled All models share: ~51M parameters, trained from scratch (not finetunes) Same data: FineWeb-Edu… See the full description on the dataset page: https://huggingface.co/datasets/Compactbot/slm-arch-scores.

sourceHugging Faceapache-2.0updated 15d agoView on Hugging Face
0likes135downloads
Dataset Card

SLM Architecture → Score (controlled ablation panel)

A small, controlled dataset of per-task zero-shot benchmark scores across different architectures, harvested from the model cards of the d0rj/tiny-llm-ablation family. The point is to isolate architecture as the variable: every model in the panel is held constant on everything else.

Why this panel is controlled

All models share:

  • —~51M parameters, trained from scratch (not finetunes)
  • —Same data: FineWeb-Edu sample-10BT, 3,932,160,000 source tokens
  • —Same budget: 15,000 optimizer steps
  • —Same tokenizer: 32,768 tokens
  • —Same eval protocol: lm-eval 0.4.12, zero-shot, 8 tasks, full official splits, 95% Wilson confidence intervals, BF16
  • —Same width/heads/FFN: d=512, 8 query / 2 KV heads, SwiGLU ffn=1792, ctx 2048

The only thing that varies is the architecture family. That is what makes an architecture→score comparison meaningful — most "which arch is best" threads confound architecture with scale, data and tokenizer.

The models

repofamilywhat varies
d0rj/q-51M-basecausal-GPTreference baseline (10 decoder layers)
d0rj/q-prefixlm-51M-baseprefix-LMbidirectional prefix context + suffix-only loss
d0rj/looped-51M-baselooped (Universal-Transformer)10 unique blocks weight-shared × 6 loops = 60 effective layers
d0rj/diffusion-51M-basediffusion / masked LMdifferent metric — see caveat
d0rj/prefixlm-51M-base is a duplicate of q-prefixlm-51M-base (identical eval scores; the only difference is whether a 32-element RoPE buffer is counted in the parameter total: 50,866,720 vs 50,866,688). It is included for completeness and flagged duplicate_of.

Headline result (AR-comparable models only)

familyHellaSwagARC-EARC-CPIQAWinoGOBQABoolQLAMBADA**macro**
causal-GPT29.1843.3124.2359.9050.0428.2059.8820.8639.45
prefix-LM28.3936.2422.7853.1049.7225.6054.8623.3536.76
looped29.6244.2822.1060.2850.1229.0061.5920.9039.74
  • —Looped (weight-shared depth) ≥ causal on 7 of 8 tasks (macro 39.74 vs 39.45); it wins most on ARC-Easy, PIQA, OBQA, BoolQ.
  • —Prefix-LM < causal on 7 of 8 tasks (macro 36.76 vs 39.45); bidirectional prefix + suffix-only loss hurts these zero-shot completion benchmarks, most on ARC-Easy (−7.1) and PIQA (−6.8). Its one win is LAMBADA (+2.5), where bidirectional context helps predict the final word.

Caveats (read before trusting this)

  1. 1.n = 3 distinct AR-comparable architectures. This is a pairwise comparison panel, not a correlation. You cannot fit an architecture→score regression on three points; the honest claim is "in this controlled panel, looped ≥ causal and prefix-LM < causal", not "deeper/shared archs correlate with score".
  2. 2.Single training seed. Differences of ~1–2 pts are within the 95% Wilson CIs on most tasks (e.g. HellaSwag causal CI [28.30, 30.07] overlaps both rivals). The directional pattern (looped up, prefix down, 7/8 tasks each) is more robust than any single-task gap.
  3. 3.Diffusion row is a different metric. diffusion-51M-base is scored with continuation perplexity (its LAMBADA 42.21 is a PLL, not AR loglikelihood), so it is excluded from the AR macro and must not be mixed into the comparison. Its card says so explicitly.
  4. 4.Zero-shot, uncorrected for contamination. One seed, no multiple-comparison correction (as the source cards state).

Source & provenance

Scores are author-reported model-index / evaluation/results.json values from the four d0rj repos, harvested 2026-09-24. This dataset is a harvest + honest-analysis artifact: it does not re-run the evals, it re-states the source numbers with the controlled-design framing and the metric caveat made explicit. To reproduce the underlying evals, see each repo's evaluation/run_core.py.

Files

  • —slm_arch_scores.jsonl — one row per model: arch features, per-task {metric, value, ci95, n}, ar_macro (null for the diffusion row), metric_type, duplicate_of.