Compactbot/slm-arch-scores
SLM Architecture → Score (controlled ablation panel) A small, controlled dataset of per-task zero-shot benchmark scores across different architectures, harvested from the model cards of the d0rj/tiny-llm-ablation family. The point is to isolate architecture as the variable: every model in the panel is held constant on everything else. Why this panel is controlled All models share: ~51M parameters, trained from scratch (not finetunes) Same data: FineWeb-Edu… See the full description on the dataset page: https://huggingface.co/datasets/Compactbot/slm-arch-scores.
Add benchmark protocol specs (from slm-architecture-benchmark-specs, now consolidated here)
Add cross-model panel findings JSON (from slm-arch-score-panel, now consolidated here)
Add cross-model panel analysis (from slm-arch-score-panel, now consolidated here)
Add cross-model score panel data (from slm-arch-score-panel, now consolidated here)
Add dataset card
Add controlled arch→score panel (d0rj ablation family)
README: reflect independent reproduction of the 5 peer rows + discrepancy note
Fill in per-task breakdown for the 5 peer models (independently reproduced, lm-eval 0.4.13)
Add dataset card (was 404): architecture-to-scores panel, 6 models (#1)
v1: 6-model same-harness arch→score panel (lm-eval 0.4.12, 0-shot, seed 1234, decontaminated)
initial commit
