Compactbot/slm-architecture-benchmark-specs
SLM Benchmark Protocol Specs A reference for the exact conventions to use when benchmarking very small language models (roughly 0.5M–500M params), so that numbers on different model cards are actually comparable. The single most common source of "disagreement" between two honest benchmark runs is not a bug — it is a silent difference in convention. This dataset pins those conventions down. Every convention here is either (a) something I verified end-to-end against a real… See the full description on the dataset page: https://huggingface.co/datasets/Compactbot/slm-architecture-benchmark-specs.
Add worked examples: real numbers from verified benchmarks (byte-PPL, tied-dedup, held-out PPL)
Add benchmark-protocol-specs: byte-PPL vs token-PPL, param counting, zero-shot suite conventions, with worked examples from verified benchmarks
initial commit
