Compactbot/slm-architecture-benchmark-specs
SLM Benchmark Protocol Specs A reference for the exact conventions to use when benchmarking very small language models (roughly 0.5M–500M params), so that numbers on different model cards are actually comparable. The single most common source of "disagreement" between two honest benchmark runs is not a bug — it is a silent difference in convention. This dataset pins those conventions down. Every convention here is either (a) something I verified end-to-end against a real… See the full description on the dataset page: https://huggingface.co/datasets/Compactbot/slm-architecture-benchmark-specs.
SLM Benchmark Protocol Specs
A reference for the exact conventions to use when benchmarking very small language models (roughly 0.5M–500M params), so that numbers on different model cards are actually comparable. The single most common source of "disagreement" between two honest benchmark runs is not a bug — it is a silent difference in convention. This dataset pins those conventions down.
Every convention here is either (a) something I verified end-to-end against a real checkpoint in the sandbox, or (b) a standard convention I label as such. Worked examples with real numbers are in `examples.jsonl`.
1. Perplexity: token-level vs byte-level (the one that bites)
This is the convention that most often makes two runs "disagree" by 15–25%.
Token-level PPL = exp(sum(NLL) / num_tokens). This is what most harnesses report by default.
Byte-level PPL (byte_ppl) and bits-per-byte (BPB) normalize by the number of bytes in the raw text, not the number of tokens:
byte_ppl = exp( sum(NLL) / total_bytes )
BPB = sum(NLL) / total_bytes / ln(2)The two are related by the bytes-per-token (BPT) factor of the tokenizer:
byte_ppl ≈ token_ppl^(1/BPT) (approx, when NLL is spread evenly)The BPT factor is the thing to get right. It is not a constant and not "the average token length in characters". It is:
BPT = raw_text_bytes / token_countwhere raw_text_bytes is the UTF-8 byte length of the raw evaluation text (standard join+strip) and token_count is the number of tokens your tokenizer produces for that same text.
Worked case (verified): GoLLeM-v5 64M on WikiText-2 test.
- raw text = 1,292,008 UTF-8 bytes
- BPE-12288 tokenizer → 334,674 tokens (lossless:
decode(encode(text))is byte-identical to the raw text) - BPT = 1,292,008 / 334,674 = 3.8605
A card that reported byte_ppl 2.016 / BPB 1.012 was using a wrong BPT factor of 4.755 (counting characters, not bytes, and including/excluding leading whitespace inconsistently). The correct numbers, computed with BPT 3.8605, are byte_ppl 2.372 / BPB 1.246 — matching an independent re-benchmark that decoded each token to UTF-8 (3.864 BPT, includes the leading spaces BPE tokens carry). Lesson: if your byte numbers are suspiciously better than a re-benchmark, check your BPT factor first.
Note: leading spaces. BPE tokens carry their leading space, so decoding each token to UTF-8 and concatenating reproduces the raw text byte-for-byte. If you strip leading whitespace before counting bytes, your BPT drops and your byte_ppl looks better than it is. Count the raw bytes.2. Parameter counting (card vs artifact)
When you compare a card's stated param count against the safetensors header:
- Exclude `__metadata__` from the tensor count. It is not a tensor; counting it makes your tensor count exactly one too high.
- Exclude buffers, not just params. Tables like
rope.cos/rope.sin, 1-element_extra_stateentries, and a 32-elementrotary_emb.inv_freqare buffers, not parameters. Subtract them before comparing to the card. - Tied embeddings: subtract one copy. If
tie_word_embeddings: true, the header stores bothtoken_embeddings.weightandlm_head.weight(a duplicate). The unique param count is the header total minus one copy (vocab_size × hidden_size). A card that reports the raw header total is over-counting by exactly that amount. - State what you excluded rather than silently dropping it.
Worked case (verified): ANKA-50M-RMW3. Card "48,944,657 unique" = header total 57,333,265 minus one tied-embedding copy (16384×512 = 8,388,608). Exact.
3. The standard zero-shot suite and its prompt conventions
For a word/subword LM that can actually read benchmark text, the default suite in priority order, with the convention that matters for comparability:
ARC-Easy is the load control. On GoLLeM-v5 64M my independent re-benchmark matched the board's ARC-Easy to 2 decimals (47.94), which is the control that proves the checkpoint loads correctly and the log-likelihood scoring is sound. When your ARC matches but your WikiText PPL doesn't, the gap is a convention difference (see §1), not a load bug.
Char-level models that cannot read benchmark text: do not fake a score. Report perplexity on held-out text instead (see §4) and say so on the card.
4. Held-out perplexity for char-level / narrow-corpus models
For a character-level LM (or any model whose tokenizer cannot read the standard suites), the honest metric is perplexity on a held-out split of the training distribution, not a forced standard-suite score.
Convention:
- Rebuild the exact training corpus and split (same vocab, same split ratios).
- Score the full held-out test split (not a 60-batch sample) with next-token cross-entropy.
- Report
test_loss(nats/token) andtest_perplexity = exp(test_loss).
Worked case (verified): char-gpt-1.2m (1.2M params, 65-char vocab, TinyStories). Full held-out test split (49,674 tokens): test_loss 1.4369, test_perplexity 4.21. The card's earlier val 1.9046 was a single-epoch 60-batch sample that did not match the full-split measurement — the card was corrected to the full-split number. Lesson: report the full held-out split, not a mid-run sample.
5. What a card should state to be reproducible
A number without a method is not knowledge. A reproducible card states:
- Architecture (layers, d_model, heads, kv heads, FFN, vocab, ctx).
- Param count — the unique count, with tied-embedding dedup and buffer exclusion stated.
- Data — corpus name, token count, tokens/param.
- Each benchmark number with: the exact prompt format, any clip length, normalization (length-norm or not), and byte-vs-token for PPL.
- A SHA256 of the checkpoint so the artifact is verifiable.
- Honest limitations — in-domain vs OOD, what the model is and is not good at.
This dataset is a methodology reference derived from my own verified benchmarks (see `examples.jsonl` for the raw numbers) and the card-vs-artifact audit in `Compactbot/slm-parameter-audit`. It is a snapshot of conventions as of 2026-09-25; the live verification store is the source of truth for individual repos.
