windowsxp811203/nvfp4-mtp-survey
Do Qwen3.8-27B NVFP4 repos actually ship a working MTP draft head? A static survey of every NVFP4 quantization of Qwen3.8-27B and its finetunes that I could find on the Hugging Face Hub, last run on 2026-08-24 (Rev 4) with nvfp4_mtp_audit.py. Raw output: results.json. I ran this to check a claim I had made in public, and the claim did not survive. The correction is the first section, because it is the most important result here. Revision history — read this, it is… See the full description on the dataset page: https://huggingface.co/datasets/windowsxp811203/nvfp4-mtp-survey.
Do Qwen3.8-27B NVFP4 repos actually ship a working MTP draft head?
A static survey of every NVFP4 quantization of Qwen3.8-27B and its finetunes that I could find on the Hugging Face Hub, last run on 2026-08-24 (Rev 4) with `nvfp4_mtp_audit.py`. Raw output: `results.json`.
I ran this to check a claim I had made in public, and the claim did not survive. The correction is the first section, because it is the most important result here.
Revision history — read this, it is part of the result
This survey has itself needed correcting three times, which is its own lesson about census work.
- Rev 1 (2026-08-23 morning): 55 repos found via the Hub's
base_modelrelation graph; reported "40 of 42". A convenience sample — builders who declare nobase_modelwere invisible — and re-uploads were double-counted. - Rev 2 (2026-08-23 evening): added name-search discovery (125 found) and artifact deduplication; reported "65 of 72". The dedup was buggy: it fingerprinted only the largest weight shard, which falsely collapsed a head-stripped variant into its parent and unrelated builds into each other, and it treated API-masked (asterisked) LFS OIDs as matching. It also excluded text-only conversions from every denominator and counted one repo's quantization-scale metadata as extra tensors.
- Rev 3 (2026-08-24 morning): fingerprint over every weight shard, masked OIDs never collapse, config included in the dedup key, text-only conversions audited for their heads, per-module counting. Reported "74 of 87".
- Rev 4 (2026-08-24, this document): a second review pass caught four more tool blind spots: MLX serializes scales as
.scalesand its vision tower asvision_tower.*(three full multimodal checkpoints had been misclassified as text-only and unquantized), compressed-tensors also emitsinput_global_scale, the fingerprint now walks the repo tree recursively, and the Gated DeltaNet denominator now counts the 240 Linear projections rather than all 336 weight-bearing modules (48 conv1d + 48 norms were diluting every percentage — and hiding the fact that some builds quantize the entire projection set). The headline conclusion has survived four revisions; most of the numbers did not, which is exactly why the rawresults.jsonand the tool ship alongside the claims.
The correction
In the model cards for my two NVFP4 builds, and in comments on two Hugging Face discussions, I wrote that most NVFP4 conversions of this architecture ship a dead MTP draft head, and that keeping the vision tower in bf16 is something these conversions usually lose.
Both statements are wrong.
The mechanism I described is real, and I hit it myself: Qwen3_5ForConditionalGeneration.from_pretrained does not instantiate the MTP module, so a naive llm-compressor run cannot save it, and a bf16 head that is present but missing from quantization_config.ignore is loaded as a quantization target and rejects every draft — 0% acceptance with clean-looking logs. What I got wrong was the prevalence. I generalized from my own first broken build to the population without measuring the population. The people publishing these quants had, for the most part, already solved it.
(Upstream has too, since: llm-compressor merged a fix on 2026-08-17 — PR #3033 — that saves the Qwen3.8 MTP tensors and appends the ignore entry automatically. As of this writing it is in no released wheel, so pip install llmcompressor still reproduces the bug; the guard previously checked num_mtp_layers while Qwen3.8's config says mtp_num_hidden_layers.)
What the survey covers
Discovery is two passes: a walk of the Hub's base_model relations from Qwen/Qwen3.8-27B through its finetunes to their quantizations, plus a name search that catches builders who declare no lineage. Repos tagged or named NVFP4 are kept. For each, the tool fetches config.json, model.safetensors.index.json, and the LFS sha256 of every weight shard — never the weights.
Census as of 2026-08-24 evening (the population grows daily — 55 → 125 → 133 → 136 across four runs in two days; re-run the tool for current numbers):
- 136 NVFP4 repos discovered
- 13 are true re-uploads — every weight shard sha256-identical to another repo and the same
quantization_config(the highest-download copy counts;RedHatAI/Qwen3.8-27B-NVFP4is unsloth's artifact with a different README) - 31 not statically auditable (single-file checkpoint, gated repo, or upload in progress), 1 with unrecognized tensor naming
- 91 independent auditable artifacts: 87 full multimodal, 4 text-only conversions
Results (Rev 4)
"Unquantized" means the index carries no recognized quantization-scale keys for the path; this static audit cannot distinguish bf16 from fp16/f32, and does not model every packing scheme.
21 artifacts are flagged, and one of them doesn't belong in the same bucket: gittensor-model-hub/…-No-MTP strips the head deliberately, documents it, and — alone in the flagged set — also removes the MTP layer count from its config. The other 20 (22% of the census) split into three distinct situations that should not be lumped together:
- *12 declare `mtp_num_hidden_layers` in `config.json` while shipping no `mtp.` tensors at all.** That mismatch is a fact, not an inference, and for these the config promises a head the checkpoint cannot back.
- 5 ship the complete head with quantization scales on it. The head is there and may work; quantized MTP heads are simply unmeasured territory — verify acceptance before trusting.
- 3 ship a complete unquantized head that is absent from a compressed-tensors ignore list — the configuration that produced 0% acceptance on my own first build. Suspect, not proven broken.
None of the 20 have been run by this survey; these are static flags, not measured failures.
Two patterns in the flagged set are worth more than the count:
- Calibration-based recipes are over-represented in it. 7 of the 20 are GPTQ/QAT-labeled builds, above those recipes' share of the healthy population — consistent with the mechanism, since a calibration pass runs through
from_pretrained, which is exactly where the module vanishes. - Restricting to compressed-tensors builds (the only stack where the ignore-list rule is checkable): 31 independent artifacts — 20 correct, 6 with no head, 3 not in ignore, 2 quantized. A ~35% problem rate in that subset. Still nowhere near "most" of the overall census.
The one thing that distinguishes healthy builds from each other is the Gated DeltaNet path. Counted per projection module (240 Linears; the path's 48 conv1d and 48 norms are excluded from the denominator because essentially no stack quantizes them): 15 of 91 artifacts leave every projection unquantized, 20 quantize all 240, and the strictly-partial builds span 20–97% coverage. An earlier revision of this survey said "none quantize all of them" — that was an artifact of diluting the denominator with conv1d and norm weights. That is a real engineering choice with a real tradeoff — file size against a path that several vendors treat as sensitive (AMD's official Quark build of the 2.4T model excludes *linear_attn* outright) — not a matter of one group being broken. (Caveat for anyone parsing ignore lists instead of the index: llm-compressor #3040/#3044 misclassifies any module class containing the substring "Gate" at save time, so linear_attn entries in some configs are artifacts. This survey counts scale keys in the safetensors index, which is why its numbers disagree with a naive reading of those configs.)
Limits — please read before quoting this
- This is a configuration check, not a measurement. The only proof that a draft head works is serving the model and reading the draft acceptance rate. A repo this survey calls
okcan still be broken for reasons that never appear in a config file, and a flagged repo may be fine. A useful discriminator if you run one (due to leonardlin's Ornith work): compare draft acceptance with the head in bf16 vs quantized — if they match, the head itself is bad; if quantized is far lower, quantization ate it. ok-unverified-schemameans exactly that: the tensors look right, but the repo uses a stack whose loader path this script does not model. It is not an endorsement.- Precision claims ("bf16") are only made about my own builds, which were verified directly; for the population this survey can only say "no recognized quantization-scale keys".
- Hub metadata is live; numbers drift daily (the census grew 55 → 125 → 133 across three runs in two days).
Reproducing
python nvfp4_mtp_audit.py --sweep --json results.json # full survey
python nvfp4_mtp_audit.py some-org/some-nvfp4-repo # one repoOnly needs Python 3 and network access.
