canada-quant/GLM-5.3-Flash-W4A16-MTP
GLM-5.3-Flash — W4A16 (INT4) + BF16 MTP
Model description
INT4 weight-only quantization of zai-org/GLM-5.3-Flash with the BF16 MTP draft head kept for speculative decoding. Only the 36,288 routed-expert GEMMs are INT4 (GPTQ, symmetric, group-size 128); attention, router, shared experts, embeddings, the vision tower and the MTP head stay in BF16. Not a new model — all capability comes from the base model.
Full benchmark grids, protocols and research notes: BENCHMARKS.md.
Uses & recommended recipes
Quick start
# 1. Download (~178 GiB)
huggingface-cli download canada-quant/GLM-5.3-Flash-W4A16-MTP --local-dir /models/glm53-flash-w4a16-mtp
# 2. Serve on 4× H100 / H200 (other hardware: see Serving recipes)
docker run --gpus '"device=0,1,2,3"' --ipc=host --network=host --rm \
-v /models:/models vllm/vllm-openai:glm53-flash-x86_64-cu130 \
vllm serve /models/glm53-flash-w4a16-mtp --served-model-name glm53-w4 \
--tensor-parallel-size 4 --enable-expert-parallel \
--max-model-len 262144 --max-num-seqs 512 \
--gpu-memory-utilization 0.92 --no-enable-prefix-caching \
--speculative-config '{"method":"mtp","num_speculative_tokens":2}' \
--reasoning-parser glm45 --tool-call-parser glm47 --enable-auto-tool-choice \
--trust-remote-code --port 8000
# 3. Call it (OpenAI-compatible)
curl http://localhost:8000/v1/chat/completions -H 'Content-Type: application/json' -d '{
"model": "glm53-w4",
"messages": [{"role": "user", "content": "Prove there are infinitely many primes."}]
}'Two things that bite. If yourconfig.jsonpredates 2026-09-08, re-download it — older copies fail in vLLM withKeyError: 'layers.0.mlp.gate_up_proj.weight'(weights are unchanged). And always pass--max-num-seqs ≤ 512— the vLLM default of 1024 exceeds this hybrid linear-attention model's 512 Mamba/KDA-state cache blocks.
Serving recipes
All configurations use expert parallelism. fp8 KV is not available on Hopper for this NoPE model.
H100 / H200, TP=4 (pinned image) — the Quick start command is the benchmarked recipe. num_speculative_tokens: 2 is the sweet spot on Hopper (52–55% acceptance; N=5 collapses acceptance to ~30%). Keep prefix caching off — it measured −2…−5% on H200. On 8× H200, two independent TP=4 replicas behind a load balancer beat one TP=8 endpoint by +28–38% aggregate at c128–c512; TP=8 wins single-stream and holds one 7.8M-token pool. On 2× H200 (89.5 GiB weights per GPU) add --tensor-parallel-size 2 --max-num-seqs 128 --max-cudagraph-capture-size 128 for 262K, or --max-model-len 1048576 --max-num-seqs 16 --gpu-memory-utilization 0.95 --max-cudagraph-capture-size 64 --max-num-batched-tokens 4096 for 1M.
8× H100, TP=8 (v2 image, bench-validated 2026-09-28) — 8192/1024 prompts. The latency recipe is TP=8 + DFlash2-G K=7: 181.2 tok/s c1, +34% over TP8 without speculation (135.4). The aggregate recipe is TP=8 without speculation: 2,418 tok/s @c256; MTP N=2 (2,330 @c256) is the balanced middle and has the best grid c1 (215.0). K=4 beats K=7 on aggregate at every concurrency. Avoid TP=4 with the DFlash2 drafter past c32: its 8 full-attention KV tensors cut the KV pool ~6× vs TP8 (311,999 vs 1,811,949 tokens) and force a structural preemption/recompute TTFT cliff. Uncontended TP=4 with MTP N=2 is the best single-stream shape measured on this rig (c1 193.18, 1,292 @c256). With any drafter, cap --max-model-len ≤ 131071: the speculative path faults on 131K-token prefills (without speculation they run). Image ghcr.io/canada-quant/vllm-glm53-flash-h100:v2-w4a16-dflash2, cold boot ≈ 12 min. Full grids: BENCHMARKS.md.
RTX PRO 6000, TP=4 — same command with the SM120 image and --max-num-seqs 64 --max-num-batched-tokens 8192 --kv-cache-dtype fp8 --enable-prefix-caching. fp8 KV is required at 262K on 96 GB cards. Keep MTP on at every concurrency here: it adds +70% at c1 and +48% at c32.
2× DGX Spark, TP=2 — prebuilt image and one-command launcher in canada-quant/vllm-glm53-flash-sm121; the launcher also ships in the drafter repo. Drafter: canada-quant/GLM-5.3-Flash-DFlash2-G, our self-trained DFlash2 drafter (Apache-2.0; 3.676 mean acceptance at K=7 on our 500-prompt holdout vs 3.632 for the incoai reference on the same hardware; drop-in successor of -F and -E). Start the worker rank first, wait 25 s, then the head rank.
# on both nodes, rank1 (worker) first, then rank0 (head) 25 s later
# 262K context (production config): 8 GiB fp8 KV -> 366,749-token pool (the launcher default since 2026-10-04;
# older launcher copies defaulted to 3 GiB, which is too small for DFlash2-E/-F/-G, so passing it explicitly is safe)
MAX_MODEL_LEN=262144 KV_CACHE_MEM=8053063680 EAGER=0 GRAPHS=1 bash launch_dflash2_tp2.sh <rank>
# long context with DFlash2-G (validated 2026-09-26/27): 800K, 16 GiB KV pin -> 888,729-token pool
MAX_MODEL_LEN=800000 KV_CACHE_MEM=17179869184 GMU=0.90 EAGER=0 GRAPHS=1 bash launch_dflash2_tp2.sh <rank>- Drafter memory (corrected 2026-10-04). DFlash2-E/-F/-G have 8 full-attention layers that each cache the whole context, about 3× the KV per token of the incoai reference drafter (2,048-token sliding window). Measured pools on this pair: 366,749 tokens at 8 GiB (262K) and 888,729 at 16 GiB (800K) with G, vs 1,360,420 at 9 GiB (1M) with the incoai drafter. 1M on 2× Spark is validated with the incoai drafter only; with G, 9 GiB holds ≈450K tokens. Earlier versions of this card paired G with 1M at 9 GiB.
- Drafter or built-in MTP head? DFlash2-G suits short-context English and code. For long contexts, non-English text or memory headroom, use the checkpoint's own BF16 MTP head (
{"method":"mtp","num_speculative_tokens":3}). Our G recipe slows at depth: 31.4 / 30.7 / 15.3 / 10.0 tok/s at 0 / 4K / 65K / 100K (llama-benchy pp2048/tg128). A community user reported 31.6 / 29.8 / 25.4 / 31.0 tok/s at the same depths with this checkpoint on eugr's b12x build + MTP (1700 MHz GPU clock, different harness; not measured by us), and G 20–30% slower than MTP on non-English text. MTP on our SM121 image is not yet validated by us; a same-pair comparison started 2026-10-04. - Notes: 7 speculative tokens is the drafter's trained and measured setting (k=5 also boots, not benchmarked). Confirm the boot log shows the mask-embedding load (
mask_token_id 154856); keep single prompts ≤ ~310K tokens; stop withdocker stop -t 30, neverrm -f. Cold boot is 6–10 minutes.
2× DGX Spark on eugr's b12x build (community-validated, not one of our validated recipes) — this checkpoint needs a fused-attention ignore-list fix there (eugr/spark-vllm-docker#403): bind-mount a config.json copy whose quantization_config.ignore adds re:.*self_attn\..*, re:.*\.in_proj_qkvgfab.* and re:.*in_proj.* (safe: nothing under self_attn is quantized), and pass --block-size 256 --moe-backend marlin --attention-backend B12X --linear-backend b12x --kv-cache-dtype fp8 with --speculative-config '{"method":"mtp","num_speculative_tokens":3,"attention_backend":"B12X"}'. We have not tested our DFlash2 drafters on that build (its b12x attention lacks a full-window non-causal path, and the drafter cache stays BF16).
Quality
On RTX PRO 6000 the raw AIME 2025 score trails the H100 result: ≈63% of the gap is a budget wall (11.7–14.2% empty answers) and ≈37% is SM120 kernel numerics. A zero-cost commit hook closes it but is not part of the published recipes. The vision tower is BF16 passthrough and was not covered by the text-only calibration, so MMMU and OCRBench are this artifact's own vision baseline (greedy, seed 42, measured 2026-09-28). Details: BENCHMARKS.md.
Benchmarks summary
Output tok/s, thinking ON, measured by the authors. Competitor comparisons against NVFP4 checkpoints were removed on 2026-10-04: earlier tables mixed two different NVFP4 checkpoints under one label and compared against NVIDIA's without its MTP head. A clean same-rig comparison will be published when it's done.
Single-stream decode is insensitive to KV length up to ≥486K on RTX PRO 6000; at batch, long-KV decode plateaus at ~2–4 tok/s per stream and long prefills serialize at a ~6–8.5K tok/s aggregate ceiling. All grids and protocols: BENCHMARKS.md.
Hardware & context limits
Each row is the largest context serving-validated on that configuration.
Known issues
- `config.json` (2026-09-08): vLLM matches
quantization_config.ignoreagainst its own fused module names, so the ignore list now carries both the HF and vLLM spellings plusre:.*\.layers\.45\..*for the MTP head. Older 765-entry copies fail at load. Weights unchanged. - DFlash2 admission wedge (SM90 research stack only, MTP recipes unaffected): with the DFlash2 drafter at block size 2304, prompts above ~15.5K tokens are never admitted. A fix was validated to 256K prompts; block size 1536 avoids it. Filed as vllm-project/vllm#55800.
- Marlin no-split-K path on SM121: deterministic illegal memory access at M=256 when forcing
split_k=1; the stock heuristic used in serving is clean. Filed as vllm-project/vllm#56064.
Quantization details
Build gates, all passing: exactly 36,288 packed tensors and nothing quantized outside routed experts; vision key set 348/348 identical to source; MTP layer present; zero dtype drift vs source; no collapsed expert scales. Loads with transformers ≥ 5.16; text generation and image captioning smoke tests pass.
Citation & license
@misc{canada_quant_glm53_flash_w4a16_mtp,
title = {GLM-5.3-Flash W4A16 (INT4) + BF16 MTP},
author = {canada-quant},
year = {2026},
url = {https://huggingface.co/canada-quant/GLM-5.3-Flash-W4A16-MTP}
}MIT, inherited from the base model. Follow the base model's usage terms.
Built, benchmarked and documented with the [Digby.ai](https://digby.ai) coding harness, developed by [CQL.ca](https://cql.ca).
