soyrsoyr/pr3118-mtp-e2e-validation
PR3118 MTP validation 19 active model repositories, matching the PR metrics table: seven earlier Qwen dense variants plus three Qwen MTP pipeline checkpoints, one Qwen MoE preservation variant, six Nemotron variants, and the full pretrained GLM-5.3 Flash and DSA checkpoints. Full-model MTP serving smoke results: GLM-5.3 (DSA): 58/70 draft tokens accepted (82.86%); one 128-token completion, four B200s, TP=4, vLLM 0.28.0, one speculative MTP token. Request, response, and… See the full description on the dataset page: https://huggingface.co/datasets/soyrsoyr/pr3118-mtp-e2e-validation.
PR3118 MTP validation
19 active model repositories, matching the PR metrics table: seven earlier Qwen dense variants plus three Qwen MTP pipeline checkpoints, one Qwen MoE preservation variant, six Nemotron variants, and the full pretrained GLM-5.3 Flash and DSA checkpoints.
Full-model MTP serving smoke results:
- GLM-5.3 (DSA): 58/70 draft tokens accepted (82.86%); one 128-token completion, four B200s, TP=4, vLLM 0.28.0, one speculative MTP token. Request, response, and counters.
- GLM-5.3-Flash: 116/137 accepted (84.67%); four 64-token completions, two B200s, TP=2, one speculative MTP token. Runtime evidence.
Both use data-free NVFP4A16 backbone MLP and MTP quantization. DSA conversion used model_free_ptq for the backbone followed by the PR MTP save integration; the linked model card documents two run-local MFPTQ adjustments. Flash exercised oneshot(output_dir=...). These generation checks establish serving and MTP execution, without measuring model quality or performance. The DSA request included first-request kernel compilation.
Qwen3.8-27B copy → FP8 → BF16: all three stages passed at PR commit 3dd54c307dd68074c8af055cb5017093a4880e54. Original BF16 MTP copied exactly with None; FP8_DYNAMIC quantized seven projections; explicit BF16 matched reference dequantization. The FP8 backbone was reused unchanged, exercising the production MTP save helper without rerunning full-model oneshot. Each checkpoint passed one 64-token completion on one B200 (TP=1), vLLM 0.28.0, with one MTP token: 30/34 (88.24%), 29/34 (85.29%), and 30/33 (90.91%) accepted, respectively. Conversion checks, requests, responses, and counters.
Active model index
Links pin the tested artifacts. Current statuses follow the PR table; earlier H100 counters are preserved separately in historical_runtime within the manifest and in the original reports.
Qwen MoE coverage is limited to unquantized MTP preservation. Nemotron MXFP4 and MXFP4-from-FP8 remain runtime failures in vLLM's MoE scale preparation/reshape path; fixing those failures is deferred to upstream vLLM. NVFP4A16 means weight-only FP4 with 16-bit activations, rather than calibrated NVFP4 W4A4.
Reproducing checks
hf-uploads.json pins all 19 active models; final-model-manifest.json records their source revisions, save paths, runtime evidence, and limitations. The existing pr3118_hf_smoke.py is a single-GPU runner for the earlier variants. Its --all-mxfp4 selection now contains four active Qwen/Nemotron variants; the two Nemotron failures above remain expected. It requires positive draft counters, without requiring nonzero acceptance.
Full GLM checkpoints require multiple GPUs. The latest DSA smoke used an activated serving environment and:
canhazgpu run --gpus 4 --timeout 90m --wait 15m -- \
vllm serve soyrsoyr/GLM-5.3-NVFP4A16-MTP-PR3118 \
--revision be5d8973114ef4425b90247f477bc70aaa60cc6d \
--tokenizer-revision be5d8973114ef4425b90247f477bc70aaa60cc6d \
--tensor-parallel-size 4 --dtype bfloat16 --max-model-len 1024 \
--max-num-seqs 2 --max-num-batched-tokens 1024 \
--gpu-memory-utilization 0.85 --enforce-eager --block-size 256 \
--host 127.0.0.1 --port 47647 \
--speculative-config '{"method":"mtp","num_speculative_tokens":1}'Send the request recorded in the runtime evidence to /v1/completions. Compare vllm:spec_decode_num_draft_tokens_total and vllm:spec_decode_num_accepted_tokens_total at /metrics before and after the request. Acceptance is accepted-token delta divided by draft-token delta. The full DSA smoke required both deltas to be positive.
Historical evidence
The DSA and Flash tiny fixtures have been retired from the active model index, collection, and runner manifest. retired-models.json records their replacements. Original per-fixture reports remain under reports/ for provenance; their model links may no longer resolve and their results are not current pretrained-model validation.
public-access-verification.json, hf-roundtrip-smoke.json, and hf-roundtrip-clean-smoke.json are historical records of earlier uploads and tiny-fixture runs. They do not define the active model set. Earlier test counts and environment details remain in repository history and the corresponding reports. Current coverage is scoped to each PR-table row, rather than every architecture/format combination.
Upstream checkpoint licenses continue to apply. This validation index grants no additional rights to source models.
