Ttimms/KAT-Coder-V2.5-Dev-REAP-50-NVFP4A16
KAT-Coder-V2.5-Dev REAP-50 NVFP4A16 (16 GB)
REAP expert-pruned (50%) + NVFP4A16 quantized build of `Kwaipilot/KAT-Coder-V2.5-Dev` (69.40 SWE-bench Verified claimed), sized and served to run as a local agentic coding model inside 16 GB of consumer VRAM — 12.45 GiB, RTX 5070 Ti (SM120), vLLM. Built with a router-renormalization fix for this architecture (contributed upstream) and a vision tower stripped of its untrained weights.
Methodology — how this and the sibling 16 GB builds are pruned, quantized, and evaluated (with the failure modes): <https://github.com/t-timms/blackwell-16gb-moe>
SWE-bench Verified: 26/50 = 52.0% resolved, via mini-swe-agent's official bash-only scaffold, at a 49K-token context ceiling — up from 40.0% at the original 32K ceiling tested on this same checkpoint. Still below the 56.4% bar set by Devstral Small (2512) under the same scaffold, but closed most of the gap. 17 of 50 instances produced no usable patch, all from hitting the context ceiling; the run must still be read as context-limited, not as an unconditional capability measurement. See "SWE-bench Verified" below before citing the headline number without that context.
Architecture
graph TD
Base["Kwaipilot/KAT-Coder-V2.5-Dev<br/>Qwen3.5 MoE - 256 experts - ~69 GB bf16 - 69.40 SWE-bench (claimed)"]
subgraph Build ["Build pipeline - RTX 5070 Ti, SM120"]
REAP["REAP expert prune 50%<br/>256 -> 128 experts + router-renormalization fix (upstreamed)"]
Strip["Strip vision tower<br/>declaration + 333 untrained tensors (0.83 GiB)"]
Quant["NVFP4A16 quantize<br/>compressed-tensors - weight-only - data-free - 82 s"]
end
subgraph HF ["Published on Hugging Face"]
A16["REAP-50-NVFP4A16 - 12.45 GiB<br/>default - Marlin NVFP4 kernel"]
W4A4["REAP-50-NVFP4-W4A4<br/>weights+activations - native FP4 kernels"]
GPTQ["REAP-50-NVFP4A16-GPTQ<br/>documented null result"]
GGUF["REAP-50-GGUF<br/>Q4_K_M / Q5_K_M / Q6_K / Q8_0"]
BF16["REAP-50-bf16<br/>pruned source for AWQ / EXL2 / MLX"]
end
subgraph Serve ["Serving & evaluation"]
vLLM["vLLM 0.20.2 - SM120<br/>CUDA graphs PIECEWISE - prefix caching (45x) - 49K ctx"]
Bench["HumanEval+ 89.0% - MBPP+ 90.5%<br/>SWE-bench Verified 52.0% (mini-swe-agent)"]
end
Base --> REAP --> Strip --> Quant --> A16
Strip --> BF16
BF16 -. re-quantize .-> W4A4
BF16 -. re-quantize .-> GPTQ
BF16 -. convert .-> GGUF
A16 --> vLLM --> BenchHighlights
Why 50 percent
Forced by arithmetic on a 16 GB card, not a tuning choice:
Supporting evidence: Half the Experts, All the Code pruned Qwen3.6-35B-A3B — this base model's size-class cousin — at 50% with no statistically detectable loss on its primary code benchmark.
SWE-bench Verified — read before citing the 52.0% figure alone
Prior measurement at this checkpoint's original 32K-token ceiling, step_limit 40: 20/50 = 40.0% (18 ContextWindowExceeded, 9 LimitsExceeded — ran out of agent turns before finishing, a different failure mode than the context ceiling — 20/22 = 90.9% resolved-of-completed). Raising both the context ceiling to 49K and the step limit to 65 (kat_overrides_sota.yaml) moved the headline number from 40.0% to 52.0%, but not primarily by reducing context-window failures — that rate barely moved (17/50 = 34% vs. 18/50 = 36%). The real lever was the step limit: LimitsExceeded failures went from 9 to 0, so 10 more instances (22→32) reached a real completion attempt instead of running out of turns first. Those newly-reachable instances resolve at a lower rate than the ones that were already completing (81.25% resolved-of-completed at 49K vs. 90.9% at 32K — consistent with them being the harder, longer problems that need the extra turns), but enough resolved anyway that the net resolved count still rose (20→26). The scaffold's context ceiling is still the safe limit this card's VRAM budget supports, not a property of the model — Devstral Small averages 86.9 LM calls/instance under the same scaffold, and instances needing more than 49K tokens of context or more than 65 agent turns still fail to submit at all. This is disclosed as a real result, not an excuse: 52.0% is the correct number to cite; the breakdown above is the correct context for interpreting it.
Config note (2026-08-23): kat_overrides_sota.yaml — the config behind the 52.0% figure above — is kept byte-identical to the run that produced it; a fresh clone reproduces this exact result by default. A candidate change (presence_penalty/top_k, completing this base model's own documented sampling recommendation) was tested full-pilot on the same 50 instances and regressed the score to 24/50 = 48.0% — more instances ran out of the turn budget exploring alternatives (LimitsExceeded 0→8) than were saved from the context ceiling (ContextWindowExceeded 17→14). Not promoted; kat_overrides_sota.yaml is unchanged. Kept in the repo as a documented negative result (kat_overrides_sota_presence_penalty.yaml), not deleted.
Prior art and scope of claims
Verified against the Hugging Face Hub on 2026-08-17:
- REAP combined with NVFP4 on
qwen3_5_moealready exists (`rene98c/Qwen3.5-397B-A17B-REAP-28-NVFP4`, March 2026, 23.1K downloads). - REAP on this specific model exists as GGUF (`gbuzhf/KAT-Coder-V2.5-Dev-REAP-205E-MTP-GGUF`).
What is distinct, and all that is claimed: a vLLM-servable KAT-Coder that is genuinely usable in 16 GB, with published SWE-bench Verified, HumanEval+, and MBPP+ numbers and their confidence intervals — none of which the prior art above publishes.
Quantization and pruning details
Usage
Requires vLLM with SM120 support (CUDA graphs are correct on this card for this model, despite past reports of SM120 CUDA-graph issues on other architectures) and native tool calling for agentic use:
vllm serve Ttimms/KAT-Coder-V2.5-Dev-REAP-50-NVFP4A16 \
--served-model-name kat-16gb \
--max-model-len 49152 --max-num-seqs 2 \
--gpu-memory-utilization 0.92 --kv-cache-dtype fp8 \
--reasoning-parser qwen3 \
--enable-auto-tool-choice --tool-call-parser qwen3_xml \
--enable-prefix-caching --max-num-batched-tokens 4096 \
--compilation-config '{"cudagraph_capture_sizes":[1,2],"cudagraph_mode":"PIECEWISE"}' \
--language-model-only--language-model-only is required: the model declares a vision tower it has no trained weights for, and without this flag vLLM profiles a 16K-token image budget through it. --enable-prefix-caching is the single largest agentic lever measured on this model — 45x on replayed history (0.21 s vs 30.74 s for a 13,130-token history). Never set --max-model-len near a measured ceiling: available KV cache swings 0.49–1.41 GiB with host desktop VRAM use, and higher values fail intermittently rather than at startup.
Sampling: temperature=1.0, top_p=0.95 is what the published SWE-bench result used and is the current recommendation. Kwaipilot/KAT-Coder-V2.5-Dev's own model card documents two more params alongside these, presence_penalty=1.5, top_k=20, for Thinking mode. Tested on this checkpoint: on a single instance it suppressed a genuine repetition-loop failure, but a full-pilot test on 50 instances found the net effect regresses the agentic score (see the SWE-bench config note above) — not recommended for agentic use on this checkpoint despite matching the base model's own documented config.
Evaluation
HumanEval+ and MBPP+ via lm-eval-harness / EvalPlus, greedy decoding, instruct framing, Wilson confidence intervals:
KAT-Coder-V2.5-Dev publishes no HumanEval/MBPP/EvalPlus numbers, so there is no published upstream figure to compare these against.
Both figures were re-measured on the released checkpoint itself and reproduced inside their intervals: HumanEval+ 90.9% [85.5, 94.4] and MBPP+ 89.9% [86.5, 92.6], same problem counts, same greedy decoding. The table reports the original measurement. The two differences run in opposite directions (+1.9 pp and -0.6 pp), which is greedy-decoding nondeterminism under vLLM's batching rather than a difference in weights. Reproduce with bash scripts/eval/eval_suite.sh.
SWE-bench Verified via the official swebench.harness.run_evaluation harness against mini-swe-agent bash-only rollouts (scaffold: SWE-bench/experiments v1.17.2 configuration) — see the dedicated section above for the full breakdown and required caveats.
Known limitations
- 49K context window (raised from an earlier 32K ceiling) is a hardware-forced limit, not a design choice — this card's KV-cache budget cannot safely support more. SWE-bench results must still be read as context-limited: 17 of 50 instances in the reported run failed purely from exceeding this ceiling, not from the model failing the task.
- No pruning-ablation baseline measured. The unpruned model is 69.3 GB bf16 and does not fit this hardware; the accuracy cost of pruning itself (independent of quantization) is not isolated here.
- SWE-bench Verified no longer accepts leaderboard submissions outside academia — these numbers are self-reported and independently reproducible from the released evaluation scripts, not a leaderboard entry.
W4A4: an alternative quantization strategy
A follow-up build, `Ttimms/KAT-Coder-V2.5-Dev-REAP-50-NVFP4-W4A4`, re-quantizes this same pruned checkpoint to NVFP4 (W4A4 — weights and activations to 4 bits) to reach native FP4 tensor-core kernels instead of this model's Marlin dequant-to-bf16 fallback. Measured properly (5 interleaved runs per arm, median + range, under both isolated eager mode and the same PIECEWISE CUDA-graph configuration this model actually serves with): W4A4 decodes at 119.2 tok/s vs this checkpoint's 142.5 tok/s (0.84x) — both figures from that interleaved A/B harness, which runs slightly under the 149.5 tok/s headline above (measured separately at 512/256); the difference is measurement configuration, not two different builds — and with no accuracy difference we can claim in either direction. The accuracy suite is not deterministic: re-running the published W4A4 checkpoint five times with byte-identical arguments spans 1.85–4.27 pp per task, which is wider than every A16-vs-W4A4 gap. Medians over 5 draws for W4A4 are 93.90 / 89.63 / 89.68 (HumanEval / HumanEval+ / MBPP+) against this checkpoint's single-draw 95.7 / 90.9 / 89.9. An earlier version of this card described a "mixed accuracy picture" with W4A4 ahead on MBPP+; that was single-draw noise and is withdrawn. This checkpoint is faster on this hardware and is what we default to; the W4A4 build is a legitimate alternative if the native FP4 execution path matters more for your use case. Only single-stream (batch=1) decode was measured. Full writeup: `ROADMAP.md`'s RESULT section in the release repo.
GPTQ: a tested rounding variant, not a recommended alternative
A third build, `Ttimms/KAT-Coder-V2.5-Dev-REAP-50-NVFP4A16-GPTQ`, re-quantizes the same pruned checkpoint to the identical NVFP4A16 scheme using GPTQ rounding instead of this release's round-to-nearest (RTN) — same size (12.4512 GiB, exact match), same kernel, only the quantization algorithm differs. Unlike W4A4 above, this is not published as an alternative worth choosing — a paired McNemar test found no statistically significant difference from this release on either HumanEval+ (p=0.68) or MBPP+ (p=0.81), and there's no distinct execution-path reason to prefer it the way there is for W4A4. It's published for completeness and independent verification of that null result, not as a recommendation. Full writeup on that card and in ROADMAP.md's RESULT section in the release repo.
License
Apache 2.0, inherited from the base model Kwaipilot/KAT-Coder-V2.5-Dev and matching the reap and llm-compressor toolchains used to build this checkpoint.
Citation
This checkpoint is derived from Kwaipilot/KAT-Coder-V2.5-Dev. If you use it, please cite the upstream technical report:
@misc{katcoder_v25_2026,
title={{KAT-Coder-V2.5 Technical Report}},
author={{KwaiKAT Team}},
year={2026},
month={July},
eprint={2607.05471},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/pdf/2607.05471}
}