tomhu/Qwen3-Coder-Next-MixedCode-Prune70-BF16
0290
Qwen3-Coder-Next Mixed-Code Prune-70 (BF16)
Experimental expert-pruned checkpoint derived from Qwen/Qwen3-Coder-Next.
Model details
- Calibration source: full HumanEval, MBPP, and BigCodeBench code datasets
- Pruning method: calibration-score-based expert pruning
- Experts per MoE layer: 512 -> 154 (69.92% removed)
- Active experts per token: 10 (unchanged)
- Layers: 48
- Weight dtype: BF16
- Weight payload: 51,166,017,024 bytes (~25.6B BF16 parameters)
- Context configuration: inherited from the base model
Experts are renumbered after pruning. expert_mapping.csv records the retained source expert for every layer and new expert ID.
Evaluation
On the associated SWE-bench Verified evaluation, this checkpoint resolved 296/500 instances (59.2%).
Usage
Use a recent Transformers, vLLM, or SGLang release with Qwen3-Next support. This is an experimental research checkpoint; validate it for your workload before deployment.
License and attribution
This derivative checkpoint follows the Apache-2.0 license of Qwen/Qwen3-Coder-Next.
