Team Ai
Modelpublic

tomhu/Qwen3-Coder-Next-MixedCode-Prune70-BF16

sourceHugging Faceapache-2.0updated 20d agoView on Hugging Face
0likes290downloads
Model Card

Qwen3-Coder-Next Mixed-Code Prune-70 (BF16)

Experimental expert-pruned checkpoint derived from Qwen/Qwen3-Coder-Next.

Model details

  • —Calibration source: full HumanEval, MBPP, and BigCodeBench code datasets
  • —Pruning method: calibration-score-based expert pruning
  • —Experts per MoE layer: 512 -> 154 (69.92% removed)
  • —Active experts per token: 10 (unchanged)
  • —Layers: 48
  • —Weight dtype: BF16
  • —Weight payload: 51,166,017,024 bytes (~25.6B BF16 parameters)
  • —Context configuration: inherited from the base model

Experts are renumbered after pruning. expert_mapping.csv records the retained source expert for every layer and new expert ID.

Evaluation

On the associated SWE-bench Verified evaluation, this checkpoint resolved 296/500 instances (59.2%).

Usage

Use a recent Transformers, vLLM, or SGLang release with Qwen3-Next support. This is an experimental research checkpoint; validate it for your workload before deployment.

License and attribution

This derivative checkpoint follows the Apache-2.0 license of Qwen/Qwen3-Coder-Next.