experimentalmachines/Qwen3-4B-ExecuTorch
Qwen3-4B for ExecuTorch
ExecuTorch exports of Qwen/Qwen3-4B (revision 1cfa9a720891) for on-device inference with the openweights Android app or any ExecuTorch 1.4.0 runtime.
Files
Windows in this repo: XNNPACK (CPU) at 2k to 32k; Vulkan (GPU) at 2k to 32k; MediaTek NeuroPilot MT6991 (Dimensity 9400) at 2k to 4k. The window is fixed inside the file: the runtime allocates the whole KV cache at load, so pick the largest window the device can hold (fits_phone_budget in each folder's config.json is the estimate against a 5 GB budget).
MediaTek folders also hold the token embedding table the NeuroPilot runner reads from disk, shared by every window: `mtk/mt6991/Qwen3-4B-neuropilot-embedding-fp32.bin`.
Tokenizer: `tokenizer.json`, copied unchanged from the source repo. Each backend folder has a config.json listing every window as a variant with the metadata the .pte reports, and an export-report-<window>.json per file with the full export record.
Memory
- XNNPACK (CPU) at 2,048 tokens: the KV cache costs 294,912 bytes per token (fp32), 603,979,776 bytes for the whole window, allocated in full when the model loads.
- XNNPACK (CPU) at 4,096 tokens: the KV cache costs 294,912 bytes per token (fp32), 1,207,959,552 bytes for the whole window, allocated in full when the model loads.
- XNNPACK (CPU) at 8,192 tokens: the KV cache costs 294,912 bytes per token (fp32), 2,415,919,104 bytes for the whole window, allocated in full when the model loads.
- XNNPACK (CPU) at 16,384 tokens: the KV cache costs 294,912 bytes per token (fp32), 4,831,838,208 bytes for the whole window, allocated in full when the model loads.
- XNNPACK (CPU) at 32,768 tokens: the KV cache costs 294,912 bytes per token (fp32), 9,663,676,416 bytes for the whole window, allocated in full when the model loads.
- Vulkan (GPU) at 2,048 tokens: the KV cache costs 294,912 bytes per token (fp32), 603,979,776 bytes for the whole window, allocated in full when the model loads.
- Vulkan (GPU) at 4,096 tokens: the KV cache costs 294,912 bytes per token (fp32), 1,207,959,552 bytes for the whole window, allocated in full when the model loads.
- Vulkan (GPU) at 8,192 tokens: the KV cache costs 294,912 bytes per token (fp32), 2,415,919,104 bytes for the whole window, allocated in full when the model loads.
- Vulkan (GPU) at 16,384 tokens: the KV cache costs 294,912 bytes per token (fp32), 4,831,838,208 bytes for the whole window, allocated in full when the model loads.
- Vulkan (GPU) at 32,768 tokens: the KV cache costs 294,912 bytes per token (fp32), 9,663,676,416 bytes for the whole window, allocated in full when the model loads.
How it was made
- XNNPACK (CPU) any arm64: ExecuTorch 1.4.0
export_llm: 8-bit dynamic activations and 4-bit weights in groups of 32, int8 per-channel embeddings, XNNPACK with extended ops, prefill chunk 2048, fp32 KV cache. Built by run 1. - Vulkan (GPU) any arm64: ExecuTorch 1.4.0
export_llm: 8-bit dynamic activations and 4-bit weights in groups of 32, int8 per-channel embeddings, the Vulkan delegate, prefill chunk 2048, fp32 KV cache. Built by run 1. - MediaTek NeuroPilot MT6991 (Dimensity 9400): ExecuTorch 1.4.0 MediaTek LLM export (
examples/mediatek,qwen.py): A16W8 (16-bit activations, 8-bit weights) calibrated on MediaTek'salpaca.txtprompts (9 of 9) in the qwen3.json chat template, cut into 4 chunks, with a 128-token prompt graph and a one-token generation graph over a 2048-token cache, compiled with MediaTek NeuroPilot Express SDK (mtkconverter 8.13.0+public, mtkneuron 8.2.23) for MT6991 (Dimensity 9400). Built by run 1. - MediaTek NeuroPilot MT6991 (Dimensity 9400): ExecuTorch 1.4.0 MediaTek LLM export (
examples/mediatek,qwen.py): A16W8 (16-bit activations, 8-bit weights) calibrated on MediaTek'salpaca.txtprompts (9 of 9) in the qwen3.json chat template, cut into 4 chunks, with a 128-token prompt graph and a one-token generation graph over a 4096-token cache, compiled with MediaTek NeuroPilot Express SDK (mtkconverter 8.13.0+public, mtkneuron 8.2.23) for MT6991 (Dimensity 9400). Built by run 1.
License
A quantized derivative of Qwen/Qwen3-4B, distributed under the same terms (apache-2.0). The upstream license files are included unchanged: `LICENSE`.
The mtk/ folders hold model binaries compiled with the MediaTek NeuroPilot Express SDK 8.0.8-build20250925 (mtkconverter 8.13.0+public, mtkneuron 8.2.23) from MediaTek Inc., used under MediaTek's license terms for that SDK. No MediaTek SDK or runtime library is included. They run on MediaTek's LLM runner from ExecuTorch (examples/mediatek/executor_runner) with the device's NeuroPilot runtime; the settings it needs are in each folder's config.json, under each variant's runner.
