Team Ai
Modelpublic

experimentalmachines/Llama-3.2-3B-Instruct-ExecuTorch

sourceHugging Facellama3.2updated 2d agoView on Hugging Face
0likes338downloads
Model Card

Built with Llama

Llama-3.2-3B-Instruct for ExecuTorch

ExecuTorch exports of meta-llama/Llama-3.2-3B-Instruct (revision 0cb88a4f764b) for on-device inference with the openweights Android app or any ExecuTorch 1.4.0 runtime.

Files

Windows in this repo: XNNPACK (CPU) at 2k to 32k; Vulkan (GPU) at 2k to 32k; Qualcomm QNN (HTP) SM8750 (Snapdragon 8 Elite) at 2k to 16k; MediaTek NeuroPilot MT6991 (Dimensity 9400) at 2k to 8k. The window is fixed inside the file: the runtime allocates the whole KV cache at load, so pick the largest window the device can hold (fits_phone_budget in each folder's config.json is the estimate against a 5 GB budget).

BackendTargetFileWindowSizeSmoke test
XNNPACK (CPU)any arm64`xnnpack/Llama-3.2-3B-Instruct-8da4w-gptq-2k.pte`2,048 tokens2.21 GBpassed ("Paris")
XNNPACK (CPU)any arm64`xnnpack/Llama-3.2-3B-Instruct-8da4w-gptq-4k.pte`4,096 tokens2.21 GBpassed ("Paris")
XNNPACK (CPU)any arm64`xnnpack/Llama-3.2-3B-Instruct-8da4w-gptq-8k.pte`8,192 tokens2.21 GBpassed ("Paris")
XNNPACK (CPU)any arm64`xnnpack/Llama-3.2-3B-Instruct-8da4w-gptq-16k.pte`16,384 tokens2.22 GBpassed ("Paris")
XNNPACK (CPU)any arm64`xnnpack/Llama-3.2-3B-Instruct-8da4w-gptq-32k.pte`32,768 tokens2.24 GBpassed ("Paris")
Vulkan (GPU)any arm64`vulkan/Llama-3.2-3B-Instruct-vulkan-8da4w-2k.pte`2,048 tokens2.82 GBstructure checked (no host NPU runtime)
Vulkan (GPU)any arm64`vulkan/Llama-3.2-3B-Instruct-vulkan-8da4w-4k.pte`4,096 tokens2.83 GBstructure checked (no host NPU runtime)
Vulkan (GPU)any arm64`vulkan/Llama-3.2-3B-Instruct-vulkan-8da4w-8k.pte`8,192 tokens2.85 GBstructure checked (no host NPU runtime)
Vulkan (GPU)any arm64`vulkan/Llama-3.2-3B-Instruct-vulkan-8da4w-16k.pte`16,384 tokens2.89 GBstructure checked (no host NPU runtime)
Vulkan (GPU)any arm64`vulkan/Llama-3.2-3B-Instruct-vulkan-8da4w-32k.pte`32,768 tokens2.97 GBstructure checked (no host NPU runtime)
Qualcomm QNN (HTP)SM8750 (Snapdragon 8 Elite)`qnn/sm8750/Llama-3.2-3B-Instruct-qnn-hybrid-2k.pte`2,048 tokens2.80 GBstructure checked (no host NPU runtime)
Qualcomm QNN (HTP)SM8750 (Snapdragon 8 Elite)`qnn/sm8750/Llama-3.2-3B-Instruct-qnn-hybrid-4k.pte`4,096 tokens2.81 GBstructure checked (no host NPU runtime)
Qualcomm QNN (HTP)SM8750 (Snapdragon 8 Elite)`qnn/sm8750/Llama-3.2-3B-Instruct-qnn-hybrid-8k.pte`8,192 tokens2.81 GBstructure checked (no host NPU runtime)
Qualcomm QNN (HTP)SM8750 (Snapdragon 8 Elite)`qnn/sm8750/Llama-3.2-3B-Instruct-qnn-hybrid-16k.pte`16,384 tokens2.84 GBstructure checked (no host NPU runtime)
MediaTek NeuroPilotMT6991 (Dimensity 9400)`mtk/mt6991/Llama-3.2-3B-Instruct-neuropilot-a16w8-2k-chunk1of4.pte`2,048 tokens0.84 GBstructure checked (no host NPU runtime)
MediaTek NeuroPilotMT6991 (Dimensity 9400)`mtk/mt6991/Llama-3.2-3B-Instruct-neuropilot-a16w8-2k-chunk2of4.pte`2,048 tokens0.84 GBstructure checked (no host NPU runtime)
MediaTek NeuroPilotMT6991 (Dimensity 9400)`mtk/mt6991/Llama-3.2-3B-Instruct-neuropilot-a16w8-2k-chunk3of4.pte`2,048 tokens0.84 GBstructure checked (no host NPU runtime)
MediaTek NeuroPilotMT6991 (Dimensity 9400)`mtk/mt6991/Llama-3.2-3B-Instruct-neuropilot-a16w8-2k-chunk4of4.pte`2,048 tokens1.23 GBstructure checked (no host NPU runtime)
MediaTek NeuroPilotMT6991 (Dimensity 9400)`mtk/mt6991/Llama-3.2-3B-Instruct-neuropilot-a16w8-4k-chunk1of4.pte`4,096 tokens0.84 GBstructure checked (no host NPU runtime)
MediaTek NeuroPilotMT6991 (Dimensity 9400)`mtk/mt6991/Llama-3.2-3B-Instruct-neuropilot-a16w8-4k-chunk2of4.pte`4,096 tokens0.84 GBstructure checked (no host NPU runtime)
MediaTek NeuroPilotMT6991 (Dimensity 9400)`mtk/mt6991/Llama-3.2-3B-Instruct-neuropilot-a16w8-4k-chunk3of4.pte`4,096 tokens0.84 GBstructure checked (no host NPU runtime)
MediaTek NeuroPilotMT6991 (Dimensity 9400)`mtk/mt6991/Llama-3.2-3B-Instruct-neuropilot-a16w8-4k-chunk4of4.pte`4,096 tokens1.23 GBstructure checked (no host NPU runtime)
MediaTek NeuroPilotMT6991 (Dimensity 9400)`mtk/mt6991/Llama-3.2-3B-Instruct-neuropilot-a16w8-8k-chunk1of4.pte`8,192 tokens0.84 GBstructure checked (no host NPU runtime)
MediaTek NeuroPilotMT6991 (Dimensity 9400)`mtk/mt6991/Llama-3.2-3B-Instruct-neuropilot-a16w8-8k-chunk2of4.pte`8,192 tokens0.84 GBstructure checked (no host NPU runtime)
MediaTek NeuroPilotMT6991 (Dimensity 9400)`mtk/mt6991/Llama-3.2-3B-Instruct-neuropilot-a16w8-8k-chunk3of4.pte`8,192 tokens0.84 GBstructure checked (no host NPU runtime)
MediaTek NeuroPilotMT6991 (Dimensity 9400)`mtk/mt6991/Llama-3.2-3B-Instruct-neuropilot-a16w8-8k-chunk4of4.pte`8,192 tokens1.23 GBstructure checked (no host NPU runtime)

MediaTek folders also hold the token embedding table the NeuroPilot runner reads from disk, shared by every window: `mtk/mt6991/Llama-3.2-3B-Instruct-neuropilot-embedding-fp32.bin`.

Tokenizer: `tokenizer.json`, copied unchanged from the source repo. Each backend folder has a config.json listing every window as a variant with the metadata the .pte reports, and an export-report-<window>.json per file with the full export record.

Memory

  • —XNNPACK (CPU) at 2,048 tokens: the KV cache costs 229,376 bytes per token (fp32), 469,762,048 bytes for the whole window, allocated in full when the model loads.
  • —XNNPACK (CPU) at 4,096 tokens: the KV cache costs 229,376 bytes per token (fp32), 939,524,096 bytes for the whole window, allocated in full when the model loads.
  • —XNNPACK (CPU) at 8,192 tokens: the KV cache costs 229,376 bytes per token (fp32), 1,879,048,192 bytes for the whole window, allocated in full when the model loads.
  • —XNNPACK (CPU) at 16,384 tokens: the KV cache costs 229,376 bytes per token (fp32), 3,758,096,384 bytes for the whole window, allocated in full when the model loads.
  • —XNNPACK (CPU) at 32,768 tokens: the KV cache costs 229,376 bytes per token (fp32), 7,516,192,768 bytes for the whole window, allocated in full when the model loads.
  • —Vulkan (GPU) at 2,048 tokens: the KV cache costs 229,376 bytes per token (fp32), 469,762,048 bytes for the whole window, allocated in full when the model loads.
  • —Vulkan (GPU) at 4,096 tokens: the KV cache costs 229,376 bytes per token (fp32), 939,524,096 bytes for the whole window, allocated in full when the model loads.
  • —Vulkan (GPU) at 8,192 tokens: the KV cache costs 229,376 bytes per token (fp32), 1,879,048,192 bytes for the whole window, allocated in full when the model loads.
  • —Vulkan (GPU) at 16,384 tokens: the KV cache costs 229,376 bytes per token (fp32), 3,758,096,384 bytes for the whole window, allocated in full when the model loads.
  • —Vulkan (GPU) at 32,768 tokens: the KV cache costs 229,376 bytes per token (fp32), 7,516,192,768 bytes for the whole window, allocated in full when the model loads.

How it was made

  • —XNNPACK (CPU) any arm64: ExecuTorch 1.4.0 export_llm: 8-bit dynamic activations and 4-bit weights in groups of 32, int8 per-channel embeddings, XNNPACK with extended ops, prefill chunk 2048, fp32 KV cache. Built by run 1.
  • —Vulkan (GPU) any arm64: ExecuTorch 1.4.0 export_llm: 8-bit dynamic activations and 4-bit weights in groups of 32, int8 per-channel embeddings, the Vulkan delegate, prefill chunk 2048, fp32 KV cache. Built by run 1.
  • —Qualcomm QNN (HTP) SM8750 (Snapdragon 8 Elite): ExecuTorch 1.5.1 Qualcomm static LLM (examples/qualcomm/oss_scripts/llama, --decoder_model llama3_2-3b_instruct): the quantization recipe ExecuTorch registers for this model, calibrated on wikitext (1 sample), hybrid prefill (128 tokens per step) and decode graphs compiled with QAIRT 2.37.0.250724 for SM8750 (Snapdragon 8 Elite). Built by run 1.
  • —MediaTek NeuroPilot MT6991 (Dimensity 9400): ExecuTorch 1.4.0 MediaTek LLM export (examples/mediatek, llama.py): A16W8 (16-bit activations, 8-bit weights) calibrated on MediaTek's alpaca.txt prompts (9 of 9) in the llama3.json chat template, cut into 4 chunks, with a 128-token prompt graph and a one-token generation graph over a 2048-token cache, compiled with MediaTek NeuroPilot Express SDK (mtkconverter 8.13.0+public, mtkneuron 8.2.23) for MT6991 (Dimensity 9400). Built by run 1.
  • —MediaTek NeuroPilot MT6991 (Dimensity 9400): ExecuTorch 1.4.0 MediaTek LLM export (examples/mediatek, llama.py): A16W8 (16-bit activations, 8-bit weights) calibrated on MediaTek's alpaca.txt prompts (9 of 9) in the llama3.json chat template, cut into 4 chunks, with a 128-token prompt graph and a one-token generation graph over a 4096-token cache, compiled with MediaTek NeuroPilot Express SDK (mtkconverter 8.13.0+public, mtkneuron 8.2.23) for MT6991 (Dimensity 9400). Built by run 1.
  • —MediaTek NeuroPilot MT6991 (Dimensity 9400): ExecuTorch 1.4.0 MediaTek LLM export (examples/mediatek, llama.py): A16W8 (16-bit activations, 8-bit weights) calibrated on MediaTek's alpaca.txt prompts (9 of 9) in the llama3.json chat template, cut into 4 chunks, with a 128-token prompt graph and a one-token generation graph over a 8192-token cache, compiled with MediaTek NeuroPilot Express SDK (mtkconverter 8.13.0+public, mtkneuron 8.2.23) for MT6991 (Dimensity 9400). Built by run 1.

License

A quantized derivative of meta-llama/Llama-3.2-3B-Instruct, distributed under the same terms (llama3.2). The upstream license files are included unchanged: `LICENSE.txt`, `USE_POLICY.md`. See `NOTICE` for the attribution the license requires.

The qnn/ folders hold QNN HTP context binaries compiled with the Qualcomm AI Runtime SDK (QAIRT) 2.37.0.250724 from Qualcomm Technologies, Inc., used under its AI Stack License. No Qualcomm SDK or runtime library is included; running them needs the matching QNN runtime (for example executorch-android-qnn 1.5.1, which depends on qnn-runtime 2.37.0).

The mtk/ folders hold model binaries compiled with the MediaTek NeuroPilot Express SDK 8.0.8-build20250925 (mtkconverter 8.13.0+public, mtkneuron 8.2.23) from MediaTek Inc., used under MediaTek's license terms for that SDK. No MediaTek SDK or runtime library is included. They run on MediaTek's LLM runner from ExecuTorch (examples/mediatek/executor_runner) with the device's NeuroPilot runtime; the settings it needs are in each folder's config.json, under each variant's runner.