Team Ai
Modelpublic

experimentalmachines/Qwen2.5-3B-Instruct-ExecuTorch

sourceHugging Faceotherupdated 7d agoView on Hugging Face
0likes146downloads
Model Card

Qwen2.5-3B-Instruct for ExecuTorch

ExecuTorch exports of Qwen/Qwen2.5-3B-Instruct (revision aa8e72537993) for on-device inference with the openweights Android app or any ExecuTorch 1.4.0 runtime.

Files

Windows in this repo: XNNPACK (CPU) at 2k to 32k; Vulkan (GPU) at 2k to 32k. The window is fixed inside the file: the runtime allocates the whole KV cache at load, so pick the largest window the device can hold (fits_phone_budget in each folder's config.json is the estimate against a 5 GB budget).

BackendTargetFileWindowSizeSmoke test
XNNPACK (CPU)any arm64`xnnpack/Qwen2.5-3B-Instruct-8da4w-gptq-2k.pte`2,048 tokens2.05 GBpassed ("Paris")
XNNPACK (CPU)any arm64`xnnpack/Qwen2.5-3B-Instruct-8da4w-gptq-4k.pte`4,096 tokens2.06 GBpassed ("Paris")
XNNPACK (CPU)any arm64`xnnpack/Qwen2.5-3B-Instruct-8da4w-gptq-8k.pte`8,192 tokens2.07 GBpassed ("Paris")
XNNPACK (CPU)any arm64`xnnpack/Qwen2.5-3B-Instruct-8da4w-gptq-16k.pte`16,384 tokens2.08 GBpassed ("Paris")
XNNPACK (CPU)any arm64`xnnpack/Qwen2.5-3B-Instruct-8da4w-gptq-32k.pte`32,768 tokens2.12 GBpassed ("Paris")
Vulkan (GPU)any arm64`vulkan/Qwen2.5-3B-Instruct-vulkan-8da4w-2k.pte`2,048 tokens2.63 GBstructure checked (no host NPU runtime)
Vulkan (GPU)any arm64`vulkan/Qwen2.5-3B-Instruct-vulkan-8da4w-4k.pte`4,096 tokens2.64 GBstructure checked (no host NPU runtime)
Vulkan (GPU)any arm64`vulkan/Qwen2.5-3B-Instruct-vulkan-8da4w-8k.pte`8,192 tokens2.65 GBstructure checked (no host NPU runtime)
Vulkan (GPU)any arm64`vulkan/Qwen2.5-3B-Instruct-vulkan-8da4w-16k.pte`16,384 tokens2.68 GBstructure checked (no host NPU runtime)
Vulkan (GPU)any arm64`vulkan/Qwen2.5-3B-Instruct-vulkan-8da4w-32k.pte`32,768 tokens2.73 GBstructure checked (no host NPU runtime)

Tokenizer: `tokenizer.json`, copied unchanged from the source repo. Each backend folder has a config.json listing every window as a variant with the metadata the .pte reports, and an export-report-<window>.json per file with the full export record.

Memory

  • —XNNPACK (CPU) at 2,048 tokens: the KV cache costs 73,728 bytes per token (fp32), 150,994,944 bytes for the whole window, allocated in full when the model loads.
  • —XNNPACK (CPU) at 4,096 tokens: the KV cache costs 73,728 bytes per token (fp32), 301,989,888 bytes for the whole window, allocated in full when the model loads.
  • —XNNPACK (CPU) at 8,192 tokens: the KV cache costs 73,728 bytes per token (fp32), 603,979,776 bytes for the whole window, allocated in full when the model loads.
  • —XNNPACK (CPU) at 16,384 tokens: the KV cache costs 73,728 bytes per token (fp32), 1,207,959,552 bytes for the whole window, allocated in full when the model loads.
  • —XNNPACK (CPU) at 32,768 tokens: the KV cache costs 73,728 bytes per token (fp32), 2,415,919,104 bytes for the whole window, allocated in full when the model loads.
  • —Vulkan (GPU) at 2,048 tokens: the KV cache costs 73,728 bytes per token (fp32), 150,994,944 bytes for the whole window, allocated in full when the model loads.
  • —Vulkan (GPU) at 4,096 tokens: the KV cache costs 73,728 bytes per token (fp32), 301,989,888 bytes for the whole window, allocated in full when the model loads.
  • —Vulkan (GPU) at 8,192 tokens: the KV cache costs 73,728 bytes per token (fp32), 603,979,776 bytes for the whole window, allocated in full when the model loads.
  • —Vulkan (GPU) at 16,384 tokens: the KV cache costs 73,728 bytes per token (fp32), 1,207,959,552 bytes for the whole window, allocated in full when the model loads.
  • —Vulkan (GPU) at 32,768 tokens: the KV cache costs 73,728 bytes per token (fp32), 2,415,919,104 bytes for the whole window, allocated in full when the model loads.

How it was made

  • —XNNPACK (CPU) any arm64: ExecuTorch 1.4.0 export_llm: 8-bit dynamic activations and 4-bit weights in groups of 32, int8 per-channel embeddings, XNNPACK with extended ops, prefill chunk 2048, fp32 KV cache. Built by run 1.
  • —Vulkan (GPU) any arm64: ExecuTorch 1.4.0 export_llm: 8-bit dynamic activations and 4-bit weights in groups of 32, int8 per-channel embeddings, the Vulkan delegate, prefill chunk 2048, fp32 KV cache. Built by run 1.

License

A quantized derivative of Qwen/Qwen2.5-3B-Instruct, distributed under the same terms (qwen-research). The upstream license files are included unchanged: `LICENSE`.