Team Ai
Modelpublic

Atomic-Germ/Qwen3.5-9B-Claude-Code-NPU2

sourceHugging Faceapache-2.0updated 23d agoView on Hugging Face
0likes333downloads
Model Card

IF YOU USE COMMUNITY QWEN MODELS DO NOT UPGRADE TO FLM v1.0.2+

Qwen3.5-9B-Claude-Code - Q4NX for FastFlowLM (AMD Ryzen AI XDNA2)

A coding-focused Qwen3.5-9B fine-tune (Claude-Opus-4.6 distillation), converted to Q4NX for FastFlowLM.

What is Q4NX?

Q4NX is FastFlowLM's native packed-quantization format - a rearranged Q4_1 layout tuned for the NPU matrix engine's tile sizes and memory access patterns. It is not a GGUF file and it does not run on llama.cpp or Ollama; it is meant exclusively for the FastFlowLM engine on AMD Ryzen AI NPUs.

Requirements

  • —FastFlowLM >= 0.9.45 (flm CLI)
  • —AMD Ryzen AI processor with XDNA2 (NPU2) - Strix Point / Ryzen AI 300 series or later
  • —Linux with the XRT NPU stack installed
  • —~16 GB of unified system memory (Q4NX weights + activations + KV cache)

Files

FilePurpose
model.q4nxQuantized Q4NX weights
config.jsonFastFlowLM model configuration
tokenizer.jsonTokenizer
tokenizer_config.jsonSpecial tokens and chat template
chat_template.jinjaChat template (optional)
flm-add.pyInstaller script - registers this model with FastFlowLM

Install and run

This repository works with flm-add, a small installer that copies the model into the FastFlowLM user directory and registers the tag. It never modifies the system FastFlowLM install.

pip install flm-add or uv tool install flm-add

bash
uv tool install flm-add
flm-add Atomic-Germ/Qwen3.5-9B-Claude-Code-NPU2 --tag qwen3.5-claude-code:9b
FLM_CONFIG_PATH="$HOME/.config/flm/model_list.json" FLM_XCLBIN_PATH="$HOME/.config/flm" flm run qwen3.5-claude-code:9b

Kernels

FastFlowLM's NPU kernels (xclbins) are closed source and are not shipped in this repository. flm-add.py links the kernels of the official `qwen3.5:9b` model (Qwen3.5-9B-NPU2), because this model shares the same engine family (qwen3.5) and architecture.

Model

  • —Registry tag: qwen3.5-code:9b
  • —Engine family: qwen3.5
  • —Kernel source: Qwen3.5-9B-NPU2
  • —Context length: 262,144 tokens (from config)
  • —model.q4nx size: 7.63 GB
  • —Base model: empero-ai/Qwen3.5-9B-Claude-Opus-4.6-Distill
  • —License: apache-2.0

GhostWriter Influence Test (Arbitrary but repeatable benchmark)

Tested on an AMD Ryzen AI 340 Framework 13 laptop.

MetricValue
Prompt Tokens9,210
Completion Tokens1,141
Total Tokens10,351
Active KV Tokens10,351
Max KV Token Capacity32,768
KV Token Occupancy31.59%
Load Duration0.000000721 seconds
Prefill Duration (TTFT)33.58 ms
Decoding Duration198.47 ms
Prefill Speed274.25 tokens/sec
Decoding Speed5.75 tokens/sec

Original model card

See the upstream model card for training details, benchmarks, and upstream usage. This repository only contains the Q4NX conversion for FastFlowLM.