VladHong/llama-cpp-K2-FAST2
llama.cpp โ K2-FAST2 fork ๐
This fork makes K2-Horizon GGUFs up to 3.1ร faster at long context. It is MBZUAI-IFM's llama.cpp model/K2Horizon branch (commit 42adf01) plus one addition: sliding-window attention (SWA) support for `k2-horizon` โ 51 lines in src/models/k2-horizon.cpp.
๐ฆ Companion model (self-contained, window embedded): K2-Horizon-MoVA-36B-A4B-FAST2 โ HuggingFace ยท ModelScope
Measured speed increase vs vanilla APEX-Mini
Same binary, same frozen 61,337-token prompt, same context window, 5 interleaved rounds, RTX 2080 Ti 22 GB, all layers in VRAM. Full data: results-head2head.json in the model repo.
At 32K context (W=8192): decode 52.64 vs 36.78 tok/s (1.43ร), prefill 1,051 vs 949. At short context FAST2 equals vanilla โ decode there is weight-bandwidth-bound and untouched.
Quality: perplexity inside the window is numerically identical to full attention (2.1814 vs 2.1814 on a 240K-token stratified corpus); needle-in-haystack at 61K passes for facts inside the window. Recall is limited to the last W tokens by design โ pick W to match your workload (W=16384 is a one-line rebuild, see the model repo's make_swa_gguf.py).
What the patch does
K2-Horizon runs full attention in all 48 layers (~192 KiB of KV per token โ 6+ GiB read per generated token at 61K context), which is why vanilla decode collapses from ~57 tok/s at 4K to ~16 tok/s at 61K. The patch lets a GGUF declare a sliding window (k2-horizon.attention.sliding_window), which:
load_arch_hparams: reads the key (from GGUF metadata or--override-kv), setsn_swa,swa_type = LLAMA_SWA_TYPE_STANDARD, and marks the attention layers SWA;graph(): routes attention through the iswa input path (build_attn_inp_kv_iswa) so each layer's KV is capped at the last W tokens โ decode cost flattens instead of growing.
Without a window key, behavior is bit-identical to stock model/K2Horizon.
Using it
# build (same as the MBZUAI fork, plus FA_ALL_QUANTS for mixed KV quant support)
cmake -B build -DGGML_CUDA=ON -DGGML_CUDA_FA_ALL_QUANTS=ON -DLLAMA_CURL=OFF \
-DCMAKE_CUDA_ARCHITECTURES=<your-sm> -DCMAKE_BUILD_TYPE=Release
cmake --build build --target llama-server -j
# option A: the FAST2 GGUF (window embedded โ no flags)
llama-server -m K2-Horizon-MoVA-36B-A4B-FAST2-SWA8K.gguf -c 65536 -ngl 99 --parallel 1 \
-fa on -b 2048 -ub 2048 -t 8 --port 8099
# option B: any k2-horizon GGUF + window at load time
llama-server -m K2-Horizon-MoVA-36B-A4B-APEX-Mini.gguf -c 65536 -ngl 99 \
--override-kv k2-horizon.attention.sliding_window=int:8192 ...Where to get what
The upstream llama.cpp README is preserved as README-upstream.md. All credit for the model/K2Horizon support belongs to MBZUAI-IFM; the FAST2 change is the 51-line diff above.
