Team Ai
Modelpublic

VladHong/llama-cpp-K2-FAST2

sourceHugging Faceapache-2.0updated 3d agoView on Hugging Face
0likes1.4kdownloads
Model Card

llama.cpp โ€” K2-FAST2 fork ๐Ÿš€

This fork makes K2-Horizon GGUFs up to 3.1ร— faster at long context. It is MBZUAI-IFM's llama.cpp model/K2Horizon branch (commit 42adf01) plus one addition: sliding-window attention (SWA) support for `k2-horizon` โ€” 51 lines in src/models/k2-horizon.cpp.

๐Ÿ“ฆ Companion model (self-contained, window embedded): K2-Horizon-MoVA-36B-A4B-FAST2 โ€” HuggingFace ยท ModelScope

Measured speed increase vs vanilla APEX-Mini

Same binary, same frozen 61,337-token prompt, same context window, 5 interleaved rounds, RTX 2080 Ti 22 GB, all layers in VRAM. Full data: results-head2head.json in the model repo.

@ 61K context (61,354-token prompt)decode tok/sprefill tok/sspeedup
vanilla APEX-Mini (q8_0 KV โ€” the only way vanilla fits 22 GB at 61K; default f16 KV OOMs)16.48692โ€”
FAST2 (this fork + embedded W=8192 window)50.721,482decode 3.1ร—, prefill 2.1ร—
FAST2 with W=1638443.901,187decode 2.7ร—

At 32K context (W=8192): decode 52.64 vs 36.78 tok/s (1.43ร—), prefill 1,051 vs 949. At short context FAST2 equals vanilla โ€” decode there is weight-bandwidth-bound and untouched.

Quality: perplexity inside the window is numerically identical to full attention (2.1814 vs 2.1814 on a 240K-token stratified corpus); needle-in-haystack at 61K passes for facts inside the window. Recall is limited to the last W tokens by design โ€” pick W to match your workload (W=16384 is a one-line rebuild, see the model repo's make_swa_gguf.py).

What the patch does

K2-Horizon runs full attention in all 48 layers (~192 KiB of KV per token โ€” 6+ GiB read per generated token at 61K context), which is why vanilla decode collapses from ~57 tok/s at 4K to ~16 tok/s at 61K. The patch lets a GGUF declare a sliding window (k2-horizon.attention.sliding_window), which:

  1. 1.load_arch_hparams: reads the key (from GGUF metadata or --override-kv), sets n_swa, swa_type = LLAMA_SWA_TYPE_STANDARD, and marks the attention layers SWA;
  2. 2.graph(): routes attention through the iswa input path (build_attn_inp_kv_iswa) so each layer's KV is capped at the last W tokens โ€” decode cost flattens instead of growing.

Without a window key, behavior is bit-identical to stock model/K2Horizon.

Using it

bash
# build (same as the MBZUAI fork, plus FA_ALL_QUANTS for mixed KV quant support)
cmake -B build -DGGML_CUDA=ON -DGGML_CUDA_FA_ALL_QUANTS=ON -DLLAMA_CURL=OFF \
      -DCMAKE_CUDA_ARCHITECTURES=<your-sm> -DCMAKE_BUILD_TYPE=Release
cmake --build build --target llama-server -j

# option A: the FAST2 GGUF (window embedded โ€” no flags)
llama-server -m K2-Horizon-MoVA-36B-A4B-FAST2-SWA8K.gguf -c 65536 -ngl 99 --parallel 1 \
    -fa on -b 2048 -ub 2048 -t 8 --port 8099

# option B: any k2-horizon GGUF + window at load time
llama-server -m K2-Horizon-MoVA-36B-A4B-APEX-Mini.gguf -c 65536 -ngl 99 \
    --override-kv k2-horizon.attention.sliding_window=int:8192 ...

Where to get what

artifactHuggingFaceModelScope
FAST2 GGUF (window embedded) + card + measurementsVladHong/K2-Horizon-MoVA-36B-A4B-FAST2Dalvlad/K2-Horizon-MoVA-36B-A4B-FAST2
this fork as a source snapshotVladHong/llama-cpp-K2-FAST2branch k2-fast2 of Dalvlad/llama-cpp-K2-FAST2
standalone patch (apply to stock MBZUAI fork)k2-fast2-swa.patch in either model reposame

The upstream llama.cpp README is preserved as README-upstream.md. All credit for the model/K2Horizon support belongs to MBZUAI-IFM; the FAST2 change is the 51-line diff above.