Team Ai
Modelpublic

efficiencyx/Jun-LoRA-12B-MTP-GGUF

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes26downloads
Model Card

Jun-LoRA-12B-MTP-GGUF

A trained MTP / speculative-decoding draft model for `efficiencyx/Jun-LoRA-12B-GGUF`. Not a standalone chat model — it only exists to propose tokens that Jun then verifies.

  • —jun-drafter-qat600-q4_k_m.gguf — 327 MB, Q4KM
  • —Initialized from google/gemma-4-12B-it-qat-q4_0-unquantized-assistant, matching Jun's own lineage (unsloth/gemma-4-12B-it-qat-q4_0-unquantized)
  • —4 layers, hidden 1024, backbone hidden 3840
  • —600 steps on a 1500-sample in-character roleplay corpus, A100 80GB

Speed

RTX 3060, ollama, Jun 12B Q4KM as the target, medians over 6 runs:

setuptok/s
Jun alone, no drafter36.28
this drafter, `draft_num_predict=1`45.48 (+25.4%)
this drafter, n=241.85
this drafter, n=340.79
this drafter, n=436.73

Use n=1. The usual "draft 2-3 tokens" advice loses on this hardware: each extra token in the verify batch costs ~10ms on a 3060, so deeper drafts pay more to verify than they save.

Acceptance

Accepted tokens per target forward, higher is better. On an A100 at bf16 the training gain is clear:

bf16 (A100)
stock drafter2.10
trained, 300 steps2.81
trained, 600 steps2.89

Acceptance improved monotonically with training. On a 3060 the end-to-end tok/s above is bounded by verify cost rather than by acceptance, so the headroom this buys shows up on hardware where the target forward is the cheaper half.

Lineage matters An earlier run initialized from the

non-qat assistant — same backbone_hidden_size, so every shape check passed — and lost to stock outright at Q4KM. A drafter reads the target's last-layer activations directly; "same size" is not "same body".

Quantization

Drafter quantization barely moves acceptance (bf16 2.18, Q80 2.16, Q4KM 2.16 accepted tokens per target forward) but strongly moves draft time (8817 / 4935 / 4352 ms for the same work). Q4K_M is the right format for a drafter.