Team Ai
Modelpublic

ZenithLLM/ZenAlta-Draft

sourceHugging Facellama3.2updated 10d agoView on Hugging Face
0likes619downloads
Model Card

โšก Zen Alta 4-Layer Speculative Decoding Draft Model (~790M)

Zen Alta Draft is a ultra-lightweight, 4-layer speculative decoding companion model engineered by ZenithLLM. Sliced from the top of the 24-layer Zen Alta architecture, it shares the exact same 128,256 BPE vocabulary and embedding/LM head, enabling lossless 2ร— speculative inference acceleration in llama.cpp, vLLM, and mobile runtimes.


๐ŸŽฏ What is Speculative Decoding?

In standard autoregressive generation, deep models calculate every single token sequentially (e.g. 24 transformer layers per token).

With Zen Alta Draft:

  1. 1.The Fast Draft (4 Layers, 790M): Quickly guesses 4โ€“5 candidate tokens in parallel in just ~20โ€“30ms.
  2. 2.The Target Model (Zen Alta 24 Layers, 2.8B): Verifies all proposed tokens in a single parallel forward pass (~40ms).
  3. 3.Result: Accepted tokens are committed simultaneously, achieving 40โ€“50+ tokens/sec on mobile chips with 0% degradation in output quality or persona.

๐Ÿ“ฆ Model Specifications

ParameterValue
Base ArchitectureLlama 3.2 (CausalLM)
Hidden Layers4 (vs 24 in Target model)
Hidden Dimension3072
Intermediate Size8192
Attention Heads24 query heads / 8 KV heads
Vocabulary Size128,256 (Identical to Llama 3.2 & Zen Alta)
Context Length131,072 tokens
RoPE Theta500,000.0

๐Ÿ“‚ Repository Contents

This consolidated repository contains both the raw PyTorch weights and the ready-to-run quantized GGUF:

FileSizeDescription
model.safetensors1.52 GBUnquantized FP16 PyTorch weights (4 layers)
zen-alta-draft-q4_k_m.gguf545.72 MBQuantized 4-bit medium GGUF for llama.cpp & mobile
config.json< 1 KB4-layer model configuration
tokenizer.json16.4 MBFast BPE tokenizer definition
chat_template.jinja< 4 KBLlama 3.2 conversational chat template

๐Ÿš€ How to Run in llama.cpp

Speculative Decoding (Target + Draft Pairing)

Download the target model from ZenithLLM/ZenAlta-1-3B-Phase2-GGUF and the draft model from this repo:

bash
# Speculative decoding command
./llama-cli \
  -m ZenAlta-1-3B-Pruned.Q4_K_M.gguf \
  -md zen-alta-draft-q4_k_m.gguf \
  --draft-max 5 \
  -p "<|start_header_id|>user<|end_header_id|>\n\nhey who are you?<|eot_id|><|start_header_id|>assistant<|end_header_id|>\n\n" \
  -n 128

Standalone Inference (Fast Preview)

bash
./llama-cli -m zen-alta-draft-q4_k_m.gguf -p "what is up" -n 64

๐Ÿ”— Related Models