ZenithLLM/ZenAlta-Draft
0619
โก Zen Alta 4-Layer Speculative Decoding Draft Model (~790M)
Zen Alta Draft is a ultra-lightweight, 4-layer speculative decoding companion model engineered by ZenithLLM. Sliced from the top of the 24-layer Zen Alta architecture, it shares the exact same 128,256 BPE vocabulary and embedding/LM head, enabling lossless 2ร speculative inference acceleration in llama.cpp, vLLM, and mobile runtimes.
๐ฏ What is Speculative Decoding?
In standard autoregressive generation, deep models calculate every single token sequentially (e.g. 24 transformer layers per token).
With Zen Alta Draft:
- The Fast Draft (4 Layers, 790M): Quickly guesses 4โ5 candidate tokens in parallel in just ~20โ30ms.
- The Target Model (Zen Alta 24 Layers, 2.8B): Verifies all proposed tokens in a single parallel forward pass (~40ms).
- Result: Accepted tokens are committed simultaneously, achieving 40โ50+ tokens/sec on mobile chips with 0% degradation in output quality or persona.
๐ฆ Model Specifications
๐ Repository Contents
This consolidated repository contains both the raw PyTorch weights and the ready-to-run quantized GGUF:
๐ How to Run in llama.cpp
Speculative Decoding (Target + Draft Pairing)
Download the target model from ZenithLLM/ZenAlta-1-3B-Phase2-GGUF and the draft model from this repo:
# Speculative decoding command
./llama-cli \
-m ZenAlta-1-3B-Pruned.Q4_K_M.gguf \
-md zen-alta-draft-q4_k_m.gguf \
--draft-max 5 \
-p "<|start_header_id|>user<|end_header_id|>\n\nhey who are you?<|eot_id|><|start_header_id|>assistant<|end_header_id|>\n\n" \
-n 128Standalone Inference (Fast Preview)
./llama-cli -m zen-alta-draft-q4_k_m.gguf -p "what is up" -n 64๐ Related Models
- Target Model (LoRA Adapter): ZenithLLM/ZenAlta-1-3B-Phase2
- Target Model (GGUF Q4_K_M): ZenithLLM/ZenAlta-1-3B-Phase2-GGUF
- Base Pruned Model (24-Layer): ZenithLLM/ZenAlta-1-3B-Pruned
