praful-goel/speculative_decoding_models
6
Model Architecture & Training
Both models are Decoder-only Transformers (GPT-style) implemented in pure PyTorch.
Key Components
- Token + Position Embeddings
- Learned token embedding table:
(vocab_size, n_embd) - Rotary Position Embedding: parameter-free sinusoidal rotations on attention queries and keys
- Decoder Blocks Each block consists of:
- Masked Self-Attention (MHA/MQA/GQA)
- Key-Value (KV) Caching for O(1) generation latency
- Rotary Positional Embedding to understand relative positions
- Feed-Forward Network (Linear → GELU → Linear → Dropout)
- Residual connection
- Pre-norm RMSNorm
- Final RMSNorm + Language Modeling Head
- Linear projection from
n_embd→vocab_sizeto produce logits
Training Strategy
Main Model
The main model was trained in 4 distinct phases to hande hardware constraints and optimize convergence
Draft Model (Small)
The draft model (small) was trained only once
max_iters = 40_000
warmup_steps = 1_000
eval_iter = 20
eval_interval = 1_000
accumulation_steps = 16
base_lr = 3e-4
weight_decay=0.1
batch_size = 16Draft Model (Medium)
The draft model (medium) was trained only once
max_iters = 40_000
warmup_steps = 2_000
eval_iters = 20
eval_interval = 2_000
accumulation_steps = 16
base_lr = 3e-4
weight_decay = 0.1