dariofiore/speculative-decoding-demo
0
Speculative Decoding Demo
Target model: `Qwen/Qwen2.5-3B-Instruct` Draft model: `Qwen/Qwen2.5-0.5B-Instruct`
Live side-by-side comparison of normal generation vs assisted (speculative) generation on ZeroGPU.
What you will see
Why this matters
Large language models are often limited by memory bandwidth, not pure compute. Every new token requires streaming the full model weights from GPU memory.
Speculative decoding lets a small draft model propose several tokens. The large model verifies them in a single parallel forward pass. When the draft is good, you keep multiple tokens for roughly the cost of one target step.
ZeroGPU notes
- Weights load on CPU at startup (no accelerate multi-device offload — those hooks often cause `AcceleratorError` on ZeroGPU).
- On each run the app probes CUDA: if assisted is faster, it stays on GPU; if not (common for 3B + 0.5B fully resident), it runs the timed comparison on CPU, where each target step is expensive and speculation shows a clear win.
- Greedy decoding (
do_sample=False) for a fair comparison. - Acceptance numbers are approximate (derived from speedup).
- Default max tokens is 64 so CPU-path runs fit the ZeroGPU time budget (
duration=180).
Hardware
This Space is designed for ZeroGPU (@spaces.GPU(duration=180)).
