Team Ai
Apppublic

dariofiore/speculative-decoding-demo

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes
App README

Speculative Decoding Demo

Target model: `Qwen/Qwen2.5-3B-Instruct` Draft model: `Qwen/Qwen2.5-0.5B-Instruct`

Live side-by-side comparison of normal generation vs assisted (speculative) generation on ZeroGPU.

What you will see

MetricDescription
TimeWall-clock seconds for each path
Tokens generatedNew tokens produced
Tokens / secondThroughput
SpeedupNormal time ÷ assisted time
Approx. avg. accepted tokensHeuristic from observed speedup
Approx. acceptance rateHeuristic vs the draft window

Why this matters

Large language models are often limited by memory bandwidth, not pure compute. Every new token requires streaming the full model weights from GPU memory.

Speculative decoding lets a small draft model propose several tokens. The large model verifies them in a single parallel forward pass. When the draft is good, you keep multiple tokens for roughly the cost of one target step.

ZeroGPU notes

  • —Weights load on CPU at startup (no accelerate multi-device offload — those hooks often cause `AcceleratorError` on ZeroGPU).
  • —On each run the app probes CUDA: if assisted is faster, it stays on GPU; if not (common for 3B + 0.5B fully resident), it runs the timed comparison on CPU, where each target step is expensive and speculation shows a clear win.
  • —Greedy decoding (do_sample=False) for a fair comparison.
  • —Acceptance numbers are approximate (derived from speedup).
  • —Default max tokens is 64 so CPU-path runs fit the ZeroGPU time budget (duration=180).

Hardware

This Space is designed for ZeroGPU (@spaces.GPU(duration=180)).