Team Ai
Apppublic

abhimittal/speculative-decoding-trace

sourceHugging Facemitupdated 3mo agoView on Hugging Face
0likes
App README

Speculative decoding, token by token

Two layers, both client-side (static Space — free tier, no ZeroGPU quota, no cold starts):

  1. 1.Interactive explainer (index.html) — drag toy draft/target distributions and watch the acceptance probability min(1, p/q), the residual norm(max(p − q, 0)), the losslessness histogram, and the E[tokens] = (1 − α^{γ+1})/(1 − α) speedup curve respond live.
  2. 2.Real inference in your browser (live.html) — quantized ONNX distilgpt2 drafts, gpt2 verifies, via transformers.js (WebGPU/WASM). The actual rejection sampling on actual logits, with the per-position math rendered for every token.

The repo also ships a Gradio app with KV-cached speculative decoding, draft-vs-target attention-trace diffing, and an optional 70B cloud cross-check — run it locally with python app.py.

Source: https://github.com/aabhimittal/attention-trace-differ-for-speculative-decoding