Pradheep1647/eagle3-speculative-decoding-policy
EAGLE3 Speculative Decoding -- Energy-Aware Policy Models
Eight models, one problem: pick (speculative_num_steps, speculative_eagle_topk, speculative_num_draft_tokens) for sglang + EAGLE3 so GPU energy utilization lands inside a 95-98% band. All trained on the `eagle3-speculative-decoding-energy-sweep` dataset. Full writeup, sweep mechanism, and live-validated results: project README.
Shared I/O contract -- input is a 4-dim state [batch_size/8.0, gpu_temp_c/100.0, gpu_mem_used_mb/8192.0, gpu_util_pct/100.0]; output is an index into the same 19-action space (RL/policy.py in the repo above), decoded to the three sglang flags.
Which one to actually use
mlp_bandit, lookup_table, and doubly_robust agree on every batch size and are the live-validated picks. cql and bcq collapsed to the non-speculative baseline past bs=1 (overly conservative default hyperparameters against this reward scale) and are not recommended -- kept here for completeness, not as a suggested pick. See the project README for the full live A/B numbers per algorithm.
Hardware this was validated on
RTX 4060 Laptop GPU (8GB), unsloth/Llama-3.2-1B-Instruct target + rescommons/SpecForge-EAGLE3-Llama-3.2-1B-Instruct draft, 80W power cap. Picks are specific to this hardware/model pair -- retrain on the linked dataset (or a fresh sweep) before trusting these on different hardware.
