BobbieBieee/Osprey-Speculative-Decoding
Osprey: Target-agnostic Pre-training Makes Stronger Drafters in Speculative Decoding
Accepted at EMNLP 2026.
Paper: https://arxiv.org/abs/2609.09338 Code: https://github.com/LeanModels/Osprey
Adapted Osprey drafters for three target models. Every checkpoint starts from the same target-agnostic pretrained backbone (Qwen3-4B pruned to 2 layers, pretrained on FineWeb) and is adapted to its target with vocabulary alignment, zero-initialized QKV expansion, and on-policy EAGLE-3 distillation.
Each directory holds config.json and model.safetensors. Serving requires the SGLang patch shipped in the code repository (sglang_patches/), which adds multi-layer EAGLE-3 drafts and keeps the drafter's own embedding at load time.
Also in this repository:
The Llama-3.3-70B-Instruct splits are drawn from frankleeeee/PerfectBlend-Regenerated-Llama-3.3-70B-Instruct (1,419,775 conversations): records are numbered in file order, shuffled once with random.seed(42), and the first 100,000 / next 512 taken as train / eval.
The MiniMax-M2.5 splits take their prompts from the code subset of NVIDIA's Nemotron-Post-Training-Dataset-v2 (CC BY 4.0) and keep that dataset's metadata fields. The generator field is inherited from the source and does not describe the assistant turns here, which were regenerated by MiniMax-M2.5.
Citation
@inproceedings{bie2026osprey,
title = {Osprey: Target-agnostic Pre-training Makes Stronger Drafters in Speculative Decoding},
author = {Bie, Fengxiang and Jian, Yuqing and Yu, Yifan and Zhou, Zhongzhu and Shao, Zelei and Athiwaratkun, Ben and Song, Shuaiwen Leon and Xu, Chenfeng and Wu, Xiaoxia and Zhang, Tianyi},
booktitle = {Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP)},
year = {2026},
}