Team Ai
Modelpublic

poolside/Laguna-S-2.1-DFlash-NVFP4

sourceHugging Faceupdated 2mo agoView on Hugging Face
32likes20kdownloads
Model Card

<p align="center"> <img alt="poolside-banner" src="https://poolside.ai/assets/laguna/laguna-s-2-1-banner.svg" width="800px"> </p>

<p align="center"> <a href="https://openrouter.ai/poolside/laguna-s-2.1"><strong>Use on OpenRouter</strong></a> · <a href="https://vercel.com/ai-gateway/models/laguna-s-2.1"><strong>Use on Vercel AI Gateway</strong></a> · <a href="https://poolside.ai/blog/introducing-laguna-s-2-1"><strong>Release blog post</strong></a> </p>

<br>

poolside/Laguna-S-2.1-DFlash-NVFP4

DFlash speculator for the NVFP4 target poolside/Laguna-S-2.1-NVFP4. The speculator is a 6-layer Laguna-style draft model (BF16); pair it with the NVFP4 base for lower-latency serving via speculative decoding.

Trained: e0630_rhiemann_baseline SFT, DFlash Stage-2, 15k steps. Recommended serving setting: num_speculative_tokens=7. DFlash upstream support is in progress (vLLM #46853, SGLang #29446, TRT-LLM #15666). Use poolside/Laguna-S-2.1-NVFP4 as the target model.

Benchmarks

Measured with TP=2, temperature=0, and num_speculative_tokens=15.

Throughput speedup

ConcurrencyGSM8KMATH-500HumanEvalMBPPMT-Bench
13.324x2.893x3.692x2.426x2.338x
42.498x2.174x2.742x1.848x1.831x
82.279x1.948x2.634x1.719x1.772x
162.302x1.965x2.626x1.731x1.672x

Acceptance length

ConcurrencyGSM8KMATH-500HumanEvalMBPPMT-Bench
15.7754.9426.4384.1714.017
45.8044.9356.2984.1614.091
85.7194.8856.5374.1934.373
165.7584.8966.4124.1233.981