ornith-ai/Ornith-1.5-397B-DFlash
<img width="600px" src="assets/ornith_logo.png">
Ornith-1.5-397B-DFlash
Ornith-1.5-397B-DFlash pairs the Ornith-1.5-397B model with a DFlash draft model for speculative decoding. This repository provides the DFlash draft checkpoint and serving configuration for accelerating Ornith-1.5-397B inference. The DFlash checkpoint is not a standalone language model. It is designed to run together with the target model under a DFlash-enabled inference setup.
Ornith-1.5-397B
๐ฆ Ornith-1.5 is a major step toward building foundation models through end-to-end self-improvement. Ornith-1.5 extends Ornith-1.0 (which was developed on top of Qwen3.5 and Gemma4 with additional continued pretraining, mid-training, and post-training) by expanding the self-improvement loop from scaffold and rollout optimization to jointly optimizing task generation, scaffold construction, and solution rollouts. Rather than relying on a fixed set of human-curated tasks and manually designed harnesses, Ornith-1.5 continuously generates new training tasks, discovers effective strategies for solving them, and improves the policy through reinforcement learning. For more details on the task, harness, and rollout reward design, please refer to the Ornith-1.5 blog.
<img style="width: 100%; max-width: 900px;" src="assets/ornith397beval.png" alt="Ornith 1.5 397B Benchmark Results" title="Ornith 1.5 397B Benchmark Results">
DFlash acceleration
DFlash provides the speculative-decoding component of this project. It uses a lightweight block-diffusion draft model to propose multiple tokens in parallel, while Ornith-1.5-397B verifies those proposals as the target model. This allows the serving stack to improve decoding throughput while preserving the target model's output distribution.
In this setup, the two models have distinct roles:
- Ornith-1.5-397B: target model responsible for verification and final generation.
- Ornith-1.5-397B-DFlash: draft model used to generate speculative token blocks for acceleration.
The sections below therefore separate target-model quality benchmarks from DFlash serving configuration and acceleration measurements.
Quickstart
<div style="border-left:4px solid #FD8E5B;background:rgba(253,142,91,0.1);border-radius:6px;padding:12px 16px;font-family:-apple-system,BlinkMacSystemFont,'Segoe UI',Roboto,sans-serif;font-size:14px;line-height:1.6"> <div style="font-weight:700;color:#FD8E5B;margin-bottom:6px">๐ NOTE</div> <p style="margin:0 0 10px"><b>Ornith-1.5-397B</b> is a <b>reasoning model</b>: by default the assistant turn opens with a <code style="background:rgba(253,142,91,0.15);padding:1px 5px;border-radius:4px"><think> โฆ </think></code> block before the final answer. The serving recipes below enable a reasoning parser so the chain-of-thought is returned in a separate <code style="background:rgba(253,142,91,0.15);padding:1px 5px;border-radius:4px">reasoningcontent</code> field, and a tool-call parser so the model's <code style="background:rgba(253,142,91,0.15);padding:1px 5px;border-radius:4px"><toolcall></code> blocks are surfaced as OpenAI-style <code style="background:rgba(253,142,91,0.15);padding:1px 5px;border-radius:4px">toolcalls</code>.</p> <p style="margin:0 0 6px">Serving Ornith-1.5-397B with DFlash requires recent runtimes:</p> <ul style="margin:0 0 10px;padding-left:20px"> <li><b>Transformers</b> โฅ 5.8.1</li> <li><b>vLLM</b> โฅ 0.20.2 (tested with 0.28.0)</li> <li><b>SGLang</b> = 0.5.18</li> </ul> <p style="margin:0 0 6px">Recommended sampling parameters:</p> <ul style="margin:0;padding-left:20px"> <li><b>For general tasks:</b> <code style="background:rgba(253,142,91,0.15);padding:1px 5px;border-radius:4px">temperature=0.6</code>, <code style="background:rgba(253,142,91,0.15);padding:1px 5px;border-radius:4px">topp=0.95</code>, <code style="background:rgba(253,142,91,0.15);padding:1px 5px;border-radius:4px">top_k=20</code></li> <li><b>To reproduce the reported benchmarks:</b> <code style="background:rgba(253,142,91,0.15);padding:1px 5px;border-radius:4px">temperature=1.0</code></li> </ul> </div>
DFlash Installation and Launch
The DFlash draft must be loaded together with the target model. The configurations below use the runtime versions validated for this deployment:
- vLLM:
>=0.20.2(tested with0.28.0) - SGLang:
0.5.18
Replace <DFlash directory> with the local directory or model repository containing this DFlash draft checkpoint.
vLLM with DFlash
vllm serve ornith-ai/Ornith-1.5-397B --served-model-name x --trust-remote-code \
--tensor-parallel-size 1 --port 8801 \
--gpu-memory-utilization 0.85 --max-model-len 32768 \
--speculative-config '{"method":"dflash","model":"<DFlash directory>","num_speculative_tokens":8}'SGLang with DFlash
python -m sglang.launch_server --model-path ornith-ai/Ornith-1.5-397B --trust-remote-code \
--tp-size 1 --port 30010 --mem-fraction-static 0.80 \
--speculative-algorithm DFLASH \
--speculative-draft-model-path <DFlash directory> \
--speculative-dflash-block-size 8The examples above use a speculative block size / token count of 8. Adjust tensor parallelism and memory settings to match the actual target-model deployment; the full 397B target model may require multi-GPU serving depending on precision and hardware.
Citation
If you find our work helpful, feel free to give us a cite.
@misc{ornith_1_5,
title = {{Ornith-1.5}: From Self-Scaffolding to Self-Improvement},
url = {https://ornith.ai/ornith_1_5.html},
author = {{Ornith Team}},
year = {2026}
}