jayyun98/MiniCPM5-2B-SFT-Draftfit-DFlash2
Draftfit — MiniCPM5-2B-SFT DFlash2
Draftfit trains and adapts speculative-decoding draft models for a frozen target LLM.
This checkpoint is the DFlash2 branch adapted to:
- Target: `openbmb/MiniCPM5-2B-SFT`
- Starting draft: `openbmb/MiniCPM5-2B-DSpark`
Each Draftfit family has separate weights. DSpark, DFlash CE, DFlash DPACE and DFlash2 are not one universal checkpoint.
Results
MLX / M5 Pro
Validation set: 53 held-out conversations
- PerfectBlend: 35
- OpenCodeInstruct: 6
- APIGen-MT: 6
- Nemotron Agentic: 6
Domain results — MLX
Full prompt-free summary: `evaluation-v4.json`
Which draft should I use?
For this MiniCPM5-SFT experiment, DSpark is the strongest practical baseline.
Draftfit DSpark reached 116.15 tok/s vs 105.38 tok/s for the official DSpark in the tested MLX setup.
These measurements do not establish a general speed advantage.
Revisions
Different revisions used different held-out sets, so their measurements should not be used to rank revisions directly.
Validation notes
The models were trained for 1,000 steps per stage using disjoint 1,000-conversation training sets.
Training mixture per stage:
Important limitations:
- DSpark/DFlash serving correctness still requires exact token/state and acceptance-length validation.
- DFlash2 was tested on MLX but not with SGLang 0.5.16.
- A successfully loaded draft checkpoint does not prove correct speculative decoding.
- Reported speed is specific to the tested target, hardware, verifier settings and prompt set.
Download
hf download jayyun98/MiniCPM5-2B-SFT-Draftfit-DFlash2 \
--revision main \
--local-dir ./draftThe repository contains:
config.json
model.safetensorsUse them with the pinned MiniCPM5-2B-SFT target and tokenizer.
Training, conversion and benchmarking workflows are documented in:
License / data notice
The training mixture includes APIGen-MT-5k, which is marked CC BY-NC 4.0 and includes additional generator-provenance restrictions.
Treat this release as research / non-commercial unless you have reviewed the source dataset terms.
See `SOURCE_TERMS.md` and release_manifest.json.
Experimental research artifact. The reported numbers are not a general production speedup or correctness claim.
