Team Ai
Modelpublic

jayyun98/MiniCPM5-2B-SFT-Draftfit-DFlash-DPACE

sourceHugging Faceotherupdated 6d agoView on Hugging Face
0likes456downloads
Model Card

Draftfit — MiniCPM5-2B-SFT DFlash DPACE

Draftfit trains and adapts speculative-decoding draft models for a frozen target LLM.

This checkpoint is the DFlash DPACE branch adapted to:

Each Draftfit family has separate weights. DSpark, DFlash CE, DFlash DPACE and DFlash2 are not one universal checkpoint.

Results

MLX / M5 Pro

DraftOverall tok/s
Official DSpark105.38
Draftfit DSpark116.15
Draftfit DFlash CE89.61
Draftfit DFlash DPACE89.97
Draftfit DFlash280.75

Validation set: 53 held-out conversations

  • —PerfectBlend: 35
  • —OpenCodeInstruct: 6
  • —APIGen-MT: 6
  • —Nemotron Agentic: 6

Domain results — MLX

DraftPerfectBlendOpenCodeAPIGen-MTNemotron Agentic
Official DSpark107.83112.19107.7985.86
Draftfit DSpark118.29122.28116.0491.99
Draftfit DFlash CE89.0299.1892.4979.31
Draftfit DFlash DPACE89.84103.2091.6276.97
Draftfit DFlash278.3792.7687.4970.41

Full prompt-free summary: `evaluation-v4.json`

Which draft should I use?

DraftBest use
DSparkBest first choice when adapting an existing DSpark draft to a new target
DFlash CESimple block-parallel DFlash baseline
DFlash DPACETest position-weighted training objectives against CE
DFlash2Research richer candidate ranking / grouped-convolution drafts

For this MiniCPM5-SFT experiment, DSpark is the strongest practical baseline.

Draftfit DSpark reached 116.15 tok/s vs 105.38 tok/s for the official DSpark in the tested MLX setup.

These measurements do not establish a general speed advantage.

Revisions

RevisionDescription
mainDFlash DPACE, primary release
v5Additional 1,000-step continual-training stage on a fresh split

Different revisions used different held-out sets, so their measurements should not be used to rank revisions directly.

Validation notes

The models were trained for 1,000 steps per stage using disjoint 1,000-conversation training sets.

Training mixture per stage:

DatasetRows
PerfectBlend deduplicated700
OpenCodeInstruct100
APIGen-MT-5k100
Nemotron-Agentic-v1100

Important limitations:

  • —DSpark/DFlash serving correctness still requires exact token/state and acceptance-length validation.
  • —DFlash2 was tested on MLX but not with SGLang 0.5.16.
  • —A successfully loaded draft checkpoint does not prove correct speculative decoding.
  • —Reported speed is specific to the tested target, hardware, verifier settings and prompt set.

Download

bash
hf download jayyun98/MiniCPM5-2B-SFT-Draftfit-DFlash-DPACE \
  --revision main \
  --local-dir ./draft

The repository contains:

text
config.json
model.safetensors

Use them with the pinned MiniCPM5-2B-SFT target and tokenizer.

Training, conversion and benchmarking workflows are documented in:

Draftfit

License / data notice

The training mixture includes APIGen-MT-5k, which is marked CC BY-NC 4.0 and includes additional generator-provenance restrictions.

Treat this release as research / non-commercial unless you have reviewed the source dataset terms.

See `SOURCE_TERMS.md` and release_manifest.json.


Experimental research artifact. The reported numbers are not a general production speedup or correctness claim.