phillipchaffee/dspark-proxy-run
DSpark proxy run on Qwen3-4B (via DeepSpec) Artifacts from an end-to-end proxy run of DeepSpec's offline DSpark pipeline (pinned 005e03b8) against a Qwen/Qwen3-4B target at 20k-sample scale, run on Modal for the DSpark-for-GLM-5.3-Flash wayfinder effort — see ticket Proxy run: DSpark on Qwen3-4B via DeepSpec, end to end and the run log in experiments/proxy-run/. Contents path what it is cache/ Target hidden-state cache from DeepSpec's… See the full description on the dataset page: https://huggingface.co/datasets/phillipchaffee/dspark-proxy-run.
DSpark proxy run on Qwen3-4B (via DeepSpec)
Artifacts from an end-to-end proxy run of DeepSpec's offline DSpark pipeline (pinned 005e03b8) against a Qwen/Qwen3-4B target at 20k-sample scale, run on Modal for the DSpark-for-GLM-5.3-Flash wayfinder effort — see ticket Proxy run: DSpark on Qwen3-4B via DeepSpec, end to end and the run log in `experiments/proxy-run/`.
Contents
Results
- Data-scale curve for this pipeline: τ 1.06 (694 samples, 10 steps) → 2.47 (20k, this repo's drafter) → 6.12 (1.3M, released
deepseek-ai/dspark_qwen3_4b_block7) on gsm8k — drafters are strongly data-hungry, with no corpus discount for a new target. - Harness validated against the released drafter: all six tasks within ±0.12 of the paper's Table 1 (alpaca exact).
- This repo's 20k drafter (τ 2.47/2.33/1.77/1.85/1.53/1.49 on gsm8k/math500/humaneval/mbpp/mt-bench/alpaca) is a net serving loss on chat/code (0.80–1.24× vs target-only baseline) and only pays on math.
- Confidence heads: ours is well-calibrated (ECE 0.005–0.008, AUC 0.91–0.95); the released one is overconfident (ECE 0.09–0.13, pred 0.83 vs observed 0.74).
Usage
from huggingface_hub import snapshot_download
snapshot_download("phillipchaffee/dspark-proxy-run", repo_type="dataset")cache/ is DeepSpec's training cache format — point prepare_target_cache.py-compatible loaders at manifest.json / samples.idx. The checkpoint loads as a drafter for Qwen3-4B DSpark stacks (e.g. SGLang's in-tree DSpark support). Intended consumer: the A/B smoke ticket (A/B smoke on Qwen: sequential-head variant vs vanilla Markov at proxy scale) reuses this exact cache so the head is the only variable.
Provenance
- Pipeline: DeepSpec @
005e03b8, orchestrated bydspark_proxy.pyon Modal (split → regen → cache → train → eval → push). Deviations from DeepSpec's defaults are documented in the run README. - Compute: ~$165 GPU total (~$3.40 regen, ~$3.62 cache, ~$35.20 train on 4×H100, rest eval).
- Source data:
mlabonne/open-perfectblend(Apache-2.0); responses generated byQwen/Qwen3-4B(Apache-2.0). Training/config code derives from DeepSpec (MIT). The Apache-2.0 tag reflects the data lineage; flip it if your use of the checkpoint needs otherwise. - Optimizer resume states (7.4 GB/rank) were deliberately not uploaded — training is complete.
