Team Ai
Datasetpublic

phillipchaffee/dspark-proxy-run

DSpark proxy run on Qwen3-4B (via DeepSpec) Artifacts from an end-to-end proxy run of DeepSpec's offline DSpark pipeline (pinned 005e03b8) against a Qwen/Qwen3-4B target at 20k-sample scale, run on Modal for the DSpark-for-GLM-5.3-Flash wayfinder effort — see ticket Proxy run: DSpark on Qwen3-4B via DeepSpec, end to end and the run log in experiments/proxy-run/. Contents path what it is cache/ Target hidden-state cache from DeepSpec's… See the full description on the dataset page: https://huggingface.co/datasets/phillipchaffee/dspark-proxy-run.

sourceHugging Faceapache-2.0updated 12d agoView on Hugging Face
0likes159downloads
Dataset Card

DSpark proxy run on Qwen3-4B (via DeepSpec)

Artifacts from an end-to-end proxy run of DeepSpec's offline DSpark pipeline (pinned 005e03b8) against a Qwen/Qwen3-4B target at 20k-sample scale, run on Modal for the DSpark-for-GLM-5.3-Flash wayfinder effort — see ticket Proxy run: DSpark on Qwen3-4B via DeepSpec, end to end and the run log in `experiments/proxy-run/`.

Contents

pathwhat it is
cache/Target hidden-state cache from DeepSpec's prepare_target_cache.py — layers [1, 9, 17, 25, 33], 19,817 valid samples (of 19,987 regenerated; 170 dropped by the min_loss_tokens=14 filter), 611 GB across 9 shards plus manifest.json and samples.idx.
data/perfectblend_train.jsonl19,999 rows shuffled (seed 42) from mlabonne/open-perfectblend.
data/regen.jsonl19,987 regenerated responses (vLLM serving Qwen3-4B: temp 0.7, topp 0.8, topk 20, thinking off). 12 failed rows are in data/regen_error.jsonl.
checkpoints/dspark_block7_qwen3_4b_full_step_380/The drafter trained in this run: block γ=7, 5 draft layers, vanilla Markov head rank 256, confidence head; CE 0.1 + TV(L1) 0.9 + confidence BCE 1.0; lr 6e-4, global batch 512, bf16. 380 steps, loss 2.64 → 1.79. model.safetensors + config.json + train_config.py.

Results

  • —Data-scale curve for this pipeline: τ 1.06 (694 samples, 10 steps) → 2.47 (20k, this repo's drafter) → 6.12 (1.3M, released deepseek-ai/dspark_qwen3_4b_block7) on gsm8k — drafters are strongly data-hungry, with no corpus discount for a new target.
  • —Harness validated against the released drafter: all six tasks within ±0.12 of the paper's Table 1 (alpaca exact).
  • —This repo's 20k drafter (τ 2.47/2.33/1.77/1.85/1.53/1.49 on gsm8k/math500/humaneval/mbpp/mt-bench/alpaca) is a net serving loss on chat/code (0.80–1.24× vs target-only baseline) and only pays on math.
  • —Confidence heads: ours is well-calibrated (ECE 0.005–0.008, AUC 0.91–0.95); the released one is overconfident (ECE 0.09–0.13, pred 0.83 vs observed 0.74).

Usage

python
from huggingface_hub import snapshot_download

snapshot_download("phillipchaffee/dspark-proxy-run", repo_type="dataset")

cache/ is DeepSpec's training cache format — point prepare_target_cache.py-compatible loaders at manifest.json / samples.idx. The checkpoint loads as a drafter for Qwen3-4B DSpark stacks (e.g. SGLang's in-tree DSpark support). Intended consumer: the A/B smoke ticket (A/B smoke on Qwen: sequential-head variant vs vanilla Markov at proxy scale) reuses this exact cache so the head is the only variable.

Provenance

  • —Pipeline: DeepSpec @ 005e03b8, orchestrated by dspark_proxy.py on Modal (split → regen → cache → train → eval → push). Deviations from DeepSpec's defaults are documented in the run README.
  • —Compute: ~$165 GPU total (~$3.40 regen, ~$3.62 cache, ~$35.20 train on 4×H100, rest eval).
  • —Source data: mlabonne/open-perfectblend (Apache-2.0); responses generated by Qwen/Qwen3-4B (Apache-2.0). Training/config code derives from DeepSpec (MIT). The Apache-2.0 tag reflects the data lineage; flip it if your use of the checkpoint needs otherwise.
  • —Optimizer resume states (7.4 GB/rank) were deliberately not uploaded — training is complete.