Kodep/jev-topic-head-example
example-repro — jev-topic-head training, reproduced from the public spec
A self-contained reimplementation of the training described in ../jev-topic-head/README.md (the scaffold), written without the original code: Qwen3-Embedding-0.6B + LoRA(r16) + prototype head (56 anchors + learned scale), on 279/85/105-item synthetic data matching the real split structure.
This validates the pipeline, not the numbers. The real labels live in the private Kodep/jev-topic-data (school-email text, not pushed yet). When they land, point config.yaml: data.dir at them and rerun; everything else stays.
Try it (no training needed)
The repo ships the seed-0 weights — everything we created, nothing we didn't:
weights/seed0/lora/adapter_model.safetensors ~39 MB 10.1M LoRA params (fixes the encoder)
weights/seed0/lora/adapter_config.json how to stack the adapter (r, targets, base)
weights/seed0/head/head.pt ~230 KB 56 prototype vectors + scale (the head)The base model is intentionally not here — Qwen/Qwen3-Embedding-0.6B is public; from_pretrained fetches it from huggingface.co on first load. Ship the delta, not the pantry.
hf download Kodep/jev-topic-head-example --local-dir .
pip install torch transformers peft # plus internet for the one-time base download
python predict.py --weights weights/seed0 "Soccer practice moved to the north field"
# {"topic": "soccer", "confidence": 0.96, "top3": [...]}Reminder: these weights trained on synthetic data — a working demo, not the real classifier.
Files
Run (on Sunny, 4090)
cd ~/jev-example
.venv/bin/python make_toy_data.py
.venv/bin/python train_head.py --config config.yaml --seed 0 # downloads base model once
.venv/bin/python evaluate.py --config config.yaml --run runs/seed0Spec fidelity notes
- Implemented from the checklist: r16/α32/dropout 0.05 on q/k/v/o + gate/up/down; protos seeded from description embeddings; scale init 20; CE + label smoothing 0.05; lora 3e-4 wd .01 / head 2e-3 no wd; warmup 10% → linear decay to 5%; batch 16; ≤12 epochs; early stop on val macro-F1 (patience 3); sampler 1/√count × hard-pair 2.
- Inferences the spec left open, marked in code: last-token pooling (official Qwen3-Embedding recipe), bf16 = bf16 base weights + fp32 head.
- Show/hide rules (
rules_table.json) are not modeled — they're a fixed downstream table, not part of the head.
Results (synthetic data, 4090, 2026-10-04)
- Trainable: 10,092,544 LoRA + 57,345 head params (1.67% of 605.9M) — matches the spec's shape: head = 56×1024 + 1 exactly.
- Early stop at epoch 4 (patience 3); best val macro-F1 0.949 (seed 0).
- Test vs labels, identical across seeds 0/1/2: top-1 0.971, macro-F1 0.956; 78/105 items ≥0.9 confidence; zero errors above 0.9 confidence — the same "errors live in the low-confidence bucket" pattern the real report claims.
- These numbers validate plumbing only: synthetic labels are clean, the real ones are Jev's noisy silver answers (real span: macro-F1 0.49–0.59).
Status
Code path fully exercised end-to-end (3 seeds + evals, artifacts in results/). Awaiting the real Kodep/jev-topic-data push → set config.yaml: data.dir to it and rerun unchanged to attempt the actual 92/105 · 0.572 · 90.5% numbers.
