Team Ai
Modelpublic

Kodep/jev-topic-head-example

sourceHugging Faceapache-2.0updated 2d agoView on Hugging Face
0likes12downloads
Model Card

example-repro — jev-topic-head training, reproduced from the public spec

A self-contained reimplementation of the training described in ../jev-topic-head/README.md (the scaffold), written without the original code: Qwen3-Embedding-0.6B + LoRA(r16) + prototype head (56 anchors + learned scale), on 279/85/105-item synthetic data matching the real split structure.

This validates the pipeline, not the numbers. The real labels live in the private Kodep/jev-topic-data (school-email text, not pushed yet). When they land, point config.yaml: data.dir at them and rerun; everything else stays.

Try it (no training needed)

The repo ships the seed-0 weights — everything we created, nothing we didn't:

weights/seed0/lora/adapter_model.safetensors   ~39 MB   10.1M LoRA params (fixes the encoder)
weights/seed0/lora/adapter_config.json                    how to stack the adapter (r, targets, base)
weights/seed0/head/head.pt                     ~230 KB  56 prototype vectors + scale (the head)

The base model is intentionally not here — Qwen/Qwen3-Embedding-0.6B is public; from_pretrained fetches it from huggingface.co on first load. Ship the delta, not the pantry.

bash
hf download Kodep/jev-topic-head-example --local-dir .
pip install torch transformers peft        # plus internet for the one-time base download
python predict.py --weights weights/seed0 "Soccer practice moved to the north field"
# {"topic": "soccer", "confidence": 0.96, "top3": [...]}

Reminder: these weights trained on synthetic data — a working demo, not the real classifier.

Files

filewhat
config.yamlevery hyperparameter from the spec checklist
make_toy_data.pysynthetic 56-topic dataset (same shapes/skew/14-unseen-topics/hard-pairs)
train_head.pythe training loop: seed protos from descriptions, LoRA, 2 AdamW groups, warmup→decay, 1/√count sampler, early stop on val macro-F1
evaluate.pyper-item dump + top-1/macro-F1/confidence buckets (eval/ schema)
predict.pyinference with the shipped seed-0 weights (no training stack needed)

Run (on Sunny, 4090)

bash
cd ~/jev-example
.venv/bin/python make_toy_data.py
.venv/bin/python train_head.py --config config.yaml --seed 0   # downloads base model once
.venv/bin/python evaluate.py --config config.yaml --run runs/seed0

Spec fidelity notes

  • —Implemented from the checklist: r16/α32/dropout 0.05 on q/k/v/o + gate/up/down; protos seeded from description embeddings; scale init 20; CE + label smoothing 0.05; lora 3e-4 wd .01 / head 2e-3 no wd; warmup 10% → linear decay to 5%; batch 16; ≤12 epochs; early stop on val macro-F1 (patience 3); sampler 1/√count × hard-pair 2.
  • —Inferences the spec left open, marked in code: last-token pooling (official Qwen3-Embedding recipe), bf16 = bf16 base weights + fp32 head.
  • —Show/hide rules (rules_table.json) are not modeled — they're a fixed downstream table, not part of the head.

Results (synthetic data, 4090, 2026-10-04)

  • —Trainable: 10,092,544 LoRA + 57,345 head params (1.67% of 605.9M) — matches the spec's shape: head = 56×1024 + 1 exactly.
  • —Early stop at epoch 4 (patience 3); best val macro-F1 0.949 (seed 0).
  • —Test vs labels, identical across seeds 0/1/2: top-1 0.971, macro-F1 0.956; 78/105 items ≥0.9 confidence; zero errors above 0.9 confidence — the same "errors live in the low-confidence bucket" pattern the real report claims.
  • —These numbers validate plumbing only: synthetic labels are clean, the real ones are Jev's noisy silver answers (real span: macro-F1 0.49–0.59).

Status

Code path fully exercised end-to-end (3 seeds + evals, artifacts in results/). Awaiting the real Kodep/jev-topic-data push → set config.yaml: data.dir to it and rerun unchanged to attempt the actual 92/105 · 0.572 · 90.5% numbers.