Team Ai
Modelpublic

jerryyan/TraceML-State-Labeler

sourceHugging Faceapache-2.0updated 9d agoView on Hugging Face
0likes390downloads
Model Card

TraceML State Labeler (Qwen3-1.7B)

Moved: both labelers now live in one repo, jerryyan/TraceML-Labelers (state/ and action/). This repo is kept only so existing links keep working.

馃搫 Paper 路 馃 Dataset 路 馃捇 Toolkit 路 馃寪 Project page 路 Action labeler

The state labeler of TraceML (NeurIPS 2026, Evaluations & Datasets Track). Given one version of an ML solution's code, it returns the ML-pipeline stages the code contains as JSON: 8 coarse tags, fine tags from a closed list of 136 with a confidence each, a one-line summary and keywords.

It is Qwen3-1.7B fine-tuned on schema-constrained labels from a larger GPT teacher model. Together with the action labeler it labeled all 151,088 code versions in TraceML.

Use

The simplest route is the TraceML toolkit. It rebuilds the exact prompts from a run, fits long code into the context window, decodes greedily and parses the output:

bash
pip install "traceml-toolkit[label] @ git+https://github.com/JerryYan123/TraceML"
traceml analyze runs/my_run    # extract every code version, label states and transitions, report against human cohorts

Direct use with transformers; the prompt builders and the output parser come from the toolkit:

python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from traceml_toolkit.labeling import render
from traceml_toolkit.labeling.parse import parse_state_output

repo = "jerryyan/TraceML-State-Labeler"
device = "cuda" if torch.cuda.is_available() else "mps" if torch.backends.mps.is_available() else "cpu"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(
    repo, torch_dtype=torch.float32 if device == "cpu" else torch.bfloat16).to(device)

code = open("train.py").read()
rec = {"comp": "commonlitreadabilityprize", "group": "my-agent", "version_number": 1,
       "code_text": code, "code_lines": code.count("\n") + 1}
prompt = render.render(tok, render.system_prompt("state"), render.user_prompt("state", rec))
inputs = tok(prompt, return_tensors="pt", add_special_tokens=False).to(device)
out = model.generate(**inputs, max_new_tokens=2000, do_sample=False, temperature=None, top_p=None, top_k=None)
print(parse_state_output(tok.decode(out[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True)))

Decode greedily with thinking disabled, as above. generation_config.json keeps Qwen3's sampling defaults, so pass the greedy settings explicitly. The labels released in TraceML were produced with vLLM 0.8.5 in bf16, greedy decoding and prefix caching disabled.

Output

json
{"coarse_tags": ["data_io", "feature_eng", "model_def", "training_cfg", "validation_cv"],
 "fine_tags": [{"tag": "...", "parent": "model_def", "confidence": "high"}],
 "summary": "...", "keywords": ["..."]}

The coarse tags are data_io, feature_eng, model_def, training_cfg, ensemble_blend, validation_cv, inference_submit and infra_util. The full vocabulary is in `manifests/schemas/` of the dataset.

Files

The same files as `models/qwen3-1.7b-state/final/` in the dataset repository, without training_args.bin.

License

Apache-2.0, inherited from Qwen3.

Citation

bibtex
@inproceedings{yan2026traceml,
  title         = {TraceML: What Auto-Research Agents Miss in Long-Horizon ML Development},
  author        = {Yan, Jiarui and Sun, Weiwei and Li, Sijie and Li, Wenhan and Yang, Yiming},
  booktitle     = {Advances in Neural Information Processing Systems (NeurIPS), Track on Evaluations and Datasets},
  year          = {2026},
  eprint        = {2608.26086},
  archivePrefix = {arXiv}
}