Team Ai
Modelpublic

benchmarkheaven/weiche-395m

sourceHugging Faceapache-2.0updated 14d agoView on Hugging Face
0likes
Model Card

Weiche-395M: a small local decision model for LLM auto-routing

Weiche (German for a railway switch) reads one user turn and decides which kind of model should handle it. It is ModernBERT-large (395M) with typed decision heads, trained for the open-source auto-model-router. One forward pass returns:

OutputMeaning
category9-way choice: coding, agentic, math, knowledge, longcontext, tooluse, design, summarisation, general
scalars[1] difficultyrubric difficulty, 0 = trivial … 1 = frontier (the router's 0–4 scale ÷ 4)
scalars[0] outcome difficultyshare of a model panel expected to fail (learned from measured outcomes)
scalars[2] stakescost of an unnoticed wrong answer, 0–1
noulP(needstools), P(needsvision), P(needslongcontext), P(follow_up)
succP(success) of a small, a mid and a strong model tier, learned from measured outcomes
routeexperimental: small / mid / strong given 14 numeric router-state features (cache warm/cold, latency budget, value of a correct answer, tier prices and latencies)

It runs on CPU through ONNX Runtime (no PyTorch needed) and in the browser through onnxruntime-web (WebGPU).

Results

All numbers come from held-out data that was never used for training. The eval scripts and raw result JSON files are in eval/ and in the GitHub repo. The baselines are the router's existing decision paths: its heuristic, the local Laya 421M classifier (backend: local) and the hosted Jev API (backend: hosted). Jev was called only for this evaluation. Its outputs were never used for training, calibration or data selection.

The "+ fitted calibration" rows give the trait-only classifiers a per-tier logistic calibration from their traits to P(success). We fitted it on a separate calibration split, which is more help than the router's fixed success curve gives them. Weiche's success heads need no calibration. Routing picks, for each request, argmax P(success_t) − λ·cost_t and sweeps λ, as the router's expected-cost policy does. APGR is the average share of the quality gap between always-small and always-strong that is recovered across 20 cost budgets.

ClassifierCategory acc. (held-out, n=1134)Macro-F1Category acc. on the router's own 70 held-out tasksneeds_tools acc.Difficulty rank corr.
Weiche-395M (this model)88.5 %0.87294.9 % (59 tasks)98.4 %0.739
Weiche-0.6B arm (Qwen3-0.6B, not released)88.5 %0.87398.3 % (59 tasks)99.2 %0.736
hosted Jev 1.13 (API)85.8 %0.84379.7 % (59 tasks)96.0 %0.753
Laya 421M (router local backend)9.3 %0.11233.9 % (59 tasks)80.5 %0.297

RouterBench held-out (n=1500; tiers Mistral-7B / Mixtral-8x7B / GPT-4-1106)

Decision pathAvg. gap recovered (APGR)Quality at 10 % of strong-tier costat 25 %Cost share to recover 80 % of the gap95 % of the gap
Router heuristic (no classifier)0.49447.2 %47.2 %100.0 %100.0 %
Laya 421M (router local backend) + fitted calibration0.81248.9 %53.8 %38.9 %79.5 %
hosted Jev 1.13 (API) + fitted calibration0.85050.3 %59.6 %28.2 %86.2 %
Weiche-395M traits + fitted calibration0.83448.7 %58.2 %41.0 %88.1 %
Weiche-395M success heads0.86251.4 %59.0 %31.4 %65.2 %
Weiche-0.6B arm success heads0.86451.8 %59.0 %30.4 %68.7 %

Always-small quality 28.4 %, always-mid 47.2 %, always-strong 68.6 %. An oracle that knows every outcome matches always-strong quality at 17.1 % of its cost and tops out at 77.2 %.

Measured router-style tasks finished after the data freeze (n=1103; tiers Mistral-Nemo-12B / Gemma-4-31B / DeepSeek-V3.2)

Decision pathAvg. gap recovered (APGR)Quality at 10 % of strong-tier costat 25 %Cost share to recover 80 % of the gap95 % of the gap
Router heuristic (no classifier)0.45942.5 %42.5 %60.8 %60.8 %
Laya 421M (router local backend) + fitted calibration0.74748.2 %64.1 %47.5 %56.2 %
hosted Jev 1.13 (API) + fitted calibration0.73448.9 %56.4 %50.5 %56.9 %
Weiche-395M traits + fitted calibration0.77252.4 %62.4 %44.0 %54.4 %
Weiche-395M success heads0.80352.9 %72.4 %44.5 %52.7 %
Weiche-0.6B arm success heads0.80554.0 %71.4 %41.9 %52.7 %

Always-small quality 42.5 %, always-mid 95.0 %, always-strong 94.0 %. An oracle that knows every outcome matches always-strong quality at 38.3 % of its cost and tops out at 98.6 %.

CPU latency per decision, 4 threads, same busy hostp50p95
Weiche-395M ONNX int4288 ms393 ms
Weiche-395M ONNX int8 (lossy)243 ms382 ms
Weiche-395M ONNX fp16386 ms545 ms
Laya 421M (PyTorch, 7 questions)11061 ms12308 ms

The difficulty rank correlation compares against the averaged rubric labels of two model families. On that measure Weiche and hosted Jev are close (0.739 vs 0.753).

Use it in auto-model-router

yaml
policy:
  classifier: {backend: local-route-head, model: benchmarkheaven/weiche-395m, variant: fp16, threads: 2}

variant is fp16 (default, matches fp32 within 1e-3), int4 (423 MB, 96 % category agreement with fp32) or int8 (lossy for this architecture, 92 % agreement; not recommended). The first request downloads the chosen ONNX file once.

Use it directly (ONNX Runtime, Python)

python
import numpy as np, onnxruntime as ort
from huggingface_hub import snapshot_download
from tokenizers import Tokenizer

d = snapshot_download("benchmarkheaven/weiche-395m", allow_patterns=["model_fp16.onnx*", "tokenizer.json", "routerhead.py"])
tok = Tokenizer.from_file(f"{d}/tokenizer.json"); tok.enable_truncation(512)
sess = ort.InferenceSession(f"{d}/model_fp16.onnx", providers=["CPUExecutionProvider"])

def render(request, context=""):   # the training-time input format (see routerhead.py)
    if len(request) > 2400:
        request = request[:1400] + "\n[...]\n" + request[-1000:]
    return f"Context: {context.strip()[:600] or '(new conversation)'}\nRequest: {request}"

ids = np.array([tok.encode(render("Fix the flaky test in tests/test_api.py and run the suite",
                                  "3 messages, 2 from the user. Tools: Bash, Read, Edit.")).ids])
category, scalars, noul, succ, route = sess.run(None, {"input_ids": ids, "attention_mask": np.ones_like(ids),
                                                       "state": np.zeros((1, 14), np.float32)})

PyTorch: encoder/ holds the fine-tuned ModernBERT (HF format) and heads.safetensors holds the heads. routerhead.py defines RouterHead and render:

python
from transformers import AutoModel, AutoTokenizer
from safetensors.torch import load_file
from routerhead import RouterHead, probs, render
enc = AutoModel.from_pretrained("benchmarkheaven/weiche-395m", subfolder="encoder")
model = RouterHead(enc, 1024).eval()
model.load_state_dict(load_file("heads.safetensors"), strict=False)

Browser: model_int4.onnx ran under onnxruntime-web 1.30 with the WebGPU execution provider (eval/webgpu_smoke_int4.txt). Its outputs matched CPU ONNX Runtime within 1e-4. That check used Chrome's software WebGPU adapter, so it proves operator coverage, not speed. model_fp16.onnx needs a GPU with shader-f16; the software adapter has none, so the fp16 browser path is not yet verified on real hardware.

Training data

We did not use any TypeSafe/Jev output, any sealed benchmark item, or any model response text. The data combines:

  • —RouterBench (MMLU, HellaSwag, GSM8K, WinoGrande and MBPP subsets only): measured correctness and cost of 11 LLMs per prompt.
  • —EmbedLLM (Apache-2.0): measured correctness of 112 open models per question.
  • —RouteLLM gpt4_dataset (Apache-2.0; only UltraChat, Anthropic-HH, FLAN and TruthfulQA prompts): GPT-4-judged adequacy of Mixtral's answer.
  • —Chatbot Arena 55k (Apache-2.0): human preference between a weaker and a stronger model.
  • —Synthetic router turns (ours, CC-BY-4.0): coding agents, IDE chat, support copilots, batch jobs and more, each with the router's own context summary. DeepSeek-V3.2 and Qwen3.8-27B wrote them. Two model families labelled them blind (author, plus Gemma-4-31B or GLM-5.1 as critic). We keep a row only when both agree on the category; any other field where they disagree is masked.
  • —Measured router-style tasks (ours, CC-BY-4.0): procedurally generated requests (maths, dates, code output, logic, tables, strings, money) whose gold answer code computes. We ran each one on Mistral-Nemo-12B, Gemma-4-31B and DeepSeek-V3.2 and recorded correctness, tokens, cost and latency.

We split by a hash of the normalised question text into train 86 %, calib 4 % and test 10 %. The same question can never appear in two splits. The pool owner checked the frozen files against a sealed router benchmark pool (exact match, 8-gram containment and MiniLM cosine) and found zero overlaps. Every source and its licence is listed in DATA-LICENCES.md.

Training: 2 epochs, AdamW (encoder 3e-5, heads 1e-3), batch 32, max 512 tokens, bf16, on one RTX PRO 6000 (33 minutes). Both arms together (this model and a Qwen3-0.6B arm) cost about USD 3.40 of GPU time.

Limits (read before relying on it)

  • —In-distribution advantage. The held-out test shares its source distributions with the training data, while Laya and Jev are zero-shot. The router's own 70 held-out tasks are the cleanest out-of-distribution check. There Weiche gets category right 94.9 % of the time (Jev 79.7 %). Its difficulty ranking on those tasks is weak (rank correlation 0.19; Jev 0.30).
  • —The success heads predict abstract tiers. Small / mid / strong mean the panels above (7B–13B, 30B–70B-class and GPT-4-class models). With a very different model line-up, use the traits, or recalibrate succ on your own outcomes.
  • —The route head is experimental. On the held-out test, exact cost arithmetic over the success heads beat the learned route head (mean regret USD 0.035 vs 0.050 per decision). On post-freeze tasks the route head was better (0.067 vs 0.083). The router computes cache and latency economics exactly, so feeding it succ is the recommended use.
  • —English-centric (ModernBERT). Other languages are untested.
  • —int8 is lossy for this architecture. Use fp16 or int4.
  • —Not a JevBench result and not evaluated on any sealed benchmark.

Licence

Apache-2.0, inherited from ModernBERT-large. The synthetic and measured data we created are CC-BY-4.0. Public datasets keep their own licences (see DATA-LICENCES.md).