Team Ai
Modelpublic

stanley-nv/toolrouter

sourceHugging Facemitupdated 3d agoView on Hugging Face
0likes
Model Card

ToolRouter

A tiny (< 8 MB), CPU-only, sub-millisecond tool router. Give it a user query and a list of tools — each with a name and a description (optionally a few example queries) — and it returns which tool to use, with per-tool probabilities and a confidence, in a strict, structured choice format:

json
{"state": "will it rain tomorrow in Paris",
 "questions": {"tool": {"type": "choice", "instructions": "Pick the single best tool.",
   "criteria": {"weather": "Forecasts.", "news": "Headlines.", "other": "No tool needed."}}}}
json
{"model": "toolrouter-7",
 "answers": {"tool": {"type": "choice", "choice": "weather",
   "probabilities": {"weather": 0.999996, "news": 4e-06, "other": 0.0}, "confidence": 0.999996}}}

It is open-set: the tools come from the request, so it works for tools it has never seen. The 7 priority tools — wiki, websearch, weather, news, files, shell, other — additionally have a trained "identity" head and are the most accurate.

  • —Try it: the playground Space lets you edit the tool list and query and see the result — it runs the model in your browser (JavaScript port in web/, verified against the Python code on 363 cases).
  • —Data: `stanley-nv/tool-routing-data` (training, validation and test sets).
  • —Code, training and evaluation: in this repo (toolrouter/).

Quick start

bash
pip install numpy huggingface_hub            # inference needs only numpy (+ huggingface_hub to download)
git clone https://huggingface.co/stanley-nv/toolrouter && cd toolrouter
python -m toolrouter.cli "what's using port 8080"
python examples/demo_tools.py --tools examples/tools.json            # route over your own tool descriptions, interactive
python
from toolrouter import ToolRouter
router = ToolRouter.from_pretrained("stanley-nv/toolrouter")           # or ToolRouter("toolrouter.npz")

tools = {
    "stock_quotes": "Real-time share prices, tickers and market indices.",
    "smart_home":   "Control lights, thermostat, locks and other connected devices.",
    "calendar":     {"description": "Read and create meetings.", "examples": ["move my 3pm to Friday"]},   # examples are optional
    "other":        "No tool needed: chit-chat, writing, math, advice.",
}
router.choose("how is Google trading today", tools)
# {'type': 'choice', 'choice': 'stock_quotes', 'probabilities': {...}, 'confidence': 0.95}
router.handle(request)    # full request -> response, as above

How it works

score(query, tool) = 7·cos(q, t_full) + 3·cos(q, t_desc) + 2·cos(q, t_name) + (.5, .5, .5, .25)·lexical_overlap + identity_head

  • —q, t_*: frozen static word-piece embeddings from `minishlab/potion-base-8M` (256-d, MIT), stored int8 in the model file, mean-pooled and normalised; t_full = "name. description". With example queries, cos(q, t_full) becomes the max over description and examples.
  • —The semantic weights are fixed (hand-set from an ablation). Only the identity head is trained: a hashed n-gram bag plus the dense query embedding, one vector per priority tool, applied to the 7 priority tools (and aliases such as web, terminal). Training is episodic (query + target + random distractor tools with random name/description styles).
  • —Softmax with a calibrated temperature gives probabilities; confidence is the probability of the chosen tool.
  • —Files: toolrouter.npz (7.66 MB, includes the backbone). Pure-python WordPiece tokenizer, numpy inference, no PyTorch needed.

Results

Accuracy; all test data are held out (never used for training). Reproduce with python -m toolrouter.eval.evaluate --data-dir <dataset> --models toolrouter.npz --full.

SettingShipped modelMean of 3 seeds
7 priority tools, independent test routing/test2 (330 queries)0.9360.935
same, reworded tool descriptions0.9360.934
7 priority tools, routing/test (210 queries)0.9330.932
same, reworded tool descriptions0.9380.935
10 tools never seen in training (tools/test), all offered at once0.8900.890
same, random 5-tool slates0.9330.937
same, differently worded descriptions / name only0.780 / 0.750
descriptions shifted onto the wrong tools (sanity check)0.060
Predictions with confidence ≥ 0.95≈ 98 % correct (170 of 210 priority queries; 56 of 100 unseen-tool queries)
Latency (CPU, 10 tools)≈ 0.10–0.16 ms warm, ~2 ms for the first request

See `docs/REPORT.md` for the full development report, ablations and negative results (what did not help: learning the semantic path, fine-tuning on ToolRet, MLP head, MaxSim, IDF pooling, larger/compressed 32M backbone).

Intended use and limitations

  • —Pre-routing for agents/assistants: pick one of a handful of tools (best with ≤ 10 tools) before calling a bigger model, with a confidence you can threshold (e.g. fall back to an LLM below ~0.8).
  • —Descriptions matter. Name-only or very short descriptions cost 10–15 points; vague catch-all tools (other, websearch, wiki) can steal queries from specific tools when many tools are offered (all ~200 tools at once: ~0.5–0.6).
  • —Very short queries (1–2 words, e.g. git status) with vague or meaningless descriptions in a tiny tool list are the weakest case: the generic other tool can win. With the default tool descriptions such queries route correctly; measured over test queries, 2-tool slates (true tool + one random priority tool) are 98.5–99 % accurate even with meaningless descriptions.
  • —English only, single-turn, one tool per query (no multi-tool plans, no arguments extraction). wiki vs websearch vs news is inherently fuzzy; labelling conventions are in the dataset card.
  • —Evaluation caveat: the test sets were written by Claude (Anthropic) / the author, and the training data were largely written by Claude agents under explicit labelling conventions, so label noise and stylistic overlap are likely; expect lower accuracy on real traffic. Sampling error is about ±3 points. Some design choices were made while looking at the older test sets (routing/test, tools/test); routing/test2 is the cleanest number.
  • —Not evaluated for safety/fairness; do not use it as the sole gate for sensitive actions.

Training and reproducing

bash
git clone https://huggingface.co/datasets/stanley-nv/tool-routing-data ../tool-routing-data
pip install -r requirements-train.txt
TOOLROUTER_FORCE_IPV4=1 python -m toolrouter.training.backbone --out backbone/potion_8m.npz   # fetch + convert the frozen backbone
python -m toolrouter.training.train --data-dir ../tool-routing-data --backbone backbone/potion_8m.npz --out toolrouter.npz   # ~4 min on a GPU
python -m toolrouter.eval.evaluate --data-dir ../tool-routing-data --models toolrouter.npz --full
python -m unittest discover tests
# browser/JS port: python -m toolrouter.export_web && python tests/make_parity_cases.py --data-dir ../tool-routing-data && node web/test_parity.mjs tests/parity_cases.json web/model.bin

License and attribution

  • —Code and weights: MIT. The model file embeds an int8-quantised copy of the potion-base-8M embeddings (MIT, © MinishLab); please keep that attribution.
  • —Training data were generated by Claude (Anthropic) and the author; see the dataset card for provenance and the terms you should review before commercial use. The released training run does not use the ToolRet corpus (an earlier experiment did; it did not help and is excluded).

Citation

bibtex
@misc{toolrouter2026,
  title  = {ToolRouter: a tiny open-set tool router with a structured choice interface},
  author = {Stanley},
  year   = {2026},
  url    = {https://huggingface.co/stanley-nv/toolrouter}
}

Related work: model2vec / potion, Bag of Tricks for Efficient Text Classification, ToolRet.