Beko2210/statim-decide-multilingual-base
Statim Decide Multilingual Base
Licence. This version was trained partly on data that is non-commercial, ShareAlike or under an unknown licence (the same training mixture as 0.7.0; findings in DATA_LICENSES.md). The weights are offered only under PolyForm Noncommercial 1.0.0. A version trained only on cleared data will follow.
A decision model for Statim, the native C++ engine for typed decisions: ask any text a choice, a score or a yes/no question and get calibrated answers from one forward pass, on CPU or GPU, without Python at runtime. Version 0.10.0, fine-tuned from `convaiinnovations/laya-multilingual` (mmBERT-base encoder).
[▶ Try it live in your browser](https://huggingface.co/spaces/Beko2210/statim): this model on a free CPU, no install and no key.
<img src="https://huggingface.co/Beko2210/statim-decide-multilingual-base/resolve/main/media/card.gif" width="100%" alt="Fourteen decision categories on held-out data: each bar grows from 0.7.0 to this version; the mean rises from 0.748 to 0.825.">
<video controls preload="none" width="100%" poster="https://beko2210.github.io/statim/images/film-16x9.webp" src="https://beko2210.github.io/statim/video/statim-flagship-60s-16x9.mp4"></video>
One support ticket, three typed answers, one forward pass: the 60-second film.
Quick start
# Statim release binary: https://github.com/BEKO2210/statim/releases
huggingface-cli download Beko2210/statim-decide-multilingual-base statim-decide-multilingual-base-q8_0.gguf --local-dir models
./statim serve -m multilingual=models/statim-decide-multilingual-base-q8_0.gguf --port 8080
curl -s localhost:8080/v1/systemone -d '{
"state": {"subject": "Duplicate charge on invoice #4411",
"body": "We were billed twice for March. Please refund the second charge."},
"questions": {
"department": {"type": "choice", "instructions": "Which department should handle this?",
"criteria": {"billing": "invoices, refunds", "technical": "bugs, outages", "sales": "pricing, contracts"}},
"urgency": {"type": "score", "instructions": "How urgent is this request?",
"criteria": ["not urgent", "soon", "critical"]},
"refund": {"type": "noul", "instructions": "Does the user explicitly request a refund?"}}}'Open http://127.0.0.1:8080/ for the playground. API reference: docs/API.md.
Files
Checksums in SHA256SUMS. The f32 file reproduces the reference implementation within 1e-4 on Statim's parity tests. q8_0 is 2.5x smaller with slightly different logits. Whether it is faster depends on the hardware: on x86-64 CPUs with AVX2 but without int8 dot-product instructions it is slower than f32; on ARM CPUs with dot-product instructions and with CUDA it is faster (measurements).
Evaluation
Measured by Statim's no-harm gate (`tools/finetune/gate.py`) on held-out test data the model selection never looked at. Against the previous release 0.7.0 (same held-out item pool), on 91 held-out suites: 6 significant gains, 85 within noise, 0 regressions (paired exact McNemar tests; gains and regressions are separately significant after Holm-Bonferroni over all suites).
Decision categories
One held-out suite per decision category, built from splits of the training sources that the mixture never loads; any text that also occurs in the training mixture is dropped. 150 items per language, macro over languages.
Published systems under the same protocol, for orientation: typed-decisions meraGPT 0.768, laya-typed-decisions 0.766, Jev 0.727; AG News zero-shot Laya 0.950, GPT-3 (CARP) 0.926, Jev 0.881; Banking77 supervised MPNet 0.941; MASSIVE XLM-R base 0.857 over 12 languages (full train set). Sources: docs/ROADMAP.md.
<details><summary>All 91 held-out suites</summary>
</details>
Reproduce these numbers: REPRODUCE.md.
Training
A uniform weight average (model soup, Wortsman et al., 2022) of 3 runs of train_multitask.py --clean from checkpoint laya-multilingual-big1 (0.4.0), which differ only in seed and batch order (best epochs 12/raw, 10/raw, 7/ema, each selected on validation data only); the temperatures were refitted on the validation items afterwards (--calibrate-only). Training data: the 0.7.0 mixture (Banking77, MASSIVE, typed-decisions replay, a tasksource mixture, Nemotron-Safety-Guard, IndicGuard, MINDS-14, SNIPS and further sources), including the sources the 2026-10-03 licence audit found non-commercial, ShareAlike or unlicensed; every source and finding is listed in DATA_LICENSES.md. Evaluation test rows were removed from the training data.
Intended use and limits
- Classification-style decisions over short texts and JSON: routing, triage, moderation, intent, yes/no checks, ordinal ratings. It does not generate text.
- Accuracy varies by task and language (see the table). Reading comprehension (Belebele) and semantic similarity are weak; do not use it for them without your own evaluation.
- Use the confidence: Statim's
min_confidenceoption marks low-confidence answers withescalate: trueso a person can review them. Do not automate high-stakes decisions about people without human review.
Licence
The weights may be used only under PolyForm Noncommercial 1.0.0 (text in LICENSE-MODEL.md); the Small Business, Free Trial and commercial licences do not apply to this version. The Statim engine is Apache-2.0.
Built on Laya (Apache-2.0) and mmBERT-base (MIT). Training data attribution: Banking77 (Casanueva et al., 2020, PolyAI), MASSIVE (FitzGerald et al., 2022, Amazon), and the CC-BY sources in DATA_LICENSES.md. Statim is independent and not affiliated with the Laya authors.
