Team Ai
Modelpublic

Beko2210/statim-decide-multilingual-base

sourceHugging Faceotherupdated 7h agoView on Hugging Face
2likes174downloads
Model Card

Statim Decide Multilingual Base

Licence. This version was trained partly on data that is non-commercial, ShareAlike or under an unknown licence (the same training mixture as 0.7.0; findings in DATA_LICENSES.md). The weights are offered only under PolyForm Noncommercial 1.0.0. A version trained only on cleared data will follow.

A decision model for Statim, the native C++ engine for typed decisions: ask any text a choice, a score or a yes/no question and get calibrated answers from one forward pass, on CPU or GPU, without Python at runtime. Version 0.10.0, fine-tuned from `convaiinnovations/laya-multilingual` (mmBERT-base encoder).

[▶ Try it live in your browser](https://huggingface.co/spaces/Beko2210/statim): this model on a free CPU, no install and no key.

<img src="https://huggingface.co/Beko2210/statim-decide-multilingual-base/resolve/main/media/card.gif" width="100%" alt="Fourteen decision categories on held-out data: each bar grows from 0.7.0 to this version; the mean rises from 0.748 to 0.825.">

<video controls preload="none" width="100%" poster="https://beko2210.github.io/statim/images/film-16x9.webp" src="https://beko2210.github.io/statim/video/statim-flagship-60s-16x9.mp4"></video>

One support ticket, three typed answers, one forward pass: the 60-second film.

Quick start

sh
# Statim release binary: https://github.com/BEKO2210/statim/releases
huggingface-cli download Beko2210/statim-decide-multilingual-base statim-decide-multilingual-base-q8_0.gguf --local-dir models
./statim serve -m multilingual=models/statim-decide-multilingual-base-q8_0.gguf --port 8080
curl -s localhost:8080/v1/systemone -d '{
  "state": {"subject": "Duplicate charge on invoice #4411",
            "body": "We were billed twice for March. Please refund the second charge."},
  "questions": {
    "department": {"type": "choice", "instructions": "Which department should handle this?",
      "criteria": {"billing": "invoices, refunds", "technical": "bugs, outages", "sales": "pricing, contracts"}},
    "urgency": {"type": "score", "instructions": "How urgent is this request?",
      "criteria": ["not urgent", "soon", "critical"]},
    "refund": {"type": "noul", "instructions": "Does the user explicitly request a refund?"}}}'

Open http://127.0.0.1:8080/ for the playground. API reference: docs/API.md.

Files

FileSizeUse
statim-decide-multilingual-base-f32.gguf0.91 GBreference precision, exact on GPU
statim-decide-multilingual-base-q8_0.gguf0.36 GBsmaller; slower than f32 on AVX2 CPUs, faster on ARM dotprod and CUDA
checkpoint/0.68 GBLaya-format checkpoint for fine-tuning and the Python reference

Checksums in SHA256SUMS. The f32 file reproduces the reference implementation within 1e-4 on Statim's parity tests. q8_0 is 2.5x smaller with slightly different logits. Whether it is faster depends on the hardware: on x86-64 CPUs with AVX2 but without int8 dot-product instructions it is slower than f32; on ARM CPUs with dot-product instructions and with CUDA it is faster (measurements).

Evaluation

Measured by Statim's no-harm gate (`tools/finetune/gate.py`) on held-out test data the model selection never looked at. Against the previous release 0.7.0 (same held-out item pool), on 91 held-out suites: 6 significant gains, 85 within noise, 0 regressions (paired exact McNemar tests; gains and regressions are separately significant after Holm-Bonferroni over all suites).

SuiteRoleThis model0.7.0Protocol
typed-decisionstrained0.77500.7630test split, first 2,000 decisions; its train split is replay data
Banking77trained0.91850.9140test split, first 2,000 rows, all 77 intents in one question
MASSIVE intentstrained0.81610.7995mean over 12 languages, 150 seeded stratified test rows each
AG Newsheld out0.92100.9295zero-shot (never trained on), first 2,000 test rows
DAIR Emotionheld out0.53050.5040zero-shot, first 2,000 test rows
HWU64 intentsheld out0.85330.7867English, 150 rows; rows overlapping MASSIVE removed
SIB-200 topicsheld out0.79670.7817zero-shot, mean over 4 languages, 150 rows each
Sentimentheld out0.63110.6050zero-shot, mean over 12 languages, 150 rows each
HateCheckheld out0.64610.6358zero-shot, mean over 11 languages, 150 rows each
Belebele readingheld out0.28000.2583zero-shot, mean over 4 languages, 150 rows each

Decision categories

One held-out suite per decision category, built from splits of the training sources that the mixture never loads; any text that also occurs in the training mixture is dropped. 150 items per language, macro over languages.

CategoryLanguagesThis model0.7.0
complainten0.8200.767
emotionde, en, es, fr, hi, pt, ru, zh0.6610.589
fact checken0.4670.313
formalityja, tr0.9530.773
intenten, nl, tr0.7930.753
nlien, ja, tr0.7690.747
piiar, de, en, es, fr, it, ja, nl, ru, sv, zh0.9090.856
readingen0.9330.927
safetyen0.8800.727
sentimenten, zh0.8770.800
similaritypt0.8800.833
stanceen0.9800.893
topicen0.6470.607
urgencyen0.9870.893

Published systems under the same protocol, for orientation: typed-decisions meraGPT 0.768, laya-typed-decisions 0.766, Jev 0.727; AG News zero-shot Laya 0.950, GPT-3 (CARP) 0.926, Jev 0.881; Banking77 supervised MPNet 0.941; MASSIVE XLM-R base 0.857 over 12 languages (full train set). Sources: docs/ROADMAP.md.

<details><summary>All 91 held-out suites</summary>

SuiteAccuracyRows
amazon_massive_intent/ar0.7400150
amazon_massive_intent/de0.8267150
amazon_massive_intent/en0.8467150
amazon_massive_intent/es0.8333150
amazon_massive_intent/fr0.8067150
amazon_massive_intent/hi0.7800150
amazon_massive_intent/it0.8133150
amazon_massive_intent/ja0.8333150
amazon_massive_intent/pl0.8333150
amazon_massive_intent/ru0.8467150
amazon_massive_intent/tr0.8133150
amazon_massive_intent/zh-CN0.8200150
belebele/ar0.2600150
belebele/de0.3200150
belebele/en0.3133150
belebele/hi0.2267150
categories:complaint/en0.8200150
categories:emotion/de0.4867150
categories:emotion/en0.6400150
categories:emotion/es0.6200150
categories:emotion/fr0.8267150
categories:emotion/hi0.8600150
categories:emotion/pt0.5600150
categories:emotion/ru0.7400150
categories:emotion/zh0.5533150
categories:fact_check/en0.4667150
categories:formality/ja0.9067150
categories:formality/tr1.0000150
categories:intent/en0.8467150
categories:intent/nl0.5333150
categories:intent/tr1.0000150
categories:nli/en0.8000150
categories:nli/ja0.6600150
categories:nli/tr0.8467150
categories:pii/ar0.9333150
categories:pii/de0.8867150
categories:pii/en0.9333150
categories:pii/es0.8867150
categories:pii/fr0.9133150
categories:pii/it0.9000150
categories:pii/ja0.9267150
categories:pii/nl0.8533150
categories:pii/ru0.9667150
categories:pii/sv0.8733150
categories:pii/zh0.9267150
categories:reading/en0.9333150
categories:safety/en0.8800150
categories:sentiment/en0.8533150
categories:sentiment/zh0.9000150
categories:similarity/pt0.8800150
categories:stance/en0.9800150
categories:topic/en0.6467150
categories:urgency/en0.9867150
farstail/fa0.6533150
go_emotions/en0.4533150
hwu64/en0.8533150
indonli/id0.7000150
multi_hatecheck/ar0.6667150
multi_hatecheck/de0.6733150
multi_hatecheck/en0.6467150
multi_hatecheck/es0.6467150
multi_hatecheck/fr0.6600150
multi_hatecheck/hi0.5800150
multi_hatecheck/it0.6200150
multi_hatecheck/nl0.6467150
multi_hatecheck/pl0.6533150
multi_hatecheck/pt0.6733150
multi_hatecheck/zh0.6400150
multilingual_sentiments/ar0.6333150
multilingual_sentiments/de0.5600150
multilingual_sentiments/en0.7200150
multilingual_sentiments/es0.6000150
multilingual_sentiments/fr0.6267150
multilingual_sentiments/hi0.5400150
multilingual_sentiments/id0.7800150
multilingual_sentiments/it0.6467150
multilingual_sentiments/ja0.6800150
multilingual_sentiments/ms0.5333150
multilingual_sentiments/pt0.6400150
multilingual_sentiments/zh0.6133150
semrel/ar0.2733150
semrel/en0.2800150
semrel/hi0.2933150
sib200/ar0.7800150
sib200/de0.8267150
sib200/en0.8467150
sib200/hi0.7333150
test/ag_news0.92102000
test/banking770.91852000
test/emotion0.53052000
test/typed_decisions0.77502000

</details>

Reproduce these numbers: REPRODUCE.md.

Training

A uniform weight average (model soup, Wortsman et al., 2022) of 3 runs of train_multitask.py --clean from checkpoint laya-multilingual-big1 (0.4.0), which differ only in seed and batch order (best epochs 12/raw, 10/raw, 7/ema, each selected on validation data only); the temperatures were refitted on the validation items afterwards (--calibrate-only). Training data: the 0.7.0 mixture (Banking77, MASSIVE, typed-decisions replay, a tasksource mixture, Nemotron-Safety-Guard, IndicGuard, MINDS-14, SNIPS and further sources), including the sources the 2026-10-03 licence audit found non-commercial, ShareAlike or unlicensed; every source and finding is listed in DATA_LICENSES.md. Evaluation test rows were removed from the training data.

Intended use and limits

  • —Classification-style decisions over short texts and JSON: routing, triage, moderation, intent, yes/no checks, ordinal ratings. It does not generate text.
  • —Accuracy varies by task and language (see the table). Reading comprehension (Belebele) and semantic similarity are weak; do not use it for them without your own evaluation.
  • —Use the confidence: Statim's min_confidence option marks low-confidence answers with escalate: true so a person can review them. Do not automate high-stakes decisions about people without human review.

Licence

The weights may be used only under PolyForm Noncommercial 1.0.0 (text in LICENSE-MODEL.md); the Small Business, Free Trial and commercial licences do not apply to this version. The Statim engine is Apache-2.0.

Built on Laya (Apache-2.0) and mmBERT-base (MIT). Training data attribution: Banking77 (Casanueva et al., 2020, PolyAI), MASSIVE (FitzGerald et al., 2022, Amazon), and the CC-BY sources in DATA_LICENSES.md. Statim is independent and not affiliated with the Laya authors.