Team Ai
Modelpublic

Mapika/decider-12b

sourceHugging Faceapache-2.0updated 7d agoView on Hugging Face
4likes247downloads
Model Card

decider-12b: typed decisions with calibrated probabilities in one forward pass, on Gemma-4-12B-it

decider-12b is google/gemma-4-12B-it (revision 707f0a3b) read through the decider readout.

  • —Input: a state and one or more typed questions (Choice, Noul yes/no, Score), each with an explicit option list.
  • —Output: a probability distribution over the options for every question, from one forward pass. There is no decoding and no parsing.
  • —Prompt: the model's chat template, with thinking off. The option letter is read at the answer slot, with the model's final-logit softcapping applied.
  • —Temperatures: one per answer type, in decider_config.json.

Version 2 (2026-09-29, this revision): a merged LoRA fine-tune (state tracking) on top of the stock weights.

  • —Only the 328 attention and MLP projection matrices differ from Google's weights. The config, tokenizer and chat template are Google's files unchanged.
  • —Version 1 is stock Gemma-4-12B-it with no training. It is kept under the tag v1: Decider("Mapika/decider-12b", revision="v1").

Usage

Needs decider-ai>=1.7.0 (1.7.0 adds Gemma's final-logit softcapping; older versions read this model too sharply).

python
from decider.infer import Decider
d = Decider("Mapika/decider-12b")
d.system_one(state, {"refund": {"type": "noul", "instructions": "Is the refund allowed under the policy?",
                                "criteria": {"true": "allowed", "false": "not allowed"}}})

Serving: DECIDER_MODEL=Mapika/decider-12b uvicorn decider.serve:app (POST /v1/systemone, the System One wire format).

Training (version 2)

  • —Method: LoRA rank 32, alpha 64, on q/k/v/o/gate/up/down. One epoch of 10,000 items (10.6M tokens), LR 5e-5 cosine, one B300, 45 minutes. The loss is cross-entropy over the option letters at the answer slot, in the same chat layout as serving.
  • —Data:
  • —6,000 generated state-tracking decisions: event logs over seven invented domains (library loans, parking permits, hotel rooms, sprint boards, vehicle fleets, licence seats, lab samples). They include rejected entries, VOID and CORRECTION entries, aliases and out-of-order exports, rendered as logs, tables, prose and JSON.
  • —The gold answers are computed by replaying the log.
  • —Of 104 yes/no rows answered blind by GLM-5.3-Flash, 101 agree with the computed gold.
  • —4,000 rows from our earlier training data, replayed:
  • —836 human-labelled public datasets;
  • —2,047 from our own generated decision families, written without reading JevBench items;
  • —1,117 teacher-written document questions.
  • —Not used for training or selection: JevBench items, public or sealed; Decision Index rows. The checkpoint and the temperatures were chosen on our own held-out sets only.

Temperatures

typev2 Tv1 Tfitted on (our own rows)
Choice1.54.0NLL and top-label ECE over 2,113 held-out Choice rows
Noul0.051.0chance-corrected v1.5 Noul score on 685 rows plus the 204-item fit half of a fresh yes/no set
Score1.03.5chance-corrected Score competence on 246 rows
  • —The fine-tune leaves the model less confident, so every temperature is lower than in v1.
  • —Noul is read sharp. JevBench v1.5 counts a yes/no answer with P(yes) between 0.2 and 0.8 as wrong, and at T 0.05 about 1% of answers fall in that band.

Measurements (decider-ai 1.8.0, one B300, each version at its own temperatures)

setv2v1
fresh 399-item yes/no set, test half (195), v1.5 chance-corrected Noul score71.658.8
same, argmax accuracy / share of answers between 0.2 and 0.8 / ECE0.856 / 1 % / 0.1440.805 / 6 % / 0.181
same set, 80 generated state-tracking logs (another generator than training), argmax0.7500.675
same set, 319 written items, argmax0.8710.859
held-out generated decisions (1,500) and teacher questions (944): Choice accuracy0.6150.588
same rows: Noul accuracy / chance-corrected Noul score0.790 / 57.10.740 / 38.7
same rows: chance-corrected Score competence61.260.1
human-labelled Choice rows (600): accuracy / ECE0.847 / 0.0150.845 / 0.024
JevBench public items, accuracy easy / standard / hard1.000 / 0.986 / 0.7121.000 / 0.986 / 0.730
JevBench public hard tier, top-label ECE0.1470.098
  • —The public JevBench items were read for measurement only.
  • —On the hard tier (111 items), v2 answers 79 correctly and v1 answers 81.
  • —Hard-tier calibration is worse in v2.

Limitations

  • —The base model's strengths and failure modes carry over.
  • —The fine-tune targets state tracking over event logs. Gains elsewhere come from the replayed rows and are 1-5 points.
  • —On the JevBench public hard tier, v2 is not better than v1: 79 vs 81 correct of 111, and ECE 0.147 vs 0.098.
  • —Per-dataset calibration still varies: knowledge multiple choice is underconfident, and hard reasoning is overconfident.

Licence

Weights: derived from google/gemma-4-12B-it, Apache-2.0 as published by Google. Readout code: Apache-2.0 (github.com/Mapika/decider).