qnaug/embeddinggemma2-decision
embeddinggemma2-decision
A decision model built on the text encoder of google/embeddinggemma-2 (~270M parameters). It reads an input, a set of questions and their options in a single forward pass and returns a calibrated probability for every option. It never generates text. The head follows the Clef design: option and question token spans are mean-pooled and scored by a small MLP.
Question types: choice (pick one of named options), noul (true / false), score (ordered levels).
Usage
from huggingface_hub import hf_hub_download
import importlib.util
spec = importlib.util.spec_from_file_location("decision_model", hf_hub_download("qnaug/embeddinggemma2-decision", "decision_model.py"))
dm_mod = importlib.util.module_from_spec(spec); spec.loader.exec_module(dm_mod)
dm = dm_mod.DecisionModel("qnaug/embeddinggemma2-decision")
dm.decide("Hi, I was charged twice for invoice #4411. Please refund the duplicate today.", {
"team": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": "invoices, payments, refunds", "technical": "bugs, outages", "sales": "pricing"}},
"refund": {"type": "noul", "instructions": "The customer asks for a refund."},
})domain="extra" (default) uses the temperatures calibrated on natural text; domain="typed" uses those calibrated on JSON state like typed-decisions. Run in float32 or bfloat16, never float16 (the base model returns NaN).
Results
typed-decisions test accuracy 74.3% (random 31.8%, majority label 47.9%).
typed-decisions (test)
Accuracy 74.3%, ECE 0.029 after calibration.
boolq, mnli, banking77, clinc150 (test)
Accuracy 80.4%, ECE 0.039 after calibration.
Limitations
- Trained on English data only. Other languages (e.g. Vietnamese) are untested and the calibration does not carry over.
- Calibration is only valid for data like the training sources. Recalibrate on a few hundred examples of your own data before relying on the probabilities.
- Yes/no reading comprehension on free text is the weakest part (boolq / mnli ≈ 70%).
- Labels in typed-decisions have low inter-annotator agreement, so accuracy has a ceiling well below 100%.
Training
Text encoder of google/embeddinggemma-2 with the vocabulary embedding frozen, trained with soft-label cross-entropy (typed-decisions) and one-hot labels (other datasets). Intent datasets show the gold label plus random distractors, so the model learns to choose among arbitrary option lists. Per-source temperature scaling on a held-out split.
Dataset licenses differ (typed-decisions Apache-2.0, BoolQ CC BY-SA 3.0, banking77 CC BY 4.0, CLINC150 CC BY 3.0, MNLI mixed); check them for your use.
