Team Ai
Modelpublic

qnaug/embeddinggemma2-decision

sourceHugging Faceapache-2.0updated 2d agoView on Hugging Face
0likes
Model Card

embeddinggemma2-decision

A decision model built on the text encoder of google/embeddinggemma-2 (~270M parameters). It reads an input, a set of questions and their options in a single forward pass and returns a calibrated probability for every option. It never generates text. The head follows the Clef design: option and question token spans are mean-pooled and scored by a small MLP.

Question types: choice (pick one of named options), noul (true / false), score (ordered levels).

Usage

python
from huggingface_hub import hf_hub_download
import importlib.util
spec = importlib.util.spec_from_file_location("decision_model", hf_hub_download("qnaug/embeddinggemma2-decision", "decision_model.py"))
dm_mod = importlib.util.module_from_spec(spec); spec.loader.exec_module(dm_mod)
dm = dm_mod.DecisionModel("qnaug/embeddinggemma2-decision")
dm.decide("Hi, I was charged twice for invoice #4411. Please refund the duplicate today.", {
    "team": {"type": "choice", "instructions": "Which team should handle this?",
             "criteria": {"billing": "invoices, payments, refunds", "technical": "bugs, outages", "sales": "pricing"}},
    "refund": {"type": "noul", "instructions": "The customer asks for a refund."},
})

domain="extra" (default) uses the temperatures calibrated on natural text; domain="typed" uses those calibrated on JSON state like typed-decisions. Run in float32 or bfloat16, never float16 (the base model returns NaN).

Results

typed-decisions test accuracy 74.3% (random 31.8%, majority label 47.9%).

typed-decisions (test)

Accuracy 74.3%, ECE 0.029 after calibration.

Confidence thresholdDecisions handledAccuracy on handled
0.594.7%76.0%
0.678.8%80.2%
0.762.4%83.9%
0.845.0%87.4%
0.926.6%91.3%

boolq, mnli, banking77, clinc150 (test)

Accuracy 80.4%, ECE 0.039 after calibration.

Confidence thresholdDecisions handledAccuracy on handled
0.590.6%83.7%
0.674.4%87.2%
0.759.5%92.6%
0.848.2%96.9%
0.940.1%98.4%

Limitations

  • —Trained on English data only. Other languages (e.g. Vietnamese) are untested and the calibration does not carry over.
  • —Calibration is only valid for data like the training sources. Recalibrate on a few hundred examples of your own data before relying on the probabilities.
  • —Yes/no reading comprehension on free text is the weakest part (boolq / mnli ≈ 70%).
  • —Labels in typed-decisions have low inter-annotator agreement, so accuracy has a ceiling well below 100%.

Training

Text encoder of google/embeddinggemma-2 with the vocabulary embedding frozen, trained with soft-label cross-entropy (typed-decisions) and one-hot labels (other datasets). Intent datasets show the gold label plus random distractors, so the model learns to choose among arbitrary option lists. Per-source temperature scaling on a held-out split.

Dataset licenses differ (typed-decisions Apache-2.0, BoolQ CC BY-SA 3.0, banking77 CC BY 4.0, CLINC150 CC BY 3.0, MNLI mixed); check them for your use.