Infosec-Consult/sage-memory-judge
SAGE Memory Judge (v15)
A small (0.87B) local judge for SAGE's memory write-gate. It answers two yes/no questions about a proposed memory, read directly off one token's log-probabilities:
- Supported — is the memory's claim supported by the supplied evidence as stated (including its time frame, certainty, and the direction of any relation)?
- Lasting — should this be stored as long-term memory (a durable fact or standing rule), rather than a remark about the current session?
It runs entirely on the machine that hosts it (via Ollama) — no cloud inference, no external judge service. It is not a chat model; its only job is to return a calibrated Y/N probability for these checks.
Model
- Base: Qwen/Qwen3.5-0.8B (Apache-2.0), full fine-tune
- Format: GGUF, q8_0 —
sha256 714ff9324133ba3b7166fe82fa362226fae5574c4f9cfbc53c054248b16f2cfd - Runtime: Ollama (OpenAI-compatible endpoint), reads the first generated token's logprobs
- License: Apache-2.0
Use it
ollama create sage-memory-judge -f ModelfileModelfile:
FROM ./sage-judge-v15-q8_0.gguf
TEMPLATE {{ .Prompt }}
RENDERER qwen3.5
PARSER qwen3.5Hard serving requirement — thinking must be OFF. The judge reads the first generated token as the label, so the model must not think first. On Ollama's OpenAI-compatible /v1/chat/completions path, only `reasoning_effort: "none"` disables thinking — think:false and chat_template_kwargs:{enable_thinking:false} silently no-op there for this architecture. The native /api/chat path uses think:false instead. If thinking is left on, the first token is a reasoning token and a correct reader fails closed (holds for review) on every call — loud, not silent, but non-functional.
The probability is P(Y)/(P(Y)+P(N)) renormalised over the label mass at the first token. Typical gate bands: accept ≥ 0.90, reject < 0.50, otherwise hold for human review. A first token that is not a label must be treated as unavailable → hold, never a guessed score.
Evaluation
Measured once on the shippable GGUF (served by Ollama), on sets authored independently of the model and frozen by digest before measurement.
- CPU latency: ~297 ms (p50) per check (single token), ≈0.3–0.8 s per memory.
- Languages: trained and usable across 15 languages. A translation-based robustness probe shows traps controlled in all 15; genuine-acceptance is strong in most, with a heavier review burden in Hindi, Thai and Bengali (lower-resource scripts — the judge sends more genuine items to review there rather than accepting).
Limitations
- Converse-voice edge: a small number of genuine passive-voice restatements ("A is outranked by B" vs "B outranks A") are over-sent to review; one confusable shape ("leases to/from") is the residual.
- Reversed-relation is validated at the class level, not per relation shape.
- Lower-resource languages (Hindi, Thai, Bengali) are the least reliable; the per-language figures come from a translation-based probe with small per-language samples, so treat them as indicative. An independently-authored multilingual acceptance set is the way to make them precise before broad reliance.
- It checks support against caller-supplied evidence; it does not authenticate the source.
Provenance
Distilled from larger instruct models using soft (probability) labels on synthetic, de-duplicated training data; the training data and the teacher models are not part of this release. Evaluation sets were written by an independent author, verified independently (structural + non-reuse checks + hand-reading of the ambiguous slices), frozen by published digest, and measured one-shot. The reader shipped with SAGE (a Go port) is checked in CI to return probabilities identical to the reference reader that produced these numbers.
