VIDraft/zero-token-confidence
Zero-Token Confidence
Measuring an LLM's confidence costs zero tokens.
Self-consistency, verbalized confidence and judge models all buy confidence with extra tokens. This demo reads the model's own judgment signal inside the same forward pass that writes the answer — no extra generation, no extra forward pass. When the signal is low, the output tokens are blocked and the question is handed to another model.
Two brains, one question, one difference: the feedback loop.
- A — conventional LLM. Answers every time. When it is wrong, it is wrong with confidence.
- B — artificial brain. Reads its own judgment neurons and stops when it should not speak.
Ask the model something it gets wrong, then raise the threshold and watch the wrong answer get cut before it reaches you.
About this build
Inference runs on a separate private engine; this app is the interface to it. Implementation details, probe construction and model internals are not part of this repository.
Confidence is calibrated on English math word problems. Outside that domain the demo says so — the ordering may still be informative, the absolute scale is not.
© VIDRAFT. All rights reserved.
