mayafree/typed-decision-leaderboard
Typed Decision Leaderboard — JEV & open-Jev answer verifiers, one identical test set
An independent, side-by-side benchmark of answer verifiers / typed-decision models / System-1 scorers — the models that read an LLM's answer and decide, with zero generated tokens, whether it can be trusted. Every system is scored on the same 2,018 items with the same labels; the scores, labels and grading code are published, and each competitor is run on its own maker's code.
If you searched for a JEV alternative, an open-Jev leaderboard, Bespoke Nimble vs JEV, CLM-8B benchmark, decider-2b AUC, a hallucination-detection / answer-verification leaderboard, or calibrated confidence (ECE) for LLM answers — this is that table.
Current ranking — weighted AUC (higher is better)
The top three sit inside the confidence interval (paired bootstrap; ZTC−JEV 95% CI [−0.034, +0.020]), so they are marked a tie. On a held-out set of answers from a never-seen model (Claude Haiku 4.5, 1,939 items), ZTC-Judge-27B v2 = 0.7752 beats JEV = 0.7521 (Δ +0.023, CI [+0.006, +0.041], excludes zero).
What is measured (the axis)
- Question: "Is the ANSWER factually correct for the QUESTION?" → a probability of true.
- Zero generation: verifiers emit a probability without writing a new answer. Token-spending LLM judges (GPT-5.2, Gemini 2.5 Flash-Lite, GPT-4o-mini, Qwen3-Next-80B) are kept in an off-axis reference, not the ranking.
- A content-blind baseline (answer length & formatting only, AUC 0.6223) sits on the same board. A verifier below it did not read the content.
- Divisions: open-weight (local, self-hostable) systems are never mixed with commercial APIs.
- Fairness: every competitor is run on its own published code / head, never marked down by a re-implementation. A label-shuffle negative control reads 0.4996 (chance), validating the harness.
- Calibration (ECE): does "0.8" mean 80% correct? Reported alongside AUC.
- The 2,018 item texts stay private per source licenses (KMMLU CC BY-ND, CLIcK, GPQA); scores, labels and grading code are fully open.
The field — the "open-Jev" ecosystem
TypeSafe's JEV ("Decisions, Not Strings") started a category that is now a whole ecosystem: a CMU paper (JEV-as-a-Judge), dozens of open reproductions (open-jev, Bespoke Nimble, CLM-8B, decider, kev, von, ZefanCai Open-Jev, APUS-OpenJev, Laya, Manchego, Eikos, JevK5, Winnow, and 20+ more), competing leaderboards, and live Spaces. This board tracks 40+ distinct families and ranks the ones that are actually measurable on a neutral, labeled test set. Systems whose axis differs (game-playing bots, routers, constrained decoders, entity-extraction encoders) are listed off-ranking with the reason.
ZTC (Zero-Token Confidence) by VIDRAFT / FINAL-Bench is an open-weight verifier that reads a model's own hidden state to judge its answer — no generated tokens, no separate API.
한국어 요약
답변 검증기(타입드 디시전·System-1 스코어러) 리더보드. LLM의 답이 맞는지 토큰 생성 없이 판정하는 모델들을 같은 2,018문항·같은 라벨로 재고, 점수·라벨·채점 코드를 공개합니다. 각 경쟁 모델은 제작자 자신의 코드로 측정합니다.
- 부문 분리: 오픈 웨이트(로컬) vs 상용 API
- 기준선 동봉: 답 길이·서식만 보는 기준선(0.6223)을 못 넘으면 내용을 못 읽는 것
- 현재 1위 그룹(무승부): JEV 0.7350 · ZTC-27B 0.7289 · ZTC-397B 0.7272
- 처음 보는 답(Haiku)에선 ZTC가 JEV를 이김 (0.7752 vs 0.7521)
- JEV·open-jev·Bespoke Nimble·CLM-8B·decider·Laya 등 40여 계열 추적
Keywords: JEV alternative, open-jev leaderboard, answer verification benchmark, typed decisions, hallucination detection, LLM judge, calibrated confidence, ECE, AUC, zero-token verifier, System-1 decision model, Bespoke Nimble, CLM-8B, decider-2b, Laya, ZTC, factuality checker, 답변 검증기, 환각 탐지, 리더보드.
