jevbench
Datasets
All datasets matching “jevbench”jev-bench
jev-bench
Real human-labeled data, reformatted into System One questions — with human label distributions wherever they exist.
22 configs · 166,054 rows · 22,773 test records · 4 calibration-gold configs · 46 models scored · v0.1.1
Code & engine · Findings · Leaderboard · Published models · Source rationale · Jev's API, verified · Other Jev evaluations
Every model on the same 22,773 test records. Down and to the right is better; the stars are open models fine-tuned on their… See the full description on the dataset page: https://huggingface.co/datasets/Praveenrajus/jev-bench.jevbench
JevBench
Cases and predictions for JevBench: An Open Evaluation Framework for Typed
Decision Models. The evaluation code is on GitHub: https://github.com/Leanmcp/jevbench JevBench evaluates models that return typed decisions
(a yes/no probability, a distribution over a set of choices, or an expected
level on an ordered rubric) on identical inputs built from public datasets.
Contents
Path
What it holds
cases/
One JSONL file per slice. Each row is the… See the full description on the dataset page: https://huggingface.co/datasets/Leanmcp/jevbench.jobfit-jevbench
JobFit-JevBench
A CV/job-post fit-scoring benchmark: 250 synthetic CVs, 404 job postings and 5,000 CV/job pairs, each judged on 17 questions with a score and a confidence. It trains and evaluates jobfit-model, a Jev-like typed-decision model (a state plus typed questions in, one typed answer per question id out).
Layout
questions.json # the 17 questions: type, name, scope, question text, 5 criteria
cvs/*.md # synthetic CVs, plain markdown… See the full description on the dataset page: https://huggingface.co/datasets/gw0/jobfit-jevbench.jevbench
JevBench
Metamorphic coherence testing for typed probabilistic decision models.
Technical report ·
Code (pip install jevbench) ·
Leaderboard
Typed probabilistic decision models answer questions about a fixed input with probability distributions over
declared answers: yes or no (a Noul), one of several options (a Choice), or a level on an ordered scale (a
Score). JevBench asks whether those probabilities stay coherent when a question is reworded, negated or logically
combined… See the full description on the dataset page: https://huggingface.co/datasets/JevBench/jevbench.jev-bench-tr
Machine translation (en → tr) of Praveenrajus/jev-bench @ b41e6b2f68a429608e1c0d354324d0bebaec7971 by qwen3.8-flash-next (prompt v1-json-bfb402e8). Included split(s): test (22,773 rows). Labels, ids, soft labels and metadata are unchanged; see manifest.json.
jev-bench
Real human-labeled data, reformatted into System One questions — with human label distributions wherever they exist.
22 configs · 166,054 rows · 22,773 test records · 4 calibration-gold configs · v0.1.1
Repo &… See the full description on the dataset page: https://huggingface.co/datasets/hayriyigit/jev-bench-tr.jev-bench
jev-bench
A small multiple-choice set for measuring a model that returns the probability of each option
instead of writing an answer — the Jev / TypeSafe System One style of API, where a request carries
a state and a question with named options and the response carries a distribution over them.
Ordinary multiple-choice benchmarks score the text a model generates. That says nothing about
whether the probability attached to the answer means anything, which is the whole point of… See the full description on the dataset page: https://huggingface.co/datasets/kishida/jev-bench.
