Team Ai
8 results

jevbench

Praveenrajus /jev-bench jev-bench Real human-labeled data, reformatted into System One questions — with human label distributions wherever they exist. 22 configs · 166,054 rows · 22,773 test records · 4 calibration-gold configs · 46 models scored · v0.1.1 Code & engine · Findings · Leaderboard · Published models · Source rationale · Jev's API, verified · Other Jev evaluations Every model on the same 22,773 test records. Down and to the right is better; the stars are open models fine-tuned on their… See the full description on the dataset page: https://huggingface.co/datasets/Praveenrajus/jev-bench.imagetext-classification100K<n<1M6 likes11k downloads9d agoHugging FaceLeanmcp /jevbench JevBench Cases and predictions for JevBench: An Open Evaluation Framework for Typed Decision Models. The evaluation code is on GitHub: https://github.com/Leanmcp/jevbench JevBench evaluates models that return typed decisions (a yes/no probability, a distribution over a set of choices, or an expected level on an ordered rubric) on identical inputs built from public datasets. Contents Path What it holds cases/ One JSONL file per slice. Each row is the… See the full description on the dataset page: https://huggingface.co/datasets/Leanmcp/jevbench.imagen<1K0 likes762 downloads12d agoHugging Facegw0 /jobfit-jevbench JobFit-JevBench A CV/job-post fit-scoring benchmark: 250 synthetic CVs, 404 job postings and 5,000 CV/job pairs, each judged on 17 questions with a score and a confidence. It trains and evaluates jobfit-model, a Jev-like typed-decision model (a state plus typed questions in, one typed answer per question id out). Layout questions.json # the 17 questions: type, name, scope, question text, 5 criteria cvs/*.md # synthetic CVs, plain markdown… See the full description on the dataset page: https://huggingface.co/datasets/gw0/jobfit-jevbench.tabulartext-classification10K<n<100K2 likes405 downloads2d agoHugging FaceJevBench /jevbench JevBench Metamorphic coherence testing for typed probabilistic decision models. Technical report · Code (pip install jevbench) · Leaderboard Typed probabilistic decision models answer questions about a fixed input with probability distributions over declared answers: yes or no (a Noul), one of several options (a Choice), or a level on an ordered scale (a Score). JevBench asks whether those probabilities stay coherent when a question is reworded, negated or logically combined… See the full description on the dataset page: https://huggingface.co/datasets/JevBench/jevbench.tabulartext-classification1K<n<10K1 likes255 downloads5d agoHugging Facehayriyigit /jev-bench-tr Machine translation (en → tr) of Praveenrajus/jev-bench @ b41e6b2f68a429608e1c0d354324d0bebaec7971 by qwen3.8-flash-next (prompt v1-json-bfb402e8). Included split(s): test (22,773 rows). Labels, ids, soft labels and metadata are unchanged; see manifest.json. jev-bench Real human-labeled data, reformatted into System One questions — with human label distributions wherever they exist. 22 configs · 166,054 rows · 22,773 test records · 4 calibration-gold configs · v0.1.1 Repo &… See the full description on the dataset page: https://huggingface.co/datasets/hayriyigit/jev-bench-tr.texttext-classification10K<n<100K3 likes248 downloads16d agoHugging Facekishida /jev-bench jev-bench A small multiple-choice set for measuring a model that returns the probability of each option instead of writing an answer — the Jev / TypeSafe System One style of API, where a request carries a state and a question with named options and the response carries a distribution over them. Ordinary multiple-choice benchmarks score the text a model generates. That says nothing about whether the probability attached to the answer means anything, which is the whole point of… See the full description on the dataset page: https://huggingface.co/datasets/kishida/jev-bench.tabularmultiple-choice1K<n<10K1 likes164 downloads17d agoHugging Face