Team Ai
Datasetpublic

Leanmcp/jevbench

JevBench Cases and predictions for JevBench: An Open Evaluation Framework for Typed Decision Models. The evaluation code is on GitHub: https://github.com/Leanmcp/jevbench JevBench evaluates models that return typed decisions (a yes/no probability, a distribution over a set of choices, or an expected level on an ordered rubric) on identical inputs built from public datasets. Contents Path What it holds cases/ One JSONL file per slice. Each row is the… See the full description on the dataset page: https://huggingface.co/datasets/Leanmcp/jevbench.

sourceHugging Faceotherupdated 8d agoView on Hugging Face
0likes550downloads
Dataset Card

JevBench

Cases and predictions for JevBench: An Open Evaluation Framework for Typed Decision Models. The evaluation code is on GitHub: https://github.com/Leanmcp/jevbench JevBench evaluates models that return typed decisions (a yes/no probability, a distribution over a set of choices, or an expected level on an ordered rubric) on identical inputs built from public datasets.

Contents

PathWhat it holds
cases/One JSONL file per slice. Each row is the exact request sent to every model, plus the gold answer, source revision and licence.
images/Images for the ScienceQA and VQA-RAD image slices.
predictions/<system>/predictions.jsonlOne decision per case with the full probability vector, latency and token counts. No source text.
analysis/Paired bootstrap results and the djev image ablation.
provenance/djev/Serving provenance for djev-0.1 (vLLM command line, container logs, configuration).
sources.jsonEvery source dataset pinned to a Hugging Face revision, with the fields sent and withheld.

Systems: jev-1.13.0 (hosted commercial model), djev-0.1 (open, DiffusionGemma 26B-A4B backbone; djev-0.1-images holds its image slices), laya-421m-en (open, ModernBERT-large encoder). Total: 24,799 decisions.

Licences

Each case row carries the licence of its source dataset in its license field; licences do not merge across sources. SST-5 states no licence on its card, so its cases are released without their sentence text: state_hash identifies each sentence, which can be recovered from the pinned source revision in sources.json.

Code

The harness that downloads the sources, builds and verifies these cases, runs models and scores them is released under the MIT Licence at https://github.com/Leanmcp/jevbench. Each predictions/<system>/scores.json holds the scores reported in the paper.