jev
Datasets
All datasets matching “jev”jev-ai-api-guide-assetsjev-bench
jev-bench
Real human-labeled data, reformatted into System One questions — with human label distributions wherever they exist.
22 configs · 166,054 rows · 22,773 test records · 4 calibration-gold configs · 46 models scored · v0.1.1
Code & engine · Findings · Leaderboard · Published models · Source rationale · Jev's API, verified · Other Jev evaluations
Every model on the same 22,773 test records. Down and to the right is better; the stars are open models fine-tuned on their… See the full description on the dataset page: https://huggingface.co/datasets/Praveenrajus/jev-bench.tasksource-jev-typed-decisions
tasksource-jev-typed-decisions
2.5 million typed decisions (choices, ratings and probabilities) from 670 sources.
Why use it
Real supervision. Labels, ratings, and annotator votes come from
established datasets, not a teacher model. Every row names its source.
Breadth. Over 300 dataset families: NLI and reasoning, QA and
commonsense, sentiment, intent and topic, toxicity and safety, preference
pairs, fact checking, entity tagging, and dozens of languages. GLUE… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/tasksource-jev-typed-decisions.Open-Jev
Open-Jev: typed decision datasets
Open-Jev turns a state and a question into a typed decision: a yes/no probability, a distribution over choices, independent label probabilities, or a discrete numeric/ordinal decision. This repository publishes twelve separate, frozen data configs from the Open-Jev project, together with original manifests, exact raw records, source code and reconstruction instructions.
These are controlled, mostly synthetic tasks and reference labels. They are… See the full description on the dataset page: https://huggingface.co/datasets/ZefanCai/Open-Jev.jev-api-article-assetsjev-decision-index-results
JEV models on the Decision Index
Complete Decision Index 0.2.1 runs of the JEV typed-decision models by AutoTrust AI, for the
Jev Decision Index board
(kit). Only engine outputs are published here (compact result rows: run
ids, statuses, answers and probabilities, timings). No benchmark inputs are included; rebuild the suite with the kit.
run
model
Decision Index 0.2.1
raw
complete
runs/jev-9b
autotrust/JEV-9B @ 4ab5dfb
43.14
56.91
yes (150,317 of 150,317 scoreable… See the full description on the dataset page: https://huggingface.co/datasets/autotrust/jev-decision-index-results.
