typesafe/evalsafe-onet
EvalSafe O*NET 150 documents · 7,500 consensus-labeled questions · 9 candidate models. Snapshot: 2026-09-29. Default reference: consensus. Only questions with an available consensus target and their corresponding documents and final model results are included. The documents are synthetic workplace examples. The reference targets are model-generated, using Astra (gpt-6-astra) and Fable (claude-fable-5-1). The default reference is their consensus. Load from datasets… See the full description on the dataset page: https://huggingface.co/datasets/typesafe/evalsafe-onet.
EvalSafe O*NET
150 documents · 7,500 consensus-labeled questions · 9 candidate models. Snapshot: 2026-09-29. Default reference: consensus.
Only questions with an available consensus target and their corresponding documents and final model results are included. The documents are synthetic workplace examples. The reference targets are model-generated, using Astra (gpt-6-astra) and Fable (claude-fable-5-1). The default reference is their consensus.
Load
from datasets import load_dataset
repo = "typesafe/evalsafe-onet"
cases = load_dataset(repo, "cases", split="test")
questions = load_dataset(repo, "questions", split="test")
results = load_dataset(repo, "run_results", split="test")Use revision="<commit SHA>" to pin the data. If access requires authentication, sign in to Hugging Face first. All three configurations have one test split.
Each document has 27 noul, 16 score, and 7 choice questions, totaling 4,050 noul, 2,400 score, and 1,050 choice questions. Coverage, scoring definitions, and model summaries are recorded in `dataset.json`.
Join tables with case_id and question_instance_id; run_id joins model results to the summaries in dataset.json. Fields ending in _json contain JSON-encoded text. Every model-result row includes its reference target.
All 7,500 questions have a consensus target. Both reference sources are available for 7,468 questions; 32 have one available source.
Metrics
- Noul/choice intelligence:
1 - Jensen-Shannon divergence, using base-2 logarithms. This uses divergence, not its square root. - Score agreement:
1 - abs(model_mean - reference_mean) / rubric_range. - Overall performance: equal-weight mean of the noul, score, and choice type means. This is a composite of two metrics, not a pure JSD score.
Failed candidates score zero. The final table has 66,528/67,500 valid candidate results. Reference models are excluded from candidate rankings. Jev is reported as jev-1.13.0; the DeepSeek models were served by Together.
Results
Dashed frontiers connect nondominated models and do not predict intermediate configurations. Cost and latency axes are logarithmic; the chart uses all nine candidates.
These are generated workplace documents and questions, not the original O*NET database tables.
Targets are model-generated rather than human-verified ground truth. This benchmark is not a held-out evaluation of generalization.
License
This dataset is licensed under the Apache License 2.0.
