Team Ai
Datasetpublic

typesafe/evalsafe-onet

EvalSafe O*NET 150 documents · 7,500 consensus-labeled questions · 9 candidate models. Snapshot: 2026-09-29. Default reference: consensus. Only questions with an available consensus target and their corresponding documents and final model results are included. The documents are synthetic workplace examples. The reference targets are model-generated, using Astra (gpt-6-astra) and Fable (claude-fable-5-1). The default reference is their consensus. Load from datasets… See the full description on the dataset page: https://huggingface.co/datasets/typesafe/evalsafe-onet.

sourceHugging Faceapache-2.0updated 8d agoView on Hugging Face
7likes796downloads
Dataset Card

EvalSafe O*NET

150 documents · 7,500 consensus-labeled questions · 9 candidate models. Snapshot: 2026-09-29. Default reference: consensus.

Only questions with an available consensus target and their corresponding documents and final model results are included. The documents are synthetic workplace examples. The reference targets are model-generated, using Astra (gpt-6-astra) and Fable (claude-fable-5-1). The default reference is their consensus.

Load

python
from datasets import load_dataset

repo = "typesafe/evalsafe-onet"
cases = load_dataset(repo, "cases", split="test")
questions = load_dataset(repo, "questions", split="test")
results = load_dataset(repo, "run_results", split="test")

Use revision="<commit SHA>" to pin the data. If access requires authentication, sign in to Hugging Face first. All three configurations have one test split.

ConfigurationRowsContents
cases150Complete documents, descriptive metadata, and question IDs
questions7,500Question/input and independent OpenAI, Anthropic, and consensus labels
run_results67,500Final prediction, embedded consensus target, score, cost and usage for each model/question

Each document has 27 noul, 16 score, and 7 choice questions, totaling 4,050 noul, 2,400 score, and 1,050 choice questions. Coverage, scoring definitions, and model summaries are recorded in `dataset.json`.

Join tables with case_id and question_instance_id; run_id joins model results to the summaries in dataset.json. Fields ending in _json contain JSON-encoded text. Every model-result row includes its reference target.

All 7,500 questions have a consensus target. Both reference sources are available for 7,468 questions; 32 have one available source.

Metrics

  • —Noul/choice intelligence: 1 - Jensen-Shannon divergence, using base-2 logarithms. This uses divergence, not its square root.
  • —Score agreement: 1 - abs(model_mean - reference_mean) / rubric_range.
  • —Overall performance: equal-weight mean of the noul, score, and choice type means. This is a composite of two metrics, not a pure JSD score.

Failed candidates score zero. The final table has 66,528/67,500 valid candidate results. Reference models are excluded from candidate rankings. Jev is reported as jev-1.13.0; the DeepSeek models were served by Together.

Results

[image]

Dashed frontiers connect nondominated models and do not predict intermediate configurations. Cost and latency axes are logarithmic; the chart uses all nine candidates.

ModelPerformanceValid resultsUSD/1,000 questionssec/question
claude-opus-50.943307,439/7,500$19.445603.184
gpt-5.6-sol0.934507,438/7,500$5.678732.720
typesafe/jev-1.13.00.933867,500/7,500$0.074630.141
gpt-5.6-terra0.917597,445/7,500$2.729392.300
claude-sonnet-50.907987,436/7,500$7.592402.125
claude-haiku-4-50.903507,468/7,500$7.9179912.008
gpt-5.6-luna0.887257,423/7,500$0.326742.281
together/deepseek-ai/DeepSeek-V4-Pro-08130.862677,245/7,500$5.601109.227
together/deepseek-ai/DeepSeek-V4-Flash-07310.835747,134/7,500$0.476917.094

These are generated workplace documents and questions, not the original O*NET database tables.

Targets are model-generated rather than human-verified ground truth. This benchmark is not a held-out evaluation of generalization.

License

This dataset is licensed under the Apache License 2.0.