datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
typed-decisions
Typed Decisions
A benchmark for typed probabilistic decisions. A model gets one piece of
unstructured state and answers five typed questions about it at once, and every
answer is a probability distribution, not a single label.
The schema follows the System One primitives (noul, choice, score) used by
TypeSafe AI, so a row replays against any API with that
shape. The benchmark is independent: it is not affiliated with TypeSafe and does
not reproduce their Jev model.… See the full description on the dataset page: https://huggingface.co/datasets/LocalLLaMA/typed-decisions.total-300-lambda00-s_signal_type6-jh-epoch4
total-300-lambda00-s_signal_type6-jh-epoch4
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.3875
Action score: 0.43125
Valid samples: 320/320
total-300-lambda10-s_signal_type6-jh-epoch4
total-300-lambda10-s_signal_type6-jh-epoch4
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.36640625
Action score: 0.41875
Valid samples: 320/320
total-300noapp-lambda02-s_signal_type6-jh-epoch4
total-300noapp-lambda02-s_signal_type6-jh-epoch4
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.36640625
Action score: 0.409375
Valid samples: 320/320
total-300-lambda02-s_signal_type6-jh-retry-epoch4
total-300-lambda02-s_signal_type6-jh-retry-epoch4
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.36953125
Action score: 0.3984375
Valid samples: 320/320
total-300-lambda02-s_signal_type6-jh-epoch4-reeval1
total-300-lambda02-s_signal_type6-jh-epoch4-reeval1
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.38828125
Action score: 0.4234375
Valid samples: 320/320
total-300-lambda02-s_signal_type6-jh-epoch4-reeval2
total-300-lambda02-s_signal_type6-jh-epoch4-reeval2
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.4125
Action score: 0.4265625
Valid samples: 320/320
total-300-lambda05-s_signal_type6-jh-epoch4
total-300-lambda05-s_signal_type6-jh-epoch4
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.35703125
Action score: 0.4375
Valid samples: 320/320
total-300-lambda08-s_signal_type6-jh-epoch4
total-300-lambda08-s_signal_type6-jh-epoch4
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.38046875
Action score: 0.4078125
Valid samples: 320/320
total-300-lambda02-s_signal_type6-jh-epoch4
total-300-lambda02-s_signal_type6-jh-epoch4
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.4046875
Action score: 0.4140625
Valid samples: 320/320
total-300app-lambda02-s_signal_type6-jh-epoch4
total-300app-lambda02-s_signal_type6-jh-epoch4
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.3625
Action score: 0.4015625
Valid samples: 320/320
total-131-lambda02-residual-s_signal_type6-jh-epoch4
total-131-lambda02-residual-s_signal_type6-jh-epoch4
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.3765625
Action score: 0.4171875
Valid samples: 320/320
procedural-typed-decisions
procedural-typed-decisions
Procedurally generated decision problems. Each row is one structured state
(JSON, or a table, CSV, key=value lines, or prose for the arithmetic,
retrieval, and aggregation configs) with several typed questions over that same state, following the
Jev / System One request shape: choice (pick one criterion), noul (a
number in [0, 1]; a probability or a yes/no), and score (an ordered rubric).
Every answer is computed exactly from the state by rules that… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/procedural-typed-decisions.evalsafe-invoice-processing
Invoice processing
Snapshot: 2026-09-28. 150 cases and 6,874 question instances.
Default reference: consensus. Labels are model-generated references.
Data
Load configuration cases, questions, or run_results; all have a test split.
cases: one row per case_id, with the complete input in input_json, descriptive
metadata_json, and openai, anthropic, and consensus labelsets. Decisions are grouped
by policy_id and contain status, actions, and primary_action.
questions:… See the full description on the dataset page: https://huggingface.co/datasets/typesafe/evalsafe-invoice-processing.evalsafe-customer-service
Customer service
Snapshot: 2026-09-28. 204 cases and 3,287 question instances.
Default reference: consensus. Labels are model-generated references.
Data
Load configuration cases, questions, or run_results; all have a test split.
cases: one row per case_id, with the complete input in input_json, descriptive
metadata_json, and openai, anthropic, and consensus labelsets. Decisions are grouped
by policy_id and contain status, actions, and primary_action.
questions:… See the full description on the dataset page: https://huggingface.co/datasets/typesafe/evalsafe-customer-service.evalsafe-security-incidents
Security incidents
Snapshot: 2026-09-28. 240 cases and 1,820 question instances.
Default reference: consensus. Labels are model-generated references.
Data
Load configuration cases, questions, or run_results; all have a test split.
cases: one row per case_id, with the complete input in input_json, descriptive
metadata_json, and openai, anthropic, and consensus labelsets. Decisions are grouped
by policy_id and contain status, actions, and primary_action.
questions:… See the full description on the dataset page: https://huggingface.co/datasets/typesafe/evalsafe-security-incidents.evalsafe-agent-trace-observability
Agent trace triage
Snapshot: 2026-09-28. 111 cases and 1,124 question instances.
Default reference: consensus. Labels are model-generated references.
Data
Load configuration cases, questions, or run_results; all have a test split.
cases: one row per case_id, with the complete input in input_json, descriptive
metadata_json, and openai, anthropic, and consensus labelsets. Decisions are grouped
by policy_id and contain status, actions, and primary_action.
questions:… See the full description on the dataset page: https://huggingface.co/datasets/typesafe/evalsafe-agent-trace-observability.evalsafe-onet
EvalSafe O*NET
150 documents · 7,500 consensus-labeled questions · 9 candidate models.
Snapshot: 2026-09-29. Default reference: consensus.
Only questions with an available consensus target and their corresponding documents
and final model results are included. The documents are synthetic workplace examples.
The reference targets are model-generated, using Astra (gpt-6-astra) and Fable
(claude-fable-5-1). The default reference is their consensus.
Load
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/typesafe/evalsafe-onet.typed-decision-bench
Typed Decision Bench v0.3
Built by Blobfish AI. A benchmark for one-pass decision models: 5,387 items, 25 tasks, 5 use-case
suites. Blobfish designed the tasks, wrote the typed questions, framed each one as a decision a business actually
delegates (use case, vertical), drew stratified seeded panels, froze them, and built the scoring, the contamination
tiers and the quality scorecard. The underlying records are drawn from 21 openly licensed public datasets plus one
generator of… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/typed-decision-bench.warn-act-notice-type-codes-crosswalk
WARN Act notice-type codes — the crosswalk
Every US state publishes WARN Act layoff notices with a free-text column saying
what kind of event it is. The statute recognises two: a plant closing and a
mass layoff. Across 48 states that column contains
552 distinct exact strings (531
once you fold case).
This dataset is the crosswalk: one row per raw string, how many notices carry
it, which states emit it, and what it normalizes to.
The finding that matters
520 of… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/warn-act-notice-type-codes-crosswalk.question-type-and-complexity
Question Type and Complexity (QTC) Dataset
Dataset Overview
The Question Type and Complexity (QTC) dataset is a comprehensive resource for linguistics/NLP research focusing on question classification and linguistic complexity analysis across multiple languages. It contains questions from two distinct sources (TyDi QA and Universal Dependencies v2.15), automatically annotated with question types (polar/content) and a set of linguistic complexity features.
Key Features:
2… See the full description on the dataset page: https://huggingface.co/datasets/rokokot/question-type-and-complexity.cav3_t-type_calcium_channels_butkiewicz
Dataset Details
Dataset Description
This dataset was initially curated from HTS data at the PubChem database.
The curation process is documented in Butkiewicz et al.
Primary screening with AID 449739 identified inhibitors of Cav3 T-type calcium channels.
Four follow-up screens were performed to confirm inhibitory effects on smaller sets of compounds
involving AID 493021, AID 493022, AID 493023, and AID 493041.
AID 489005 was performed as counter screen validating active… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/cav3_t-type_calcium_channels_butkiewicz.typescript-codefour_types_weightedHist-Pile-NER-Type
Hist-Pile-NER-Type
A multilingual historical-newspaper silver annotation dataset. This dataset combines the passage-to-conversation construction approach of UniversalNER / Pile-NER-type with an operational NERC adaptation of the Impresso / HIPE-2020 annotation guidelines v2.2.0.
This is not human-reviewed gold data, an official HIPE release, or an exact reproduction of the original open-type Pile-NER-type dataset. No HistGliNER model is included.
Contents… See the full description on the dataset page: https://huggingface.co/datasets/emanuelaboros/Hist-Pile-NER-Type.africa-mauritius-number-of-driving-licence-issued-by-year-and-by-type-of-li-40eebb1e
Number of Driving Licence Issued by Year and by Type of Li | Africa (MDPA)
72 rows - 1 Africa country/area - 2008-2019 - 1 indicator - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 72 rows from MDPA, covering Number of Driving Licence Issued by Year and by Type of Li. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.
What This Dataset Measures… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-mauritius-number-of-driving-licence-issued-by-year-and-by-type-of-li-40eebb1e.calibrated-typed-decisions
calibrated-typed-decisions (v1)
Typed decisions (multiple choice and yes/no) over short English texts, each with a human
gold label and a calibrated soft label from a teacher LLM. Built for training and
evaluating small decision models whose confidence scores mean something.
Why
A model that says 0.9 should be right about 90% of the time. Instruction-tuned LLMs read
as classifiers tend to be overconfident. This dataset pairs every decision with a teacher… See the full description on the dataset page: https://huggingface.co/datasets/m-lagnajit/calibrated-typed-decisions.fenic-codebaseafrica-mauritius-number-of-driving-licence-issued-by-year-and-by-type-of-li-b4a8c34f
Number of Driving Licence Issued by Year and by Type of Li | Africa (MDPA)
45 rows - 1 Africa country/area - 2010-2014 - 9 indicators - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 45 rows from MDPA, covering Number of Driving Licence Issued by Year and by Type of Li. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.
What This Dataset Measures… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-mauritius-number-of-driving-licence-issued-by-year-and-by-type-of-li-b4a8c34f.stack_edu_typescript
