tasksource/procedural-typed-decisions
procedural-typed-decisions Procedurally generated decision problems. Each row is one structured state (JSON, or a table, CSV, key=value lines, or prose for the arithmetic, retrieval, and aggregation configs) with several typed questions over that same state, following the Jev / System One request shape: choice (pick one criterion), noul (a number in [0, 1]; a probability or a yes/no), and score (an ordered rubric). Every answer is computed exactly from the state by rules that… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/procedural-typed-decisions.
procedural-typed-decisions
Procedurally generated decision problems. Each row is one structured state (JSON, or a table, CSV, key=value lines, or prose for the arithmetic, retrieval, and aggregation configs) with several typed questions over that same state, following the Jev / System One request shape: choice (pick one criterion), noul (a number in [0, 1]; a probability or a yes/no), and score (an ordered rubric). Every answer is computed exactly from the state by rules that the state itself spells out, so the labels are noise-free. Several configs vary the number of options (4 to 60), to balance the binary and 4–6-option questions that dominate the rest of Jev.
This is an independent dataset. It is not an official TypeSafe Jev dataset and is not produced by or affiliated with TypeSafe or OpenJev.
Configs
Schema
States are unique within a split, and validation/test states never occur in train. In each config, the first 1,000 train rows cycle through the levels (easiest first) for browsing; the rest of the split is shuffled.
Difficulty by level
Level 0 is meant to be easy for a strong decision model and level 4 hard. The table gives Jev's chance-adjusted accuracy, kappa = (accuracy − chance) / (1 − chance), on 40 fresh states per level (every question of each state; typesafe/jev-1.13-20260917, September 2026). 1 is perfect, 0 is chance.
Probability answers are scored above by their rounding to yes/no; Jev's mean absolute error on the exact probability grows from 0.16 (level 0) to 0.29 (level 4) on incident_real, and stays around 0.33 on the posterior access_allowed of policy_under_uncertainty. policy_under_uncertainty and partial_observation_calibration (exact posteriors) are hard from level 0 on; table_lookup and multi_view_adjudication remain the easiest at level 4.
A second model, upstage/solar-decide (10 states per level; it takes at most 26 options, so the longest lists are left out), shows the same easy-to-hard slope on most configs, e.g. 0.95 → 0.53 on event_state_reconstruction, 1.00 → 0.48 on evidence_sufficiency, 0.88 → 0.42 on policy_applicability; arithmetic, table_lookup, and state_perturbation stay easy for it (about 0.8–0.9 at every level). Rerun with scripts/calibrate_procedural_levels.py (--model for another model of the OpenRouter decisions API).
Use
As a multi-question Jev request, send {"state": row["state"], "questions": json.loads(row["questions"])} (parsing the state first when it is JSON) and compare with row["answers"]. The same rows are included, grouped by state, in `tasksource/tasksource-jev-typed-decisions`.
Reproduction
Generation is deterministic (row i of a split is seeded by task:split:i). From a tasksource checkout:
PYTHONPATH=.:src python scripts/build_procedural_jev.py --output build/procedural-typed-decisions --uploadGenerators live in src/tasksource/jev/procedural/.
Citation
Generated with tasksource; please cite:
@inproceedings{sileo-2024-tasksource,
title = "tasksource: A Large Collection of {NLP} tasks with a Structured Dataset Preprocessing Framework",
author = "Sileo, Damien",
booktitle = "Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)",
month = may,
year = "2024",
address = "Torino, Italia",
publisher = "ELRA and ICCL",
url = "https://aclanthology.org/2024.lrec-main.1361",
pages = "15655--15684",
}