Team Ai
Datasetpublic

tasksource/procedural-typed-decisions

procedural-typed-decisions Procedurally generated decision problems. Each row is one structured state (JSON, or a table, CSV, key=value lines, or prose for the arithmetic, retrieval, and aggregation configs) with several typed questions over that same state, following the Jev / System One request shape: choice (pick one criterion), noul (a number in [0, 1]; a probability or a yes/no), and score (an ordered rubric). Every answer is computed exactly from the state by rules that… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/procedural-typed-decisions.

sourceHugging Faceapache-2.0updated 8d agoView on Hugging Face
4likes1.5kdownloads
Dataset Card

procedural-typed-decisions

Procedurally generated decision problems. Each row is one structured state (JSON, or a table, CSV, key=value lines, or prose for the arithmetic, retrieval, and aggregation configs) with several typed questions over that same state, following the Jev / System One request shape: choice (pick one criterion), noul (a number in [0, 1]; a probability or a yes/no), and score (an ordered rubric). Every answer is computed exactly from the state by rules that the state itself spells out, so the labels are noise-free. Several configs vary the number of options (4 to 60), to balance the binary and 4–6-option questions that dominate the rest of Jev.

This is an independent dataset. It is not an official TypeSafe Jev dataset and is not produced by or affiliated with TypeSafe or OpenJev.

Configs

configquestions
all (default)Every config below in one table, with a task column and the shared fields only (no flat label columns); the first 1,000 train rows cycle through levels and tasks, the rest is shuffled
arithmeticAn order with a discount/shipping rule, an account ledger, or a schedule; each state asks 2–5 of: amount_due / final_balance / finish_time (choice among the result and typical slips), within_budget, went_negative, done_by_deadline (noul), random_line_bulk, random_is_deposit, random_is_long (noul, exact probability k/n), budget_use, net_change (score, descriptive levels), lines_above, withdrawal_count, starts_before_noon (score), largest_line, lowest_day, longest_task (choice)
entity_belief_trackingworld_location (choice), agent_belief_location (choice), belief_matches_world (noul), from level 2 nested_belief_location (choice: where A thinks B believes an object is); 4 to 16 locations
event_state_reconstructioncurrent_owner (choice), is_open (noul), current_severity (score); the log is shuffled from level 2 and has voided entries from level 3
evidence_sufficiencyclaim_supported (noul), has_conflict (noul), strongest_support_origin (choice); retractions from level 2, mirrored (non-independent) origins from level 3, validity by collection day at level 4
multi_view_adjudicationintent (choice), is_urgent (noul), workflow_impact (score); near-threshold signals from level 2, auth failures counted from login events from level 3, deadlines as clock times at level 4
needle_retrievalvalue_of_id (choice), id_has_value (noul), id_listed (noul); up to ~300 records whose ids differ from the target by one or two digits, and from level 2 a chain of one to three id reissues to follow; 6 to 20 options
partial_observation_calibrationincident_real (noul, exact Bayesian posterior); 1 to 6 sensors
policy_applicabilityaccess_allowed (noul), governing_policy (choice), review_risk (score); 2 one-constraint policies at level 0, about 12 policies of up to 5 constraints, many of them near misses, at level 4
policy_under_uncertaintyaccess_allowed (noul), governing_policy (choice), requester_role (choice); exact posteriors over a role known through history counts and reports of stated reliability
record_aggregationcount_in_category (score), largest_quantity (choice), any_out_of_stock (noul), total_above (noul), from level 2 count_filtered (score, quantity and stock filters)
state_perturbationmaterial_change (noul), changed_dimension (choice), risk_direction (score); 1 to 8 records with up to 4 simultaneous changes whose risk effects can offset, and look-alike non-material fields
table_lookupfind_person (choice, two-condition filter, through the manager from level 2 and with a start-year condition from level 3; 6 to 40 options, capped by the table), manager_of (choice, join), started_before (noul), count_matching (score)
taxonomy_routingroute (choice among the 4–60 categories of a routing guide drawn fresh per state; many rules share a condition with the right one), belongs_to (noul), conditions_met (score, 0–4)

Schema

fieldmeaning
idtask:split:index
levelDifficulty level (0–4), calibrated against Jev (see below).
stateThe state: a JSON string, or rendered text for the retrieval and aggregation configs.
questionsJSON object of named System One questions (type, instructions, criteria).
answersJSON object of reference answers, in the System One answers shape.
one column per questionFlat label, for browsing and filtering: a ClassLabel for choice, score, and yes/no noul questions; a float for graded noul (incident_real, random_*); the option text for open numeric choices (amount_due, final_balance, finish_time). Null when the state does not ask that question (arithmetic, and level-dependent questions).

States are unique within a split, and validation/test states never occur in train. In each config, the first 1,000 train rows cycle through the levels (easiest first) for browsing; the rest of the split is shuffled.

Difficulty by level

Level 0 is meant to be easy for a strong decision model and level 4 hard. The table gives Jev's chance-adjusted accuracy, kappa = (accuracy − chance) / (1 − chance), on 40 fresh states per level (every question of each state; typesafe/jev-1.13-20260917, September 2026). 1 is perfect, 0 is chance.

configlevel 01234
arithmetic0.830.640.590.540.59
entity_belief_tracking0.910.930.780.740.66
event_state_reconstruction0.970.990.920.730.60
evidence_sufficiency0.990.970.780.720.78
multi_view_adjudication0.820.870.820.730.74
needle_retrieval1.000.990.780.830.36
partial_observation_calibration0.730.370.600.180.23
policy_applicability0.730.600.530.640.35
policy_under_uncertainty0.350.420.380.570.39
record_aggregation0.990.960.840.750.70
state_perturbation0.950.890.440.630.50
table_lookup1.000.980.970.900.87
taxonomy_routing1.000.980.920.780.64

Probability answers are scored above by their rounding to yes/no; Jev's mean absolute error on the exact probability grows from 0.16 (level 0) to 0.29 (level 4) on incident_real, and stays around 0.33 on the posterior access_allowed of policy_under_uncertainty. policy_under_uncertainty and partial_observation_calibration (exact posteriors) are hard from level 0 on; table_lookup and multi_view_adjudication remain the easiest at level 4.

A second model, upstage/solar-decide (10 states per level; it takes at most 26 options, so the longest lists are left out), shows the same easy-to-hard slope on most configs, e.g. 0.95 → 0.53 on event_state_reconstruction, 1.00 → 0.48 on evidence_sufficiency, 0.88 → 0.42 on policy_applicability; arithmetic, table_lookup, and state_perturbation stay easy for it (about 0.8–0.9 at every level). Rerun with scripts/calibrate_procedural_levels.py (--model for another model of the OpenRouter decisions API).

Use

As a multi-question Jev request, send {"state": row["state"], "questions": json.loads(row["questions"])} (parsing the state first when it is JSON) and compare with row["answers"]. The same rows are included, grouped by state, in `tasksource/tasksource-jev-typed-decisions`.

Reproduction

Generation is deterministic (row i of a split is seeded by task:split:i). From a tasksource checkout:

bash
PYTHONPATH=.:src python scripts/build_procedural_jev.py --output build/procedural-typed-decisions --upload

Generators live in src/tasksource/jev/procedural/.

Citation

Generated with tasksource; please cite:

bibtex
@inproceedings{sileo-2024-tasksource,
    title = "tasksource: A Large Collection of {NLP} tasks with a Structured Dataset Preprocessing Framework",
    author = "Sileo, Damien",
    booktitle = "Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)",
    month = may,
    year = "2024",
    address = "Torino, Italia",
    publisher = "ELRA and ICCL",
    url = "https://aclanthology.org/2024.lrec-main.1361",
    pages = "15655--15684",
}