blazeofchi/system-one-eval
System One Eval Structured State Reasoning Probe, v0.1. This is a small, openly keyed diagnostic set, not a population benchmark or a blinded leaderboard. It contains 60 originally authored synthetic text tasks and 70 named questions covering policy precedence, cross-row aggregation, joins, boundaries, scheduling, access control, evidence limits, and multi-step arithmetic. There are no images or borrowed public-benchmark items in this package. What is in the data… See the full description on the dataset page: https://huggingface.co/datasets/blazeofchi/system-one-eval.
System One Eval
Structured State Reasoning Probe, v0.1.
This is a small, openly keyed diagnostic set, not a population benchmark or a blinded leaderboard. It contains 60 originally authored synthetic text tasks and 70 named questions covering policy precedence, cross-row aggregation, joins, boundaries, scheduling, access control, evidence limits, and multi-step arithmetic. There are no images or borrowed public-benchmark items in this package.
What is in the data
Each line of test.jsonl is one task. state_json, questions_json, and gold_json are JSON strings so the Hub viewer can display a consistent schema across varied task states. Parse those fields before evaluation. questions_json uses named choice questions (option names in criteria) and noul true/false questions. prompt_preview is only for browsing; the complete question objects are in questions_json.
The task is the scoring unit. A task passes only if every named answer matches its key. Choice names require exact matching; noul outputs are true at probability ≥0.5. The public scorer also accepts direct Boolean answers. The dataset includes the answer key and rationale openly, so it must not be treated as an unseen holdout after publication.
Construction and provenance
The first twelve tasks began in a personal DiffusionGemma evaluation. V2 added twenty-eight cases and repaired the wording of one sibling-dependent question (T03). V3 revised twelve of those cases for extra rule interactions and added twenty new cases. The V3 revisions were chosen after the authors saw V2 model outcomes. That adaptivity and several related task families reduce the set's value for general model ranking. All state values, prompts, choices, and keys were authored for this probe; they were not copied from ChartQA, MMMU, or another benchmark.
The keys received local arithmetic and rule checks. Independent external adjudication has not been completed. This dataset is English-only, synthetic, multiple choice or Boolean, and strongly weighted toward operational rules and structured JSON states. It does not measure open-ended generation, real user traffic, or broad factual knowledge.
Reproduce scoring
python validate.py validates the public file. Produce a JSONL file with one row per task, shaped like:
{"task_id":"T01","answers":{"priority":"P2","is_p1":false}}Then run python score.py predictions.jsonl. The scorer requires all 60 tasks and all named questions, and prints whole-task and question accuracy plus per-task checks. Feed models only state_json and questions_json; do not send gold_json or rationale.
The included prediction files contain sanitized named predictions from Jev, DiffusionGemma, direct CLM, GPT-6 Luna, Decider-4B v2, Decider-2B-Vision, and GLiNER2.5-Decide runs. They contain no API credentials, account identifiers, endpoint URLs, or raw provider traces. Run python score.py jev_predictions.jsonl (or any other model's prediction file) to recompute scores.
Reported runs and limits
The included model results, if present, are one attempt per model per task, with no retries. Jev and DiffusionGemma used different hosted services and input encodings required by their APIs. Request latency therefore includes different network and service effects and is not a controlled hardware comparison. Cost estimates use different billing bases and should not be interpreted as a price ratio. Scores on this selected, openly keyed challenge set do not establish a general model ranking.
All 62 choice questions have null option descriptions. CLM's released API embeds a bare option key when no description is supplied, so its direct result measures this exact input shape. The separately reported answer-text checks were chosen after the direct result and are exploratory.
License
The authored dataset and accompanying benchmark files are published under CC BY 4.0. Model names remain attributed to their respective providers; this package does not include model weights or third-party images.
