kaivoss/system-one-270m-data
system-one-270m-data 25,002 synthetic typed decisions: a piece of state, a question, a caller-supplied option set, and a soft target distribution over those options. Built to train kaivoss/system-one-270m, an open take on the System One model class (TypeSafe Jev, Laya). Schema Field Type Meaning prompt string the full rendered prompt, state + question + lettered options letters list[string] the option letters in play, ["A", "B", ...] target… See the full description on the dataset page: https://huggingface.co/datasets/kaivoss/system-one-270m-data.
system-one-270m-data
25,002 synthetic typed decisions: a piece of state, a question, a caller-supplied option set, and a soft target distribution over those options.
Built to train `kaivoss/system-one-270m`, an open take on the System One model class (TypeSafe Jev, Laya).
Schema
target is the point of the dataset. One-hot labels train argmax accuracy and destroy calibration; these are distributions.
Splitting
Split on `state_id`, not per row. Each state carries 2–5 questions (3.32 average) that share the entire prompt body, so a row-level split leaks the state into both sides.
How it was made
Two stages per state with openai/gpt-oss-20b via OpenRouter, ~$0.000275 per question:
- Seed — sample domain, length, target ambiguity and cardinality; generate one realistic state plus 2–5 typed questions with prose criteria per option.
- Label — an ensemble of permuted reads. Each read shuffles the option order and asks for a single letter, so the answer is one token and
top_logprobsreturns a distribution. Reads stop early at 3 when unanimous.
The ensemble exists because gpt-oss reasons before answering and cannot be told not to — a post-reasoning distribution is effectively one-hot (measured: top option at -0.0 with no competitor in the top 8). The uncertainty lives in the variance across reads, and shuffling also cancels letter-position bias.
Gates reject: states under 200 chars, option counts below 60% of requested, malformed yes/no pairs, distributions that do not sum to 1, and the keyword giveaway (a state naming the winning option and none of its competitors). State-level acceptance was ~60%.
Composition
ambiguity_requested vs ambiguity is published deliberately: asking for an ambiguous state yields a confident answer about 70% of the time. That is a fact about the teacher, and hiding it would have meant paying to regenerate good rows.
Limitations
- Skewed to easy. 84.6% clear means graded uncertainty is under-represented. This is the dataset's main weakness; a targeted pass keeping only rows where the ensemble splits would improve it.
- Labels are gpt-oss-20b's judgement, not human. Its systematic biases are baked in, and the quality ceiling is its own ability.
- High cardinality is thin — only 2.6% at 9–14 options and almost nothing above 15, despite both Jev and Laya degrading in exactly that range.
- No human validation. No sample has been checked by a person for label correctness.
- English only, and the states are model-invented, not drawn from real logs or tickets.
Generated with openai/gpt-oss-20b (Apache-2.0). Code: https://github.com/Akicou/system-one-270m
