Team Ai
Datasetpublic

kaivoss/system-one-270m-data

system-one-270m-data 25,002 synthetic typed decisions: a piece of state, a question, a caller-supplied option set, and a soft target distribution over those options. Built to train kaivoss/system-one-270m, an open take on the System One model class (TypeSafe Jev, Laya). Schema Field Type Meaning prompt string the full rendered prompt, state + question + lettered options letters list[string] the option letters in play, ["A", "B", ...] target… See the full description on the dataset page: https://huggingface.co/datasets/kaivoss/system-one-270m-data.

sourceHugging Faceapache-2.0updated 19d agoView on Hugging Face
0likes300downloads
Dataset Card

system-one-270m-data

25,002 synthetic typed decisions: a piece of state, a question, a caller-supplied option set, and a soft target distribution over those options.

Built to train `kaivoss/system-one-270m`, an open take on the System One model class (TypeSafe Jev, Laya).

Schema

FieldTypeMeaning
promptstringthe full rendered prompt, state + question + lettered options
letterslist[string]the option letters in play, ["A", "B", ...]
targetlist[float]soft target distribution over those options, sums to 1
labelstringargmax option's label
qtypestringchoice, noul (yes/no) or score (ordinal)
n_optionsintoption count, 2–20
domainstringone of 28 source domains
ambiguitystringmeasured band: clear / borderline / ambiguous
ambiguity_requestedstringwhat the generator was asked for
entropyfloatnormalised entropy of target
scorefloat\nullordinal expectation, score questions only
state_idstringgroups questions sharing one state — split on this

target is the point of the dataset. One-hot labels train argmax accuracy and destroy calibration; these are distributions.

Splitting

Split on `state_id`, not per row. Each state carries 2–5 questions (3.32 average) that share the entire prompt body, so a row-level split leaks the state into both sides.

How it was made

Two stages per state with openai/gpt-oss-20b via OpenRouter, ~$0.000275 per question:

  1. 1.Seed — sample domain, length, target ambiguity and cardinality; generate one realistic state plus 2–5 typed questions with prose criteria per option.
  2. 2.Label — an ensemble of permuted reads. Each read shuffles the option order and asks for a single letter, so the answer is one token and top_logprobs returns a distribution. Reads stop early at 3 when unanimous.

The ensemble exists because gpt-oss reasons before answering and cannot be told not to — a post-reasoning distribution is effectively one-hot (measured: top option at -0.0 with no competitor in the top 8). The uncertainty lives in the variance across reads, and shuffling also cancels letter-position bias.

Gates reject: states under 200 chars, option counts below 60% of requested, malformed yes/no pairs, distributions that do not sum to 1, and the keyword giveaway (a state naming the winning option and none of its competitors). State-level acceptance was ~60%.

Composition

Questions / states25,002 / 7,537
Type49.3% choice, 33.4% noul, 17.3% score
Measured band84.6% clear, 10.3% borderline, 5.1% ambiguous
Soft targets (confidence < 0.9)19.2%
Option count86.6% at 2–5, 10.8% at 6–8, 2.6% at 9–14
Domains28, roughly balanced (728–955 rows each)
Labels from logprobs~100% (6 rows fell back to hard voting)

ambiguity_requested vs ambiguity is published deliberately: asking for an ambiguous state yields a confident answer about 70% of the time. That is a fact about the teacher, and hiding it would have meant paying to regenerate good rows.

Limitations

  • —Skewed to easy. 84.6% clear means graded uncertainty is under-represented. This is the dataset's main weakness; a targeted pass keeping only rows where the ensemble splits would improve it.
  • —Labels are gpt-oss-20b's judgement, not human. Its systematic biases are baked in, and the quality ceiling is its own ability.
  • —High cardinality is thin — only 2.6% at 9–14 options and almost nothing above 15, despite both Jev and Laya degrading in exactly that range.
  • —No human validation. No sample has been checked by a person for label correctness.
  • —English only, and the states are model-invented, not drawn from real logs or tickets.

Generated with openai/gpt-oss-20b (Apache-2.0). Code: https://github.com/Akicou/system-one-270m