DavidHatley/system-one-mini-data
System One Mini Synthetic Diagnostics Deterministic synthetic controlled-intervention summaries used by DavidHatley/system-one-mini. The records contain no real coding-agent traces, private repositories, personal data, external model outputs, or teacher labels. This dataset and model are independent research artifacts, not reproductions of Jev or RLCD. Splits Split Rows train 20,000 validation 2,000 calibration 2,000 development_renderer 2,000… See the full description on the dataset page: https://huggingface.co/datasets/DavidHatley/system-one-mini-data.
System One Mini Synthetic Diagnostics
Deterministic synthetic controlled-intervention summaries used by DavidHatley/system-one-mini. The records contain no real coding-agent traces, private repositories, personal data, external model outputs, or teacher labels.
This dataset and model are independent research artifacts, not reproductions of Jev or RLCD.
Splits
development_renderer is the former sealed renderer split that was subsequently used to compare candidates. It must not be described as an unbiased final test. test is a separate renderer written and frozen after model selection, then evaluated once for the initial release.
Fields
id: SHA-256 fingerprint of the latent synthetic situationstate: rendered diagnostic summarylabels: five numeric labels in schema order- Five named string label fields
latent: complete generated situation used to compute labelsprovenance: generator, split, renderer, and seed
Generation
Generation code and manifests are included in generator/ and provenance/. The policy makes labels mechanically identifiable: every component trial starts independently from the same original failing snapshot, and exactly one clearing intervention identifies the operational failure source.
Limitations
- All scenarios are synthetic and templated.
- The label taxonomy covers only four component classes plus
unknown. - Recurring logical patterns across splits do not represent natural incident prevalence.
- Public test labels make the split unsuitable as a permanently sealed benchmark for future model selection.
