Team Ai
Datasetpublic

DavidHatley/system-one-mini-data

System One Mini Synthetic Diagnostics Deterministic synthetic controlled-intervention summaries used by DavidHatley/system-one-mini. The records contain no real coding-agent traces, private repositories, personal data, external model outputs, or teacher labels. This dataset and model are independent research artifacts, not reproductions of Jev or RLCD. Splits Split Rows train 20,000 validation 2,000 calibration 2,000 development_renderer 2,000… See the full description on the dataset page: https://huggingface.co/datasets/DavidHatley/system-one-mini-data.

sourceHugging Faceapache-2.0updated 20d agoView on Hugging Face
0likes255downloads
Dataset Card

System One Mini Synthetic Diagnostics

Deterministic synthetic controlled-intervention summaries used by DavidHatley/system-one-mini. The records contain no real coding-agent traces, private repositories, personal data, external model outputs, or teacher labels.

This dataset and model are independent research artifacts, not reproductions of Jev or RLCD.

Splits

SplitRows
train20,000
validation2,000
calibration2,000
development_renderer2,000
test2,000

development_renderer is the former sealed renderer split that was subsequently used to compare candidates. It must not be described as an unbiased final test. test is a separate renderer written and frozen after model selection, then evaluated once for the initial release.

Fields

  • —id: SHA-256 fingerprint of the latent synthetic situation
  • —state: rendered diagnostic summary
  • —labels: five numeric labels in schema order
  • —Five named string label fields
  • —latent: complete generated situation used to compute labels
  • —provenance: generator, split, renderer, and seed

Generation

Generation code and manifests are included in generator/ and provenance/. The policy makes labels mechanically identifiable: every component trial starts independently from the same original failing snapshot, and exactly one clearing intervention identifies the operational failure source.

Limitations

  • —All scenarios are synthetic and templated.
  • —The label taxonomy covers only four component classes plus unknown.
  • —Recurring logical patterns across splits do not represent natural incident prevalence.
  • —Public test labels make the split unsuitable as a permanently sealed benchmark for future model selection.