ouroboroscollective/evidence-bound-css
Evidence-Bound CSS A provenance-first German CSS/HTML learning, debugging and repair corpus for LLM training. Release 0.4.1 The public release exposes 601 unique structured CSS knowledge records, an alternate 601-record SFT view, 100 source-grounded debugging records, the original 6-record browser-verified seed, plus a v0.3 quality layer with 120 runtime-verified repairs, 120 verifier-confirmed hard negatives, and 120 verifier-backed preference pairs across 30… See the full description on the dataset page: https://huggingface.co/datasets/ouroboroscollective/evidence-bound-css.
Evidence-Bound CSS
A provenance-first German CSS/HTML learning, debugging and repair corpus for LLM training.
Release 0.4.1
The public release exposes 601 unique structured CSS knowledge records, an alternate 601-record SFT view, 100 source-grounded debugging records, the original 6-record browser-verified seed, plus a v0.3 quality layer with 120 runtime-verified repairs, 120 verifier-confirmed hard negatives, and 120 verifier-backed preference pairs across 30 CSS/DOM behavior families.
The configs are alternate views of overlapping source material. Do not add their row counts together as independent training examples.
Unique source inventory
Configs
knowledge_de — canonical knowledge view
Structured German CSS knowledge records with stable IDs, task/category metadata, explanatory content, code where applicable, provenance state and integrity hashes.
sft_de — supervised fine-tuning view
The same 601 underlying items represented as instruction/response training examples. This is a view, not 601 additional independent facts.
recipes_de — compact compatibility view
The original 32 robust CSS solution recipes published in v0.1.0.
debug_reasoning_de — 100 source-grounded debugging records
Deterministically parsed from the 100 existing debugging cases. Each row preserves diagnosis, targeted fix, limitations, counter-check, source record hash and family grouping. These records are source-grounded but are marked runtime_status: not_replayed unless separately browser-tested.
browser_verified_repairs_de — runtime-admitted repair pairs
A deliberately small high-confidence layer of synthetic defect/repair pairs. Each admitted row must demonstrate both conditions in Chromium: the failure predicate is true before the fix and the success predicate is true after the fix.
The first suite admitted 6 / 8 candidates on Chromium 144.0.7559.96. Two candidates were rejected because the broken state was not reproducible enough in that runtime. Rejected candidates are not published as ground truth.
Source provenance and publication authorization
The source package is CSS_Wissenspaket_DE.zip, SHA-256:
12353f528865470a753efc860889a0d82156455a9cc1145a28ce3abd525f24e7
On 2026-10-04, the publisher explicitly attested that the PDFs and derived material in this package were privately created for them using a research/learning GPT workflow, translated and structured for their project, and explicitly authorized publication of the full corpus here.
This repository records that statement as a publisher authorization / provenance attestation. It is not an independent legal finding about copyright ownership.
Downstream reuse terms
No explicit standard open-source/content license (MIT, Apache-2.0, CC BY, etc.) was named in that attestation. Therefore this release does not invent one. The provenance records distinguish:
publication_status: publisher_authorizeddownstream_license: unspecified
This keeps the corpus publicly inspectable and publishable under the publisher's authorization without silently choosing legal terms on their behalf. A later release can add an explicit downstream license without changing the data lineage.
Reproduced package validation
Before publication the supplied package was re-checked:
- 233 / 233 SHA-256 manifest entries matched.
- 130 / 130 local file and fragment targets passed.
- 158 / 158 local Chromium browser checks passed.
- Local validation browser: Chromium 144.0.7559.96.
A separate Hugging Face compute smoke check successfully launched Chromium, Firefox and WebKit engines. That engine smoke is not represented as a 158-test cross-browser pass.
Truth boundary
The 601-item corpus is training/knowledge material, not a claim of state-of-the-art model performance and not a contamination-free benchmark.
We intentionally do not manufacture random validation/test splits from this single source package. Future benchmark splits should be source- and family-disjoint, near-deduplicated against training sources, and certified by browser replay.
Third-party intake policy
Third-party sources remain fail-closed:
TheOdinProject/css-exercises: MIT-licensed candidate/reference family; retain immutable source revision and attribution.sarvinoz23/html_css_website1: not ingested; explicit repository and asset rights are still unresolved.
The publisher authorization for CSS_Wissenspaket_DE.zip does not automatically extend to unrelated third-party repositories.
Recommended use
Use knowledge_de for retrieval, curriculum construction, classification and knowledge-grounded CSS training. Use sft_de for supervised instruction tuning. Use recipes_de only when a small focused recipe subset is preferred.
The quality layer now follows:
defect → repair candidate → browser replay → failure-before gate → success-after gate → evidence receipt → training admission
Current runtime-admitted seed: 6 browser-verified repairs. The suite is intentionally small; unproven mutations are rejected rather than padded into the dataset.
The goal is not raw-code volume. The goal is evidence-bound CSS supervision with inspectable provenance and reproducible acceptance rules.
Quick start
from datasets import load_dataset
knowledge = load_dataset(
"ouroboroscollective/evidence-bound-css",
"knowledge_de",
split="train",
)
sft = load_dataset(
"ouroboroscollective/evidence-bound-css",
"sft_de",
split="train",
)
print(len(knowledge)) # 601
print(len(sft)) # 601, alternate view of the same underlying corpusRelease evidence
See source_registry.json for the machine-readable rights, source and validation boundary.
v0.3 quality layer
browser_verified_repairs_v2_de — 120 runtime-verified repairs
The v0.3 suite covers 30 distinct CSS/DOM behavior families with four parameter variants each. Every admitted row satisfies the same three runtime gates in Chromium 144.0.7559.96:
- the broken state fails the success predicate;
- the proposed repair passes the success predicate;
- a plausible hard-negative repair still fails the same predicate.
The suite covers intrinsic flex/grid sizing, box sizing, percentage-plus-gap overflow, wrapping and ellipsis, cross-axis alignment, containing blocks, scroll panels, intrinsic image ratio, responsive sizing, min-width conflicts, preformatted text, table scrolling, media queries, custom-property fallbacks, specificity, hover and checked state selectors, :has(), transform centering, z-index positioning, pointer-event overlays, calc() syntax, equal flex columns, reduced motion, dark color scheme, print media and grid-content overflow.
hard_negatives_de — 120 verifier-confirmed wrong repairs
Each row is paired to the exact HTML, defect family, split group and success predicate of a runtime-verified positive. The negative candidate was replayed and failed the predicate. These are hard negatives, not merely model-generated alternatives.
preference_pairs_de — 120 runtime-derived preferences
Each row stores the passing repair as chosen_css and its verifier-confirmed failing alternative as rejected_css. Preference labels therefore come from runtime evidence rather than subjective model ranking.
Admission evidence
Final generator result: 120 candidates → 120 admitted → 0 rejected, across 30 families × 4 variants. All broken states were observed to fail, all chosen repairs passed, and all hard negatives failed. Execution was split into four batches only to stay inside tool runtime limits; the admission predicate was identical in every batch.
The v0.3 task views overlap by construction and must not be counted as 360 new independent CSS facts.
v0.4 benchmark layer
v0.4 introduces a family-disjoint benchmark derived from the 30 runtime-verified v0.3 CSS/DOM behavior families.
Split policy
Families are assigned deterministically after lexicographic ordering:
- train: 20 families / 80 rows
- validation: 5 families / 20 rows
- test: 5 families / 20 rows
All four parameter variants of a family stay in the same split.
Leakage checks
The published benchmark manifest reports:
- family overlap train↔validation: 0
- family overlap train↔test: 0
- family overlap validation↔test: 0
- normalized exact HTML/CSS/predicate collisions across splits: 0
Test replay
The 20 repair rows in the test split were replayed separately on Chromium 144.0.7559.96.
Result: 20 / 20 PASS.
For all 20 rows:
- the broken state failed the success predicate;
- the repaired state passed the same success predicate.
The test families are:
specificity_override, table_scroll_wrapper, text_ellipsis, two_columns_gap, z_index_positioning.
Benchmark truth boundary
This benchmark is family-disjoint within this dataset. It is not claimed contamination-free against the public web or external model pretraining corpora. The manifest and replay receipt make that boundary explicit.
v0.4.1 split materialization fix
The benchmark rows were already labeled with the correct family-disjoint benchmark_split, but the Dataset Card mapped train, validation and test to the same 120-row JSONL file. This made the Dataset Viewer expose 120 rows for every split.
v0.4.1 materializes separate physical files for each benchmark config:
- train: 80 rows
- validation: 20 rows
- test: 20 rows
The underlying v0.4 family assignment, leakage checks and 20/20 test replay remain unchanged.
