Team Ai
Datasetpublic

ouroboroscollective/evidence-bound-css

Evidence-Bound CSS A provenance-first German CSS/HTML learning, debugging and repair corpus for LLM training. Release 0.4.1 The public release exposes 601 unique structured CSS knowledge records, an alternate 601-record SFT view, 100 source-grounded debugging records, the original 6-record browser-verified seed, plus a v0.3 quality layer with 120 runtime-verified repairs, 120 verifier-confirmed hard negatives, and 120 verifier-backed preference pairs across 30… See the full description on the dataset page: https://huggingface.co/datasets/ouroboroscollective/evidence-bound-css.

sourceHugging Faceupdated 5d agoView on Hugging Face
1likes169downloads
Dataset Card

Evidence-Bound CSS

A provenance-first German CSS/HTML learning, debugging and repair corpus for LLM training.

Release 0.4.1

The public release exposes 601 unique structured CSS knowledge records, an alternate 601-record SFT view, 100 source-grounded debugging records, the original 6-record browser-verified seed, plus a v0.3 quality layer with 120 runtime-verified repairs, 120 verifier-confirmed hard negatives, and 120 verifier-backed preference pairs across 30 CSS/DOM behavior families.

The configs are alternate views of overlapping source material. Do not add their row counts together as independent training examples.

Unique source inventory

FamilyRecords
CSS properties279
Selector / state patterns74
CSS functions78
At-rules18
Debugging / error cases100
Solution recipes32
Complete offline examples20
Total unique records601

Configs

knowledge_de — canonical knowledge view

Structured German CSS knowledge records with stable IDs, task/category metadata, explanatory content, code where applicable, provenance state and integrity hashes.

sft_de — supervised fine-tuning view

The same 601 underlying items represented as instruction/response training examples. This is a view, not 601 additional independent facts.

recipes_de — compact compatibility view

The original 32 robust CSS solution recipes published in v0.1.0.

debug_reasoning_de — 100 source-grounded debugging records

Deterministically parsed from the 100 existing debugging cases. Each row preserves diagnosis, targeted fix, limitations, counter-check, source record hash and family grouping. These records are source-grounded but are marked runtime_status: not_replayed unless separately browser-tested.

browser_verified_repairs_de — runtime-admitted repair pairs

A deliberately small high-confidence layer of synthetic defect/repair pairs. Each admitted row must demonstrate both conditions in Chromium: the failure predicate is true before the fix and the success predicate is true after the fix.

The first suite admitted 6 / 8 candidates on Chromium 144.0.7559.96. Two candidates were rejected because the broken state was not reproducible enough in that runtime. Rejected candidates are not published as ground truth.

Source provenance and publication authorization

The source package is CSS_Wissenspaket_DE.zip, SHA-256:

12353f528865470a753efc860889a0d82156455a9cc1145a28ce3abd525f24e7

On 2026-10-04, the publisher explicitly attested that the PDFs and derived material in this package were privately created for them using a research/learning GPT workflow, translated and structured for their project, and explicitly authorized publication of the full corpus here.

This repository records that statement as a publisher authorization / provenance attestation. It is not an independent legal finding about copyright ownership.

Downstream reuse terms

No explicit standard open-source/content license (MIT, Apache-2.0, CC BY, etc.) was named in that attestation. Therefore this release does not invent one. The provenance records distinguish:

  • —publication_status: publisher_authorized
  • —downstream_license: unspecified

This keeps the corpus publicly inspectable and publishable under the publisher's authorization without silently choosing legal terms on their behalf. A later release can add an explicit downstream license without changing the data lineage.

Reproduced package validation

Before publication the supplied package was re-checked:

  • —233 / 233 SHA-256 manifest entries matched.
  • —130 / 130 local file and fragment targets passed.
  • —158 / 158 local Chromium browser checks passed.
  • —Local validation browser: Chromium 144.0.7559.96.

A separate Hugging Face compute smoke check successfully launched Chromium, Firefox and WebKit engines. That engine smoke is not represented as a 158-test cross-browser pass.

Truth boundary

The 601-item corpus is training/knowledge material, not a claim of state-of-the-art model performance and not a contamination-free benchmark.

We intentionally do not manufacture random validation/test splits from this single source package. Future benchmark splits should be source- and family-disjoint, near-deduplicated against training sources, and certified by browser replay.

Third-party intake policy

Third-party sources remain fail-closed:

  • —TheOdinProject/css-exercises: MIT-licensed candidate/reference family; retain immutable source revision and attribution.
  • —sarvinoz23/html_css_website1: not ingested; explicit repository and asset rights are still unresolved.

The publisher authorization for CSS_Wissenspaket_DE.zip does not automatically extend to unrelated third-party repositories.

Recommended use

Use knowledge_de for retrieval, curriculum construction, classification and knowledge-grounded CSS training. Use sft_de for supervised instruction tuning. Use recipes_de only when a small focused recipe subset is preferred.

The quality layer now follows:

defect → repair candidate → browser replay → failure-before gate → success-after gate → evidence receipt → training admission

Current runtime-admitted seed: 6 browser-verified repairs. The suite is intentionally small; unproven mutations are rejected rather than padded into the dataset.

The goal is not raw-code volume. The goal is evidence-bound CSS supervision with inspectable provenance and reproducible acceptance rules.

Quick start

python
from datasets import load_dataset

knowledge = load_dataset(
    "ouroboroscollective/evidence-bound-css",
    "knowledge_de",
    split="train",
)

sft = load_dataset(
    "ouroboroscollective/evidence-bound-css",
    "sft_de",
    split="train",
)

print(len(knowledge))  # 601
print(len(sft))        # 601, alternate view of the same underlying corpus

Release evidence

See source_registry.json for the machine-readable rights, source and validation boundary.

v0.3 quality layer

browser_verified_repairs_v2_de — 120 runtime-verified repairs

The v0.3 suite covers 30 distinct CSS/DOM behavior families with four parameter variants each. Every admitted row satisfies the same three runtime gates in Chromium 144.0.7559.96:

  1. 1.the broken state fails the success predicate;
  2. 2.the proposed repair passes the success predicate;
  3. 3.a plausible hard-negative repair still fails the same predicate.

The suite covers intrinsic flex/grid sizing, box sizing, percentage-plus-gap overflow, wrapping and ellipsis, cross-axis alignment, containing blocks, scroll panels, intrinsic image ratio, responsive sizing, min-width conflicts, preformatted text, table scrolling, media queries, custom-property fallbacks, specificity, hover and checked state selectors, :has(), transform centering, z-index positioning, pointer-event overlays, calc() syntax, equal flex columns, reduced motion, dark color scheme, print media and grid-content overflow.

hard_negatives_de — 120 verifier-confirmed wrong repairs

Each row is paired to the exact HTML, defect family, split group and success predicate of a runtime-verified positive. The negative candidate was replayed and failed the predicate. These are hard negatives, not merely model-generated alternatives.

preference_pairs_de — 120 runtime-derived preferences

Each row stores the passing repair as chosen_css and its verifier-confirmed failing alternative as rejected_css. Preference labels therefore come from runtime evidence rather than subjective model ranking.

Admission evidence

Final generator result: 120 candidates → 120 admitted → 0 rejected, across 30 families × 4 variants. All broken states were observed to fail, all chosen repairs passed, and all hard negatives failed. Execution was split into four batches only to stay inside tool runtime limits; the admission predicate was identical in every batch.

The v0.3 task views overlap by construction and must not be counted as 360 new independent CSS facts.

v0.4 benchmark layer

v0.4 introduces a family-disjoint benchmark derived from the 30 runtime-verified v0.3 CSS/DOM behavior families.

Split policy

Families are assigned deterministically after lexicographic ordering:

  • —train: 20 families / 80 rows
  • —validation: 5 families / 20 rows
  • —test: 5 families / 20 rows

All four parameter variants of a family stay in the same split.

Leakage checks

The published benchmark manifest reports:

  • —family overlap train↔validation: 0
  • —family overlap train↔test: 0
  • —family overlap validation↔test: 0
  • —normalized exact HTML/CSS/predicate collisions across splits: 0

Test replay

The 20 repair rows in the test split were replayed separately on Chromium 144.0.7559.96.

Result: 20 / 20 PASS.

For all 20 rows:

  • —the broken state failed the success predicate;
  • —the repaired state passed the same success predicate.

The test families are:

specificity_override, table_scroll_wrapper, text_ellipsis, two_columns_gap, z_index_positioning.

Benchmark truth boundary

This benchmark is family-disjoint within this dataset. It is not claimed contamination-free against the public web or external model pretraining corpora. The manifest and replay receipt make that boundary explicit.

v0.4.1 split materialization fix

The benchmark rows were already labeled with the correct family-disjoint benchmark_split, but the Dataset Card mapped train, validation and test to the same 120-row JSONL file. This made the Dataset Viewer expose 120 rows for every split.

v0.4.1 materializes separate physical files for each benchmark config:

  • —train: 80 rows
  • —validation: 20 rows
  • —test: 20 rows

The underlying v0.4 family assignment, leakage checks and 20/20 test replay remain unchanged.