BlackLansky/cwi-corpus
CWI Public Corpus Version 1.0.0 — built by the CWI dataset-pipeline skill (FineWeb-shaped: trafilatura extraction → quality filters → MinHash dedup → BPE tokenization). Contents structured-record: 163 docs text-file: 9 docs unknown: 9 docs Provenance All documents come from public CWI sources only: the public BlackLansky/cwi-catalog Hugging Face dataset, public GitHub Pages (cumulativewebinc.github.io), and CWI's own published articles. Every row… See the full description on the dataset page: https://huggingface.co/datasets/BlackLansky/cwi-corpus.
CWI Public Corpus
Version 1.0.0 — built by the CWI dataset-pipeline skill (FineWeb-shaped: trafilatura extraction → quality filters → MinHash dedup → BPE tokenization).
Contents
structured-record: 163 docstext-file: 9 docsunknown: 9 docs
Provenance
All documents come from public CWI sources only: the public BlackLansky/cwi-catalog Hugging Face dataset, public GitHub Pages (cumulativewebinc.github.io), and CWI's own published articles. Every row carries source_url and collected_at. No private-repo or non-public data.
Pipeline
Validation
See validation-report.json / validation-report.md in this repo.
Intended use
Pretraining / continued-pretraining and retrieval corpora for CWI music-catalog understanding, playlist-fit scoring, and trust-verdict classification. Not for production decisions without human review.
