Team Ai
Datasetpublic

BlackLansky/cwi-corpus

CWI Public Corpus Version 1.0.0 — built by the CWI dataset-pipeline skill (FineWeb-shaped: trafilatura extraction → quality filters → MinHash dedup → BPE tokenization). Contents structured-record: 163 docs text-file: 9 docs unknown: 9 docs Provenance All documents come from public CWI sources only: the public BlackLansky/cwi-catalog Hugging Face dataset, public GitHub Pages (cumulativewebinc.github.io), and CWI's own published articles. Every row… See the full description on the dataset page: https://huggingface.co/datasets/BlackLansky/cwi-corpus.

sourceHugging Faceotherupdated 8d agoView on Hugging Face
0likes162downloads
Dataset Card

CWI Public Corpus

Version 1.0.0 — built by the CWI dataset-pipeline skill (FineWeb-shaped: trafilatura extraction → quality filters → MinHash dedup → BPE tokenization).

Contents

  • —structured-record: 163 docs
  • —text-file: 9 docs
  • —unknown: 9 docs

Provenance

All documents come from public CWI sources only: the public BlackLansky/cwi-catalog Hugging Face dataset, public GitHub Pages (cumulativewebinc.github.io), and CWI's own published articles. Every row carries source_url and collected_at. No private-repo or non-public data.

Pipeline

stageinout
collect—181
extract (trafilatura)181172
quality filters172171
MinHash dedup (thr=0.85)171159
tokenize159159 (31,879 tokens)

Validation

See validation-report.json / validation-report.md in this repo.

Intended use

Pretraining / continued-pretraining and retrieval corpora for CWI music-catalog understanding, playlist-fit scoring, and trust-verdict classification. Not for production decisions without human review.