Tinuade/common-crawl-docx-sample
Common Crawl DOCX Sample A sample of normalized text extracted from DOCX records in Common Crawl. Source Common Crawl release: CC-MAIN-YYYY-NN Source index: Common Crawl URL Index Pipeline: marin-community/marin Pipeline revision: REPLACE_WITH_GIT_SHA Records were selected using declared DOCX MIME type, detected DOCX MIME type, or a .docx URL suffix. Only successful, non-truncated index records were eligible. Processing The pipeline:… See the full description on the dataset page: https://huggingface.co/datasets/Tinuade/common-crawl-docx-sample.
# Common Crawl DOCX Sample
A sample of normalized text extracted from DOCX records in Common Crawl.
## Source
- Common Crawl release:
CC-MAIN-YYYY-NN - Source index: Common Crawl URL Index
- Pipeline:
marin-community/marin - Pipeline revision:
REPLACE_WITH_GIT_SHA
Records were selected using declared DOCX MIME type, detected DOCX MIME type, or a .docx URL suffix. Only successful, non-truncated index records were eligible.
## Processing
The pipeline:
- discovered records in the Common Crawl URL Index;
- retrieved and verified the corresponding WARC records;
- validated the DOCX containers;
- extracted text and tables with Docling 2.99.0;
- identified language with Lingua 2.2.0;
- normalized and exactly deduplicated the extracted text.
## Sampling
Describe the sampling method here, including:
- the number of source documents;
- the number of published documents;
- the random seed or deterministic selection rule;
- whether the sample was stratified by language, table presence, or selection reason.
## Schema
id: normalized document identifiertext: normalized extracted textcrawl_id: Common Crawl releaseurl: source URLlanguage: ISO 639-1 language prediction orunknownlanguage_score: detector confidenceword_count: extracted word counttable_count: number of extracted tablesimage_count: number of detected imagesselection_reason: URL Index signal that selected the recordextractor: extraction implementation and versionlanguage_detector: language detector and version
The Parquet files contain additional WARC and extraction provenance fields.
## Licensing and limitations
Common Crawl provides access to crawled content but does not grant a uniform license over the underlying documents. Individual documents may remain subject to copyright, privacy rights, and source-site terms. This dataset is therefore published without assigning a single license to the underlying text.
The sample may contain extraction errors, personal information, biased content, or material unsuitable for downstream use. Users are responsible for reviewing the applicable rights and restrictions.
