Team Ai
Datasetpublic

Tinuade/common-crawl-docx-sample

Common Crawl DOCX Sample A sample of normalized text extracted from DOCX records in Common Crawl. Source Common Crawl release: CC-MAIN-YYYY-NN Source index: Common Crawl URL Index Pipeline: marin-community/marin Pipeline revision: REPLACE_WITH_GIT_SHA Records were selected using declared DOCX MIME type, detected DOCX MIME type, or a .docx URL suffix. Only successful, non-truncated index records were eligible. Processing The pipeline:… See the full description on the dataset page: https://huggingface.co/datasets/Tinuade/common-crawl-docx-sample.

sourceHugging Faceotherupdated 24d agoView on Hugging Face
0likes71downloads
Dataset Card

# Common Crawl DOCX Sample

A sample of normalized text extracted from DOCX records in Common Crawl.

## Source

  • —Common Crawl release: CC-MAIN-YYYY-NN
  • —Source index: Common Crawl URL Index
  • —Pipeline: marin-community/marin
  • —Pipeline revision: REPLACE_WITH_GIT_SHA

Records were selected using declared DOCX MIME type, detected DOCX MIME type, or a .docx URL suffix. Only successful, non-truncated index records were eligible.

## Processing

The pipeline:

  1. 1.discovered records in the Common Crawl URL Index;
  2. 2.retrieved and verified the corresponding WARC records;
  3. 3.validated the DOCX containers;
  4. 4.extracted text and tables with Docling 2.99.0;
  5. 5.identified language with Lingua 2.2.0;
  6. 6.normalized and exactly deduplicated the extracted text.

## Sampling

Describe the sampling method here, including:

  • —the number of source documents;
  • —the number of published documents;
  • —the random seed or deterministic selection rule;
  • —whether the sample was stratified by language, table presence, or selection reason.

## Schema

  • —id: normalized document identifier
  • —text: normalized extracted text
  • —crawl_id: Common Crawl release
  • —url: source URL
  • —language: ISO 639-1 language prediction or unknown
  • —language_score: detector confidence
  • —word_count: extracted word count
  • —table_count: number of extracted tables
  • —image_count: number of detected images
  • —selection_reason: URL Index signal that selected the record
  • —extractor: extraction implementation and version
  • —language_detector: language detector and version

The Parquet files contain additional WARC and extraction provenance fields.

## Licensing and limitations

Common Crawl provides access to crawled content but does not grant a uniform license over the underlying documents. Individual documents may remain subject to copyright, privacy rights, and source-site terms. This dataset is therefore published without assigning a single license to the underlying text.

The sample may contain extraction errors, personal information, biased content, or material unsuitable for downstream use. Users are responsible for reviewing the applicable rights and restrictions.