datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
host-index-testing-v2
Common Crawl Host Index v2
GitHub: https://github.com/commoncrawl/cc-host-index
Each crawl, we generate a Host Index, which aggregates information about each web hosted visited during the crawl. The
information is aggregated from the Common Crawl columnar index,
web graph, and raw crawler logs.
Quickstart
The dataset is Hive-partitioned on crawl (data/crawl=CC-MAIN-2025-18/*.parquet). Open the whole
dataset once, then filter with WHERE crawl = '...': because… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/host-index-testing-v2.common-crawl-sample
Common Crawl sample
A small unofficial random subset of the famous Common Crawl dataset.
60 random segment WET files were downloaded from Common Crawl on 2024-05-12.
Lines between 500 and 5000 characters long (inclusive) were kept.
Only unique texts were kept.
No other filtering.
Languages
Each text was assigned to one of the language codes using the GCLD3 Python package.
The Chinese texts were classified as either simplified, traditional, or Cantonese using the… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/common-crawl-sample.CommonCrawl-CreativeCommons
The Common Crawl Creative Commons Corpus (C5)
Raw CommonCrawl crawls, annotated with Creative Commons license information
C5 is an effort to collect Creative Commons-licensed web data in one place.
The licensing information is extracted from the web pages based on whether they link to Creative Commons licenses either overtly in a tags (like in the footer of Wikipedia) or in metadata fields indicating deliberate Creative Commons publication. However, false positives may occur! See… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/CommonCrawl-CreativeCommons.CommonCrawl-CreativeCommons-fine
Common Crawl Creative Commons Corpus Fine (C5f)
A filtered version of the Common Crawl Creative Commons Corpus (C5), only retaining samples that are also present in the FineWeb or FineWeb-2 datasets. As such, this dataset contains a high-quality subset of C5.
Created with this script.
For more information, see C5.
Progress
In the v1 release, the following crawls are included
CC-MAIN-2019-30
CC-MAIN-2020-05CC-MAIN-2023-06
CC-MAIN-2024-51
CC-MAIN-2024-46
CC-MAIN-2025-05… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/CommonCrawl-CreativeCommons-fine.CommonCrawl-CreativeCommons-strict
Common Crawl Creative Commons Corpus Strict (C5s)
A filtered version of the Common Crawl Creative Commons Corpus (C5), only retaining samples that:
are also present in the FineWeb or FineWeb-2 datasets;
have no license disagreement (all found licenses have the same type; version number might differ);
are not "non-commercial" ("nc" in license);
are not "cc-unknown";
do not have "wiki" in their name (the idea is that you should include Wikipedia and other Wikidata from other… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/CommonCrawl-CreativeCommons-strict.common-crawl-english-filtered
🧠 FineWeb-English-Filtered
📘 Dataset Summary
FineWeb-English-Filtered is a large-scale, cleaned, English-only text dataset derived from Common Crawl’s WET archives.It contains 940 million documents of publicly available web text, converted into Apache Parquet format with a consistent schema for fast and efficient data loading.
The dataset was generated using a custom AWS Glue pipeline that processed, filtered, and merged .wet files across multiple terabytes of… See the full description on the dataset page: https://huggingface.co/datasets/anandjh8/common-crawl-english-filtered.commoncrawl_dataset
Kyrgyz CommonCrawl Dataset
A 271 MB text corpus of Kyrgyz language data extracted from CommonCrawl — one of the largest openly available Kyrgyz text collections for NLP research.
Dataset Description
This dataset contains Kyrgyz-language web text scraped from CommonCrawl archives, filtered by the Kyrgyz language tag (ky). The data covers a wide range of domains including news, blogs, government sites, educational content, and general web pages.
Why this matters: Kyrgyz is… See the full description on the dataset page: https://huggingface.co/datasets/Zarinaaa/commoncrawl_dataset.common-crawl-zhtw
Dataset Card for Common Crawl Traditional Chinese
De-duplicated version of jed351/Traditional-Chinese-Common-Crawl-Filtered.
De-duplicated with MinHash
Is suggested to filter the dataset with NLU models before any serious use.
common-crawl-docx-sample
Common Crawl DOCX Sample
A sample of normalized text extracted from DOCX records in Common Crawl.
Source
Common Crawl release: CC-MAIN-YYYY-NN
Source index: Common Crawl URL Index
Pipeline: marin-community/marin
Pipeline revision: REPLACE_WITH_GIT_SHA
Records were selected using declared DOCX MIME type, detected DOCX MIME type,
or a .docx URL suffix. Only successful, non-truncated index records were
eligible.
Processing
The pipeline:… See the full description on the dataset page: https://huggingface.co/datasets/Tinuade/common-crawl-docx-sample.common-crawl-20-warcs-en-te
Common Crawl, 20 WARCs: what an English + Telugu pipeline keeps and removes
430,279 web records from Common Crawl CC-MAIN-2026-39 went through extraction, URL filtering and language
identification. 139,643 English or Telugu pages were kept, and every page the language filter removed is here too,
with the reason, so you can see exactly what a filter throws away.
Processed with dataflow on Hugging Face Jobs: 2 cpu-upgrade Jobs,
75 Job-minutes, $0.04. The run's live dashboard:… See the full description on the dataset page: https://huggingface.co/datasets/kuluruvineeth/common-crawl-20-warcs-en-te.GeoGPT_Training_Data_from_Geoscience_Subset_of_CommonCrawl
Description
This dataset is a geoscience-specific subset of CommonCrawl used for GeoGPT training. CommonCrawl is a free and open repository of web crawl data with over 250 billion web pages and is widely used by leading large language models. We apply data mining algorithms to extract geoscience-related content from this vast dataset.
This dataset comprises 12,414,268 samples, each containing the following metadata to trace the data source within CommonCrawl:
id (string):… See the full description on the dataset page: https://huggingface.co/datasets/GeoGPT-Research-Project/GeoGPT_Training_Data_from_Geoscience_Subset_of_CommonCrawl.
