datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
common-crawl-character-countswdc-common-crawl-embedded-jsonldgneissweb-annotation-url-testing-v1
GneissWeb Annotations
GneissWeb Annotations, powered by IBM Research's GneissWeb methodology, is a dataset of quality and category annotations applied to the Common Crawl corpus.
This dataset enables precise filtering of web content across medical, educational, technology, and scientific domains, making it easier to build high-quality corpora for research projects, language models, and specialized applications.
Learn more about the annotation process and methodology in our… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/gneissweb-annotation-url-testing-v1.commoncrawl-tr
Dataset Card for "commoncrawl-tr"
More Information needed
thai-commoncrawl-index
Thai Common Crawl Index (2019–2026)
An index of every page Common Crawl detected as Thai across 70 monthly crawls, from
January 2019 (CC-MAIN-2019-04) to August 2026 (CC-MAIN-2026-30).
932,874,727 page captures · 450,971,497 unique URLs · 6,997,185 hosts · 6,674,969 domains
Each row records where the page lives inside Common Crawl's WARC archives — file name,
byte offset, and record length — so you can fetch exactly the pages you want with HTTP
range requests, without scanning… See the full description on the dataset page: https://huggingface.co/datasets/wayu-ai/thai-commoncrawl-index.host-index-testing-v2
Common Crawl Host Index v2
GitHub: https://github.com/commoncrawl/cc-host-index
Each crawl, we generate a Host Index, which aggregates information about each web hosted visited during the crawl. The
information is aggregated from the Common Crawl columnar index,
web graph, and raw crawler logs.
Quickstart
The dataset is Hive-partitioned on crawl (data/crawl=CC-MAIN-2025-18/*.parquet). Open the whole
dataset once, then filter with WHERE crawl = '...': because… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/host-index-testing-v2.common-crawl-sample
Common Crawl sample
A small unofficial random subset of the famous Common Crawl dataset.
60 random segment WET files were downloaded from Common Crawl on 2024-05-12.
Lines between 500 and 5000 characters long (inclusive) were kept.
Only unique texts were kept.
No other filtering.
Languages
Each text was assigned to one of the language codes using the GCLD3 Python package.
The Chinese texts were classified as either simplified, traditional, or Cantonese using the… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/common-crawl-sample.sea-commoncrawlCommonCrawl-CreativeCommons
The Common Crawl Creative Commons Corpus (C5)
Raw CommonCrawl crawls, annotated with Creative Commons license information
C5 is an effort to collect Creative Commons-licensed web data in one place.
The licensing information is extracted from the web pages based on whether they link to Creative Commons licenses either overtly in a tags (like in the footer of Wikipedia) or in metadata fields indicating deliberate Creative Commons publication. However, false positives may occur! See… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/CommonCrawl-CreativeCommons.Traditional-Chinese-Common-Crawl-Filtered
Traditional Chinese C4
Dataset Summary
Data obtained from 2013~2025 Common Crawl.
Downloaded and processed using code based on another project attempting to recreate the C4 dataset.
The resultant dataset contains both simplified and traditional Chinese, which could be found here.
It was then filtered using a modified list of simplified Chinese characters to obtain this traditional Chinese dataset.
Unfortunately, I don't have enough funding to run a deduplication across… See the full description on the dataset page: https://huggingface.co/datasets/jed351/Traditional-Chinese-Common-Crawl-Filtered.Traditional-Chinese-Common-Crawl-NOT-CleanedCommon Crawl Dumps that were briefly filtered by keywords to remove bad words and simplified Chinese.
The hash based cleaned dataset can be found here.
Files here are for future usage (downloading from Common Crawl and keyword filtering are very slow)
common_crawl_pointers_by_collectionsea-commoncrawl-high-qualityCommonCrawl-CreativeCommons-fine
Common Crawl Creative Commons Corpus Fine (C5f)
A filtered version of the Common Crawl Creative Commons Corpus (C5), only retaining samples that are also present in the FineWeb or FineWeb-2 datasets. As such, this dataset contains a high-quality subset of C5.
Created with this script.
For more information, see C5.
Progress
In the v1 release, the following crawls are included
CC-MAIN-2019-30
CC-MAIN-2020-05CC-MAIN-2023-06
CC-MAIN-2024-51
CC-MAIN-2024-46
CC-MAIN-2025-05… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/CommonCrawl-CreativeCommons-fine.FRENCH-ONLY-Common-Crawl-2026-25common_crawl_pointers_by_collectionstatistics
Common Crawl Statistics
Number of pages, distribution of top-level domains, crawl overlaps, etc. - basic metrics about Common Crawl Monthly Crawl Archives, for more detailed information and graphs please visit our official statistics page. Here you can find the following statistics files:
Charsets
The character set or encoding of HTML pages only is identified by Tika's AutoDetectReader. The table shows the percentage how character sets have been used to encode… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/statistics.CommonCrawl-CreativeCommons-strict
Common Crawl Creative Commons Corpus Strict (C5s)
A filtered version of the Common Crawl Creative Commons Corpus (C5), only retaining samples that:
are also present in the FineWeb or FineWeb-2 datasets;
have no license disagreement (all found licenses have the same type; version number might differ);
are not "non-commercial" ("nc" in license);
are not "cc-unknown";
do not have "wiki" in their name (the idea is that you should include Wikipedia and other Wikidata from other… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/CommonCrawl-CreativeCommons-strict.Cantonese_Common_Crawl_Filtered
Cantonese Chinese C4
Dataset Summary
Downloaded and processed using code based on another project attempting to recreate the C4 dataset.
The resultant traditional Chinese dataset can be found here.
This dataset contains data processed with CantoneseDetect.
In CantoneseDetect, you can choose whether to include quotes (i.e. categorise the data as Cantonese even if Cantonese appeared only in quotes).
And I found that a lot of entries came from Wikipedia and LIHKG. If you… See the full description on the dataset page: https://huggingface.co/datasets/jed351/Cantonese_Common_Crawl_Filtered.CommonCrawl_wet_v2commonlid-results
CommonLID results
This dataset contains the results of the CommonLID leaderboard as summaries (aggregated scores like F1) and raw predictions for each dataset-model combination.
See https://github.com/commoncrawl/commonlid-eval/
Chinese-Common-Crawl-Filtered
Traditional Chinese C4
Dataset Summary
Data obtained from 2025-18 and 2025-13 Common Crawl.
Downloaded and processed using code based on another project attempting to recreate the C4 dataset.
The resultant dataset contains both simplified and traditional Chinese.
It was then filtered using a modified list of simplified Chinese characters to obtain another traditional Chinese dataset.
I am still ironning out the process of filtering.
The 2025-13 dataset was… See the full description on the dataset page: https://huggingface.co/datasets/jed351/Chinese-Common-Crawl-Filtered.gneissweb-annotation-host-testing-v1
GneissWeb Annotations
GneissWeb Annotations, powered by IBM Research's GneissWeb methodology, is a dataset of quality and category annotations applied to the Common Crawl corpus.
This dataset enables precise filtering of web content across medical, educational, technology, and scientific domains, making it easier to build high-quality corpora for research projects, language models, and specialized applications.
Learn more about the annotation process and methodology in our… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/gneissweb-annotation-host-testing-v1.common-crawl-2026-21
Common Crawl SEO & AEO/GEO Dataset — CC-MAIN-2026-21
Web pages from the May 2026 Common Crawl (CC-MAIN-2026-21) that mention SEO, Answer Engine Optimization (AEO), or Generative Engine Optimization (GEO) — filtered to English content and ranked by Common Crawl Web Graph harmonic centrality and PageRank.
🔎 Interactive Explorer
A Gradio app to search, filter, and analyze this dataset:
https://huggingface.co/spaces/metehan777/cc-seo-explorer
Tabs: SEO raw search ·… See the full description on the dataset page: https://huggingface.co/datasets/metehan777/common-crawl-2026-21.common_crawl_pointer_indicescommon_crawl_meta_indexesweb-graph-embeddings
Web Graph Embedings from Common Crawl's Host-level Hyperlink Graph
Dense 128-dimensional embeddings for 52,913,544 web hosts, learned by link prediction on the
Common Crawl host-level hyperlink graph release cc-main-2025-26-nov-dec-jan. Vectors are L2-normalized and served in float16;
similarity is cosine (a dot product on the unit vectors).
The dataset contains hosts with link degree >= 8 (total in+out degree), a 52.9 M / ~19 % induced subgraph that carries ~97 % of the edges.… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/web-graph-embeddings.llms.txt
llms.txt files extracted from the Common Crawl corpus
This dataset contains llms.txt files extracted from the Common Crawl corpus.
Specifically, we extracted response records matching */llms.txt and */llms-full.txt with status 200 and MIME-type text/plain or text/markdown.
What is llms.txt?
The /llms.txt file: A proposal to standardise on using an /llms.txt file to provide information to help LLMs use a website at inference time.
Our own analysis of this dataset is available… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/llms.txt.BBT_CommonCrawl_2020
Context
BigBanyanTree is an initiative to empower colleges to set up their data engineering clusters, and drive interest towards data processing and analysis using tools such as Apache Spark. The data provided here is the direct result of this initiative. The data was processed by Gautam and Suchit, under the guidance of Harsh Singhal.
Content
Each arrow file contains a table with fields extracted from Common Crawl WARC files.
The datasets provided are derived from… See the full description on the dataset page: https://huggingface.co/datasets/big-banyan-tree/BBT_CommonCrawl_2020.dead-web-commoncrawl
Dead-Web Common Crawl — a longitudinal host-reachability panel (2018–2026)
· Hugging Face
· Kaggle
· License: CC BY 4.0
An open dataset labeling the reachability of 172,959,928 registered domains across 80 monthly
Common Crawl archives (CC-MAIN-2018-05 → CC-MAIN-2026-21), built from Common Crawl's robotstxt
subset and calibrated against a live re-probe to separate genuinely-dead domains from those merely
blocking crawlers or dropped from the… See the full description on the dataset page: https://huggingface.co/datasets/crawlora-net/dead-web-commoncrawl.
