Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01yvfu /common-crawl-character-counts0 likes37k downloads11mo agoHugging Face02permutans /wdc-common-crawl-embedded-jsonldtext10B<n<100B4 likes9.7k downloads2y agoHugging Face03commoncrawl /gneissweb-annotation-url-testing-v1 GneissWeb Annotations GneissWeb Annotations, powered by IBM Research's GneissWeb methodology, is a dataset of quality and category annotations applied to the Common Crawl corpus. This dataset enables precise filtering of web content across medical, educational, technology, and scientific domains, making it easier to build high-quality corpora for research projects, language models, and specialized applications. Learn more about the annotation process and methodology in our… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/gneissweb-annotation-url-testing-v1.tabular10B<n<100B0 likes9.3k downloads10mo agoHugging Face04musabg /commoncrawl-tr Dataset Card for "commoncrawl-tr" More Information needed text10M<n<100M4 likes9.2k downloads3y agoHugging Face05wayu-ai /thai-commoncrawl-index Thai Common Crawl Index (2019–2026) An index of every page Common Crawl detected as Thai across 70 monthly crawls, from January 2019 (CC-MAIN-2019-04) to August 2026 (CC-MAIN-2026-30). 932,874,727 page captures · 450,971,497 unique URLs · 6,997,185 hosts · 6,674,969 domains Each row records where the page lives inside Common Crawl's WARC archives — file name, byte offset, and record length — so you can fetch exactly the pages you want with HTTP range requests, without scanning… See the full description on the dataset page: https://huggingface.co/datasets/wayu-ai/thai-commoncrawl-index.tabular100M<n<1B0 likes7.9k downloads1mo agoHugging Face06commoncrawl /host-index-testing-v2 Common Crawl Host Index v2 GitHub: https://github.com/commoncrawl/cc-host-index Each crawl, we generate a Host Index, which aggregates information about each web hosted visited during the crawl. The information is aggregated from the Common Crawl columnar index, web graph, and raw crawler logs. Quickstart The dataset is Hive-partitioned on crawl (data/crawl=CC-MAIN-2025-18/*.parquet). Open the whole dataset once, then filter with WHERE crawl = '...': because… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/host-index-testing-v2.tabulartext-generation1B<n<10B0 likes6.8k downloads21d agoHugging Face07agentlans /common-crawl-sample Common Crawl sample A small unofficial random subset of the famous Common Crawl dataset. 60 random segment WET files were downloaded from Common Crawl on 2024-05-12. Lines between 500 and 5000 characters long (inclusive) were kept. Only unique texts were kept. No other filtering. Languages Each text was assigned to one of the language codes using the GCLD3 Python package. The Chinese texts were classified as either simplified, traditional, or Cantonese using the… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/common-crawl-sample.texttext-generation1M<n<10M8 likes6.6k downloads2y agoHugging Face08sailor2 /sea-commoncrawltext100M<n<1B1 likes4.7k downloads2y agoHugging Face09BramVanroy /CommonCrawl-CreativeCommons The Common Crawl Creative Commons Corpus (C5) Raw CommonCrawl crawls, annotated with Creative Commons license information C5 is an effort to collect Creative Commons-licensed web data in one place. The licensing information is extracted from the web pages based on whether they link to Creative Commons licenses either overtly in a tags (like in the footer of Wikipedia) or in metadata fields indicating deliberate Creative Commons publication. However, false positives may occur! See… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/CommonCrawl-CreativeCommons.texttext-generation100M<n<1B41 likes3.2k downloads1y agoHugging Face10jed351 /Traditional-Chinese-Common-Crawl-Filtered Traditional Chinese C4 Dataset Summary Data obtained from 2013~2025 Common Crawl. Downloaded and processed using code based on another project attempting to recreate the C4 dataset. The resultant dataset contains both simplified and traditional Chinese, which could be found here. It was then filtered using a modified list of simplified Chinese characters to obtain this traditional Chinese dataset. Unfortunately, I don't have enough funding to run a deduplication across… See the full description on the dataset page: https://huggingface.co/datasets/jed351/Traditional-Chinese-Common-Crawl-Filtered.text100M<n<1B26 likes2.8k downloads1y agoHugging Face11jed351 /Traditional-Chinese-Common-Crawl-NOT-CleanedCommon Crawl Dumps that were briefly filtered by keywords to remove bad words and simplified Chinese. The hash based cleaned dataset can be found here. Files here are for future usage (downloading from Common Crawl and keyword filtering are very slow) text100M<n<1B0 likes2k downloads1y agoHugging Face12endomorphosis /common_crawl_pointers_by_collection1 likes1.4k downloads8mo agoHugging Face13sailor2 /sea-commoncrawl-high-qualitytext10M<n<100M0 likes1.4k downloads2y agoHugging Face14BramVanroy /CommonCrawl-CreativeCommons-fine Common Crawl Creative Commons Corpus Fine (C5f) A filtered version of the Common Crawl Creative Commons Corpus (C5), only retaining samples that are also present in the FineWeb or FineWeb-2 datasets. As such, this dataset contains a high-quality subset of C5. Created with this script. For more information, see C5. Progress In the v1 release, the following crawls are included CC-MAIN-2019-30 CC-MAIN-2020-05CC-MAIN-2023-06 CC-MAIN-2024-51 CC-MAIN-2024-46 CC-MAIN-2025-05… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/CommonCrawl-CreativeCommons-fine.texttext-generation10M<n<100M5 likes1k downloads1y agoHugging Face15PhysiQuanty /FRENCH-ONLY-Common-Crawl-2026-25tabular1M<n<10M3 likes987 downloads3mo agoHugging Face16Publicus /common_crawl_pointers_by_collection0 likes857 downloads7mo agoHugging Face17commoncrawl /statistics Common Crawl Statistics Number of pages, distribution of top-level domains, crawl overlaps, etc. - basic metrics about Common Crawl Monthly Crawl Archives, for more detailed information and graphs please visit our official statistics page. Here you can find the following statistics files: Charsets The character set or encoding of HTML pages only is identified by Tika's AutoDetectReader. The table shows the percentage how character sets have been used to encode… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/statistics.tabular100K<n<1M30 likes842 downloads16d agoHugging Face18BramVanroy /CommonCrawl-CreativeCommons-strict Common Crawl Creative Commons Corpus Strict (C5s) A filtered version of the Common Crawl Creative Commons Corpus (C5), only retaining samples that: are also present in the FineWeb or FineWeb-2 datasets; have no license disagreement (all found licenses have the same type; version number might differ); are not "non-commercial" ("nc" in license); are not "cc-unknown"; do not have "wiki" in their name (the idea is that you should include Wikipedia and other Wikidata from other… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/CommonCrawl-CreativeCommons-strict.texttext-generation10M<n<100M2 likes778 downloads1y agoHugging Face19jed351 /Cantonese_Common_Crawl_Filtered Cantonese Chinese C4 Dataset Summary Downloaded and processed using code based on another project attempting to recreate the C4 dataset. The resultant traditional Chinese dataset can be found here. This dataset contains data processed with CantoneseDetect. In CantoneseDetect, you can choose whether to include quotes (i.e. categorise the data as Cantonese even if Cantonese appeared only in quotes). And I found that a lot of entries came from Wikipedia and LIHKG. If you… See the full description on the dataset page: https://huggingface.co/datasets/jed351/Cantonese_Common_Crawl_Filtered.text1M<n<10M4 likes631 downloads1y agoHugging Face20hatakeyama-llm-team /CommonCrawl_wet_v2text100K<n<1M1 likes620 downloads3y agoHugging Face21commoncrawl /commonlid-results CommonLID results This dataset contains the results of the CommonLID leaderboard as summaries (aggregated scores like F1) and raw predictions for each dataset-model combination. See https://github.com/commoncrawl/commonlid-eval/ text-classification2 likes574 downloads13d agoHugging Face22jed351 /Chinese-Common-Crawl-Filtered Traditional Chinese C4 Dataset Summary Data obtained from 2025-18 and 2025-13 Common Crawl. Downloaded and processed using code based on another project attempting to recreate the C4 dataset. The resultant dataset contains both simplified and traditional Chinese. It was then filtered using a modified list of simplified Chinese characters to obtain another traditional Chinese dataset. I am still ironning out the process of filtering. The 2025-13 dataset was… See the full description on the dataset page: https://huggingface.co/datasets/jed351/Chinese-Common-Crawl-Filtered.text10M<n<100M18 likes563 downloads1y agoHugging Face23commoncrawl /gneissweb-annotation-host-testing-v1 GneissWeb Annotations GneissWeb Annotations, powered by IBM Research's GneissWeb methodology, is a dataset of quality and category annotations applied to the Common Crawl corpus. This dataset enables precise filtering of web content across medical, educational, technology, and scientific domains, making it easier to build high-quality corpora for research projects, language models, and specialized applications. Learn more about the annotation process and methodology in our… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/gneissweb-annotation-host-testing-v1.tabular100M<n<1B1 likes510 downloads10mo agoHugging Face24metehan777 /common-crawl-2026-21 Common Crawl SEO & AEO/GEO Dataset — CC-MAIN-2026-21 Web pages from the May 2026 Common Crawl (CC-MAIN-2026-21) that mention SEO, Answer Engine Optimization (AEO), or Generative Engine Optimization (GEO) — filtered to English content and ranked by Common Crawl Web Graph harmonic centrality and PageRank. 🔎 Interactive Explorer A Gradio app to search, filter, and analyze this dataset: https://huggingface.co/spaces/metehan777/cc-seo-explorer Tabs: SEO raw search ·… See the full description on the dataset page: https://huggingface.co/datasets/metehan777/common-crawl-2026-21.texttext-classification1M<n<10M0 likes497 downloads4mo agoHugging Face25Publicus /common_crawl_pointer_indicestext1M<n<10M0 likes451 downloads7mo agoHugging Face26endomorphosis /common_crawl_meta_indexes0 likes354 downloads8mo agoHugging Face27commoncrawl /web-graph-embeddings Web Graph Embedings from Common Crawl's Host-level Hyperlink Graph Dense 128-dimensional embeddings for 52,913,544 web hosts, learned by link prediction on the Common Crawl host-level hyperlink graph release cc-main-2025-26-nov-dec-jan. Vectors are L2-normalized and served in float16; similarity is cosine (a dot product on the unit vectors). The dataset contains hosts with link degree >= 8 (total in+out degree), a 52.9 M / ~19 % induced subgraph that carries ~97 % of the edges.… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/web-graph-embeddings.textfeature-extraction10M<n<100M5 likes308 downloads28d agoHugging Face28commoncrawl /llms.txt llms.txt files extracted from the Common Crawl corpus This dataset contains llms.txt files extracted from the Common Crawl corpus. Specifically, we extracted response records matching */llms.txt and */llms-full.txt with status 200 and MIME-type text/plain or text/markdown. What is llms.txt? The /llms.txt file: A proposal to standardise on using an /llms.txt file to provide information to help LLMs use a website at inference time. Our own analysis of this dataset is available… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/llms.txt.text100K<n<1M2 likes290 downloads28d agoHugging Face29big-banyan-tree /BBT_CommonCrawl_2020 Context BigBanyanTree is an initiative to empower colleges to set up their data engineering clusters, and drive interest towards data processing and analysis using tools such as Apache Spark. The data provided here is the direct result of this initiative. The data was processed by Gautam and Suchit, under the guidance of Harsh Singhal. Content Each arrow file contains a table with fields extracted from Common Crawl WARC files. The datasets provided are derived from… See the full description on the dataset page: https://huggingface.co/datasets/big-banyan-tree/BBT_CommonCrawl_2020.text10M<n<100M2 likes276 downloads2y agoHugging Face30crawlora-net /dead-web-commoncrawl Dead-Web Common Crawl — a longitudinal host-reachability panel (2018–2026) &nbsp;·&nbsp; Hugging Face &nbsp;·&nbsp; Kaggle &nbsp;·&nbsp; License: CC BY 4.0 An open dataset labeling the reachability of 172,959,928 registered domains across 80 monthly Common Crawl archives (CC-MAIN-2018-05 → CC-MAIN-2026-21), built from Common Crawl's robotstxt subset and calibrated against a live re-probe to separate genuinely-dead domains from those merely blocking crawlers or dropped from the… See the full description on the dataset page: https://huggingface.co/datasets/crawlora-net/dead-web-commoncrawl.tabular100M<n<1B1 likes267 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.