Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01musabg /commoncrawl-tr Dataset Card for "commoncrawl-tr" More Information needed text10M<n<100M4 likes7.8k downloads3y agoHugging Face02commoncrawl /gneissweb-annotation-url-testing-v1 GneissWeb Annotations GneissWeb Annotations, powered by IBM Research's GneissWeb methodology, is a dataset of quality and category annotations applied to the Common Crawl corpus. This dataset enables precise filtering of web content across medical, educational, technology, and scientific domains, making it easier to build high-quality corpora for research projects, language models, and specialized applications. Learn more about the annotation process and methodology in our… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/gneissweb-annotation-url-testing-v1.tabular10B<n<100B0 likes7.8k downloads10mo agoHugging Face03commoncrawl /host-index-testing-v2 Common Crawl Host Index v2 GitHub: https://github.com/commoncrawl/cc-host-index Each crawl, we generate a Host Index, which aggregates information about each web hosted visited during the crawl. The information is aggregated from the Common Crawl columnar index, web graph, and raw crawler logs. Quickstart The dataset is Hive-partitioned on crawl (data/crawl=CC-MAIN-2025-18/*.parquet). Open the whole dataset once, then filter with WHERE crawl = '...': because… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/host-index-testing-v2.tabulartext-generation1B<n<10B0 likes6.4k downloads25d agoHugging Face04agentlans /common-crawl-sample Common Crawl sample A small unofficial random subset of the famous Common Crawl dataset. 60 random segment WET files were downloaded from Common Crawl on 2024-05-12. Lines between 500 and 5000 characters long (inclusive) were kept. Only unique texts were kept. No other filtering. Languages Each text was assigned to one of the language codes using the GCLD3 Python package. The Chinese texts were classified as either simplified, traditional, or Cantonese using the… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/common-crawl-sample.texttext-generation1M<n<10M8 likes5.1k downloads2y agoHugging Face05wayu-ai /thai-commoncrawl-index Thai Common Crawl Index (2019–2026) An index of every page Common Crawl detected as Thai across 70 monthly crawls, from January 2019 (CC-MAIN-2019-04) to August 2026 (CC-MAIN-2026-30). 932,874,727 page captures · 450,971,497 unique URLs · 6,997,185 hosts · 6,674,969 domains Each row records where the page lives inside Common Crawl's WARC archives — file name, byte offset, and record length — so you can fetch exactly the pages you want with HTTP range requests, without scanning… See the full description on the dataset page: https://huggingface.co/datasets/wayu-ai/thai-commoncrawl-index.tabular100M<n<1B1 likes4.8k downloads2mo agoHugging Face06sailor2 /sea-commoncrawltext100M<n<1B1 likes4k downloads2y agoHugging Face07BramVanroy /CommonCrawl-CreativeCommons The Common Crawl Creative Commons Corpus (C5) Raw CommonCrawl crawls, annotated with Creative Commons license information C5 is an effort to collect Creative Commons-licensed web data in one place. The licensing information is extracted from the web pages based on whether they link to Creative Commons licenses either overtly in a tags (like in the footer of Wikipedia) or in metadata fields indicating deliberate Creative Commons publication. However, false positives may occur! See… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/CommonCrawl-CreativeCommons.texttext-generation100M<n<1B42 likes2.5k downloads1y agoHugging Face08sailor2 /sea-commoncrawl-high-qualitytext10M<n<100M0 likes984 downloads2y agoHugging Face09BramVanroy /CommonCrawl-CreativeCommons-fine Common Crawl Creative Commons Corpus Fine (C5f) A filtered version of the Common Crawl Creative Commons Corpus (C5), only retaining samples that are also present in the FineWeb or FineWeb-2 datasets. As such, this dataset contains a high-quality subset of C5. Created with this script. For more information, see C5. Progress In the v1 release, the following crawls are included CC-MAIN-2019-30 CC-MAIN-2020-05CC-MAIN-2023-06 CC-MAIN-2024-51 CC-MAIN-2024-46 CC-MAIN-2025-05… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/CommonCrawl-CreativeCommons-fine.texttext-generation10M<n<100M5 likes911 downloads1y agoHugging Face10commoncrawl /statistics Common Crawl Statistics Number of pages, distribution of top-level domains, crawl overlaps, etc. - basic metrics about Common Crawl Monthly Crawl Archives, for more detailed information and graphs please visit our official statistics page. Here you can find the following statistics files: Charsets The character set or encoding of HTML pages only is identified by Tika's AutoDetectReader. The table shows the percentage how character sets have been used to encode… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/statistics.tabular100K<n<1M30 likes861 downloads21d agoHugging Face11BramVanroy /CommonCrawl-CreativeCommons-strict Common Crawl Creative Commons Corpus Strict (C5s) A filtered version of the Common Crawl Creative Commons Corpus (C5), only retaining samples that: are also present in the FineWeb or FineWeb-2 datasets; have no license disagreement (all found licenses have the same type; version number might differ); are not "non-commercial" ("nc" in license); are not "cc-unknown"; do not have "wiki" in their name (the idea is that you should include Wikipedia and other Wikidata from other… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/CommonCrawl-CreativeCommons-strict.texttext-generation10M<n<100M2 likes794 downloads1y agoHugging Face12hatakeyama-llm-team /CommonCrawl_wet_v2text100K<n<1M1 likes524 downloads3y agoHugging Face13metehan777 /common-crawl-2026-21 Common Crawl SEO & AEO/GEO Dataset — CC-MAIN-2026-21 Web pages from the May 2026 Common Crawl (CC-MAIN-2026-21) that mention SEO, Answer Engine Optimization (AEO), or Generative Engine Optimization (GEO) — filtered to English content and ranked by Common Crawl Web Graph harmonic centrality and PageRank. 🔎 Interactive Explorer A Gradio app to search, filter, and analyze this dataset: https://huggingface.co/spaces/metehan777/cc-seo-explorer Tabs: SEO raw search ·… See the full description on the dataset page: https://huggingface.co/datasets/metehan777/common-crawl-2026-21.texttext-classification1M<n<10M0 likes506 downloads4mo agoHugging Face14commoncrawl /gneissweb-annotation-host-testing-v1 GneissWeb Annotations GneissWeb Annotations, powered by IBM Research's GneissWeb methodology, is a dataset of quality and category annotations applied to the Common Crawl corpus. This dataset enables precise filtering of web content across medical, educational, technology, and scientific domains, making it easier to build high-quality corpora for research projects, language models, and specialized applications. Learn more about the annotation process and methodology in our… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/gneissweb-annotation-host-testing-v1.tabular100M<n<1B1 likes401 downloads10mo agoHugging Face15Publicus /common_crawl_pointer_indicestext1M<n<10M0 likes319 downloads7mo agoHugging Face16commoncrawl /llms.txt llms.txt files extracted from the Common Crawl corpus This dataset contains llms.txt files extracted from the Common Crawl corpus. Specifically, we extracted response records matching */llms.txt and */llms-full.txt with status 200 and MIME-type text/plain or text/markdown. What is llms.txt? The /llms.txt file: A proposal to standardise on using an /llms.txt file to provide information to help LLMs use a website at inference time. Our own analysis of this dataset is available… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/llms.txt.text100K<n<1M2 likes290 downloads1mo agoHugging Face17qikp /commoncrawl-2022-33text1M<n<10M0 likes251 downloads5mo agoHugging Face18commoncrawl /web-graph-embeddings Web Graph Embedings from Common Crawl's Host-level Hyperlink Graph Dense 128-dimensional embeddings for 52,913,544 web hosts, learned by link prediction on the Common Crawl host-level hyperlink graph release cc-main-2025-26-nov-dec-jan. Vectors are L2-normalized and served in float16; similarity is cosine (a dot product on the unit vectors). The dataset contains hosts with link degree >= 8 (total in+out degree), a 52.9 M / ~19 % induced subgraph that carries ~97 % of the edges.… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/web-graph-embeddings.textfeature-extraction10M<n<100M5 likes249 downloads1mo agoHugging Face19crawlora-net /dead-web-commoncrawl Dead-Web Common Crawl — a longitudinal host-reachability panel (2018–2026) &nbsp;·&nbsp; Hugging Face &nbsp;·&nbsp; Kaggle &nbsp;·&nbsp; License: CC BY 4.0 An open dataset labeling the reachability of 172,959,928 registered domains across 80 monthly Common Crawl archives (CC-MAIN-2018-05 → CC-MAIN-2026-21), built from Common Crawl's robotstxt subset and calibrated against a live re-probe to separate genuinely-dead domains from those merely blocking crawlers or dropped from the… See the full description on the dataset page: https://huggingface.co/datasets/crawlora-net/dead-web-commoncrawl.tabular100M<n<1B1 likes235 downloads3mo agoHugging Face20anandjh8 /common-crawl-english-filtered 🧠 FineWeb-English-Filtered 📘 Dataset Summary FineWeb-English-Filtered is a large-scale, cleaned, English-only text dataset derived from Common Crawl’s WET archives.It contains 940 million documents of publicly available web text, converted into Apache Parquet format with a consistent schema for fast and efficient data loading. The dataset was generated using a custom AWS Glue pipeline that processed, filtered, and merged .wet files across multiple terabytes of… See the full description on the dataset page: https://huggingface.co/datasets/anandjh8/common-crawl-english-filtered.texttext-generation100M<n<1B2 likes222 downloads4mo agoHugging Face21commoncrawl /citations Common Crawl Citations Overview This dataset contains citations referencing Common Crawl Foundation and its datasets, pulled from Google Scholar. Please note that these citations are not curated, so they will include some false positives. An annotated subset of these citations with additional fields can be found at cc-citations. text1K<n<10K5 likes207 downloads6mo agoHugging Face22big-banyan-tree /BBT_CommonCrawl_2023 Context BigBanyanTree is an initiative to empower colleges to set up their data engineering clusters, and drive interest towards data processing and analysis using tools such as Apache Spark. The data provided here is the direct result of this initiative. The data was processed by Gautam and Suchit, under the guidance of Harsh Singhal. Content Each arrow file contains a table with fields extracted from Common Crawl WARC files. The datasets provided are derived from… See the full description on the dataset page: https://huggingface.co/datasets/big-banyan-tree/BBT_CommonCrawl_2023.text10M<n<100M2 likes200 downloads2y agoHugging Face23anton-l /commoncrawl_tex Dataset Card for "commoncrawl_tex" More Information needed tabular100K<n<1M1 likes180 downloads3y agoHugging Face24xerozee /nigeriaai-commoncrawltext1M<n<10M0 likes180 downloads7mo agoHugging Face25commoncrawl /CommonLIDgated CommonLID: Language Identification for the Real Web CommonLID is a community-created language identification (LID) benchmark. CommonLID consists of web text manually annotated for the language that it is written in. CommonLID contains annotations for 109 languages, where 78 of those languages have at least 100 lines of data. The number of lines available for each language is provided in Appendix A of the paper. ➡️ CommonLID Leaderboard on Hugging Face 🤗 ➡️ Source Code on… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/CommonLID.texttext-classification100K<n<1M56 likes174 downloads18d agoHugging Face26liswei /common-crawl-zhtw Dataset Card for Common Crawl Traditional Chinese De-duplicated version of jed351/Traditional-Chinese-Common-Crawl-Filtered. De-duplicated with MinHash Is suggested to filter the dataset with NLU models before any serious use. texttext-generation1M<n<10M6 likes173 downloads2y agoHugging Face27big-banyan-tree /BBT_CommonCrawl_2020 Context BigBanyanTree is an initiative to empower colleges to set up their data engineering clusters, and drive interest towards data processing and analysis using tools such as Apache Spark. The data provided here is the direct result of this initiative. The data was processed by Gautam and Suchit, under the guidance of Harsh Singhal. Content Each arrow file contains a table with fields extracted from Common Crawl WARC files. The datasets provided are derived from… See the full description on the dataset page: https://huggingface.co/datasets/big-banyan-tree/BBT_CommonCrawl_2020.text10M<n<100M2 likes170 downloads2y agoHugging Face28big-banyan-tree /BBT_CommonCrawl_2019 Context BigBanyanTree is an initiative to empower colleges to set up their data engineering clusters, and drive interest towards data processing and analysis using tools such as Apache Spark. The data provided here is the direct result of this initiative. The data was processed by Gautam and Suchit, under the guidance of Harsh Singhal. Content Each arrow file contains a table with fields extracted from Common Crawl WARC files. The datasets provided are derived from… See the full description on the dataset page: https://huggingface.co/datasets/big-banyan-tree/BBT_CommonCrawl_2019.text10M<n<100M2 likes160 downloads2y agoHugging Face29jtl11 /commoncrawl-feb-2025tabular1K<n<10K0 likes147 downloads2y agoHugging Face30big-banyan-tree /BBT_CommonCrawl_2018 Context BigBanyanTree is an initiative to empower colleges to set up their data engineering clusters, and drive interest towards data processing and analysis using tools such as Apache Spark. The data provided here is the direct result of this initiative. The data was processed by Gautam and Suchit, under the guidance of Harsh Singhal. Content Each arrow file contains a table with fields extracted from Common Crawl WARC files. The datasets provided are derived from… See the full description on the dataset page: https://huggingface.co/datasets/big-banyan-tree/BBT_CommonCrawl_2018.text10M<n<100M3 likes140 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.