Team Ai
20 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01commoncrawl /gneissweb-annotation-url-testing-v1 GneissWeb Annotations GneissWeb Annotations, powered by IBM Research's GneissWeb methodology, is a dataset of quality and category annotations applied to the Common Crawl corpus. This dataset enables precise filtering of web content across medical, educational, technology, and scientific domains, making it easier to build high-quality corpora for research projects, language models, and specialized applications. Learn more about the annotation process and methodology in our… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/gneissweb-annotation-url-testing-v1.tabular10B<n<100B0 likes7.8k downloads10mo agoHugging Face02commoncrawl /host-index-testing-v2 Common Crawl Host Index v2 GitHub: https://github.com/commoncrawl/cc-host-index Each crawl, we generate a Host Index, which aggregates information about each web hosted visited during the crawl. The information is aggregated from the Common Crawl columnar index, web graph, and raw crawler logs. Quickstart The dataset is Hive-partitioned on crawl (data/crawl=CC-MAIN-2025-18/*.parquet). Open the whole dataset once, then filter with WHERE crawl = '...': because… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/host-index-testing-v2.tabulartext-generation1B<n<10B0 likes6.4k downloads25d agoHugging Face03wayu-ai /thai-commoncrawl-index Thai Common Crawl Index (2019–2026) An index of every page Common Crawl detected as Thai across 70 monthly crawls, from January 2019 (CC-MAIN-2019-04) to August 2026 (CC-MAIN-2026-30). 932,874,727 page captures · 450,971,497 unique URLs · 6,997,185 hosts · 6,674,969 domains Each row records where the page lives inside Common Crawl's WARC archives — file name, byte offset, and record length — so you can fetch exactly the pages you want with HTTP range requests, without scanning… See the full description on the dataset page: https://huggingface.co/datasets/wayu-ai/thai-commoncrawl-index.tabular100M<n<1B1 likes4.8k downloads2mo agoHugging Face04commoncrawl /statistics Common Crawl Statistics Number of pages, distribution of top-level domains, crawl overlaps, etc. - basic metrics about Common Crawl Monthly Crawl Archives, for more detailed information and graphs please visit our official statistics page. Here you can find the following statistics files: Charsets The character set or encoding of HTML pages only is identified by Tika's AutoDetectReader. The table shows the percentage how character sets have been used to encode… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/statistics.tabular100K<n<1M30 likes861 downloads21d agoHugging Face05PhysiQuanty /FRENCH-ONLY-Common-Crawl-2026-25tabular1M<n<10M3 likes849 downloads4mo agoHugging Face06commoncrawl /gneissweb-annotation-host-testing-v1 GneissWeb Annotations GneissWeb Annotations, powered by IBM Research's GneissWeb methodology, is a dataset of quality and category annotations applied to the Common Crawl corpus. This dataset enables precise filtering of web content across medical, educational, technology, and scientific domains, making it easier to build high-quality corpora for research projects, language models, and specialized applications. Learn more about the annotation process and methodology in our… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/gneissweb-annotation-host-testing-v1.tabular100M<n<1B1 likes401 downloads10mo agoHugging Face07crawlora-net /dead-web-commoncrawl Dead-Web Common Crawl — a longitudinal host-reachability panel (2018–2026) &nbsp;·&nbsp; Hugging Face &nbsp;·&nbsp; Kaggle &nbsp;·&nbsp; License: CC BY 4.0 An open dataset labeling the reachability of 172,959,928 registered domains across 80 monthly Common Crawl archives (CC-MAIN-2018-05 → CC-MAIN-2026-21), built from Common Crawl's robotstxt subset and calibrated against a live re-probe to separate genuinely-dead domains from those merely blocking crawlers or dropped from the… See the full description on the dataset page: https://huggingface.co/datasets/crawlora-net/dead-web-commoncrawl.tabular100M<n<1B1 likes235 downloads3mo agoHugging Face08anton-l /commoncrawl_tex Dataset Card for "commoncrawl_tex" More Information needed tabular100K<n<1M1 likes180 downloads3y agoHugging Face09jtl11 /commoncrawl-feb-2025tabular1K<n<10K0 likes147 downloads2y agoHugging Face10Publicus /common_crawl_meta_indexestabular1B<n<10B0 likes105 downloads8mo agoHugging Face11sc2qa /sc2qa_commoncrawl\tabular10K<n<100K0 likes102 downloads5y agoHugging Face12Tinuade /common-crawl-docx-sample Common Crawl DOCX Sample A sample of normalized text extracted from DOCX records in Common Crawl. Source Common Crawl release: CC-MAIN-YYYY-NN Source index: Common Crawl URL Index Pipeline: marin-community/marin Pipeline revision: REPLACE_WITH_GIT_SHA Records were selected using declared DOCX MIME type, detected DOCX MIME type, or a .docx URL suffix. Only successful, non-truncated index records were eligible. Processing The pipeline:… See the full description on the dataset page: https://huggingface.co/datasets/Tinuade/common-crawl-docx-sample.tabulartext-generation1K<n<10K0 likes71 downloads24d agoHugging Face13davanstrien /commoncrawl-jobs-demo Common Crawl on Jobs — datatrove JobsPipelineExecutor demo 228,088 English web documents (~820 MB compressed, 60 jsonl.gz shards) extracted from 1,237,374 Common Crawl pages — the output of a test run of datatrove's experimental JobsPipelineExecutor, which fans a datatrove pipeline out across a pool of Hugging Face Jobs instead of a Slurm cluster. This is a pipeline demo artifact, not a curated corpus: one segment slice of one crawl, shared as the verifiable receipt for the… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/commoncrawl-jobs-demo.tabular100K<n<1M11 likes65 downloads3mo agoHugging Face14endomorphosis /common_crawl_state_indextabular100M<n<1B0 likes31 downloads8mo agoHugging Face15commoncrawl /eot2024_hostlevel_logsgatedThis dataset is a host-level summary of the initial crawl logs for the End of Term 2024 dataset. Since this project will not finish until January 2025, please do not ask for access unless you are directly involved in this effort. Organizations involved are the Library of Congress, the Internet Archive, the University of North Texas Libraries, Stanford University Libraries, the US Government Publishing Office, the US National Archives, and the Common Crawl Foundation. tabular100K<n<1M1 likes22 downloads2y agoHugging Face16endomorphosis /common_crawl_municipal_indextabular10M<n<100M0 likes22 downloads8mo agoHugging Face17amazingvince /common-crawl-diverse-sampletabular10K<n<100K0 likes20 downloads2y agoHugging Face18endomorphosis /common_crawl_federal_indextabular100M<n<1B0 likes14 downloads8mo agoHugging Face19yongxin2020 /tdf-olmo3-common_crawltabular1K<n<10K0 likes11 downloads5mo agoHugging Face20amazingvince /common_crawl_diverse_sample-no-complexity-filtertabular10K<n<100K0 likes5 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.