Team Ai
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01commoncrawl /host-index-testing-v2 Common Crawl Host Index v2 GitHub: https://github.com/commoncrawl/cc-host-index Each crawl, we generate a Host Index, which aggregates information about each web hosted visited during the crawl. The information is aggregated from the Common Crawl columnar index, web graph, and raw crawler logs. Quickstart The dataset is Hive-partitioned on crawl (data/crawl=CC-MAIN-2025-18/*.parquet). Open the whole dataset once, then filter with WHERE crawl = '...': because… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/host-index-testing-v2.tabulartext-generation1B<n<10B0 likes6.4k downloads26d agoHugging Face02agentlans /common-crawl-sample Common Crawl sample A small unofficial random subset of the famous Common Crawl dataset. 60 random segment WET files were downloaded from Common Crawl on 2024-05-12. Lines between 500 and 5000 characters long (inclusive) were kept. Only unique texts were kept. No other filtering. Languages Each text was assigned to one of the language codes using the GCLD3 Python package. The Chinese texts were classified as either simplified, traditional, or Cantonese using the… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/common-crawl-sample.texttext-generation1M<n<10M8 likes5.1k downloads2y agoHugging Face03BramVanroy /CommonCrawl-CreativeCommons The Common Crawl Creative Commons Corpus (C5) Raw CommonCrawl crawls, annotated with Creative Commons license information C5 is an effort to collect Creative Commons-licensed web data in one place. The licensing information is extracted from the web pages based on whether they link to Creative Commons licenses either overtly in a tags (like in the footer of Wikipedia) or in metadata fields indicating deliberate Creative Commons publication. However, false positives may occur! See… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/CommonCrawl-CreativeCommons.texttext-generation100M<n<1B42 likes2.5k downloads1y agoHugging Face04BramVanroy /CommonCrawl-CreativeCommons-fine Common Crawl Creative Commons Corpus Fine (C5f) A filtered version of the Common Crawl Creative Commons Corpus (C5), only retaining samples that are also present in the FineWeb or FineWeb-2 datasets. As such, this dataset contains a high-quality subset of C5. Created with this script. For more information, see C5. Progress In the v1 release, the following crawls are included CC-MAIN-2019-30 CC-MAIN-2020-05CC-MAIN-2023-06 CC-MAIN-2024-51 CC-MAIN-2024-46 CC-MAIN-2025-05… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/CommonCrawl-CreativeCommons-fine.texttext-generation10M<n<100M5 likes911 downloads1y agoHugging Face05BramVanroy /CommonCrawl-CreativeCommons-strict Common Crawl Creative Commons Corpus Strict (C5s) A filtered version of the Common Crawl Creative Commons Corpus (C5), only retaining samples that: are also present in the FineWeb or FineWeb-2 datasets; have no license disagreement (all found licenses have the same type; version number might differ); are not "non-commercial" ("nc" in license); are not "cc-unknown"; do not have "wiki" in their name (the idea is that you should include Wikipedia and other Wikidata from other… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/CommonCrawl-CreativeCommons-strict.texttext-generation10M<n<100M2 likes794 downloads1y agoHugging Face06anandjh8 /common-crawl-english-filtered 🧠 FineWeb-English-Filtered 📘 Dataset Summary FineWeb-English-Filtered is a large-scale, cleaned, English-only text dataset derived from Common Crawl’s WET archives.It contains 940 million documents of publicly available web text, converted into Apache Parquet format with a consistent schema for fast and efficient data loading. The dataset was generated using a custom AWS Glue pipeline that processed, filtered, and merged .wet files across multiple terabytes of… See the full description on the dataset page: https://huggingface.co/datasets/anandjh8/common-crawl-english-filtered.texttext-generation100M<n<1B2 likes222 downloads4mo agoHugging Face07Zarinaaa /commoncrawl_dataset Kyrgyz CommonCrawl Dataset A 271 MB text corpus of Kyrgyz language data extracted from CommonCrawl — one of the largest openly available Kyrgyz text collections for NLP research. Dataset Description This dataset contains Kyrgyz-language web text scraped from CommonCrawl archives, filtered by the Kyrgyz language tag (ky). The data covers a wide range of domains including news, blogs, government sites, educational content, and general web pages. Why this matters: Kyrgyz is… See the full description on the dataset page: https://huggingface.co/datasets/Zarinaaa/commoncrawl_dataset.text-generation100M<n<1B0 likes221 downloads8mo agoHugging Face08liswei /common-crawl-zhtw Dataset Card for Common Crawl Traditional Chinese De-duplicated version of jed351/Traditional-Chinese-Common-Crawl-Filtered. De-duplicated with MinHash Is suggested to filter the dataset with NLU models before any serious use. texttext-generation1M<n<10M6 likes173 downloads2y agoHugging Face09Tinuade /common-crawl-docx-sample Common Crawl DOCX Sample A sample of normalized text extracted from DOCX records in Common Crawl. Source Common Crawl release: CC-MAIN-YYYY-NN Source index: Common Crawl URL Index Pipeline: marin-community/marin Pipeline revision: REPLACE_WITH_GIT_SHA Records were selected using declared DOCX MIME type, detected DOCX MIME type, or a .docx URL suffix. Only successful, non-truncated index records were eligible. Processing The pipeline:… See the full description on the dataset page: https://huggingface.co/datasets/Tinuade/common-crawl-docx-sample.tabulartext-generation1K<n<10K0 likes71 downloads24d agoHugging Face10kuluruvineeth /common-crawl-20-warcs-en-te Common Crawl, 20 WARCs: what an English + Telugu pipeline keeps and removes 430,279 web records from Common Crawl CC-MAIN-2026-39 went through extraction, URL filtering and language identification. 139,643 English or Telugu pages were kept, and every page the language filter removed is here too, with the reason, so you can see exactly what a filter throws away. Processed with dataflow on Hugging Face Jobs: 2 cpu-upgrade Jobs, 75 Job-minutes, $0.04. The run's live dashboard:… See the full description on the dataset page: https://huggingface.co/datasets/kuluruvineeth/common-crawl-20-warcs-en-te.texttext-generation100K<n<1M0 likes40 downloads8d agoHugging Face11GeoGPT-Research-Project /GeoGPT_Training_Data_from_Geoscience_Subset_of_CommonCrawl Description This dataset is a geoscience-specific subset of CommonCrawl used for GeoGPT training. CommonCrawl is a free and open repository of web crawl data with over 250 billion web pages and is widely used by leading large language models. We apply data mining algorithms to extract geoscience-related content from this vast dataset. This dataset comprises 12,414,268 samples, each containing the following metadata to trace the data source within CommonCrawl: id (string):… See the full description on the dataset page: https://huggingface.co/datasets/GeoGPT-Research-Project/GeoGPT_Training_Data_from_Geoscience_Subset_of_CommonCrawl.texttext-generation10M<n<100M1 likes34 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.