Team Ai
20 results

commoncrawl

yvfu /common-crawl-character-counts0 likes41k downloads11mo agoHugging Facemusabg /commoncrawl-tr Dataset Card for "commoncrawl-tr" More Information needed text10M<n<100M4 likes7.8k downloads3y agoHugging Facecommoncrawl /gneissweb-annotation-url-testing-v1 GneissWeb Annotations GneissWeb Annotations, powered by IBM Research's GneissWeb methodology, is a dataset of quality and category annotations applied to the Common Crawl corpus. This dataset enables precise filtering of web content across medical, educational, technology, and scientific domains, making it easier to build high-quality corpora for research projects, language models, and specialized applications. Learn more about the annotation process and methodology in our… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/gneissweb-annotation-url-testing-v1.tabular10B<n<100B0 likes7.8k downloads10mo agoHugging Facecommoncrawl /host-index-testing-v2 Common Crawl Host Index v2 GitHub: https://github.com/commoncrawl/cc-host-index Each crawl, we generate a Host Index, which aggregates information about each web hosted visited during the crawl. The information is aggregated from the Common Crawl columnar index, web graph, and raw crawler logs. Quickstart The dataset is Hive-partitioned on crawl (data/crawl=CC-MAIN-2025-18/*.parquet). Open the whole dataset once, then filter with WHERE crawl = '...': because… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/host-index-testing-v2.tabulartext-generation1B<n<10B0 likes6.4k downloads25d agoHugging Faceagentlans /common-crawl-sample Common Crawl sample A small unofficial random subset of the famous Common Crawl dataset. 60 random segment WET files were downloaded from Common Crawl on 2024-05-12. Lines between 500 and 5000 characters long (inclusive) were kept. Only unique texts were kept. No other filtering. Languages Each text was assigned to one of the language codes using the GCLD3 Python package. The Chinese texts were classified as either simplified, traditional, or Cantonese using the… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/common-crawl-sample.texttext-generation1M<n<10M8 likes5.1k downloads2y agoHugging Facewayu-ai /thai-commoncrawl-index Thai Common Crawl Index (2019–2026) An index of every page Common Crawl detected as Thai across 70 monthly crawls, from January 2019 (CC-MAIN-2019-04) to August 2026 (CC-MAIN-2026-30). 932,874,727 page captures · 450,971,497 unique URLs · 6,997,185 hosts · 6,674,969 domains Each row records where the page lives inside Common Crawl's WARC archives — file name, byte offset, and record length — so you can fetch exactly the pages you want with HTTP range requests, without scanning… See the full description on the dataset page: https://huggingface.co/datasets/wayu-ai/thai-commoncrawl-index.tabular100M<n<1B1 likes4.8k downloads2mo agoHugging Face