Team Ai
Datasetpublic

commoncrawl/statistics

Common Crawl Statistics Number of pages, distribution of top-level domains, crawl overlaps, etc. - basic metrics about Common Crawl Monthly Crawl Archives, for more detailed information and graphs please visit our official statistics page. Here you can find the following statistics files: Charsets The character set or encoding of HTML pages only is identified by Tika's AutoDetectReader. The table shows the percentage how character sets have been used to encode… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/statistics.

sourceHugging Faceupdated 16d agoView on Hugging Face
30likes842downloads
README.md102 linesDownload Raw Back to root
1---2pretty_name: Common Crawl Statistics3 4configs:5- config_name: Charsets6  data_files: "charsets.csv"7- config_name: Duplicates8  data_files: "crawlduplicates.txt"9  sep: \s+10  header: 011  names:12  - id13  - crawl14  - page15  - url16  - digest estim.17  - 1-(urls/pages)18  - 1-(digests/pages)19- config_name: Crawlmetrics20  data_files: "crawlmetrics.csv"21- config_name: Crawl metrics by type22  data_files: "crawlmetricsbytype.csv"23- config_name: Crawl overlaps digest24  data_files: "crawloverlap_digest.csv"25- config_name: Crawl overlaps URL26  data_files: "crawloverlap_url.csv"27- config_name: Crawl Similarity Digest28  data_files: "crawlsimilarity_digest.csv"29- config_name: Crawl Similarity URL30  data_files: "crawlsimilarity_url.csv"31- config_name: Crawl Size32  data_files: "crawlsize.csv"33- config_name: Crawl Size by Type34  data_files: "crawlsizebytype.csv"35- config_name: Domains top 50036  data_files: "domains-top-500.csv"37- config_name: Languages38  data_files: "languages.csv"39- config_name: MIME types detected40  data_files: "mimetypes_detected.csv"41- config_name: MIME Types42  data_files: "mimetypes.csv"43- config_name: Top-level domains44  data_files: "tlds.csv"45---46 47# Common Crawl Statistics48 49Number of pages, distribution of top-level domains, crawl overlaps, etc. - basic metrics about Common Crawl Monthly Crawl Archives, for more detailed information and graphs please visit our [official statistics page](https://commoncrawl.github.io/cc-crawl-statistics/). Here you can find the following statistics files:50 51## Charsets52 53The [character set or encoding](https://en.wikipedia.org/wiki/Character_encoding) of HTML pages only is identified by [Tika](https://tika.apache.org/)'s [AutoDetectReader](https://tika.apache.org/1.25/api/org/apache/tika/detect/AutoDetectReader.html). The table shows the percentage how character sets have been used to encode HTML pages crawled by the latest monthly crawls.54 55## Crawl Metrics56 57Crawler-related metrics are extracted from the crawler log files and include58 59- the size of the URL database (CrawlDb)60- the fetch list size (number of URLs scheduled for fetching)61- the response status of the fetch:62  - success63  - redirect64  - denied (forbidden by HTTP 403 or robots.txt)65  - failed (404, host not found, etc.)66- usage of http/https URL protocols (schemes)67 68## Crawl Overlaps69 70Overlaps between monthly crawl archives are calculated and plotted as [Jaccard similarity](https://en.wikipedia.org/wiki/Jaccard_index) of unique URLs or content digests. The cardinality of the monthly crawls and the union of two crawls are [Hyperloglog](https://en.wikipedia.org/wiki/HyperLogLog) estimates.71 72Note, that the content overlaps are small and in the same order of magnitude as the 1% error rate of the Hyperloglog cardinality estimates.73 74## Crawl Size75 76The number of released pages per month fluctuates varies over time due to changes to the number of available seeds, scheduling policy for page revists and crawler operating issues. Because of duplicates the numbers of unique URLs or unique content digests (here Hyperloglog estimates) are lower than the number of page captures.77 78The size on various aggregation levels (host, domain, top-level domain / public suffix) is shown in the next plot. Note that the scale differs per level of aggregation, see the exponential notation behind the labels.79 80## Domains Top 50081 82The shows the top 500 registered domains (in terms of page captures) of the last main/monthly crawl.83 84Note that the ranking by page captures only partially corresponds to the importance of domains, as the crawler respects the robots.txt and tries hard not to overload web servers. Highly ranked domains tend to be underrepresented. If you're looking for a list of domain or host names ranked by page rank or harmonic centrality, consider using one of the [webgraph datasets](https://github.com/commoncrawl/cc-webgraph#exploring-webgraph-data-sets) instead.85 86## Languages87 88The language of a document is identified by [Compact Language Detector 2 (CLD2)](https://github.com/CLD2Owners/cld2). It is able to identify 160 different languages and up to 3 languages per document. The table lists the percentage covered by the primary language of a document (returned first by CLD2). So far, only HTML pages are passed to the language detector.89 90## MIME Types91 92The crawled content is dominated by HTML pages and contains only a small percentage of other document formats. The tables show the percentage of the top 100 media or MIME types of the latest monthly crawls.93 94While the first table is based the `Content-Type` HTTP header, the second uses the MIME type detected by [Apache Tika](https://tika.apache.org/) based on the actual content.95 96## Top-level Domains97 98[Top-level domains](https://en.wikipedia.org/wiki/Top-level_domain) (abbrev. "TLD"/"TLDs") are a significant indicator for the representativeness of the data, whether the data set or particular crawl is biased towards certain countries, regions or languages.99 100Note, that top-level domain is defined here as the left-most element of a host name (`com` in `www.example.com`). [Country-code second-level domains](https://en.wikipedia.org/wiki/Second-level_domain#Country-code_second-level_domains) ("ccSLD") and [public suffixes](https://en.wikipedia.org/wiki/Public_Suffix_List) are not covered by this metrics.101 102