commoncrawl/llms.txt-report
1
What is actually inside llms.txt
A content analysis of the 584,107 non-empty llms.txt and llms-full.txt documents in the Common Crawl `commoncrawl/llms.txt` dataset, config CC-MAIN-2026-30 — spec conformance, which tool generated the file, AI-usage policy and named-crawler rules, manipulation and spam, and language, length, topics and ingestion cost.
Every grouped result links five example documents, each to the live URL and to its row in the dataset viewer, so any figure can be traced back to the exact bytes it was computed from.
Generated by `analyze.py` in one streaming pass over the corpus.
