Team Ai
Datasetpublic

OpenTransformer/web-crawl-2026

Web Crawl 2026 A large-scale web crawl dataset for language model pretraining, collected by the OpenTransformer project. Dataset Description This dataset contains text extracted from web pages crawled directly from the internet using custom high-throughput crawlers. All data is freshly scraped. Data Format Each record is a JSON line (gzipped) with fields: text: extracted text content (200-200,000 chars) url: source URL domain: source domain… See the full description on the dataset page: https://huggingface.co/datasets/OpenTransformer/web-crawl-2026.

sourceHugging Faceapache-2.0updated 6mo agoView on Hugging Face
1likes7.8kdownloads
filechunk1_10GB.jsonl.gz10.49 GBdownload
filechunk1_11GB.jsonl.gz10.54 GBdownload
filechunk1_18GB.jsonl.gz17.80 GBdownload
filechunk1_5.6GB.jsonl.gz6.44 GBdownload
filechunk2_21GB.jsonl.gz21.19 GBdownload
filechunk2_22GB.jsonl.gz21.67 GBdownload
fileclean_v2_2GB.jsonl.gz2.08 GBdownload
fileclean_v2_chunk1_3GB.jsonl.gz3.05 GBdownload

OpenTransformer/web-crawl-2026 · main · files are served by the source, never re-hosted here