Team Ai
Datasetpublic

OpenTransformer/web-crawl-2026

Web Crawl 2026 A large-scale web crawl dataset for language model pretraining, collected by the OpenTransformer project. Dataset Description This dataset contains text extracted from web pages crawled directly from the internet using custom high-throughput crawlers. All data is freshly scraped. Data Format Each record is a JSON line (gzipped) with fields: text: extracted text content (200-200,000 chars) url: source URL domain: source domain… See the full description on the dataset page: https://huggingface.co/datasets/OpenTransformer/web-crawl-2026.

sourceHugging Faceapache-2.0updated 6mo agoView on Hugging Face
1likes7.8kdownloads

Nothing at this path on main. The folder may be empty, or the revision may not exist.

OpenTransformer/web-crawl-2026 · main · files are served by the source, never re-hosted here