Team Ai
Datasetpublic

OpenTransformer/web-crawl-2026

Web Crawl 2026 A large-scale web crawl dataset for language model pretraining, collected by the OpenTransformer project. Dataset Description This dataset contains text extracted from web pages crawled directly from the internet using custom high-throughput crawlers. All data is freshly scraped. Data Format Each record is a JSON line (gzipped) with fields: text: extracted text content (200-200,000 chars) url: source URL domain: source domain… See the full description on the dataset page: https://huggingface.co/datasets/OpenTransformer/web-crawl-2026.

sourceHugging Faceapache-2.0updated 6mo agoView on Hugging Face
1likes7.8kdownloads
settings

This repository belongs to OpenTransformer on Hugging Face.

Team Ai never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameweb-crawl-2026
visibilitypublic
licenceapache-2.0
gatedno
ownerOpenTransformer
Account settings
OpenTransformer/web-crawl-2026 · Team Ai