Team Ai
Datasetpublic

Alignment-Lab-AI/Open-Web-Math

Keiran Paster*, Marco Dos Santos*, Zhangir Azerbayev, Jimmy Ba GitHub | ArXiv | PDF OpenWebMath is a dataset containing the majority of the high-quality, mathematical text from the internet. It is filtered and extracted from over 200B HTML files on Common Crawl down to a set of 6.3 million documents containing a total of 14.7B tokens. OpenWebMath is intended for use in pretraining and finetuning large language models. You can download the dataset using Hugging Face: from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/Alignment-Lab-AI/Open-Web-Math.

sourceHugging Faceupdated 2y agoView on Hugging Face
7likes998downloads
README.md111 linesDownload Raw Back to root
1---2dataset_info:3  features:4    - name: url5      dtype: string6    - name: text7      dtype: string8    - name: date9      dtype: string10    - name: metadata11      dtype: string12  splits:13    - name: train14      num_bytes: 5665199505715      num_examples: 631523316  download_size: 1637068992517  dataset_size: 5665199505718  license: odc-by19  task_categories:20    - text-generation21  language:22    - en23  pretty_name: OpenWebMath24  size_categories:25    - 10B<n<100B26---27 28<img src="imgs/OpenWebMath-left.png" width="300">29 30[Keiran Paster](https://keirp.com)\*, [Marco Dos Santos](https://marco-dossantos.github.io/)\*, [Zhangir Azerbayev](https://zhangir-azerbayev.github.io/), [Jimmy Ba](https://jimmylba.github.io/)31 32[GitHub ](https://github.com/keirp/OpenWebMath) | [ArXiv](https://arxiv.org/abs/2310.06786)33| [PDF](https://arxiv.org/pdf/2310.06786.pdf)34 35**OpenWebMath** is a dataset containing the majority of the high-quality, mathematical text from the internet. It is filtered and extracted from over 200B HTML files on Common Crawl down to a set of **6.3 million documents** containing a total of **14.7B tokens**. OpenWebMath is intended for use in _pretraining_ and _finetuning_ large language models.36 37You can download the dataset using Hugging Face:38 39```python40from datasets import load_dataset41ds = load_dataset("open-web-math/open-web-math")42```43 44# OpenWebMath Contents45 46The dataset is structured as follows:47 48```python49{50  "text": ...,  # document text.51  "url": ...,  # document url.52  "date": ...,  # date the page was crawled.53  "metadata": ...,  # JSON containing information from the extraction process.54}55```56 57OpenWebMath contains documents from over 130k different domains, including data from forums, educational pages, and blogs. The dataset contains documents covering mathematics, physics, statistics, computer science, and more. The following table shows the most common domains in OpenWebMath by character count.58 59| Domain            | # Characters  | % Characters |60| ----------------- | ------------- | ------------ |61| stackexchange.com | 4,655,132,784 | 9.55%        |62| nature.com        | 1,529,935,838 | 3.14%        |63| wordpress.com     | 1,294,166,938 | 2.66%        |64| physicsforums.com | 1,160,137,919 | 2.38%        |65| github.io         | 725,689,722   | 1.49%        |66| zbmath.org        | 620,019,503   | 1.27%        |67| wikipedia.org     | 618,024,754   | 1.27%        |68| groundai.com      | 545,214,990   | 1.12%        |69| blogspot.com      | 520,392,333   | 1.07%        |70| mathoverflow.net  | 499,102,560   | 1.02%        |71 72# OpenWebMath Pipeline73 74<img src="imgs/pipeline.png" alt="Overview of the OpenWebMath Pipeline">75 76OpenWebMath builds on the massive [Common Crawl](https://commoncrawl.org/) dataset, which contains over 200B HTML documents. We filtered the data to only include documents that are: (1) in English, (2) contain mathematical content, and (3) are of high quality. We also put a strong emphasis on extracting LaTeX content from the HTML documents as well as reducing boilerplate in comparison to other web datasets.77 78The OpenWebMath pipeline consists of five steps:79 801. **Prefiltering HTML Documents**:81   - We apply a simple prefilter to all HTML documents in Common Crawl in order to skip documents without mathematical content to unnecessary processing time.822. **Text Extraction**:83   - Extract text, including LaTeX content, from the HTML documents while removing boilerplate.843. **Content Classification and Filtering**:85   - Apply a [FastText language identification model](https://fasttext.cc/docs/en/language-identification.html) to keep only English documents.86   - Filter high perplexity documents using a [KenLM](https://github.com/kpu/kenlm) model trained on [Proof-Pile](https://huggingface.co/datasets/hoskinson-center/proof-pile).87   - Filter non-mathematical documents using our own _MathScore_ model.884. **Deduplication**:89   - Deduplicate the dataset using SimHash in [text-dedup](https://github.com/ChenghaoMou/text-dedup).905. **Manual Inspection**:91   - Inspect the documents gathered from previous steps and remove low quality pages.92 93For a detailed discussion on the processing pipeline, please refer to our paper.94 95# License96 97OpenWebMath is made available under an ODC-By 1.0 license; users should also abide by the CommonCrawl ToU: [https://commoncrawl.org/terms-of-use/](https://commoncrawl.org/terms-of-use/). We do not alter the license of any of the underlying data.98 99# Citation Information100 101```102@misc{paster2023openwebmath,103      title={OpenWebMath: An Open Dataset of High-Quality Mathematical Web Text},104      author={Keiran Paster and Marco Dos Santos and Zhangir Azerbayev and Jimmy Ba},105      year={2023},106      eprint={2310.06786},107      archivePrefix={arXiv},108      primaryClass={cs.AI}109}110```111