allenai/c4
C4 Dataset Summary A colossal, cleaned version of Common Crawl's web crawl corpus. Based on Common Crawl dataset: "https://commoncrawl.org". This is the processed version of Google's C4 dataset We prepared five variants of the data: en, en.noclean, en.noblocklist, realnewslike, and multilingual (mC4). For reference, these are the sizes of the variants: en: 305GB en.noclean: 2.3TB en.noblocklist: 380GB realnewslike: 15GB multilingual (mC4): 9.7TB (108 subsets, one… See the full description on the dataset page: https://huggingface.co/datasets/allenai/c4.
🤗 Datasets support (#6)
Update README.md
Merge branch 'main' of https://huggingface.co/datasets/allenai/c4 into main
Updated readme
Updated the description to refer to "four" datasets
Added docs for the `noblocklist` data.
Adds the multilingual set
We can't say badwords
Adds the en.withbadwords dataset
Updated readme
Copy&Paste error in the README
Actually adds a README
Create README.md
Fix line endings between records
Adds all the datafiles
Adds one datafile to check whether .gitattributes works
Put gz files into LFS
initial commit
