Team Ai
Datasetpublic

allenai/c4

C4 Dataset Summary A colossal, cleaned version of Common Crawl's web crawl corpus. Based on Common Crawl dataset: "https://commoncrawl.org". This is the processed version of Google's C4 dataset We prepared five variants of the data: en, en.noclean, en.noblocklist, realnewslike, and multilingual (mC4). For reference, these are the sizes of the variants: en: 305GB en.noclean: 2.3TB en.noblocklist: 380GB realnewslike: 15GB multilingual (mC4): 9.7TB (108 subsets, one… See the full description on the dataset page: https://huggingface.co/datasets/allenai/c4.

sourceHugging Faceodc-byupdated 3y agoView on Hugging Face
675likes999kdownloads
18 commits on main
1588ec43y ago

🤗 Datasets support (#6)

dirkgr, lhoestq
607bd4c5y ago

Update README.md

Sasha Luccioni
f3b95a15y ago

Merge branch 'main' of https://huggingface.co/datasets/allenai/c4 into main

dirkgr
34c61cd5y ago

Updated readme

dirkgr
e42b2d45y ago

Updated the description to refer to "four" datasets

dirkgr
a8922de5y ago

Added docs for the `noblocklist` data.

dirkgr
1ddc9175y ago

Adds the multilingual set

dirkgr
28d473d5y ago

We can't say badwords

dirkgr
44979765y ago

Adds the en.withbadwords dataset

dirkgr
ba9859b6y ago

Updated readme

dirkgr
f888b0f6y ago

Copy&Paste error in the README

dirkgr
31d802a6y ago

Actually adds a README

dirkgr
75d87b36y ago

Create README.md

dirkgr
aa4cb216y ago

Fix line endings between records

dirkgr
f998d2c6y ago

Adds all the datafiles

dirkgr
1bf6ba16y ago

Adds one datafile to check whether .gitattributes works

dirkgr
456f8a76y ago

Put gz files into LFS

dirkgr
22a1da26y ago

initial commit

system