Team Ai
Datasetpublic

agentlans/common-crawl-sample

Common Crawl sample A small unofficial random subset of the famous Common Crawl dataset. 60 random segment WET files were downloaded from Common Crawl on 2024-05-12. Lines between 500 and 5000 characters long (inclusive) were kept. Only unique texts were kept. No other filtering. Languages Each text was assigned to one of the language codes using the GCLD3 Python package. The Chinese texts were classified as either simplified, traditional, or Cantonese using… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/common-crawl-sample.

sourceHugging Faceupdated 2y agoView on Hugging Face
8likes6.6kdownloads
../
filetest.json.gz1.3 MBdownload
filetrain.json.gz11.9 MBdownload

agentlans/common-crawl-sample · main · files are served by the source, never re-hosted here