Team Ai
Datasetpublic

agentlans/common-crawl-sample

Common Crawl sample A small unofficial random subset of the famous Common Crawl dataset. 60 random segment WET files were downloaded from Common Crawl on 2024-05-12. Lines between 500 and 5000 characters long (inclusive) were kept. Only unique texts were kept. No other filtering. Languages Each text was assigned to one of the language codes using the GCLD3 Python package. The Chinese texts were classified as either simplified, traditional, or Cantonese using… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/common-crawl-sample.

sourceHugging Faceupdated 2y agoView on Hugging Face
8likes6.6kdownloads
settings

This repository belongs to agentlans on Hugging Face.

Team Ai never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namecommon-crawl-sample
visibilitypublic
licencenot set
gatedno
owneragentlans
Account settings
agentlans/common-crawl-sample · Team Ai