Team Ai
Datasetpublicgated

commoncrawl/CommonLID

CommonLID: Language Identification for the Real Web CommonLID is a community-created language identification (LID) benchmark. CommonLID consists of web text manually annotated for the language that it is written in. CommonLID contains annotations for 109 languages, where 78 of those languages have at least 100 lines of data. The number of lines available for each language is provided in Appendix A of the paper. ➡️ CommonLID Leaderboard on Hugging Face 🤗 ➡️ Source Code on… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/CommonLID.

sourceHugging Faceotherupdated 14d agoView on Hugging Face
56likes195downloads
Dataset Card

No card is published for this repository, or it could not be fetched from Hugging Face right now.