Team Ai
Datasetpublicgated

commoncrawl/CommonLID

CommonLID: Language Identification for the Real Web CommonLID is a community-created language identification (LID) benchmark. CommonLID consists of web text manually annotated for the language that it is written in. CommonLID contains annotations for 109 languages, where 78 of those languages have at least 100 lines of data. The number of lines available for each language is provided in Appendix A of the paper. ➡️ CommonLID Leaderboard on Hugging Face 🤗 ➡️ Source Code on… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/CommonLID.

sourceHugging Faceotherupdated 18d agoView on Hugging Face
56likes174downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

Team Ai shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
commoncrawl/CommonLID · Team Ai