commoncrawl/CommonLID
CommonLID: Language Identification for the Real Web CommonLID is a community-created language identification (LID) benchmark. CommonLID consists of web text manually annotated for the language that it is written in. CommonLID contains annotations for 109 languages, where 78 of those languages have at least 100 lines of data. The number of lines available for each language is provided in Appendix A of the paper. ➡️ CommonLID Leaderboard on Hugging Face 🤗 ➡️ Source Code on… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/CommonLID.
Conversations for this repository live on Hugging Face.
Team Ai shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face