commoncrawl/CommonLID
CommonLID: Language Identification for the Real Web CommonLID is a community-created language identification (LID) benchmark. CommonLID consists of web text manually annotated for the language that it is written in. CommonLID contains annotations for 109 languages, where 78 of those languages have at least 100 lines of data. The number of lines available for each language is provided in Appendix A of the paper. ➡️ CommonLID Leaderboard on Hugging Face 🤗 ➡️ Source Code on… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/CommonLID.
56195
No card is published for this repository, or it could not be fetched from Hugging Face right now.
