langid
Datasets
All datasets matching “langid”nordic_langidAutomatic language identification is a challenging problem. Discriminating
between closely related languages is especially difficult. This paper presents
a machine learning approach for automatic language identification for the
Nordic languages, which often suffer miscategorisation by existing
state-of-the-art tools. Concretely we will focus on discrimination between six
Nordic languages: Danish, Swedish, Norwegian (Nynorsk), Norwegian (Bokmål),
Faroese and Icelandic.
This is the data for the tasks. Two variants are provided: 10K and 50K, with
holding 10,000 and 50,000 examples for each language respectively.lti_langid_corpusThe LTI LangID corpus is a dataset for language identification.
The most recent version, v5, contains training data for 1266 languages, and some (possibly very tiny) amount of text for a total of 1706 languages.
This dataloader can only be executed in a BASH environment at the moment. (See https://github.com/SEACrowd/seacrowd-datahub/pull/405)sinhala-script-lid
Sinhala-script language identification: Sinhala, Pali, Sanskrit
Sentence-level instances of three languages written in Sinhala script, with a
leakage-free document-blocked split. Labels: sin_Sinh, pli_Sinh, san_Sinh.
Rows
split
sin_Sinh
pli_Sinh
san_Sinh
total
train
83105
58309
10565
151979
validation
10345
7317
1337
18999
test
10567
7105
1310
18982
Construction
Built by scripts/maintainer/resplit_target.py (config: target_split… See the full description on the dataset page: https://huggingface.co/datasets/script-langid/sinhala-script-lid.scandi-langid
Dataset Card for "scandi-langid"
More Information needed
sindhi-corpus-langid-cleanlang_ident
