Team Ai
Datasetpublic

TalTechNLP/voxlingua107_wds

VoxLingua107 VoxLingua107 is a speech dataset for training spoken language identification models. The dataset consists of short speech segments automatically extracted from YouTube videos and labeled according the language of the video title and description, with some post-processing steps to filter out false positives. VoxLingua107 contains data for 107 languages. The total amount of speech in the training set is 6628 hours. The average amount of data per language is 62… See the full description on the dataset page: https://huggingface.co/datasets/TalTechNLP/voxlingua107_wds.

sourceHugging Facecc-by-4.0updated 1y agoView on Hugging Face
4likes6.9kdownloads
Dataset Card

VoxLingua107

VoxLingua107 is a speech dataset for training spoken language identification models. The dataset consists of short speech segments automatically extracted from YouTube videos and labeled according the language of the video title and description, with some post-processing steps to filter out false positives.

VoxLingua107 contains data for 107 languages. The total amount of speech in the training set is 6628 hours. The average amount of data per language is 62 hours. However, the real amount per language varies a lot. There is also a seperate development set containing 1609 speech segments from 33 languages, validated by at least two volunteers to really contain the given language.

For more information, see the paper Jörgen Valk, Tanel Alumäe. _VoxLingua107: a Dataset for Spoken Language Recognition_. Proc. SLT 2021.

Why

VoxLingua107 can be used for training spoken language recognition models that work well with real-world, varying speech data. You can try a demo system trained on this dataset here.

How

We extracted audio data from YouTube videos that are retrieved using language-specific search phrases (random phrases from Wikipedia of the particular language). If the language of the video title and description matched with the language of the search phrase, the audio in the video was deemed likely to be in that particular language. This allowed to collect large amounts of somewhat noisy data relatively cheaply. Speech/non-speech detection and speaker diarization was used to segment the videos into short sentence-like utterances. A data-driven post-filtering step was applied to remove clips that were very different from other clips in this language's dataset, and thus likely not in the given language. Due to the automatic data collection process, there are still clips in the dataset that are not in the given language or contain non-speech (around 2% overall), especially for some languages (like Welsh).

License and copyright

The VoxLingua107 dataset is distributed under the Creative Commons Attribution 4.0 International License. The copyright remains with the original owners of the video.

We also point out that the distribution of languages, accents, dialects, genders, races and societal factors in this dataset is not representative of the global population. Using this dataset for training and deploying models may thus introduce unintended biases.

Notice and take down policy

Notice: Should you consider that our data contains material that is owned by you and should therefore not be reproduced here, please:

  • —Clearly identify yourself, with detailed contact data such as an address, telephone number or email address at which you can be contacted.
  • —Clearly identify the copyrighted work claimed to be infringed.
  • —Clearly identify the material that is claimed to be infringing and information reasonably sufficient to allow us to locate the material.
  • —Send the request to Tanel Alumäe

Take down: We will comply to legitimate requests by removing the affected sources from the corpus.

Languages and sizes

Language CodeLanguage NameHours
abAbkhazian10
afAfrikaans108
amAmharic81
arArabic59
asAssamese155
azAzerbaijani58
baBashkir58
beBelarusian133
bgBulgarian50
bnBengali55
boTibetan101
brBreton44
bsBosnian105
caCatalan88
cebCebuano6
csCzech67
cyWelsh76
daDanish28
deGerman39
elGreek66
enEnglish49
eoEsperanto10
esSpanish39
etEstonian38
euBasque29
faPersian56
fiFinnish33
foFaroese67
frFrench67
glGalician72
gnGuarani2
guGujarati46
gvManx4
haHausa106
hawHawaiian12
hiHindi81
hrCroatian118
htHaitian96
huHungarian73
hyArmenian69
iaInterlingua3
idIndonesian40
isIcelandic92
itItalian51
iwHebrew96
jaJapanese56
jwJavanese53
kaGeorgian98
kkKazakh78
kmCentral Khmer41
knKannada46
koKorean77
laLatin67
lbLuxembourgish75
lnLingala90
loLao42
ltLithuanian82
lvLatvian42
mgMalagasy109
miMaori34
mkMacedonian112
mlMalayalam47
mnMongolian71
mrMarathi85
msMalay83
mtMaltese66
myBurmese41
neNepali72
nlDutch40
nnNorwegian Nynorsk57
noNorwegian107
ocOccitan15
paPanjabi54
plPolish80
psPushto47
ptPortuguese64
roRomanian65
ruRussian73
saSanskrit15
scoScots3
sdSindhi84
siSinhala67
skSlovak40
slSlovenian121
snShona30
soSomali103
sqAlbanian71
srSerbian50
suSundanese64
svSwedish34
swSwahili64
taTamil51
teTelugu77
tgTajik64
thThai61
tkTurkmen85
tlTagalog93
trTurkish59
ttTatar103
ukUkrainian52
urUrdu42
uzUzbek45
viVietnamese64
warWaray11
yiYiddish46
yoYoruba94
zhMandarin Chinese44

Usage

Although webdataset can be used in a streaming fashion, it is recommended to first make a local copy fo the dataset using git clone.

git lfs install git clone git@hf.co:datasets/TalTechNLP/voxlingua107_wds

Then you can use the Python webdataset library to create an iterable dataset out of it:

import webdataset as wds import random trainfiles = glob.glob("voxlingua107wds/train/*/.tar") # you can also limit the training data to selected languages devfiles = glob.glob("voxlingua107wds/dev/*.tar")

random.shuffle(train_files)

def mapper(sample): # "audio" field represents 16 kHz raw audio return {"audio": sample[0], "lang": sample[1]["lang"]}

# Since each shard contains 500 samples for a single language, it is good to use a reasonably large buffer size to get nicely shuffled samples buffersize = 100000 dataset = wds.WebDataset(trainurls, shardshuffle=2000).shuffle(buffersize, initial=buffersize).decode(wds.torchaudio).totuple("wav","json").map(mapper) devdataset = wds.WebDataset(devurls).decode(wds.torchaudio).totuple("wav","json").map(mapper)

trainiter = iter(dataset) print(next(trainiter)) {'audio': (tensor([[-9.7656e-04, -8.5449e-04, -3.0518e-05, ..., 2.7466e-03, 3.7842e-03, 5.1880e-03]]), 16000), 'lang': 'kn'}

Citing

@inproceedings{valk2021slt,
  title={{VoxLingua107}: a Dataset for Spoken Language Recognition},
  author={J{\"o}rgen Valk and Tanel Alum{\"a}e},
  booktitle={Proc. IEEE SLT Workshop},
  year={2021},
}