Team Ai
Datasetpublic

LEMAS-Project/LEMAS-Dataset-train

Overview This dataset is part of LEMAS-Project (lemas-project.github.io/LEMAS-Project). It contains a large-scale training set (150k+ hours) and a curated evaluation set (500 utterances per language) covering 10 languages, all with word-level alignment. Fields key: unique utterance identifier; the first two characters indicate the language ID audio: relative path to the MP3 audio file (in the eval set, this key is renamed to "file_name" for compatibility with the… See the full description on the dataset page: https://huggingface.co/datasets/LEMAS-Project/LEMAS-Dataset-train.

sourceHugging Facecc-by-4.0updated 6mo agoView on Hugging Face
90likes6.5kdownloads
../
filees000.tar.gz21.53 GBdownload
filees001.tar.gz21.53 GBdownload
filees002.tar.gz21.53 GBdownload
filees003.tar.gz21.53 GBdownload
filees004.tar.gz21.53 GBdownload
filees005.tar.gz21.53 GBdownload
filees006.tar.gz21.53 GBdownload
filees007.tar.gz21.53 GBdownload
filees008.tar.gz21.53 GBdownload
filees009.tar.gz21.53 GBdownload
filees010.tar.gz21.53 GBdownload
filees011.tar.gz21.53 GBdownload
filees012.tar.gz21.53 GBdownload
filees013.tar.gz21.53 GBdownload
filees014.tar.gz21.53 GBdownload
filees015.tar.gz21.53 GBdownload
filees016.tar.gz21.53 GBdownload
filees017.tar.gz21.53 GBdownload
filees018.tar.gz21.53 GBdownload
filees019.tar.gz21.53 GBdownload
filees020.tar.gz21.54 GBdownload
filees021.tar.gz4.81 GBdownload

LEMAS-Project/LEMAS-Dataset-train · main · files are served by the source, never re-hosted here