Team Ai
Datasetpublicgated

djelia/multilingual-asr

multilingual-asr CoVoST 2 and Common Voice 17.0 Swahili and Hausa, re-packaged under one feature schema so the configs can be concatenated into a single multi-task training mix. Four configs, ~206 hours of distinct audio, 8.44 GB of Parquet. No Bambara. Load from datasets import load_dataset asr = load_dataset("djelia/multilingual-asr", "covost2-transcription", split="train") sw_test = load_dataset("djelia/multilingual-asr", "swahili", split="test")… See the full description on the dataset page: https://huggingface.co/datasets/djelia/multilingual-asr.

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes24downloads
Dataset Card

No card is published for this repository, or it could not be fetched from Hugging Face right now.