Team Ai
Datasetpublic

AAdonis/multilingual_audio_alignments

Multilingual MFA-Aligned Speech Dataset A large-scale multilingual speech dataset with word-level and phoneme-level alignments produced using the Montreal Forced Aligner (MFA). Dataset Description This dataset consolidates multiple speech corpora across various languages, all processed through MFA to provide precise phoneme and word alignments. Each sample includes the original audio, transcript, and detailed timing information for both words and phonemes.… See the full description on the dataset page: https://huggingface.co/datasets/AAdonis/multilingual_audio_alignments.

sourceHugging Facecc-by-4.0updated 5mo agoView on Hugging Face
27likes615downloads
1 commits on main
f79258b5mo ago

squash: remove stale deleted-data history

AAdonis