datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mms-multilingual-audio-5to30min
Multilingual Audio Dataset (5-30min)
This dataset contains continuous speech audio files for various languages (ranging from 5 to 30 minutes in length per language) collected from diverse sources including Hugging Face and YouTube.
Dataset Statistics
Total Languages: 100
Sources: HF Omnilingual ASR Corpus, YouTube
Language Details
Language Code
Language Name
Source
Duration (seconds)
jpn
Japanese
youtube
1459.84
eng
English
youtube… See the full description on the dataset page: https://huggingface.co/datasets/aoiandroid/mms-multilingual-audio-5to30min.sbpn-mms-final-vote-5episodes-review
SBPN + MMS final-vote five-episode review set
This dataset contains 474 final selected clips (0.59 hours)
from five source episodes. It is intentionally marked pending manual review.
Selection policy
One winner is retained per transcript chunk using:
frozen_vote
+ 0.10 * mean_rank(MMS_forward, SBPN_forward)
+ 0.10 * mean_rank(-SBPN_previous_context_gain, -SBPN_next_context_gain)
Ranks are calculated only among candidate cuts for the same chunk. This policy was… See the full description on the dataset page: https://huggingface.co/datasets/NancyT/sbpn-mms-final-vote-5episodes-review.
