datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
common_voice_16_1_semisupervisedp2-etf-graph-laplacian-semisupervised-resultssemisupervised-audiobook
Pseudolabel Youtube Malay audiobooks using Whisper Large V3
Notebooks at https://github.com/mesolitica/malaysian-dataset/tree/master/speech-to-text-semisupervised/youtube-audiobook
Split based on 10 seconds utterances using WebRTC VAD.
how-to
Download files,
wget https://huggingface.co/datasets/mesolitica/semisupervised-audiobook/resolve/main/bukan-kerana-aku-5secs-noisy.tar.gz
wget… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/semisupervised-audiobook.semisupervised-abstractive-summarization-ms-newsSemi_supervised_datasetSemi_Supervised_Natural_FoSSIL
