datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
chichewa_english_code_switch_dataset
Chichewa-English Code-Switched Speech Dataset
Dataset Description
A speech dataset containing 247 audio recordings of Chichewa-English code-switched phrases. Code-switching — the practice of alternating between two or more languages within a single conversation — is extremely common in Malawi and across multilingual African communities. This dataset captures that natural linguistic behavior in spoken form.
Purpose
This dataset is designed to support research… See the full description on the dataset page: https://huggingface.co/datasets/suru8-ai/chichewa_english_code_switch_dataset.nepali-podcast-code-switch-asr-datasetLegal_audio_dataset_aws_free_code_campasha_twi_dataset_200nepali-podcast-code-switch-asr-datasetaudio_dataset_aws_free_code_camp__audioLegal_audio_dataset_aws_free_code_camp_trimmed3nep_eng_code-mixed_asr_datasetengland-phoneme-dataset
British English Phonetic Dataset
Introduction
This dataset is an extension of Common Voice, from which 6 subsets were selected (Common Voice Corpus 1, Common Voice Corpus 2, Common Voice Corpus 3, Common Voice Corpus 4, Common Voice Corpus 18.0, Common Voice Corpus 19.0). All data containing the England accent from these 6 subsets were extracted and phonetically annotated accordingly.
Description
Key fields explanation:
sentence: The English sentence… See the full description on the dataset page: https://huggingface.co/datasets/zdm-code/england-phoneme-dataset.nep_eng_code-mixed_asr_datasetLegal_audio_dataset_aws_free_code_camp_trimmednep_eng_code-mixed_asr_dataseten-id-code-switch-dataset
