datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
khmer-english-codeswitch-tts-llm
Khmer–English Code-Switch Synthetic Speech (LLM-authored)
19,825 utterances / 21.7 hours of synthetic Khmer–English code-switched speech at
16 kHz, generated with VoxCPM2 from code-switch
sentences written by an LLM and validated programmatically.
⚠️ This is synthetic speech, not recordings of people. Every utterance was produced by a TTS
model, and every sentence was written by a language model — they are not transcripts of anything
a person said. It is intended as an… See the full description on the dataset page: https://huggingface.co/datasets/Panhapich/khmer-english-codeswitch-tts-llm.Fleurs-KnThis is a filtered version of the Fleurs dataset only containing samples of Kannada language.
The dataset contains total of 2283 training, 368 validation and 838 test samples.
Data Sample:
{'id': 1053,
'num_samples': 226560,
'path': '/home/ravi.naik/.cache/huggingface/datasets/downloads/extracted/e7c8b501d4e6892673b6dc291d42de48e7987b0d2aa6471066a671f686224ed1/10000267636955490843.wav',
'audio': {'path': 'train/10000267636955490843.wav',
'array': array([ 0. , 0.… See the full description on the dataset page: https://huggingface.co/datasets/Indic-LLM-Labs/Fleurs-Kn.Fleurs-KnThis is a filtered version of the Fleurs dataset only containing samples of Kannada language.
The dataset contains total of 2283 training, 368 validation and 838 test samples.
Data Sample:
{'id': 1053,
'num_samples': 226560,
'path': '/home/ravi.naik/.cache/huggingface/datasets/downloads/extracted/e7c8b501d4e6892673b6dc291d42de48e7987b0d2aa6471066a671f686224ed1/10000267636955490843.wav',
'audio': {'path': 'train/10000267636955490843.wav',
'array': array([ 0. , 0.… See the full description on the dataset page: https://huggingface.co/datasets/Kannada-LLM-Labs/Fleurs-Kn.Egyptian-Arabic-Data-LLMllmPapers
