Team Ai
Datasetpublic

simpra/xh-tts-samples

isiXhosa TTS — reference audio and training samples Two very different kinds of audio live here. Check the folder before judging anything. folder what it is source speakers/ REAL HUMAN speech — 20 s excerpts of ViXSD readers, for choosing a voice ViXSD recordings samples_vixsd/ MODEL OUTPUT — what the VITS model generates at a given training step generated speakers/ — ground truth male_xho_reader_008_20s.wav etc. Excerpts taken from the middle of… See the full description on the dataset page: https://huggingface.co/datasets/simpra/xh-tts-samples.

sourceHugging Faceotherupdated 26d agoView on Hugging Face
0likes601downloads
Dataset Card

isiXhosa TTS — reference audio and training samples

Two very different kinds of audio live here. Check the folder before judging anything.

folderwhat it issource
speakers/REAL HUMAN speech — 20 s excerpts of ViXSD readers, for choosing a voiceViXSD recordings
samples_vixsd/MODEL OUTPUT — what the VITS model generates at a given training stepgenerated

speakers/ — ground truth

male_xho_reader_008_20s.wav etc. Excerpts taken from the middle of each reader's longest recording, not the start, so you hear steady-state delivery. These are the target the model is trying to imitate — not model output.

samples_vixsd/ — training progress

step_NNNNNNN.wav is the model's rendering of one fixed sentence at that training step:

Molo, ndicofa phi xa ndirhalela iqanda eliqhekekileyo?

It exercises all three isiXhosa click letters — c (ndicofa), q (iqanda, eliqhekekileyo) and x (xa).

Every file currently in this folder is speaker id 4 (`xho_reader_005`). Files written from now on carry the speaker in the name (step_NNNNNNN_spkN.wav), because id 4 turned out to be the wrong default: it has the most audio (108.9 min) but the worst signal-to-noise in the corpus, about 18 dB against 43-66 dB for the others. It is quiet and noisy, which is why step_0008000.wav sounds distorted.

The default is now id 7 (`xho_reader_008`) — male, 38.7 min, 66 dB SNR, the cleanest voice available.

Expected progression

stepexpect
0noise (random init)
2,000speech rhythm, no words
~8,000words forming, clicks audible, still rough
20,000+clearer, less robotic

Judge quality against speakers/, not against commercial TTS: the model trains on 7.4 h, where VITS's reference corpus LJSpeech has ~24 h.

Related