simpra/xh-tts-samples
isiXhosa TTS — reference audio and training samples Two very different kinds of audio live here. Check the folder before judging anything. folder what it is source speakers/ REAL HUMAN speech — 20 s excerpts of ViXSD readers, for choosing a voice ViXSD recordings samples_vixsd/ MODEL OUTPUT — what the VITS model generates at a given training step generated speakers/ — ground truth male_xho_reader_008_20s.wav etc. Excerpts taken from the middle of… See the full description on the dataset page: https://huggingface.co/datasets/simpra/xh-tts-samples.
isiXhosa TTS — reference audio and training samples
Two very different kinds of audio live here. Check the folder before judging anything.
speakers/ — ground truth
male_xho_reader_008_20s.wav etc. Excerpts taken from the middle of each reader's longest recording, not the start, so you hear steady-state delivery. These are the target the model is trying to imitate — not model output.
samples_vixsd/ — training progress
step_NNNNNNN.wav is the model's rendering of one fixed sentence at that training step:
Molo, ndicofa phi xa ndirhalela iqanda eliqhekekileyo?
It exercises all three isiXhosa click letters — c (ndicofa), q (iqanda, eliqhekekileyo) and x (xa).
Every file currently in this folder is speaker id 4 (`xho_reader_005`). Files written from now on carry the speaker in the name (step_NNNNNNN_spkN.wav), because id 4 turned out to be the wrong default: it has the most audio (108.9 min) but the worst signal-to-noise in the corpus, about 18 dB against 43-66 dB for the others. It is quiet and noisy, which is why step_0008000.wav sounds distorted.
The default is now id 7 (`xho_reader_008`) — male, 38.7 min, 66 dB SNR, the cleanest voice available.
Expected progression
Judge quality against speakers/, not against commercial TTS: the model trains on 7.4 h, where VITS's reference corpus LJSpeech has ~24 h.
Related
- training clips: simpra/xh-tts-vixsd
- alignment spot-checks: simpra/xh-tts-check
