Team Ai
Datasetpublic

Porameht/processed-voice-th-169k

processed-voice-th-169k Cleaned open-source Thai speech dataset: 169,550 utterances (149,953 train / 7,614 dev / 11,983 test) with transcripts, prepared for ASR fine-tuning. Format Field Description sentence Transcript in Thai audio Audio clip Usage from datasets import load_dataset ds = load_dataset("Porameht/processed-voice-th-169k") Used to fine-tune Porameht/whisper-tiny-thai. See also the smaller, cleaner… See the full description on the dataset page: https://huggingface.co/datasets/Porameht/processed-voice-th-169k.

sourceHugging Facecc-by-sa-4.0updated 1mo agoView on Hugging Face
17likes103downloads
Dataset Card

processed-voice-th-169k

Cleaned open-source Thai speech dataset: 169,550 utterances (149,953 train / 7,614 dev / 11,983 test) with transcripts, prepared for ASR fine-tuning.

Format

FieldDescription
sentenceTranscript in Thai
audioAudio clip

Usage

python
from datasets import load_dataset

ds = load_dataset("Porameht/processed-voice-th-169k")

Used to fine-tune Porameht/whisper-tiny-thai. See also the smaller, cleaner processed-cv-17-th-130k derived from Common Voice 17.