datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
STT_MODEL
Multilingual STT Dataset
Audio and transcript pairs for 50 languages. Each language is a Dataset Viewer configuration with train, validation, and test splits.
Language configurations
amharic: Amharic
arabic_msa: Arabic MSA
assamese: Assamese
bengali: Bengali
czech: Czech
dutch: Dutch
egyptian_arabic: Egyptian Arabic
english: English
farsi_persian: Farsi - Persian
filipino_tagalog: Filipino - Tagalog
french: French
german: German
greek: Greek
gujarati: Gujarati… See the full description on the dataset page: https://huggingface.co/datasets/RidheshBhati/STT_MODEL.spanish-slang-stt-data
Spanish Regional Speech-to-Text Dataset
A multilingual Spanish speech recognition dataset covering 4 regional dialects for fine-tuning Whisper and other ASR models.
Dataset Description
This dataset contains ~39,000 audio samples with transcriptions across 4 Spanish-speaking regions:
Region
Samples
Description
Mexico
17,725
Mexican Spanish including CIEMPIESS corpus
Spain
11,360
Castilian Spanish from TEDx and Common Voice
Argentina
5,839
Rioplatense Spanish… See the full description on the dataset page: https://huggingface.co/datasets/shraavb/spanish-slang-stt-data.Malaysian-STT-Whisper
Malaysian STT Whisper format
Heavy postprocessing and post-translation to improve pseudolabeled Whisper Large V3. Also include word level timestamp.
Postprocessing
Check repetitive trigrams.
Verify Voice Activity using Silero-VAD.
Verify scores using Force Alignment.
Post-translation
We use mesolitica/nanot5-base-malaysian-translation-v2.1.
Dataset involved
Malaysian context v2
Singaporean context
Indonesian context
Mandarin audio
Tamil audio… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-STT-Whisper.hi-stt-preprocessed-webdatasetramanv-stt-filteredstt-benchmark-dataDataset for Pipecat Speech-to-Text benchmarks:
https://github.com/pipecat-ai/stt-benchmark
Uzbek-STT-Dataset-780h
Uzbek STT Dataset (~780 hours)
A large Uzbek speech-to-text dataset for training and fine-tuning automatic
speech recognition (ASR) models such as Whisper.
Dataset summary
Language
Uzbek (uz)
Examples
122,464
Total audio
~780 hours
Clip length
up to 30 seconds each
Columns
audio, transcription
Audio
embedded in parquet, original sample rates (e.g. 44.1 kHz / 16 kHz), mono/stereo as recorded
Split
single train split (split it yourself as… See the full description on the dataset page: https://huggingface.co/datasets/Abduqayum/Uzbek-STT-Dataset-780h.talkbank_4_stt
Dataset Card
Dataset Description
This dataset is a benchmark based on the TalkBank[1] corpus—a large multilingual repository of conversational speech that captures real-world, unstructured interactions. We use CA-Bank [2], which focuses on phone conversations between adults, which include natural speech phenomena such as laughter, pauses, and interjections. To ensure the dataset is highly accurate and suitable for benchmarking conversational ASR systems, we employ… See the full description on the dataset page: https://huggingface.co/datasets/diabolocom/talkbank_4_stt.kb_stt_data
Dataset Card for "kb_stt_data"
More Information needed
malaya-speech-malay-stt
Malaya-Speech Speech-to-Text dataset
This dataset combined from semisupervised Google Speech-to-Text and private datasets.
Processing script https://github.com/mesolitica/malaya-speech/blob/master/pretrained-model/prepare-stt/prepare-malay-stt-train.ipynb
This repository is to centralize the dataset for https://malaya-speech.readthedocs.io/
IMDA-STT
IMDA National Speech Corpus (NSC) Speech-to-Text
Originally from https://www.imda.gov.sg/how-we-can-help/national-speech-corpus, this repository simply a mirror. This dataset associated with Singapore Open Data Licence, https://www.sla.gov.sg/newsroom/statistics/singapore-open-data-licence
We uploaded mp3 files and compressed using 7z,
7za x part1-mp3.7z.001
All notebooks at https://github.com/mesolitica/malaysian-dataset/tree/master/speech-to-text/imda
total lengths… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/IMDA-STT.pendakwah_teknologi_yt_stt_datasetMalaysian-STT
Malaysian-STT
Prepare streaming and whole mode Speech-to-Text Malaysian context dataset, suitable to train streaming LLM base or Encoder-Decoder such as Whisper.
Merged 30 seconds chunk into one audio file, can up to 10 minutes.
Segmentize based on silent at least 0.3 seconds.
Reject low score based on force alignment.
Reject timestamp anomaly based on force alignment.
Dataset involved
Dialects
IMDA
Malaysian context
Malaysia Parliament
Science context… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/Malaysian-STT.stt-benchMalaysian-STT-Whisper-Stage2
Malaysian STT Whisper Stage 2
Extra dataset to compliment mesolitica/Malaysian-STT-Whisper.
This dataset is stronger in confidence and suitable for second stage / annealing finetuning.
how to prepare the dataset
huggingface-cli download \
mesolitica/Malaysian-STT-Whisper-Stage2 \
--include "*.zip" \
--repo-type "dataset" \
--local-dir './'
huggingface-cli download \
mesolitica/Malaysian-Multiturn-Chat-Assistant \
--include "*.zip" \
--exclude "voice/*.zip" \
--repo-type… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-STT-Whisper-Stage2.ramanv-stt-augmentedstt-vibravox-fr-test
VibraVox FR — test split (mirror of Cnam-LMSSC/vibravox)
Mirror public des splits test de VibraVox (CNAM-LMSSC, Paris) pour
benchmark ASR français multi-capteur sur audio standard ET non-standard
(bone-conduction, in-ear, throat, accéléromètre).
Ce repo contient uniquement les configs speech_clean + speech_noisy
(les seules avec transcription). Les configs speechless_* upstream sont
exclues car sans texte → pas de WER possible.
Configs
Config
Test rows
Test… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/stt-vibravox-fr-test.TAID-Dataset
TAID-Dataset
Terrain intrinsic decomposition dataset with 16,000 scenes and one row per scene.
Columns and numeric spaces
Input, A, S, V: 8-bit RGB PNG. Byte values represent linear values in
[0, 1], quantized as round(clamp(x, 0, 1) * 255). No sRGB/gamma transfer
function is applied.
D: NumPy .npy bytes (float32, HWC RGB), in linear HDR space [0, 5].
water_mask: 8-bit one-hot RGB PNG (R=water, G=terrain, B=sky).
D_filename: original-style filename for the… See the full description on the dataset page: https://huggingface.co/datasets/sttkw/TAID-Dataset.Meta_STT_HI_Set1
Meta Speech Recognition Hindi Dataset (Set 1)
This dataset contains both metadata and audio files for Hindi speech recognition samples, curated from multiple sources.
Dataset Sources and Credits
This dataset combines samples from the following sources:
AI4Bharat Indic Speech Dataset
Source: https://ai4bharat.org/indic-speech-dataset
License: CC-BY 4.0
Citation: Please cite the original paper if you use this data
Common Voice Hindi
Source:… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/Meta_STT_HI_Set1.pendakwah_teknologi_yt_stt_datasetSTT_datasetAgri_STT_Benchmarking_Dataset
Agri STT Benchmarking Dataset
10,808 farmer voice queries in Hindi, Telugu and Odia, with reference transcripts, for benchmarking automatic speech recognition in agricultural contexts. The audio is included in this repository.
Every recording is a smallholder farmer speaking a question to Farmer.Chat, an AI advisory service run by Digital Green. Reference transcripts were produced by human annotators. Nothing here is read from a script or recorded in a studio, so the audio… See the full description on the dataset page: https://huggingface.co/datasets/DigiGreen/Agri_STT_Benchmarking_Dataset.TAID-AtmosEdit
TAID-AtmosEdit
Atmospheric editing dataset with 10,000 rows. Each seed identifies a scene
and p_idx identifies one of its atmospheric conditions.
Columns and numeric spaces
S, V: 8-bit RGB PNG. Byte values contain linearly quantized values using
round(clamp(x, 0, 1) * 255). No sRGB/gamma transfer function is applied.
D: NumPy .npy bytes (float32, HWC RGB) in linear HDR space [0, 5].
s_density, s_aerosol, s_ozone: atmospheric parameters as float32.
seed, p_idx:… See the full description on the dataset page: https://huggingface.co/datasets/sttkw/TAID-AtmosEdit.slr54-part1-prepared-stt-v3augmented-dataset-part2-prepared-stt-v3STT_Korean_Datasetpure_pixel_yt_stt_datasetMeta_STT_EN_Set2
Meta Speech Recognition English Dataset (Set 2)
This dataset contains both metadata and audio files for English speech recognition samples.
Dataset Statistics
Splits and Sample Counts
train: 42961 samples
valid: 2387 samples
test: 2387 samples
Example Samples
train
{
"audio_filepath": "/external1/datasets/asr-himanshu/avspeech-data/audio/AzSutepklXI_2.wav",
"text": "To Jesus, so God is faithful, because when he keeps, you know, when… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/Meta_STT_EN_Set2.augmented-dataset-part1-prepared-stt-v3darija-stt-mixThe Darija Speech To Text Dataset is a comprehensive collection designed to support speech recognition tasks for the Darija dialect, it includes audio data totaling 8.23 GB and consists of 13,178 rows of transcribed speech.
This dataset covers a variety of dialects, primarily focusing on Algerian and Moroccan Darija, and also includes slang from other Arabic-speaking countries.
The data has been meticulously gathered from diverse resources to ensure a rich and varied representation of spoken… See the full description on the dataset page: https://huggingface.co/datasets/ayoubkirouane/darija-stt-mix.
