datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
STT_MODEL
Multilingual STT Dataset
Audio and transcript pairs for 50 languages. Each language is a Dataset Viewer configuration with train, validation, and test splits.
Language configurations
amharic: Amharic
arabic_msa: Arabic MSA
assamese: Assamese
bengali: Bengali
czech: Czech
dutch: Dutch
egyptian_arabic: Egyptian Arabic
english: English
farsi_persian: Farsi - Persian
filipino_tagalog: Filipino - Tagalog
french: French
german: German
greek: Greek
gujarati: Gujarati… See the full description on the dataset page: https://huggingface.co/datasets/RidheshBhati/STT_MODEL.spanish-slang-stt-data
Spanish Regional Speech-to-Text Dataset
A multilingual Spanish speech recognition dataset covering 4 regional dialects for fine-tuning Whisper and other ASR models.
Dataset Description
This dataset contains ~39,000 audio samples with transcriptions across 4 Spanish-speaking regions:
Region
Samples
Description
Mexico
17,725
Mexican Spanish including CIEMPIESS corpus
Spain
11,360
Castilian Spanish from TEDx and Common Voice
Argentina
5,839
Rioplatense Spanish… See the full description on the dataset page: https://huggingface.co/datasets/shraavb/spanish-slang-stt-data.STT_ABMalaysian-STT-Whisper
Malaysian STT Whisper format
Heavy postprocessing and post-translation to improve pseudolabeled Whisper Large V3. Also include word level timestamp.
Postprocessing
Check repetitive trigrams.
Verify Voice Activity using Silero-VAD.
Verify scores using Force Alignment.
Post-translation
We use mesolitica/nanot5-base-malaysian-translation-v2.1.
Dataset involved
Malaysian context v2
Singaporean context
Indonesian context
Mandarin audio
Tamil audio… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-STT-Whisper.ramanv-stt-filteredhi-stt-preprocessed-webdatasetstt-benchmark-dataDataset for Pipecat Speech-to-Text benchmarks:
https://github.com/pipecat-ai/stt-benchmark
Uzbek-STT-Dataset-780h
Uzbek STT Dataset (~780 hours)
A large Uzbek speech-to-text dataset for training and fine-tuning automatic
speech recognition (ASR) models such as Whisper.
Dataset summary
Language
Uzbek (uz)
Examples
122,464
Total audio
~780 hours
Clip length
up to 30 seconds each
Columns
audio, transcription
Audio
embedded in parquet, original sample rates (e.g. 44.1 kHz / 16 kHz), mono/stereo as recorded
Split
single train split (split it yourself as… See the full description on the dataset page: https://huggingface.co/datasets/Abduqayum/Uzbek-STT-Dataset-780h.stt-pseudo-labeled-whisper-large-v3-multilingualThis collection includes over 189,000 hours of speech-to-text data in seven languages: English, French, Spanish, Portuguese, Italian, German, and Dutch
All segments were initially sorted by their IDs (timestamps). Adjacent segments from the same source were concatenated into 30-second chunks before being decoded using Whisper-Large-V3. The only exception was Common Voice, where segments were decoded individually before concatenation.
In total, over 288,000 hours of audio data were collected… See the full description on the dataset page: https://huggingface.co/datasets/bofenghuang/stt-pseudo-labeled-whisper-large-v3-multilingual.talkbank_4_stt
Dataset Card
Dataset Description
This dataset is a benchmark based on the TalkBank[1] corpus—a large multilingual repository of conversational speech that captures real-world, unstructured interactions. We use CA-Bank [2], which focuses on phone conversations between adults, which include natural speech phenomena such as laughter, pauses, and interjections. To ensure the dataset is highly accurate and suitable for benchmarking conversational ASR systems, we employ… See the full description on the dataset page: https://huggingface.co/datasets/diabolocom/talkbank_4_stt.IMDA-STT
IMDA National Speech Corpus (NSC) Speech-to-Text
Originally from https://www.imda.gov.sg/how-we-can-help/national-speech-corpus, this repository simply a mirror. This dataset associated with Singapore Open Data Licence, https://www.sla.gov.sg/newsroom/statistics/singapore-open-data-licence
We uploaded mp3 files and compressed using 7z,
7za x part1-mp3.7z.001
All notebooks at https://github.com/mesolitica/malaysian-dataset/tree/master/speech-to-text/imda
total lengths… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/IMDA-STT.malaya-speech-malay-stt
Malaya-Speech Speech-to-Text dataset
This dataset combined from semisupervised Google Speech-to-Text and private datasets.
Processing script https://github.com/mesolitica/malaya-speech/blob/master/pretrained-model/prepare-stt/prepare-malay-stt-train.ipynb
This repository is to centralize the dataset for https://malaya-speech.readthedocs.io/
kb_stt_data
Dataset Card for "kb_stt_data"
More Information needed
stt-benchovos-stt-bench-stt-sampler-v1
OVOS stt bench — stt-sampler-v1
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
OpenVoiceOS/stt-sampler-v1.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-stt-sampler-v1.ramanv-stt-all-rawMalaysian-STT
Malaysian-STT
Prepare streaming and whole mode Speech-to-Text Malaysian context dataset, suitable to train streaming LLM base or Encoder-Decoder such as Whisper.
Merged 30 seconds chunk into one audio file, can up to 10 minutes.
Segmentize based on silent at least 0.3 seconds.
Reject low score based on force alignment.
Reject timestamp anomaly based on force alignment.
Dataset involved
Dialects
IMDA
Malaysian context
Malaysia Parliament
Science context… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/Malaysian-STT.pendakwah_teknologi_yt_stt_datasetMeta_STT_HI_Set1
Meta Speech Recognition Hindi Dataset (Set 1)
This dataset contains both metadata and audio files for Hindi speech recognition samples, curated from multiple sources.
Dataset Sources and Credits
This dataset combines samples from the following sources:
AI4Bharat Indic Speech Dataset
Source: https://ai4bharat.org/indic-speech-dataset
License: CC-BY 4.0
Citation: Please cite the original paper if you use this data
Common Voice Hindi
Source:… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/Meta_STT_HI_Set1.Malaysian-STT-Whisper-Stage2
Malaysian STT Whisper Stage 2
Extra dataset to compliment mesolitica/Malaysian-STT-Whisper.
This dataset is stronger in confidence and suitable for second stage / annealing finetuning.
how to prepare the dataset
huggingface-cli download \
mesolitica/Malaysian-STT-Whisper-Stage2 \
--include "*.zip" \
--repo-type "dataset" \
--local-dir './'
huggingface-cli download \
mesolitica/Malaysian-Multiturn-Chat-Assistant \
--include "*.zip" \
--exclude "voice/*.zip" \
--repo-type… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-STT-Whisper-Stage2.STT_datasetramanv-stt-domainspendakwah_teknologi_yt_stt_datasetramanv-stt-augmentedramanv-stt-stage2-datastt-vibravox-fr-test
VibraVox FR — test split (mirror of Cnam-LMSSC/vibravox)
Mirror public des splits test de VibraVox (CNAM-LMSSC, Paris) pour
benchmark ASR français multi-capteur sur audio standard ET non-standard
(bone-conduction, in-ear, throat, accéléromètre).
Ce repo contient uniquement les configs speech_clean + speech_noisy
(les seules avec transcription). Les configs speechless_* upstream sont
exclues car sans texte → pas de WER possible.
Configs
Config
Test rows
Test… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/stt-vibravox-fr-test.Processed-STT-SALT
Dataset Card for "Processed-STT-SALT"
More Information needed
TAID-Dataset
TAID-Dataset
Terrain intrinsic decomposition dataset with 16,000 scenes and one row per scene.
Columns and numeric spaces
Input, A, S, V: 8-bit RGB PNG. Byte values represent linear values in
[0, 1], quantized as round(clamp(x, 0, 1) * 255). No sRGB/gamma transfer
function is applied.
D: NumPy .npy bytes (float32, HWC RGB), in linear HDR space [0, 5].
water_mask: 8-bit one-hot RGB PNG (R=water, G=terrain, B=sky).
D_filename: original-style filename for the… See the full description on the dataset page: https://huggingface.co/datasets/sttkw/TAID-Dataset.slr54-part1-prepared-stt-v3augmented-dataset-part2-prepared-stt-v3
