Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01RidheshBhati /STT_MODEL Multilingual STT Dataset Audio and transcript pairs for 50 languages. Each language is a Dataset Viewer configuration with train, validation, and test splits. Language configurations amharic: Amharic arabic_msa: Arabic MSA assamese: Assamese bengali: Bengali czech: Czech dutch: Dutch egyptian_arabic: Egyptian Arabic english: English farsi_persian: Farsi - Persian filipino_tagalog: Filipino - Tagalog french: French german: German greek: Greek gujarati: Gujarati… See the full description on the dataset page: https://huggingface.co/datasets/RidheshBhati/STT_MODEL.audio1M<n<10M1 likes9k downloads2mo agoHugging Face02shraavb /spanish-slang-stt-data Spanish Regional Speech-to-Text Dataset A multilingual Spanish speech recognition dataset covering 4 regional dialects for fine-tuning Whisper and other ASR models. Dataset Description This dataset contains ~39,000 audio samples with transcriptions across 4 Spanish-speaking regions: Region Samples Description Mexico 17,725 Mexican Spanish including CIEMPIESS corpus Spain 11,360 Castilian Spanish from TEDx and Common Voice Argentina 5,839 Rioplatense Spanish… See the full description on the dataset page: https://huggingface.co/datasets/shraavb/spanish-slang-stt-data.audioautomatic-speech-recognition10K<n<100K0 likes3.4k downloads9mo agoHugging Face03openpecha /STT_AB0 likes2.9k downloads3y agoHugging Face04mesolitica /Malaysian-STT-Whisper Malaysian STT Whisper format Heavy postprocessing and post-translation to improve pseudolabeled Whisper Large V3. Also include word level timestamp. Postprocessing Check repetitive trigrams. Verify Voice Activity using Silero-VAD. Verify scores using Force Alignment. Post-translation We use mesolitica/nanot5-base-malaysian-translation-v2.1. Dataset involved Malaysian context v2 Singaporean context Indonesian context Mandarin audio Tamil audio… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-STT-Whisper.audioautomatic-speech-recognition10M<n<100M5 likes2.3k downloads1y agoHugging Face05guruawe /ramanv-stt-filteredgatedtext10K<n<100K5 likes2.3k downloads4d agoHugging Face06collabora /hi-stt-preprocessed-webdatasettext100K<n<1M1 likes2k downloads1y agoHugging Face07pipecat-ai /stt-benchmark-dataDataset for Pipecat Speech-to-Text benchmarks: https://github.com/pipecat-ai/stt-benchmark audio1K<n<10K13 likes1.4k downloads7d agoHugging Face08Abduqayum /Uzbek-STT-Dataset-780h Uzbek STT Dataset (~780 hours) A large Uzbek speech-to-text dataset for training and fine-tuning automatic speech recognition (ASR) models such as Whisper. Dataset summary Language Uzbek (uz) Examples 122,464 Total audio ~780 hours Clip length up to 30 seconds each Columns audio, transcription Audio embedded in parquet, original sample rates (e.g. 44.1 kHz / 16 kHz), mono/stereo as recorded Split single train split (split it yourself as… See the full description on the dataset page: https://huggingface.co/datasets/Abduqayum/Uzbek-STT-Dataset-780h.audioautomatic-speech-recognition100K<n<1M2 likes1k downloads3mo agoHugging Face09bofenghuang /stt-pseudo-labeled-whisper-large-v3-multilingualThis collection includes over 189,000 hours of speech-to-text data in seven languages: English, French, Spanish, Portuguese, Italian, German, and Dutch All segments were initially sorted by their IDs (timestamps). Adjacent segments from the same source were concatenated into 30-second chunks before being decoded using Whisper-Large-V3. The only exception was Common Voice, where segments were decoded individually before concatenation. In total, over 288,000 hours of audio data were collected… See the full description on the dataset page: https://huggingface.co/datasets/bofenghuang/stt-pseudo-labeled-whisper-large-v3-multilingual.4 likes1k downloads2y agoHugging Face10diabolocom /talkbank_4_stt Dataset Card Dataset Description This dataset is a benchmark based on the TalkBank[1] corpus—a large multilingual repository of conversational speech that captures real-world, unstructured interactions. We use CA-Bank [2], which focuses on phone conversations between adults, which include natural speech phenomena such as laughter, pauses, and interjections. To ensure the dataset is highly accurate and suitable for benchmarking conversational ASR systems, we employ… See the full description on the dataset page: https://huggingface.co/datasets/diabolocom/talkbank_4_stt.audioautomatic-speech-recognition100K<n<1M2 likes912 downloads1y agoHugging Face11mesolitica /IMDA-STT IMDA National Speech Corpus (NSC) Speech-to-Text Originally from https://www.imda.gov.sg/how-we-can-help/national-speech-corpus, this repository simply a mirror. This dataset associated with Singapore Open Data Licence, https://www.sla.gov.sg/newsroom/statistics/singapore-open-data-licence We uploaded mp3 files and compressed using 7z, 7za x part1-mp3.7z.001 All notebooks at https://github.com/mesolitica/malaysian-dataset/tree/master/speech-to-text/imda total lengths… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/IMDA-STT.text1M<n<10M6 likes911 downloads1y agoHugging Face12mesolitica /malaya-speech-malay-stt Malaya-Speech Speech-to-Text dataset This dataset combined from semisupervised Google Speech-to-Text and private datasets. Processing script https://github.com/mesolitica/malaya-speech/blob/master/pretrained-model/prepare-stt/prepare-malay-stt-train.ipynb This repository is to centralize the dataset for https://malaya-speech.readthedocs.io/ audio1M<n<10M9 likes839 downloads3y agoHugging Face13BaekRok /kb_stt_data Dataset Card for "kb_stt_data" More Information needed audio100K<n<1M1 likes742 downloads3y agoHugging Face14kalpalabs /stt-benchaudio100K<n<1M0 likes647 downloads11mo agoHugging Face15OpenVoiceOS /ovos-stt-bench-stt-sampler-v1 OVOS stt bench — stt-sampler-v1 Per-clip transcripts predictions of the registered OVOS Plugin Arena stt fighters over OpenVoiceOS/stt-sampler-v1. One dedicated repo per modality; one dataset split per language; one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow the arena §3.2 contract (pinned dataset_revision, plugin_version, latency_ms). Produced by the reproducible benchmark script in the arena repo; the arena's assemble workflow turns… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-stt-sampler-v1.0 likes560 downloads25d agoHugging Face16lingamvamshikrishnareddy /ramanv-stt-all-raw0 likes510 downloads27d agoHugging Face17malaysia-ai /Malaysian-STT Malaysian-STT Prepare streaming and whole mode Speech-to-Text Malaysian context dataset, suitable to train streaming LLM base or Encoder-Decoder such as Whisper. Merged 30 seconds chunk into one audio file, can up to 10 minutes. Segmentize based on silent at least 0.3 seconds. Reject low score based on force alignment. Reject timestamp anomaly based on force alignment. Dataset involved Dialects IMDA Malaysian context Malaysia Parliament Science context… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/Malaysian-STT.text10M<n<100M2 likes492 downloads1y agoHugging Face18cheelam /pendakwah_teknologi_yt_stt_datasetaudio100K<n<1M0 likes477 downloads2y agoHugging Face19WhissleAI /Meta_STT_HI_Set1 Meta Speech Recognition Hindi Dataset (Set 1) This dataset contains both metadata and audio files for Hindi speech recognition samples, curated from multiple sources. Dataset Sources and Credits This dataset combines samples from the following sources: AI4Bharat Indic Speech Dataset Source: https://ai4bharat.org/indic-speech-dataset License: CC-BY 4.0 Citation: Please cite the original paper if you use this data Common Voice Hindi Source:… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/Meta_STT_HI_Set1.audioautomatic-speech-recognition100K<n<1M0 likes473 downloads1y agoHugging Face20mesolitica /Malaysian-STT-Whisper-Stage2 Malaysian STT Whisper Stage 2 Extra dataset to compliment mesolitica/Malaysian-STT-Whisper. This dataset is stronger in confidence and suitable for second stage / annealing finetuning. how to prepare the dataset huggingface-cli download \ mesolitica/Malaysian-STT-Whisper-Stage2 \ --include "*.zip" \ --repo-type "dataset" \ --local-dir './' huggingface-cli download \ mesolitica/Malaysian-Multiturn-Chat-Assistant \ --include "*.zip" \ --exclude "voice/*.zip" \ --repo-type… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-STT-Whisper-Stage2.text10M<n<100M2 likes428 downloads1y agoHugging Face21Cathle /STT_datasetaudio1K<n<10K0 likes424 downloads1y agoHugging Face22guruawe /ramanv-stt-domainsgatedtext100K<n<1M0 likes424 downloads17d agoHugging Face23skilledu /pendakwah_teknologi_yt_stt_datasetaudio100K<n<1M0 likes387 downloads4mo agoHugging Face24guruawe /ramanv-stt-augmentedgatedtext100K<n<1M1 likes385 downloads23d agoHugging Face25guruawe /ramanv-stt-stage2-data0 likes384 downloads15d agoHugging Face26ggfox00000 /stt-vibravox-fr-test VibraVox FR — test split (mirror of Cnam-LMSSC/vibravox) Mirror public des splits test de VibraVox (CNAM-LMSSC, Paris) pour benchmark ASR français multi-capteur sur audio standard ET non-standard (bone-conduction, in-ear, throat, accéléromètre). Ce repo contient uniquement les configs speech_clean + speech_noisy (les seules avec transcription). Les configs speechless_* upstream sont exclues car sans texte → pas de WER possible. Configs Config Test rows Test… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/stt-vibravox-fr-test.audioautomatic-speech-recognition1K<n<10K0 likes378 downloads5mo agoHugging Face27mekaneeky /Processed-STT-SALT Dataset Card for "Processed-STT-SALT" More Information needed 10K<n<100K0 likes366 downloads3y agoHugging Face28sttkw /TAID-Dataset TAID-Dataset Terrain intrinsic decomposition dataset with 16,000 scenes and one row per scene. Columns and numeric spaces Input, A, S, V: 8-bit RGB PNG. Byte values represent linear values in [0, 1], quantized as round(clamp(x, 0, 1) * 255). No sRGB/gamma transfer function is applied. D: NumPy .npy bytes (float32, HWC RGB), in linear HDR space [0, 5]. water_mask: 8-bit one-hot RGB PNG (R=water, G=terrain, B=sky). D_filename: original-style filename for the… See the full description on the dataset page: https://huggingface.co/datasets/sttkw/TAID-Dataset.image10K<n<100K0 likes360 downloads1mo agoHugging Face29awajai /slr54-part1-prepared-stt-v3audio10K<n<100K0 likes336 downloads2y agoHugging Face30awajai /augmented-dataset-part2-prepared-stt-v3audio10K<n<100K0 likes328 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.