Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01RidheshBhati /STT_MODEL Multilingual STT Dataset Audio and transcript pairs for 50 languages. Each language is a Dataset Viewer configuration with train, validation, and test splits. Language configurations amharic: Amharic arabic_msa: Arabic MSA assamese: Assamese bengali: Bengali czech: Czech dutch: Dutch egyptian_arabic: Egyptian Arabic english: English farsi_persian: Farsi - Persian filipino_tagalog: Filipino - Tagalog french: French german: German greek: Greek gujarati: Gujarati… See the full description on the dataset page: https://huggingface.co/datasets/RidheshBhati/STT_MODEL.audio1M<n<10M1 likes8k downloads2mo agoHugging Face02shraavb /spanish-slang-stt-data Spanish Regional Speech-to-Text Dataset A multilingual Spanish speech recognition dataset covering 4 regional dialects for fine-tuning Whisper and other ASR models. Dataset Description This dataset contains ~39,000 audio samples with transcriptions across 4 Spanish-speaking regions: Region Samples Description Mexico 17,725 Mexican Spanish including CIEMPIESS corpus Spain 11,360 Castilian Spanish from TEDx and Common Voice Argentina 5,839 Rioplatense Spanish… See the full description on the dataset page: https://huggingface.co/datasets/shraavb/spanish-slang-stt-data.audioautomatic-speech-recognition10K<n<100K0 likes4.4k downloads9mo agoHugging Face03mesolitica /Malaysian-STT-Whisper Malaysian STT Whisper format Heavy postprocessing and post-translation to improve pseudolabeled Whisper Large V3. Also include word level timestamp. Postprocessing Check repetitive trigrams. Verify Voice Activity using Silero-VAD. Verify scores using Force Alignment. Post-translation We use mesolitica/nanot5-base-malaysian-translation-v2.1. Dataset involved Malaysian context v2 Singaporean context Indonesian context Mandarin audio Tamil audio… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-STT-Whisper.audioautomatic-speech-recognition10M<n<100M5 likes2.3k downloads1y agoHugging Face04collabora /hi-stt-preprocessed-webdatasettext100K<n<1M1 likes2.1k downloads1y agoHugging Face05guruawe /ramanv-stt-filteredgatedtext10K<n<100K5 likes1.8k downloads8d agoHugging Face06pipecat-ai /stt-benchmark-dataDataset for Pipecat Speech-to-Text benchmarks: https://github.com/pipecat-ai/stt-benchmark audio1K<n<10K13 likes1.5k downloads11d agoHugging Face07Abduqayum /Uzbek-STT-Dataset-780h Uzbek STT Dataset (~780 hours) A large Uzbek speech-to-text dataset for training and fine-tuning automatic speech recognition (ASR) models such as Whisper. Dataset summary Language Uzbek (uz) Examples 122,464 Total audio ~780 hours Clip length up to 30 seconds each Columns audio, transcription Audio embedded in parquet, original sample rates (e.g. 44.1 kHz / 16 kHz), mono/stereo as recorded Split single train split (split it yourself as… See the full description on the dataset page: https://huggingface.co/datasets/Abduqayum/Uzbek-STT-Dataset-780h.audioautomatic-speech-recognition100K<n<1M3 likes1k downloads3mo agoHugging Face08diabolocom /talkbank_4_stt Dataset Card Dataset Description This dataset is a benchmark based on the TalkBank[1] corpus—a large multilingual repository of conversational speech that captures real-world, unstructured interactions. We use CA-Bank [2], which focuses on phone conversations between adults, which include natural speech phenomena such as laughter, pauses, and interjections. To ensure the dataset is highly accurate and suitable for benchmarking conversational ASR systems, we employ… See the full description on the dataset page: https://huggingface.co/datasets/diabolocom/talkbank_4_stt.audioautomatic-speech-recognition100K<n<1M2 likes711 downloads1y agoHugging Face09BaekRok /kb_stt_data Dataset Card for "kb_stt_data" More Information needed audio100K<n<1M1 likes669 downloads3y agoHugging Face10mesolitica /malaya-speech-malay-stt Malaya-Speech Speech-to-Text dataset This dataset combined from semisupervised Google Speech-to-Text and private datasets. Processing script https://github.com/mesolitica/malaya-speech/blob/master/pretrained-model/prepare-stt/prepare-malay-stt-train.ipynb This repository is to centralize the dataset for https://malaya-speech.readthedocs.io/ audio1M<n<10M9 likes643 downloads3y agoHugging Face11mesolitica /IMDA-STT IMDA National Speech Corpus (NSC) Speech-to-Text Originally from https://www.imda.gov.sg/how-we-can-help/national-speech-corpus, this repository simply a mirror. This dataset associated with Singapore Open Data Licence, https://www.sla.gov.sg/newsroom/statistics/singapore-open-data-licence We uploaded mp3 files and compressed using 7z, 7za x part1-mp3.7z.001 All notebooks at https://github.com/mesolitica/malaysian-dataset/tree/master/speech-to-text/imda total lengths… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/IMDA-STT.text1M<n<10M6 likes622 downloads1y agoHugging Face12cheelam /pendakwah_teknologi_yt_stt_datasetaudio100K<n<1M0 likes505 downloads2y agoHugging Face13malaysia-ai /Malaysian-STT Malaysian-STT Prepare streaming and whole mode Speech-to-Text Malaysian context dataset, suitable to train streaming LLM base or Encoder-Decoder such as Whisper. Merged 30 seconds chunk into one audio file, can up to 10 minutes. Segmentize based on silent at least 0.3 seconds. Reject low score based on force alignment. Reject timestamp anomaly based on force alignment. Dataset involved Dialects IMDA Malaysian context Malaysia Parliament Science context… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/Malaysian-STT.text10M<n<100M2 likes474 downloads1y agoHugging Face14kalpalabs /stt-benchaudio100K<n<1M0 likes435 downloads1y agoHugging Face15mesolitica /Malaysian-STT-Whisper-Stage2 Malaysian STT Whisper Stage 2 Extra dataset to compliment mesolitica/Malaysian-STT-Whisper. This dataset is stronger in confidence and suitable for second stage / annealing finetuning. how to prepare the dataset huggingface-cli download \ mesolitica/Malaysian-STT-Whisper-Stage2 \ --include "*.zip" \ --repo-type "dataset" \ --local-dir './' huggingface-cli download \ mesolitica/Malaysian-Multiturn-Chat-Assistant \ --include "*.zip" \ --exclude "voice/*.zip" \ --repo-type… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-STT-Whisper-Stage2.text10M<n<100M2 likes421 downloads1y agoHugging Face16guruawe /ramanv-stt-augmentedgatedtext100K<n<1M1 likes385 downloads27d agoHugging Face17ggfox00000 /stt-vibravox-fr-test VibraVox FR — test split (mirror of Cnam-LMSSC/vibravox) Mirror public des splits test de VibraVox (CNAM-LMSSC, Paris) pour benchmark ASR français multi-capteur sur audio standard ET non-standard (bone-conduction, in-ear, throat, accéléromètre). Ce repo contient uniquement les configs speech_clean + speech_noisy (les seules avec transcription). Les configs speechless_* upstream sont exclues car sans texte → pas de WER possible. Configs Config Test rows Test… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/stt-vibravox-fr-test.audioautomatic-speech-recognition1K<n<10K0 likes377 downloads6mo agoHugging Face18sttkw /TAID-Dataset TAID-Dataset Terrain intrinsic decomposition dataset with 16,000 scenes and one row per scene. Columns and numeric spaces Input, A, S, V: 8-bit RGB PNG. Byte values represent linear values in [0, 1], quantized as round(clamp(x, 0, 1) * 255). No sRGB/gamma transfer function is applied. D: NumPy .npy bytes (float32, HWC RGB), in linear HDR space [0, 5]. water_mask: 8-bit one-hot RGB PNG (R=water, G=terrain, B=sky). D_filename: original-style filename for the… See the full description on the dataset page: https://huggingface.co/datasets/sttkw/TAID-Dataset.image10K<n<100K0 likes351 downloads1mo agoHugging Face19WhissleAI /Meta_STT_HI_Set1 Meta Speech Recognition Hindi Dataset (Set 1) This dataset contains both metadata and audio files for Hindi speech recognition samples, curated from multiple sources. Dataset Sources and Credits This dataset combines samples from the following sources: AI4Bharat Indic Speech Dataset Source: https://ai4bharat.org/indic-speech-dataset License: CC-BY 4.0 Citation: Please cite the original paper if you use this data Common Voice Hindi Source:… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/Meta_STT_HI_Set1.audioautomatic-speech-recognition100K<n<1M0 likes338 downloads1y agoHugging Face20skilledu /pendakwah_teknologi_yt_stt_datasetaudio100K<n<1M0 likes310 downloads4mo agoHugging Face21Cathle /STT_datasetaudio1K<n<10K0 likes273 downloads1y agoHugging Face22DigiGreen /Agri_STT_Benchmarking_Dataset Agri STT Benchmarking Dataset 10,808 farmer voice queries in Hindi, Telugu and Odia, with reference transcripts, for benchmarking automatic speech recognition in agricultural contexts. The audio is included in this repository. Every recording is a smallholder farmer speaking a question to Farmer.Chat, an AI advisory service run by Digital Green. Reference transcripts were produced by human annotators. Nothing here is read from a script or recorded in a studio, so the audio… See the full description on the dataset page: https://huggingface.co/datasets/DigiGreen/Agri_STT_Benchmarking_Dataset.audioautomatic-speech-recognition10K<n<100K3 likes260 downloads2mo agoHugging Face23sttkw /TAID-AtmosEdit TAID-AtmosEdit Atmospheric editing dataset with 10,000 rows. Each seed identifies a scene and p_idx identifies one of its atmospheric conditions. Columns and numeric spaces S, V: 8-bit RGB PNG. Byte values contain linearly quantized values using round(clamp(x, 0, 1) * 255). No sRGB/gamma transfer function is applied. D: NumPy .npy bytes (float32, HWC RGB) in linear HDR space [0, 5]. s_density, s_aerosol, s_ozone: atmospheric parameters as float32. seed, p_idx:… See the full description on the dataset page: https://huggingface.co/datasets/sttkw/TAID-AtmosEdit.image10K<n<100K0 likes254 downloads1mo agoHugging Face24awajai /slr54-part1-prepared-stt-v3audio10K<n<100K0 likes238 downloads2y agoHugging Face25awajai /augmented-dataset-part2-prepared-stt-v3audio10K<n<100K0 likes220 downloads2y agoHugging Face26Junhoee /STT_Korean_Datasetaudio100K<n<1M6 likes219 downloads2y agoHugging Face27cheelam /pure_pixel_yt_stt_datasetaudio100K<n<1M0 likes215 downloads2y agoHugging Face28WhissleAI /Meta_STT_EN_Set2 Meta Speech Recognition English Dataset (Set 2) This dataset contains both metadata and audio files for English speech recognition samples. Dataset Statistics Splits and Sample Counts train: 42961 samples valid: 2387 samples test: 2387 samples Example Samples train { "audio_filepath": "/external1/datasets/asr-himanshu/avspeech-data/audio/AzSutepklXI_2.wav", "text": "To Jesus, so God is faithful, because when he keeps, you know, when… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/Meta_STT_EN_Set2.audioaudio-classification10K<n<100K0 likes213 downloads1y agoHugging Face29awajai /augmented-dataset-part1-prepared-stt-v3audio10K<n<100K0 likes210 downloads2y agoHugging Face30ayoubkirouane /darija-stt-mixThe Darija Speech To Text Dataset is a comprehensive collection designed to support speech recognition tasks for the Darija dialect, it includes audio data totaling 8.23 GB and consists of 13,178 rows of transcribed speech. This dataset covers a variety of dialects, primarily focusing on Algerian and Moroccan Darija, and also includes slang from other Arabic-speaking countries. The data has been meticulously gathered from diverse resources to ensure a rich and varied representation of spoken… See the full description on the dataset page: https://huggingface.co/datasets/ayoubkirouane/darija-stt-mix.audioautomatic-speech-recognition10K<n<100K6 likes196 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.