Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01openslr /librispeech_asr Dataset Card for librispeech_asr Dataset Summary LibriSpeech is a corpus of approximately 1000 hours of 16kHz read English speech, prepared by Vassil Panayotov with the assistance of Daniel Povey. The data is derived from read audiobooks from the LibriVox project, and has been carefully segmented and aligned. Supported Tasks and Leaderboards automatic-speech-recognition, audio-speaker-identification: The dataset can be used to train a model for Automatic… See the full description on the dataset page: https://huggingface.co/datasets/openslr/librispeech_asr.audioautomatic-speech-recognition100K<n<1M246 likes57k downloads1y agoHugging Face02huseyin-karaca /hit-asr HIT-ASR — data and results The data behind HIT-ASR: Hierarchical Transformer Routing for Adaptive ASR Expert Selection (Huseyin Karaca, A. Samil Namli, Suleyman S. Kozat — Bilkent University): what the pretrained ASR experts of the paper produce on its four English corpora, and the stored results every notebook of the code repository reads. Code and notebooks: github.com/huseyin-karaca/hit-asr Documentation: huseyin-karaca.github.io/hit-asr What is here… See the full description on the dataset page: https://huggingface.co/datasets/huseyin-karaca/hit-asr.audioautomatic-speech-recognition100K<n<1M0 likes14k downloads9d agoHugging Face03nithinraok /asr-leaderboard-datasets ASR Leaderboard Datasets This repository contains test splits from multiple speech corpora, including FLEURS, Common Voice (MCV), and Multilingual LibriSpeech (MLS). How to Load To load a specific subset, use load_dataset with the corresponding config_name in the format <set>_<lang>. from datasets import load_dataset # Load the FLEURS dataset for Bulgarian fleurs_bg = load_dataset("nithinraok/asr-leaderboard-datasets", "fleurs_bg") print(fleurs_bg) # Load the MCV… See the full description on the dataset page: https://huggingface.co/datasets/nithinraok/asr-leaderboard-datasets.audioautomatic-speech-recognition100K<n<1M4 likes5.4k downloads1y agoHugging Face04aman-hf /indic_asr Indic ASR Unified Dataset Unified collection of Indian language ASR datasets for pretraining. Stats Total hours: 10,278 Total samples: 4,732,705 Languages: 1 Audio: 16kHz mono (mixed flac/mp3/wav) Languages Language Hours Samples hi2 10,278 4,732,705 Usage from datasets import load_dataset # Load all languages (streaming) ds = load_dataset("aman-hf/indic_asr", streaming=True, split="train") # Load specific language ds_hi =… See the full description on the dataset page: https://huggingface.co/datasets/aman-hf/indic_asr.automatic-speech-recognition10M<n<100M0 likes4.9k downloads7mo agoHugging Face05zihan-audio /librispeech_asr Dataset Card for librispeech_asr Dataset Summary LibriSpeech is a corpus of approximately 1000 hours of 16kHz read English speech, prepared by Vassil Panayotov with the assistance of Daniel Povey. The data is derived from read audiobooks from the LibriVox project, and has been carefully segmented and aligned. Supported Tasks and Leaderboards automatic-speech-recognition, audio-speaker-identification: The dataset can be used to train a model for… See the full description on the dataset page: https://huggingface.co/datasets/zihan-audio/librispeech_asr.audioautomatic-speech-recognition100K<n<1M0 likes4.6k downloads1mo agoHugging Face06Digital-Divide-Data /Luhya-ASR-Data-subset-642H Luhya ASR Data Subset 642H Luhya speech dataset for automatic speech recognition. audioautomatic-speech-recognition100K<n<1M1 likes4.1k downloads2mo agoHugging Face07facebook /omnilingual-asr-corpus Meta Omnilingual ASR Corpus The Omnilingual ASR Corpus is a collection of spontaneous speech recordings and their transcriptions for 348 under-served languages. The corpus was collected as part of Meta FAIR’s Omnilingual ASR project (blog, model, paper) for the purposes of training automatic speech recognition (ASR) and spoken language identification models. Data schema { `language`: "lij_Latn", `iso_639_3`: "lij", `iso_15924`: "Latn", `glottocode`:… See the full description on the dataset page: https://huggingface.co/datasets/facebook/omnilingual-asr-corpus.audioautomatic-speech-recognition100K<n<1M213 likes3.9k downloads11mo agoHugging Face08Digital-Divide-Data /Somali-ASR-Subset-68H Somali ASR Subset 68H Somali speech dataset for automatic speech recognition. audioautomatic-speech-recognition100K<n<1M5 likes3.6k downloads2mo agoHugging Face09RyeAI /danish-asr-leaderboard Open Danish ASR Leaderboard — Results Benchmark results backing the Open Danish ASR Leaderboard — an open, reproducible comparison of Danish speech-to-text models. Every model is transcribed and scored identically on the same five independent public Danish test sets, so the numbers compare directly: Word Error Rate (WER) and Character Error Rate (CER) — lower is better — plus speed. Open-weight Danish speech recognition models you can run yourself and hosted transcription APIs… See the full description on the dataset page: https://huggingface.co/datasets/RyeAI/danish-asr-leaderboard.tabularautomatic-speech-recognition1M<n<10M5 likes3.3k downloads3h agoHugging Face10Digital-Divide-Data /Gusii-ASR-Data-Subset-470H Gusii ASR Data Subset 470H Gusii speech dataset for automatic speech recognition. audioautomatic-speech-recognition100K<n<1M0 likes2.9k downloads2mo agoHugging Face11syvai /danish-asr-unified Danish ASR Unified Dataset Unified Danish speech recognition dataset combining 7 sources (~3.5M samples, ~16k hours): Source Samples Description VoxPopuli 1,775,578 European Parliament recordings ftspeech 995,677 Danish Parliament (Folketinget) CoRal-v3 read_aloud 299,255 Read-aloud Danish speech nst-da 182,605 NST Danish speech CoRal-v3 conversation 147,249 Conversational Danish speech nota 98,600 Danish broadcast media Common Voice 17 3,484 Crowd-sourced… See the full description on the dataset page: https://huggingface.co/datasets/syvai/danish-asr-unified.audioautomatic-speech-recognition1M<n<10M6 likes2.8k downloads2mo agoHugging Face12Vinidapooh /indic-dialect-asr Indic Dialect ASR Dataset A multilingual ASR dataset covering 30 Indic dialect/languages with 2.8M+ samples. Usage from datasets import load_dataset # Load a specific language ds = load_dataset("grushaaaaa/indic-dialect-asr", "assamese", split="train") Features audio: 16kHz WAV audio sentence: Transcription text language: Language name source: Source dataset audioautomatic-speech-recognition1M<n<10M0 likes2.7k downloads28d agoHugging Face13Digital-Divide-Data /Kamba-ASR-Data-Subset-484H Kamba ASR Data Subset 484H Kamba speech dataset for automatic speech recognition. audioautomatic-speech-recognition100K<n<1M2 likes2.6k downloads2mo agoHugging Face14grushaaaaa /indic-dialect-asr Indic Dialect ASR Dataset A multilingual ASR dataset covering 30 Indic dialect/languages with 2.8M+ samples. Usage from datasets import load_dataset # Load a specific language ds = load_dataset("grushaaaaa/indic-dialect-asr", "assamese", split="train") Features audio: 16kHz WAV audio sentence: Transcription text language: Language name source: Source dataset audioautomatic-speech-recognition1M<n<10M4 likes2.5k downloads8mo agoHugging Face15Olicorne /UltiMed-ASR-FR-v1 UltiMed-ASR-FR-v1 [!IMPORTANT] Want faster improvements? This project is entirely self-funded on my minimum-wage salary, and every training run competes for a single consumer GPU. If you or your organisation can donate an RTX 5090, or the money to buy one second-hand, it would directly speed up the next versions of the dataset and the fine-tuned models. Reach out via olicorne.org or open a discussion on this page. [!TIP] More improvements are planned for October 2026, stay… See the full description on the dataset page: https://huggingface.co/datasets/Olicorne/UltiMed-ASR-FR-v1.audioautomatic-speech-recognition100K<n<1M0 likes2.4k downloads5d agoHugging Face16opedromartins /ASR-datasets-ptbr 📚 Datasets de Áudio em Português (PT-BR) Este repositório reúne diversos corpora públicos de fala em português do Brasil, combinados em um único dataset para facilitar treinamentos e pesquisas em ASR (Automatic Speech Recognition). O objetivo é fornecer um recurso amplo, padronizado e de fácil acesso para a comunidade. 📂 Datasets Integrados A tabela abaixo lista todos os datasets incluídos, com suas informações: Dataset Config Name TOTAL train test validation… See the full description on the dataset page: https://huggingface.co/datasets/opedromartins/ASR-datasets-ptbr.audioautomatic-speech-recognition1M<n<10M13 likes2.3k downloads1y agoHugging Face17syvai /danish-asr-unified-hviske-v5-tinygated danish-asr-unified — two-model labels and a quality manifest Transcriptions, per-token confidences, and a per-row quality verdict for every row of syvai/danish-asr-unified (3,414,589 rows, 8 sources). Two independently trained models labelled the whole corpus: model architecture vocabulary syvai/hviske-v5-tiny encoder-decoder 16,384 BPE 3dio-ai/svale-110M RNN-T (Parakeet) 44 characters Both decode greedily. Shards mirror the source data/train-*.parquet by name… See the full description on the dataset page: https://huggingface.co/datasets/syvai/danish-asr-unified-hviske-v5-tiny.tabularautomatic-speech-recognition1M<n<10M0 likes2.1k downloads24d agoHugging Face18ghanaopenai /twi-health-asr-gemini-500hrs This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Twi Health Speech Dataset Gemini (500 hours) A domain-specific speech recognition dataset for Twi, one of Ghana's most widely spoken languages, sourced from publicly available video content on health and wellness. Created by… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/twi-health-asr-gemini-500hrs.audioautomatic-speech-recognition10K<n<100K0 likes1.8k downloads2mo agoHugging Face19hf-audio /open-asr-leaderboard-multilingual-datasets ASR Leaderboard Datasets This repository contains test splits from multiple speech corpora, including FLEURS, Common Voice (MCV), and Multilingual LibriSpeech (MLS). How to Load To load a specific subset, use load_dataset with the corresponding config_name in the format <set>_<lang>. from datasets import load_dataset # Load the FLEURS dataset for Bulgarian fleurs_bg = load_dataset("nithinraok/asr-leaderboard-datasets", "fleurs_bg") print(fleurs_bg) # Load the… See the full description on the dataset page: https://huggingface.co/datasets/hf-audio/open-asr-leaderboard-multilingual-datasets.audioautomatic-speech-recognition100K<n<1M4 likes1.5k downloads3mo agoHugging Face20VoiceArena /MonsoonASR-Open-ASR-leaderboard-en-IN Voice Arena Monsoon en-IN (public test) Part of the Open ASR Leaderboard, in the main board's default column set, so it contributes to the headline Average WER for every model listed. A conversational Indian English ASR test set that records who was speaking, not only what was said. Every clip carries twelve speaker attributes — gender, age, native district and state, education, occupation, income band, handset — so a difference between two systems can be traced to a group of… See the full description on the dataset page: https://huggingface.co/datasets/VoiceArena/MonsoonASR-Open-ASR-leaderboard-en-IN.audioautomatic-speech-recognition1K<n<10K4 likes1.3k downloads1mo agoHugging Face21ekacare /eka-medical-asr-evaluation-dataset Eka Medical ASR Evaluation Dataset Dataset Overview and Sourcing The Eka Medical ASR Evaluation Dataset enables comprehensive evaluation of automatic speech recognition systems designed to transcribe medical speech into accurate text—a fundamental component of any medical scribe system. This dataset captures the unique challenges of processing medical terminology, particularly branded drugs, which is specific to the Indian context. The dataset comprises over 3,900+… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/eka-medical-asr-evaluation-dataset.audioautomatic-speech-recognition1K<n<10K19 likes1.2k downloads1y agoHugging Face22aman4014 /translated-german-english-asr Translated German-English ASR Dataset A large-scale, multi-source German speech dataset with paired English translations, designed for training and evaluating German Automatic Speech Recognition (ASR), Speech Translation, and Text-to-Speech (TTS) systems. This dataset is a curated mixture of well-established open-source German and multilingual speech corpora, all unified under a common schema with German audio, original German transcriptions, and English translations.… See the full description on the dataset page: https://huggingface.co/datasets/aman4014/translated-german-english-asr.audioautomatic-speech-recognition1M<n<10M4 likes1.2k downloads5mo agoHugging Face23flozi00 /german-asr-mixed-whisper Dataset Card Dataset Sources and Licensing This dataset is a mixture of several German and multilingual speech datasets. For each dataset, the license of the original author applies. Please consult the linked sources for detailed licensing information and terms of use. Dataset Name Original Source / Author Link TUDA-De German Speech Corpus LT Group at UHH / TU Darmstadt https://huggingface.co/datasets/uhhlt/Tuda-De Mozilla Common Voice Mozilla Foundation… See the full description on the dataset page: https://huggingface.co/datasets/flozi00/german-asr-mixed-whisper.audioautomatic-speech-recognition1M<n<10M4 likes1.2k downloads1y agoHugging Face24ghanaopenai /ghana-english-asr-2700hrs This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. 🇬🇭 Ghana English ASR Dataset A speech dataset of Ghanaian English extracted from Ghanaian news media broadcasts, designed for training and fine-tuning Automatic Speech Recognition (ASR) models on West African English accents.… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-english-asr-2700hrs.audioautomatic-speech-recognition100K<n<1M8 likes1.2k downloads4mo agoHugging Face25Perle-ai /ASR_Code_Switch ASR Code-Switching Benchmark A curated benchmark of 1,200 code-switching utterances (300 per language pair) for evaluating commercial ASR systems on multilingual speech with intra-sentential language switching. Paper Benchmarking Commercial ASR Systems on Code-Switching Speech: Arabic, Persian, and German arXiv link Language pairs Split Language pair Samples Scripts egyptian_arabic_english Egyptian Arabic–English 300 Arabic + Latin… See the full description on the dataset page: https://huggingface.co/datasets/Perle-ai/ASR_Code_Switch.audioautomatic-speech-recognition1K<n<10K12 likes1.1k downloads5mo agoHugging Face26KathleenKunLiu /omnilingual-asr-corpus Meta Omnilingual ASR Corpus The Omnilingual ASR Corpus is a collection of spontaneous speech recordings and their transcriptions for 348 under-served languages. The corpus was collected as part of Meta FAIR’s Omnilingual ASR project (blog, model, paper) for the purposes of training automatic speech recognition (ASR) and spoken language identification models. Data schema { `language`: "lij_Latn", `iso_639_3`: "lij", `iso_15924`: "Latn"… See the full description on the dataset page: https://huggingface.co/datasets/KathleenKunLiu/omnilingual-asr-corpus.audioautomatic-speech-recognition100K<n<1M0 likes1k downloads4mo agoHugging Face27Digital-Divide-Data /khm-asr-cultural Khmer ASR Cultural Dataset 134.6 hours manually curated speech-text pairs by native speakers in Khmer language about Cambodian cultural topics. On average, each recording is 8.54 seconds with the standard deviation of 3.37. Speaker metadata (gender, age group, and origin city) is provided. Language: Khmer (khm). Source(s): Native speakers from Cambodia (4 females, 4 males). The utterances were manually generated based on topics and subtopics listed in metadata. Domain(s):… See the full description on the dataset page: https://huggingface.co/datasets/Digital-Divide-Data/khm-asr-cultural.audioautomatic-speech-recognition10K<n<100K9 likes853 downloads6mo agoHugging Face28carlosdanielhernandezmena /ravnursson_asr Dataset Card for ravnursson_asr Dataset Summary The corpus "RAVNURSSON FAROESE SPEECH AND TRANSCRIPTS" (or RAVNURSSON Corpus for short) is a collection of speech recordings with transcriptions intended for Automatic Speech Recognition (ASR) applications in the language that is spoken at the Faroe Islands (Faroese). It was curated at the Reykjavík University (RU) in 2022. The RAVNURSSON Corpus is an extract of the "Basic Language Resource Kit 1.0" (BLARK 1.0) [1] developed… See the full description on the dataset page: https://huggingface.co/datasets/carlosdanielhernandezmena/ravnursson_asr.audioautomatic-speech-recognition10K<n<100K3 likes828 downloads1y agoHugging Face29LocalDoc /azerbaijani_asr Azerbaijani ASR Dataset Dataset Description This dataset contains Azerbaijani speech data for Automatic Speech Recognition (ASR) tasks. Dataset Summary Language: Azerbaijani (az) Task: Automatic Speech Recognition Total Duration: ~328 hours Total Samples: ~345,643 audio-text pairs Audio Format: WAV, 16kHz sampling rate License: CC-BY-4.0 Dataset Structure Each audio segment is specially numbered so that you can merge them if you… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/azerbaijani_asr.audioautomatic-speech-recognition100K<n<1M5 likes808 downloads3mo agoHugging Face30MTSAIR /mws-reson-asr-osd MWS-RESON-ASR-OSD 🇷🇺 Русскоязычное описание ниже / Russian summary below. RESON ASR OSD is a Russian-language benchmark for automatic speech recognition in the telephone channel. The release contains 12,114 recordings and 37 hours 52 minutes of 8 kHz audio. It is an evaluation set: every configuration has a single test split. The benchmark has two domains, each in two acoustic conditions: 🗣️ General — conversational telephone speech: support requests, automated… See the full description on the dataset page: https://huggingface.co/datasets/MTSAIR/mws-reson-asr-osd.audioautomatic-speech-recognition10K<n<100K2 likes791 downloads10d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.