Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nineninesix /multilingual-tts-benchmark Multilingual Speech Benchmark for Zero-Shot TTS A voice-cloning and intelligibility benchmark for 8 language subsets, with 10,100 examples selected from Common Voice 17.0. Each example supplies a speaker reference and an independently selected target text, with human recordings as WER/CER and speaker-similarity anchors when the corresponding audio is available. Corpus WER and the WavLM-FT evaluation follow seed-tts-eval. Version 3.1. Adds ru, kk through the same S1–S4 selection… See the full description on the dataset page: https://huggingface.co/datasets/nineninesix/multilingual-tts-benchmark.audiotext-to-speech100K<n<1M0 likes1.8k downloads4d agoHugging Face02Krisp-AI /VoiceIsolation-Benchmark-Dataset Voice Isolation Benchmark Dataset 265 real-world recordings for measuring how a second voice breaks speech-to-text, and how much Krisp Voice Isolation fixes it. Three scenarios, 47 speakers, real rooms, real headsets. No synthetic mixing. 265 recordings · 47 speakers · 3 scenarios · 65 scripts Why this dataset exists Modern STT engines handle noise well. They still fail when a second person talks near the microphone: they transcribe the wrong speaker, and voice… See the full description on the dataset page: https://huggingface.co/datasets/Krisp-AI/VoiceIsolation-Benchmark-Dataset.audioautomatic-speech-recognitionn<1K10 likes1.4k downloads10d agoHugging Face03Quran-Lab /quranic-asr-benchmarkgated Quranic ASR Benchmark - leakage-free, held-out A small, leakage-free benchmark (600 clips) for evaluating Arabic ASR on Quranic recitation (Hafs riwayah). Every clip is verified absent from our training data, so it measures generalization, not memorization. Same clips + same scoring for every model. 📊 Live leaderboard: https://huggingface.co/spaces/Muno459/quranic-asr-leaderboard The set (600 clips, 200 per source) Source n What it is everyayah_heldout… See the full description on the dataset page: https://huggingface.co/datasets/Quran-Lab/quranic-asr-benchmark.audioautomatic-speech-recognitionn<1K5 likes666 downloads2mo agoHugging Face04MR3z4 /persian-accents-benchmark Persian Accents Benchmark Dataset Summary A benchmark for Persian automatic speech recognition (ASR): 279 short utterances of informal Persian (Farsi) dialect speech across 16 regional accents, released as a fixed evaluation set. Total audio duration is approximately 4.4 hours. The primary label is the transcription; each utterance also carries an accent label (usable for accent classification as a secondary task) and an emotion label as auxiliary metadata. This… See the full description on the dataset page: https://huggingface.co/datasets/MR3z4/persian-accents-benchmark.audioautomatic-speech-recognitionn<1K2 likes450 downloads2mo agoHugging Face05QUD-Technologies /quran-alignment-benchmark Quran Recitation Alignment Benchmark Audio recordings of Quran recitation with a reviewed word-level ground truth: every recited word, in the order it was recited, with its start and end time, plus the reviewed segmentation and non-Quran regions. This is the corpus behind the Quran Recitation Alignment Benchmark; the task, scoring rules, leaderboard and submission format are documented there, not here. 16 recordings · 357 minutes · 18,421 recited words · Hafs ʿan ʿĀṣim ·… See the full description on the dataset page: https://huggingface.co/datasets/QUD-Technologies/quran-alignment-benchmark.audioautomatic-speech-recognitionn<1K0 likes398 downloads6d agoHugging Face06modulate /entity-transcription-benchmark Entity Transcription Benchmark Measures whether a speech recognition system transcribes named entities correctly — as distinct from word error rate. WER weights every token equally. The tokens that matter for redaction, lookup, routing and search are proper nouns, and they are a small fraction of any transcript. A system can improve WER while getting worse at exactly the words a downstream consumer needs, and nothing in the standard evaluation will show it. 2,151 clips, 6.0… See the full description on the dataset page: https://huggingface.co/datasets/modulate/entity-transcription-benchmark.audioautomatic-speech-recognition1K<n<10K4 likes373 downloads29d agoHugging Face07DigiGreen /Agri_STT_Benchmarking_Dataset Agri STT Benchmarking Dataset 10,808 farmer voice queries in Hindi, Telugu and Odia, with reference transcripts, for benchmarking automatic speech recognition in agricultural contexts. The audio is included in this repository. Every recording is a smallholder farmer speaking a question to Farmer.Chat, an AI advisory service run by Digital Green. Reference transcripts were produced by human annotators. Nothing here is read from a script or recorded in a studio, so the audio… See the full description on the dataset page: https://huggingface.co/datasets/DigiGreen/Agri_STT_Benchmarking_Dataset.audioautomatic-speech-recognition10K<n<100K3 likes260 downloads3mo agoHugging Face08ARTPARK-IISc /Vaani-Benchmark-V1.0gated Vaani-Benchmark-V1.0 A curated Hindi ASR evaluation set collected as part of the Vaani project at IISc Bengaluru. This is a separate, held-out collection — distinct from the publicly released Vaani dataset — built specifically for benchmarking. This benchmark contains 5,050 audio segments from 1,103 speakers across 104 Indian districts, each with three independent human transcriptions. Evaluation Toolkit A standalone toolkit implementing this benchmark's scoring… See the full description on the dataset page: https://huggingface.co/datasets/ARTPARK-IISc/Vaani-Benchmark-V1.0.audioautomatic-speech-recognition1K<n<10K7 likes228 downloads5d agoHugging Face09Kimyayd /vocal-money-codeswitch-asr-benchmark Vocal Money — Yoruba–English Code-Switched ASR Benchmark A 210-clip evaluation subset used to benchmark five speech recognition systems on naturally code-switched Yoruba–English speech, together with the reference transcriptions and the output of every system on every clip, so that the published results can be recomputed or contradicted. Produced for the MLC (Africa) × Intron Agentic Voice AI Challenge, Deep Learning Indaba 2026. Team Vocal Money — Hospice Hounfodji, Mohamed… See the full description on the dataset page: https://huggingface.co/datasets/Kimyayd/vocal-money-codeswitch-asr-benchmark.audioautomatic-speech-recognitionn<1K0 likes162 downloads2mo agoHugging Face10sophia8888 /clipquill-asr-benchmark Measuring whisper-tiny vs whisper-base in a browser tab Word error rate, wall-clock timing, transfer size and peak memory for two quantised Whisper tiers running entirely client-side in a real Chrome window, with the scripts that produced every number. If you are building an in-browser transcription page, the two results worth knowing before you pick a model tier: On clean synthetic audio the two tiers tie. If that is all you test, you will conclude the tier does not matter… See the full description on the dataset page: https://huggingface.co/datasets/sophia8888/clipquill-asr-benchmark.tabularautomatic-speech-recognitionn<1K0 likes152 downloads22d agoHugging Face11woodygan /humans-benchmark HUMANS Benchmark Dataset Authors: Woody Haosheng Gan¹, William Held²'³, Diyi Yang² ¹University of Southern California, ²Stanford University, ³OpenAthena This dataset is part of the Putting HUMANS first: Efficient LAM Evaluation with Human Preference Alignment paper. HUMANS (HUman-aligned Minimal Audio evaluatioN Subsets for Large Audio Models) Benchmark is designed to efficiently evaluate Large Audio Models using minimal subsets while predicting human preferences through learned… See the full description on the dataset page: https://huggingface.co/datasets/woodygan/humans-benchmark.audioaudio-classificationn<1K0 likes133 downloads27d agoHugging Face12C1Tech /Persian-ASR-BenchmarkThis dataset consists of 3 hours of 16kHz audio collected from diverse environments to better represent real-world scenarios. The recordings were sourced from audiobooks, YouTube, and other public sources, ensuring a wide variety of speech styles and acoustic conditions. One key advantage of this dataset is that it was collected from recent sources within the last few months, ensuring no overlap with training data and fairness for evaluating other STT models. To enable a robust and fair… See the full description on the dataset page: https://huggingface.co/datasets/C1Tech/Persian-ASR-Benchmark.audioautomatic-speech-recognition1K<n<10K4 likes128 downloads3mo agoHugging Face13yusasif /intron-stt_tts-benchmark Code-switched benchmark audio The exact 16 recordings used to benchmark speech models for the Sahara CodeSwitch Africa Challenge. Published so the reported numbers can be checked against the audio that produced them. Every clip is intra-sentential code-switching — one speaker moving between a Nigerian language and English inside a single utterance, which is the case the challenge is judged on. Language Clips Hausa 4 Igbo 4 Nigerian Pidgin 4 Yoruba 4… See the full description on the dataset page: https://huggingface.co/datasets/yusasif/intron-stt_tts-benchmark.audioautomatic-speech-recognitionn<1K0 likes125 downloads26d agoHugging Face14SaarAI /asr-benchmark-outputsgated SaarAI ASR Benchmark Outputs Raw per-utterance model outputs (transcription manifests) produced by the gsma-asr-bench runners on SaarAI/asr-leaderboard-datasets. files: 570 utterances: 4615160 languages: 7 models: 50 Layout data/<language_name>/<split>__<dataset_config>__<model_slug>.jsonl index.jsonl # one record per file (language, split, model, rows, sha256, ...) index.csv Directories categorise by language name; the file name begins with the split name… See the full description on the dataset page: https://huggingface.co/datasets/SaarAI/asr-benchmark-outputs.tabularautomatic-speech-recognition1M<n<10M1 likes113 downloads5d agoHugging Face15NightPrince /Fasih-TTS-Benchmark Fasih-TTS-Benchmark Evaluation audio and objective scores for the Fasih-TTS-V1 Arabic (MSA / Fusha) text-to-speech model. Every clip is the model's own output, paired with its reference text, its Whisper-large-v3 transcription, and per-clip WER / CER. Contents Split (test_set) Clips Purpose Mean CER silma_msa 10 SILMA open-source Arabic TTS benchmark (MSA sentences) 2.0% samples 3 General showcase (greeting, fiqh, reflection) 0.6% consistency 8… See the full description on the dataset page: https://huggingface.co/datasets/NightPrince/Fasih-TTS-Benchmark.audiotext-to-speechn<1K2 likes88 downloads3mo agoHugging Face16AirCaps /a5sv2-asr-benchmark-dataset A5Sv2 ASR Benchmark Dataset Public references, saved predictions, scores, and provenance for the A5Sv2 ASR benchmark. The benchmark evaluates streaming English ASR on four fixed public corpora with approximately equal normalized reference word counts. Corpus Fixed selection Reference words Audio in this repository Mega-ASR / Voices-in-the-Wild-2M 1,250 utterances, 250 per acoustic condition 32,928 Yes AMI 7 scenario-only unseen-evaluation meetings 32,928 Yes DiPCo… See the full description on the dataset page: https://huggingface.co/datasets/AirCaps/a5sv2-asr-benchmark-dataset.audioautomatic-speech-recognition1K<n<10K0 likes88 downloads1mo agoHugging Face17gametime-benchmark /gametime Gametime Benchmark The Gametime dataset provides lightweight, streaming-friendly splits for TTS/ASR/SpokenLM prototyping.For full details, please refer to the paper:👉 Game-Time: Evaluating Temporal Dynamics in Spoken Language Models 📦 Download Options 1️⃣ Recommended — Full ZIP Download If you prefer the original folder layout you can download one of the ZIPs packaged in gametime/download/. There are two kinds available in this repository:… See the full description on the dataset page: https://huggingface.co/datasets/gametime-benchmark/gametime.audioautomatic-speech-recognition1K<n<10K2 likes86 downloads5mo agoHugging Face18themechanism /script-fidelity-benchmark Script fidelity benchmark Anonymous supplement for the paper "Script collapse in multilingual ASR: A reference-free metric and 100-pair benchmark." Script Fidelity Rate (SFR) measures the fraction of ASR hypothesis characters that belong to the expected target script. WER measures word edits, while SFR checks whether the output is written in the target orthography. Related resources: PyPI package: https://pypi.org/project/script-fidelity/ Hugging Face Evaluate metric:… See the full description on the dataset page: https://huggingface.co/datasets/themechanism/script-fidelity-benchmark.tabularautomatic-speech-recognition10K<n<100K0 likes81 downloads5mo agoHugging Face19Revolab /ASR-Benchmark-Publicgated Revolab ASR Benchmark (Public Split) The Revolab ASR Benchmark is a human-annotated evaluation dataset for Malaysian Malay speech recognition. This is the public split - a downloadable subset that anyone can use to run their own evaluation. A larger private split is used to maintain the official leaderboard. Dataset Details Audio samples: 820 Duration: ~2.2 hours Languages: Bahasa Malaysia and English (including code-switching) Sample rate: 16kHz License:… See the full description on the dataset page: https://huggingface.co/datasets/Revolab/ASR-Benchmark-Public.audioaudio-to-audion<1K3 likes76 downloads2mo agoHugging Face20Fat-Shork /nepali-homophone-asr-benchmark Nepali Confusable-Pair Voice Dataset A small research dataset for studying Nepali speech-to-text (ASR) errors on context-dependent confusable / homophonic words. Dataset size 54 total WAV recordings preserved 50 recordings provisionally mapped to canonical sentence prompts 3 valid speech recordings intentionally left without a transcript because their exact prompt is unresolved 1 intentional near-silent recording retained as a no-speech / hallucination challenge… See the full description on the dataset page: https://huggingface.co/datasets/Fat-Shork/nepali-homophone-asr-benchmark.audioautomatic-speech-recognitionn<1K0 likes74 downloads10d agoHugging Face21aranemini /kurdish-multidialect-asr-benchmark Kurdish Dialect Speech Corpus This project aims to provide a multi-dialect speech recognition benchmark for the Kurdish language. The Central Kurdish portion is the same as the Asosoft benchmark. The sentences were originally written in Central Kurdish (CKB), translated into other Kurdish dialects, and then recorded by native speakers. The current version includes three Kurdish dialects: Central Kurdish, Northern Kurdish, and Southern Kurdish. A Hawrami version and the Badini… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/kurdish-multidialect-asr-benchmark.audioautomatic-speech-recognition1K<n<10K0 likes59 downloads2mo agoHugging Face22RinggAI /ASR-Benchmarking-Dataset Hindi STT Benchmarking Eval Overview This dataset packages the Hindi eval split used for STT benchmarking across six Vistaar-derived parts: IndicTTS, FLEURS, CommonVoice, Kathbath, Kathbath noisy, and MUCS. Each row contains the audio, original reference transcript, and raw plus normalized transcripts from Ringg, ElevenLabs, Deepgram, and Sarvam. The dataset contains 10,000 utterances and about 15.5 hours of 16 kHz mono WAV audio. The dataset is published as part-specific… See the full description on the dataset page: https://huggingface.co/datasets/RinggAI/ASR-Benchmarking-Dataset.audioautomatic-speech-recognition10K<n<100K1 likes58 downloads5mo agoHugging Face23zadterishi /darija-asr-benchmark-6speaker Darija ASR 6-Speaker Benchmark A fixed, paired 20-utterance benchmark read identically by 6 held-out speakers (3 female: F1, F2, F3; 3 male: M1, M2, M3 -- none present in any training corpus), used to evaluate cross-speaker generalization for a Moroccan Darija (Arabizi) Whisper fine-tuning pipeline (paper not yet published -- citation forthcoming). Consent and anonymization Written informed consent was obtained from all six speakers for the recording and… See the full description on the dataset page: https://huggingface.co/datasets/zadterishi/darija-asr-benchmark-6speaker.audioautomatic-speech-recognitionn<1K0 likes56 downloads16d agoHugging Face24abnajlae /darija-asr-benchmark-6speaker Darija ASR 6-Speaker Benchmark A fixed, paired 20-utterance benchmark read identically by 6 held-out speakers (3 female: F1, F2, F3; 3 male: M1, M2, M3 -- none present in any training corpus), used to evaluate cross-speaker generalization for a Moroccan Darija (Arabizi) Whisper fine-tuning pipeline (paper not yet published -- citation forthcoming). Consent and anonymization Written informed consent was obtained from all six speakers for the recording and… See the full description on the dataset page: https://huggingface.co/datasets/abnajlae/darija-asr-benchmark-6speaker.audioautomatic-speech-recognitionn<1K0 likes54 downloads1mo agoHugging Face25AWANNABY /French-Medical-Transcription-Benchmark 🩺 French Medical Transcription Evaluation Dataset Ce dataset a été créé et ouvert à la communauté dans le cadre du développement R&D de LucioleScribe, la plateforme souveraine de transcription IA 100% locale, spécifiquement conçue pour les milieux médicaux et juridiques (compatibilité RGPD, HDS, et architectures Air-Gapped). 🔗 Découvrir LucioleScribe Édition Santé | ⚙️ Voir le Pipeline Technologique Local 📊 Présentation du Dataset L'évaluation des modèles de… See the full description on the dataset page: https://huggingface.co/datasets/AWANNABY/French-Medical-Transcription-Benchmark.textautomatic-speech-recognition1K<n<10K1 likes53 downloads7mo agoHugging Face26Reza2kn /visualears-benchmark-269-gold 🗂️ visualears-benchmark-269-gold English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission 🌟 At a glance | معرفی سریع English فارسی 🎯 Purpose 269-record gold/noisy benchmark dataset. معیار طلایی ۲۶۹ نمونه‌ای برای بررسی سریع خطاهای گفتار نویزی و مقایسهٔ نسخه‌های مدل. 🧩 Role evaluation and benchmarking asset مصنوع ارزیابی و بنچمارک 📦 Snapshot 276 files; approximately 44.35 MB 276 فایل؛ حدود 44.35 MB 🧱 Packaging 1 Parquet… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/visualears-benchmark-269-gold.audioautomatic-speech-recognitionn<1K1 likes53 downloads2mo agoHugging Face27sudoping01 /bam-asr-benchmark Bambara ASR Benchmark The first standardized evaluation set for Automatic Speech Recognition in Bambara (Bamanankan). One hour of studio-quality constitutional text, transcribed and validated by linguists from Mali's Direction Nationale de l'Éducation Non Formelle et des Langues Nationales (DNENF-LN). This benchmark accompanies the paper "Where Are We at with Automatic Speech Recognition for the Bambara Language?" and the public leaderboard at MALIBA-AI/bambara-asr-leaderboard.… See the full description on the dataset page: https://huggingface.co/datasets/sudoping01/bam-asr-benchmark.audioautomatic-speech-recognitionn<1K0 likes50 downloads8mo agoHugging Face28Squadstack /conversational-streaming-asr-benchmark SquadStack Conversational Streaming ASR Benchmark (8 kHz) Version 1.0.0 · maintained by SquadStack Schema · Leaderboard · Latency · Submit a system · Licence · Terms of use Key takeaways What this is. 863 real Hindi–English telesales calls (5.53 hours of customer speech, 8 kHz phone audio), human-transcribed turn by turn, and 11 speech recognisers scored on them. The question it answers: which recogniser should run inside an Indian voice agent, judged on… See the full description on the dataset page: https://huggingface.co/datasets/Squadstack/conversational-streaming-asr-benchmark.audioautomatic-speech-recognition1K<n<10K1 likes49 downloads1d agoHugging Face29sumanpaudel1997 /nepali-asr-benchmark Nepali ASR Benchmark Per-utterance reference, hypothesis, WER, and CER for the six released Nepali ASR checkpoints evaluated on three independent test sets. Released alongside the paper Comparative Analysis of Multilingual Pre-trained Models for Nepali Automatic Speech Recognition. Contents Field Type Description utterance_id string stable identifier {test_set}-{index} reference string NFC-normalised gold transcription (Devanagari) hypothesis string… See the full description on the dataset page: https://huggingface.co/datasets/sumanpaudel1997/nepali-asr-benchmark.tabularautomatic-speech-recognition10K<n<100K0 likes44 downloads4mo agoHugging Face30HUMANSBenchmark /humans-benchmark HUMANS Benchmark Dataset (Anonymous, Under Review) This dataset is part of the HUMANS (HUman-aligned Minimal Audio evaluatioN Subsets for Large Audio Models) Benchmark, designed to efficiently evaluate Large Audio Models using minimal subsets while predicting human preferences through learned regression weights. Installation Install the HUMANS evaluation package from GitHub (our anonymous repo): # Option 1: Install via pip pip install… See the full description on the dataset page: https://huggingface.co/datasets/HUMANSBenchmark/humans-benchmark.audioaudio-classificationn<1K0 likes40 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.