Team Ai
27 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01huseyin-karaca /hit-asr HIT-ASR — data and results The data behind HIT-ASR: Hierarchical Transformer Routing for Adaptive ASR Expert Selection (Huseyin Karaca, A. Samil Namli, Suleyman S. Kozat — Bilkent University): what the pretrained ASR experts of the paper produce on its four English corpora, and the stored results every notebook of the code repository reads. Code and notebooks: github.com/huseyin-karaca/hit-asr Documentation: huseyin-karaca.github.io/hit-asr What is here… See the full description on the dataset page: https://huggingface.co/datasets/huseyin-karaca/hit-asr.audioautomatic-speech-recognition100K<n<1M0 likes14k downloads9d agoHugging Face02RyeAI /danish-asr-leaderboard Open Danish ASR Leaderboard — Results Benchmark results backing the Open Danish ASR Leaderboard — an open, reproducible comparison of Danish speech-to-text models. Every model is transcribed and scored identically on the same five independent public Danish test sets, so the numbers compare directly: Word Error Rate (WER) and Character Error Rate (CER) — lower is better — plus speed. Open-weight Danish speech recognition models you can run yourself and hosted transcription APIs… See the full description on the dataset page: https://huggingface.co/datasets/RyeAI/danish-asr-leaderboard.tabularautomatic-speech-recognition1M<n<10M5 likes3.3k downloads6h agoHugging Face03syvai /danish-asr-unified-hviske-v5-tinygated danish-asr-unified — two-model labels and a quality manifest Transcriptions, per-token confidences, and a per-row quality verdict for every row of syvai/danish-asr-unified (3,414,589 rows, 8 sources). Two independently trained models labelled the whole corpus: model architecture vocabulary syvai/hviske-v5-tiny encoder-decoder 16,384 BPE 3dio-ai/svale-110M RNN-T (Parakeet) 44 characters Both decode greedily. Shards mirror the source data/train-*.parquet by name… See the full description on the dataset page: https://huggingface.co/datasets/syvai/danish-asr-unified-hviske-v5-tiny.tabularautomatic-speech-recognition1M<n<10M0 likes2.1k downloads24d agoHugging Face04abnajlae /darija-asr-corpus Darija ASR Corpus (dataset-core) Arabizi (Latin-script) transcriptions of Moroccan Darija speech, produced for a Whisper fine-tuning pipeline (paper not yet published -- citation forthcoming). This repo contains four source subsets: DODa, DVoice, Wiki, and YouTube. Each subset carries its own upstream license/terms -- see below -- because they are drawn from four different original projects. Subsets Config Rows Audio bundled? Upstream license Upstream source… See the full description on the dataset page: https://huggingface.co/datasets/abnajlae/darija-asr-corpus.audioautomatic-speech-recognition10K<n<100K0 likes518 downloads1mo agoHugging Face05syvai /danish-asr-verified danish-asr-verified ALL rows of syvai/danish-asr-unified transcribed by the syv-transcribe ensemble (hviske-v5.3 + hviske-v5, confidence-weighted ROVER), each annotated with: verified — True when the ensemble independently reproduced the reference exactly (compared after lowercasing, punctuation-strip, whitespace-collapse). Two independent witnesses agree => near-certain label. wer_teacher_vs_ref / cer_teacher_vs_ref — word/character error rate between normalized teacher output… See the full description on the dataset page: https://huggingface.co/datasets/syvai/danish-asr-verified.tabularautomatic-speech-recognition1M<n<10M0 likes301 downloads2mo agoHugging Face06Edge0 /ark-asr-3b-open-asr-leaderboard-results ARK-ASR-3B Open ASR Leaderboard Results Raw JSONL manifests for AutoArk-AI/ARK-ASR-3B on the public English short-form hf-audio/open-asr-leaderboard splits. These manifests were generated on a local 8x RTX 4090 machine and scored with the shared Open ASR Leaderboard scorer: PYTHONPATH=. python - <<'PY' from normalizer.eval_utils import score_results score_results( 'ark_asr/results.AutoArk-AI-ARK-ASR-3B_20260622_official', 'AutoArk-AI/ARK-ASR-3B', ) PY Important:… See the full description on the dataset page: https://huggingface.co/datasets/Edge0/ark-asr-3b-open-asr-leaderboard-results.tabularautomatic-speech-recognition10K<n<100K12 likes202 downloads4mo agoHugging Face07benderrodriguez /hebrew-asr-vn Hebrew ASR three-source training dataset Derived from ivrit.ai crowd-transcribe-v5, crowd-recital and knesset-committees. Original data and transcripts are credited to ivrit.ai and its contributors. Pinned revisions and preparation rules are in metadata/sources.json and metadata/preparation-config.json. VoxKnesset is excluded by user decision. Source/split Clips Hours crowd-recital/test 1,557 1.071 crowd-recital/train 45,372 33.258 crowd-recital/validation 1,033… See the full description on the dataset page: https://huggingface.co/datasets/benderrodriguez/hebrew-asr-vn.tabularautomatic-speech-recognition1M<n<10M0 likes188 downloads27d agoHugging Face08sanchit-gandhi /librispeech_asr_dummy Dataset Card for librispeech_asr_dummy Dataset Summary This is a truncated version of the LibriSpeech dataset. It contains 20 samples from each of the splits. To view the full dataset, visit: https://huggingface.co/datasets/librispeech_asr LibriSpeech is a corpus of approximately 1000 hours of 16kHz read English speech, prepared by Vassil Panayotov with the assistance of Daniel Povey. The data is derived from read audiobooks from the LibriVox project, and has been… See the full description on the dataset page: https://huggingface.co/datasets/sanchit-gandhi/librispeech_asr_dummy.audioautomatic-speech-recognitionn<1K0 likes185 downloads3y agoHugging Face09Atika88 /Indonesian-ASR-11-Class-Dataset Indonesian ASR 11-Class Dataset Public Hugging Face repository for an Indonesian ASR corpus and its paper-supporting benchmark artifacts. Dataset summary Audio files: 104,500 WAV files Real/human recordings: 104,368 Synthetic repair files: 132 Sentence classes: 11 Indonesian sentence categories Canonical balanced sentence slots: 209 (11 categories × 19 retained slots) Public speaker labels: M1..M12, F1..F8, plus synthetic labels Ms*/Fs* Audio format: 16 kHz… See the full description on the dataset page: https://huggingface.co/datasets/Atika88/Indonesian-ASR-11-Class-Dataset.tabularautomatic-speech-recognition100K<n<1M0 likes170 downloads1mo agoHugging Face10sophia8888 /clipquill-asr-benchmark Measuring whisper-tiny vs whisper-base in a browser tab Word error rate, wall-clock timing, transfer size and peak memory for two quantised Whisper tiers running entirely client-side in a real Chrome window, with the scripts that produced every number. If you are building an in-browser transcription page, the two results worth knowing before you pick a model tier: On clean synthetic audio the two tiers tie. If that is all you test, you will conclude the tier does not matter… See the full description on the dataset page: https://huggingface.co/datasets/sophia8888/clipquill-asr-benchmark.tabularautomatic-speech-recognitionn<1K0 likes152 downloads22d agoHugging Face11danielrosehill /Tech-Sentences-For-ASR-Training TechVoice Dataset Work in Progress – This dataset is actively being expanded with new recordings. Dataset Statistics Metric Current Target Progress Duration 38m 43s 5h 0m 0s ██░░░░░░░░░░░░░░░░░░ 12.9% Words 10,412 50,000 ████░░░░░░░░░░░░░░░░ 20.8% Total Recordings: 205 samples Total Characters: 74,312 A specialized speech dataset for fine-tuning Automatic Speech Recognition (ASR) models on technical and developer vocabulary. Contains human-recorded… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Tech-Sentences-For-ASR-Training.audioautomatic-speech-recognitionn<1K2 likes120 downloads11mo agoHugging Face12SaarAI /asr-benchmark-outputsgated SaarAI ASR Benchmark Outputs Raw per-utterance model outputs (transcription manifests) produced by the gsma-asr-bench runners on SaarAI/asr-leaderboard-datasets. files: 570 utterances: 4615160 languages: 7 models: 50 Layout data/<language_name>/<split>__<dataset_config>__<model_slug>.jsonl index.jsonl # one record per file (language, split, model, rows, sha256, ...) index.csv Directories categorise by language name; the file name begins with the split name… See the full description on the dataset page: https://huggingface.co/datasets/SaarAI/asr-benchmark-outputs.tabularautomatic-speech-recognition1M<n<10M1 likes113 downloads5d agoHugging Face13Peacockery /tajik-asr-youtube tajik-asr-youtube Every Tajik-language clip from this project's YouTube scrape — 41 channels of news, talk shows, podcasts, audiobooks, and learning content — with machine transcripts and the verification scores left in as columns instead of applied as a filter. Pick your own quality threshold; the training corpus this project actually ships (tajik-asr-corpus-v3) is the gated subset. Layout Parquet shards under data/ with an audio struct column (16 kHz mono FLAC… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/tajik-asr-youtube.tabularautomatic-speech-recognition100K<n<1M0 likes100 downloads4mo agoHugging Face14s512757 /polish-tedx-asr-eval Polish-TEDx-ASR-Eval A dataset for evaluating automatic speech recognition (ASR) systems for Polish in the domain of TEDx public talks. Contains audio segments from Polish TEDx talks available on YouTube (CC BY-NC-ND 4.0) and synthetic speech generated with KugelAudio (MIT), with manually created and cross-verified transcriptions. Created as part of the course "Workshops on Evaluation of Speech Recognition Systems" (ZWESUI, AMU 2026) by Group 1. Statistics… See the full description on the dataset page: https://huggingface.co/datasets/s512757/polish-tedx-asr-eval.audioautomatic-speech-recognitionn<1K0 likes94 downloads4mo agoHugging Face15Reza2kn /persian-asr-text-2.69M-deduped 🗂️ persian-asr-text-2.69M-deduped English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission 🌟 At a glance | معرفی سریع English فارسی 🎯 Purpose Deduplicated Persian ASR text dataset used by the training stack. پیکرهٔ متنی فارسیِ حذف‌تکرارشده برای ساخت واژگان، مدل‌سازی زبانی و پشتیبانی از آموزش ASR. 🧩 Role Persian text and linguistic asset مصنوع متنی و زبانی فارسی 📦 Snapshot 4 files; approximately 109.64 MB 4 فایل؛ حدود 109.64… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-asr-text-2.69M-deduped.tabularautomatic-speech-recognition1M<n<10M0 likes65 downloads2mo agoHugging Face16bengaliAI /ben10-asr-results Ben-10 Regional ASR — public results Score rows for the maintainer-run Ben-10 regional dialect ASR leaderboard. Field Meaning model_id Hub id or slug model_url Link to weights / paper wer Corpus Word Error Rate on private ben-10-test (lower better) wer_by_region JSON map region → WER backend Decode stack used by maintainers scorer_commit / decode_commit Git SHAs in BengaliAI/reg-speech-aacl evaluated_at ISO date requested_by Who asked, or maintainer if… See the full description on the dataset page: https://huggingface.co/datasets/bengaliAI/ben10-asr-results.tabularautomatic-speech-recognitionn<1K0 likes61 downloads3mo agoHugging Face17kvest /Swedia-ASR-Dataset Swedia ASR Dataset This repository contains a small Swedish ASR evaluation dataset based on speech transcriptions from Swedia 2000. It was assembled to compare automatic speech-recognition output against manually corrected reference transcriptions for Swedish dialectal speech. The dataset is useful for quick experiments with Swedish ASR systems, especially when you want to inspect recognition quality on spontaneous speech from different regions, speakers, ages, and genders.… See the full description on the dataset page: https://huggingface.co/datasets/kvest/Swedia-ASR-Dataset.tabularautomatic-speech-recognitionn<1K1 likes46 downloads5mo agoHugging Face18nhatminh /korean-asr korean-asr — Korean ASR pseudo-labels for YODAS2 This repository contains transcripts and segment metadata only. It does not contain audio. Every row points into espnet/yodas2 by (shard, video_id, start, end), so you fetch the audio from the upstream dataset and cut it yourself. See Reconstructing the audio. split utterances hours train 1,034,181 6,974.3 heldout 46,542 314.6 dev (subset of heldout) 3,000 20.4 Total labelled: 1,080,723 utterances / 7,288.9… See the full description on the dataset page: https://huggingface.co/datasets/nhatminh/korean-asr.tabularautomatic-speech-recognition1M<n<10M0 likes45 downloads1mo agoHugging Face19sumanpaudel1997 /nepali-asr-benchmark Nepali ASR Benchmark Per-utterance reference, hypothesis, WER, and CER for the six released Nepali ASR checkpoints evaluated on three independent test sets. Released alongside the paper Comparative Analysis of Multilingual Pre-trained Models for Nepali Automatic Speech Recognition. Contents Field Type Description utterance_id string stable identifier {test_set}-{index} reference string NFC-normalised gold transcription (Devanagari) hypothesis string… See the full description on the dataset page: https://huggingface.co/datasets/sumanpaudel1997/nepali-asr-benchmark.tabularautomatic-speech-recognition10K<n<100K0 likes44 downloads4mo agoHugging Face20AdityK2409 /vistaar_small_asr_eval Vistaar Small ASR Eval Dataset Description The Vistaar Small ASR Eval dataset is a multilingual automatic speech recognition evaluation dataset containing 9,486 audio samples across 12 Indian languages. This dataset represents a subset of the larger Vistaar dataset published by AI4Bharat, designed specifically for evaluating ASR model performance on diverse Indian language speech data. A smaller evaluation dataset was created for the use-cases where a quick benchmarking… See the full description on the dataset page: https://huggingface.co/datasets/AdityK2409/vistaar_small_asr_eval.audioautomatic-speech-recognition10K<n<100K0 likes41 downloads1y agoHugging Face21anuj-inavlabs /kupe-asr-en-data kupe-asr-en-mini-150m — data Two loadable configs, packed into ~20-25 bunch_*.parquet files each (Hub-quota friendly): raw — 24 kHz mono English audio (flac bytes) + text. The encode stage reads this. mimi — Mimi c0..c7 codes (12.5 Hz) + text. Training reads this. Ledgers under ledger/ (data.json, mimi.json) track collected/encoded hours and resume state. from datasets import load_dataset ds = load_dataset("anuj-inavlabs/kupe-asr-en-data", "mimi", split="train") tabularautomatic-speech-recognition1M<n<10M0 likes39 downloads1mo agoHugging Face22speech-uk /asr-evaluationstabularautomatic-speech-recognition10K<n<100K0 likes35 downloads2y agoHugging Face23Quran-Lab /quranic-asr-cloud-rawdatagated Quranic ASR Provider Benchmark Results Professional benchmark artifacts for comparing commercial and official ASR providers on the Quranic ASR benchmark hosted at Quran-Lab/quranic-asr-benchmark. This repository contains metadata, normalized result tables, raw provider responses, unchanged run scripts, scoring outputs, Tarteel streaming probes, and reports. It does not duplicate the source audio. What Is Included Area Path Purpose Benchmark split… See the full description on the dataset page: https://huggingface.co/datasets/Quran-Lab/quranic-asr-cloud-rawdata.tabularautomatic-speech-recognition1K<n<10K1 likes35 downloads2mo agoHugging Face24uam-wmi-asr-eval-labs /2026-dwesui-g02-kulinarna DWESUI 2026 - Grupa 2 - kulinarna (PIEROGA) Robocza/archiwalna kopia zbioru ewaluacyjnego ASR zbudowanego przez studentow kursu Warsztaty z ewaluacji systemow rozpoznawania mowy (UAM WMI), edycja 2026, tryb dzienny. Zespol (atrybucja): Grupa 2 (DWESUI 2026) Zrodlo oryginalne: https://huggingface.co/datasets/s479246/dwesui-grupa-2-kulinarna Domena: kulinarna Licencja zrodla: nagrania YouTube CC-BY/CC-BY-SA + TTS Status: kopia publiczna w organizacji kursowej (zespół opublikował… See the full description on the dataset page: https://huggingface.co/datasets/uam-wmi-asr-eval-labs/2026-dwesui-g02-kulinarna.audioautomatic-speech-recognitionn<1K0 likes30 downloads2mo agoHugging Face25knesset-asr /knesset-committees-inference Knesset Committees Inference Transcriptions of Knesset committee audio by two models, with the protocol reference alongside, for the personalised-ASR study (Stage 1: per-speaker WER, general model vs Hebrew fine-tune). Audio and references come from Hadasy/knesset-committees-chunks; speaker identities from Dolevabudi/knesset-committees-speakers. No audio is included. arm model served by language A openai/whisper-large-v3 HF Inference (deepinfra) forced he B… See the full description on the dataset page: https://huggingface.co/datasets/knesset-asr/knesset-committees-inference.tabularautomatic-speech-recognition1M<n<10M0 likes19 downloads27d agoHugging Face26cillegio /az-asr-voa-305hgated Labelling field value label_origin script speech_register broadcast channel wideband-16k provenance inferred Editorial broadcast text aligned to VOA audio. Terminal punctuation at 77.4% (against 98.1% for LocalDoc) is consistent with segments cut from continuous broadcast rather than at sentence ends. Adding this to a call model made it worse. Chinar-F8 v4 mixed in 147.8 h of it and strict WER on human-transcribed calls moved 44.23% -> 56.69%. Register, not… See the full description on the dataset page: https://huggingface.co/datasets/cillegio/az-asr-voa-305h.tabularautomatic-speech-recognition100K<n<1M0 likes19 downloads21d agoHugging Face27pavanyellow /librispeech_asr Dataset Card for librispeech_asr LibriSpeech ASR 2s Splits Dataset Version of LibriSpeech ASR corpus split into 2s clips. Usage from datasets import load_dataset # Load the dataset from the Hub dataset = load_dataset("pavanyellow/librispeech_asr") # Or load a specific split dataset = load_dataset("pavanyellow/librispeech_asr", split="train") # Access the data for example in dataset['train'][:5]: audio = example['audio'] text = example['text'] tabularautomatic-speech-recognition10K<n<100K0 likes12 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.