datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hit-asr
HIT-ASR — data and results
The data behind HIT-ASR: Hierarchical Transformer Routing for Adaptive ASR Expert Selection (Huseyin Karaca,
A. Samil Namli, Suleyman S. Kozat — Bilkent University): what the pretrained ASR experts of the paper produce on its
four English corpora, and the stored results every notebook of the code repository reads.
Code and notebooks: github.com/huseyin-karaca/hit-asr
Documentation: huseyin-karaca.github.io/hit-asr
What is here… See the full description on the dataset page: https://huggingface.co/datasets/huseyin-karaca/hit-asr.danish-asr-leaderboard
Open Danish ASR Leaderboard — Results
Benchmark results backing the Open Danish ASR Leaderboard — an open, reproducible comparison of Danish speech-to-text models.
Every model is transcribed and scored identically on the same five independent public Danish test sets, so the numbers compare directly: Word Error Rate (WER) and Character Error Rate (CER) — lower is better — plus speed. Open-weight Danish speech recognition models you can run yourself and hosted transcription APIs… See the full description on the dataset page: https://huggingface.co/datasets/RyeAI/danish-asr-leaderboard.danish-asr-unified-hviske-v5-tiny
danish-asr-unified — two-model labels and a quality manifest
Transcriptions, per-token confidences, and a per-row quality verdict for every
row of syvai/danish-asr-unified
(3,414,589 rows, 8 sources).
Two independently trained models labelled the whole corpus:
model
architecture
vocabulary
syvai/hviske-v5-tiny
encoder-decoder
16,384 BPE
3dio-ai/svale-110M
RNN-T (Parakeet)
44 characters
Both decode greedily. Shards mirror the source data/train-*.parquet by name… See the full description on the dataset page: https://huggingface.co/datasets/syvai/danish-asr-unified-hviske-v5-tiny.darija-asr-corpus
Darija ASR Corpus (dataset-core)
Arabizi (Latin-script) transcriptions of Moroccan Darija speech, produced for a
Whisper fine-tuning pipeline (paper not yet published -- citation forthcoming).
This repo contains four source subsets: DODa, DVoice, Wiki, and
YouTube. Each subset carries its own upstream license/terms -- see below --
because they are drawn from four different original projects.
Subsets
Config
Rows
Audio bundled?
Upstream license
Upstream source… See the full description on the dataset page: https://huggingface.co/datasets/abnajlae/darija-asr-corpus.danish-asr-verified
danish-asr-verified
ALL rows of syvai/danish-asr-unified transcribed by the
syv-transcribe ensemble (hviske-v5.3 + hviske-v5, confidence-weighted ROVER),
each annotated with:
verified — True when the ensemble independently reproduced the reference
exactly (compared after lowercasing, punctuation-strip, whitespace-collapse).
Two independent witnesses agree => near-certain label.
wer_teacher_vs_ref / cer_teacher_vs_ref — word/character error rate
between normalized teacher output… See the full description on the dataset page: https://huggingface.co/datasets/syvai/danish-asr-verified.ark-asr-3b-open-asr-leaderboard-results
ARK-ASR-3B Open ASR Leaderboard Results
Raw JSONL manifests for AutoArk-AI/ARK-ASR-3B on the public English
short-form hf-audio/open-asr-leaderboard splits.
These manifests were generated on a local 8x RTX 4090 machine and scored with
the shared Open ASR Leaderboard scorer:
PYTHONPATH=. python - <<'PY'
from normalizer.eval_utils import score_results
score_results(
'ark_asr/results.AutoArk-AI-ARK-ASR-3B_20260622_official',
'AutoArk-AI/ARK-ASR-3B',
)
PY
Important:… See the full description on the dataset page: https://huggingface.co/datasets/Edge0/ark-asr-3b-open-asr-leaderboard-results.hebrew-asr-vn
Hebrew ASR three-source training dataset
Derived from ivrit.ai crowd-transcribe-v5, crowd-recital and knesset-committees.
Original data and transcripts are credited to ivrit.ai and its contributors.
Pinned revisions and preparation rules are in metadata/sources.json and
metadata/preparation-config.json. VoxKnesset is excluded by user decision.
Source/split
Clips
Hours
crowd-recital/test
1,557
1.071
crowd-recital/train
45,372
33.258
crowd-recital/validation
1,033… See the full description on the dataset page: https://huggingface.co/datasets/benderrodriguez/hebrew-asr-vn.librispeech_asr_dummy
Dataset Card for librispeech_asr_dummy
Dataset Summary
This is a truncated version of the LibriSpeech dataset. It contains 20 samples from each of the splits. To view the full dataset, visit: https://huggingface.co/datasets/librispeech_asr
LibriSpeech is a corpus of approximately 1000 hours of 16kHz read English speech, prepared by Vassil Panayotov with the assistance of Daniel Povey. The data is derived from read audiobooks from the LibriVox project, and has been… See the full description on the dataset page: https://huggingface.co/datasets/sanchit-gandhi/librispeech_asr_dummy.Indonesian-ASR-11-Class-Dataset
Indonesian ASR 11-Class Dataset
Public Hugging Face repository for an Indonesian ASR corpus and its paper-supporting benchmark artifacts.
Dataset summary
Audio files: 104,500 WAV files
Real/human recordings: 104,368
Synthetic repair files: 132
Sentence classes: 11 Indonesian sentence categories
Canonical balanced sentence slots: 209 (11 categories × 19 retained slots)
Public speaker labels: M1..M12, F1..F8, plus synthetic labels Ms*/Fs*
Audio format: 16 kHz… See the full description on the dataset page: https://huggingface.co/datasets/Atika88/Indonesian-ASR-11-Class-Dataset.clipquill-asr-benchmark
Measuring whisper-tiny vs whisper-base in a browser tab
Word error rate, wall-clock timing, transfer size and peak memory for two
quantised Whisper tiers running entirely client-side in a real Chrome window,
with the scripts that produced every number.
If you are building an in-browser transcription page, the two results worth
knowing before you pick a model tier:
On clean synthetic audio the two tiers tie. If that is all you test, you
will conclude the tier does not matter… See the full description on the dataset page: https://huggingface.co/datasets/sophia8888/clipquill-asr-benchmark.Tech-Sentences-For-ASR-Training
TechVoice Dataset
Work in Progress – This dataset is actively being expanded with new recordings.
Dataset Statistics
Metric
Current
Target
Progress
Duration
38m 43s
5h 0m 0s
██░░░░░░░░░░░░░░░░░░ 12.9%
Words
10,412
50,000
████░░░░░░░░░░░░░░░░ 20.8%
Total Recordings: 205 samples
Total Characters: 74,312
A specialized speech dataset for fine-tuning Automatic Speech Recognition (ASR) models on technical and developer vocabulary. Contains human-recorded… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Tech-Sentences-For-ASR-Training.asr-benchmark-outputs
SaarAI ASR Benchmark Outputs
Raw per-utterance model outputs (transcription manifests) produced by the gsma-asr-bench runners on SaarAI/asr-leaderboard-datasets.
files: 570
utterances: 4615160
languages: 7
models: 50
Layout
data/<language_name>/<split>__<dataset_config>__<model_slug>.jsonl
index.jsonl # one record per file (language, split, model, rows, sha256, ...)
index.csv
Directories categorise by language name; the file name begins with the split name… See the full description on the dataset page: https://huggingface.co/datasets/SaarAI/asr-benchmark-outputs.tajik-asr-youtube
tajik-asr-youtube
Every Tajik-language clip from this project's YouTube scrape — 41 channels of news, talk
shows, podcasts, audiobooks, and learning content — with machine transcripts and the
verification scores left in as columns instead of applied as a filter. Pick your own
quality threshold; the training corpus this project actually ships
(tajik-asr-corpus-v3)
is the gated subset.
Layout
Parquet shards under data/ with an audio struct column (16 kHz mono FLAC… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/tajik-asr-youtube.polish-tedx-asr-eval
Polish-TEDx-ASR-Eval
A dataset for evaluating automatic speech recognition (ASR) systems for Polish in the domain of TEDx public talks.
Contains audio segments from Polish TEDx talks available on YouTube (CC BY-NC-ND 4.0) and synthetic speech generated with KugelAudio (MIT), with manually created and cross-verified transcriptions. Created as part of the course "Workshops on Evaluation of Speech Recognition Systems" (ZWESUI, AMU 2026) by Group 1.
Statistics… See the full description on the dataset page: https://huggingface.co/datasets/s512757/polish-tedx-asr-eval.persian-asr-text-2.69M-deduped
🗂️ persian-asr-text-2.69M-deduped
English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission
🌟 At a glance | معرفی سریع
English
فارسی
🎯 Purpose
Deduplicated Persian ASR text dataset used by the training stack.
پیکرهٔ متنی فارسیِ حذفتکرارشده برای ساخت واژگان، مدلسازی زبانی و پشتیبانی از آموزش ASR.
🧩 Role
Persian text and linguistic asset
مصنوع متنی و زبانی فارسی
📦 Snapshot
4 files; approximately 109.64 MB
4 فایل؛ حدود 109.64… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-asr-text-2.69M-deduped.ben10-asr-results
Ben-10 Regional ASR — public results
Score rows for the maintainer-run Ben-10 regional dialect ASR leaderboard.
Field
Meaning
model_id
Hub id or slug
model_url
Link to weights / paper
wer
Corpus Word Error Rate on private ben-10-test (lower better)
wer_by_region
JSON map region → WER
backend
Decode stack used by maintainers
scorer_commit / decode_commit
Git SHAs in BengaliAI/reg-speech-aacl
evaluated_at
ISO date
requested_by
Who asked, or maintainer if… See the full description on the dataset page: https://huggingface.co/datasets/bengaliAI/ben10-asr-results.Swedia-ASR-Dataset
Swedia ASR Dataset
This repository contains a small Swedish ASR evaluation dataset based on speech
transcriptions from Swedia 2000. It was assembled to compare automatic
speech-recognition output against manually corrected reference transcriptions
for Swedish dialectal speech.
The dataset is useful for quick experiments with Swedish ASR systems, especially
when you want to inspect recognition quality on spontaneous speech from
different regions, speakers, ages, and genders.… See the full description on the dataset page: https://huggingface.co/datasets/kvest/Swedia-ASR-Dataset.korean-asr
korean-asr — Korean ASR pseudo-labels for YODAS2
This repository contains transcripts and segment metadata only. It does not contain audio.
Every row points into espnet/yodas2 by
(shard, video_id, start, end), so you fetch the audio from the upstream dataset and cut it
yourself. See Reconstructing the audio.
split
utterances
hours
train
1,034,181
6,974.3
heldout
46,542
314.6
dev (subset of heldout)
3,000
20.4
Total labelled: 1,080,723 utterances / 7,288.9… See the full description on the dataset page: https://huggingface.co/datasets/nhatminh/korean-asr.nepali-asr-benchmark
Nepali ASR Benchmark
Per-utterance reference, hypothesis, WER, and CER for the six released Nepali ASR
checkpoints evaluated on three independent test sets. Released alongside the paper
Comparative Analysis of Multilingual Pre-trained Models for Nepali Automatic Speech
Recognition.
Contents
Field
Type
Description
utterance_id
string
stable identifier {test_set}-{index}
reference
string
NFC-normalised gold transcription (Devanagari)
hypothesis
string… See the full description on the dataset page: https://huggingface.co/datasets/sumanpaudel1997/nepali-asr-benchmark.vistaar_small_asr_eval
Vistaar Small ASR Eval
Dataset Description
The Vistaar Small ASR Eval dataset is a multilingual automatic speech recognition evaluation dataset containing 9,486 audio samples across 12 Indian languages. This dataset represents a subset of the larger Vistaar dataset published by AI4Bharat, designed specifically for evaluating ASR model performance on diverse Indian language speech data. A smaller evaluation dataset was created for the use-cases where a quick benchmarking… See the full description on the dataset page: https://huggingface.co/datasets/AdityK2409/vistaar_small_asr_eval.kupe-asr-en-data
kupe-asr-en-mini-150m — data
Two loadable configs, packed into ~20-25 bunch_*.parquet files each (Hub-quota friendly):
raw — 24 kHz mono English audio (flac bytes) + text. The encode stage reads this.
mimi — Mimi c0..c7 codes (12.5 Hz) + text. Training reads this.
Ledgers under ledger/ (data.json, mimi.json) track collected/encoded hours and resume state.
from datasets import load_dataset
ds = load_dataset("anuj-inavlabs/kupe-asr-en-data", "mimi", split="train")
asr-evaluationsquranic-asr-cloud-rawdata
Quranic ASR Provider Benchmark Results
Professional benchmark artifacts for comparing commercial and official ASR providers on the Quranic ASR benchmark hosted at Quran-Lab/quranic-asr-benchmark.
This repository contains metadata, normalized result tables, raw provider responses, unchanged run scripts, scoring outputs, Tarteel streaming probes, and reports. It does not duplicate the source audio.
What Is Included
Area
Path
Purpose
Benchmark split… See the full description on the dataset page: https://huggingface.co/datasets/Quran-Lab/quranic-asr-cloud-rawdata.2026-dwesui-g02-kulinarna
DWESUI 2026 - Grupa 2 - kulinarna (PIEROGA)
Robocza/archiwalna kopia zbioru ewaluacyjnego ASR zbudowanego przez studentow kursu
Warsztaty z ewaluacji systemow rozpoznawania mowy (UAM WMI), edycja 2026, tryb dzienny.
Zespol (atrybucja): Grupa 2 (DWESUI 2026)
Zrodlo oryginalne: https://huggingface.co/datasets/s479246/dwesui-grupa-2-kulinarna
Domena: kulinarna
Licencja zrodla: nagrania YouTube CC-BY/CC-BY-SA + TTS
Status: kopia publiczna w organizacji kursowej (zespół opublikował… See the full description on the dataset page: https://huggingface.co/datasets/uam-wmi-asr-eval-labs/2026-dwesui-g02-kulinarna.knesset-committees-inference
Knesset Committees Inference
Transcriptions of Knesset committee audio by two models, with the protocol
reference alongside, for the personalised-ASR study (Stage 1: per-speaker
WER, general model vs Hebrew fine-tune). Audio and references come from
Hadasy/knesset-committees-chunks;
speaker identities from
Dolevabudi/knesset-committees-speakers.
No audio is included.
arm
model
served by
language
A
openai/whisper-large-v3
HF Inference (deepinfra)
forced he
B… See the full description on the dataset page: https://huggingface.co/datasets/knesset-asr/knesset-committees-inference.az-asr-voa-305h
Labelling
field
value
label_origin
script
speech_register
broadcast
channel
wideband-16k
provenance
inferred
Editorial broadcast text aligned to VOA audio. Terminal punctuation at 77.4% (against 98.1% for LocalDoc) is consistent with segments cut from continuous broadcast rather than at sentence ends.
Adding this to a call model made it worse. Chinar-F8 v4 mixed in 147.8 h of it and strict WER on human-transcribed calls moved 44.23% -> 56.69%. Register, not… See the full description on the dataset page: https://huggingface.co/datasets/cillegio/az-asr-voa-305h.librispeech_asr
Dataset Card for librispeech_asr
LibriSpeech ASR 2s Splits Dataset
Version of LibriSpeech ASR corpus split into 2s clips.
Usage
from datasets import load_dataset
# Load the dataset from the Hub
dataset = load_dataset("pavanyellow/librispeech_asr")
# Or load a specific split
dataset = load_dataset("pavanyellow/librispeech_asr", split="train")
# Access the data
for example in dataset['train'][:5]:
audio = example['audio']
text = example['text']
