datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
common-voice-scripted-speech-26
Common Voice Scripted Speech
A row-normalized multilingual ASR dataset built from Mozilla Data Collective
Common Voice Scripted Speech. Each upstream archive is converted to appendable
parquet shards under data/<upstream_split>/, one shard per source archive and
split, with audio bytes embedded in an audio struct column.
Status
Manifest languages: 60
Languages uploaded: 18
Columns
audio (bytes, path)
sentence, locale, language, upstream_split… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/common-voice-scripted-speech-26.script-fidelity-benchmark
Script fidelity benchmark
Anonymous supplement for the paper "Script collapse in multilingual ASR:
A reference-free metric and 100-pair benchmark."
Script Fidelity Rate (SFR) measures the fraction of ASR hypothesis characters
that belong to the expected target script. WER measures word edits, while SFR
checks whether the output is written in the target orthography.
Related resources:
PyPI package: https://pypi.org/project/script-fidelity/
Hugging Face Evaluate metric:… See the full description on the dataset page: https://huggingface.co/datasets/themechanism/script-fidelity-benchmark.f1-team-radio
F1 Team Radio Dataset
A comprehensive dataset of Formula 1 team radio communications with transcriptions.
Dataset Description
This dataset contains team radio audio clips from Formula 1 races along with their text transcriptions. Team radio communications are the real-time messages exchanged between F1 drivers and their pit wall engineers during race weekends.
Dataset Statistics
Metric
Value
Total audio clips
14,681
Grand Prix events
149
Unique… See the full description on the dataset page: https://huggingface.co/datasets/scriptaudio/f1-team-radio.cv-en-scripted-test-500
Common Voice English Scripted Test Set — 500 clips
n = 500 utterances · private eval set for ASR benchmarking
Source
Derived from Mozilla Common Voice Scripted Speech 25.0 — English (test split), downloaded via the Mozilla Data Collective API (dataset ID cmndapwry02jnmh07dyo46mot, 94 GB tarball).
Construction
Starting from the full CV 25.0 English test split (16,398 rows), a stratified 500-clip subset was produced using the same recipe as… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/cv-en-scripted-test-500.Thai-human-to-machine-call-center-audio-with-scriptThis dataset features natural Thai-language conversations between human speakers and machine agents, simulating real-world call center interactions across a variety of customer service domains. All dialogues are non-scripted and performed as role-play scenarios, capturing spontaneous, realistic exchanges.
🗣️ Speech Type: Human-to-machine conversations, simulating AI agents and human customers' dialogues.
🎭 Style: Spontaneous, unscripted role-playing, designed to reflect actual customer… See the full description on the dataset page: https://huggingface.co/datasets/AIxBlock/Thai-human-to-machine-call-center-audio-with-script.
