datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dia-alimeeting-test
AliMeeting — test split (far + near, speaker diarization)
Copie du split test d'AliMeeting (M2MeT challenge) en deux vues :
far-field : 1 mix WAV par session (8-mic array, channel 1)
near-field : 1 WAV par participant (headset microphones)
Les TextGrid sources ont été convertis en RTTM standard pyannote par
scripts/hf/upload_alimeeting.py du projet STTSTAGE.
Contenu
Vue
Sessions
WAV
RTTM
far
20
20 (1 mix/session)
20
near
20
60 (≈3 speakers/session)
20… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/dia-alimeeting-test.NPSC_test
Dataset Card for NBAiLab/NPSC
The Norwegian Parliament Speech Corpus (NPSC) is a corpus for training a Norwegian ASR (Automatic Speech Recognition) models. The corpus is created by Språkbanken at the National Library in Norway.
NPSC is based on sound recording from meeting in the Norwegian Parliament. These talks are orthographically transcribed to either Norwegian Bokmål or Norwegian Nynorsk. In addition to the data actually included in this dataset, there is a significant amount… See the full description on the dataset page: https://huggingface.co/datasets/NbAiLab/NPSC_test.asr-testset-kw-ja-v1
日本語 ASR アノテーション v1
重要:評価結果を報告する際の規約
本テストセットで評価結果を報告する際は、事前学習を含む学習にYouTubeの音声を使用したかどうかと、次の評価区分を必ず明記してください。
使用した場合:in-domain評価
使用していない場合:out-of-domain評価
この区分は、音声認識の評価結果を比較する上で重要です。最終的な追加学習だけでなく、使用するモデルの事前学習も含めて判断してください。
元データセットの音声に、人手で区間ごとの転記・タグ・採否を付けたデータです。音声の内容、ディレクトリ構成、ファイル名は元のままです。
ファイルと表示
**annotations.jsonl**:提出済みの全結果。1行が1音声です。
**metadata.jsonl**:HF表示用に自動生成したデータ。音声全体が不使用の行を除き、audio を file_name に置き換えています。
**data/**:採用した音声ファイル。
HFのビューアーは… See the full description on the dataset page: https://huggingface.co/datasets/bandad/asr-testset-kw-ja-v1.dia-voxconverse-test
VoxConverse — test split (speaker diarization)
Copie du split test de VoxConverse v0.3 mise en forme pour les benchmarks
de diarisation (pyannote, NeMo, etc.). 232 fichiers audio + 232 RTTM de référence.
Contenu
232 enregistrements (TV/YouTube anglais, multi-speakers, réunions & débats)
Audio : WAV 16 kHz, mono
Annotations : RTTM (Rich Transcription Time Marked)
Langue : anglais (en)
Licence : CC-BY-4.0 (identique à VoxConverse upstream)
Structure… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/dia-voxconverse-test.dia-MsdWild-test
MSDWild — validation / test split (speaker diarization in the wild)
Copie du split d'évaluation de MSDWild (Liu et al., Interspeech 2022), un
corpus de diarisation in the wild multimodal construit à partir de vidéos
réelles (talk-shows, interviews, débats…). MSDWild fournit deux sets de
validation utilisés comme test de facto par la communauté :
few : 2 à 4 locuteurs par session (490 fichiers) — tâche "classique"
many : 5+ locuteurs par session (177 fichiers) — tâche… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/dia-MsdWild-test.stt-vibravox-fr-test
VibraVox FR — test split (mirror of Cnam-LMSSC/vibravox)
Mirror public des splits test de VibraVox (CNAM-LMSSC, Paris) pour
benchmark ASR français multi-capteur sur audio standard ET non-standard
(bone-conduction, in-ear, throat, accéléromètre).
Ce repo contient uniquement les configs speech_clean + speech_noisy
(les seules avec transcription). Les configs speechless_* upstream sont
exclues car sans texte → pas de WER possible.
Configs
Config
Test rows
Test… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/stt-vibravox-fr-test.speechio_test
SpeechIO ASR Test Sets (parquet)
Parquet repackaging of the SpeechColab SpeechIO Mandarin ASR benchmark,
re-exported from yuekai/speechio (Lhotse cuts) into standard
HuggingFace parquet with embedded 16 kHz audio.
27 test sets: SPEECHIO_ASR_ZH00000 ... SPEECHIO_ASR_ZH00026, each a config with a single test split.
~43k utterances, ~66 hours total, evaluation only.
Columns
column
type
note
segment_id
string
utterance id
speaker
string
speaker id… See the full description on the dataset page: https://huggingface.co/datasets/yuekai/speechio_test.dia-aishell4-test
AISHELL-4 — test split (meeting diarization)
Copie du split test d'AISHELL-4, un corpus de réunions en mandarin capturé
par un array de 8 micros (on garde ici la version single-channel extraite
pour les benchmarks diarisation).
Contenu
20 sessions de réunion (3–7 speakers / session, durée variable)
Audio : FLAC mono
Annotations : RTTM par session
Langue : mandarin (zh)
Licence : Apache-2.0 (upstream)
Structure
dia-aishell4-test/
├── audio/test/<file_id>.flac… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/dia-aishell4-test.fleurs_test
FLEURS Test Dataset with Enhanced Metadata
This dataset is an enhanced version of the FLEURS (Few-shot Learning Evaluation of Universal Representations of Speech) test set, restructured with complete metadata for easier use in automatic speech recognition (ASR) and multilingual speech processing tasks.
Dataset Description
FLEURS is a multilingual speech benchmark dataset designed to evaluate universal speech representations. This particular version focuses on 25 European… See the full description on the dataset page: https://huggingface.co/datasets/rasgaard/fleurs_test.h-test@misc{humair025/h-test,
title = {h-test},
author = {Humair Munir},
year = {2025},
howpublished = {\url{https://huggingface.co/datasets/humair025/h-test}},
note = {Synthetic dataset of speech (with emotions). Licensed under CC-BY 4.0.}
}
ciempiess_testThe CIEMPIESS TEST Corpus is a gender balanced corpus destined to test acoustic models for the speech recognition task. The corpus was manually transcribed and it contains audio recordings from 10 male and 10 female speakers. The CIEMPIESS TEST is one of the three corpora included at the LDC's \"CIEMPIESS Experimentation\" (LDC2019S07).dhravani-mit-testCheck out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
Dataset Preparation Interface for Fine-tuning Whisper
A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication.
Features
🔐 User authentication via Pocketbase
☁️ Cloud storage… See the full description on the dataset page: https://huggingface.co/datasets/shrikanth-19/dhravani-mit-test.Speech-MASSIVE-test
Speech-MASSIVE Test Split
This dataset repository is only for test split of Speech-MASSIVE.
train and dev splits are available in the separate dataset repository. https://huggingface.co/datasets/FBK-MT/Speech-MASSIVE
Dataset Description
Speech-MASSIVE is a multilingual Spoken Language Understanding (SLU) dataset comprising the speech counterpart for a portion of the MASSIVE textual corpus. Speech-MASSIVE covers 12 languages (Arabic, German, Spanish, French, Hungarian… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/Speech-MASSIVE-test.test-fixtures
podscripter test fixtures
Small, curated audio fixtures used by the
podscripter project's Tier 1 regression
tests (tests/test_audio_fixtures.py). Each audio file pairs with an .expected.json
metadata file checked into the podscripter repo at
tests/fixtures/audio/<lang>/<name>.expected.json.
The repo pins a specific revision of this dataset in
tests/fixtures/audio/download.py, so audio + tests stay in lockstep.
Aggregate license
CC-BY 4.0 — the most restrictive… See the full description on the dataset page: https://huggingface.co/datasets/podscripter-project/test-fixtures.dhravani-iitpatna-testCheck out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
Dataset Preparation Interface for Fine-tuning Whisper
A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication.
Features
🔐 User authentication via Pocketbase
☁️ Cloud storage… See the full description on the dataset page: https://huggingface.co/datasets/shrikanth-19/dhravani-iitpatna-test.dhravani-IIT_Guwahati-testCheck out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
Dataset Preparation Interface for Fine-tuning Whisper
A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication.
Features
🔐 User authentication via Pocketbase
☁️ Cloud storage… See the full description on the dataset page: https://huggingface.co/datasets/shrikanth-19/dhravani-IIT_Guwahati-test.dhravani-iitdelhi-testCheck out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
Dataset Preparation Interface for Fine-tuning Whisper
A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication.
Features
🔐 User authentication via Pocketbase
☁️ Cloud storage… See the full description on the dataset page: https://huggingface.co/datasets/shrikanth-19/dhravani-iitdelhi-test.dhravani-IGDTUW_Delhi-testCheck out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
Dataset Preparation Interface for Fine-tuning Whisper
A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication.
Features
🔐 User authentication via Pocketbase
☁️ Cloud storage… See the full description on the dataset page: https://huggingface.co/datasets/shrikanth-19/dhravani-IGDTUW_Delhi-test.dhravani-iiitdelhi-testCheck out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
Dataset Preparation Interface for Fine-tuning Whisper
A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication.
Features
🔐 User authentication via Pocketbase
☁️ Cloud storage… See the full description on the dataset page: https://huggingface.co/datasets/shrikanth-19/dhravani-iiitdelhi-test.stt-cefc-fr-test
CEFC-Orfeo FR — long-form oral test mirror
Mirror non-officiel du Corpus d'Études du Français Contemporain (CEFC)
agrégé par le projet Orfeo (ANR), tel que distribué sur le portail
projet-orfeo.fr (release 13).
Long-form : 1 row = 1 fichier audio entier (30-60 min en moyenne).
12 sous-corpus oraux du français contemporain, 303 heures au total,
901 fichiers. Idéal pour bench Whisper / Canary en conditions réelles
(chunked decoding, dérive temporelle, multi-locuteurs).… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/stt-cefc-fr-test.RS-test-fix2NVVSpeech-Challenge-Track1-Test-Set
NVVSpeech Challenge Track 1 Test Set
Track 1 test set for the NVVSpeech Challenge at ISCSLP 2026.
Task
Given a speech recording, produce a transcript that contains the spoken content and the non-verbal vocalization (NVV) tags at their corresponding positions.
Dataset Summary
Language
Samples
Chinese
985
English
961
Total
1,946
Files
.
├── README.md
├── SUBMISSION_GUIDE.txt
├── test.jsonl
├── ground_truth.jsonl
├──… See the full description on the dataset page: https://huggingface.co/datasets/NVVSpeech-Challenge/NVVSpeech-Challenge-Track1-Test-Set.testThe Snow Mountain dataset contains the audio recordings (in .mp3 format) and the corresponding text of The Bible
in 11 Indian languages. The recordings were done in a studio setting by native speakers. Each language has a single
speaker in the dataset. Most of these languages are geographically concentrated in the Northern part of India around
the state of Himachal Pradesh. Being related to Hindi they all use the Devanagari script for transcription.voxtral-synthetic-eng-test
Voxtral Synthetic English (ASR)
Synthetic speech dataset for fine-tuning Voxtral ASR models. English utterances generated with ElevenLabs TTS from the CohereLabs/aya_collection_language_split (english, targets column). All audio is 16 kHz mono WAV.
Dataset structure
Column
Type
Description
audio_path
string
Path to the audio file in this repo (e.g. audio/utt_000000.wav)
text
string
Ground-truth transcript for the audio
Audio: 16 kHz, mono, WAV, stored… See the full description on the dataset page: https://huggingface.co/datasets/shakods/voxtral-synthetic-eng-test.RS-teststt-summre-fr-test
SUMM-RE — French test split (mirror of linagora/SUMM-RE)
Mirror public du split test de SUMM-RE (LINAGORA / Aix-Marseille
LPL), pour benchmark ASR français conversationnel (parole de réunion,
3-4 locuteurs, ~20 min par session).
⚠ Ce repo ne contient que le split test (124 tracks individuelles =
37 réunions × 3-4 micros). Pour les splits train / dev, voir le repo
upstream linagora/SUMM-RE.
Contenu
124 pistes audio individuelles (1 piste = 1 microphone d'un locuteur… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/stt-summre-fr-test.X-Voice-TestsetX-Voice Multilingual Test Set
High-Fidelity Test Set for Multilingual Text-to-Speech across 30 Languages
This test set is built as part of the research: X-Voice: One Speaker, 30+ Languages with Zero-Shot Voice Cloning, serving as the evaluation benchmark for our model.
Dataset Summary
30 languages
European: bg (Bulgarian), cs (Czech), da (Danish), de (German), el (Greek), en (English), es (Spanish), et (Estonian), fi (Finnish), fr (French), hr (Croatian), hu (Hungarian), it… See the full description on the dataset page: https://huggingface.co/datasets/XRXRX/X-Voice-Testset.common-voice-17-tr-test
common-voice-17-tr-test
Turkish test split of Common Voice 17.0 (tr), re-hosted for Turkish STT benchmarking.
Rows: 11290
Columns: client_id, path, audio, sentence, up_votes, down_votes, age, gender, accent, locale, segment, variant
Source: https://commonvoice.mozilla.org
License: cc0-1.0 (inherited from source)
Only the Turkish test split is included, extracted as-is from the source dataset.
audio-testing
audio-testing
Overview
This is a small, open dataset designed for quick validation of audio-related pipelines and applications, especially for Text-to-Speech (TTS) and Speech-to-Text (STT) systems.
It provides a few short, diverse audio clips and corresponding text transcripts, allowing developers to verify input/output handling, audio processing, and transcription logic without downloading large datasets.
Contents
3 short audio samples (.mp3, .wav)… See the full description on the dataset page: https://huggingface.co/datasets/JacobLinCool/audio-testing.testtrdataset-1
TestTRDataset_1
This is a merged speech dataset containing 11930 audio segments from 24 source datasets.
Dataset Information
Total Segments: 11930
Speakers: 69
Languages: tr
Emotions: angry, happy, neutral, sad
Original Datasets: 24
Dataset Structure
Each example contains:
audio: Audio file (WAV format, 16kHz sampling rate)
text: Transcription of the audio
speaker_id: Unique speaker identifier (made unique across all merged datasets)
emotion: Detected… See the full description on the dataset page: https://huggingface.co/datasets/Codyfederer/testtrdataset-1.
