Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ASLP-lab /Easy-Turn-Testset Easy Turn: Integrating Acoustic and Linguistic Modalities for Robust Turn-Taking in Full-Duplex Spoken Dialogue Systems Guojian Li1, Chengyou Wang1, Hongfei Xue1, Shuiyuan Wang1, Dehui Gao1, Zihan Zhang2, Yuke Lin2, Wenjie Li2, Longshuai Xiao2, Zhonghua Fu1,╀, Lei Xie1,╀ 1 Audio, Speech and Language Processing Group (ASLP@NPU), Northwestern Polytechnical University 2 Huawei Technologies, China 🎤 Demo Page 🤖 Easy Turn Model 📑 Paper 🌐 Huggingface… See the full description on the dataset page: https://huggingface.co/datasets/ASLP-lab/Easy-Turn-Testset.automatic-speech-recognition8 likes1.5k downloads1y agoHugging Face02ggfox00000 /dia-alimeeting-test AliMeeting — test split (far + near, speaker diarization) Copie du split test d'AliMeeting (M2MeT challenge) en deux vues : far-field : 1 mix WAV par session (8-mic array, channel 1) near-field : 1 WAV par participant (headset microphones) Les TextGrid sources ont été convertis en RTTM standard pyannote par scripts/hf/upload_alimeeting.py du projet STTSTAGE. Contenu Vue Sessions WAV RTTM far 20 20 (1 mix/session) 20 near 20 60 (≈3 speakers/session) 20… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/dia-alimeeting-test.audioautomatic-speech-recognitionn<1K2 likes622 downloads6mo agoHugging Face03NbAiLab /NPSC_test Dataset Card for NBAiLab/NPSC The Norwegian Parliament Speech Corpus (NPSC) is a corpus for training a Norwegian ASR (Automatic Speech Recognition) models. The corpus is created by Språkbanken at the National Library in Norway. NPSC is based on sound recording from meeting in the Norwegian Parliament. These talks are orthographically transcribed to either Norwegian Bokmål or Norwegian Nynorsk. In addition to the data actually included in this dataset, there is a significant amount… See the full description on the dataset page: https://huggingface.co/datasets/NbAiLab/NPSC_test.audioautomatic-speech-recognitionn<1K0 likes507 downloads4y agoHugging Face04bandad /asr-testset-kw-ja-v1 日本語 ASR アノテーション v1 重要:評価結果を報告する際の規約 本テストセットで評価結果を報告する際は、事前学習を含む学習にYouTubeの音声を使用したかどうかと、次の評価区分を必ず明記してください。 使用した場合:in-domain評価 使用していない場合:out-of-domain評価 この区分は、音声認識の評価結果を比較する上で重要です。最終的な追加学習だけでなく、使用するモデルの事前学習も含めて判断してください。 元データセットの音声に、人手で区間ごとの転記・タグ・採否を付けたデータです。音声の内容、ディレクトリ構成、ファイル名は元のままです。 ファイルと表示 **annotations.jsonl**:提出済みの全結果。1行が1音声です。 **metadata.jsonl**:HF表示用に自動生成したデータ。音声全体が不使用の行を除き、audio を file_name に置き換えています。 **data/**:採用した音声ファイル。 HFのビューアーは… See the full description on the dataset page: https://huggingface.co/datasets/bandad/asr-testset-kw-ja-v1.audioautomatic-speech-recognitionn<1K0 likes494 downloads6d agoHugging Face05ggfox00000 /dia-voxconverse-test VoxConverse — test split (speaker diarization) Copie du split test de VoxConverse v0.3 mise en forme pour les benchmarks de diarisation (pyannote, NeMo, etc.). 232 fichiers audio + 232 RTTM de référence. Contenu 232 enregistrements (TV/YouTube anglais, multi-speakers, réunions & débats) Audio : WAV 16 kHz, mono Annotations : RTTM (Rich Transcription Time Marked) Langue : anglais (en) Licence : CC-BY-4.0 (identique à VoxConverse upstream) Structure… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/dia-voxconverse-test.audioautomatic-speech-recognitionn<1K0 likes464 downloads6mo agoHugging Face06ggfox00000 /dia-MsdWild-test MSDWild — validation / test split (speaker diarization in the wild) Copie du split d'évaluation de MSDWild (Liu et al., Interspeech 2022), un corpus de diarisation in the wild multimodal construit à partir de vidéos réelles (talk-shows, interviews, débats…). MSDWild fournit deux sets de validation utilisés comme test de facto par la communauté : few : 2 à 4 locuteurs par session (490 fichiers) — tâche "classique" many : 5+ locuteurs par session (177 fichiers) — tâche… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/dia-MsdWild-test.audioautomatic-speech-recognitionn<1K0 likes423 downloads6mo agoHugging Face07ggfox00000 /stt-vibravox-fr-test VibraVox FR — test split (mirror of Cnam-LMSSC/vibravox) Mirror public des splits test de VibraVox (CNAM-LMSSC, Paris) pour benchmark ASR français multi-capteur sur audio standard ET non-standard (bone-conduction, in-ear, throat, accéléromètre). Ce repo contient uniquement les configs speech_clean + speech_noisy (les seules avec transcription). Les configs speechless_* upstream sont exclues car sans texte → pas de WER possible. Configs Config Test rows Test… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/stt-vibravox-fr-test.audioautomatic-speech-recognition1K<n<10K0 likes377 downloads6mo agoHugging Face08yuekai /speechio_test SpeechIO ASR Test Sets (parquet) Parquet repackaging of the SpeechColab SpeechIO Mandarin ASR benchmark, re-exported from yuekai/speechio (Lhotse cuts) into standard HuggingFace parquet with embedded 16 kHz audio. 27 test sets: SPEECHIO_ASR_ZH00000 ... SPEECHIO_ASR_ZH00026, each a config with a single test split. ~43k utterances, ~66 hours total, evaluation only. Columns column type note segment_id string utterance id speaker string speaker id… See the full description on the dataset page: https://huggingface.co/datasets/yuekai/speechio_test.audioautomatic-speech-recognition10K<n<100K0 likes287 downloads3mo agoHugging Face09ggfox00000 /dia-aishell4-test AISHELL-4 — test split (meeting diarization) Copie du split test d'AISHELL-4, un corpus de réunions en mandarin capturé par un array de 8 micros (on garde ici la version single-channel extraite pour les benchmarks diarisation). Contenu 20 sessions de réunion (3–7 speakers / session, durée variable) Audio : FLAC mono Annotations : RTTM par session Langue : mandarin (zh) Licence : Apache-2.0 (upstream) Structure dia-aishell4-test/ ├── audio/test/<file_id>.flac… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/dia-aishell4-test.audioautomatic-speech-recognitionn<1K0 likes271 downloads6mo agoHugging Face101xg /Easy-Turn-Testset Easy Turn: Integrating Acoustic and Linguistic Modalities for Robust Turn-Taking in Full-Duplex Spoken Dialogue Systems Guojian Li1, Chengyou Wang1, Hongfei Xue1, Shuiyuan Wang1, Dehui Gao1, Zihan Zhang2, Yuke Lin2, Wenjie Li2, Longshuai Xiao2, Zhonghua Fu1,╀, Lei Xie1,╀ 1 Audio, Speech and Language Processing Group (ASLP@NPU), Northwestern Polytechnical University 2 Huawei Technologies, China 🎤 Demo Page 🤖 Easy Turn Model 📑 Paper 🌐 Huggingface… See the full description on the dataset page: https://huggingface.co/datasets/1xg/Easy-Turn-Testset.automatic-speech-recognition0 likes242 downloads2mo agoHugging Face11rasgaard /fleurs_test FLEURS Test Dataset with Enhanced Metadata This dataset is an enhanced version of the FLEURS (Few-shot Learning Evaluation of Universal Representations of Speech) test set, restructured with complete metadata for easier use in automatic speech recognition (ASR) and multilingual speech processing tasks. Dataset Description FLEURS is a multilingual speech benchmark dataset designed to evaluate universal speech representations. This particular version focuses on 25 European… See the full description on the dataset page: https://huggingface.co/datasets/rasgaard/fleurs_test.audioautomatic-speech-recognition10K<n<100K0 likes227 downloads8mo agoHugging Face12humair025 /h-test@misc{humair025/h-test, title = {h-test}, author = {Humair Munir}, year = {2025}, howpublished = {\url{https://huggingface.co/datasets/humair025/h-test}}, note = {Synthetic dataset of speech (with emotions). Licensed under CC-BY 4.0.} } audiotext-to-speechn<1K0 likes200 downloads10mo agoHugging Face13ciempiess /ciempiess_testThe CIEMPIESS TEST Corpus is a gender balanced corpus destined to test acoustic models for the speech recognition task. The corpus was manually transcribed and it contains audio recordings from 10 male and 10 female speakers. The CIEMPIESS TEST is one of the three corpora included at the LDC's \"CIEMPIESS Experimentation\" (LDC2019S07).audioautomatic-speech-recognition1K<n<10K3 likes175 downloads3y agoHugging Face14shrikanth-19 /dhravani-mit-testCheck out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference Dataset Preparation Interface for Fine-tuning Whisper A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication. Features 🔐 User authentication via Pocketbase ☁️ Cloud storage… See the full description on the dataset page: https://huggingface.co/datasets/shrikanth-19/dhravani-mit-test.audioautomatic-speech-recognitionn<1K0 likes172 downloads3d agoHugging Face15FBK-MT /Speech-MASSIVE-testgated Speech-MASSIVE Test Split This dataset repository is only for test split of Speech-MASSIVE. train and dev splits are available in the separate dataset repository. https://huggingface.co/datasets/FBK-MT/Speech-MASSIVE Dataset Description Speech-MASSIVE is a multilingual Spoken Language Understanding (SLU) dataset comprising the speech counterpart for a portion of the MASSIVE textual corpus. Speech-MASSIVE covers 12 languages (Arabic, German, Spanish, French, Hungarian… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/Speech-MASSIVE-test.audioaudio-classification10K<n<100K9 likes165 downloads1y agoHugging Face16podscripter-project /test-fixtures podscripter test fixtures Small, curated audio fixtures used by the podscripter project's Tier 1 regression tests (tests/test_audio_fixtures.py). Each audio file pairs with an .expected.json metadata file checked into the podscripter repo at tests/fixtures/audio/<lang>/<name>.expected.json. The repo pins a specific revision of this dataset in tests/fixtures/audio/download.py, so audio + tests stay in lockstep. Aggregate license CC-BY 4.0 — the most restrictive… See the full description on the dataset page: https://huggingface.co/datasets/podscripter-project/test-fixtures.audioautomatic-speech-recognitionn<1K1 likes162 downloads4mo agoHugging Face17shrikanth-19 /dhravani-iitpatna-testCheck out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference Dataset Preparation Interface for Fine-tuning Whisper A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication. Features 🔐 User authentication via Pocketbase ☁️ Cloud storage… See the full description on the dataset page: https://huggingface.co/datasets/shrikanth-19/dhravani-iitpatna-test.audioautomatic-speech-recognitionn<1K0 likes147 downloads3d agoHugging Face18shrikanth-19 /dhravani-IIT_Guwahati-testCheck out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference Dataset Preparation Interface for Fine-tuning Whisper A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication. Features 🔐 User authentication via Pocketbase ☁️ Cloud storage… See the full description on the dataset page: https://huggingface.co/datasets/shrikanth-19/dhravani-IIT_Guwahati-test.audioautomatic-speech-recognitionn<1K0 likes146 downloads3d agoHugging Face19shrikanth-19 /dhravani-iitdelhi-testCheck out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference Dataset Preparation Interface for Fine-tuning Whisper A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication. Features 🔐 User authentication via Pocketbase ☁️ Cloud storage… See the full description on the dataset page: https://huggingface.co/datasets/shrikanth-19/dhravani-iitdelhi-test.audioautomatic-speech-recognitionn<1K0 likes134 downloads3d agoHugging Face20shrikanth-19 /dhravani-IGDTUW_Delhi-testCheck out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference Dataset Preparation Interface for Fine-tuning Whisper A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication. Features 🔐 User authentication via Pocketbase ☁️ Cloud storage… See the full description on the dataset page: https://huggingface.co/datasets/shrikanth-19/dhravani-IGDTUW_Delhi-test.audioautomatic-speech-recognitionn<1K0 likes131 downloads3d agoHugging Face21shrikanth-19 /dhravani-iiitdelhi-testCheck out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference Dataset Preparation Interface for Fine-tuning Whisper A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication. Features 🔐 User authentication via Pocketbase ☁️ Cloud storage… See the full description on the dataset page: https://huggingface.co/datasets/shrikanth-19/dhravani-iiitdelhi-test.audioautomatic-speech-recognitionn<1K0 likes130 downloads3d agoHugging Face22ggfox00000 /stt-cefc-fr-test CEFC-Orfeo FR — long-form oral test mirror Mirror non-officiel du Corpus d'Études du Français Contemporain (CEFC) agrégé par le projet Orfeo (ANR), tel que distribué sur le portail projet-orfeo.fr (release 13). Long-form : 1 row = 1 fichier audio entier (30-60 min en moyenne). 12 sous-corpus oraux du français contemporain, 303 heures au total, 901 fichiers. Idéal pour bench Whisper / Canary en conditions réelles (chunked decoding, dérive temporelle, multi-locuteurs).… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/stt-cefc-fr-test.audioautomatic-speech-recognitionn<1K0 likes125 downloads6mo agoHugging Face23Yehor /RS-test-fix2audioautomatic-speech-recognition10K<n<100K0 likes113 downloads3mo agoHugging Face24NVVSpeech-Challenge /NVVSpeech-Challenge-Track1-Test-Set NVVSpeech Challenge Track 1 Test Set Track 1 test set for the NVVSpeech Challenge at ISCSLP 2026. Task Given a speech recording, produce a transcript that contains the spoken content and the non-verbal vocalization (NVV) tags at their corresponding positions. Dataset Summary Language Samples Chinese 985 English 961 Total 1,946 Files . ├── README.md ├── SUBMISSION_GUIDE.txt ├── test.jsonl ├── ground_truth.jsonl ├──… See the full description on the dataset page: https://huggingface.co/datasets/NVVSpeech-Challenge/NVVSpeech-Challenge-Track1-Test-Set.audioautomatic-speech-recognition1K<n<10K0 likes106 downloads12d agoHugging Face25HiTZ /benchmark_eseu_testsets Benchmark Test-sets for evaluations on Spanish and Basque This test-sets are a reduced version of public available datasets. The datasets are balanced with more or less the same amount of hours in each dataset, for equal evaluation tasks. Test splits: mozilla-foundation/common_voice_18_0/es: a small split made from the official "test" split for spanish. mozilla-foundation/common_voice_18_0/eu: a small split made from the official "test" split for basque. openslr/es: a… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/benchmark_eseu_testsets.automatic-speech-recognition0 likes103 downloads1y agoHugging Face26anjalyjayakrishnan /testThe Snow Mountain dataset contains the audio recordings (in .mp3 format) and the corresponding text of The Bible in 11 Indian languages. The recordings were done in a studio setting by native speakers. Each language has a single speaker in the dataset. Most of these languages are geographically concentrated in the Northern part of India around the state of Himachal Pradesh. Being related to Hindi they all use the Devanagari script for transcription.audioautomatic-speech-recognition1K<n<10K0 likes100 downloads4y agoHugging Face27shakods /voxtral-synthetic-eng-test Voxtral Synthetic English (ASR) Synthetic speech dataset for fine-tuning Voxtral ASR models. English utterances generated with ElevenLabs TTS from the CohereLabs/aya_collection_language_split (english, targets column). All audio is 16 kHz mono WAV. Dataset structure Column Type Description audio_path string Path to the audio file in this repo (e.g. audio/utt_000000.wav) text string Ground-truth transcript for the audio Audio: 16 kHz, mono, WAV, stored… See the full description on the dataset page: https://huggingface.co/datasets/shakods/voxtral-synthetic-eng-test.audioautomatic-speech-recognitionn<1K1 likes96 downloads7mo agoHugging Face28Yehor /RS-testaudioautomatic-speech-recognition10K<n<100K0 likes96 downloads3mo agoHugging Face29ggfox00000 /stt-summre-fr-test SUMM-RE — French test split (mirror of linagora/SUMM-RE) Mirror public du split test de SUMM-RE (LINAGORA / Aix-Marseille LPL), pour benchmark ASR français conversationnel (parole de réunion, 3-4 locuteurs, ~20 min par session). ⚠ Ce repo ne contient que le split test (124 tracks individuelles = 37 réunions × 3-4 micros). Pour les splits train / dev, voir le repo upstream linagora/SUMM-RE. Contenu 124 pistes audio individuelles (1 piste = 1 microphone d'un locuteur… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/stt-summre-fr-test.audioautomatic-speech-recognitionn<1K0 likes93 downloads6mo agoHugging Face30speechcolab /gigaspeech2-test Dataset Card for GigaSpeech 2 TEST Dataset Description GigaSpeech 2 is an evolving, large-scale, multi-domain, and multilingual ASR corpus focusing on low-resource languages. GigaSpeech 2 raw comprises about 30,000 hours of automatically transcribed speech, across Thai, Indonesian, and Vietnamese. GigaSpeech 2 refine consists of 10,000 hours of Thai, 6,000 hours each for Indonesian and Vietnamese. Repository: https://github.com/SpeechColab/GigaSpeech2 Paper:… See the full description on the dataset page: https://huggingface.co/datasets/speechcolab/gigaspeech2-test.automatic-speech-recognition1M<n<10M0 likes92 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.