Team Ai
25 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ASLP-lab /Easy-Turn-Testset Easy Turn: Integrating Acoustic and Linguistic Modalities for Robust Turn-Taking in Full-Duplex Spoken Dialogue Systems Guojian Li1, Chengyou Wang1, Hongfei Xue1, Shuiyuan Wang1, Dehui Gao1, Zihan Zhang2, Yuke Lin2, Wenjie Li2, Longshuai Xiao2, Zhonghua Fu1,╀, Lei Xie1,╀ 1 Audio, Speech and Language Processing Group (ASLP@NPU), Northwestern Polytechnical University 2 Huawei Technologies, China 🎤 Demo Page 🤖 Easy Turn Model 📑 Paper 🌐 Huggingface… See the full description on the dataset page: https://huggingface.co/datasets/ASLP-lab/Easy-Turn-Testset.automatic-speech-recognition8 likes1.5k downloads1y agoHugging Face02bandad /asr-testset-kw-ja-v1 日本語 ASR アノテーション v1 重要:評価結果を報告する際の規約 本テストセットで評価結果を報告する際は、事前学習を含む学習にYouTubeの音声を使用したかどうかと、次の評価区分を必ず明記してください。 使用した場合:in-domain評価 使用していない場合:out-of-domain評価 この区分は、音声認識の評価結果を比較する上で重要です。最終的な追加学習だけでなく、使用するモデルの事前学習も含めて判断してください。 元データセットの音声に、人手で区間ごとの転記・タグ・採否を付けたデータです。音声の内容、ディレクトリ構成、ファイル名は元のままです。 ファイルと表示 **annotations.jsonl**:提出済みの全結果。1行が1音声です。 **metadata.jsonl**:HF表示用に自動生成したデータ。音声全体が不使用の行を除き、audio を file_name に置き換えています。 **data/**:採用した音声ファイル。 HFのビューアーは… See the full description on the dataset page: https://huggingface.co/datasets/bandad/asr-testset-kw-ja-v1.audioautomatic-speech-recognitionn<1K0 likes494 downloads5d agoHugging Face031xg /Easy-Turn-Testset Easy Turn: Integrating Acoustic and Linguistic Modalities for Robust Turn-Taking in Full-Duplex Spoken Dialogue Systems Guojian Li1, Chengyou Wang1, Hongfei Xue1, Shuiyuan Wang1, Dehui Gao1, Zihan Zhang2, Yuke Lin2, Wenjie Li2, Longshuai Xiao2, Zhonghua Fu1,╀, Lei Xie1,╀ 1 Audio, Speech and Language Processing Group (ASLP@NPU), Northwestern Polytechnical University 2 Huawei Technologies, China 🎤 Demo Page 🤖 Easy Turn Model 📑 Paper 🌐 Huggingface… See the full description on the dataset page: https://huggingface.co/datasets/1xg/Easy-Turn-Testset.automatic-speech-recognition0 likes242 downloads2mo agoHugging Face04NVVSpeech-Challenge /NVVSpeech-Challenge-Track1-Test-Set NVVSpeech Challenge Track 1 Test Set Track 1 test set for the NVVSpeech Challenge at ISCSLP 2026. Task Given a speech recording, produce a transcript that contains the spoken content and the non-verbal vocalization (NVV) tags at their corresponding positions. Dataset Summary Language Samples Chinese 985 English 961 Total 1,946 Files . ├── README.md ├── SUBMISSION_GUIDE.txt ├── test.jsonl ├── ground_truth.jsonl ├──… See the full description on the dataset page: https://huggingface.co/datasets/NVVSpeech-Challenge/NVVSpeech-Challenge-Track1-Test-Set.audioautomatic-speech-recognition1K<n<10K0 likes106 downloads12d agoHugging Face05HiTZ /benchmark_eseu_testsets Benchmark Test-sets for evaluations on Spanish and Basque This test-sets are a reduced version of public available datasets. The datasets are balanced with more or less the same amount of hours in each dataset, for equal evaluation tasks. Test splits: mozilla-foundation/common_voice_18_0/es: a small split made from the official "test" split for spanish. mozilla-foundation/common_voice_18_0/eu: a small split made from the official "test" split for basque. openslr/es: a… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/benchmark_eseu_testsets.automatic-speech-recognition0 likes103 downloads1y agoHugging Face06XRXRX /X-Voice-TestsetX-Voice Multilingual Test Set High-Fidelity Test Set for Multilingual Text-to-Speech across 30 Languages This test set is built as part of the research: X-Voice: One Speaker, 30+ Languages with Zero-Shot Voice Cloning, serving as the evaluation benchmark for our model. Dataset Summary 30 languages European: bg (Bulgarian), cs (Czech), da (Danish), de (German), el (Greek), en (English), es (Spanish), et (Estonian), fi (Finnish), fr (French), hr (Croatian), hu (Hungarian), it… See the full description on the dataset page: https://huggingface.co/datasets/XRXRX/X-Voice-Testset.audiotext-to-speech10K<n<100K4 likes87 downloads5mo agoHugging Face07Yehor /cv10-uk-testset-clean The cleaned Common Voice 10 (test set) that has been checked by a human for Ukrainian 🇺🇦 Overview This repository contains the archive of Common Voice 10 (test set) with checked Ukrainian transcriptions and audios. All audios have been checked by a human to be sure that they are correct. This archive is used to test all ASR models listed here: https://github.com/egorsmkv/speech-recognition-uk Community Discord: https://bit.ly/discord-uds Speech… See the full description on the dataset page: https://huggingface.co/datasets/Yehor/cv10-uk-testset-clean.audioautomatic-speech-recognition1K<n<10K3 likes61 downloads2y agoHugging Face08danielrosehill /Audio-Understanding-Test-Set Audio Understanding Test Set A structured dataset for evaluating audio understanding capabilities of multimodal AI models. Contains 137 test prompts across 22 categories, paired with a 20-minute voice sample and 49 completed model outputs from Gemini 3.1 Flash Lite. Overview Property Value Total prompts 137 Implemented (with prompt text) 49 Suggested (description only) 88 Completed outputs 49 Categories 22 Model under test… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Audio-Understanding-Test-Set.audioaudio-classificationn<1K0 likes28 downloads7mo agoHugging Face09heimayuan /wuw_testset1gated Yougen/wuw_testset1 Wake-Up-Word (WUW) speech dataset, packed as WebDataset tar shards. The input is a Kaldi-style data directory (wav.scp, text, utt2spk, utt2dur, segments), where multiple utterances share a long recording via the segments file. To avoid duplicating audio, each tar sample corresponds to one full recording. The utterance-level metadata (id / start / end / text / spk / duration) is stored in a JSON list inside that sample. Downstream consumers slice the decoded… See the full description on the dataset page: https://huggingface.co/datasets/heimayuan/wuw_testset1.audioautomatic-speech-recognition0 likes22 downloads2mo agoHugging Face10xuyaya /ASR-Testset Chinese ASR Open-Source Testsets (中文语音识别开源测试集汇总) 本仓库汇总了当前中文语音识别(ASR)领域最主流的开源测试集,涵盖了从通用场景、会议场景到多方言场景的全面评估维度。 1. 基础通用领域测试集 这些是学术界和工业界最常用的标准测试集: Aishell-1: 经典的中文开源语音数据库,由 400 人录制,测试集包含了清晰的普通话。 Aishell-2: 规模比 Aishell-1 更大,录制环境更严谨(使用 iPhone 录制),是评估中文模型性能的基石。 Aidatatatang_200zh: 由数据堂开源的 200 小时数据集的测试部分,包含多种移动端录制场景。 2. 行业与特定场景测试集 针对办公、远场及专业场景的针对性评估: ws-net / ws-meeting: 侧重于会议场景、职场交流,通常包含较多的环境噪声或多人交谈。 Test_Ali_near / far:… See the full description on the dataset page: https://huggingface.co/datasets/xuyaya/ASR-Testset.automatic-speech-recognition0 likes17 downloads9mo agoHugging Face11hongshaoyo /Easy-Turn-Testset Easy Turn: Integrating Acoustic and Linguistic Modalities for Robust Turn-Taking in Full-Duplex Spoken Dialogue Systems Guojian Li1, Chengyou Wang1, Hongfei Xue1, Shuiyuan Wang1, Dehui Gao1, Zihan Zhang2, Yuke Lin2, Wenjie Li2, Longshuai Xiao2, Zhonghua Fu1,╀, Lei Xie1,╀ 1 Audio, Speech and Language Processing Group (ASLP@NPU), Northwestern Polytechnical University 2 Huawei Technologies, China 🎤 Demo Page 🤖 Easy Turn Model 📑 Paper 🌐 Huggingface… See the full description on the dataset page: https://huggingface.co/datasets/hongshaoyo/Easy-Turn-Testset.automatic-speech-recognition0 likes17 downloads6mo agoHugging Face12ygyuan /kws_testset_ug_huitinggated ygyuan/kws_testset_ug_huiting Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards. The input is a Kaldi-style data directory (wav.scp, text, utt2spk, utt2dur, segments), where each utterance is packed as a single tar sample. Layout data/ <split>/ metadata.csv audio/ <split>-000.tar <split>-001.tar ... Shard counts: test: 1 tar shard(s) Inside each tar, every sample is a pair sharing a unique key: <key>.wav #… See the full description on the dataset page: https://huggingface.co/datasets/ygyuan/kws_testset_ug_huiting.audioaudio-classification0 likes14 downloads2mo agoHugging Face13ygyuan /kws_testset_bo_sphgated ygyuan/kws_testset_bo_sph Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards. The input is a Kaldi-style data directory (wav.scp, text, utt2spk, utt2dur, segments), where each utterance is packed as a single tar sample. Layout data/ <split>/ metadata.csv audio/ <split>-000.tar <split>-001.tar ... Shard counts: test: 1 tar shard(s) Inside each tar, every sample is a pair sharing a unique key: <key>.wav # raw… See the full description on the dataset page: https://huggingface.co/datasets/ygyuan/kws_testset_bo_sph.audioaudio-classification0 likes13 downloads2mo agoHugging Face14xingluran /Easy-Turn-Testset Easy Turn: Integrating Acoustic and Linguistic Modalities for Robust Turn-Taking in Full-Duplex Spoken Dialogue Systems Guojian Li1, Chengyou Wang1, Hongfei Xue1, Shuiyuan Wang1, Dehui Gao1, Zihan Zhang2, Yuke Lin2, Wenjie Li2, Longshuai Xiao2, Zhonghua Fu1,╀, Lei Xie1,╀ 1 Audio, Speech and Language Processing Group (ASLP@NPU), Northwestern Polytechnical University 2 Huawei Technologies, China 🎤 Demo Page 🤖 Easy Turn Model 📑 Paper 🌐 Huggingface… See the full description on the dataset page: https://huggingface.co/datasets/xingluran/Easy-Turn-Testset.automatic-speech-recognition0 likes12 downloads8mo agoHugging Face15ygyuan /kws_testset_bo_yalugated ygyuan/kws_testset_bo_yalu Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards. The input is a Kaldi-style data directory (wav.scp, text, utt2spk, utt2dur, segments), where each utterance is packed as a single tar sample. Layout data/ <split>/ metadata.csv audio/ <split>-000.tar <split>-001.tar ... Shard counts: test: 2 tar shard(s) Inside each tar, every sample is a pair sharing a unique key: <key>.wav # raw… See the full description on the dataset page: https://huggingface.co/datasets/ygyuan/kws_testset_bo_yalu.audioaudio-classification0 likes12 downloads2mo agoHugging Face16ygyuan /kws_testset_digated ygyuan/kws_testset_di Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards. The input is a Kaldi-style data directory (wav.scp, text, utt2spk, utt2dur, segments), where each utterance is packed as a single tar sample. Layout data/ <split>/ metadata.csv audio/ <split>-000.tar <split>-001.tar ... Shard counts: test: 1 tar shard(s) Inside each tar, every sample is a pair sharing a unique key: <key>.wav # raw… See the full description on the dataset page: https://huggingface.co/datasets/ygyuan/kws_testset_di.audioaudio-classification0 likes11 downloads2mo agoHugging Face17ygyuan /kws_testset_bo_huitinggated ygyuan/kws_testset_bo_huiting Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards. The input is a Kaldi-style data directory (wav.scp, text, utt2spk, utt2dur, segments), where each utterance is packed as a single tar sample. Layout data/ <split>/ metadata.csv audio/ <split>-000.tar <split>-001.tar ... Shard counts: test: 2 tar shard(s) Inside each tar, every sample is a pair sharing a unique key: <key>.wav #… See the full description on the dataset page: https://huggingface.co/datasets/ygyuan/kws_testset_bo_huiting.audioaudio-classification0 likes11 downloads2mo agoHugging Face18ygyuan /kws_testset_ct_sphgated ygyuan/kws_testset_ct_sph Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards. The input is a Kaldi-style data directory (wav.scp, text, utt2spk, utt2dur, segments), where each utterance is packed as a single tar sample. Layout data/ <split>/ metadata.csv audio/ <split>-000.tar <split>-001.tar ... Shard counts: test: 1 tar shard(s) Inside each tar, every sample is a pair sharing a unique key: <key>.wav # raw… See the full description on the dataset page: https://huggingface.co/datasets/ygyuan/kws_testset_ct_sph.audioaudio-classification0 likes10 downloads2mo agoHugging Face19ygyuan /kws_testset_kkgated ygyuan/kws_testset_kk Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards. The input is a Kaldi-style data directory (wav.scp, text, utt2spk, utt2dur, segments), where each utterance is packed as a single tar sample. Layout data/ <split>/ metadata.csv audio/ <split>-000.tar <split>-001.tar ... Shard counts: test_mht: 1 tar shard(s) test_thu: 1 tar shard(s) Inside each tar, every sample is a pair sharing a unique… See the full description on the dataset page: https://huggingface.co/datasets/ygyuan/kws_testset_kk.audioaudio-classification0 likes9 downloads2mo agoHugging Face20ygyuan /kws_testset_mn_huitinggated ygyuan/kws_testset_mn_huiting Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards. The input is a Kaldi-style data directory (wav.scp, text, utt2spk, utt2dur, segments), where each utterance is packed as a single tar sample. Layout data/ <split>/ metadata.csv audio/ <split>-000.tar <split>-001.tar ... Shard counts: test: 1 tar shard(s) Inside each tar, every sample is a pair sharing a unique key: <key>.wav #… See the full description on the dataset page: https://huggingface.co/datasets/ygyuan/kws_testset_mn_huiting.audioaudio-classification1 likes9 downloads2mo agoHugging Face21ygyuan /kws_testset_mn_sphgated ygyuan/kws_testset_mn_sph Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards. The input is a Kaldi-style data directory (wav.scp, text, utt2spk, utt2dur, segments), where each utterance is packed as a single tar sample. Layout data/ <split>/ metadata.csv audio/ <split>-000.tar <split>-001.tar ... Shard counts: test: 1 tar shard(s) Inside each tar, every sample is a pair sharing a unique key: <key>.wav # raw… See the full description on the dataset page: https://huggingface.co/datasets/ygyuan/kws_testset_mn_sph.audioaudio-classification0 likes9 downloads2mo agoHugging Face22ygyuan /kws_testset_ct_huitinggated ygyuan/kws_testset_ct_huiting Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards. The input is a Kaldi-style data directory (wav.scp, text, utt2spk, utt2dur, segments), where each utterance is packed as a single tar sample. Layout data/ <split>/ metadata.csv audio/ <split>-000.tar <split>-001.tar ... Shard counts: test: 3 tar shard(s) Inside each tar, every sample is a pair sharing a unique key: <key>.wav #… See the full description on the dataset page: https://huggingface.co/datasets/ygyuan/kws_testset_ct_huiting.audioaudio-classification0 likes9 downloads2mo agoHugging Face23ygyuan /kws_testset_ug_sphgated ygyuan/kws_testset_ug_sph Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards. The input is a Kaldi-style data directory (wav.scp, text, utt2spk, utt2dur, segments), where each utterance is packed as a single tar sample. Layout data/ <split>/ metadata.csv audio/ <split>-000.tar <split>-001.tar ... Shard counts: test: 1 tar shard(s) Inside each tar, every sample is a pair sharing a unique key: <key>.wav # raw… See the full description on the dataset page: https://huggingface.co/datasets/ygyuan/kws_testset_ug_sph.audioaudio-classification0 likes9 downloads2mo agoHugging Face24ygyuan /kws_testset_hmgated ygyuan/kws_testset_hm Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards. The input is a Kaldi-style data directory (wav.scp, text, utt2spk, utt2dur, segments), where each utterance is packed as a single tar sample. Layout data/ <split>/ metadata.csv audio/ <split>-000.tar <split>-001.tar ... Shard counts: test: 3 tar shard(s) Inside each tar, every sample is a pair sharing a unique key: <key>.wav # raw… See the full description on the dataset page: https://huggingface.co/datasets/ygyuan/kws_testset_hm.audioaudio-classification0 likes8 downloads2mo agoHugging Face25ygyuan /kws_testset_zh_s2tgated ygyuan/kws_testset_zh_s2t Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards. The input is a Kaldi-style data directory (wav.scp, text, utt2spk, utt2dur, segments), where each utterance is packed as a single tar sample. Layout data/ <split>/ metadata.csv audio/ <split>-000.tar <split>-001.tar ... Shard counts: test: 12 tar shard(s) Inside each tar, every sample is a pair sharing a unique key: <key>.wav # raw… See the full description on the dataset page: https://huggingface.co/datasets/ygyuan/kws_testset_zh_s2t.audioaudio-classification0 likes8 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.