Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MiniMaxAI /TTS-Multilingual-Test-Set Overview To assess the multilingual zero-shot voice cloning capabilities of TTS models, we have constructed a test set encompassing 24 languages. This dataset provides both audio samples for voice cloning and corresponding test texts. Specifically, the test set for each language includes: 100 distinct test sentences. Audio samples from two speakers (one male and one female) carefully selected from the Mozilla Common Voice (MCV) dataset, intended for voice cloning. Researchers can… See the full description on the dataset page: https://huggingface.co/datasets/MiniMaxAI/TTS-Multilingual-Test-Set.audiotext-to-speechn<1K47 likes1.1k downloads1y agoHugging Face02masumtechnonext /test-data-set-Arabic-letteraudio10K<n<100K0 likes585 downloads2mo agoHugging Face03bandad /asr-testset-kw-ja-v1 日本語 ASR アノテーション v1 重要:評価結果を報告する際の規約 本テストセットで評価結果を報告する際は、事前学習を含む学習にYouTubeの音声を使用したかどうかと、次の評価区分を必ず明記してください。 使用した場合:in-domain評価 使用していない場合:out-of-domain評価 この区分は、音声認識の評価結果を比較する上で重要です。最終的な追加学習だけでなく、使用するモデルの事前学習も含めて判断してください。 元データセットの音声に、人手で区間ごとの転記・タグ・採否を付けたデータです。音声の内容、ディレクトリ構成、ファイル名は元のままです。 ファイルと表示 **annotations.jsonl**:提出済みの全結果。1行が1音声です。 **metadata.jsonl**:HF表示用に自動生成したデータ。音声全体が不使用の行を除き、audio を file_name に置き換えています。 **data/**:採用した音声ファイル。 HFのビューアーは… See the full description on the dataset page: https://huggingface.co/datasets/bandad/asr-testset-kw-ja-v1.audioautomatic-speech-recognitionn<1K0 likes494 downloads5d agoHugging Face04CocoBro /MMEdit-TestSet MMEdit Test Set A paired audio editing test set for text-guided audio manipulation evaluation, released with MMEdit. Overview This dataset contains 3,317 aligned triplets: Component Description raw/ Source audio before editing target/ Target audio after editing content.jsonl Editing instruction (caption) keyed by audio_id Each sample is linked by a shared audio_id. For example, sample add_017221 corresponds to: raw/add_017221.wav — original… See the full description on the dataset page: https://huggingface.co/datasets/CocoBro/MMEdit-TestSet.audioaudio-to-audio1K<n<10K0 likes389 downloads4mo agoHugging Face05Eureka-Leo /MCABSA_testsetaudion<1K2 likes306 downloads1y agoHugging Face06AnonyData /Continuo-Testset Continuo-Testset Continuo-Testset is a benchmark for long-form, multi-speaker zero-shot speech generation. Each case asks a system to synthesize a complete speaker-attributed script as one recording, given a reference prompt for every speaker. The output is scored against a human-verified target recording. This dataset accompanies an anonymous ICLR 2027 submission. Statistics Chinese (zh) English (en) Total Cases 59 59 118 Single-speaker long-form 18… See the full description on the dataset page: https://huggingface.co/datasets/AnonyData/Continuo-Testset.audiotext-to-speechn<1K0 likes156 downloads14d agoHugging Face07bandad /asr-testset-kw-jaaudio1K<n<10K0 likes144 downloads14d agoHugging Face08NVVSpeech-Challenge /NVVSpeech-Challenge-Track1-Test-Set NVVSpeech Challenge Track 1 Test Set Track 1 test set for the NVVSpeech Challenge at ISCSLP 2026. Task Given a speech recording, produce a transcript that contains the spoken content and the non-verbal vocalization (NVV) tags at their corresponding positions. Dataset Summary Language Samples Chinese 985 English 961 Total 1,946 Files . ├── README.md ├── SUBMISSION_GUIDE.txt ├── test.jsonl ├── ground_truth.jsonl ├──… See the full description on the dataset page: https://huggingface.co/datasets/NVVSpeech-Challenge/NVVSpeech-Challenge-Track1-Test-Set.audioautomatic-speech-recognition1K<n<10K0 likes106 downloads12d agoHugging Face09XRXRX /X-Voice-TestsetX-Voice Multilingual Test Set High-Fidelity Test Set for Multilingual Text-to-Speech across 30 Languages This test set is built as part of the research: X-Voice: One Speaker, 30+ Languages with Zero-Shot Voice Cloning, serving as the evaluation benchmark for our model. Dataset Summary 30 languages European: bg (Bulgarian), cs (Czech), da (Danish), de (German), el (Greek), en (English), es (Spanish), et (Estonian), fi (Finnish), fr (French), hr (Croatian), hu (Hungarian), it… See the full description on the dataset page: https://huggingface.co/datasets/XRXRX/X-Voice-Testset.audiotext-to-speech10K<n<100K4 likes87 downloads5mo agoHugging Face10YoonSeon /TTS-Multilingual-Test-Set Overview To assess the multilingual zero-shot voice cloning capabilities of TTS models, we have constructed a test set encompassing 24 languages. This dataset provides both audio samples for voice cloning and corresponding test texts. Specifically, the test set for each language includes: 100 distinct test sentences. Audio samples from two speakers (one male and one female) carefully selected from the Mozilla Common Voice (MCV) dataset, intended for voice cloning. Researchers can… See the full description on the dataset page: https://huggingface.co/datasets/YoonSeon/TTS-Multilingual-Test-Set.audiotext-to-speechn<1K0 likes71 downloads7mo agoHugging Face11KYAGABA /kinyarwanda_cleaned_testset_verified_20HRSaudio10K<n<100K0 likes68 downloads2y agoHugging Face12DimQ1 /sortformer-diarization-test-set Sortformer Diarization Test Set 100 real speech samples extracted from LibriSpeech test-clean for speaker diarization testing and benchmarking with NVIDIA Sortformer 4spk-v2 ONNX models. Usage with Sortformer ONNX from huggingface_hub import snapshot_download import soundfile as sf # Download the test set dataset_path = snapshot_download("DimQ1/sortformer-diarization-test-set") # Load audio audio, sr = sf.read(f"{dataset_path}/audio/ls_real_000.wav")… See the full description on the dataset page: https://huggingface.co/datasets/DimQ1/sortformer-diarization-test-set.audioaudio-classificationn<1K0 likes65 downloads3mo agoHugging Face13NVVSpeech-Challenge /NVVSpeech-Challenge-Track2-Test-Set NVVSpeech Challenge Track 2 Test Set Track 2 test set for the NVVSpeech Challenge at ISCSLP 2026. Task Given a transcript containing one or more non-verbal vocalization (NVV) tags, synthesize speech that naturally realizes the requested NVVs while preserving intelligibility, naturalness, and audio quality. Dataset Summary Language Samples Chinese 800 English 800 Total 1,600 Files . ├── README.md ├──… See the full description on the dataset page: https://huggingface.co/datasets/NVVSpeech-Challenge/NVVSpeech-Challenge-Track2-Test-Set.texttext-to-speech1K<n<10K0 likes63 downloads12d agoHugging Face14UmaSubhashiniRavuri1 /BandFlex_Testsetcat > README.md << 'EOF' BandFlex Testset Test sets for BandFlex bandwidth extension, built from VCTK and Expresso. Download VCTK Testset (~2.4 GB): included in this GitHub repo (VCTK_Testset/) Full dataset (VCTK + Expresso, ~22 GB): hosted on Hugging Face 👉 https://huggingface.co/datasets/UmaSubhashiniRavuri1/BandFlex_Testset Download from Hugging Face: pip install -U huggingface_hub hf download UmaSubhashiniRavuri1/BandFlex_Testset --repo-type=dataset… See the full description on the dataset page: https://huggingface.co/datasets/UmaSubhashiniRavuri1/BandFlex_Testset.audio100K<n<1M1 likes63 downloads15d agoHugging Face15Yehor /cv10-uk-testset-clean The cleaned Common Voice 10 (test set) that has been checked by a human for Ukrainian 🇺🇦 Overview This repository contains the archive of Common Voice 10 (test set) with checked Ukrainian transcriptions and audios. All audios have been checked by a human to be sure that they are correct. This archive is used to test all ASR models listed here: https://github.com/egorsmkv/speech-recognition-uk Community Discord: https://bit.ly/discord-uds Speech… See the full description on the dataset page: https://huggingface.co/datasets/Yehor/cv10-uk-testset-clean.audioautomatic-speech-recognition1K<n<10K3 likes61 downloads2y agoHugging Face16wanasash /corpus-siarad-test-setaudio1K<n<10K0 likes40 downloads2y agoHugging Face17jeshica /TTS-Multilingual-Test-Set Overview To assess the multilingual zero-shot voice cloning capabilities of TTS models, we have constructed a test set encompassing 24 languages. This dataset provides both audio samples for voice cloning and corresponding test texts. Specifically, the test set for each language includes: 100 distinct test sentences. Audio samples from two speakers (one male and one female) carefully selected from the Mozilla Common Voice (MCV) dataset, intended for voice cloning. Researchers can… See the full description on the dataset page: https://huggingface.co/datasets/jeshica/TTS-Multilingual-Test-Set.audiotext-to-speechn<1K0 likes40 downloads7mo agoHugging Face18KYAGABA /amharic_cleaned_testset_verifiedaudio10K<n<100K1 likes37 downloads2y agoHugging Face19KYAGABA /kinyarwanda_cleaned_testset_verified_200HRSaudio100K<n<1M0 likes37 downloads2y agoHugging Face20otozz /MSA_test_setPre-processed MSA data based on https://huggingface.co/datasets/mozilla-foundation/common_voice_16_1. audio10K<n<100K0 likes35 downloads2y agoHugging Face21hoangducanh1865 /vi-en-ast-testsetaudio1K<n<10K0 likes32 downloads1mo agoHugging Face22KYAGABA /amharic_cleaned_testset_fleurs_currentaudio1K<n<10K0 likes29 downloads2y agoHugging Face23LGB666 /SageLM_testset_audioaudion<1K0 likes29 downloads1y agoHugging Face24prvInSpace /eval_framework_testsetaudio1K<n<10K0 likes28 downloads1y agoHugging Face25danielrosehill /Audio-Understanding-Test-Set Audio Understanding Test Set A structured dataset for evaluating audio understanding capabilities of multimodal AI models. Contains 137 test prompts across 22 categories, paired with a 20-minute voice sample and 49 completed model outputs from Gemini 3.1 Flash Lite. Overview Property Value Total prompts 137 Implemented (with prompt text) 49 Suggested (description only) 88 Completed outputs 49 Categories 22 Model under test… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Audio-Understanding-Test-Set.audioaudio-classificationn<1K0 likes28 downloads7mo agoHugging Face26hoangducanh1865 /en-vi-ast-testsetaudio1K<n<10K0 likes23 downloads1mo agoHugging Face27heimayuan /wuw_testset1gated Yougen/wuw_testset1 Wake-Up-Word (WUW) speech dataset, packed as WebDataset tar shards. The input is a Kaldi-style data directory (wav.scp, text, utt2spk, utt2dur, segments), where multiple utterances share a long recording via the segments file. To avoid duplicating audio, each tar sample corresponds to one full recording. The utterance-level metadata (id / start / end / text / spk / duration) is stored in a JSON list inside that sample. Downstream consumers slice the decoded… See the full description on the dataset page: https://huggingface.co/datasets/heimayuan/wuw_testset1.audioautomatic-speech-recognition0 likes22 downloads2mo agoHugging Face28KYAGABA /kinyarwanda_cleaned_testset_verifiedaudio100K<n<1M0 likes21 downloads2y agoHugging Face29Sammau /test_setaudio1K<n<10K0 likes20 downloads1y agoHugging Face30KYAGABA /kinyarwanda_cleaned_testset_verified_10HRSaudio1K<n<10K0 likes19 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.