Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01malaiwah /qwen3-tts-customvoice-ab-clips qwen3-tts: full 5-way cloning comparison + cross-row diagnostic Generated 2026-04-14 on RTX 4080 SUPER. Directories original/ CustomVoice.generate_custom_voice(speaker=X) -> the ground truth voice clone/ Base.generate_voice_clone(ref_audio=original.wav, ref_text=...) -> full ICL clone via Base's own speaker encoder transplant/ Base.generate_voice_clone(voice_clone_prompt=[row]) x_vector_only_mode=True… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qwen3-tts-customvoice-ab-clips.audion<1K0 likes700 downloads6mo agoHugging Face02masuidrive /cv-corpus-17.0-zh-CN-client_id-grouped cv-corpus-17.0-zh-CN-client_id-grouped This dataset is a subset of the Common Voice dataset, filtered and grouped based on the client ID (treated as speaker ID). Dataset Details The dataset is derived from the Common Voice dataset. The original dataset is available at Common Voice Dataset. The dataset is grouped by client ID, which is treated as the speaker ID for this dataset. Each group is filtered to include only client IDs with a minimum of 30 samples and a maximum… See the full description on the dataset page: https://huggingface.co/datasets/masuidrive/cv-corpus-17.0-zh-CN-client_id-grouped.audioautomatic-speech-recognition100K<n<1M3 likes475 downloads2y agoHugging Face03masuidrive /cv-corpus-17.0-zh-TW-client_id-grouped cv-corpus-17.0-zh-TW-client_id-grouped This dataset is a subset of the Common Voice dataset, filtered and grouped based on the client ID (treated as speaker ID). Dataset Details The dataset is derived from the Common Voice dataset. The original dataset is available at Common Voice Dataset. The dataset is grouped by client ID, which is treated as the speaker ID for this dataset. Each group is filtered to include only client IDs with a minimum of 30 samples and a maximum… See the full description on the dataset page: https://huggingface.co/datasets/masuidrive/cv-corpus-17.0-zh-TW-client_id-grouped.audioautomatic-speech-recognition10K<n<100K1 likes384 downloads2y agoHugging Face04masuidrive /cv-corpus-17.0-ja-client_id-grouped cv-corpus-17.0-ja-client_id-grouped This dataset is a subset of the Common Voice dataset, filtered and grouped based on the client ID (treated as speaker ID). Dataset Details The dataset is derived from the Common Voice dataset. The original dataset is available at Common Voice Dataset. The dataset is grouped by client ID, which is treated as the speaker ID for this dataset. Each group is filtered to include only client IDs with a minimum of 30 samples and a maximum of… See the full description on the dataset page: https://huggingface.co/datasets/masuidrive/cv-corpus-17.0-ja-client_id-grouped.audioautomatic-speech-recognition10K<n<100K2 likes158 downloads2y agoHugging Face05imtiyaz517 /aisha-urdu-voice-clipsaudion<1K0 likes142 downloads6d agoHugging Face06erzhanbakanbayev /kk-yt-clipsaudio1K<n<10K0 likes136 downloads1y agoHugging Face07GiusMagi /spyken-clipsaudio10K<n<100K0 likes104 downloads25d agoHugging Face08sieep55 /wildlands-clientaudio10K<n<100K0 likes89 downloads4mo agoHugging Face09masuidrive /cv-corpus-1.0-en-client_id-grouped cv-corpus-1.0-en-client_id-grouped This dataset is a subset of the Common Voice dataset, filtered and grouped based on the client ID (treated as speaker ID). Dataset Details The dataset is derived from the Common Voice dataset. The original dataset is available at Common Voice Dataset. The dataset is grouped by client ID, which is treated as the speaker ID for this dataset. Each group is filtered to include only client IDs with a minimum of 60 samples and a maximum of… See the full description on the dataset page: https://huggingface.co/datasets/masuidrive/cv-corpus-1.0-en-client_id-grouped.audioautomatic-speech-recognition100K<n<1M1 likes85 downloads2y agoHugging Face10Blasteur19 /data-clipaudion<1K0 likes68 downloads4mo agoHugging Face11knoriy /OE-DCT-Movie-clipstabular10K<n<100K0 likes55 downloads3y agoHugging Face12khamidov17 /clinical-conversations-anon-benchmarkaudion<1K0 likes53 downloads3mo agoHugging Face13Jingya /video-clips-and-imgsaudion<1K0 likes52 downloads2mo agoHugging Face14HamdanXI /uclass_clipped_labeled Dataset Card for "uclass_clipped_labeled" More Information needed audio1K<n<10K0 likes46 downloads2y agoHugging Face15yongjian /music-clips-50There are 50 music clips(of 3~5 seconds). You can load them by the following code: from datasets import load_dataset dataset = load_dataset('yongjian/music-clips-50') clips = dataset['train'] # all 50 music clips music_1_np_array = clips[0]['audio']['array'] # numpy array of shape=[N,] Or you can directly download them from Google Drive: music-clips-50.tar.gz. audion<1K3 likes45 downloads4y agoHugging Face16Harmonic-Frontier-Audio /Mouth_Clicks_Smacks_and_Pop_Articulations_Preview Harmonic Frontier Audio – Mouth Clicks, Smacks and Pop Articulations (Preview, v0.9) A high-fidelity human vocal dataset designed for AI training, speech research, and expressive voice modeling. Mouth Clicks, Smacks and Pop Articulations (Preview), created by Harmonic Frontier Audio, provides a compact reference set demonstrating the quality, formatting, and metadata conventions used in the Harmonic Frontier Audio Human Vocality Primitives series. 🔎 Summary This… See the full description on the dataset page: https://huggingface.co/datasets/Harmonic-Frontier-Audio/Mouth_Clicks_Smacks_and_Pop_Articulations_Preview.audioothern<1K2 likes31 downloads7mo agoHugging Face17isabelarvelo /sep28k-train-4-second-clips Dataset Card for "sep28k-train-4-second-clips" More Information needed audio1K<n<10K0 likes29 downloads2y agoHugging Face18JBJoyce /DENTAL_CLICK Dataset Card for "DENTAL_CLICK" More Information needed audio1K<n<10K1 likes26 downloads4y agoHugging Face19BonusLockSMith /1776-track-clips 1776 Track Clips Fast, wordless, loop-safe background music built for Shorts, Reels, and edits. This dataset contains 1776 unique audio clips designed for creators, editors, and developers who need clean, reusable background music without lyrics. What’s included 1776 unique clips No lyrics (wordless hooks) Clean, loop-safe structure Optimized for short-form video Editor-first design Preview See the preview video(s) in this repository for examples across… See the full description on the dataset page: https://huggingface.co/datasets/BonusLockSMith/1776-track-clips.audio1K<n<10K0 likes25 downloads10mo agoHugging Face20noflm /jdd_topic1_20251224-cliponly_sample100audio1K<n<10K0 likes24 downloads10mo agoHugging Face21clifemall /audios-inglesaudion<1K0 likes24 downloads5d agoHugging Face22isabelarvelo /sep28k-train-5-second-clips Dataset Card for "sep28k-train-5-second-clips" More Information needed audio1K<n<10K0 likes23 downloads2y agoHugging Face23burakozdelen /CommonVoice_TR_clips_16khz_g711_White_noiseaudio0 likes23 downloads2mo agoHugging Face24isabelarvelo /sep28k-dev-4-second-clips Dataset Card for "sep28k-dev-0120-4-second-clips" More Information needed audio1K<n<10K0 likes21 downloads2y agoHugging Face25isabelarvelo /sep28k-test-5-second-clips Dataset Card for "sep28k-test-5-second-clips" More Information needed audio1K<n<10K0 likes20 downloads2y agoHugging Face26CLiC-UB /rapnic-examplegated RAPNIC Dataset (example) Dataset Description This is an example of the full dataset, yet to be published, with 10 audio examples for 72 speakers. RAPNIC (Reconeixement Automàtic de la Parla No Intel·ligible en Català) is a Catalan speech corpus collected from individuals with speech disorders, specifically cerebral palsy and Down syndrome. This dataset was collected to develop and improve automatic speech recognition (ASR) systems that are accessible to people with speech… See the full description on the dataset page: https://huggingface.co/datasets/CLiC-UB/rapnic-example.audioautomatic-speech-recognition1K<n<10K0 likes20 downloads5mo agoHugging Face27isabelarvelo /sep28k-dev-5-second-clips Dataset Card for "sep28k-dev-5-second-clips" More Information needed audio1K<n<10K0 likes18 downloads2y agoHugging Face28isabelarvelo /fluencybank-3-second-clips Dataset Card for "fluencybank-3-second-clips" More Information needed audio1K<n<10K0 likes18 downloads2y agoHugging Face29isabelarvelo /sep28k-train-3-second-clips-full-agreement Dataset Card for "sep28k-train-3-second-clips-full-agreement" More Information needed audio1K<n<10K0 likes18 downloads2y agoHugging Face30Diffusion-ASR /worst100-testclean-clips Worst-100 test-clean clips — audio, transcripts, and the vocabulary finding The 100 LibriSpeech test-clean clips where the block-4 production model (4.60% WER) made the most word errors — with audio embedded so the failures can be listened to, plus the model's transcript next to the reference for each clip. The finding this dataset produced 47% of the word errors in these clips are on words that never appeared in the 30-hour training vocabulary at all (20,066… See the full description on the dataset page: https://huggingface.co/datasets/Diffusion-ASR/worst100-testclean-clips.audion<1K0 likes18 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.