datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
qwen3-tts-customvoice-ab-clips
qwen3-tts: full 5-way cloning comparison + cross-row diagnostic
Generated 2026-04-14 on RTX 4080 SUPER.
Directories
original/ CustomVoice.generate_custom_voice(speaker=X)
-> the ground truth voice
clone/ Base.generate_voice_clone(ref_audio=original.wav, ref_text=...)
-> full ICL clone via Base's own speaker encoder
transplant/ Base.generate_voice_clone(voice_clone_prompt=[row])
x_vector_only_mode=True… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qwen3-tts-customvoice-ab-clips.cv-corpus-17.0-zh-CN-client_id-grouped
cv-corpus-17.0-zh-CN-client_id-grouped
This dataset is a subset of the Common Voice dataset, filtered and grouped based on the client ID (treated as speaker ID).
Dataset Details
The dataset is derived from the Common Voice dataset.
The original dataset is available at Common Voice Dataset.
The dataset is grouped by client ID, which is treated as the speaker ID for this dataset.
Each group is filtered to include only client IDs with a minimum of 30 samples and a maximum… See the full description on the dataset page: https://huggingface.co/datasets/masuidrive/cv-corpus-17.0-zh-CN-client_id-grouped.cv-corpus-17.0-zh-TW-client_id-grouped
cv-corpus-17.0-zh-TW-client_id-grouped
This dataset is a subset of the Common Voice dataset, filtered and grouped based on the client ID (treated as speaker ID).
Dataset Details
The dataset is derived from the Common Voice dataset.
The original dataset is available at Common Voice Dataset.
The dataset is grouped by client ID, which is treated as the speaker ID for this dataset.
Each group is filtered to include only client IDs with a minimum of 30 samples and a maximum… See the full description on the dataset page: https://huggingface.co/datasets/masuidrive/cv-corpus-17.0-zh-TW-client_id-grouped.cv-corpus-17.0-ja-client_id-grouped
cv-corpus-17.0-ja-client_id-grouped
This dataset is a subset of the Common Voice dataset, filtered and grouped based on the client ID (treated as speaker ID).
Dataset Details
The dataset is derived from the Common Voice dataset.
The original dataset is available at Common Voice Dataset.
The dataset is grouped by client ID, which is treated as the speaker ID for this dataset.
Each group is filtered to include only client IDs with a minimum of 30 samples and a maximum of… See the full description on the dataset page: https://huggingface.co/datasets/masuidrive/cv-corpus-17.0-ja-client_id-grouped.aisha-urdu-voice-clipskk-yt-clipsspyken-clipswildlands-clientcv-corpus-1.0-en-client_id-grouped
cv-corpus-1.0-en-client_id-grouped
This dataset is a subset of the Common Voice dataset, filtered and grouped based on the client ID (treated as speaker ID).
Dataset Details
The dataset is derived from the Common Voice dataset.
The original dataset is available at Common Voice Dataset.
The dataset is grouped by client ID, which is treated as the speaker ID for this dataset.
Each group is filtered to include only client IDs with a minimum of 60 samples and a maximum of… See the full description on the dataset page: https://huggingface.co/datasets/masuidrive/cv-corpus-1.0-en-client_id-grouped.data-clipOE-DCT-Movie-clipsclinical-conversations-anon-benchmarkvideo-clips-and-imgsuclass_clipped_labeled
Dataset Card for "uclass_clipped_labeled"
More Information needed
music-clips-50There are 50 music clips(of 3~5 seconds).
You can load them by the following code:
from datasets import load_dataset
dataset = load_dataset('yongjian/music-clips-50')
clips = dataset['train'] # all 50 music clips
music_1_np_array = clips[0]['audio']['array'] # numpy array of shape=[N,]
Or you can directly download them from Google Drive: music-clips-50.tar.gz.
Mouth_Clicks_Smacks_and_Pop_Articulations_Preview
Harmonic Frontier Audio – Mouth Clicks, Smacks and Pop Articulations (Preview, v0.9)
A high-fidelity human vocal dataset designed for AI training, speech research, and expressive voice modeling.
Mouth Clicks, Smacks and Pop Articulations (Preview), created by Harmonic Frontier Audio, provides a compact reference set demonstrating the quality, formatting, and metadata conventions used in the Harmonic Frontier Audio Human Vocality Primitives series.
🔎 Summary
This… See the full description on the dataset page: https://huggingface.co/datasets/Harmonic-Frontier-Audio/Mouth_Clicks_Smacks_and_Pop_Articulations_Preview.sep28k-train-4-second-clips
Dataset Card for "sep28k-train-4-second-clips"
More Information needed
DENTAL_CLICK
Dataset Card for "DENTAL_CLICK"
More Information needed
1776-track-clips
1776 Track Clips
Fast, wordless, loop-safe background music built for Shorts, Reels, and edits.
This dataset contains 1776 unique audio clips designed for creators, editors, and developers who need clean, reusable background music without lyrics.
What’s included
1776 unique clips
No lyrics (wordless hooks)
Clean, loop-safe structure
Optimized for short-form video
Editor-first design
Preview
See the preview video(s) in this repository for examples across… See the full description on the dataset page: https://huggingface.co/datasets/BonusLockSMith/1776-track-clips.jdd_topic1_20251224-cliponly_sample100audios-inglessep28k-train-5-second-clips
Dataset Card for "sep28k-train-5-second-clips"
More Information needed
CommonVoice_TR_clips_16khz_g711_White_noisesep28k-dev-4-second-clips
Dataset Card for "sep28k-dev-0120-4-second-clips"
More Information needed
sep28k-test-5-second-clips
Dataset Card for "sep28k-test-5-second-clips"
More Information needed
rapnic-example
RAPNIC Dataset (example)
Dataset Description
This is an example of the full dataset, yet to be published, with 10 audio examples for 72 speakers.
RAPNIC (Reconeixement Automàtic de la Parla No Intel·ligible en Català) is a Catalan speech corpus collected from individuals with speech disorders, specifically cerebral palsy and Down syndrome.
This dataset was collected to develop and improve automatic speech recognition (ASR) systems that are accessible to people with speech… See the full description on the dataset page: https://huggingface.co/datasets/CLiC-UB/rapnic-example.sep28k-dev-5-second-clips
Dataset Card for "sep28k-dev-5-second-clips"
More Information needed
fluencybank-3-second-clips
Dataset Card for "fluencybank-3-second-clips"
More Information needed
sep28k-train-3-second-clips-full-agreement
Dataset Card for "sep28k-train-3-second-clips-full-agreement"
More Information needed
worst100-testclean-clips
Worst-100 test-clean clips — audio, transcripts, and the vocabulary finding
The 100 LibriSpeech test-clean clips where the block-4 production model (4.60% WER) made
the most word errors — with audio embedded so the failures can be listened to, plus the
model's transcript next to the reference for each clip.
The finding this dataset produced
47% of the word errors in these clips are on words that never appeared in the 30-hour
training vocabulary at all (20,066… See the full description on the dataset page: https://huggingface.co/datasets/Diffusion-ASR/worst100-testclean-clips.
