datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Codemixed_New
Codemixed ASR Dataset
Unified collection of code-mixed ASR datasets.
fluent_speech_commands_synth
Dataset Card for "fluent_speech_commands_synth"
More Information needed
CodecFake
CodecFake: Enhancing Anti-Spoofing Models Against Deepfake Audios from Codec-Based Speech Synthesis Systems
Paper,
Code,
Project Page
Interspeech 2024
TL;DR: We show that better detection of deepfake speech from codec-based TTS systems can be achieved by training models on speech re-synthesized with neural audio codecs.
This dataset is released for this purpose.
See our paper and Github for more details on using our dataset.
Acknowledgement… See the full description on the dataset page: https://huggingface.co/datasets/rogertseng/CodecFake.voice-code-bench
VoiceCodeBench
VoiceCodeBench is a test-only benchmark for evaluating whether automatic
speech recognition (ASR) systems preserve exact structured values in English
workplace speech.
Paper: VoiceCodeBench: Evaluating Exact Structured-Token Recovery in Automatic Speech Recognition
The benchmark targets cases where a transcript is software input: callback
numbers, email addresses, command-line flags, file paths, URLs, account
identifiers, dates, measurements, and similar values… See the full description on the dataset page: https://huggingface.co/datasets/besimple-ai/voice-code-bench.WSC-EvalASR_Code_Switch
ASR Code-Switching Benchmark
A curated benchmark of 1,200 code-switching utterances (300 per language pair)
for evaluating commercial ASR systems on multilingual speech with intra-sentential
language switching.
Paper
Benchmarking Commercial ASR Systems on Code-Switching Speech: Arabic, Persian, and German
arXiv link
Language pairs
Split
Language pair
Samples
Scripts
egyptian_arabic_english
Egyptian Arabic–English
300
Arabic + Latin… See the full description on the dataset page: https://huggingface.co/datasets/Perle-ai/ASR_Code_Switch.crema_d_synth
Dataset Card for "crema_d_synth"
More Information needed
maestro_synth
Dataset Card for "maestro_synth"
More Information needed
codecfake-audio
Codecfake Dataset
Overview
The Codecfake dataset is a large-scale dataset designed for the detection of Audio Language Model (ALM)-based deepfake audio. This dataset includes millions of audio samples across two languages and various test conditions, tailored specifically for ALM-based audio detection.
Conversion
The original dataset was downloaded from Zenodo and converted to FLAC format to maintain audio quality while reducing file size. The dataset has been… See the full description on the dataset page: https://huggingface.co/datasets/ajaykarthick/codecfake-audio.voxceleb1_synthsmash-battlefield-fox-codec-20fps
Smash Battlefield Fox — codec-ready videos
Lossless 20 FPS preprocessing of DhruvBhatia0/smash-battlefield-fox at revision 2c94351b82c2c65a31fb39fe52a34ff905b6abcf. This dataset contains 512 complete replays selected deterministically with seed 28.
Every third decoded frame is resized to 252×208 with PyAV's training-time resize, then stored losslessly as RGB FFV1 in a streaming NUT container. Decoding the processed files reproduces the preprocessed RGB tensors bit-for-bit.… See the full description on the dataset page: https://huggingface.co/datasets/DhruvBhatia0/smash-battlefield-fox-codec-20fps.vocal_imitation_synth
Dataset Card for "vocal_imitation_synth"
More Information needed
opensinger_synthtorgo_synthspeech_accent_archive_synthfsd50k_synthlibrispeech_asr_test_48k_synthGhana_English-Twi_Code-switching_Speech
Dataset Card for KasaSpeech
Dataset Summary
KasaSpeech is a large-scale English–Twi code-switching speech dataset developed to advance research in speech technologies for English and Twi.
The dataset comprises 54,855 transcribed speech recordings collected from speakers across Ghana and is designed to capture natural code-switching between English and Twi across a diverse range of everyday topics and communication scenarios
With over 95 hours of manually… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/Ghana_English-Twi_Code-switching_Speech.vox_lingua_top10_synthsample_100vocalset_synthspeech_accent_archive_englishfluent_speech_commands_femalelibrispeech_asr_test_synthvocalset_synth
Dataset Card for "vocalset_synth"
More Information needed
snips_test_valid_synth
Dataset Card for "snips_test_valid_synth"
More Information needed
speech_accent_archive_othereasycall_synthlibrispeech_synth
Dataset Card for "librispeech_synth"
More Information needed
cv_13_zh_tw_synth
