datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
test_librispeech_parquetnsynth-parquetfsdkaggle2019-parquet
FSDKaggle2019
FSDKaggle2019[1] is an audio dataset containing 29,266 audio files annotated with 80 labels of the AudioSet Ontology.
FSDKaggle2019 has been used for the DCASE Challenge 2019 Task 2, which was run as a Kaggle competition titled Freesound Audio Tagging 2019.
All audio clips are provided as uncompressed PCM 16 bit, 44.1 kHz, mono audio files.
This version of database could be found and downloaded from here.
Data Split Statistics
Curated
Noisy
Test… See the full description on the dataset page: https://huggingface.co/datasets/mteb/fsdkaggle2019-parquet.esc50-parquetfsdkaggle2019-parquet
FSDKaggle2019
FSDKaggle2019[1] is an audio dataset containing 29,266 audio files annotated with 80 labels of the AudioSet Ontology.
FSDKaggle2019 has been used for the DCASE Challenge 2019 Task 2, which was run as a Kaggle competition titled Freesound Audio Tagging 2019.
All audio clips are provided as uncompressed PCM 16 bit, 44.1 kHz, mono audio files.
This version of database could be found and downloaded from here.
Data Split Statistics
Curated
Noisy
Test… See the full description on the dataset page: https://huggingface.co/datasets/confit/fsdkaggle2019-parquet.cremad-parquetgtzan-parquet
GTZAN Music Genre Classification
GTZAN consists of 100 30-second recording excerpts in each of 10 categories, and is the most-used public dataset in music information retrieval (MIR) research.
Following Kereliuk et al. (2015), we use the "fault-filtered" partitioning version of GTZAN, which is constructed by hand to include 443/197/290 excerpts.
This version of database could be found and downloaded from here.
Citations
@article{kereliuk2015deep,
title={Deep… See the full description on the dataset page: https://huggingface.co/datasets/confit/gtzan-parquet.librispeech_parquetDDD-Cambodia-khmer-speech-dataset-parquet-000-159-en-translateDisclaimer: The original dataset can be found here.
It is published by Digital Divide Data Cambodia (DDD-Cambodia).
License:
Khmer ASR Cultural Dataset's license is Creative Commons Attribution Share Alike 4.0 International (CC-BY-SA-4.0).
Please attribute Digital Divide Data if you use this dataset in any way.
Objective of this dataset
Add English translation: a new column en_translate is added to the original dataset (only from parquet 000 to 159 of the original… See the full description on the dataset page: https://huggingface.co/datasets/KrorngAI/DDD-Cambodia-khmer-speech-dataset-parquet-000-159-en-translate.emilia-yodas-fr-parquetsimchoir-parquet
FastMSS synthetic multi-speaker meetings - parquet edition
Streaming-friendly parquet shards of the FastMSS synthetic multi-speaker conversational corpus. Each row is one mixture with the audio bytes embedded inline (16 kHz mono WAV) plus per-segment diarization timestamps, per-word transcript and the full lhotse cut as a JSON blob. See fastmss/hf_dataset.py for the schema docstring.
Subsets and splits
debug — splits: train — 1 mixtures, 1.6 min total, 6 unique speakers… See the full description on the dataset page: https://huggingface.co/datasets/arda-argmax/simchoir-parquet.librispeech-sid-parquet
LibriSpeech Speaker Identification
LibriSpeech is a corpus of approximately 1000 hours of 16kHz read English speech, prepared by Vassil Panayotov with the assistance of Daniel Povey.
The data is derived from read audiobooks from the LibriVox project, and has been carefully segmented and aligned.
However, although LibriSpeech is very popular in ASR tasks, we use LibriSpeech database as a speaker identification task.
We follow SincNet paper official split for training and evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/confit/librispeech-sid-parquet.ravdess-parquetmunch-1-latent-NEW-parquet
🎙️ Urdu TTS Latent Dataset — munch-1-latent-NEW-parquet
Pre-computed DACVAE latent representations for 51,021 Urdu utterances, ready for TTS model training. No audio decoding required at training time — load the dataset, reshape the binary blob, and train.
Source
Field
Value
Source audio
Humair332/Urdu-munch-1
Codec
Aratako/Semantic-DACVAE-Japanese-32dim
Codec sample rate
48,000 Hz
Encoder hop size
1,920 samples
Latent frame rate
25.0 Hz
Latent dim… See the full description on the dataset page: https://huggingface.co/datasets/zuhri025/munch-1-latent-NEW-parquet.OpenSLR54-Nepali-ASR-parquet
OpenSLR 54: Large Nepali ASR training data set (unmodified parquet repackaging)
This is an unofficial repackaging of the official OpenSLR 54 release
(SLR54, https://www.openslr.org/54/), converted to parquet so it can be streamed with 🤗 datasets.
It is not affiliated with or endorsed by OpenSLR or the original authors.
All credit for the data belongs to the original creators (see Citation).
What's inside
157,905 utterances, 16 shards: one per original zip… See the full description on the dataset page: https://huggingface.co/datasets/JeevanDai/OpenSLR54-Nepali-ASR-parquet.wmms-parquet
Watkins Marine Mammal Sound (WMMS) Database
Sound files on this website are free to download for personal or academic (not commercial) use.
Sound files and associated metadata are credited as follows: "Watkins Marine Mammal Sound Database, Woods Hole Oceanographic Institution and the New Bedford Whaling Museum."
Database could be found and downloaded from here.
In this database version, the audio archive includes sounds of 32 species:
Atlantic_Spotted_Dolphin
Bearded_Seal
Beluga… See the full description on the dataset page: https://huggingface.co/datasets/confit/wmms-parquet.mswc-parquetukrainian-tts-audiobook-pani-nina-parquet
Ukrainian TTS audiobook dataset Pani Nina (Parquet)
Segmented Ukrainian audiobook speech with aligned text, prepared for training and evaluating Text-to-Speech (TTS) models.
The dataset is published as Hugging Face-compatible Parquet shards so the Hub Dataset Preview can render an audio column.
The dataset was prepared using whisper and ffmpeg:
Whisper was used for transcription and approximate segment timing.
FFmpeg was used to slice audio into short utterances (roughly 2-10… See the full description on the dataset page: https://huggingface.co/datasets/rishchen/ukrainian-tts-audiobook-pani-nina-parquet.one_voice_FACEBOOK_PARQUET
Artificial Omnivoice Hungarian Speaker Dataset
Ez egy teljesen szintetikus magyar nyelvű beszédadatbázis, amely kiváló minőségű szövegfelolvasó (TTS) és beszédfelismerő (ASR) modellek tanításához és finomhangolásához készült.
Adatforrás és Referencia Hang
A dataset alapjául szolgáló referencia hang (speaker identity) egy 20 másodperces részlet az alábbi YouTube videóból:
Forrás: Hogyan legyél tökéletes magyar várvédő tutorial
Licenc: A videó CC (Creative Commons)… See the full description on the dataset page: https://huggingface.co/datasets/fablevi/one_voice_FACEBOOK_PARQUET.OpenSLR43-Nepali-TTS-parquet
OpenSLR 43: High quality TTS data for Nepali (unmodified parquet repackaging)
This is an unofficial repackaging of the official OpenSLR 43 release (SLR43, https://www.openslr.org/43/),
converted to parquet for 🤗 datasets. It is not affiliated with or endorsed by OpenSLR, Google or the original authors.
All credit for the data belongs to the original creators (see Citation).
What's inside
2,064 utterances, 18 female speakers, ~2.8 h, 48 kHz mono; recorded in… See the full description on the dataset page: https://huggingface.co/datasets/JeevanDai/OpenSLR43-Nepali-TTS-parquet.emodb-parquet
EmoDB
The EmoDB is the freely available German emotional database, containing a total of 535 utterances.
It comprises of seven emotions: 1) anger; 2) boredom; 3) anxiety; 4) happiness; 5) sadness; 6) disgust; and 7) neutral.
The data was recorded at a 48-kHz sampling rate and then down-sampled to 16-kHz.
We follow the unofficial speaker-independent train/test split from here.
Citations
@inproceedings{burkhardt2005database,
title={A database of German emotional… See the full description on the dataset page: https://huggingface.co/datasets/confit/emodb-parquet.pianos-parquet
Pianos Sound Quality Dataset
This version of dataset comprises seven models of pianos:
Kawai upright piano
Kawai grand piano
Young Change upright piano
Hsinghai upright piano
Grand Theatre Steinway piano
Steinway grand piano
Pearl River upright piano.
Note: the paper (Zhou et al., 2023) only uses the first 7 piano classes in the dataset, its future work has finished the 8-class evaluation.
License
MIT License
Copyright (c) CCMUSIC
Permission is hereby granted… See the full description on the dataset page: https://huggingface.co/datasets/confit/pianos-parquet.ukrainian-tts-audiobook-pani-nina-parquet-old
Ukrainian TTS audiobook dataset (Parquet)
Segmented Ukrainian audiobook speech with aligned text, prepared for training and evaluating Text-to-Speech (TTS) models. The dataset is published as Hugging Face-compatible Parquet shards so the Hub Dataset Preview can render an audio column.
That was mabe by using whisper (https://github.com/openai/whisper) and ffmpeg (https://www.ffmpeg.org/), where with whisper we set start and end of voices + transcribe it and using ffmpeg slice into… See the full description on the dataset page: https://huggingface.co/datasets/rishchen/ukrainian-tts-audiobook-pani-nina-parquet-old.wmms-parquet
Watkins Marine Mammal Sound (WMMS) Database
Sound files on this website are free to download for personal or academic (not commercial) use.
Sound files and associated metadata are credited as follows: "Watkins Marine Mammal Sound Database, Woods Hole Oceanographic Institution and the New Bedford Whaling Museum."
Database could be found and downloaded from here.
In this database version, the audio archive includes sounds of 32 species:
Atlantic_Spotted_Dolphin
Bearded_Seal
Beluga… See the full description on the dataset page: https://huggingface.co/datasets/qdskipper/wmms-parquet.medical-tts-parquet-2-16khz
IntelMedica Medical TTS Dataset v2 (16kHz)
Description
Synthetic medical speech dataset for fine-tuning Whisper-based ASR models on clinical and nursing terminology. Contains 101,475 audio-text pairs totaling 184.1 hours of speech at 16 kHz mono, generated using Kokoro-82M TTS with 19 voices across three English accent groups.
This is v2 -- a companion to the v1 dataset (125,500 samples, ~257 hours). v2 focuses on terms from additional data sources (RxNorm API, FDA… See the full description on the dataset page: https://huggingface.co/datasets/intelmedica/medical-tts-parquet-2-16khz.my_parquet_dataset_2dhivehi-javaabu-speech-parquetlibrispeech_asr_parquetwhiser_parquet
Dataset Description
Motivation
Enable both categorical emotion recognition and dimensional affect regression on speech clips with multiple crowd-sourced annotations per instance.
Study annotator variability, label aggregation, and uncertainty.
Composition
Audio: WAV files in wavs/
Annotations:
Raw: WHiSER/Labels/labels.txt (per-file summary + multiple worker lines)
Aggregated: WHiSER/Labels/labels_consensus.{csv,json}
Per-annotator:… See the full description on the dataset page: https://huggingface.co/datasets/Lab-MSP/whiser_parquet.bengali-tts-folderized-parquet-stage2
Bengali TTS Folderized Parquet Stage 2
Final filtered (keep=True) Bengali TTS dataset, with merged/combined chunks.
Layout
<speaker>.parquet (single shard, audio <= ~1GB)
<speaker>_001.parquet, _002.parquet, ... (multiple shards, split by audio byte size)
Audio sourcing convention
COMBINED == False -> sourced from extracted_audio/<speaker>/<video_id>/<chunk_file>
COMBINED == True -> sourced from… See the full description on the dataset page: https://huggingface.co/datasets/dipit099/bengali-tts-folderized-parquet-stage2.
