datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ASMR-Archive-Processed-SFW
ASMR-Archive-Processed-SFW
Overview
This dataset is an “educational” subset of the original OmniAICreator/ASMR-Archive-Processed dataset.
We filtered the original dataset to include only records where the nsfw metadata flag is false.
To maintain the randomness and anonymity of the entries, multiple directories were combined and shuffled.
The nsfw tag in the original dataset is inherited from the tags of the original audio works before they were passed through the… See the full description on the dataset page: https://huggingface.co/datasets/noxwano/ASMR-Archive-Processed-SFW.PROCESS-2
PROCESS-2: Remote Speech Dataset for Cognitive Assessment
Dataset Summary
PROCESS-2, the successor to the PROCESS Challenge [Tao, F., Mirheidari, B., Pahar, M., Young, S., Xiao, Y., Elghazaly, H., Peters, F., Illingworth, C., Braun, D., O’Malley, R., & Bell, S. (2025). Early dementia detection using multiple spontaneous speech prompts: The PROCESS Challenge. In ICASSP 2025 – IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (pp.… See the full description on the dataset page: https://huggingface.co/datasets/CognoSpeak/PROCESS-2.librispeech-arpabet-processed
LibriSpeech ARPAbet Processed Dataset
Pre-processed dataset for training ARPAbet phoneme recognition models using CTC loss.
Dataset Description
This dataset is derived from LibriSpeech (train-clean-100 split) with the following preprocessing:
Audio: Resampled to 16kHz, normalized using Wav2Vec2 feature extractor
Labels: Text transcriptions converted to ARPAbet phoneme sequences using CMU Pronouncing Dictionary
Filtering: Samples with out-of-vocabulary words (not in CMU… See the full description on the dataset page: https://huggingface.co/datasets/davidggphy/librispeech-arpabet-processed.AISHELL-1-processedprocessed-voice-th-169k
processed-voice-th-169k
Cleaned open-source Thai speech dataset: 169,550 utterances (149,953 train / 7,614 dev / 11,983 test) with transcripts, prepared for ASR fine-tuning.
Format
Field
Description
sentence
Transcript in Thai
audio
Audio clip
Usage
from datasets import load_dataset
ds = load_dataset("Porameht/processed-voice-th-169k")
Used to fine-tune Porameht/whisper-tiny-thai. See also the smaller, cleaner… See the full description on the dataset page: https://huggingface.co/datasets/Porameht/processed-voice-th-169k.genshin-voice-v3.5-mandarin-processedcommon-voice-urdu-processed
🎙️ Common Voice Urdu (Processed)
Ready-to-use Urdu speech dataset for fine-tuning ASR models
Mozilla Common Voice → Preprocessed → Whisper-Ready ✨
📊 Dataset at a Glance
Split
Samples
Use
🏋️ Train
7,339
Model training
🔧 Validation
5,046
Hyperparameter tuning
🧪 Test
5,091
Final evaluation
Total
17,476
💡 Audio is pre-resampled to 16kHz — plug directly into Whisper!
🚀 Quick Start
from datasets import load_dataset
#… See the full description on the dataset page: https://huggingface.co/datasets/khawajaaliarshad/common-voice-urdu-processed.processed-cv-17-th-130k
processed-cv-17-th-130k
Cleaned Thai split of Mozilla Common Voice 17: 130,551 utterances (117,536 train / 3,950 dev / 9,065 test) with transcripts, ready for ASR training.
Format
Field
Description
sentence
Transcript in Thai
audio
Audio clip
Usage
from datasets import load_dataset
ds = load_dataset("Porameht/processed-cv-17-th-130k")
Source and license
Derived from Mozilla Common Voice 17.0 (Thai), released under… See the full description on the dataset page: https://huggingface.co/datasets/Porameht/processed-cv-17-th-130k.arknights_voices_zh-processedProcessed_TTS_Multilingual_Data
Processed TTS Multilingual Data
Validated and quality-checked multilingual speech datasets for TTS training, covering 12+ Indian languages.
Datasets Included
Subset
Samples
Hours
Description
indic_voices_r
239,684
548.8h
Indic Voices_R — IVR recordings
rasa
201,509
361.2h
RASA — read speech (wiki, conv, book, news)
indictts_iitm
155,236
253.6h
Indic TTS (IIT Madras) — studio TTS recordings at 48kHz
Total
596,429
1,163.6h
Languages… See the full description on the dataset page: https://huggingface.co/datasets/PalakEngineerMaster/Processed_TTS_Multilingual_Data.akan_audio_processedprocessed-smarthome-th
processed-smarthome-th
Cleaned Thai speech dataset for smart-home commands: 9,600 utterances (7,680 train / 960 dev / 960 test) with transcripts.
Format
Field
Description
sentence
Transcript in Thai
audio
Audio clip
Usage
from datasets import load_dataset
ds = load_dataset("Porameht/processed-smarthome-th")
Used to fine-tune Porameht/whisper-tiny-smarthome-thai (WER 24.375 on the eval split).
ViSEC-processed
ViSEC Processed
Processed ViSEC speaker audio generated for the Meddies ASR collection.
Contents
processed_audio_by_id/: 147 WAV files named by speaker id.
metadata.csv: per-speaker metadata with duration, clip count, emotion coverage, and source-duration summary.
Schema
metadata.csv contains:
speaker_id: integer speaker identifier.
output_path: relative path to the processed WAV file.
duration_seconds: duration of the processed audio file.… See the full description on the dataset page: https://huggingface.co/datasets/Meddies/ViSEC-processed.
