datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ASMR-Archive-Processed-SFW
ASMR-Archive-Processed-SFW
Overview
This dataset is an “educational” subset of the original OmniAICreator/ASMR-Archive-Processed dataset.
We filtered the original dataset to include only records where the nsfw metadata flag is false.
To maintain the randomness and anonymity of the entries, multiple directories were combined and shuffled.
The nsfw tag in the original dataset is inherited from the tags of the original audio works before they were passed through the… See the full description on the dataset page: https://huggingface.co/datasets/noxwano/ASMR-Archive-Processed-SFW.ASMR-Archive-Processed
ASMR-Archive-Processed (WIP)
Update (2026-04-03): This dataset has reached the Hugging Face Public Storage Limit. After contacting support, we were informed that the only option is to pay for a storage expansion. Consequently, updates to this dataset are now suspended.
Work in Progress — expect breaking changes while the pipeline and data layout stabilize.
This dataset contains ASMR audio data sourced from DeliberatorArchiver/asmr-archive-data-01 and… See the full description on the dataset page: https://huggingface.co/datasets/OmniAICreator/ASMR-Archive-Processed.PROCESS-2
PROCESS-2: Remote Speech Dataset for Cognitive Assessment
Dataset Summary
PROCESS-2, the successor to the PROCESS Challenge [Tao, F., Mirheidari, B., Pahar, M., Young, S., Xiao, Y., Elghazaly, H., Peters, F., Illingworth, C., Braun, D., O’Malley, R., & Bell, S. (2025). Early dementia detection using multiple spontaneous speech prompts: The PROCESS Challenge. In ICASSP 2025 – IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (pp.… See the full description on the dataset page: https://huggingface.co/datasets/CognoSpeak/PROCESS-2.librispeech-arpabet-processed
LibriSpeech ARPAbet Processed Dataset
Pre-processed dataset for training ARPAbet phoneme recognition models using CTC loss.
Dataset Description
This dataset is derived from LibriSpeech (train-clean-100 split) with the following preprocessing:
Audio: Resampled to 16kHz, normalized using Wav2Vec2 feature extractor
Labels: Text transcriptions converted to ARPAbet phoneme sequences using CMU Pronouncing Dictionary
Filtering: Samples with out-of-vocabulary words (not in CMU… See the full description on the dataset page: https://huggingface.co/datasets/davidggphy/librispeech-arpabet-processed.AISHELL-1-processedprocessed-voice-th-169k
processed-voice-th-169k
Cleaned open-source Thai speech dataset: 169,550 utterances (149,953 train / 7,614 dev / 11,983 test) with transcripts, prepared for ASR fine-tuning.
Format
Field
Description
sentence
Transcript in Thai
audio
Audio clip
Usage
from datasets import load_dataset
ds = load_dataset("Porameht/processed-voice-th-169k")
Used to fine-tune Porameht/whisper-tiny-thai. See also the smaller, cleaner… See the full description on the dataset page: https://huggingface.co/datasets/Porameht/processed-voice-th-169k.genshin-voice-v3.5-mandarin-processedcommon-voice-urdu-processed
🎙️ Common Voice Urdu (Processed)
Ready-to-use Urdu speech dataset for fine-tuning ASR models
Mozilla Common Voice → Preprocessed → Whisper-Ready ✨
📊 Dataset at a Glance
Split
Samples
Use
🏋️ Train
7,339
Model training
🔧 Validation
5,046
Hyperparameter tuning
🧪 Test
5,091
Final evaluation
Total
17,476
💡 Audio is pre-resampled to 16kHz — plug directly into Whisper!
🚀 Quick Start
from datasets import load_dataset
#… See the full description on the dataset page: https://huggingface.co/datasets/khawajaaliarshad/common-voice-urdu-processed.processed-cv-17-th-130k
processed-cv-17-th-130k
Cleaned Thai split of Mozilla Common Voice 17: 130,551 utterances (117,536 train / 3,950 dev / 9,065 test) with transcripts, ready for ASR training.
Format
Field
Description
sentence
Transcript in Thai
audio
Audio clip
Usage
from datasets import load_dataset
ds = load_dataset("Porameht/processed-cv-17-th-130k")
Source and license
Derived from Mozilla Common Voice 17.0 (Thai), released under… See the full description on the dataset page: https://huggingface.co/datasets/Porameht/processed-cv-17-th-130k.arknights_voices_zh-processedakan_audio_processedprocessed-smarthome-th
processed-smarthome-th
Cleaned Thai speech dataset for smart-home commands: 9,600 utterances (7,680 train / 960 dev / 960 test) with transcripts.
Format
Field
Description
sentence
Transcript in Thai
audio
Audio clip
Usage
from datasets import load_dataset
ds = load_dataset("Porameht/processed-smarthome-th")
Used to fine-tune Porameht/whisper-tiny-smarthome-thai (WER 24.375 on the eval split).
ViSEC-processed
ViSEC Processed
Processed ViSEC speaker audio generated for the Meddies ASR collection.
Contents
processed_audio_by_id/: 147 WAV files named by speaker id.
metadata.csv: per-speaker metadata with duration, clip count, emotion coverage, and source-duration summary.
Schema
metadata.csv contains:
speaker_id: integer speaker identifier.
output_path: relative path to the processed WAV file.
duration_seconds: duration of the processed audio file.… See the full description on the dataset page: https://huggingface.co/datasets/Meddies/ViSEC-processed.
