datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
asmr-yt-chaptersASMR-Archive-Processed-SFW
ASMR-Archive-Processed-SFW
Overview
This dataset is an “educational” subset of the original OmniAICreator/ASMR-Archive-Processed dataset.
We filtered the original dataset to include only records where the nsfw metadata flag is false.
To maintain the randomness and anonymity of the entries, multiple directories were combined and shuffled.
The nsfw tag in the original dataset is inherited from the tags of the original audio works before they were passed through the… See the full description on the dataset page: https://huggingface.co/datasets/noxwano/ASMR-Archive-Processed-SFW.ASMR-Archive-Processed
ASMR-Archive-Processed (WIP)
Update (2026-04-03): This dataset has reached the Hugging Face Public Storage Limit. After contacting support, we were informed that the only option is to pay for a storage expansion. Consequently, updates to this dataset are now suspended.
Work in Progress — expect breaking changes while the pipeline and data layout stabilize.
This dataset contains ASMR audio data sourced from DeliberatorArchiver/asmr-archive-data-01 and… See the full description on the dataset page: https://huggingface.co/datasets/OmniAICreator/ASMR-Archive-Processed.asmr-archive-data-01
ASMR Media Archive Storage
This repository contains an archive of ASMR works.
All data in this repository is uploaded for educational and research purposes only. All use is at your own risk.
[!IMPORTANT]
This repository contains >= 64 TiB of files.Git LFS consumes twice as much disk space because of the way it works, so git clone is not recommended. Hugging Face CLI or Python libraries allow you to select and download only a subset of files.
>>> CLICK HERE or on the IMAGE BELOW… See the full description on the dataset page: https://huggingface.co/datasets/DeliberatorArchiver/asmr-archive-data-01.asmr-archive-data-02
ASMR Media Archive Storage
This repository contains an archive of ASMR works.
All data in this repository is uploaded for educational and research purposes only. All use is at your own risk.
[!IMPORTANT]
This repository contains >= 64 TiB of files.Git LFS consumes twice as much disk space because of the way it works, so git clone is not recommended. Hugging Face CLI or Python libraries allow you to select and download only a subset of files.
>>> CLICK HERE or on the IMAGE BELOW… See the full description on the dataset page: https://huggingface.co/datasets/DeliberatorArchiver/asmr-archive-data-02.asmr-zh-r18
asmr-zh-r18
Chinese R18 ASMR audio dataset for TTS/voice cloning fine-tuning.
Works: 5876 RJ-coded works
Total size: ~1.1TB raw (10 compressed parts)
Format: MP3/WAV/FLAC, 3 tracks per work
Source: asmr.one API, Chinese R18 category
Extract
for f in packs/*.tar.zst; do
zstd -d "$f" --stdout | tar -xf -
done
asmr
Dataset Card for ASMR Audio Dataset
Dataset Summary
This dataset contains a large collection of ASMR (Autonomous Sensory Meridian Response) audio clips with corresponding machine-generated transcriptions. The dataset includes approximately 283,132 audio segments totaling over 307 hours of content, with an average duration of 3.92 seconds per clip. All audio files are provided in WAV format at 24 kHz sampling rate, making them suitable for various audio processing and… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/asmr.japanese_asmrasmr-mamiASMRjapanese_asmrasmr-api-backup-outputASMR_Datasetcrawl progress: 57% 121249/213330
pipeline:
Manual cleaning, The following situations will be discarded:
no Chinese subtitles
timestamps of each sentence in the subtitles are connected at the end
timestamps are not aligned
mel band roformer remove SE, BGM, saliva sound, etc.
parsec subtitle & cut audio
Split the left and right channels and use anime-whisper to identify the loudest channel(beta version, maybe there is a better way)
create index
asmr_archive_qwentts_encodeduk_UA-ASMR
Ukrainian ASMR TTS Dataset
A Ukrainian text-to-speech dataset for training single-speaker ASMR-style voice models using Piper.
Dataset Details
Property
Value
Language
Ukrainian (uk_UA)
Speakers
1
Segments
7,318
Audio Format
16-bit WAV, 22050 Hz, Mono
License
CC0
Dataset Structure
Prerequisites
# Install Piper training dependencies
git clone https://github.com/kontextox/piper1-gpl.git
cd piper1-gpl
python3 -m venv .venv
source… See the full description on the dataset page: https://huggingface.co/datasets/kontextox/uk_UA-ASMR.asmr-archive-data-metaasm_ryoutube_en_asmr_rawhumanoid-asmr-datayoutube_en_asmrwhisper-asmrpersonal-asmr
