aoxo/t2a-daddy
t2a-daddy Male-voice ASMR corpus for the text2asmr project. Previously published as aoxo/audios3. Companion repos: aoxo/t2a-mommy (female voice), aoxo/t2a-audios-v1 (the original v1 corpus). Layout path what <creator>/<title>.m4a source audio, 48 kHz AAC, one folder per creator <creator>/<title>.json word-level Whisper large-v3 alignment ([] = skipped: near-silent or undecodable) labels/qwen3omni.jsonl non-speech ontology labels for gap clips… See the full description on the dataset page: https://huggingface.co/datasets/aoxo/t2a-daddy.
t2a-daddy
Male-voice ASMR corpus for the text2asmr project. Previously published as aoxo/audios3.
Companion repos: `aoxo/t2a-mommy` (female voice), `aoxo/t2a-audios-v1` (the original v1 corpus).
Layout
Label rows are {uid, raw, label, labeler}, where uid is <source>.m4a_<start_ms>.
Ontology: whispering, normal speech, breathing, mouth sounds, moaning, kissing, silence, plus the physical tail (tapping, scratching, crinkling, brushing, liquid, page turning, other sound). A label covers at most 3–4 s of audio.
Audio is sourced from soundgasm; each creator keeps their own folder, so attribution and creator-split evaluation are preserved.
