Team Ai
Datasetpublic

kontextox/uk_UA-ASMR

Ukrainian ASMR TTS Dataset A Ukrainian text-to-speech dataset for training single-speaker ASMR-style voice models using Piper. Dataset Details Property Value Language Ukrainian (uk_UA) Speakers 1 Segments 7,318 Audio Format 16-bit WAV, 22050 Hz, Mono License CC0 Dataset Structure Prerequisites # Install Piper training dependencies git clone https://github.com/kontextox/piper1-gpl.git cd piper1-gpl python3 -m… See the full description on the dataset page: https://huggingface.co/datasets/kontextox/uk_UA-ASMR.

sourceHugging Facecc0-1.0updated 6mo agoView on Hugging Face
0likes20downloads
Dataset Card

Ukrainian ASMR TTS Dataset

A Ukrainian text-to-speech dataset for training single-speaker ASMR-style voice models using Piper.

Dataset Details

PropertyValue
LanguageUkrainian (uk_UA)
Speakers1
Segments7,318
Audio Format16-bit WAV, 22050 Hz, Mono
LicenseCC0

Dataset Structure

Prerequisites

bash
# Install Piper training dependencies
git clone https://github.com/kontextox/piper1-gpl.git
cd piper1-gpl
python3 -m venv .venv
source .venv/bin/activate
python3 -m pip install -e '.[train]'
./build_monotonic_align.sh
# If use `OHF-voice/piper1-gpl` (fixed in `kontextox/piper1-gpl`):
# pip install scikit-build
python3 setup.py build_ext --inplace

# CRITICAL FIX for custom text phonemes in Piper `OHF-voice/piper1-gpl` (fixed in `kontextox/piper1-gpl`):
# This patches dataset.py to properly use the custom phoneme map loaded via --data.phonemes_path
# sed -i 's/phonemes_to_ids(sentence_phonemes)/phonemes_to_ids(sentence_phonemes, id_map=self.piper_config.phoneme_id_map)/g' src/piper/train/vits/dataset.py
bash
# 1. Download the dataset
hf download kontextox/uk_UA-ASMR \
  --repo-type dataset --local-dir uk_UA-ASMR
tar -xzf uk_UA-ASMR/clear_audio.tar.gz -C uk_UA-ASMR/audio

# 2. Download the base checkpoint AND its configuration
hf download rhasspy/piper-checkpoints uk/uk_UA/ukrainian_tts/medium/epoch=2090-step=1166778.ckpt \
  --repo-type dataset --local-dir uk_UA-ASMR/checkpoints

hf download rhasspy/piper-checkpoints uk/uk_UA/ukrainian_tts/medium/config.json \
  --repo-type dataset --local-dir uk_UA-ASMR/checkpoints

# 3. Extract the exact phoneme map from the base config to use for training
python3 -c "import json; d=json.load(open('uk_UA-ASMR/checkpoints/uk/uk_UA/ukrainian_tts/medium/config.json')); json.dump(d['phoneme_id_map'], open('uk_UA-ASMR/phonemes.json','w'), ensure_ascii=False, indent=2)"
text
uk_UA-ASMR/
├── README.md
├── metadata.csv          # Metadata
├── phonemes.json         # Automatically extracted Ukrainian phoneme map
├── audio/                # Audio files (22050 Hz, mono, 16-bit)
│   ├── utt_0001.wav
│   ├── utt_0002.wav
│   └── ...
└── checkpoints/uk/uk_UA/ukrainian_tts/medium/
    ├── config.json
    └── epoch=2090-step=1166778.ckpt

Audio Specifications

  • —Sample Rate: 22050 Hz
  • —Channels: Mono
  • —Bit Depth: 16-bit
  • —Format: WAV

Training

Training Command

bash
python3 -m piper.train fit \
  --data.voice_name "uk_asmr" \
  --data.csv_path uk_UA-ASMR/metadata.csv \
  --data.audio_dir uk_UA-ASMR/audio \
  --data.espeak_voice "uk" \
  --model.sample_rate 22050 \
  --data.phoneme_type "text" \
  --data.dataset_type "text" \
  --data.phonemes_path uk_UA-ASMR/phonemes.json \
  --data.cache_dir uk_UA-ASMR/cache \
  --data.config_path uk_UA-ASMR/output/uk_UA-asmr-medium.onnx.json \
  --data.batch_size 32 \
  --data.num_workers 8 \
  --model.vocoder_warmstart_ckpt uk_UA-ASMR/checkpoints/uk/uk_UA/ukrainian_tts/medium/epoch=2090-step=1166778.ckpt \
  --trainer.max_epochs 500 \
  --trainer.check_val_every_n_epoch 1 \
  --trainer.default_root_dir uk_UA-ASMR/output

**Note**: `--trainer.defaultrootdir` ensures PyTorch Lightning saves logs and checkpoints cleanly to `ukUA-ASMR/output/lightninglogs/`

**Note**: The NVIDIA driver on your system is too old (found version 12080) or NVIDIA GeForce RTX 5090 with CUDA capability `sm120` is not compatible with the current PyTorch installation:_

  • —Check: python -c "import torch; print(torch.__version__); print(torch.cuda.get_arch_list()); print(torch.randn(1).cuda())"
  • —Run: pip install --upgrade torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu128

**Note**: Check CPU process `find ukUA-ASMR/cache -name "*.pt" | wc -l`_

Hardware Configuration

GPUVRAMBatch SizeNum WorkersSpeedEpoch Time
H100 NVL93 GB12816~1.0 it/s~60s
L40S46 GB328~2.6 it/s~90s
RTX 309024 GB248~2.0 it/s~120s
RTX 306012 GB164~1.5 it/s~180s
A10080 GB9616~0.9 it/s~70s
Continue from latest checkpoint
bash
python3 -m piper.train fit \
  --data.voice_name "uk_asmr" \
  --data.csv_path uk_UA-ASMR/metadata.csv \
  --data.audio_dir uk_UA-ASMR/audio \
  --data.espeak_voice "uk" \
  --model.sample_rate 22050 \
  --data.phoneme_type "text" \
  --data.dataset_type "text" \
  --data.phonemes_path uk_UA-ASMR/phonemes.json \
  --data.cache_dir uk_UA-ASMR/cache \
  --data.config_path uk_UA-ASMR/output/uk_UA-asmr-medium.onnx.json \
  --data.batch_size 32 \
  --model.vocoder_warmstart_ckpt uk_UA-ASMR/checkpoints/uk/uk_UA/ukrainian_tts/medium/epoch=2090-step=1166778.ckpt \
  --trainer.max_epochs 500 \
  --trainer.check_val_every_n_epoch 1 \
  --trainer.default_root_dir uk_UA-ASMR/output \
  --ckpt_path uk_UA-ASMR/output/lightning_logs/version_0/checkpoints/epoch=35-step=14832.ckpt

(Check your `ukUA-ASMR/output/lightninglogs/` folder for the exact `.ckpt` filename)

**Note**: Find checkpoints `find /workspace -name "*.ckpt" 2>/dev/null | head -5`

Exporting

bash
# 1. Export the ONNX model from your best/latest checkpoint
python3 -m piper.train.export_onnx \
  --checkpoint uk_UA-ASMR/output/lightning_logs/version_0/checkpoints/epoch=14-step=6180.ckpt \
  --output-file uk_UA-ASMR/output/uk_UA-asmr-medium.onnx

Model Output

After training and export, you will have:

FileDescription
uk_UA-asmr-medium.onnxONNX model for inference
uk_UA-asmr-medium.onnx.jsonModel configuration

Usage with Piper

bash
# Install piper
pip install piper-tts

# Generate speech
# (Pipe the text using 'echo' to avoid CLI parsing errors with raw text modes)
echo "привіт, як справи?" | python3 -m piper \
  --model uk_UA-ASMR/output/uk_UA-asmr-medium.onnx \
  --output_file audio.wav

Phoneme Type

This dataset uses phoneme_type: "text", meaning raw Ukrainian characters are used directly without espeak-ng phonemization. The model uses a character-based phoneme map with Ukrainian Cyrillic characters.

Valid characters:

а б в г ґ д е є ж з и і ї й к л м н о п р с т у ф х ц ч ш щ ь ю я

Plus punctuation: space ! ' , - . : ; ? _ ^ $ — + diacritics

Metadata:

utt_4197.wav|про що ти хочеш мене попросити?
utt_4198.wav|запитала вона підозріло.

Citation

If you use this dataset, please cite:

bibtex
@misc{uk_ua_asmr,
  title={Ukrainian ASMR TTS Dataset},
  author={Kontextox},
  year={2026},
  url={https://huggingface.co/datasets/kontextox/uk_UA-ASMR}
}

License

CC0 - Public Domain

Acknowledgments