Team Ai
Modelpublic

Synaptics/Piper-TTS

sourceHugging Facemitupdated 9d agoView on Hugging Face
0likes392downloads
Model Card

Piper TTS for Synaptics Torq (SL2619)

Piper (VITS) text-to-speech split across CPU and NPU for the Synaptics SL2619 board, in three voices:

VoiceLanguageSample rateSpeakersAssets
en_US-libritts_r-medium (default)English (US)22.05 kHz904top level
en_US-lessac-lowEnglish (US)16 kHz1en_US-lessac-low/
es_MX-ald-mediumSpanish (Mexico)22.05 kHz1es_MX-ald-medium/

The VITS graph is cut where the per-phoneme durations are ceiled and summed — the point at which the output length becomes exact:

text --[espeak]--> phoneme ids --> [partA]  (CPU, onnxruntime)  --> z [1,192,F] (+ g [1,512,1])
                                           z (,g) --> [partB]   (NPU, bf16 vmfb) --> audio [F*256]

partA holds 85% of the nodes (small shape/attention ops) but partB — the HiFi-GAN vocoder — holds most of the time and is pure convolution, which is what the NPU accelerates. Because partA yields the exact frame count F, the right static vocoder window is known before the vocoder runs. g is the speaker embedding: the single-speaker voices have none, so their vocoder takes z alone.

Performance (SL2619 board)

A 3-sentence sample, text in to all audio synthesized (phonemization included, model load excluded), median of 5 runs:

VoiceAudioCPU onlyCPU + NPUSpeedupFirst audio (CPU / CPU + NPU)
en_US-libritts_r-medium7.70 s5.55 s (1.39× RT)2.83 s (2.72× RT)1.96×0.95 s / 0.77 s
en_US-lessac-low9.82 s5.25 s (1.87× RT)2.51 s (3.92× RT)2.09×0.89 s / 0.68 s
es_MX-ald-medium12.63 s8.66 s (1.46× RT)4.07 s (3.10× RT)2.13×1.81 s / 1.38 s

"CPU only" is the full Piper voice in onnxruntime (2 threads), one sentence at a time; "CPU + NPU" is partA on the CPU pipelined with the vocoder on the NPU. The English and Spanish samples are the same three sentences, translated.

The vocoder alone on the NPU, per window:

windowlibritts / Spanish (22.05 kHz)lessac-low (16 kHz)
1 s170 ms (5.9× RT)126 ms (7.8× RT)
1.5 s—186 ms (8.1× RT)
2 s335 ms (6.0× RT)249 ms (8.0× RT)
3 s—365 ms (8.2× RT)
4 s688 ms (5.8× RT)497 ms (8.0× RT)
5 s—628 ms (8.0× RT)
6 s1032 ms (5.8× RT)749 ms (8.0× RT)
8 s1356 ms (5.9× RT)1027 ms (7.8× RT)

The CPU+NPU pipeline overlaps the halves: the CPU encodes sentence n+1 while the NPU vocodes sentence n and the speaker plays sentence n-1.

Accuracy: each bf16 NPU vocoder window matches the fp32 onnxruntime vocoder on the same latent at 37.3–40.7 dB SNR (correlation ≥ 0.9999).

Contents

Per voice (English at the top level, the others under their voice-named folder):

pathwhat it is
onnx/partA.onnxText encoder + duration predictor. Runs on the CPU under onnxruntime.
vmfb/partB_static_<N>s.vmfbThe vocoder compiled for the Torq NPU (NSS-only, bf16), one per window: 1, 2, 4, 6, 8 s (lessac-low adds 1.5, 3, 5 s).
voice/<voice>.onnx.jsonVoice config: phoneme→id map, sample rate, espeak language.
onnx/partB_static_<N>s{.bf16io,}.onnxThe statically shaped bf16-I/O vocoder each VMFB was compiled from.

Shared, and English-only extras:

pathwhat it is
espeak/phonemizerdSmall resident espeak-ng phonemizer daemon (aarch64) mirroring libpiper's phonemization.
espeak/espeak-ng-data.tar.gzespeak-ng dictionaries; one copy covers every language.
onnx/en_US-libritts_r-medium.onnxThe original monolithic Piper voice, for reference.
tflite/partB_static_4s.{int8,int16x8}.tfliteQuantized vocoder conversions (int8 was slower than bf16 on this target; kept for reference).

The vocoder ships as one VMFB per window because the NPU model is statically shaped: 1, 2, 4, 6 and 8 seconds of audio, plus 1.5, 3 and 5 s for en_US-lessac-low. Per sentence, pick the smallest window that fits, edge-pad the latent up to it, and trim the output back to F × 256 samples. lessac-low's finer windows cut padding from 26% to 16% of vocoder frames (NPU vocoder time −12.5% on a 34-sentence mix; 1–3% end to end, since that voice is CPU-bound once the vocoder is on the NPU).

How these were built

Every voice is produced by one command in torq-tools:

sh
torq-export-model piper -v es_MX-ald-medium     # any rhasspy/piper-voices key

It downloads the voice from rhasspy/piper-voices (pinned to v1.0.0), splits it, pins the five vocoder windows (sized from the voice's own sample rate), converts them to bf16 with bf16 I/O, and compiles each NSS-only (--torq-disable-css --torq-disable-host). The VMFBs here were compiled with the Torq compiler including the HiFi-GAN vocoder fixes (not yet on main; these builds are about 30% faster on the NPU than the first release of this repo, with bit-identical output).

Usage

These files are consumed by the piper_tts demo in synaptics-torq/torq-examples, which writes a .wav and plays it on the board's speaker:

sh
cd piper_tts
pip install -r requirements.txt
cd .. && python setup_demos.py piper_tts
cd piper_tts
python src/infer.py --text "Hello from the Synaptics board."
python src/infer.py --voice en_US-lessac-low --interactive
python src/infer.py --voice es_MX-ald-medium --text "Buenos días."

Only the requested voice is downloaded. The demo reads each window's frame width from the VMFB signature and each voice's sample rate from its config, so recompiling with a different set of windows needs no code change.

Licensing

The Piper voices and models are MIT. espeak/phonemizerd links espeak-ng (GPLv3) statically; its source ships with the demo at piper_tts/piper_core/phonemizerd.c in torq-examples, with the build command in its header comment.