Synaptics/Piper-TTS
Piper TTS for Synaptics Torq (SL2619)
Piper (VITS) text-to-speech split across CPU and NPU for the Synaptics SL2619 board, in three voices:
The VITS graph is cut where the per-phoneme durations are ceiled and summed — the point at which the output length becomes exact:
text --[espeak]--> phoneme ids --> [partA] (CPU, onnxruntime) --> z [1,192,F] (+ g [1,512,1])
z (,g) --> [partB] (NPU, bf16 vmfb) --> audio [F*256]partA holds 85% of the nodes (small shape/attention ops) but partB — the HiFi-GAN vocoder — holds most of the time and is pure convolution, which is what the NPU accelerates. Because partA yields the exact frame count F, the right static vocoder window is known before the vocoder runs. g is the speaker embedding: the single-speaker voices have none, so their vocoder takes z alone.
Performance (SL2619 board)
A 3-sentence sample, text in to all audio synthesized (phonemization included, model load excluded), median of 5 runs:
"CPU only" is the full Piper voice in onnxruntime (2 threads), one sentence at a time; "CPU + NPU" is partA on the CPU pipelined with the vocoder on the NPU. The English and Spanish samples are the same three sentences, translated.
The vocoder alone on the NPU, per window:
The CPU+NPU pipeline overlaps the halves: the CPU encodes sentence n+1 while the NPU vocodes sentence n and the speaker plays sentence n-1.
Accuracy: each bf16 NPU vocoder window matches the fp32 onnxruntime vocoder on the same latent at 37.3–40.7 dB SNR (correlation ≥ 0.9999).
Contents
Per voice (English at the top level, the others under their voice-named folder):
Shared, and English-only extras:
The vocoder ships as one VMFB per window because the NPU model is statically shaped: 1, 2, 4, 6 and 8 seconds of audio, plus 1.5, 3 and 5 s for en_US-lessac-low. Per sentence, pick the smallest window that fits, edge-pad the latent up to it, and trim the output back to F × 256 samples. lessac-low's finer windows cut padding from 26% to 16% of vocoder frames (NPU vocoder time −12.5% on a 34-sentence mix; 1–3% end to end, since that voice is CPU-bound once the vocoder is on the NPU).
How these were built
Every voice is produced by one command in torq-tools:
torq-export-model piper -v es_MX-ald-medium # any rhasspy/piper-voices keyIt downloads the voice from rhasspy/piper-voices (pinned to v1.0.0), splits it, pins the five vocoder windows (sized from the voice's own sample rate), converts them to bf16 with bf16 I/O, and compiles each NSS-only (--torq-disable-css --torq-disable-host). The VMFBs here were compiled with the Torq compiler including the HiFi-GAN vocoder fixes (not yet on main; these builds are about 30% faster on the NPU than the first release of this repo, with bit-identical output).
Usage
These files are consumed by the piper_tts demo in synaptics-torq/torq-examples, which writes a .wav and plays it on the board's speaker:
cd piper_tts
pip install -r requirements.txt
cd .. && python setup_demos.py piper_tts
cd piper_tts
python src/infer.py --text "Hello from the Synaptics board."
python src/infer.py --voice en_US-lessac-low --interactive
python src/infer.py --voice es_MX-ald-medium --text "Buenos días."Only the requested voice is downloaded. The demo reads each window's frame width from the VMFB signature and each voice's sample rate from its config, so recompiling with a different set of windows needs no code change.
Licensing
The Piper voices and models are MIT. espeak/phonemizerd links espeak-ng (GPLv3) statically; its source ships with the demo at piper_tts/piper_core/phonemizerd.c in torq-examples, with the build command in its header comment.
