litert-community/Audio8-TTS-Preview-0.6b
Audio8-TTS-Preview-0.6b — LiteRT
Edge0/Audio8-TTS-Preview-0.6b (Apache-2.0) converted to LiteRT (.tflite) for on-device text-to-speech with zero-shot voice cloning in 11 languages, at the model's native 44.1 kHz.
Audio8 TTS is a DualAR speech LM in the Fish Audio S2 Pro lineage: a 24-layer "slow" transformer predicts one semantic token per 46 ms frame, a 4-layer "fast" transformer predicts the frame's ten codec codebooks one at a time, and a neural codec (RVQ, an 8-layer windowed transformer and a causal upsampling ConvNet) turns the frames into audio. This repository holds the four graphs plus a Python host loop that reproduces the vendor's generation loop and sampler.
Quick start (Python, desktop)
pip install ai-edge-litert tokenizers soundfile scipy numpy huggingface_hub
hf download litert-community/Audio8-TTS-Preview-0.6b --local-dir audio8
cd audio8
python audio8_tts_litert.py --model-dir . --text "Hello from LiteRT, fully on device." \
--voice voices/en_librispeech_1272 --out hello.wav
# Japanese with the bundled Japanese voice
python audio8_tts_litert.py --model-dir . --text "今日は天気が良いので、公園まで散歩に行きましょう。" \
--voice voices/ja_funasr_example --out konnichiwa.wav
# clone your own voice from a 0.5-10 s clip and its exact transcript
python audio8_tts_litert.py --model-dir . --register-voice me.wav --ref-text "exact transcript of me.wav" --voice-out voices/me
python audio8_tts_litert.py --model-dir . --text "..." --voice voices/me --out out.wav
# generation without a reference voice
python audio8_tts_litert.py --model-dir . --text "This utterance does not use a reference voice." --out noref.wav--slow slow_ar_int4.tflite selects the smaller slow AR. --threads sets the CPU thread count (default 4).
Android
Sample app (Android): `android/` — type a sentence, pick one of the two sample voices or record your own for 10 s, tap Speak. Galaxy S26, the two autoregressive graphs on the CPU (4 threads): RTF 1.12 for an 18-token English sentence with the app on screen (6.0 s of speech in 6.7 s) and 1.42–1.50 in the app's device check, 2026-10-10. On that phone the fp16 codec decoder on the GPU returned the same output for every input (correlation 0.0097 with the CPU int8 decoder's output), so the app runs the int8 decoder on the CPU there.
Files
The sampler and the prompt format are the vendor's, so the loop accepts the same generation parameters (temperature 0.7, top-p 0.9, top-k 50, max 512 frames by default).
Accuracy
Measured against the vendor's PyTorch implementation (transformers 4.57, CPU fp32) on 14 seeded sentences (6 English and 6 Japanese with a cloned reference voice, one of each without a reference), with the same random draws:
- The fp32 graphs reproduce the reference code sequence frame for frame on all 14 sentences; the codec decoder is bit-exact at torch level and within 2e-6 as
.tflite. - Speech recognition of the output (whisper large-v3-turbo) against the input text, and speaker similarity (TitaNet-L cosine) against the reference clip:
The one English error is shared with the reference (the model drops the first word of one sentence). Quantized graphs sample a different but valid trajectory; per-frame agreement with the reference is not a meaningful metric for a sampled model, so the gate is the transcript and the voice.
Performance
Measured, not estimated. The host loop is Python; the Android sample under android/ runs the same graphs from Kotlin through the LiteRT CompiledModel API (its numbers are in the Android section above).
Apple silicon Mac (CPU, 4 threads, ai-edge-litert 2.2.0, audio8_tts_litert.py)
Galaxy S26 (SM-S942Q, Snapdragon 8 Elite Gen 5, Android 16; LiteRT benchmark_model, CPU 4 threads; frequency caps verified absent before each row)
Per generated frame the two AR graphs cost 12.3 + 9.7 = 22 ms on the CPU, i.e. an autoregressive real-time factor of about 0.5 at the codec's 21.5 frames/s, before the codec. A 256-frame decoder does not prepare on the S26 GPU (Dilated im2col buffer size overflowed) and is not shipped; int8 convolutions do not run on that GPU; the fp32 decoder delegates 971/971 ops but runs at the same 924 ms with a 1462 MB footprint; the AR graphs stay on the CPU (their codebook gather is not delegated). Memory figures are the benchmark tool's footprint; the multi-signature slow graph packs its weights once per signature.
Limitations
- The codec decoder is causal but its transformer stacks eight 128-frame windows, so decoding in windows is not sample-exact against a single offline decode; the host loop makes one T128 call for utterances up to 5.9 s, one T192 call up to 8.9 s, and beyond that T192 windows with 128 frames of left context (the same context the vendor's ONNX runtime uses).
- Voice registration accepts up to 10 s of reference audio (one static bucket); the reference transcript must match the audio.
- Prompt length + generated frames must stay under 2048 positions; the default cap is 512 frames (about 24 s) per call.
- Preview checkpoint: the vendor documents limited dialect coverage and sensitivity to noisy or mis-transcribed references. Generated speech can be misused for impersonation; obtain consent before cloning a voice and disclose synthetic audio.
License
Apache-2.0, inherited from the base model by Edge0.
