Team Ai
Modelpublic

litert-community/Audio8-TTS-Preview-0.6b

sourceHugging Faceapache-2.0updated 1d agoView on Hugging Face
1likes187downloads
Model Card

Audio8-TTS-Preview-0.6b — LiteRT

Edge0/Audio8-TTS-Preview-0.6b (Apache-2.0) converted to LiteRT (.tflite) for on-device text-to-speech with zero-shot voice cloning in 11 languages, at the model's native 44.1 kHz.

Audio8 TTS is a DualAR speech LM in the Fish Audio S2 Pro lineage: a 24-layer "slow" transformer predicts one semantic token per 46 ms frame, a 4-layer "fast" transformer predicts the frame's ten codec codebooks one at a time, and a neural codec (RVQ, an 8-layer windowed transformer and a causal upsampling ConvNet) turns the frames into audio. This repository holds the four graphs plus a Python host loop that reproduces the vendor's generation loop and sampler.

Quick start (Python, desktop)

bash
pip install ai-edge-litert tokenizers soundfile scipy numpy huggingface_hub
hf download litert-community/Audio8-TTS-Preview-0.6b --local-dir audio8
cd audio8
python audio8_tts_litert.py --model-dir . --text "Hello from LiteRT, fully on device." \
    --voice voices/en_librispeech_1272 --out hello.wav
# Japanese with the bundled Japanese voice
python audio8_tts_litert.py --model-dir . --text "今日は天気が良いので、公園まで散歩に行きましょう。" \
    --voice voices/ja_funasr_example --out konnichiwa.wav
# clone your own voice from a 0.5-10 s clip and its exact transcript
python audio8_tts_litert.py --model-dir . --register-voice me.wav --ref-text "exact transcript of me.wav" --voice-out voices/me
python audio8_tts_litert.py --model-dir . --text "..." --voice voices/me --out out.wav
# generation without a reference voice
python audio8_tts_litert.py --model-dir . --text "This utterance does not use a reference voice." --out noref.wav

--slow slow_ar_int4.tflite selects the smaller slow AR. --threads sets the CPU thread count (default 4).

Android

Sample app (Android): `android/` — type a sentence, pick one of the two sample voices or record your own for 10 s, tap Speak. Galaxy S26, the two autoregressive graphs on the CPU (4 threads): RTF 1.12 for an 18-token English sentence with the app on screen (6.0 s of speech in 6.7 s) and 1.42–1.50 in the app's device check, 2026-10-10. On that phone the fp16 codec decoder on the GPU returned the same output for every input (correlation 0.0097 with the CPU int8 decoder's output), so the app runs the int8 decoder on the CPU there.

Files

FileSizeRole
slow_ar_int8.tflite552 MBSlow AR (24 layers). Signatures prefill_256 and decode; KV cache 2048 (the model's maxseqlen) as graph I/O; outputs the 4097 logits the sampler can pick (4096 semantic tokens + end-of-speech) and the normalized hidden state that conditions the fast AR. Dynamic int8 (per-channel) projections, int8 embedding tables.
slow_ar_int4.tflite386 MBSame graph, blockwise-32 OCTAV int4 projections.
fast_ar_int8.tflite68 MBFast AR step (4 layers, 10-slot KV cache as I/O). Called 10 times per frame: position 0 takes the slow hidden state, positions 1-9 take the previous codebook token. Dynamic int8.
codec_decoder_fp16_T128.tflite262 MBCodec decoder for up to 128 frames (5.9 s) per call; fp16 weights. Runs on the mobile GPU.
codec_decoder_fp16_T192.tflite262 MBSame decoder for up to 192 frames (8.9 s) per call; the host loop uses it for longer text as 128-frame-context windows. Runs on the mobile GPU.
codec_decoder_int8_T128.tflite132 MBCodec decoder, all convolutions and projections int8 (export-time PT2E, fp32 codebooks); the CPU option.
codec_encoder_fp16_10s.tflite419 MBCodec encoder for voice registration: 10.03 s of 44.1 kHz audio -> 10 x 216 codes.
tokenizer.json12 MBThe vendor's tokenizer (Qwen2 BPE + the `<\semantic:N\>` and role tokens), unchanged.
voices/*/codes.npy, meta.json6 KBTwo bundled reference voices (codec codes + transcript): an English LibriSpeech dev-clean speaker (CC BY 4.0) and the Japanese example clip from FunAudioLLM/Fun-ASR-Nano-2512 (Apache-2.0).
audio8_tts_litert.py14 KBHost loop: prompt construction, chunked prefill, the vendor's sampler (top-k / top-p / temperature, repetition-aware re-draw), fast-AR loop, windowed codec decode, voice registration.

The sampler and the prompt format are the vendor's, so the loop accepts the same generation parameters (temperature 0.7, top-p 0.9, top-k 50, max 512 frames by default).

Accuracy

Measured against the vendor's PyTorch implementation (transformers 4.57, CPU fp32) on 14 seeded sentences (6 English and 6 Japanese with a cloned reference voice, one of each without a reference), with the same random draws:

  • —The fp32 graphs reproduce the reference code sequence frame for frame on all 14 sentences; the codec decoder is bit-exact at torch level and within 2e-6 as .tflite.
  • —Speech recognition of the output (whisper large-v3-turbo) against the input text, and speaker similarity (TitaNet-L cosine) against the reference clip:
configurationen WERja CERspeaker cosine en / ja
PyTorch reference1.1%0.0%0.66 / 0.74
slow_ar_int8 + fast_ar_int8 + codec_decoder_fp161.1%0.0%0.68 / 0.76
slow_ar_int4 + fast_ar_int8 + codec_decoder_fp161.1%0.0%0.63 / 0.77
codec_decoder_int8 (reference codes decoded)1.1%0.0%0.67 / 0.73

The one English error is shared with the reference (the model drops the first word of one sentence). Quantized graphs sample a different but valid trajectory; per-frame agreement with the reference is not a meaningful metric for a sampled model, so the gate is the transcript and the voice.

Performance

Measured, not estimated. The host loop is Python; the Android sample under android/ runs the same graphs from Kotlin through the LiteRT CompiledModel API (its numbers are in the Android section above).

Apple silicon Mac (CPU, 4 threads, ai-edge-litert 2.2.0, audio8_tts_litert.py)

configurationslow AR / framefast AR / frame (10 calls)codec (T128 call)RTF (median of 14)
int8 / int8 / fp169.0 ms9.2 ms1.43 s0.81 (0.73-0.98)
int4 / int8 / fp1612.2 ms9.2 ms1.43 s0.92 (0.81-1.07)
int8 / int8 / int8 codec8.9 ms9.4 ms0.96 s0.70 (0.64-0.82)

Galaxy S26 (SM-S942Q, Snapdragon 8 Elite Gen 5, Android 16; LiteRT benchmark_model, CPU 4 threads; frequency caps verified absent before each row)

graphsignaturebackendinference (avg)init / overall memory
slow_ar_int8.tflitedecodeCPU12.3 ms1077 / 1176 MB
slow_ar_int8.tfliteprefill_256CPU241 ms1077 / 1249 MB
slow_ar_int4.tflitedecodeCPU10.0 ms610 / 709 MB
slow_ar_int4.tfliteprefill_256CPU417 ms610 / 783 MB
fast_ar_int8.tflitestepCPU0.97 ms127 / 127 MB
codec_decoder_fp16_T128.tflitedecodeGPU (OpenCL, 936 of 1069 ops)923 ms per 5.9 s1050 MB
codec_decoder_fp16_T192.tflitedecodeGPU (OpenCL)1425 ms per 8.9 s1109 MB
codec_decoder_fp16_T128.tflitedecodeCPU3664 ms per 5.9 s862 / 1433 MB
codec_decoder_int8_T128.tflitedecodeCPU1959 ms per 5.9 s259 / 836 MB
codec_encoder_fp16_10s.tfliteencodeCPU2365 ms per 10 s1206 / 1899 MB

Per generated frame the two AR graphs cost 12.3 + 9.7 = 22 ms on the CPU, i.e. an autoregressive real-time factor of about 0.5 at the codec's 21.5 frames/s, before the codec. A 256-frame decoder does not prepare on the S26 GPU (Dilated im2col buffer size overflowed) and is not shipped; int8 convolutions do not run on that GPU; the fp32 decoder delegates 971/971 ops but runs at the same 924 ms with a 1462 MB footprint; the AR graphs stay on the CPU (their codebook gather is not delegated). Memory figures are the benchmark tool's footprint; the multi-signature slow graph packs its weights once per signature.

Limitations

  • —The codec decoder is causal but its transformer stacks eight 128-frame windows, so decoding in windows is not sample-exact against a single offline decode; the host loop makes one T128 call for utterances up to 5.9 s, one T192 call up to 8.9 s, and beyond that T192 windows with 128 frames of left context (the same context the vendor's ONNX runtime uses).
  • —Voice registration accepts up to 10 s of reference audio (one static bucket); the reference transcript must match the audio.
  • —Prompt length + generated frames must stay under 2048 positions; the default cap is 512 frames (about 24 s) per call.
  • —Preview checkpoint: the vendor documents limited dialect coverage and sensitivity to noisy or mis-transcribed references. Generated speech can be misused for impersonation; obtain consent before cloning a voice and disclose synthetic audio.

License

Apache-2.0, inherited from the base model by Edge0.