acul3/chatterbox-executorch
Chatterbox Multilingual TTS โ ExecuTorch Models
Pre-exported .pte model files for running Resemble AI's Chatterbox Multilingual TTS fully on-device using ExecuTorch.
๐ฆ Code & export scripts: acul3/chatterbox-executorch on GitHub
What's Here
9 ExecuTorch .pte files covering the complete TTS pipeline โ from text input to 24kHz waveform โ with zero PyTorch runtime required:
Quick Download
from huggingface_hub import snapshot_download
snapshot_download(
"acul3/chatterbox-executorch",
local_dir="et_models",
repo_type="model"
)Pipeline Overview
Text โ MTLTokenizer โ text tokens
Reference Audio โ VoiceEncoder + CAMPPlus โ speaker conditioning
โ
T3 Prefill (LlamaModel, conditioned)
โ
T3 Decode (autoregressive, ~100 tokens)
โ
S3Gen Encoder (Conformer)
โ
CFM Step ร 2 (flow matching)
โ
HiFiGAN (vocoder, chunked)
โ
24kHz PCM waveform ๐ตKey Technical Notes
- T3 Decode uses a manually unrolled 30-layer Llama forward pass with static KV cache (
torch.wherewrites) โ bypasses HFDynamicCachefortorch.exportcompatibility - HiFiGAN uses a manual real-valued DFT (cosine/sine matrix multiply) โ replaces
torch.stft/torch.istftwhich XNNPACK doesn't support - T3 models are FP16 (XNNPACK half-precision kernels) โ ~half the size of FP32 with near-identical quality
- Fixed shapes: CFM expects
T_MEL=2200, HiFiGAN expectsT_MEL=300(use chunked processing for longer audio)
Usage
See the GitHub repo for full inference code: acul3/chatterbox-executorch
# Clone code
git clone https://github.com/acul3/chatterbox-executorch.git
cd chatterbox-executorch
# Download models (this repo)
python -c "
from huggingface_hub import snapshot_download
snapshot_download('acul3/chatterbox-executorch', local_dir='et_models', repo_type='model')
"
# Run full PTE inference
python test_true_full_pte.pyAndroid Integration
These models are designed for Android deployment via the ExecuTorch Android SDK. Load with:
val module = Module.load(context.filesDir.path + "/t3_prefill.pte")With QNN/NPU delegation on a Snapdragon device, expect 10โ50ร speedup over the CPU timings below.
Performance (Jetson AGX Orin, CPU only)
License
Model weights are derived from Resemble AI's Chatterbox. The export pipeline code is MIT licensed. Please refer to the original Chatterbox license for model weights usage terms.
