Team Ai
Modelpublic

acul3/chatterbox-executorch

sourceHugging Faceapache-2.0updated 7mo agoView on Hugging Face
1likes421downloads
Model Card

Chatterbox Multilingual TTS โ€” ExecuTorch Models

Pre-exported .pte model files for running Resemble AI's Chatterbox Multilingual TTS fully on-device using ExecuTorch.

๐Ÿ“ฆ Code & export scripts: acul3/chatterbox-executorch on GitHub


What's Here

9 ExecuTorch .pte files covering the complete TTS pipeline โ€” from text input to 24kHz waveform โ€” with zero PyTorch runtime required:

FileSizeBackendPrecisionStage
voice_encoder.pte7 MBportableFP32Speaker embedding
xvector_encoder.pte27 MBportableFP32X-vector conditioning
t3_cond_speech_emb.pte49 MBportableFP32Speech token embedding
t3_cond_enc.pte18 MBportableFP32Text/conditioning encoder
t3_prefill.pte1010 MBXNNPACKFP16T3 Transformer prefill
t3_decode.pte1002 MBXNNPACKFP16T3 Transformer decode
s3gen_encoder.pte178 MBportableFP32S3Gen Conformer encoder
cfm_step.pte274 MBXNNPACKFP32CFM flow matching step
hifigan.pte84 MBXNNPACKFP32HiFiGAN vocoder
Total~2.6 GB

Quick Download

python
from huggingface_hub import snapshot_download

snapshot_download(
    "acul3/chatterbox-executorch",
    local_dir="et_models",
    repo_type="model"
)

Pipeline Overview

Text โ†’ MTLTokenizer โ†’ text tokens
Reference Audio โ†’ VoiceEncoder + CAMPPlus โ†’ speaker conditioning
                          โ†“
              T3 Prefill (LlamaModel, conditioned)
                          โ†“
              T3 Decode (autoregressive, ~100 tokens)
                          โ†“
              S3Gen Encoder (Conformer)
                          โ†“
              CFM Step ร— 2 (flow matching)
                          โ†“
              HiFiGAN (vocoder, chunked)
                          โ†“
              24kHz PCM waveform ๐ŸŽต

Key Technical Notes

  • โ€”T3 Decode uses a manually unrolled 30-layer Llama forward pass with static KV cache (torch.where writes) โ€” bypasses HF DynamicCache for torch.export compatibility
  • โ€”HiFiGAN uses a manual real-valued DFT (cosine/sine matrix multiply) โ€” replaces torch.stft/torch.istft which XNNPACK doesn't support
  • โ€”T3 models are FP16 (XNNPACK half-precision kernels) โ€” ~half the size of FP32 with near-identical quality
  • โ€”Fixed shapes: CFM expects T_MEL=2200, HiFiGAN expects T_MEL=300 (use chunked processing for longer audio)

Usage

See the GitHub repo for full inference code: acul3/chatterbox-executorch

bash
# Clone code
git clone https://github.com/acul3/chatterbox-executorch.git
cd chatterbox-executorch

# Download models (this repo)
python -c "
from huggingface_hub import snapshot_download
snapshot_download('acul3/chatterbox-executorch', local_dir='et_models', repo_type='model')
"

# Run full PTE inference
python test_true_full_pte.py

Android Integration

These models are designed for Android deployment via the ExecuTorch Android SDK. Load with:

kotlin
val module = Module.load(context.filesDir.path + "/t3_prefill.pte")

With QNN/NPU delegation on a Snapdragon device, expect 10โ€“50ร— speedup over the CPU timings below.

Performance (Jetson AGX Orin, CPU only)

StageTime
Voice encoding~1s
T3 prefill~22s
T3 decode (~100 tokens)~800s total (~8s/token)
S3Gen encoder~2s
CFM (2 steps)~40s
HiFiGAN~10s/chunk

License

Model weights are derived from Resemble AI's Chatterbox. The export pipeline code is MIT licensed. Please refer to the original Chatterbox license for model weights usage terms.