Team Ai
Modelpublic

FluidInference/supertonic-3-coreml

sourceHugging Faceopenrail++updated 3d agoView on Hugging Face
8likes3kdownloads
1# Supertonic-3 — CoreML conversion2 3Hand-port of [Supertone Supertonic-3 v1.7.3](https://huggingface.co/Supertone/supertonic-3)4from ONNX to PyTorch to CoreML. 31 languages, 44.1 kHz, flow-matching5diffusion (8 denoising steps, classifier-free guidance baked into the6ONNX graph via batch-2 duplication).7 8End-to-end pipeline:9 10```11text → UnicodeProcessor → token_ids, text_mask12   ├── duration_predictor → duration_sec13   └── text_encoder       → text_emb [B, 256, T]14                          ↓15        sample_noisy_latent(duration_sec) → noisy [B, 144, L], latent_mask16                          ↓17        for 8 steps: vector_estimator(noisy, text_emb, style, masks, step, total)18                          ↓19                       vocoder(denoised_latent) → wav [B, 512*6*L]20```21 22Audio chunk granularity:23- AE / vocoder frame: 512 / 44100 ≈ **11.6 ms**24- TTL latent slot (model "tick"): 512 × 6 / 44100 ≈ **69.7 ms**25 26## Layout27 28```29models/tts/supertonic-3/30├── README.md31├── pyproject.toml          # uv project (Python 3.11, torch + coremltools 8)32└── coreml/33    ├── trials.md           # numerical-parity bug log (4 vector_estimator gotchas)34    ├── __init__.py35    ├── common.py           # ONNX-graph loader utilities (assign_param, etc.)36    ├── text_encoder.py     # PyTorch port: build_text_encoder_from_onnx37    ├── duration_predictor.py38    ├── vector_estimator.py39    ├── vocoder.py40    ├── convert_coreml.py   # PyTorch trace -> .mlpackage for all 4 modules41    ├── validate.py         # ONNX vs PyTorch parity check42    ├── verify_coreml.py    # CoreML vs PyTorch parity check43    ├── infer.py            # end-to-end PyTorch TTS driver (text -> wav)44    └── infer_coreml.py     # end-to-end CoreML TTS driver (text -> wav)45```46 47## Setup48 49```bash50cd models/tts/supertonic-3/51uv sync52 53# Fetch upstream ONNX + style + tokenizer assets54mkdir -p build/_onnx build/voice_styles55HF=https://huggingface.co/Supertone/supertonic-3/resolve/main56for f in text_encoder duration_predictor vector_estimator vocoder; do57    curl -L $HF/_onnx/${f}.onnx -o build/_onnx/${f}.onnx58done59curl -L $HF/_onnx/tts.json -o build/_onnx/tts.json60curl -L $HF/_onnx/unicode_indexer.json -o build/_onnx/unicode_indexer.json61curl -L $HF/voice_styles/M1.json -o build/voice_styles/M1.json62```63 64## Convert65 66```bash67# FP32 (numerical reference; ALL modules fall back to CPU on ANE)68uv run python -m coreml.convert_coreml build/_onnx --out-dir build/_mlpackage69 70# FP16 (required for ANE residency; 3/4 modules land on ANE — see Profile below)71uv run python -m coreml.convert_coreml build/_onnx --fp16 --out-dir build/_mlpackage_fp1672 73# Fixed-shape VectorEstimator variant for ANE profiling (RangeDim/Enum hit74# ANE shape limits — see trials.md "Dynamic shapes vs ANE"):75uv run python -m coreml.convert_ve_fixed \76    --onnx build/_onnx/vector_estimator.onnx \77    --out  build/_mlpackage_fp16_fixed/VectorEstimator_L128.mlpackage \78    --L 128 --T 12879```80 81Produces four `.mlpackage` bundles (FP32 ~380 MB, FP16 ~190 MB; mlprogram,82iOS 18+):83 84| Module             | FP32  | FP16   | Variable axes                             |85| ------------------ | ----- | ------ | ----------------------------------------- |86| vocoder            | 97 MB | 48 MB  | `latent.L_ttl` = RangeDim(4..512)         |87| text_encoder       | 35 MB | 17 MB  | fixed `text.T = 128`                      |88| duration_predictor | 3.5 MB| 1.8 MB | fixed `text.T = 128`                      |89| vector_estimator   | 244 MB| 122 MB | `latent.L` & `text.T` = RangeDim(17..512) |90 91## Validate92 93```bash94# ONNX vs PyTorch port (per module)95uv run python -m coreml.validate96 97# CoreML vs PyTorch port (per module)98uv run python -m coreml.verify_coreml99 100# End-to-end PyTorch (writes WAV)101uv run python -m coreml.infer \102    --onnx-dir build/_onnx \103    --voice-style build/voice_styles/M1.json \104    --text "Hello world."105 106# End-to-end CoreML (writes WAV)107uv run python -m coreml.infer_coreml \108    --mlpackage-dir build/_mlpackage \109    --tts-json build/_onnx/tts.json \110    --unicode-indexer build/_onnx/unicode_indexer.json \111    --voice-style build/voice_styles/M1.json \112    --text "Hello world."113```114 115Final parity vs ONNX-Runtime CPU:116 117| Module             | PyTorch vs ONNX max_abs | CoreML vs PyTorch max_abs |118| ------------------ | ----------------------- | ------------------------- |119| vocoder            | 2.53e-4                 | 1.41e-6                   |120| text_encoder       | 9.77e-2 (relaxed tol)   | 2.33e-4                   |121| duration_predictor | 3.04e-6                 | 3.82e-6                   |122| vector_estimator   | 1.21e-3                 | 2.96e-5                   |123 124End-to-end CoreML on M-series CPU+ANE: **~0.74 s** to synthesize1256.32 s of audio for a single English sentence (RTFx ≈ 8.5x), 8126denoising steps. ASR-verified against FluidAudio Parakeet TDT.127 128## Profile (FP16, Apple M2, macOS 26.5, `cpu_and_neural_engine`)129 130| Module                              | CPU% | GPU% | ANE% | Predict | Notes |131| ----------------------------------- | ---- | ---- | ---- | ------- | ----- |132| duration_predictor                  | 100  | 0    | 0    | 0.82 ms | tiny, CPU-bound |133| text_encoder (T=128)                | 38   | 0    | 62   | 2.15 ms | partial ANE |134| vocoder (RangeDim L 4..512)         | 0    | 0    | 100  | 1.17 ms | full ANE, 4× vs FP32 |135| vector_estimator (RangeDim 17..512) | —    | —    | —    | —       | dynamic shapes crash on ANE — must bucket to fixed L |136| vector_estimator (fixed L=128 T=128)| 6    | 0    | 94   | 3.8 ms  | **lands on ANE** (M5 Pro): NE 3.82 ms vs CPU-only 14.20 ms = 3.7×. `ANECCompile FAILED` msg is non-fatal — see trials.md "M5 Pro re-evaluation" |137| vector_estimator (fixed L=256/512)  | 4    | 0    | 96   | 8.4 / 16.4 ms | ANE holds across buckets; int8 halves size (64.5 MB) at same speed/parity 41.5 dB |138 139See `coreml/trials.md` → "ANE residency profiling" for the full breakdown,140the float-mask refactor that eliminated the bool-tile blocker, the141residual opaque `ANECCompile() FAILED (11)`, and the EnumeratedShapes142runtime stride gotcha.143 144## Critical gotchas145 146See `coreml/trials.md` for the full log. Highlights:147 1481. **CFG via batch-2 duplication** — the ONNX vector_estimator tiles149   inputs to batch=2, runs cond + uncond in parallel, then combines150   with `(noisy + (1/total)*(4*cond - 3*uncond)) * mask`. The cond151   style key is **not** the user `style_ttl` — it is a learned152   initializer at `/vector_estimator/Expand_output_0`.1532. **Rotary is length-normalized** — `angles = (pos / sum(mask)) * theta`,154   divisor differs for Q (latent_mask) and K (text_mask).1553. **Attention divisor is 16.0**, not `sqrt(dk)=8`. Off-by-2x in scoring.1564. **Style attention applies `tanh(K)`** before the score matmul; text157   attention does not.1585. **Replicate-pad lower bound** — ConvNeXt depthwise pads scale with159   dilation: `pad = (K-1)*D/2`. CoreML enforces `pad ≤ dim-1` at load160   time, hence `RangeDim.lower_bound = 17` for vector_estimator and161   `4` for vocoder.1626. **int32 vs int64 tokens** — CoreML wants int32, PyTorch indexes int64.163   Wrap modules with a tiny `_Int32Wrapper` that casts inside the164   traced graph so the external input stays int32.1657. **Python 3.14 has no BlobWriter** — pin `requires-python = ">=3.11,<3.13"`.1668. **Float masking, not bool masking** — `masked_fill(mask==0, -inf)` and167   `where(mask==0, 0, attn)` compile to `bool tile`/`select` ops that ANE168   rejects. Use `scores - (1.0 - mask) * 1e4` (additive) and `attn * mask`169   (multiplicative) instead. Lifts vector_estimator from 89.6% → 93.0%170   ANE-eligible (though the residual opaque `ANECCompile() FAILED (11)`171   still blocks final ANE landing — see trials.md).1729. **coremltools `_int` cast with (1,) tensor** — `aten::Int` on a173   (1,)-shape int tensor trips `TypeError: only 0-dimensional arrays can174   be converted to Python scalars` inside coremltools' `_cast` handler.175   `convert_coreml.py` monkey-patches `_cast` (`_patch_int_cast`) to176   squeeze (1,) → scalar before forwarding.177 178## Upstream + downstream179 180- Upstream: <https://huggingface.co/Supertone/supertonic-3>181- Reference Python driver: <https://github.com/supertone-inc/supertonic/blob/main/py/helper.py>182- Republished CoreML: `FluidInference/supertonic-3-coreml` (HuggingFace)183- FluidAudio Swift integration: `Sources/FluidAudio/TTS/Supertonic3/`184