roman220220/ACE-Step1.5-sft-MLX-8bit
<p align="center"> <img src="llmtray-banner.png" alt="LLMTray" width="100%"> </p>
ACE-Step 1.5 sft for MLX: 8-bit
### ▶ Run it locally in LLMTray A free, native macOS app for local AI on Apple Silicon — chat, images, music, agents and an OpenAI-compatible API. Point it at this model; nothing leaves your Mac.  
ACE-Step 1.5 sft (ACE-Step team, MIT, trained on licensed and royalty-free data) is converted for Apple Silicon here: songs with sung, intelligible lyrics from a style description, 48 kHz stereo.
DiT and condition encoders at 8-bit (round-to-nearest, group 64): practically indistinguishable from bf16.
sft vs turbo
ACE-Step 1.5 ships two DiTs. We compared them on 3 songs × 4 seeds: pop, rock and a ballad, 30 s each. The score is Whisper large-v3-turbo's word error rate against the lyrics that were asked for (lower = the lyrics come through more clearly).
By ear: turbo's mix sounds fuller and more like a finished song, while sft puts the voice upfront and gets the words across. On electronic music sft sounded best.
Quantization
Teacher-forced error against bf16: every decoder call of the bf16 run, fed to each variant.
Lyrics intelligibility, the same 3 songs × 4 seeds (Whisper WER, lower is better; 12 tracks, so differences of a few points are noise):
Pick bf16 or 8-bit if you have the memory, GPTQ 4-bit on a 16 GB Mac.
⚠️ How to run it: the stock mlx-audio settings sing gibberish
Pulled through mlx-audio's defaults (branch `pc/add-ace`, commit 1e8264a), the sft DiT produces cacophony. Two causes, both found and measured (findings):
- The 5 Hz LM hints. mlx-audio feeds its planner's audio codes into the DiT, and that breaks the sft model. sft doesn't need a planner: run it with
use_lm=False. It is also ~15 s faster. - The unconditional CFG branch. mlx-audio encodes all-zero text there. The official pipeline uses the trained
null_condition_emb. The patch below uses it.
import mlx.core as mx, mlx.nn as nn
import mlx_audio.utils as u
from mlx_audio.tts import load
# a local folder doesn't tell mlx-audio which model class it is
pick = u.get_model_class
u.get_model_class = lambda model_type, model_name, category, model_remapping: pick("ace_step", None, category, model_remapping)
model = load("roman220220/ACE-Step1.5-sft-MLX-8bit")
class NullAwareEncoder(nn.Module): # CFG's unconditional branch, as the official pipeline does it
def __init__(self, inner, null):
super().__init__(); self.inner, self._null = inner, null
def __call__(self, text_hidden_states=None, lyric_hidden_states=None, **kw):
out, mask = self.inner(text_hidden_states=text_hidden_states, lyric_hidden_states=lyric_hidden_states, **kw)
if not mx.any(text_hidden_states).item() and not mx.any(lyric_hidden_states).item():
out = mx.broadcast_to(self._null.astype(out.dtype), out.shape)
return out, mask
model.encoder = NullAwareEncoder(model.encoder, model.null_condition_emb)
result = list(model.generate(
text="upbeat pop song with female vocals, bright synths, driving beat",
lyrics="[Verse]\nCity lights are calling out my name\n...",
duration=30.0, seed=1, vocal_language="en",
use_lm=False, num_steps=50, guidance_scale=7.0, shift=1.0, guidance_interval=1.0, cfg_type="apg",
))[-1] # result.audio: [2, samples] at 48 kHzThe IPSupport local-AI stack
Local-first AI tools for macOS by IPSupport — nothing leaves your Mac.
- [LLMTray](https://www.ipsupport.us/llmtray/) — your local AI workstation for macOS: chat with local LLMs, generate and edit images, make music, run agents, and serve an OpenAI-compatible API. Downloads models from Hugging Face in-app, with per-model profiles.
- [IPSupport Code](https://ipsupport-llc.github.io/ipsupport-code/) — your AI coding agent for real repositories: analyze, fix, test, report.
LLMTray uses this model for in-chat music generation.
License
Licensed under the MIT License, the same license as the ACE-Step 1.5 model. See `LICENSE`.
This is a modified version of ACE-Step/acestep-v15-sft, together with the ACE-Step 1.5 VAE and its Qwen3-Embedding-0.6B text encoder:
- converted to mlx-audio's MLX layout: the decoder's Conv1d weights transposed and the rotary caches dropped;
- the DiT and condition encoders quantized to 8-bit.
