Team Ai
Modelpublic

daksh-neo/Moss_tts_cpu_optimised

sourceHugging Faceapache-2.0updated 6mo agoView on Hugging Face
2likes23downloads
Model Card

MOSS-TTS (CPU Optimized)

๐Ÿš€ CPU Optimized Version: This repository contains a specialized build of MOSS-TTS that has been specifically optimized for high-performance execution on CPU-only environments.

This optimization and packaging process was performed autonomously by [NEO](https://heyneo.so/), an autonomous ML engineering agent.

Overview

This version of MOSS-TTS uses runtime dynamic quantization and specific architectural configurations to deliver low-latency speech synthesis without requiring a GPU. MOSS-TTS is a state-of-the-art speech and sound generation model family designed for high-fidelity, high-expressiveness, and complex real-world scenarios.

Key Optimizations by NEO:

  • โ€”Dynamic INT8 Quantization: Reduces memory footprint and accelerates inference on modern CPUs.
  • โ€”Thread Scaling: Configured for optimal multi-threaded performance.
  • โ€”CPU-Friendly Tensors: Ensured all weights and buffers are optimized for FP32/INT8 execution paths.
  • โ€”Autonomous Validation: Verified functionality in resource-constrained environments.

๐Ÿ›  Usage

Installation

bash
pip install transformers torch torchaudio

Quick Start

python
from transformers import AutoModel, AutoProcessor
import torch

# Load the CPU-optimized model
model_name = "daksh-neo/MOSS-TTS"
processor = AutoProcessor.from_pretrained(model_name, trust_remote_code=True)
model = AutoModel.from_pretrained(
    model_name, 
    trust_remote_code=True,
    torch_dtype=torch.float32 
)

# Inference (Example)
text = "This is a CPU-optimized speech synthesis by NEO."
inputs = processor(text=[text], mode="generation")
outputs = model.generate(**inputs)

๐Ÿ“Š Capabilities

  • โ€”Zero-shot Voice Cloning: Clone voices from short reference clips.
  • โ€”Multilingual Support: High-quality synthesis across 20+ languages.
  • โ€”Long-form Stability: Synthesize stable audio for durations up to 1 hour.
  • โ€”Fine-grained Control: Phoneme-level and duration-level control for precise prosody.

๐Ÿ— Architecture

This specific export is based on the MossTTSDelay architecture, optimized for sequential stability and CPU throughput.

FeatureSpecification
Optimization EngineNEO (Autonomous ML Agent)
Device TargetCPU (x86_64 / ARM64)
QuantizationDynamic INT8
Sampling Rate24kHz / 44.1kHz (Configurable)

๐Ÿ“œ License

This model is released under the Apache-2.0 License.

๐Ÿค Acknowledgments

Original model by MOSI.AI and the OpenMOSS Team. CPU Optimization and Hugging Face packaging by [NEO](https://heyneo.so/).