Team Ai
Apppublic

WilliamCass/voice-synthesis

sourceHugging Facemitupdated 11d agoView on Hugging Face
0likes
App README

Voice Synthesis with Voice Cloning

AI-powered Text-to-Speech application with multilingual voice cloning using XTTS v2.

Features

  • โ€”๐ŸŽค Voice Cloning: Upload any audio file or use sample voices in English, German, and Chinese
  • โ€”๐ŸŒ Multilingual: Synthesizes English, German, and Simplified Chinese (zh-cn)
  • โ€”โšก Real-time Synthesis: Fast speech generation
  • โ€”๐ŸŽจ Modern Web Interface: Responsive HTML/JS frontend
  • โ€”๐Ÿ“ Multiple Audio Formats: WAV, MP3, M4A, FLAC, OGG support

Technical Details

  • โ€”TTS Model: XTTS v2 (Coqui AI) - Multilingual zero-shot voice cloning
  • โ€”Backend: FastAPI 0.104.1 + Uvicorn
  • โ€”Audio Processing: librosa + FFmpeg + soundfile
  • โ€”Deep Learning: PyTorch 2.10+ with CUDA support

API Endpoints

GET /health

Check if TTS model is ready

GET /languages

Get supported synthesis languages (returns {en, de, zh-cn})

GET /samples

Get available sample audio files

POST /synthesize

Synthesize speech with voice cloning

Parameters:

  • โ€”text (string, required): Text to synthesize
  • โ€”language (string): Language code (en, de, or zh-cn, default: en)
  • โ€”reference_audio (file, optional): Upload custom audio for voice cloning
  • โ€”sample_audio (string, optional): Use a sample audio filename returned by /samples

Response: WAV audio file

For Chinese synthesis, select zh-cn and choose one of the Chinese sample voices listed by /samples, or upload a custom reference recording.

POST /synthesize-batch

Batch synthesize multiple texts (separated by newlines)

Usage

  1. 1.Using Sample Voices:
  2. 2.Select a sample voice from the available options
  3. 3.Enter text in your chosen language
  4. 4.Click "Synthesize"
  1. 1.Using Custom Voice:
  2. 2.Upload your audio file (min 2 seconds, any format)
  3. 3.Enter text
  4. 4.Click "Synthesize"
  1. 1.API Usage:
bash
   curl -X POST "http://localhost:7860/synthesize" \
     -F "text=Hello world" \
     -F "language=en" \
     -F "sample_audio=arnold_schwarzenegger.wav"

Model Information

  • โ€”XTTS v2: Multilingual TTS model supporting voice cloning
  • โ€”Automatic Sentence Splitting: Disabled for natural long-text synthesis
  • โ€”Audio Preprocessing: Gentle silence trimming (top_db=50) to preserve voice characteristics
  • โ€”Device: Automatically detects and uses CUDA GPU if available, falls back to CPU

Installation (Local Development)

bash
# Create environment
conda create -n voice python=3.11
conda activate voice

# Install FFmpeg
conda install ffmpeg -y

# Install dependencies
pip install -r requirements.txt

# Run server
python -m uvicorn app:app --host 127.0.0.1 --port 8000 --reload

Environment Variables

  • โ€”PORT: Server port (default: 7860 for HuggingFace Spaces, 8000 for local)

Performance Notes

  • โ€”First synthesis takes ~15-20 seconds (model initialization)
  • โ€”Subsequent syntheses are faster (~10-13 seconds depending on text length)
  • โ€”Longer texts may take additional time
  • โ€”GPU recommended for better performance

Deployment on HuggingFace Spaces

  1. 1.Fork this repository to your HuggingFace account
  2. 2.Create new Space with Docker runtime
  3. 3.Set environment: PORT=7860
  4. 4.Space will automatically build and deploy

License

This project uses:

  • โ€”XTTS v2 by Coqui AI (Apache 2.0)
  • โ€”FastAPI (MIT)
  • โ€”PyTorch (BSD)

Note: This application requires significant computational resources. The XTTS v2 model is ~1.87GB and benefits from GPU acceleration.

Check out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference