WilliamCass/voice-synthesis
0
Voice Synthesis with Voice Cloning
AI-powered Text-to-Speech application with multilingual voice cloning using XTTS v2.
Features
- ๐ค Voice Cloning: Upload any audio file or use sample voices in English, German, and Chinese
- ๐ Multilingual: Synthesizes English, German, and Simplified Chinese (
zh-cn) - โก Real-time Synthesis: Fast speech generation
- ๐จ Modern Web Interface: Responsive HTML/JS frontend
- ๐ Multiple Audio Formats: WAV, MP3, M4A, FLAC, OGG support
Technical Details
- TTS Model: XTTS v2 (Coqui AI) - Multilingual zero-shot voice cloning
- Backend: FastAPI 0.104.1 + Uvicorn
- Audio Processing: librosa + FFmpeg + soundfile
- Deep Learning: PyTorch 2.10+ with CUDA support
API Endpoints
GET /health
Check if TTS model is ready
GET /languages
Get supported synthesis languages (returns {en, de, zh-cn})
GET /samples
Get available sample audio files
POST /synthesize
Synthesize speech with voice cloning
Parameters:
text(string, required): Text to synthesizelanguage(string): Language code (en,de, orzh-cn, default:en)reference_audio(file, optional): Upload custom audio for voice cloningsample_audio(string, optional): Use a sample audio filename returned by/samples
Response: WAV audio file
For Chinese synthesis, select zh-cn and choose one of the Chinese sample voices listed by /samples, or upload a custom reference recording.
POST /synthesize-batch
Batch synthesize multiple texts (separated by newlines)
Usage
- Using Sample Voices:
- Select a sample voice from the available options
- Enter text in your chosen language
- Click "Synthesize"
- Using Custom Voice:
- Upload your audio file (min 2 seconds, any format)
- Enter text
- Click "Synthesize"
- API Usage:
curl -X POST "http://localhost:7860/synthesize" \
-F "text=Hello world" \
-F "language=en" \
-F "sample_audio=arnold_schwarzenegger.wav"Model Information
- XTTS v2: Multilingual TTS model supporting voice cloning
- Automatic Sentence Splitting: Disabled for natural long-text synthesis
- Audio Preprocessing: Gentle silence trimming (top_db=50) to preserve voice characteristics
- Device: Automatically detects and uses CUDA GPU if available, falls back to CPU
Installation (Local Development)
# Create environment
conda create -n voice python=3.11
conda activate voice
# Install FFmpeg
conda install ffmpeg -y
# Install dependencies
pip install -r requirements.txt
# Run server
python -m uvicorn app:app --host 127.0.0.1 --port 8000 --reloadEnvironment Variables
PORT: Server port (default: 7860 for HuggingFace Spaces, 8000 for local)
Performance Notes
- First synthesis takes ~15-20 seconds (model initialization)
- Subsequent syntheses are faster (~10-13 seconds depending on text length)
- Longer texts may take additional time
- GPU recommended for better performance
Deployment on HuggingFace Spaces
- Fork this repository to your HuggingFace account
- Create new Space with Docker runtime
- Set environment:
PORT=7860 - Space will automatically build and deploy
License
This project uses:
- XTTS v2 by Coqui AI (Apache 2.0)
- FastAPI (MIT)
- PyTorch (BSD)
Note: This application requires significant computational resources. The XTTS v2 model is ~1.87GB and benefits from GPU acceleration.
Check out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
