Team Ai
Apppublic

Abdul-Rehman-99/Browser-Voice_Agent-Backend

sourceHugging Faceupdated 3mo agoView on Hugging Face
1likes
App README

Browser Voice Agent Backend

Real-time voice AI backend with STT (Whisper/Groq), agent (Llama 3.3 / OpenAI Agents SDK), and TTS (Edge-TTS).

WebSocket Endpoint

wss://<host>/ws/voice

Message Protocol

Text frames (JSON control messages):

DirectionTypeDescription
Client → Serverstart_sessionBegin conversation
Server → Clientsession_startedSession confirmed
Client → Serveraudio_end + mime_typeMark audio complete
Server → ClientprocessingAgent is thinking
Server → Clientresponse_startTTS audio begins
Server → Clientresponse_endTTS audio ends
Server → ClientheartbeatKeepalive (respond with pong)
Server → Clientturn_completeFull turn transcript + response

Binary frames: Raw audio chunks (bidirectional).

Environment Variables

VariableDefaultDescription
GROQ_API_KEY—Groq API key (required)
GROQ_BASE_URLhttps://api.groq.com/openai/v1Groq API base URL
WS_HOST0.0.0.0Bind address
WS_PORT8000Listen port
AGENT_MODELllama-3.3-70b-versatileGroq model for agent
STT_MODELwhisper-large-v3-turboGroq model for STT
STT_LANGUAGE(auto)Force STT language (en/ur/empty)
TTS_VOICEur-PK-UzmaNeuralEdge-TTS voice
MAX_TURNS15Conversation history turns

Local Development

bash
uv sync
uv run uvicorn app.main:app --reload

Deploy

Build and run with Docker:

bash
docker build -t browser-voice-agent-backend .
docker run -p 7860:7860 -e GROQ_API_KEY=your_key browser-voice-agent-backend