alex-codecub/voice-ai-assistant
Voice AI Assistant - Streamlit Chatbot
A voice-enabled AI chatbot application built with Streamlit and Google Gemini API, featuring speech-to-text input, text-to-speech output, and multiple AI personalities.
Features
- ๐ค AI-Powered Chat: Powered by Google Gemini 2.5 Flash model
- ๐ค Voice Input: Record your voice and get automatic speech-to-text conversion
- ๐ Voice Output: Text-to-speech (TTS) for AI responses with automatic audio generation
- ๐ญ Multiple Personalities: Choose from 4 different AI personalities:
- General Assistant - Versatile helper for various topics
- Study Buddy - Patient learning companion
- Fitness Coach - Motivational fitness and wellness guide
- Gaming Helper - Knowledgeable gaming companion
- ๐ฌ Chat History: Maintains conversation context throughout the session
- โจ๏ธ Dual Input: Support for both voice and text input
- ๐จ Clean UI: User-friendly interface with polished layout and helpful feedback
Prerequisites
- Python 3.8 or higher
- Google Gemini API key (get yours at Google AI Studio)
- Microphone (for voice input feature)
- Internet connection (for speech recognition and AI responses)
Installation
- Clone the repository
git clone https://github.com/CodeCubCA/voice-ai-assistant-Alex-CodeCub.git
cd voice-ai-assistant-Alex-CodeCub- Install dependencies
pip install -r requirements.txt- Set up environment variables
Create a .env file in the project root:
cp .env.example .env Edit .env and add your Gemini API key:
GEMINI_API_KEY=your_api_key_hereUsage
- Start the application
streamlit run app.py- Access the application
Open your browser and navigate to:
- Local URL:
http://localhost:8501 - Network URL:
http://your-ip:8501
- Using Voice Input
- Click the microphone button to start recording
- Speak clearly into your microphone
- Click the button again to stop recording
- The audio will be automatically transcribed and sent to the AI
- Using Text Input
- Type your message in the text input box at the bottom
- Press Enter to send
- Listening to AI Responses
- Audio is automatically generated for each AI response
- Audio player appears below each assistant message with a divider
- Use browser controls to play, pause, adjust speed, or control volume
- Long messages may take a moment to generate audio
- Switching Personalities
- Use the sidebar to select different AI personalities
- Chat history will be cleared when switching personalities
Project Structure
voice-ai-assistant/
โโโ app.py # Main application file
โโโ requirements.txt # Python dependencies
โโโ .env.example # Environment variables template
โโโ .env # Your API keys (git-ignored)
โโโ .gitignore # Git ignore configuration
โโโ README.md # This fileDependencies
streamlit>=1.31.0- Web framework for the UIgoogle-generativeai>=0.3.2- Google Gemini API clientpython-dotenv>=1.0.0- Environment variable managementaudio-recorder-streamlit>=0.0.8- Audio recording componentSpeechRecognition>=3.10.0- Speech-to-text conversiongtts>=2.3.0- Google Text-to-Speech for audio output
Technical Details
Voice Recognition Implementation
The voice input feature uses a two-step process to ensure reliable speech recognition:
- Audio Recording: Uses
audio-recorder-streamlitto capture audio from the microphone - Speech-to-Text: Converts audio to text using Google Speech Recognition API
- Audio is saved to a temporary WAV file
sr.AudioFile()is used to properly handle audio format- This approach ensures correct sample rate and audio format parsing
Important Note: The implementation uses sr.AudioFile() instead of sr.AudioData() constructor to avoid format compatibility issues.
Text-to-Speech Implementation
The TTS feature provides automatic audio generation for AI responses:
- Automatic Generation: Every AI response is automatically converted to speech
- Audio Playback: Audio players are displayed below each AI message with playback controls
- Smart Features:
- Warning alerts for long messages that may take time to process
- Automatic truncation for extremely long messages (>1000 characters)
- Browser-native playback controls (play, pause, speed adjustment, volume)
- Graceful error handling - chat continues even if audio generation fails
- Implementation Details:
- Uses Google Text-to-Speech (gTTS) library
- Audio saved as temporary MP3 files
- Audio players rendered outside chat message containers for Streamlit compatibility
- Anti-loop protections prevent duplicate audio generation
AI Model
- Model: Google Gemini 2.5 Flash (
gemini-2.5-flash) - Context: Maintains conversation history with system prompts
- Personalities: Each personality has a custom system prompt that influences AI behavior
Configuration
Customizing AI Personalities
Edit the PERSONALITIES dictionary in app.py to add or modify personalities:
PERSONALITIES = {
"Your Personality Name": {
"name": "Display Name",
"icon": "๐ฏ",
"system_prompt": "Your custom system prompt here...",
"description": "Brief description"
}
}Adjusting UI Colors
Modify the audio_recorder parameters in app.py:
audio_bytes = audio_recorder(
recording_color="#e74c3c", # Color when recording
neutral_color="#3498db", # Color when idle
icon_name="microphone",
icon_size="2x",
)Troubleshooting
Voice Recognition Issues
Problem: "Could not understand audio" error
Solutions:
- Speak more clearly and at a moderate pace
- Reduce background noise
- Check your microphone is working properly
- Ensure you have a stable internet connection
Problem: Audio processing fails
Solutions:
- Verify the
SpeechRecognitionlibrary is properly installed - Check that temporary files can be created in your system
- Ensure Google Speech Recognition API is accessible
Text-to-Speech Issues
Problem: Audio not generating for AI responses
Solutions:
- Check that
gttslibrary is properly installed - Verify internet connection (gTTS requires online access)
- Check browser console for any errors
- Ensure temporary files can be created in your system
Problem: Audio playback not working
Solutions:
- Try a different browser (Chrome, Firefox, Edge recommended)
- Check browser audio/autoplay settings
- Verify system volume is not muted
- Try refreshing the page
API Issues
Problem: API key errors
Solutions:
- Verify your
.envfile exists and contains the correct API key - Check that your Gemini API key is valid
- Ensure you haven't exceeded API rate limits
Problem: Gemini model errors
Solutions:
- Confirm you're using the correct model name:
gemini-2.5-flash - Check your internet connection
- Verify your API key has access to the Gemini API
Security Notes
- Never commit your
.envfile to version control - Keep your API keys secure and private
- The
.gitignorefile is configured to exclude.envautomatically - Use
.env.exampleas a template for other users
Contributing
Feel free to submit issues, fork the repository, and create pull requests for any improvements.
License
This project is for educational purposes.
Acknowledgments
- Built with Streamlit
- Powered by Google Gemini API
- Speech recognition by Google Speech Recognition
- Text-to-speech by gTTS (Google Text-to-Speech)
- Audio recording component by audio-recorder-streamlit
Author
Created as part of CodeCub's AI development course.
Powered by Google Gemini 2.5 Flash | Built with Streamlit
