Team Ai
Apppublic

alex-codecub/voice-ai-assistant

sourceHugging Faceupdated 11mo agoView on Hugging Face
0likes
App README

Voice AI Assistant - Streamlit Chatbot

A voice-enabled AI chatbot application built with Streamlit and Google Gemini API, featuring speech-to-text input, text-to-speech output, and multiple AI personalities.

Features

  • โ€”๐Ÿค– AI-Powered Chat: Powered by Google Gemini 2.5 Flash model
  • โ€”๐ŸŽค Voice Input: Record your voice and get automatic speech-to-text conversion
  • โ€”๐Ÿ”Š Voice Output: Text-to-speech (TTS) for AI responses with automatic audio generation
  • โ€”๐ŸŽญ Multiple Personalities: Choose from 4 different AI personalities:
  • โ€”General Assistant - Versatile helper for various topics
  • โ€”Study Buddy - Patient learning companion
  • โ€”Fitness Coach - Motivational fitness and wellness guide
  • โ€”Gaming Helper - Knowledgeable gaming companion
  • โ€”๐Ÿ’ฌ Chat History: Maintains conversation context throughout the session
  • โ€”โŒจ๏ธ Dual Input: Support for both voice and text input
  • โ€”๐ŸŽจ Clean UI: User-friendly interface with polished layout and helpful feedback

Prerequisites

  • โ€”Python 3.8 or higher
  • โ€”Google Gemini API key (get yours at Google AI Studio)
  • โ€”Microphone (for voice input feature)
  • โ€”Internet connection (for speech recognition and AI responses)

Installation

  1. 1.Clone the repository
bash
   git clone https://github.com/CodeCubCA/voice-ai-assistant-Alex-CodeCub.git
   cd voice-ai-assistant-Alex-CodeCub
  1. 1.Install dependencies
bash
   pip install -r requirements.txt
  1. 1.Set up environment variables

Create a .env file in the project root:

bash
   cp .env.example .env

Edit .env and add your Gemini API key:

   GEMINI_API_KEY=your_api_key_here

Usage

  1. 1.Start the application
bash
   streamlit run app.py
  1. 1.Access the application

Open your browser and navigate to:

  • โ€”Local URL: http://localhost:8501
  • โ€”Network URL: http://your-ip:8501
  1. 1.Using Voice Input
  2. 2.Click the microphone button to start recording
  3. 3.Speak clearly into your microphone
  4. 4.Click the button again to stop recording
  5. 5.The audio will be automatically transcribed and sent to the AI
  1. 1.Using Text Input
  2. 2.Type your message in the text input box at the bottom
  3. 3.Press Enter to send
  1. 1.Listening to AI Responses
  2. 2.Audio is automatically generated for each AI response
  3. 3.Audio player appears below each assistant message with a divider
  4. 4.Use browser controls to play, pause, adjust speed, or control volume
  5. 5.Long messages may take a moment to generate audio
  1. 1.Switching Personalities
  2. 2.Use the sidebar to select different AI personalities
  3. 3.Chat history will be cleared when switching personalities

Project Structure

voice-ai-assistant/
โ”œโ”€โ”€ app.py                 # Main application file
โ”œโ”€โ”€ requirements.txt       # Python dependencies
โ”œโ”€โ”€ .env.example          # Environment variables template
โ”œโ”€โ”€ .env                  # Your API keys (git-ignored)
โ”œโ”€โ”€ .gitignore           # Git ignore configuration
โ””โ”€โ”€ README.md            # This file

Dependencies

  • โ€”streamlit>=1.31.0 - Web framework for the UI
  • โ€”google-generativeai>=0.3.2 - Google Gemini API client
  • โ€”python-dotenv>=1.0.0 - Environment variable management
  • โ€”audio-recorder-streamlit>=0.0.8 - Audio recording component
  • โ€”SpeechRecognition>=3.10.0 - Speech-to-text conversion
  • โ€”gtts>=2.3.0 - Google Text-to-Speech for audio output

Technical Details

Voice Recognition Implementation

The voice input feature uses a two-step process to ensure reliable speech recognition:

  1. 1.Audio Recording: Uses audio-recorder-streamlit to capture audio from the microphone
  2. 2.Speech-to-Text: Converts audio to text using Google Speech Recognition API
  3. 3.Audio is saved to a temporary WAV file
  4. 4.sr.AudioFile() is used to properly handle audio format
  5. 5.This approach ensures correct sample rate and audio format parsing

Important Note: The implementation uses sr.AudioFile() instead of sr.AudioData() constructor to avoid format compatibility issues.

Text-to-Speech Implementation

The TTS feature provides automatic audio generation for AI responses:

  1. 1.Automatic Generation: Every AI response is automatically converted to speech
  2. 2.Audio Playback: Audio players are displayed below each AI message with playback controls
  3. 3.Smart Features:
  4. 4.Warning alerts for long messages that may take time to process
  5. 5.Automatic truncation for extremely long messages (>1000 characters)
  6. 6.Browser-native playback controls (play, pause, speed adjustment, volume)
  7. 7.Graceful error handling - chat continues even if audio generation fails
  8. 8.Implementation Details:
  9. 9.Uses Google Text-to-Speech (gTTS) library
  10. 10.Audio saved as temporary MP3 files
  11. 11.Audio players rendered outside chat message containers for Streamlit compatibility
  12. 12.Anti-loop protections prevent duplicate audio generation

AI Model

  • โ€”Model: Google Gemini 2.5 Flash (gemini-2.5-flash)
  • โ€”Context: Maintains conversation history with system prompts
  • โ€”Personalities: Each personality has a custom system prompt that influences AI behavior

Configuration

Customizing AI Personalities

Edit the PERSONALITIES dictionary in app.py to add or modify personalities:

python
PERSONALITIES = {
    "Your Personality Name": {
        "name": "Display Name",
        "icon": "๐ŸŽฏ",
        "system_prompt": "Your custom system prompt here...",
        "description": "Brief description"
    }
}

Adjusting UI Colors

Modify the audio_recorder parameters in app.py:

python
audio_bytes = audio_recorder(
    recording_color="#e74c3c",  # Color when recording
    neutral_color="#3498db",    # Color when idle
    icon_name="microphone",
    icon_size="2x",
)

Troubleshooting

Voice Recognition Issues

Problem: "Could not understand audio" error

Solutions:

  • โ€”Speak more clearly and at a moderate pace
  • โ€”Reduce background noise
  • โ€”Check your microphone is working properly
  • โ€”Ensure you have a stable internet connection

Problem: Audio processing fails

Solutions:

  • โ€”Verify the SpeechRecognition library is properly installed
  • โ€”Check that temporary files can be created in your system
  • โ€”Ensure Google Speech Recognition API is accessible

Text-to-Speech Issues

Problem: Audio not generating for AI responses

Solutions:

  • โ€”Check that gtts library is properly installed
  • โ€”Verify internet connection (gTTS requires online access)
  • โ€”Check browser console for any errors
  • โ€”Ensure temporary files can be created in your system

Problem: Audio playback not working

Solutions:

  • โ€”Try a different browser (Chrome, Firefox, Edge recommended)
  • โ€”Check browser audio/autoplay settings
  • โ€”Verify system volume is not muted
  • โ€”Try refreshing the page

API Issues

Problem: API key errors

Solutions:

  • โ€”Verify your .env file exists and contains the correct API key
  • โ€”Check that your Gemini API key is valid
  • โ€”Ensure you haven't exceeded API rate limits

Problem: Gemini model errors

Solutions:

  • โ€”Confirm you're using the correct model name: gemini-2.5-flash
  • โ€”Check your internet connection
  • โ€”Verify your API key has access to the Gemini API

Security Notes

  • โ€”Never commit your .env file to version control
  • โ€”Keep your API keys secure and private
  • โ€”The .gitignore file is configured to exclude .env automatically
  • โ€”Use .env.example as a template for other users

Contributing

Feel free to submit issues, fork the repository, and create pull requests for any improvements.

License

This project is for educational purposes.

Acknowledgments

Author

Created as part of CodeCub's AI development course.


Powered by Google Gemini 2.5 Flash | Built with Streamlit