Team Ai
Apppublic

Leo-codecub/voice-ai-assistant

sourceHugging Faceupdated 10mo agoView on Hugging Face
1likes
App README

AI Voice Chatbot with Streamlit and Hugging Face

A powerful AI chatbot web application built with Streamlit and Hugging Face Llama 3.2, featuring voice input capabilities, multi-language support, and customizable AI personalities.

๐Ÿš€ Live Deployments

Try the app online without any installation:

Note: The app is deployed on Hugging Face Spaces for easy access. Simply click the link above to start chatting with AI using voice or text!

Features

Core Functionality

  • โ€”AI-Powered Conversations: Leverages Hugging Face Llama 3.2 for intelligent, context-aware responses
  • โ€”Voice Input: Speak to the chatbot using your microphone with real-time speech-to-text transcription
  • โ€”Voice Output (TTS): Multi-language text-to-speech with native speaker voices and automatic playback
  • โ€”Multi-Language Support: Voice recognition in 10 languages including English, Spanish, French, German, Chinese, Japanese, Korean, Italian, Portuguese, and Russian
  • โ€”Voice Commands: Hands-free control with commands like "clear chat", "switch personality", and TTS speed control
  • โ€”Multiple AI Personalities: Choose from 4 distinct personalities tailored for different use cases

AI Personalities

  1. 1.General Assistant - A versatile assistant ready to help with any topic
  2. 2.Study Buddy - Your learning companion for academic success
  3. 3.Fitness Coach - Motivational coach for health and fitness goals
  4. 4.Gaming Helper - Your guide to gaming strategies and tips

User Experience

  • โ€”Visual Feedback: Real-time status indicators for voice processing, transcription, and errors
  • โ€”Error Handling: Comprehensive error messages with retry options for microphone permissions, silent recordings, and connection issues
  • โ€”Responsive Design: Clean, intuitive interface with clear organization
  • โ€”Chat History: Maintains conversation context throughout your session
  • โ€”Quick Response Mode: Streamlined voice conversation workflow for faster interactions

Installation

Prerequisites

  • โ€”Python 3.8 or higher
  • โ€”A Hugging Face API token with write permissions (Get one here)
  • โ€”FFmpeg (required for audio processing)

FFmpeg Installation

macOS:

bash
brew install ffmpeg

Windows: Download from FFmpeg official website and add to PATH

Linux:

bash
sudo apt-get install ffmpeg

Setup Steps

  1. 1.Clone the repository:
bash
git clone https://github.com/CodeCubCA/voice-ai-assistant-Leo-CodeCub.git
cd voice-ai-assistant
  1. 1.Create a virtual environment:
bash
python -m venv venv
source venv/bin/activate  # On Windows: venv\Scripts\activate
  1. 1.Install dependencies:
bash
pip install -r requirements.txt
  1. 1.Set up your API token:
  2. 2.Create a .env file in the project root
  3. 3.Add your Hugging Face API token:
HUGGINGFACE_TOKEN=your_huggingface_token_here

Usage

  1. 1.Start the application:
bash
streamlit run app.py
  1. 1.Access the chatbot:
  2. 2.The app will automatically open in your default browser
  3. 3.If not, navigate to http://localhost:8501
  1. 1.Interact with the chatbot:
  2. 2.Text Input: Type your message in the text box at the bottom
  3. 3.Voice Input: Click the microphone icon, speak clearly, then click again to stop recording
  4. 4.Voice Output: Audio responses play automatically after each AI message
  5. 5.Change Personality: Use the sidebar dropdown or voice command "switch to [personality]"

Voice Commands

Use these voice commands for hands-free control:

Chat Control

  • โ€”"Clear chat" or "clear history" - Clears the conversation history
  • โ€”"Help" or "show commands" - Display all available voice commands

Personality Switching

  • โ€”"Change to study" or "switch to study" - Switches to Study Buddy personality
  • โ€”"Change to fitness" or "switch to fitness" - Switches to Fitness Coach personality
  • โ€”"Change to gaming" or "switch to gaming" - Switches to Gaming Helper personality
  • โ€”"Change to general" or "switch to general" - Switches to General Assistant personality

Audio Control

  • โ€”"Speak faster" or "speed up" - Increase TTS speaking speed
  • โ€”"Speak slower" or "slow down" - Decrease TTS speaking speed
  • โ€”"Normal speed" - Reset to default speaking speed
  • โ€”"Stop talking" - Stop current audio (note: browser limitations apply)

Wake Word Support

All commands can optionally start with "Hey Assistant", "Hey Chatbot", or "OK Assistant"

Supported Languages

Voice input supports the following languages:

  • โ€”English (en-US)
  • โ€”Spanish (es-ES)
  • โ€”French (fr-FR)
  • โ€”German (de-DE)
  • โ€”Chinese/Mandarin (zh-CN)
  • โ€”Japanese (ja-JP)
  • โ€”Korean (ko-KR)
  • โ€”Italian (it-IT)
  • โ€”Portuguese (pt-BR)
  • โ€”Russian (ru-RU)

Select your preferred language from the sidebar dropdown.

Technical Stack

  • โ€”Frontend Framework: Streamlit
  • โ€”AI Model: Google Gemini 2.5 Flash
  • โ€”Speech Recognition: Google Speech Recognition API
  • โ€”Audio Processing: PyDub
  • โ€”Voice Recording: audio-recorder-streamlit
  • โ€”Environment Management: python-dotenv

Project Structure

voice-ai-assistant/
โ”œโ”€โ”€ app.py                 # Main application file
โ”œโ”€โ”€ requirements.txt       # Python dependencies
โ”œโ”€โ”€ .env                  # API key (not committed to Git)
โ”œโ”€โ”€ .env.example          # Template for API key
โ”œโ”€โ”€ .gitignore           # Git ignore rules
โ””โ”€โ”€ README.md            # This file

Tips for Best Voice Recognition

  • โ€”Speak clearly and at a moderate pace
  • โ€”Use a quiet environment to minimize background noise
  • โ€”Keep your device close to the microphone
  • โ€”Click the microphone once to start recording, once to stop
  • โ€”Allow microphone permissions when prompted by your browser

Troubleshooting

Microphone Not Working

  • โ€”Ensure browser has microphone permissions
  • โ€”Check system microphone settings
  • โ€”Try the "Try Again" button after granting permissions

No Speech Detected

  • โ€”Speak louder and more clearly
  • โ€”Check that your microphone is working properly
  • โ€”Ensure you're speaking after clicking the microphone button

Connection Errors

  • โ€”Verify your internet connection
  • โ€”Check that your Gemini API key is valid
  • โ€”Ensure the API key is properly set in the .env file

API Key Issues

  • โ€”Make sure you've created a .env file (not .env.example)
  • โ€”Verify the API key is correct and active
  • โ€”Check that there are no extra spaces in the .env file

Security Notes

  • โ€”The .env file containing your API key is excluded from version control
  • โ€”Never commit your actual API key to the repository
  • โ€”Use .env.example as a template for other users

Contributing

Contributions are welcome! Please feel free to submit a Pull Request.

License

This project is open source and available for educational and personal use.

Acknowledgments


Powered by Google Gemini 2.5 Flash