Team Ai
Apppublic

sinhapiyush86/convAI

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes
App README

πŸ€– Conversational AI RAG System

A comprehensive Retrieval-Augmented Generation (RAG) system with advanced guard rails, built with Streamlit, FAISS, and Hugging Face models.

πŸš€ Features

  • β€”Hybrid Search: Combines dense (FAISS) and sparse (BM25) retrieval for optimal results
  • β€”Advanced Guard Rails: Comprehensive safety and security measures
  • β€”Multiple Models: Support for Qwen 2.5 1.5B and distilgpt2 fallback
  • β€”PDF Processing: Intelligent document chunking and processing
  • β€”Real-time Monitoring: Performance metrics and system health checks
  • β€”Docker Support: Containerized deployment with Docker Compose
  • β€”Hugging Face Spaces Ready: Optimized for HF Spaces deployment

πŸ—οΈ Architecture

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚   Streamlit UI  │───▢│   RAG System    │───▢│  Guard Rails    β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                              β”‚
                              β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚  PDF Processor  β”‚    β”‚   FAISS Index   β”‚    β”‚  Language Model β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

πŸ› οΈ Technology Stack

Core Technologies

  • β€”πŸ” Vector Database: FAISS for efficient similarity search
  • β€”πŸ“ Sparse Retrieval: BM25 for keyword-based search
  • β€”πŸ§  Embedding Model: all-MiniLM-L6-v2 for document embeddings
  • β€”πŸ€– Generative Model: Qwen 2.5 1.5B for answer generation
  • β€”πŸŒ UI Framework: Streamlit for interactive interface
  • β€”πŸ³ Containerization: Docker for deployment

Supporting Libraries

  • β€”πŸ“Š Data Processing: Pandas, NumPy for data manipulation
  • β€”πŸ“„ PDF Handling: PyPDF for document processing
  • β€”πŸ”§ ML Utilities: Scikit-learn for preprocessing
  • β€”πŸ“ Logging: Loguru for structured logging
  • β€”βš‘ Optimization: Accelerate for model optimization

πŸš€ Quick Start

Local Development

  1. 1.Clone and Setup:
bash
git clone <repository-url>
cd convAI
pip install -r requirements.txt
  1. 1.Run the Application:
bash
streamlit run app.py
  1. 1.Upload PDFs and Start Chatting!

Docker Deployment

  1. 1.Build and Run:
bash
docker-compose up --build
  1. 1.Access at: http://localhost:8501

🌟 Hugging Face Spaces Deployment

This application is optimized for deployment on Hugging Face Spaces. The system automatically:

  • β€”Uses /tmp directories for cache storage (writable in HF Spaces)
  • β€”Configures environment variables for HF Spaces compatibility
  • β€”Handles permission issues automatically
  • β€”Optimizes model loading for HF Spaces environment

HF Spaces Configuration

The application includes:

  • β€”Cache Management: All model caches stored in /tmp directories
  • β€”Permission Handling: Automatic fallback to writable directories
  • β€”Environment Detection: Adapts to HF Spaces runtime environment
  • β€”Resource Optimization: Efficient memory and CPU usage

Deploy to HF Spaces

  1. 1.Create a new Space on Hugging Face
  2. 2.Choose Docker as the SDK
  3. 3.Upload all files from this repository
  4. 4.The system will automatically:
  5. 5.Set up cache directories in /tmp
  6. 6.Download and cache models
  7. 7.Initialize the RAG system with guard rails
  8. 8.Start the Streamlit interface

HF Spaces Environment Variables

The system automatically configures:

bash
HF_HOME=/tmp/huggingface
TRANSFORMERS_CACHE=/tmp/huggingface/transformers
TORCH_HOME=/tmp/torch
XDG_CACHE_HOME=/tmp
HF_HUB_CACHE=/tmp/huggingface/hub

πŸ“– Usage Guide

Document Upload

  • β€”Automatic Loading: PDF documents in the container are loaded automatically
  • β€”Manual Upload: Use the sidebar to upload additional PDF documents
  • β€”Supported Formats: PDF files with text content

Search Methods

  • β€”πŸ”€ Hybrid: Combines vector similarity and keyword matching (recommended)
  • β€”πŸŽ― Dense: Uses only vector similarity search
  • β€”πŸ“ Sparse: Uses only keyword-based BM25 search

Query Interface

  • β€”Natural Language: Ask questions in plain English
  • β€”Context Awareness: System uses retrieved documents for context
  • β€”Confidence Scores: See how confident the system is in its answers
  • β€”Source Citations: View which documents were used for the answer

βš™οΈ Configuration

Environment Variables

bash
# Model Configuration
EMBEDDING_MODEL=all-MiniLM-L6-v2
GENERATIVE_MODEL=Qwen/Qwen2.5-1.5B-Instruct

# Chunk Sizes
CHUNK_SIZES=100,400

# Vector Store Path
VECTOR_STORE_PATH=./vector_store

# Streamlit Configuration
STREAMLIT_SERVER_PORT=8501
STREAMLIT_SERVER_ADDRESS=0.0.0.0

Performance Tuning

  • β€”Chunk Sizes: Adjust for different document types (smaller for technical docs, larger for narratives)
  • β€”Top-k Results: Increase for more comprehensive answers, decrease for faster responses
  • β€”Model Selection: Choose between Qwen 2.5 1.5B and distilgpt2 based on performance needs

πŸ“Š Performance

Optimization Features

  • β€”Parallel Processing: Documents are loaded concurrently for faster initialization
  • β€”Optimized Search: Hybrid retrieval combines the best of vector and keyword search
  • β€”Memory Efficient: Uses CPU-optimized models for deployment compatibility
  • β€”Caching: FAISS index and metadata are cached for faster subsequent queries

Expected Performance

  • β€”Document Loading: ~2-5 seconds per PDF (depending on size)
  • β€”Query Response: ~1-3 seconds for typical questions
  • β€”Memory Usage: ~2-4GB RAM for typical document collections
  • β€”Storage: ~100MB per 1000 document chunks

πŸ”§ Development

Project Structure

convAI/
β”œβ”€β”€ app.py                 # Main Streamlit application
β”œβ”€β”€ rag_system.py          # Core RAG system implementation
β”œβ”€β”€ pdf_processor.py       # PDF processing utilities
β”œβ”€β”€ requirements.txt       # Python dependencies
β”œβ”€β”€ Dockerfile            # Container configuration
β”œβ”€β”€ docker-compose.yml    # Multi-container setup
β”œβ”€β”€ README.md             # This file
β”œβ”€β”€ DEPLOYMENT_GUIDE.md   # Detailed deployment instructions
β”œβ”€β”€ test_deployment.py    # Deployment testing script
β”œβ”€β”€ test_docker.py        # Docker testing script
└── src/
    └── streamlit_app.py  # Sample Streamlit app

Testing

bash
# Test deployment readiness
python test_deployment.py

# Test Docker configuration
python test_docker.py

# Run local tests
streamlit run app.py

πŸ› Troubleshooting

Common Issues

  1. 1.Model Loading Errors
  2. 2.Check internet connectivity for model downloads
  3. 3.Verify sufficient disk space
  4. 4.Try the fallback model (distilgpt2)
  1. 1.Memory Issues
  2. 2.Reduce chunk sizes
  3. 3.Use smaller embedding models
  4. 4.Limit the number of documents
  1. 1.Performance Issues
  2. 2.Adjust top-k parameter
  3. 3.Use sparse search for keyword-heavy queries
  4. 4.Consider hardware upgrades
  1. 1.Docker Issues
  2. 2.Check Docker installation
  3. 3.Verify port availability
  4. 4.Check container logs

Getting Help

  • β€”Check the logs in your Space's "Logs" tab
  • β€”Review the deployment guide for common solutions
  • β€”Create an issue in the project repository

🀝 Contributing

We welcome contributions! Please see our contributing guidelines for:

  • β€”Code style and standards
  • β€”Testing requirements
  • β€”Documentation updates
  • β€”Feature requests and bug reports

πŸ“„ License

This project is licensed under the MIT License - see the LICENSE file for details.

πŸ™ Acknowledgments

  • β€”Hugging Face for providing the platform and models
  • β€”FAISS team for the efficient vector search library
  • β€”Streamlit team for the excellent web framework
  • β€”OpenAI for inspiring the RAG architecture

Built with ❀️ for efficient document question-answering

Ready to explore your documents? Start asking questions! πŸš€