Team Ai
Apppublic

kazakiakayami/ai-document-assistant

sourceHugging Faceupdated 7mo agoView on Hugging Face
0likes
App README

๐Ÿ“„ AI Document Assistant

An intelligent document assistant powered by RAG (Retrieval-Augmented Generation) that allows you to upload PDF documents and ask questions about their content in natural language.

Python LangChain Streamlit Groq

๐Ÿš€ Live Demo

๐Ÿค— Try it on HuggingFace Spaces


โœจ Features

  • โ€”๐Ÿ“ Upload any PDF document
  • โ€”๐Ÿ’ฌ Chat with your document in natural language
  • โ€”๐Ÿ” Accurate answers based strictly on document content
  • โ€”๐Ÿ“„ Source page references for every answer
  • โ€”๐ŸŒ Supports both English and Indonesian questions
  • โ€”๐Ÿ”„ Switch documents anytime without refreshing

๐Ÿ› ๏ธ Tech Stack

ComponentTechnology
FrameworkLangChain
LLMGroq (Llama 3.1 8B) - Free
EmbeddingHuggingFace all-MiniLM-L6-v2
Vector DBChromaDB
UIStreamlit
PDF ParsingPyPDF

๐Ÿ”„ RAG Pipeline

๐Ÿ“„ PDF Upload
     โ†“
๐Ÿ“ฅ Document Loader (PyPDF)
     โ†“
โœ‚๏ธ Text Chunking (chunk_size=1000, overlap=100)
     โ†“
๐Ÿ”ข Embedding (HuggingFace sentence-transformers)
     โ†“
๐Ÿ—„๏ธ Vector Store (ChromaDB)
     โ†“
โ“ User Query โ†’ Similarity Search โ†’ Top 3 Chunks
     โ†“
๐Ÿค– LLM (Groq/Llama3.1) โ†’ Answer + Sources

๐Ÿ“ Project Structure

ai-document-assistant/
โ”‚
โ”œโ”€โ”€ app.py                  # Main Streamlit UI
โ”œโ”€โ”€ requirements.txt        # Dependencies
โ”œโ”€โ”€ .env                    # API Keys (not committed)
โ”œโ”€โ”€ .gitignore
โ”‚
โ””โ”€โ”€ src/
    โ”œโ”€โ”€ document_loader.py  # PDF loading
    โ”œโ”€โ”€ chunker.py          # Text chunking
    โ”œโ”€โ”€ embedder.py         # Embedding + VectorDB
    โ””โ”€โ”€ rag_chain.py        # RAG pipeline + LLM

โš™๏ธ Installation & Setup

1. Clone the repository

bash
git clone https://github.com/YOUR_USERNAME/ai-document-assistant.git
cd ai-document-assistant

2. Create virtual environment

bash
conda create -n rag_env python=3.11
conda activate rag_env

3. Install dependencies

bash
pip install -r requirements.txt

4. Setup environment variables

Create a .env file in the root directory:

GROQ_API_KEY=your_groq_api_key_here

Get your free Groq API key at groq.com

5. Run the app

bash
streamlit run app.py

๐Ÿงช Testing Results

Test TypeQuestionResult
โœ… Positive"What is the offside rule?"Accurate answer with sources
โœ… Positive"How many players in a football team?""Maximum 11 players"
โœ… Negative"What is the NBA rules?""Cannot find answer in document"
โœ… Edge Case"Tell me everything about football"General summary from context
โœ… Boundary"List all offences in football"Partial list with honest disclaimer

๐Ÿ”‘ Key Design Decisions

  • โ€”Chunk size 1000, overlap 100 โ†’ optimal for legal/rule documents
  • โ€”k=3 retrieval โ†’ balance between context and token efficiency
  • โ€”Temperature 0.2 โ†’ low temperature for factual accuracy
  • โ€”In-memory VectorDB โ†’ no persistence needed for demo deployment
  • โ€”Prompt engineering โ†’ strict instruction to not hallucinate beyond context

๐Ÿ“ Lessons Learned

  • โ€”RAG accuracy heavily depends on chunk size and overlap configuration
  • โ€”PDF parsing can introduce artifacts (headers, footers) that affect chunk quality
  • โ€”Filtering chunks shorter than 50 characters significantly improves retrieval quality
  • โ€”LLM temperature should be low (0.1-0.3) for document Q&A tasks

๐Ÿ‘จโ€๐Ÿ’ป Author

Ahmad Mustofa Z Hacktiv8 Data Science Bootcamp - Phase 1 Project