Team Ai
Apppublic

Inzeera/Multi_Document_Summarizer_Plus_QA_Bot

sourceHugging Faceupdated 20d agoView on Hugging Face
0likes
App README

Features

  • —Upload PDF, DOCX, XLS, and XLSX files
  • —Automatic document parsing
  • —Intelligent text chunking
  • —Semantic search using embeddings
  • —Retrieval-Augmented Generation (RAG)
  • —AI-powered document summarization
  • —Question Answering with source citations
  • —Download summaries as TXT or PDF
  • —Modern glassmorphism-based user interface
  • —Supports multiple document uploads

Architecture

text
User Uploads Documents
          │
          ▼
Document Loaders
(PDF / DOCX / XLSX)
          │
          ▼
Text Chunking
          │
          ▼
HuggingFace Embeddings
          │
          ▼
Chroma Vector Database
          │
          ▼
RetrievalQA Chain
          │
          ▼
Groq LLM
          │
          ▼
Summary / Answer

Project Structure

text
document_qa_bot/
│
├── app.py
│
├── config/
│   └── settings.py
│
├── loaders/
│   ├── pdf_loader.py
│   ├── docx_loader.py
│   ├── excel_loader.py
│   └── document_loader.py
│
├── rag/
│   ├── embeddings.py
│   ├── llm.py
│   ├── vectordb.py
│   ├── qa_chain.py
│   └── summarizer.py
│
├── utils/
│   ├── hash_utils.py
│   └── pdf_export.py
│
├── temp_docs/
│
├── requirements.txt
│
└── README.md

Technologies Used

Frontend

  • —Streamlit

AI Framework

  • —LangChain

Vector Database

  • —ChromaDB

Embedding Model

  • —sentence-transformers/all-MiniLM-L6-v2

Large Language Model

  • —Groq API
  • —openai/gpt-oss-120b

File Processing

  • —PyPDF2
  • —python-docx
  • —pandas

PDF Generation

  • —ReportLab

Installation

Clone Repository

bash
git clone https://github.com/your-username/document-qa-bot.git
cd document-qa-bot

Create Virtual Environment

bash
python -m venv venv

Activate:

Windows

bash
venv\Scripts\activate

Linux / Mac

bash
source venv/bin/activate

Install Dependencies

bash
pip install -r requirements.txt

Environment Variables

Create a .env file in the project root.

env
GROQ_API_KEY=your_groq_api_key_here

Running the Application

bash
streamlit run app.py

The application will open in your browser.


Workflow

Document Upload

Users upload one or more files:

  • —PDF
  • —DOCX
  • —XLS
  • —XLSX

Document Processing

The system:

  1. 1.Extracts text
  2. 2.Creates LangChain documents
  3. 3.Splits text into chunks
  4. 4.Generates embeddings
  5. 5.Stores vectors in ChromaDB

Summarization

The uploaded document content is summarized using a Groq-hosted Large Language Model.

Question Answering

User questions are:

  1. 1.Converted into embeddings
  2. 2.Matched against document embeddings
  3. 3.Relevant chunks are retrieved
  4. 4.Groq generates answers based on retrieved content
  5. 5.Source citations are displayed

Example Questions

  • —What is the main objective of the document?
  • —What conclusions were reached?
  • —Summarize the methodology section.
  • —What are the key findings?
  • —List the recommendations mentioned.

RAG Pipeline

text
Document
   │
   ▼
Chunking
   │
   ▼
Embeddings
   │
   ▼
Vector Database
   │
   ▼
User Query
   │
   ▼
Query Embedding
   │
   ▼
Similarity Search
   │
   ▼
Relevant Chunks
   │
   ▼
LLM
   │
   ▼
Answer + Citations

Performance Optimizations

  • —Cached embedding model
  • —Cached LLM initialization
  • —Session state management
  • —Chunk-based retrieval
  • —Efficient vector search

Future Enhancements

  • —OCR support for scanned PDFs
  • —Conversation memory
  • —Multi-document comparison
  • —Hybrid search (BM25 + Vector Search)
  • —Citation highlighting
  • —User authentication
  • —Cloud deployment
  • —Database-backed document storage

License

This project is intended for educational and portfolio purposes.


Author

Developed by Inzeera Z