Inzeera/Multi_Document_Summarizer_Plus_QA_Bot
0
Features
- Upload PDF, DOCX, XLS, and XLSX files
- Automatic document parsing
- Intelligent text chunking
- Semantic search using embeddings
- Retrieval-Augmented Generation (RAG)
- AI-powered document summarization
- Question Answering with source citations
- Download summaries as TXT or PDF
- Modern glassmorphism-based user interface
- Supports multiple document uploads
Architecture
User Uploads Documents
│
▼
Document Loaders
(PDF / DOCX / XLSX)
│
▼
Text Chunking
│
▼
HuggingFace Embeddings
│
▼
Chroma Vector Database
│
▼
RetrievalQA Chain
│
▼
Groq LLM
│
▼
Summary / AnswerProject Structure
document_qa_bot/
│
├── app.py
│
├── config/
│ └── settings.py
│
├── loaders/
│ ├── pdf_loader.py
│ ├── docx_loader.py
│ ├── excel_loader.py
│ └── document_loader.py
│
├── rag/
│ ├── embeddings.py
│ ├── llm.py
│ ├── vectordb.py
│ ├── qa_chain.py
│ └── summarizer.py
│
├── utils/
│ ├── hash_utils.py
│ └── pdf_export.py
│
├── temp_docs/
│
├── requirements.txt
│
└── README.mdTechnologies Used
Frontend
- Streamlit
AI Framework
- LangChain
Vector Database
- ChromaDB
Embedding Model
- sentence-transformers/all-MiniLM-L6-v2
Large Language Model
- Groq API
- openai/gpt-oss-120b
File Processing
- PyPDF2
- python-docx
- pandas
PDF Generation
- ReportLab
Installation
Clone Repository
git clone https://github.com/your-username/document-qa-bot.git
cd document-qa-botCreate Virtual Environment
python -m venv venvActivate:
Windows
venv\Scripts\activateLinux / Mac
source venv/bin/activateInstall Dependencies
pip install -r requirements.txtEnvironment Variables
Create a .env file in the project root.
GROQ_API_KEY=your_groq_api_key_hereRunning the Application
streamlit run app.pyThe application will open in your browser.
Workflow
Document Upload
Users upload one or more files:
- DOCX
- XLS
- XLSX
Document Processing
The system:
- Extracts text
- Creates LangChain documents
- Splits text into chunks
- Generates embeddings
- Stores vectors in ChromaDB
Summarization
The uploaded document content is summarized using a Groq-hosted Large Language Model.
Question Answering
User questions are:
- Converted into embeddings
- Matched against document embeddings
- Relevant chunks are retrieved
- Groq generates answers based on retrieved content
- Source citations are displayed
Example Questions
- What is the main objective of the document?
- What conclusions were reached?
- Summarize the methodology section.
- What are the key findings?
- List the recommendations mentioned.
RAG Pipeline
Document
│
▼
Chunking
│
▼
Embeddings
│
▼
Vector Database
│
▼
User Query
│
▼
Query Embedding
│
▼
Similarity Search
│
▼
Relevant Chunks
│
▼
LLM
│
▼
Answer + CitationsPerformance Optimizations
- Cached embedding model
- Cached LLM initialization
- Session state management
- Chunk-based retrieval
- Efficient vector search
Future Enhancements
- OCR support for scanned PDFs
- Conversation memory
- Multi-document comparison
- Hybrid search (BM25 + Vector Search)
- Citation highlighting
- User authentication
- Cloud deployment
- Database-backed document storage
License
This project is intended for educational and portfolio purposes.
Author
Developed by Inzeera Z
