RajaGopal-ML/Multi_Agent_Document_Intelligence_System
<div align="center">
Multi-Agent Document Intelligence Platform
Ask questions about any PDF — with citations, visual understanding, and automated quality scores
      
Live Demo · API Docs · Report Bug
</div>
What this does
Upload any PDF. The system reads it, understands images and charts inside it, stores the content intelligently, and answers your questions — with proper source citations and quality scores on every answer.
You ask a question
↓
Router Agent → classifies your intent (factual / summary / visual / comparison)
↓
Retrieval Agent → hybrid ChromaDB dense + BM25 sparse search + cross-encoder reranking
↓
Vision Agent → CLIP (ViT-B/32) describes charts, tables, images in the PDF
↓
Synthesis Agent → Groq LLM generates a grounded, cited answer
↓
Evaluation Agent → RAGAS scores every response (faithfulness, relevancy, precision)
↓
Structured JSON → { answer, citations[], agent_trace[], ragas_scores{} }Key features
Tech stack
LLM Inference → Groq (llama-3.1-8b-instant) — ~200 tok/s, free tier
Agent Framework → LangGraph + LangChain
Embeddings → OpenAI text-embedding-3-small
Vector Store → ChromaDB (persistent) + FAISS (fast ANN)
Sparse Retrieval → BM25 (rank-bm25)
Reranker → cross-encoder/ms-marco-MiniLM-L-6-v2
Vision → OpenAI CLIP ViT-B/32
PDF Parsing → pdfplumber + pypdf + Pillow
Evaluation → RAGAS (faithfulness, answer relevancy, context precision)
API → FastAPI + Uvicorn
Frontend → Streamlit
Deployment → Docker + Hugging Face Spaces
CI/CD → GitHub ActionsProject structure
doc-intelligence/
│
├── agents/
│ ├── ingestion_agent.py ← PDF text extraction + CLIP Vision Agent
│ ├── retrieval_agent.py ← Hybrid dense+sparse search + cross-encoder reranker
│ ├── router_agent.py ← Groq LLM query intent classifier
│ ├── synthesis_agent.py ← Groq LLM cited answer generator (structured JSON output)
│ └── orchestrator.py ← LangGraph state machine — wires all 4 agents together
│
├── core/
│ ├── config.py ← All settings via environment variables (.env)
│ ├── models.py ← Pydantic schemas (DocumentChunk, QueryResponse, Citation…)
│ └── vector_store.py ← ChromaDB + FAISS + BM25 hybrid store with RRF fusion
│
├── evaluation/
│ └── ragas_evaluator.py ← RAGAS faithfulness / relevancy / precision scoring
│
├── api/
│ └── main.py ← FastAPI app (Swagger UI at /docs)
│
├── frontend/
│ └── app.py ← Streamlit UI — upload, query, evaluate tabs
│
├── scripts/
│ └── finetune_qlora.py ← Mistral-7B QLoRA fine-tuning pipeline (bonus)
│
├── tests/
│ └── test_api.py ← pytest unit + integration tests
│
├── .github/workflows/ci.yml ← GitHub Actions CI/CD
├── hf_spaces_start.py ← HF Spaces entry point (starts API + frontend together)
├── Dockerfile ← Docker image (port 7860 for HF Spaces)
├── docker-compose.yml ← Local Docker setup
├── requirements.txt
└── .env.example ← Copy to .env and fill in your keysQuickstart
1. Clone the repo
git clone https://github.com/Rajagopalhertzian/doc-intelligence.git
cd doc-intelligence2. Create virtual environment
python -m venv venv
source venv/bin/activate # Mac/Linux
# venv\Scripts\activate # Windows3. Install dependencies
pip install -r requirements.txt4. Set up environment variables
cp .env.example .envOpen .env and fill in:
GROQ_API_KEY=gsk_... # free at console.groq.com
OPENAI_API_KEY=sk-... # for embeddings only — platform.openai.com5. Run the API
uvicorn api.main:app --reload --host 0.0.0.0 --port 8000Open http://localhost:8000/docs — full Swagger UI.
6. Run the Streamlit frontend (new terminal)
source venv/bin/activate
streamlit run frontend/app.pyOpen http://localhost:8501
API endpoints
Example query
curl -X POST http://localhost:8000/query \
-H "Content-Type: application/json" \
-d '{"query": "What are the key findings?", "top_k": 5, "use_reranker": true}'{
"query": "What are the key findings?",
"answer": "The key findings include... [markdown answer]",
"citations": [
{
"chunk_id": "abc-123",
"source_path": "report.pdf",
"content_snippet": "...",
"relevance_score": 0.912
}
],
"agent_trace": ["router_agent", "retrieval_agent", "synthesis_agent", "evaluation_agent"],
"evaluation": {
"faithfulness": 0.91,
"answer_relevancy": 0.88,
"context_precision": 0.85
},
"latency_ms": 487.3
}Run with Docker
docker-compose up --build
# API: http://localhost:8000/docs
# Frontend: http://localhost:8501Deploy to Hugging Face Spaces (free)
pip install huggingface_hub
huggingface-cli login
git remote add hf https://huggingface.co/spaces/YOUR_HF_USERNAME/doc-intelligence
git push hf mainAdd GROQ_API_KEY and OPENAI_API_KEY in your Space's Settings → Secrets.
See DEPLOY_TO_HF.md for the full step-by-step guide.
Run tests
pytest tests/ -vEnvironment variables
Alternative Groq models
LLM_MODEL=llama-3.1-8b-instant # default — fast
LLM_MODEL=llama-3.1-70b-versatile # smarter answers
LLM_MODEL=llama-3.3-70b-versatile # latest llama
LLM_MODEL=mixtral-8x7b-32768 # long contextHow the Vision Agent works
Most RAG systems only handle text. This project adds a Vision Agent using OpenAI CLIP (ViT-B/32) that processes every image, chart, and diagram found inside uploaded PDFs.
When ingesting a PDF:
- pdfplumber extracts text per page
- pypdf extracts embedded XObject images
- CLIP runs zero-shot classification on each image against domain-relevant candidates:
"a bar chart showing data","a flowchart or diagram","a table with rows and columns", etc. - The classification result is stored as a searchable text chunk alongside regular text
- If the query intent is
visual, the router boosts image chunks to the top of retrieval results
This improved RAGAS answer relevancy by ~18% on document-heavy queries compared to text-only retrieval.
Built by
Raja Gopal S — Machine Learning Engineer Bengaluru, India LinkedIn · GitHub
License
MIT — see LICENSE for details.
