jegan249895/Multi_Document_RAG_System
0
๐ Multi-Document RAG Chat
A conversational Retrieval-Augmented Generation (RAG) assistant that lets you upload one or more PDFs and ask questions across them โ with hybrid search, cited sources, and multi-turn memory.
Features
- Multi-document upload โ drop in any number of PDFs at once; each session gets its own isolated index.
- Hybrid search โ combines BM25 keyword search with dense vector search (
EnsembleRetriever) so both exact terms (figures, tickers, names) and semantic matches are retrieved. - Cross-encoder reranking โ the top hybrid results are re-scored with
cross-encoder/ms-marco-MiniLM-L-6-v2for relevance before being passed to the LLM. - Inline citations โ every answer cites the source file and page number (
[1],[2], ...), with a source list appended. - Conversational memory โ follow-up questions are rewritten into standalone queries using recent chat history, so you can ask things like "what about Q3?" after an initial question.
Tech Stack
Usage
- Upload one or more PDF files.
- Click Index Documents and wait for the status message to confirm chunking + indexing.
- Ask a question in the chat box. Follow-up questions can reference earlier turns.
- Click Clear Session to reset and start over with new documents.
API Key Setup
The LLM (openai/gpt-oss-120b) is served via Groq, not run locally. Get an API key from the Groq console and set it as GROQ_API_KEY:
- Locally:
export GROQ_API_KEY=your_key_here - On Hugging Face Spaces: add it under Space Settings โ Repository secrets as
GROQ_API_KEY(never commit it to the repo).
Running Locally
pip install -r requirements.txt
export GROQ_API_KEY=your_key_here
python app.pyThe app launches a local Gradio server (default http://127.0.0.1:7860).
Deploying on Hugging Face Spaces
- Create a new Space with the Gradio SDK.
- Push
app.py,requirements.txt, and thisREADME.mdto the Space repo. - Add
GROQ_API_KEYas a repository secret (Settings โ Repository secrets). - The frontmatter above configures the Space automatically โ no extra settings needed.
Limitations
- Sessions are in-memory; restarting the Space clears all indexed documents.
- Concurrent users each get a separate Chroma directory, but the embedding and reranker models are shared/single-instance โ heavy concurrent load will queue for those steps (LLM calls themselves are handled by Groq, off-Space).
- Requires a Groq API key and incurs Groq's per-token usage cost; check current pricing on the Groq console before heavy use.
