Team Ai
Apppublic

jegan249895/Multi_Document_RAG_System

sourceHugging Facemitupdated 1mo agoView on Hugging Face
0likes
App README

๐Ÿ“š Multi-Document RAG Chat

A conversational Retrieval-Augmented Generation (RAG) assistant that lets you upload one or more PDFs and ask questions across them โ€” with hybrid search, cited sources, and multi-turn memory.

Features

  • โ€”Multi-document upload โ€” drop in any number of PDFs at once; each session gets its own isolated index.
  • โ€”Hybrid search โ€” combines BM25 keyword search with dense vector search (EnsembleRetriever) so both exact terms (figures, tickers, names) and semantic matches are retrieved.
  • โ€”Cross-encoder reranking โ€” the top hybrid results are re-scored with cross-encoder/ms-marco-MiniLM-L-6-v2 for relevance before being passed to the LLM.
  • โ€”Inline citations โ€” every answer cites the source file and page number ([1], [2], ...), with a source list appended.
  • โ€”Conversational memory โ€” follow-up questions are rewritten into standalone queries using recent chat history, so you can ask things like "what about Q3?" after an initial question.

Tech Stack

ComponentTool
UIGradio (Blocks)
Document loadingPyPDFLoader (LangChain)
ChunkingRecursiveCharacterTextSplitter
EmbeddingsBAAI/bge-small-en-v1.5
Vector storeChromaDB
Keyword searchBM25 (rank_bm25)
Hybrid retrievalLangChain EnsembleRetriever
Rerankercross-encoder/ms-marco-MiniLM-L-6-v2
LLMopenai/gpt-oss-120b via Groq API

Usage

  1. 1.Upload one or more PDF files.
  2. 2.Click Index Documents and wait for the status message to confirm chunking + indexing.
  3. 3.Ask a question in the chat box. Follow-up questions can reference earlier turns.
  4. 4.Click Clear Session to reset and start over with new documents.

API Key Setup

The LLM (openai/gpt-oss-120b) is served via Groq, not run locally. Get an API key from the Groq console and set it as GROQ_API_KEY:

  • โ€”Locally: export GROQ_API_KEY=your_key_here
  • โ€”On Hugging Face Spaces: add it under Space Settings โ†’ Repository secrets as GROQ_API_KEY (never commit it to the repo).

Running Locally

bash
pip install -r requirements.txt
export GROQ_API_KEY=your_key_here
python app.py

The app launches a local Gradio server (default http://127.0.0.1:7860).

Deploying on Hugging Face Spaces

  1. 1.Create a new Space with the Gradio SDK.
  2. 2.Push app.py, requirements.txt, and this README.md to the Space repo.
  3. 3.Add GROQ_API_KEY as a repository secret (Settings โ†’ Repository secrets).
  4. 4.The frontmatter above configures the Space automatically โ€” no extra settings needed.

Limitations

  • โ€”Sessions are in-memory; restarting the Space clears all indexed documents.
  • โ€”Concurrent users each get a separate Chroma directory, but the embedding and reranker models are shared/single-instance โ€” heavy concurrent load will queue for those steps (LLM calls themselves are handled by Groq, off-Space).
  • โ€”Requires a Groq API key and incurs Groq's per-token usage cost; check current pricing on the Groq console before heavy use.