Jagadeep24/RAG-System-with-Evaluation-Framework
Agentic SEC RAG Explorer (Production-Grade)
    
An enterprise-ready, low-latency Retrieval-Augmented Generation (RAG) system specialized in analyzing and comparing official SEC 10-K filings. The platform features query decomposition routing, hybrid dense/sparse indexing, cross-encoder reranking, and a native REST batch-embedding engine designed to bypass rate limits under free-tier API quotas.
A beautiful dashboard UI is served locally at http://localhost:8000/.
๐๏ธ System Architecture & Data Flow
flowchart TD
subgraph subgraph_ingestion ["Data Ingestion"]
A["Raw SEC 10-K filings in data/"] --> B("BSHTMLLoader + html.parser")
B --> C["RecursiveCharacterTextSplitter / 15k char window"]
C --> D["HuggingFaceEmbeddings / all-MiniLM-L6-v2"]
D -->|Batch Upsert| E[("(Local Qdrant Dense Index)")]
C --> F[("(In-Memory BM25 Sparse Index)")]
end
subgraph subgraph_routing ["Agentic Query Routing"]
Q["User Query"] --> G("Agent Router / gemini-2.5-flash")
G -->|Decompose & Parallelize| H["Sub-Query 1: Apple metrics"]
G -->|Decompose & Parallelize| I["Sub-Query 2: Microsoft metrics"]
end
subgraph subgraph_retrieval ["Hybrid Retrieval & Reranking"]
H --> J["Parallel Hybrid Retriever"]
I --> J
J -->|Semantic Match| E
J -->|Keyword Match| F
E --> K("Reciprocal Rank Fusion - RRF")
F --> K
K --> L["Cross-Encoder Reranker"]
end
subgraph subgraph_generation ["Generation & Consensus"]
L --> M["System QA Prompt / Strict Citation Rules"]
M --> N("Self-Consistency Voter / Batch Temp 0.7")
N --> O["Final Synthesized Answer w/ UI Citation Badges"]
end
subgraph subgraph_presentation ["User Presentation"]
O --> P["FastAPI Server"]
P -->|Mounts public/ directory| UI["Interactive HTML/CSS/JS Dashboard"]
end
subgraph subgraph_evaluation ["Offline Evaluation Pipeline"]
GGS["generate_golden_set.py"] -->|Generates| GDS["data/golden_dataset.json (50 Q&A Pairs)"]
GDS -->|Evaluated by| RE["run_evaluation.py / test_ragas_eval.py"]
RE -->|Queries RAG Pipeline| G
RE -->|Runs evaluation via| RAGAS["Ragas Evaluation Framework"]
RAGAS -->|Calculates metrics| METRICS["Faithfulness, Answer Relevancy, Context Precision/Recall"]
end
style subgraph_ingestion fill:#0d1117,stroke:#21262d,stroke-width:2px,color:#c9d1d9
style subgraph_routing fill:#0d1117,stroke:#21262d,stroke-width:2px,color:#c9d1d9
style subgraph_retrieval fill:#0d1117,stroke:#21262d,stroke-width:2px,color:#c9d1d9
style subgraph_generation fill:#0d1117,stroke:#21262d,stroke-width:2px,color:#c9d1d9
style subgraph_presentation fill:#161b22,stroke:#30363d,stroke-width:2px,color:#c9d1d9
style subgraph_evaluation fill:#161b22,stroke:#30363d,stroke-width:2px,color:#c9d1d9๐ ๏ธ Core Engineering Features
1. High-Performance Local & Batch Embeddings
To entirely bypass external API rate limits and network latency during document ingestion and search, the system runs local semantic embeddings using Hugging Face's all-MiniLM-L6-v2 model, persisting vector embeddings to a local disk-based Qdrant instance.
Additionally, for cases where native Google embeddings are preferred, we implement a custom RESTBatchEmbeddings class inside src/retrieval.py to communicate directly with the Google REST endpoint batchEmbedContents. This packages up to 100 document chunks in a single payload/network request (avoiding the concurrent thread-pool rate limits of standard wrappers).
2. Native SEC HTML Loader
Since SEC EDGAR files are natively published in HTML format, our ingestion pipeline in src/ingest.py reads .htm / .html documents directly using BSHTMLLoader wrapped with the built-in python "html.parser". Old, low-accuracy mock PDF files have been deprecated.
3. Agentic Query Decomposer & Router
Complex comparison queries are split into single-topic sub-queries using agent_router.py. Each sub-query runs parallel semantic (Qdrant) and lexical (BM25) searches, which are fused using Reciprocal Rank Fusion (RRF): $$\text{RRF Score}(d \in D) = \sum{m \in M} \frac{1}{k + rm(d)}$$ A Cross-Encoder Reranker (ms-marco-MiniLM-L-6-v2) evaluates joint token representations of the query and candidate passages, filtering out noise and bubbles the most relevant sections to the generation node.
4. Hallucination Detection & Strict Citation Persona
Our prompt guidelines enforce strict bracketed citations linked directly to raw HTML files (e.g. [Document: AAPL_10K.html, Page: 22]). We implement self-consistency checks using a higher temperature batch execution (temperature=0.7, 3 paths). If generating divergent facts, a warning is raised.
๐ Project Structure
Production-RAG/
โโโ data/ # Raw SEC 10-K HTML filings (AAPL, MSFT)
โโโ public/ # Static web dashboard resources
โ โโโ index.html # Front-end structure
โ โโโ style.css # Stylesheets (custom themes, dark mode)
โ โโโ script.js # Frontend chat interface & citation parser
โโโ src/ # Core application logic
โ โโโ api.py # FastAPI router endpoints & mounting
โ โโโ ingest.py # Document loading & character chunking
โ โโโ retrieval.py # REST batching, Qdrant, BM25, RRF, Reranker
โ โโโ agent_router.py # Query decomposition & routing logic
โ โโโ generation.py # LCEL chain, prompt context, self-consistency
โโโ tests/ # Verification suite
โโโ requirements-dev.txt # Core RAG, LLM, and developer dependencies
โโโ requirements_ui.txt # FastAPI/Uvicorn dashboard dependenciesโ๏ธ Configuration & Environment
Create a .env file in the root directory:
GOOGLE_API_KEY=AIzaSy... # Your Google AI Studio API Key
CHUNK_TYPE=fixed # fixed or semantic๐ Running the Application
1. Install Dependencies
pip install -r requirements_ui.txt
pip install -r requirements-dev.txt2. Start the FastAPI Application
Execute uvicorn to start the local server. Ingestion, chunking, and index construction will trigger automatically on the first chat query request:
python -m uvicorn src.api:app --host 0.0.0.0 --port 80003. Access the Dashboard
Open your browser and navigate to: ๐ [http://localhost:8000/](http://localhost:8000/)
๐งช Testing the API
You can test the RAG server programmatically.
cURL Request:
curl -X POST "http://localhost:8000/api/chat" \
-H "Content-Type: application/json" \
-d '{"query": "Compare Apple and Microsoft R&D spending in 2023"}'Response Schema:
{
"answer": "In 2023, Apple's Research and development (R&D) expense was $29,915 million [Document: AAPL_10K.html, Page: 22]. Microsoft's R&D spending was $27,195 million [Document: MSFT_10K.html, Page: 47]. This represents approximately $2.72 billion more spent by Apple.",
"steps": [
"Decomposing complex query...",
"Sub-Query 1: 'What was Apple's R&D spending in 2023?'",
"Sub-Query 2: 'What was Microsoft's R&D spending in 2023?'",
"Running parallel dense/sparse hybrid retrieval...",
"Reciprocal Rank Fusion (RRF) & Cross-Encoder reranking...",
"Synthesizing final multi-hop response."
]
}๐ Evaluation & Observability
RAGAS Integration
We evaluate the quality of responses across four major metrics:
- Faithfulness: Verifies if the answer is derived strictly from context.
- Answer Relevancy: Verifies if the answer directly addresses the user query.
- Context Precision: Measures whether the retrieved documents match ground truth ordering.
- Context Recall: Verifies if the retrieval system recovered all necessary information fragments.
LangSmith Tracing
To trace latency, LLM paths, and RRF rank details, set the following environment variables:
LANGCHAIN_TRACING_V2=true
LANGCHAIN_PROJECT=sec-qa-production
LANGCHAIN_API_KEY=your-langsmith-key๐ก๏ธ Production Deployment Guidelines
For deploying this application in production:
- WSGI/ASGI Server: Run uvicorn behind a process manager like Gunicorn with Uvicorn workers:
gunicorn src.api:app -w 4 -k uvicorn.workers.UvicornWorker -b 0.0.0.0:8000- Reverse Proxy: Place the application behind Nginx to handle SSL termination, rate-limiting, and static file caching for the
public/folder. - Containerization: Use a multi-stage Dockerfile containing caching for Python wheels and lightweight base images (
python:3.12-slim).
