Mahrufa/clinical-rag-evaluation-framework
Evidence-first Clinical RAG Evaluation Framework
An evidence-first clinical RAG evaluation framework for healthcare AI research. The project helps prevent healthcare AI from becoming a black-box answer generator by forcing the system to retrieve traceable source evidence first, then evaluate whether that evidence is useful.
This repository intentionally stops at retrieval. It does not call an LLM or generate medical answers. Retrieved evidence is inspectable through the CLI, API, and hosted playground before any clinical response is considered.
Motivation
Healthcare AI systems need traceable evidence. Before a model can answer a clinical or biomedical question responsibly, it should be able to retrieve the source passages that support an answer. Semantic retrieval helps bridge the gap between natural-language questions and relevant passages that may use different wording.
Examples:
- A user asks about "heart attack discharge medication"; the document may discuss "post-myocardial infarction therapy."
- A user asks about "kidney function monitoring"; the document may refer to "renal labs" or "eGFR."
This project provides the retrieval layer needed for later RAG experiments while keeping the implementation auditable and easy to extend.
Pipeline Overview
- Add PDF files to
data/. clinical_rag_eval.ingestreads PDFs withpypdf.- Text is extracted page by page when page information is available.
- Text is split into overlapping chunks using token-aware chunking by default.
- Reference-heavy and low-value chunks are filtered out.
- Chunks are embedded with
BAAI/bge-small-en-v1.5by default. - Embeddings, text, and metadata are persisted in ChromaDB under
chroma_db/. clinical_rag_eval.queryembeds a user question and prints the top retrieved chunks with source metadata.
Project Structure
clinical-rag-evaluation-framework/
|-- src/
| `-- clinical_rag_eval/
| |-- __init__.py
| |-- config.py
| |-- ingest.py
| |-- query.py
| |-- api.py
| |-- langchain_retriever.py
| |-- rerank.py
| `-- evaluate.py
|-- compare_embeddings.py
|-- eval_queries.json
|-- results/
|-- tests/
|-- pyproject.toml
|-- requirements.txt
|-- README.md
|-- .gitignore
|-- data/ # local only, ignored except .gitkeep
`-- chroma_db/ # generated locally after ingestiondata/ and chroma_db/ are created automatically when needed. They are ignored by Git so the repository does not accidentally include documents or local vector indexes.
Installation
Use Python 3.10 or newer.
python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -r requirements.txt
pip install -e .The first ingestion or query run will download the embedding model from Hugging Face if it is not already cached locally.
Add PDFs
Place PDF files in the data/ folder:
data/
|-- public_guideline.pdf
`-- synthetic_notes.pdfPrivacy requirement: use only public, de-identified, or synthetic documents. Do not store real patient data in this repository or in the local Chroma database.
Keep only one copy of each document in data/. If both an old filename and a renamed filename are present, both will be indexed as separate sources. During ingestion, the script prints all PDFs it found and warns about similar-looking filenames so you can catch duplicates before querying.
Run Ingestion
python -m clinical_rag_eval.ingestOptional arguments:
python -m clinical_rag_eval.ingest --data-dir data --persist-dir chroma_dbpython -m clinical_rag_eval.ingest rebuilds the Chroma collection from scratch by default. Old chunks are deleted before the current PDFs are embedded, which prevents stale results after deleting or renaming files. After ingestion, the script prints the unique source filenames stored in the collection.
The default embedding model is BAAI/bge-small-en-v1.5. It was selected because it performed best in the expanded token-aware benchmark. If you change the embedding model in src/clinical_rag_eval/config.py, rerun python -m clinical_rag_eval.ingest so the vector store is rebuilt with embeddings from the new model.
By default, ingestion uses token-aware chunking with the embedding model tokenizer. This makes chunk sizes closer to what the embedding model actually sees and avoids overly large or inconsistent chunks. Character-based chunking remains available as a simple fallback:
python -m clinical_rag_eval.ingest --chunking-method charThe default chunking settings live in src/clinical_rag_eval/config.py:
CHUNKING_METHOD = "token"
CHUNK_SIZE_TOKENS = 256
CHUNK_OVERLAP_TOKENS = 50
CHUNK_SIZE_CHARS = 900
CHUNK_OVERLAP_CHARS = 150If the tokenizer cannot be loaded, ingestion falls back to character-based chunking and prints a warning.
During ingestion, the pipeline skips chunks that are likely to be bibliography, references, acknowledgements, URL-heavy text, arXiv-heavy text, or dense citation lists. This keeps the vector store focused on explanatory content and improves retrieval quality for conceptual questions.
Run Semantic Retrieval
Interactive mode:
python -m clinical_rag_eval.queryThen type questions at the prompt. Use exit or quit to stop.
Ask a single question:
python -m clinical_rag_eval.query "What does the document say about anticoagulation?"Set the number of results:
python -m clinical_rag_eval.query "How should renal function be monitored?" --top-k 5Use optional cross-encoder reranking:
python -m clinical_rag_eval.query "How should renal function be monitored?" --top-k 5 --rerank --candidate-k 10You can also set --top-k for interactive mode:
python -m clinical_rag_eval.query --top-k 5Use reranking in interactive mode:
python -m clinical_rag_eval.query --top-k 5 --rerank --candidate-k 10Control chunk preview length:
python -m clinical_rag_eval.query "What is retrieval augmented generation?" --preview-chars 1200Each result includes the source filename, page number, chunk index, vector distance, approximate similarity, and a cleaned text preview. The Evidence Summary uses rule-based extractive sentence selection from retrieved chunks; it is not an LLM-generated answer and does not invent new information.
By default, retrieval is vector-only: ChromaDB returns the nearest chunks from the embedding space. With --rerank, the system first retrieves --candidate-k vector candidates, then uses cross-encoder/ms-marco-MiniLM-L-6-v2 to score each query-chunk pair and reorder the final top-k results. Reranking can improve evidence-level retrieval, but it is slower because every candidate is scored by a cross-encoder.
FastAPI API Layer
The CLI remains the main transparent research workflow, but the project also includes a small FastAPI layer for portfolio-ready API integration. Start the API after installing the package:
python -m uvicorn clinical_rag_eval.api:app --reloadAvailable endpoints:
GET http://127.0.0.1:8000/provides basic project info.GET http://127.0.0.1:8000/healthreturns API status.GET http://127.0.0.1:8000/docsprovides interactive API documentation.POST http://127.0.0.1:8000/uploadaccepts a.pdffile and saves it todata/.POST http://127.0.0.1:8000/queryaccepts{"question": "...", "top_k": 5}and returns retrieved chunks from the existing ChromaDB collection.
The upload endpoint does not automatically ingest documents. After uploading PDFs, run python -m clinical_rag_eval.ingest to rebuild the vector store. API logs are written to logs/api.log; uploaded file contents are not logged.
Optional API Key Protection
API key protection is optional. If API_KEY is not set, /upload and /query remain open for local development and portfolio demos. Set API_KEY in the environment to protect those endpoints while keeping /health, /docs, and OpenAPI routes public.
PowerShell example:
$env:API_KEY="change-me"
python -m uvicorn clinical_rag_eval.api:app --reloadInclude the key in protected requests:
X-API-Key: change-meAudit Logging
Normal API logs are written to logs/api.log. Structured audit logs are written to logs/audit.log as JSON lines, recording important API events and errors.
Audit logs intentionally avoid API keys, uploaded file contents, full query text, and retrieved chunk contents.
API Documentation
Health Check Response
Query Endpoint Response
Docker
Build the image:
docker build -t clinical-rag-evaluation-framework .Run the API:
docker run -p 8000:8000 clinical-rag-evaluation-frameworkLocal Docker defaults to port 8000. Cloud deployment platforms may set a PORT environment variable at runtime; the Dockerfile respects that value while keeping EXPOSE 8000 for local usage.
Then open:
http://127.0.0.1:8000/healthhttp://127.0.0.1:8000/docs
Local PDFs and ChromaDB indexes are not committed to Git. For real usage, mount local folders for data/, chroma_db/, and logs/ so documents, vector indexes, and API logs persist outside the container:
docker run -p 8000:8000 -v ${PWD}/data:/app/data -v ${PWD}/chroma_db:/app/chroma_db -v ${PWD}/logs:/app/logs clinical-rag-evaluation-frameworkDeployment Notes
This demo deployment exposes only retrieval endpoints. No real patient data should be uploaded.
data/ and chroma_db/ are local/generated directories and are not committed. On free cloud deployments, /query may not work until documents are uploaded and ingestion is run or a vector store is provided.
Live Demo
The FastAPI Docker app is deployed on Hugging Face Spaces.
- Landing page: https://mahrufa-clinical-rag-evaluation-framework.hf.space/
- Interactive API playground: https://mahrufa-clinical-rag-evaluation-framework.hf.space/demo
- Health check: https://mahrufa-clinical-rag-evaluation-framework.hf.space/health
- API documentation: https://mahrufa-clinical-rag-evaluation-framework.hf.space/docs
- Hugging Face Space: https://huggingface.co/spaces/Mahrufa/clinical-rag-evaluation-framework
The deployed demo exposes the API layer of the project. / opens the evidence-first landing page, /demo lets visitors try evidence retrieval, and /docs opens developer API documentation. The playground includes sample retrieval that works immediately with a built-in public/synthetic corpus and keeps answer generation disabled so retrieved evidence remains inspectable. It also includes upload-and-index support so public or synthetic PDFs can be indexed for querying. The full /query endpoint requires an ingested ChromaDB vector store. Upload and query actions may require an API key if API_KEY is enabled. Do not upload PHI, PII, or real patient records.
This deployment is intended as a portfolio demonstration of Dockerized FastAPI deployment for a retrieval-first healthcare AI project. Do not upload real patient data or sensitive clinical documents.
Privacy and Compliance Notes
This project is a research and portfolio prototype. It is not a production clinical system, does not provide medical advice, and does not generate clinical answers with an LLM. The design is retrieval-first and evidence-focused: it indexes documents, retrieves relevant passages, and supports evaluation of retrieval quality.
Use only public, synthetic, or properly de-identified documents. Do not upload protected health information, personally identifiable information, or real patient records.
The project demonstrates awareness of healthcare privacy and compliance requirements, including HIPAA and GDPR considerations, but it is not certified HIPAA compliant or GDPR compliant. Production use would require additional legal, technical, and organizational controls, including:
- authentication and authorization
- encryption in transit and at rest
- access control
- audit trails
- retention and deletion policies
- consent or another legal basis for data processing
- data processing agreements
- incident response process
- clinical validation and human oversight
API keys provide basic demo-level protection only; they are not complete security. Audit logs intentionally avoid API keys, full query text, uploaded contents, and retrieved chunk contents.
Optional LangChain Retrieval Demo
The main retrieval pipeline is implemented directly with ChromaDB and SentenceTransformers for transparency and easier debugging. LangChain is included as an optional wrapper using langchain-chroma and langchain-huggingface to demonstrate framework integration over the same persisted Chroma collection. This is still retrieval-only; it does not add LLM-based RAG answer generation.
Run the demo after ingestion:
python -m clinical_rag_eval.langchain_retriever "What is MIMIC-IV?"Set the number of retrieved chunks:
python -m clinical_rag_eval.langchain_retriever "What is retrieval augmented generation?" --top-k 5Each result prints the source filename, page number, chunk index, vector distance, and a readable text preview.
Run Retrieval Evaluation
After running ingestion, evaluate retrieval quality with the bundled query set:
python -m clinical_rag_eval.evaluateSet the number of retrieved chunks used for each evaluation query:
python -m clinical_rag_eval.evaluate --top-k 5Evaluate with cross-encoder reranking:
python -m clinical_rag_eval.evaluate --rerankYou can also control candidate depth:
python -m clinical_rag_eval.evaluate --rerank --candidate-k 10The evaluation file, eval_queries.json, contains 25 manually written validation queries with expected source documents, expected evidence keywords, and expected evidence phrases for questions about MIMIC-IV.pdf and retrieval_augmented_generation.pdf.
Metrics:
- Source Recall@K measures whether the expected source document appears within the top K retrieved chunks. For example, Source Recall@3 passes for a query if any of the top 3 chunks comes from the expected PDF.
- MRR, or Mean Reciprocal Rank, rewards systems that return the expected source earlier. A match at rank 1 gets
1.0, rank 2 gets0.5, rank 3 gets0.333, and missing results get0.0. - Average keyword hit rate measures how many expected evidence keywords appear across the retrieved chunks.
- Evidence Phrase Recall@K measures whether any expected evidence phrase appears in the top K retrieved chunks using exact normalized phrase matching. Matching is case-insensitive and whitespace-normalized, but it may undercount useful evidence when PDF extraction changes punctuation or when chunk boundaries split a phrase.
Retrieval evaluation matters for healthcare AI and RAG systems because the generation layer, if added later, can only be as trustworthy as the evidence it receives. Source-level evaluation checks whether the system finds the right document, but it can be easy when there are only a few documents. Evidence-level evaluation is stricter because it checks whether the retrieved chunks contain specific useful passages, not just the correct PDF filename.
Run Tests
Run the lightweight utility test suite:
pip install -e .
python -m pytestThe tests cover text normalization, character chunking, low-value chunk filtering, keyword matching, and evidence phrase matching. They do not require PDFs, ChromaDB, or model downloads.
Compare Embedding Models
Run the same retrieval benchmark across multiple local sentence embedding models:
python compare_embeddings.pyThe script compares:
sentence-transformers/all-MiniLM-L6-v2sentence-transformers/multi-qa-MiniLM-L6-cos-v1BAAI/bge-small-en-v1.5
Set the retrieval depth used for evaluation:
python compare_embeddings.py --top-k 5For each model, the script builds a temporary Chroma vector store from the same PDFs in data/, using the same fixed token-aware chunk set for controlled comparison. It then runs eval_queries.json, prints a comparison table, and saves results to:
results/embedding_comparison.json
results/embedding_comparison.csvEmbedding model choice matters because semantic retrieval depends on how well a model maps questions and document chunks into the same vector space. A model that performs well for general sentence similarity may not retrieve the best evidence for question answering, clinical text, or technical research documents. Comparing models on the same benchmark helps make retrieval design decisions empirical instead of relying on defaults.
Results
Latest retrieval benchmark setup:
- 25 evaluation queries.
- Default embedding model:
BAAI/bge-small-en-v1.5. - Token-aware chunking: 256 tokens with 50-token overlap.
Reranking preserved perfect source-level retrieval and improved evidence-level retrieval by +0.080 on both average keyword hit rate and evidence phrase recall. Reranking is optional because it is slower than vector-only retrieval.
Embedding model comparison can also be run on the expanded 25-query benchmark. The latest model comparison showed that all three embedding models retrieved the correct source document at rank 1, while BAAI/bge-small-en-v1.5 performed best on evidence-level retrieval, with the strongest keyword hit rate and evidence phrase recall. It is now the default embedding model.
Exact comparison metrics are written to results/embedding_comparison.json and results/embedding_comparison.csv when python compare_embeddings.py is run. The phrase-level score is intentionally strict and is sensitive to chunk boundaries and PDF text normalization. Because the benchmark currently uses only two PDFs, these results should be interpreted as an initial validation experiment rather than a broad generalization claim.
Design Decisions
pypdfis used for lightweight local PDF extraction.- Token-aware chunking is supported by default, with character-based chunking kept as a fallback for simplicity and robustness.
- Simple heuristics remove reference-heavy chunks before embedding, reducing retrieval noise before vector search or optional reranking.
- Cross-encoder reranking is optional and disabled by default because it is slower than vector-only retrieval.
- LangChain is included as an optional retrieval integration layer, while the direct ChromaDB implementation remains the main transparent pipeline.
BAAI/bge-small-en-v1.5is the default embedding model because it performed best in the expanded retrieval benchmark.- ChromaDB provides persistent local vector storage without requiring an external database service.
- Metadata is stored with each chunk: source file name, page number when available, and chunk index.
- The project avoids OpenAI or hosted model APIs at this stage to keep the retrieval layer reproducible and private by default.
Current Limitations
- PDF extraction quality depends on the source PDF. Scanned PDFs need OCR, which is not included.
- Token-aware chunking depends on loading the embedding model tokenizer; if that fails, the pipeline falls back to character-based chunking.
- The 25-query evaluation set is intended as a small validation benchmark, not a broad benchmark.
- Cross-encoder reranking is implemented as an optional slower mode.
- Hybrid BM25/vector retrieval and query rewriting are not included.
- The terminal interface is intended for research and debugging, not production use.
- Retrieved passages are not medical advice and should not be treated as clinical guidance.
Future Work
- Additional API hardening, including request validation, rate limiting, and deployment configuration.
- Broader retrieval evaluation metrics such as nDCG and graded relevance labels.
- Authentication for API access.
- Audit logging for document ingestion and retrieval events.
- RAG answer generation with citation-aware responses.
- OCR support for scanned PDFs.
- More advanced chunking strategies, such as section-aware or sentence-aware chunking.
- Hybrid retrieval experiments.
Research Portfolio Notes
This repository is structured to demonstrate the retrieval foundation of a healthcare AI system: data ingestion, metadata preservation, local embeddings, vector persistence, and evidence-first querying. The implementation is intentionally small, but the boundaries are designed so future API, evaluation, and generation layers can be added without rewriting the core pipeline.
