Team Ai
Apppublic

ajnx014/Context-Aware-QA

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes
App README

Check out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference

πŸ“š Context-Aware QA from a Webpage (FAISS + HuggingFace)

This project allows users to ask context-aware questions using the content of any publicly accessible webpage. It uses FAISS for efficient retrieval and HuggingFace's Zephyr-7B model to generate human-like answers based on context.


πŸš€ How It Works

  1. 1.Enter a webpage URL (like a blog, article, or documentation).
  2. 2.Click "Load Webpage" – this fetches and embeds the content using bge-base-en-v1.5 embeddings.
  3. 3.Ask a question related to the webpage.
  4. 4.The model generates a response strictly based on the context. If it doesn't know the answer, it'll say: "I don't know."

🧠 Tech Stack

  • β€”Gradio – for the interactive UI
  • β€”LangChain – for chaining components and context handling
  • β€”HuggingFace Hub – for Zephyr-7B LLM and embedding models
  • β€”FAISS – for vector similarity search

πŸ’‘ What’s happening behind the scenes? It's powered by RAG (Retrieval-Augmented Generation)

Here’s how I implemented RAG in this project:

Retrieval πŸ“š Using FAISS and BAAI/bge-base-en-v1.5 embeddings, I split the webpage into chunks and retrieve the top relevant ones based on the user's question.

Augmentation 🧩 The retrieved chunks are passed as "context" into the prompt β€” this ensures the LLM stays grounded in factual, source-based information.

Generation πŸ€– Using zephyr-7b-alpha from Hugging Face, the model answers based only on the provided context. If no answer is found, it says: "I don't know." β€” no guessing!

⚠️ Known Limitations

  • β€”Only works with public webpages (not behind logins or paywalls)
  • β€”Cannot process PDFs, images, or JavaScript-heavy content
  • β€”Limited understanding of math formulas or code-heavy content

πŸ’‘ Example Use Cases

  • β€”Ask questions from blog posts or documentation
  • β€”Extract info from a product or research article
  • β€”Generate context-aware summaries