ajnx014/Context-Aware-QA
Check out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
π Context-Aware QA from a Webpage (FAISS + HuggingFace)
This project allows users to ask context-aware questions using the content of any publicly accessible webpage. It uses FAISS for efficient retrieval and HuggingFace's Zephyr-7B model to generate human-like answers based on context.
π How It Works
- Enter a webpage URL (like a blog, article, or documentation).
- Click "Load Webpage" β this fetches and embeds the content using
bge-base-en-v1.5embeddings. - Ask a question related to the webpage.
- The model generates a response strictly based on the context. If it doesn't know the answer, it'll say: "I don't know."
π§ Tech Stack
- Gradio β for the interactive UI
- LangChain β for chaining components and context handling
- HuggingFace Hub β for Zephyr-7B LLM and embedding models
- FAISS β for vector similarity search
π‘ Whatβs happening behind the scenes? It's powered by RAG (Retrieval-Augmented Generation)
Hereβs how I implemented RAG in this project:
Retrieval π Using FAISS and BAAI/bge-base-en-v1.5 embeddings, I split the webpage into chunks and retrieve the top relevant ones based on the user's question.
Augmentation π§© The retrieved chunks are passed as "context" into the prompt β this ensures the LLM stays grounded in factual, source-based information.
Generation π€ Using zephyr-7b-alpha from Hugging Face, the model answers based only on the provided context. If no answer is found, it says: "I don't know." β no guessing!
β οΈ Known Limitations
- Only works with public webpages (not behind logins or paywalls)
- Cannot process PDFs, images, or JavaScript-heavy content
- Limited understanding of math formulas or code-heavy content
π‘ Example Use Cases
- Ask questions from blog posts or documentation
- Extract info from a product or research article
- Generate context-aware summaries
