Team Ai
Apppublic

kernelmind/ai-resume-screener

sourceHugging Facemitupdated 7mo agoView on Hugging Face
0likes
App README

๐Ÿš€ AI Resume Screener

An intelligent resume screening system that uses Gemini AI for data extraction, ChromaDB for semantic search, and a hybrid scoring algorithm to rank candidates against job descriptions.

Python FastAPI Streamlit Gemini


โœจ Features

  • โ€”PDF Resume Parsing โ€” Extracts text from PDF resumes with magic-byte validation for security
  • โ€”AI-Powered Extraction โ€” Uses Gemini 2.5 Flash to extract structured data (name, skills, experience)
  • โ€”Vector Search โ€” Stores resume embeddings in ChromaDB for semantic matching
  • โ€”Hybrid Scoring Algorithm โ€” Combines semantic similarity (40%) + skill match (40%) + experience (20%)
  • โ€”Bulk Upload โ€” Ingest multiple resumes in a single API call
  • โ€”Beautiful Dashboard โ€” Streamlit UI with real-time progress, candidate cards, and CSV export

๐Ÿ—๏ธ Architecture

โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”     โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”     โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚  Streamlit   โ”‚โ”€โ”€โ”€โ”€โ–ถโ”‚   FastAPI     โ”‚โ”€โ”€โ”€โ”€โ–ถโ”‚  Gemini AI    โ”‚
โ”‚  Frontend    โ”‚     โ”‚   Backend     โ”‚     โ”‚  (Extraction) โ”‚
โ”‚  (app.py)    โ”‚โ—€โ”€โ”€โ”€โ”€โ”‚  (main.py)    โ”‚     โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜     โ”‚               โ”‚
                    โ”‚               โ”‚โ”€โ”€โ”€โ”€โ–ถโ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
                    โ”‚               โ”‚     โ”‚   ChromaDB     โ”‚
                    โ”‚               โ”‚โ—€โ”€โ”€โ”€โ”€โ”‚ (Vector Store) โ”‚
                    โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜     โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
FileRole
main.pyFastAPI endpoints (/upload/, /upload-bulk/, /match/)
app.pyStreamlit dashboard UI
gemini_service.pyGemini AI structured data extraction with retry logic
embedding_service.pyChromaDB vector storage & semantic search
ranking_service.pyHybrid scoring algorithm
pdf_parser.pyPDF validation & text cleaning
schemas.pyPydantic models (ExtractedResume, JobDescription)

๐Ÿš€ Getting Started

Prerequisites

Installation

bash
# Clone the repo
git clone https://github.com/THEN01EXPLORER/ai-resume-screener.git
cd ai-resume-screener

# Create virtual environment
python -m venv venv
source venv/Scripts/activate   # Windows
# source venv/bin/activate     # Mac/Linux

# Install dependencies
pip install -r requirements.txt

Configuration

Create a .env file in the project root:

env
GEMINI_API_KEY=your_api_key_here

Running

Open two terminals:

bash
# Terminal 1 โ€” Start the FastAPI backend
uvicorn main:app --reload

# Terminal 2 โ€” Start the Streamlit frontend
streamlit run app.py

Then open http://localhost:8501 in your browser.


๐Ÿ“ก API Endpoints

MethodEndpointDescription
POST/upload/Upload a single PDF resume
POST/upload-bulk/Upload multiple PDF resumes
POST/match/Rank all resumes against a job description

๐Ÿงฎ Scoring Algorithm

The hybrid score is calculated as:

Final Score = (Semantic Similarity ร— 0.40) + (Skill Match ร— 0.40) + (Experience ร— 0.20)
  • โ€”Semantic Similarity โ€” Cosine distance from ChromaDB, converted to similarity
  • โ€”Skill Match โ€” Percentage of required skills found in the resume
  • โ€”Experience โ€” Full marks if meets minimum, 50% penalty otherwise

๐Ÿ› ๏ธ Tech Stack

  • โ€”Backend: FastAPI, Uvicorn
  • โ€”Frontend: Streamlit
  • โ€”AI: Google Gemini 2.5 Flash
  • โ€”Vector DB: ChromaDB with SentenceTransformer (all-MiniLM-L6-v2)
  • โ€”PDF Parsing: pdfplumber
  • โ€”Data Validation: Pydantic

๐Ÿ“„ License

This project is open source and available under the MIT License.