projectAnish/developersjob-backend
DevelopersJob
AI-powered job search and resume optimization platform built with React 19, TypeScript, and FastAPI. Upload your resume, search across multiple job boards, get ATS compatibility scores, and receive personalized resume improvement suggestions — all in one place.
Live: developersjob.onrender.com · Backend / HF Space: huggingface.co/spaces/projectAnish/developersjob-backend
✨ Features
For Job Seekers
- AI Resume Parsing — Extract skills, experience, education, and contact info from PDF resumes using LLM-powered extraction
- Multi-Source Job Search — Search across 6 paid + free aggregators (Adzuna, JSearch, LinkedIn, SerpAPI, JobSpy, Jooble) plus 141 ATS career pages (Greenhouse, Lever, Ashby, SmartRecruiters, Workday) — covers companies that don't post on aggregators (Stripe, Anthropic, Airbnb, Mastercard, Qualys, Cred, PhonePe, etc.)
- Per-User Paid-API Rotation — Smart rotation across paid providers with per-user daily quotas; cache hits don't consume slots; falls back to free providers automatically
- ATS Score Checker — Get detailed ATS compatibility scores (0-100) with breakdown across 7 scoring factors, with sparse-JD calibration for short ATS career-page listings
- Resume Suggestions — LLM-generated, section-by-section editing guidance tailored to each job description
- Fit Score on Every Job — Each result shows a match percentage to prioritize applications
- Advanced Filters — Filter by skills, title, location, experience level, and posting recency. Two-tier filter (strict role/title + skill-density override) catches consultancy/services postings with generic titles
- Resume Storage — Cloudflare R2 persistent storage so you never re-upload the same resume
- Session Persistence — Search results, ATS scores, and form state persist across page refreshes
Technical Highlights
- Multi-LLM Architecture — Supports Ollama Cloud, HuggingFace (free), and OpenAI (paid fallback)
- 7-Factor ATS Scoring — Keyword match, hard/soft skills, semantic similarity, experience fit, education match, title relevance
- Clerk SSO Integration — Secure authentication with JWT session management
- Responsive Design — Mobile-first UI with Tailwind CSS 4
🏗️ Architecture
┌─────────────────────────────────────────────────────────────────┐
│ Frontend (React 19) │
│ ┌─────────────┐ ┌──────────────┐ ┌─────────────────────────┐ │
│ │ Clerk Auth │ │ SWR Caching │ │ Tailwind CSS 4 Styling │ │
│ └─────────────┘ └──────────────┘ └─────────────────────────┘ │
│ ┌─────────────────────────────────────────────────────────────┐ │
│ │ Components: ResumeSearch, ATSScore, FilterSearch, JobCard │ │
│ └─────────────────────────────────────────────────────────────┘ │
└─────────────────────────────────────────────────────────────────┘
↓ HTTP/REST
┌─────────────────────────────────────────────────────────────────┐
│ Backend (FastAPI + Python) │
│ ┌─────────────┐ ┌──────────────┐ ┌─────────────────────────┐ │
│ │ JWT Auth │ │ Rate Limiting│ │ CORS Middleware │ │
│ └─────────────┘ └──────────────┘ └─────────────────────────┘ │
│ ┌─────────────────────────────────────────────────────────────┐ │
│ │ Agents: ATS Scorer, Job Search, Skill Selector, Customizer │ │
│ └─────────────────────────────────────────────────────────────┘ │
│ ┌─────────────┐ ┌──────────────┐ ┌─────────────────────────┐ │
│ │ Cloudflare │ │ SQLAlchemy │ │ Multi-LLM Router │ │
│ │ R2 Storage │ │ (PostgreSQL) │ │ (Ollama/HF/OpenAI) │ │
│ └─────────────┘ └──────────────┘ └─────────────────────────┘ │
└─────────────────────────────────────────────────────────────────┘🚀 Quick Start
Prerequisites
- Python 3.12+
- Node.js 18+ and npm
- At least one LLM credential:
HF_TOKEN(free),OPENAI_API_KEY, orOLLAMA_API_KEY
1. Clone and Install
git clone https://github.com/MrAnishRajwani/application-tracker.git
cd application-tracker
# Backend
python -m venv .venv
# Windows
.\.venv\Scripts\Activate.ps1
# macOS/Linux
source .venv/bin/activate
pip install -r requirements.txt
# Frontend
cd frontend && npm install && cd ..2. Configure Environment
Copy the template and fill in your keys:
cp env.example .envRequired Variables:
# LLM Provider (at least one)
HF_TOKEN=hf_xxx # Free HuggingFace token
# OPENAI_API_KEY=sk-xxx # Paid OpenAI fallback
# OLLAMA_API_KEY=xxx # Self-hosted or Ollama Cloud
# Auth (Clerk)
CLERK_SECRET_KEY=sk_test_xxx
CLERK_PUBLISHABLE_KEY=pk_test_xxx
SECRET_KEY=your-random-secret-keyOptional (Enhanced Job Search):
ADZUNA_APP_ID=
ADZUNA_APP_KEY=
RAPIDAPI_KEY=
SERPAPI_KEY=Frontend — set Clerk key in `frontend/.env`:
VITE_CLERK_PUBLISHABLE_KEY=pk_test_xxx3. Run Development Servers
Terminal 1 — Backend:
uvicorn src.api_server:app --host 127.0.0.1 --port 8001 --reloadTerminal 2 — Frontend:
cd frontend && npm run devOpen http://localhost:5173 (dev) or http://localhost:8001 (production).
4. Production Build
cd frontend && npm run build && cd ..
uvicorn src.api_server:app --host 0.0.0.0 --port 8001📁 Project Structure
application-tracker/
├── src/
│ ├── api_server.py # FastAPI app — endpoints, CORS, logging
│ ├── config.py # LLM routing, rate limits, feature flags
│ ├── auth.py # JWT tokens, password hashing, rate limiting
│ ├── clerk_auth.py # Clerk JWT verification, user sync
│ ├── storage.py # Cloudflare R2 resume storage (S3-compatible)
│ ├── generate_ats_report.py # TXT/XLSX report generation
│ ├── agents/
│ │ ├── ats_scorer.py # 7-factor ATS scoring engine (TF-IDF + NLP)
│ │ ├── llm_ats_scorer.py # LLM-based ATS scoring (optional)
│ │ ├── customizer_agent.py # Resume suggestion generator
│ │ └── skill_selector_agent.py# Intelligent skill extraction
│ ├── tools/
│ │ ├── job_search.py # Provider fan-out, rotation orchestrator, fit scoring
│ │ │ # prewarm_company_careers — fills cache, no filters
│ │ │ # query_company_careers — cache-only, runtime filters
│ │ ├── ats_scrapers.py # 6 ATS scrapers (Greenhouse/Lever/Ashby/Workable/SmartRecruiters/Workday)
│ │ ├── company_ats_registry.py# Loads merged registry from config/{ats}_companies.json
│ │ ├── scraped_jobs.py # Adapter for web-scraped Workday jobs (prewarm + query)
│ │ ├── cache_prewarm.py # Background pre-warm thread + 30-min refresh
│ │ ├── cache.py # Per-provider response cache (DB-backed, 6h paid / 1h free)
│ │ ├── quota.py # Per-user-per-provider daily counter
│ │ ├── circuit_breaker.py # 429/401 cooldown for paid providers
│ │ └── errors.py # Typed exceptions for provider failures
│ ├── scraping_jobs/ # Web-scraping package (Workday, anti-throttle)
│ │ ├── workday.py # Paginated POST scraper with UA rotation + retry
│ │ ├── user_agents.py # Rotated browser User-Agent pool
│ │ ├── helpers.py # Location/recruiter/seniority classifiers
│ │ ├── metadata.py # Per-job stamping (scraped_at, source)
│ │ └── registry.py # Derives WORKDAY_SLUGS from JSON config
│ ├── database/
│ │ └── tracker.py # SQLAlchemy models (User, Resume, Session, ProviderQueryCache, UserProviderUsage)
│ └── utils/
│ └── experience_calculator.py# Date parsing + overlap merging
├── config/ # Per-ATS company registries (JSON, single source of truth)
│ ├── greenhouse_companies.json # 59 entries
│ ├── lever_companies.json # 6 entries
│ ├── ashby_companies.json # 25 entries
│ ├── smartrecruiters_companies.json # 3 entries
│ └── workday_companies.json # 48 entries (canonical {tenant, site, host} schema)
├── data/
│ └── *_scraped.json.xz # Per-ATS caches (xz; NOT committed — published as
│ # assets on the rolling `scraped-cache` release)
├── scripts/ # Manual ops tools (not deployed)
│ ├── probe_ats_slugs.py # Verify new Greenhouse/Lever/Ashby/SmartRecruiters slugs
│ ├── probe_workday_slugs.py # Verify new Workday tenants
│ └── scrape_workday_jobs.py # CLI runner for the Workday scraper
├── .github/workflows/
│ ├── sync-to-hf.yml # Push to Hugging Face Space on every commit to main
│ └── refresh-scraped-cache.yml # 3x-daily cron: scrape every ATS, publish xz caches
│ # as release assets
├── frontend/
│ ├── src/
│ │ ├── App.tsx # Root component, Clerk auth, tab routing
│ │ ├── api.ts # Backend HTTP client with token refresh
│ │ ├── types.ts # Shared TypeScript interfaces
│ │ ├── components/
│ │ │ ├── ResumeSearch.tsx # Resume upload + job search page
│ │ │ ├── ATSScore.tsx # ATS scoring page with file upload
│ │ │ ├── FilterSearch.tsx # Advanced filter search page
│ │ │ ├── JobCard.tsx # Job result card with fit score
│ │ │ ├── ScoreCard.tsx # ATS score visualization (7-factor)
│ │ │ ├── SuggestionsModal.tsx# Resume suggestions with markdown
│ │ │ ├── AuthPanel.tsx # Login/signup forms with Clerk
│ │ │ └── Footer.tsx # App footer with links
│ │ └── context/
│ │ └── AppSessionContext.tsx# Session state + auth management
│ ├── package.json
│ └── vite.config.ts
├── tests/ # pytest test suite
├── env.example # Environment variable template
├── requirements.txt # Python dependencies
└── Dockerfile # HF Spaces / Docker deployment🔌 API Endpoints
Authentication
Resume
Job Search
ATS Scoring
🔍 Job Search Architecture
A single user search runs through four stages in _search_all_providers:
┌── Stage 0: Paid rotation (per-user, cached, circuit-broken) ────────────┐
│ Tries adzuna → linkedin → jsearch → serpapi(admin-only) in order. │
│ Each provider checked: cache → quota → API call. Cache hit returns │
│ immediately at no quota cost. 429 trips a 60s cooldown; 401 trips 24h. │
│ First non-empty result wins; rotation stops. │
└─────────────────────────────────────────────────────────────────────────┘
┌── Stage 1: JobSpy ──────────────────────────────────────────────────────┐
│ Free scraper for Indeed / LinkedIn / Google Jobs / Naukri. │
└─────────────────────────────────────────────────────────────────────────┘
┌── Stage 2: Jooble (free aggregator) ────────────────────────────────────┐
│ POST API; silently skipped without JOOBLE_API_KEY. │
└─────────────────────────────────────────────────────────────────────────┘
┌── Stage 3: ATS company careers (141 companies, parallel fan-out) ───────┐
│ Greenhouse, Lever, Ashby, SmartRecruiters, Workday public endpoints. │
│ Returns ALL of each company's open jobs; filtered locally. │
└─────────────────────────────────────────────────────────────────────────┘
↓
Two-tier filter (per job, every source)
↓
┌─ Tier 1 STRICT ─────────────────────────────────────┐
│ Title-token + role-class + experience + location. │
│ Catches "Cyber Security Engineer" leaking past a │
│ "Data Engineer Databricks" query. │
└─────────────────────────────────────────────────────┘
↓ (if strict fails)
┌─ Tier 2 SKILL-DENSITY OVERRIDE ─────────────────────┐
│ Bypass title-token + role-class IF │
│ ≥SKILL_OVERRIDE_MIN_MATCHES (default 2) primary │
│ skills appear in title+description. │
│ Location + experience are NEVER bypassed. │
│ Catches Accenture "Customer Success Lead" with │
│ React stack, Deloitte "Senior Consultant" with │
│ Python/Spark/Kafka. │
└─────────────────────────────────────────────────────┘
↓
Dedup by URL → fit-score → sort → trim to max_resultsWhy rotation + cache + per-user quotas? Free-tier paid APIs have tight daily limits (Adzuna ~25/day fleet, SerpAPI ~3/day). The orchestrator distributes user load across providers, caches responses for 6 hours, and falls back to free sources automatically when paid quota exhausts.
Pre-warm + cache-only user path
ATS company-board fetching (Stage 3) and the scraped-Workday source are never run synchronously in the user-search hot path. Instead:
Server startup (lifespan hook)
↓
cache_prewarm.start_prewarm()
↓
┌─ daemon thread ─────────────────────────────────────────┐
│ prewarm_company_careers() │
│ → fans out across config/{ats}_companies.json (141 │
│ companies), populates _company_careers_cache │
│ with raw, location-agnostic jobs. No filtering. │
│ prewarm_scraped_workday() │
│ → 1) reads data/workday_scraped.json.xz (GHA cron │
│ output) if < 36h old — instant population │
│ 2) else live-scrapes the 48 Workday tenants │
│ via src/scraping_jobs/workday.py (UA rotation, │
│ retry, jitter, mid-pagination throttle detect) │
│ Reschedules itself every 30 minutes. │
└─────────────────────────────────────────────────────────┘
↓
User search arrives at /jobs/search/filters
↓
┌─ query_company_careers / query_scraped_workday ────────┐
│ CACHE-ONLY READ — no network calls. Cold cache │
│ returns []; the next pre-warm fills it. │
│ │
│ Location + experience are HARD gates — a job outside │
│ the requested location, or demanding more years than │
│ the candidate has, is dropped outright. │
│ │
│ Relevance is applied by RANKING, not by a cut-off: │
│ the surviving jobs are ordered by skill density across │
│ all the user's skills and the best are returned. A │
│ fixed "must match N skills" threshold used to decide │
│ this, which discarded the tail of a corpus we already │
│ hold in full and matched on ~2 skills regardless of │
│ resume breadth. Pass an explicit `min_skill_score` to │
│ get the old filter behaviour. │
│ │
│ Ranking uses the cheap string-match score; the real │
│ ATS score (~130-180 ms/job) is computed only for the │
│ handful actually returned — see add_fit_scores. │
└─────────────────────────────────────────────────────────┘
↓
┌─ Empty-result fallback (loose) ─────────────────────────┐
│ If every source returned 0 matches, retry with │
│ min_skill_score=0 to surface up to 10 related jobs. │
│ Frontend shows a banner: "We couldn't find exact │
│ matches — here are some related jobs." │
└─────────────────────────────────────────────────────────┘Production deployment (Hugging Face Space):
- Pattern 1 — server pre-warms the in-memory cache on startup; live ATS sources are fast (sub-second per call).
- Pattern 2 — the 3x-daily GHA workflow
.github/workflows/refresh-scraped-cache.ymlruns the scrapers from a clean GitHub-Actions IP and publishesdata/*_scraped.json.xzas assets on the rollingscraped-cacheGitHub Release; the Space downloads them at startup viaGH_TOKEN— Workday's anti-scraping never sees production server traffic.
The caches are not committed. Committing them grew git history by ~105 MB/day (3 runs x ~35 MB) with no way to reclaim it. They are xz rather than gzip because LZMA nearly halves them (greenhouse: 35.1 -> 18.4 MB) for ~95s of compression per run.
🎯 ATS Scoring Breakdown
The ATS engine computes a weighted composite score (0–100):
Score >= 45 is marked as ATS Pass.
Sparse-JD calibration: ATS career-page postings often omit explicit qualifications/experience sections, leaving those score components stuck at the neutral 80 default. When the JD is sparse AND the candidate has actual matching signal (kw + hard + semantic > 30), a targeted boost of up to +15 is added so the displayed score reflects the true relevance rather than the JD's authoring style. Full-JD scoring (Adzuna/JSearch/etc.) is unaffected.
🧪 Testing
# Run all unit tests
python -m pytest tests/ -q
# Live ATS endpoints (skipped by default — hits real boards)
RUN_ATS_LIVE_TESTS=1 python -m pytest tests/test_ats_live.py::test_hr_ba_po_company_returns_jobs -v
# Sweep the entire registry against live ATS endpoints
RUN_ATS_LIVE_TESTS_FULL=1 python -m pytest tests/test_ats_live.py::test_full_registry_returns_jobs -v
# End-to-end flow with the test resume PDF (resume parse → ATS score → filter search)
RUN_INTEGRATION_E2E=1 python -m pytest tests/test_integration_e2e.py -v -s🗂️ Company Registry
The set of companies fetched in Stage 3 lives in config/ as five JSON files — one per ATS:
config/greenhouse_companies.json # 59 companies
config/lever_companies.json # 6
config/ashby_companies.json # 25
config/smartrecruiters_companies.json # 3
config/workday_companies.json # 48 (with tenant/site/host)Each entry is a dict with {name, ats, slug}; Workday entries also need {tenant, site, host} and use a pipe-delimited slug ("tenant|wd#|site") so multiple sites of the same tenant are unique.
Adding a company: edit the appropriate JSON, restart the server. No Python changes required.
Verifying new candidates (before adding to the registry):
# Greenhouse / Lever / Ashby / SmartRecruiters
python scripts/probe_ats_slugs.py
# Workday (registered tenants OR ad-hoc via --slug)
python scripts/probe_workday_slugs.py
python scripts/probe_workday_slugs.py --slug "Microsoft|microsoft|wd1|External_Site|microsoft.wd1.myworkdayjobs.com"⚙️ Environment Variables Reference
LLM Configuration
Authentication
Job Search Providers
Paid-Provider Rotation
When a search hits paid APIs, the orchestrator rotates through them in priority order. Each user has a per-day cap per provider. Cache hits do NOT consume slots. Admin tier gets a bonus on top of regular caps.
Provider Cache
Filter Tuning
Resume Storage (Cloudflare R2)
Rate Limiting
See `env.example` for the full list.
🛠️ Tech Stack
📄 License
MIT
🤝 Contributing
- Fork the repository
- Create a feature branch (
git checkout -b feature/amazing-feature) - Commit your changes (
git commit -m 'Add amazing feature') - Push to the branch (
git push origin feature/amazing-feature) - Open a Pull Request
📞 Support
For issues, questions, or contributions, please open a GitHub issue or contact the maintainers.
