booleanbeyond/jobfetch
<!-- The YAML block above configures the Hugging Face Space (see "Deploying"). GitHub renders it harmlessly; Hugging Face requires it in README.md. -->
JobFetch
Paste any company careers URL. Get every open role back as structured JSON.
POST /api/scrape {"url": "https://acme.com/careers"}Why the old scraper got confused (and what replaced it)
Two failure modes were reported. Both are addressed structurally, not with patches.
1. "It goes to the footer job and gets confused"
A keyword scraper looks for links containing job or career, and the footer is full of them: Careers, Jobs at Acme, See all openings, Job alerts — often pointing at genuinely job-shaped URLs like /careers/jobs. Filtering by keyword cannot separate those from real openings, because they share every surface feature.
Four independent mechanisms handle it:
Why the gate exists. Scoring alone was not enough, and this was found by failing on a real site. Scoring is additive, so a WordPress/Elementor footer column — 14 title-cased links to product pages, in a container with nofooterin its tag or class — accumulated a passing score without a single job-specific signal, and was returned as 14 jobs while the real board (loaded client-side) was never reached. The gate makes evidence a precondition rather than a bonus: with it, static extraction correctly returns zero, the pipeline falls through to a render, and the real openings are found.tests/fixtures/quick_links_trap.htmlreproduces it exactly, anddiagnostics.dom.rejected[]now reports which signal each rejected candidate was missing.
On top of that, normalize.py rejects any row whose title is navigation vocabulary, a bare location (Austin, TX), or a bare contract type (Full-time) — whichever tier produced it. The vocabulary is `app/vocab.py`, and the regression suite asserts none of it ever appears in output.
2. "Sometimes it doesn't read"
Usually one of five things. Each has a dedicated tier:
Architecture
URL ─► SSRF guard ─► tier cascade ─► normalise ─► dedupe ─► JSONTiers stop at the first one that produces rows:
diagnostics.tiers reports what every tier tried, found and why it was rejected — so when a site does fail, you can see exactly where.
ATS coverage
Full API integration with pagination: Greenhouse, Lever, Ashby, Workday, SmartRecruiters, Workable, Recruitee, BambooHR, Personio, Breezy, Rippling, Pinpoint, JazzHR, Eightfold, Oracle Cloud HCM.
Fingerprinted and reported, then handled by later tiers (no public listing API): iCIMS, Taleo, SuccessFactors, Teamtailor, Jobvite, Phenom, Avature, Zoho Recruit, Freshteam, Paylocity, Dayforce, ADP, Comeet, Polymer, Wellfound.
Adding one is a ~30-line class in `app/extract/ats.py`: give it URL/HTML regexes and a fetch(); the registry and the pipeline pick it up automatically.
The LLM cannot hallucinate a job into the output
Tier 5 runs only when everything deterministic returned nothing, and its output is verified against the page before it is accepted:
- a job is kept only if its title occurs verbatim in the page's visible text;
apply_urlis kept only if that exact URL was in the page's anchor set — otherwise the field is nulled, never invented;location,department,employment_typeandposted_atare dropped unless they also appear on the page;- navigation vocabulary and echoed duplicates are removed.
diagnostics.tiers[].note and the verification block report exactly how many rows were dropped and why. Set LLM_ENABLED=false to run fully deterministically.
Providers
Tier 5 runs through OpenRouter or the Anthropic API — set one key:
OPENROUTER_API_KEY=sk-or-v1-... # https://openrouter.ai/keys
OPENROUTER_MODEL=anthropic/claude-sonnet-4.5LLM_PROVIDER=auto (the default) picks OpenRouter when its key is present and falls back to Anthropic; force one with openrouter or anthropic. Any tool-calling model on OpenRouter works, and models without tool support are handled too — the client falls back to parsing a JSON response body. Extraction is a structured-output task, so a cheaper model is usually sufficient; check /healthz to confirm which provider and model are live. The verification layer is identical whichever provider you use, so a weaker model cannot introduce jobs that are not on the page.
Safety
The service fetches URLs supplied by untrusted users — a textbook SSRF sink. `app/security.py` enforces:
- scheme allow-list (
http/httpsonly) and port allow-list; - cloud-metadata host deny-list;
- DNS resolution before connecting, rejecting any host that resolves to a private, loopback, link-local, CGNAT, multicast or reserved address — which also defeats obfuscated literals (
2130706433,0x7f.0.0.1, IPv4-mapped IPv6, 6to4, Teredo); - DNS pinning: the connection goes to the validated IP with the real hostname in
Hostand TLS SNI, closing the rebinding window between validation and connect; - manual redirect following, so every hop is re-validated — a public URL that 302s to
169.254.169.254is refused at the second hop; - the same guard re-applied to browser subresource requests, which bypass the HTTP client entirely.
Every scrape also runs under a hard envelope: wall-clock deadline, request count, response size, total bytes, redirect depth, page count and job count. Server-wide there is a concurrency gate, a bounded TTL cache and a bounded token-bucket rate limiter (bounded key space — an unbounded dict keyed by user input is itself a DoS vector). Chromium runs as one shared process with a semaphore-bounded context pool; every context is closed in a finally, and an idle watchdog reaps the browser.
robots.txt is respected by default (RESPECT_ROBOTS=true) for HTML fetches. Documented vendor JSON APIs are called directly.
Response shape
Every key is always present; unknown values are null, never missing.
{
"ok": true,
"company": "Acme Robotics",
"source_url": "https://acme.com/careers",
"fetched_at": "2026-08-12T05:58:06.412Z",
"job_count": 6,
"confidence": 0.98,
"jobs": [
{
"id": "4512901",
"requisition_id": "R-1182",
"title": "Senior Backend Engineer",
"department": "Engineering",
"team": "Payments",
"location": "Bengaluru, India",
"locations": ["Bengaluru, India", "Remote - India"],
"workplace_type": "hybrid", // remote | hybrid | onsite | null
"employment_type": "Full-time",
"seniority": "senior",
"experience": "1.5 to 3 years",
"salary": {"min": 180000, "max": 240000, "currency": "USD", "interval": "year"},
"posted_at": "2026-07-14T00:00:00Z", // ISO-8601, or null — never a guess
"updated_at": null,
"apply_url": "https://acme.com/careers/jobs/senior-backend-engineer",
"source": "ats_api",
"confidence": 0.98
}
],
"warnings": [],
"diagnostics": {
"detected_ats": "greenhouse (html:acmecorp)",
"strategy": "ats_api",
"tiers": [ /* every tier: what it found, how long it took, why it was skipped */ ],
"pages_fetched": 3,
"used_browser": false,
"used_llm": false,
"robots_allowed": true,
"duration_ms": 842,
"trace_id": "e5dbb1d38d0d"
}
}Relative dates ("Posted 30+ Days Ago", "yesterday") return null rather than a fabricated timestamp — a deliberate choice, since a wrong date is worse than a missing one.
Endpoints
Errors return {ok: false, error, detail, trace_id} with 400 (bad/unsafe URL), 429 (rate limited), 504 (budget exceeded) or 500. Internal exception detail is logged, never sent to the client.
Running it
python3 -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python -m playwright install chromium # skip if BROWSER_ENABLED=false
cp .env.example .env # add OPENROUTER_API_KEY for tier 5
python -m uvicorn app.main:app --port 8000.env is read automatically at startup and is gitignored. Real environment variables take precedence over it, so container config is never overridden.
Open <http://localhost:8000>.
docker build -t jobfetch .
docker run --rm -p 8000:8000 --env-file .env jobfetchcurl -s 'http://localhost:8000/api/scrape?url=https://boards.greenhouse.io/stripe' | jq '.job_count'Deploying
This service needs a container host, not a serverless platform. The headless browser is not optional: ATS fingerprints on most careers pages are injected by JavaScript, so _tier_ats only sees them after a render. Measured against eight live careers pages with BROWSER_ENABLED=false, the scraper returned zero jobs on all eight — including Greenhouse- and Ashby-backed boards that look like they should work statically. Vercel, Lambda and Cloudflare Workers cannot run Chromium within their bundle limits, so they are not viable targets.
Size the instance for Chromium: 2GB RAM with BROWSER_MAX_CONTEXTS=2. On a 512MB instance the container is OOM-killed mid-scrape rather than erroring cleanly.
Hugging Face Spaces — the only free tier with enough memory, and it needs no credit card. The free CPU tier is 2 vCPU / 16GB, against the 512MB that Koyeb, Render Free and similar offer; Chromium alone is ~400MB resident before FastAPI starts, so those tiers OOM-kill mid-scrape rather than failing cleanly.
The YAML block at the top of this README is the Space config (sdk: docker, app_port: 8000). The Dockerfile runs as UID 1000 because Spaces mandate it.
git remote add hf https://huggingface.co/spaces/<user>/jobfetch
git push hf mainThen add OPENROUTER_API_KEY under the Space's Settings → Secrets.
A Space is public. This app has no authentication — only a rate limit — so anyone who finds the URL can spend your LLM credits. Mint a separate OpenRouter key with a hard credit limit for the Space and keep your unlimited key for local work. Spaces also sleep after 48h idle; the next request wakes it after a cold start.
Render (no CLI needed — deploys from the GitHub remote):
Dashboard -> New -> Blueprint -> pick this repo -> paste OPENROUTER_API_KEYrender.yaml pins the Docker runtime, the standard plan and /healthz. The API key is declared sync: false, so it is prompted for once and never stored in the repo.
Fly.io:
brew install flyctl && fly auth login
fly launch --no-deploy --copy-config
fly secrets set OPENROUTER_API_KEY=sk-or-v1-...
fly deployBoth configs keep ALLOW_PRIVATE_HOSTS=false. That is load-bearing in a hosted environment: this service fetches user-supplied URLs, and disabling the guard would let a caller reach the provider's internal network and metadata endpoint.
Putting the UI on Vercel
The scraper cannot run on Vercel, but the front end is a single self-contained HTML file, which is exactly what Vercel is good at. The split is: Vercel serves the page, the container host runs the API, and the page is told where to find it at build time.
- Deploy the API to Render or Fly as above, and note its origin.
- Import the repo into Vercel.
vercel.jsonalready points atscripts/build-vercel.sh, which stamps the API origin into the page. - Set one Vercel environment variable:
API_BASE=https://jobfetch.onrender.com The build fails loudly if it is missing or not https:// — a UI whose every request 404s against Vercel's static host is worse than a failed deploy, and an http:// API origin is silently blocked as mixed content.
- Allow the Vercel origin on the API host:
CORS_ORIGINS=https://your-app.vercel.appNothing changes for local development or the single-container deploy: the meta tag ships empty, so API_BASE resolves to the empty string and the page calls its own origin exactly as before.
Set CORS_ORIGINS to your front end's origin if you serve the UI separately; leave it empty for same-origin only.
Tests
pytest -q # 179 tests, no network requiredThe suite serves fixture pages from 127.0.0.1 that reproduce the real failure modes: a careers page with five separate footer/nav/CTA traps, JSON-LD mixed with breadcrumbs, a Next.js payload alongside a nav array, a table layout, a JS-only board that is invisible without a browser, pagination, an empty board, a homepage that must be followed to its careers page, and a WordPress-style "Quick Links" footer column that must never be mistaken for a board, and cards that glue title/experience/location/type into a single text node. Plus SSRF rejection (including redirect hops and obfuscated IPs), ATS field mapping and pagination for five vendors, and the LLM verifier's hallucination guards.
Two of them are the regression tests that matter most:
def assert_no_chrome(result):
for job in result.jobs:
assert job.title.strip().lower() not in NAV_WORDS
assert not is_nav_text(job.title)async def test_no_openings_page_returns_zero_not_footer_links(server):
result = await scrape(f"{server}/no_openings.html")
assert result.job_count == 0 # the footer's "Careers" must not fill the gapTuning
Layout
app/
main.py FastAPI app, error handling, rate limiting
pipeline.py tier orchestration, pagination, discovery recursion
security.py SSRF guard: scheme/port/DNS/redirect/pinning
net.py SafeClient: budgets, size caps, manual redirects, robots
normalize.py the single place raw rows become validated Jobs
vocab.py navigation vocabulary, job-URL patterns, role words
htmlutil.py parsing, boilerplate stripping with rescue rules
browser.py Chromium pool, network capture, load-more/scroll
discovery.py careers-page discovery, "no openings" detection
cache.py bounded TTL cache, bounded rate limiter
extract/
ats.py 15 vendor APIs + 15 fingerprint-only vendors
structured.py JSON-LD, microdata, RSS/Atom
embedded.py JSON islands in <script>
jsonfind.py find job arrays in arbitrary JSON
xhrsniff.py captured network payloads
domheur.py the scored DOM extractor
llm.py Claude fallback + verification
static/index.html the UI
tests/ 170 tests + fixture pages