Team Ai
Apppublic

booleanbeyond/jobfetch

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes
App README

<!-- The YAML block above configures the Hugging Face Space (see "Deploying"). GitHub renders it harmlessly; Hugging Face requires it in README.md. -->

JobFetch

Paste any company careers URL. Get every open role back as structured JSON.

POST /api/scrape  {"url": "https://acme.com/careers"}

Why the old scraper got confused (and what replaced it)

Two failure modes were reported. Both are addressed structurally, not with patches.

1. "It goes to the footer job and gets confused"

A keyword scraper looks for links containing job or career, and the footer is full of them: Careers, Jobs at Acme, See all openings, Job alerts — often pointing at genuinely job-shaped URLs like /careers/jobs. Filtering by keyword cannot separate those from real openings, because they share every surface feature.

Four independent mechanisms handle it:

#MechanismWhat it rules out
1Prefer data over markup. Six tiers run before any HTML scraping: ATS APIs, JSON-LD, embedded JSON, RSS feeds, captured XHR.On most sites the DOM is never scraped at all, so there is no footer to get wrong.
2Chrome removal. <header>/<nav>/<footer>/<aside>, ARIA landmarks, hidden nodes and chrome-named containers are deleted before candidate detection — with rescue rules so a real list inside <aside class="job-list"> survives.The footer is physically gone by the time anything is scored.
3Repetition requirement. A job list is a set of siblings sharing a structural signature. A lone footer "Careers" link has no same-signature sibling, so it can never form a candidate.Single stray links, hero CTAs, "View all jobs" buttons.
4Evidence gate. Before scoring counts for anything, a candidate must clear at least one job-specific signal: job-shaped URLs (>=50%), locations (>=40%), employment types (>=40%), job-list container naming (>=50%), or job-title nouns in the titles (>=50%). A group sitting under a footer-column heading ("Quick Links", "Company", "Products") is vetoed outright.Long, plausible link columns that carry no job evidence at all.
5Scored single winner. Surviving candidates score on href shape, href distinctness, title validity, location/type presence and container naming, and are penalised for navigation vocabulary, all-identical hrefs, and self-links. Exactly one candidate wins.Footer links can never be mixed into an otherwise good result set.
Why the gate exists. Scoring alone was not enough, and this was found by failing on a real site. Scoring is additive, so a WordPress/Elementor footer column — 14 title-cased links to product pages, in a container with no footer in its tag or class — accumulated a passing score without a single job-specific signal, and was returned as 14 jobs while the real board (loaded client-side) was never reached. The gate makes evidence a precondition rather than a bonus: with it, static extraction correctly returns zero, the pipeline falls through to a render, and the real openings are found. tests/fixtures/quick_links_trap.html reproduces it exactly, and diagnostics.dom.rejected[] now reports which signal each rejected candidate was missing.

On top of that, normalize.py rejects any row whose title is navigation vocabulary, a bare location (Austin, TX), or a bare contract type (Full-time) — whichever tier produced it. The vocabulary is `app/vocab.py`, and the regression suite asserts none of it ever appears in output.

2. "Sometimes it doesn't read"

Usually one of five things. Each has a dedicated tier:

CauseHandling
Board is an embedded ATS iframe/scriptATS fingerprints are matched against the HTML, not just the URL, so acme.com/careers resolves to its Greenhouse/Lever/Ashby board and its API is called
Page is a JS-only SPAData usually ships inline (__NEXT_DATA__, __NUXT__, Apollo cache) — read statically, no browser needed. If not, Chromium renders it
Jobs load over XHRThe page's own API response is captured during render and parsed directly
Paginated / "Load more" / infinite scrollATS APIs paginate via their own params; static pagers are crawled; client-side pagers are clicked and scrolled to exhaustion
Homepage pasted instead of careers pageCareers links are discovered, ranked and followed (one hop)
Genuinely no openingsDetected and reported as job_count: 0 with a warning — not as a failure, and never backfilled with footer links

Architecture

URL ─► SSRF guard ─► tier cascade ─► normalise ─► dedupe ─► JSON

Tiers stop at the first one that produces rows:

#TierSourceConfidence
1ats_apiVendor JSON API0.98
2structured_dataschema.org JobPosting (JSON-LD / microdata)0.95
2bembedded_json__NEXT_DATA__, __NUXT__, Apollo state~0.80
2cjob_feedDeclared RSS/Atom jobs feed0.90
3xhr_captureThe SPA's own API, seen through Chromium~0.80
4dom_heuristicScored, chrome-stripped repeated structure~0.50
5llm_fallbackClaude, verified line-by-line against the page0.60

diagnostics.tiers reports what every tier tried, found and why it was rejected — so when a site does fail, you can see exactly where.

ATS coverage

Full API integration with pagination: Greenhouse, Lever, Ashby, Workday, SmartRecruiters, Workable, Recruitee, BambooHR, Personio, Breezy, Rippling, Pinpoint, JazzHR, Eightfold, Oracle Cloud HCM.

Fingerprinted and reported, then handled by later tiers (no public listing API): iCIMS, Taleo, SuccessFactors, Teamtailor, Jobvite, Phenom, Avature, Zoho Recruit, Freshteam, Paylocity, Dayforce, ADP, Comeet, Polymer, Wellfound.

Adding one is a ~30-line class in `app/extract/ats.py`: give it URL/HTML regexes and a fetch(); the registry and the pipeline pick it up automatically.


The LLM cannot hallucinate a job into the output

Tier 5 runs only when everything deterministic returned nothing, and its output is verified against the page before it is accepted:

  • —a job is kept only if its title occurs verbatim in the page's visible text;
  • —apply_url is kept only if that exact URL was in the page's anchor set — otherwise the field is nulled, never invented;
  • —location, department, employment_type and posted_at are dropped unless they also appear on the page;
  • —navigation vocabulary and echoed duplicates are removed.

diagnostics.tiers[].note and the verification block report exactly how many rows were dropped and why. Set LLM_ENABLED=false to run fully deterministically.

Providers

Tier 5 runs through OpenRouter or the Anthropic API — set one key:

bash
OPENROUTER_API_KEY=sk-or-v1-...          # https://openrouter.ai/keys
OPENROUTER_MODEL=anthropic/claude-sonnet-4.5

LLM_PROVIDER=auto (the default) picks OpenRouter when its key is present and falls back to Anthropic; force one with openrouter or anthropic. Any tool-calling model on OpenRouter works, and models without tool support are handled too — the client falls back to parsing a JSON response body. Extraction is a structured-output task, so a cheaper model is usually sufficient; check /healthz to confirm which provider and model are live. The verification layer is identical whichever provider you use, so a weaker model cannot introduce jobs that are not on the page.


Safety

The service fetches URLs supplied by untrusted users — a textbook SSRF sink. `app/security.py` enforces:

  • —scheme allow-list (http/https only) and port allow-list;
  • —cloud-metadata host deny-list;
  • —DNS resolution before connecting, rejecting any host that resolves to a private, loopback, link-local, CGNAT, multicast or reserved address — which also defeats obfuscated literals (2130706433, 0x7f.0.0.1, IPv4-mapped IPv6, 6to4, Teredo);
  • —DNS pinning: the connection goes to the validated IP with the real hostname in Host and TLS SNI, closing the rebinding window between validation and connect;
  • —manual redirect following, so every hop is re-validated — a public URL that 302s to 169.254.169.254 is refused at the second hop;
  • —the same guard re-applied to browser subresource requests, which bypass the HTTP client entirely.

Every scrape also runs under a hard envelope: wall-clock deadline, request count, response size, total bytes, redirect depth, page count and job count. Server-wide there is a concurrency gate, a bounded TTL cache and a bounded token-bucket rate limiter (bounded key space — an unbounded dict keyed by user input is itself a DoS vector). Chromium runs as one shared process with a semaphore-bounded context pool; every context is closed in a finally, and an idle watchdog reaps the browser.

robots.txt is respected by default (RESPECT_ROBOTS=true) for HTML fetches. Documented vendor JSON APIs are called directly.


Response shape

Every key is always present; unknown values are null, never missing.

jsonc
{
  "ok": true,
  "company": "Acme Robotics",
  "source_url": "https://acme.com/careers",
  "fetched_at": "2026-08-12T05:58:06.412Z",
  "job_count": 6,
  "confidence": 0.98,
  "jobs": [
    {
      "id": "4512901",
      "requisition_id": "R-1182",
      "title": "Senior Backend Engineer",
      "department": "Engineering",
      "team": "Payments",
      "location": "Bengaluru, India",
      "locations": ["Bengaluru, India", "Remote - India"],
      "workplace_type": "hybrid",          // remote | hybrid | onsite | null
      "employment_type": "Full-time",
      "seniority": "senior",
      "experience": "1.5 to 3 years",
      "salary": {"min": 180000, "max": 240000, "currency": "USD", "interval": "year"},
      "posted_at": "2026-07-14T00:00:00Z", // ISO-8601, or null — never a guess
      "updated_at": null,
      "apply_url": "https://acme.com/careers/jobs/senior-backend-engineer",
      "source": "ats_api",
      "confidence": 0.98
    }
  ],
  "warnings": [],
  "diagnostics": {
    "detected_ats": "greenhouse (html:acmecorp)",
    "strategy": "ats_api",
    "tiers": [ /* every tier: what it found, how long it took, why it was skipped */ ],
    "pages_fetched": 3,
    "used_browser": false,
    "used_llm": false,
    "robots_allowed": true,
    "duration_ms": 842,
    "trace_id": "e5dbb1d38d0d"
  }
}

Relative dates ("Posted 30+ Days Ago", "yesterday") return null rather than a fabricated timestamp — a deliberate choice, since a wrong date is worse than a missing one.

Endpoints

MethodPathNotes
GET/The UI
POST/api/scrape{url, force_refresh?, use_browser?, use_llm?, max_pages?}
GET/api/scrape?url=…Same, for quick checks and cURL
GET/healthzBrowser/LLM/cache status
GET/api/configEffective settings, credentials redacted
GET/docsOpenAPI

Errors return {ok: false, error, detail, trace_id} with 400 (bad/unsafe URL), 429 (rate limited), 504 (budget exceeded) or 500. Internal exception detail is logged, never sent to the client.


Running it

bash
python3 -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python -m playwright install chromium      # skip if BROWSER_ENABLED=false
cp .env.example .env                       # add OPENROUTER_API_KEY for tier 5
python -m uvicorn app.main:app --port 8000

.env is read automatically at startup and is gitignored. Real environment variables take precedence over it, so container config is never overridden.

Open <http://localhost:8000>.

bash
docker build -t jobfetch .
docker run --rm -p 8000:8000 --env-file .env jobfetch
bash
curl -s 'http://localhost:8000/api/scrape?url=https://boards.greenhouse.io/stripe' | jq '.job_count'

Deploying

This service needs a container host, not a serverless platform. The headless browser is not optional: ATS fingerprints on most careers pages are injected by JavaScript, so _tier_ats only sees them after a render. Measured against eight live careers pages with BROWSER_ENABLED=false, the scraper returned zero jobs on all eight — including Greenhouse- and Ashby-backed boards that look like they should work statically. Vercel, Lambda and Cloudflare Workers cannot run Chromium within their bundle limits, so they are not viable targets.

Size the instance for Chromium: 2GB RAM with BROWSER_MAX_CONTEXTS=2. On a 512MB instance the container is OOM-killed mid-scrape rather than erroring cleanly.

Hugging Face Spaces — the only free tier with enough memory, and it needs no credit card. The free CPU tier is 2 vCPU / 16GB, against the 512MB that Koyeb, Render Free and similar offer; Chromium alone is ~400MB resident before FastAPI starts, so those tiers OOM-kill mid-scrape rather than failing cleanly.

The YAML block at the top of this README is the Space config (sdk: docker, app_port: 8000). The Dockerfile runs as UID 1000 because Spaces mandate it.

bash
git remote add hf https://huggingface.co/spaces/<user>/jobfetch
git push hf main

Then add OPENROUTER_API_KEY under the Space's Settings → Secrets.

A Space is public. This app has no authentication — only a rate limit — so anyone who finds the URL can spend your LLM credits. Mint a separate OpenRouter key with a hard credit limit for the Space and keep your unlimited key for local work. Spaces also sleep after 48h idle; the next request wakes it after a cold start.

Render (no CLI needed — deploys from the GitHub remote):

Dashboard -> New -> Blueprint -> pick this repo -> paste OPENROUTER_API_KEY

render.yaml pins the Docker runtime, the standard plan and /healthz. The API key is declared sync: false, so it is prompted for once and never stored in the repo.

Fly.io:

bash
brew install flyctl && fly auth login
fly launch --no-deploy --copy-config
fly secrets set OPENROUTER_API_KEY=sk-or-v1-...
fly deploy

Both configs keep ALLOW_PRIVATE_HOSTS=false. That is load-bearing in a hosted environment: this service fetches user-supplied URLs, and disabling the guard would let a caller reach the provider's internal network and metadata endpoint.

Putting the UI on Vercel

The scraper cannot run on Vercel, but the front end is a single self-contained HTML file, which is exactly what Vercel is good at. The split is: Vercel serves the page, the container host runs the API, and the page is told where to find it at build time.

  1. 1.Deploy the API to Render or Fly as above, and note its origin.
  2. 2.Import the repo into Vercel. vercel.json already points at scripts/build-vercel.sh, which stamps the API origin into the page.
  3. 3.Set one Vercel environment variable:
   API_BASE=https://jobfetch.onrender.com

The build fails loudly if it is missing or not https:// — a UI whose every request 404s against Vercel's static host is worse than a failed deploy, and an http:// API origin is silently blocked as mixed content.

  1. 1.Allow the Vercel origin on the API host:
   CORS_ORIGINS=https://your-app.vercel.app

Nothing changes for local development or the single-container deploy: the meta tag ships empty, so API_BASE resolves to the empty string and the page calls its own origin exactly as before.

Set CORS_ORIGINS to your front end's origin if you serve the UI separately; leave it empty for same-origin only.

Tests

bash
pytest -q     # 179 tests, no network required

The suite serves fixture pages from 127.0.0.1 that reproduce the real failure modes: a careers page with five separate footer/nav/CTA traps, JSON-LD mixed with breadcrumbs, a Next.js payload alongside a nav array, a table layout, a JS-only board that is invisible without a browser, pagination, an empty board, a homepage that must be followed to its careers page, and a WordPress-style "Quick Links" footer column that must never be mistaken for a board, and cards that glue title/experience/location/type into a single text node. Plus SSRF rejection (including redirect hops and obfuscated IPs), ATS field mapping and pagination for five vendors, and the LLM verifier's hallucination guards.

Two of them are the regression tests that matter most:

python
def assert_no_chrome(result):
    for job in result.jobs:
        assert job.title.strip().lower() not in NAV_WORDS
        assert not is_nav_text(job.title)
python
async def test_no_openings_page_returns_zero_not_footer_links(server):
    result = await scrape(f"{server}/no_openings.html")
    assert result.job_count == 0        # the footer's "Careers" must not fill the gap

Tuning

SymptomKnob
Junk rows on unusual layoutsRaise DOM_MIN_SCORE (default 3.0)
A real board comes back emptyLower DOM_MIN_SCORE; check diagnostics.tiers first
Timeouts on large boardsRaise TOTAL_BUDGET_S, MAX_FETCHES_PER_SCRAPE, MAX_PAGES
Memory pressureLower BROWSER_MAX_CONTEXTS, MAX_CONCURRENT_SCRAPES
Blocked by robots.txtRESPECT_ROBOTS=false — a policy decision, so it is explicit

Layout

app/
  main.py          FastAPI app, error handling, rate limiting
  pipeline.py      tier orchestration, pagination, discovery recursion
  security.py      SSRF guard: scheme/port/DNS/redirect/pinning
  net.py           SafeClient: budgets, size caps, manual redirects, robots
  normalize.py     the single place raw rows become validated Jobs
  vocab.py         navigation vocabulary, job-URL patterns, role words
  htmlutil.py      parsing, boilerplate stripping with rescue rules
  browser.py       Chromium pool, network capture, load-more/scroll
  discovery.py     careers-page discovery, "no openings" detection
  cache.py         bounded TTL cache, bounded rate limiter
  extract/
    ats.py         15 vendor APIs + 15 fingerprint-only vendors
    structured.py  JSON-LD, microdata, RSS/Atom
    embedded.py    JSON islands in <script>
    jsonfind.py    find job arrays in arbitrary JSON
    xhrsniff.py    captured network payloads
    domheur.py     the scored DOM extractor
    llm.py         Claude fallback + verification
static/index.html  the UI
tests/             170 tests + fixture pages