Heit39/LLM_FullTextScreener
Scientific Article Screener
This private Gradio app assists researchers with full-text screening and structured data extraction in PRISMA-like reviews. It reads an Excel review table and text-based article PDFs, asks Azure OpenAI for field suggestions and optional eligibility assessment, and requires a researcher to verify and accept each row.
AI output is a suggestion, not a final review decision. Researchers remain responsible for checking the article, evidence snippets, extracted values, and include/exclude decision.
Use the hosted Hugging Face Space
The Space is intended for invited colleagues. Each authenticated user gets an isolated persistent session containing their own workbook, uploaded PDFs, progress, and exports. Live shared projects are not supported in this version; exchange the exported Excel workbook for handoff.
Administrator setup
- Create a private Gradio Space using this repository.
- Attach persistent Space storage. The app stores data under
/data/llm_fulltextscreener. - Add these Repository secrets under
Settings -> Repository secrets:
AZURE_OPENAI_ENDPOINTAZURE_OPENAI_API_KEYAZURE_OPENAI_DEPLOYMENT- optionally
AZURE_OPENAI_API_VERSION(default:2024-08-01-preview) - optionally
AZURE_OPENAI_TIMEOUT_SECONDSandWORKFLOW_TIMEOUT_SECONDS
AZURE_OPENAI_ENDPOINT can be a classic Azure OpenAI resource root (for example, https://<resource>.openai.azure.com/) or an Azure OpenAI v1 endpoint ending in /openai/v1. The app selects the matching client automatically; the dated API version is used only with the classic endpoint.
- Configure login users. Use one secret per account, for example:
USER1=("alice","strong-password")USER2=("bob","another-strong-password")
Alternatively use SPACE_APP_USERNAME and SPACE_APP_PASSWORD for one account. Authentication is required automatically when running in a Space. Never commit credentials to the repository.
Optional usage reporting:
- Set
USAGE_ADMIN_USERSto a comma-separated list of login usernames that can view the Usage Analytics panel, for examplediogo. If unset, the panel stays hidden. USAGE_DB_PATHselects the SQLite usage database. By default it is created next toAPP_STORAGE_DIR, outside the directory Gradio serves for downloads. Set an absolute path on the persistent/datavolume if you configure custom storage or want to choose its location.
- Review the storage and Azure OpenAI privacy/retention policy before allowing sensitive research documents.
Researcher workflow
- Prepare an
.xlsxworkbook with one header row. Include an article URL column where possible, plus metadata columns such asTitle,Authors, andYear. - Upload the workbook. Leave Fields to extract empty to use all non-URL, non-review columns, or enter a comma-separated list.
- Add a short description for each field. For example:
Number of participants included in the final analysis. - Optionally upload a
criteria.ymlfile. See `examples/criteria.example.yaml`. - When a row loads, the app automatically tries to collect its PDF: the article URL directly, then an open-access copy via Unpaywall using a DOI, then a PubMed Central copy. The DOI is read from the URL first and, if none is found there, from a dedicated
DOIcolumn if the workbook has one; PMC is looked up from aPMCID/PMIDcolumn if present, or otherwise resolved from that same DOI. This only works for openly accessible copies, so it will regularly find nothing — use Try auto-fetch PDF to retry, or upload the PDF manually, then choose Process PDF. - Check the extracted fields and evidence snippets. Edit the fields, decision, and rationale as needed, then choose Accept reviewed row. Use Discard suggestion to clear an AI result or Skip row to move on without accepting it.
- Use the row number control or table preview to navigate. Choose Refresh download after edits to produce the current Excel export.
- Use Delete my saved session when the account’s uploaded documents and progress should be permanently removed.
The app limits uploads and workbooks to protect Space resources. Defaults are 25 MB per Excel file, 50 MB per PDF, 10,000 rows, 100 columns, and 2 million extracted PDF characters. These can be changed with environment variables, but increasing them should be tested against the Space hardware.
Supported documents and limitations
.xlsxworkbooks are supported. The first row must contain unique, non-blank column names.- Text/selectable PDFs are supported. Scanned or image-only PDFs are rejected; OCR is intentionally outside the first hosted release.
- Extracted evidence snippets are shown as read-only reference text and are not currently written to Excel.
- Automatic PDF collection is best-effort and legal-only (direct URL, Unpaywall, or PubMed Central open-access copies); paywalled articles will not be fetched and always fall back to manual upload. Set
AUTO_FETCH_PDF_ENABLED=0to disable fetching on row load/navigation while keeping the manual button. - The Unpaywall fallback requires
UNPAYWALL_EMAIL(a real contact email — Unpaywall rejects placeholder addresses) to be set as a Space variable/secret. Without it, auto-fetch only succeeds for URLs that serve a PDF directly (PMC lookups still work, since they don't require this). - The PMC fallback fetches full-text XML via Europe PMC's REST API, retries transient server errors, and then tries NCBI's PMC OAI endpoint. It re-renders the result as a plain-text PDF — the PMC and Europe PMC article websites are themselves behind bot-challenge walls (AWS WAF / Cloudflare / a proof-of-work check) that this tool does not attempt to bypass. The result has no original figures, tables, or layout — just the article's text. Set
AUTO_FETCH_PMC_ENABLED=0to disable it. - Review decisions and rationales are saved in the
DecisionandInclusion Rationalecolumns. - Sessions are isolated by login username and persist only while the attached Space storage remains available.
- The configured Azure OpenAI endpoint receives the extracted article text and review criteria. Do not use the app for documents prohibited by your organization’s data policy.
Usage Analytics
Each Process PDF attempt for a selected workbook row is recorded with the authenticated username, UTC timestamps, completion status, model request/response counts, and provider-reported token totals. Extraction uses one model workflow; a configured criteria.yml adds a second workflow. The panel aggregates these values by username and includes an Overall row. It is visible only to usernames listed in USAGE_ADMIN_USERS, and the report endpoint checks that access again when refreshed.
The database stores usage counts and timestamps without article text, extracted values, PDFs, or review criteria. Model requests count one OpenAI SDK call per workflow; SDK retries are not counted separately. Token counts are exact when Azure OpenAI returns usage metadata. The Responses with token usage column counts returned responses with at least one token field.
The default database path is a sibling of APP_STORAGE_DIR, so it stays outside Gradio's allowed download path. In a Space, that sibling is on /data when the default storage root is used. Usage history persists only while the volume containing USAGE_DB_PATH persists.
Run locally
- Use Python 3.11 and create a virtual environment.
- Install dependencies:
pip install -r requirements.txt- Create local configuration:
cp .env.example .env Fill in the Azure OpenAI values. .env must remain uncommitted.
- Start the app:
python run_local.py Local sessions are stored in .local_data/ and the app opens at http://127.0.0.1:7860.
Development checks
Run the test suite from the repository root after installing requirements:
python -m unittest discover -vFor a deployment smoke test, set the required Azure variables and run python app.py; the launcher validates the configuration before serving the Space.
