fmegahed/structured_text_extraction
1
1---2title: Structured Text Extraction3emoji: 📚4colorFrom: blue5colorTo: yellow6sdk: docker7pinned: false8license: mit9short_description: A structured text extraction using the ellmer package10---11 12# AI-Powered Text Extraction Tool13 14A Shiny app that turns free text, PDFs, scanned PDFs and images into a table. You define15the columns (a label, a description and a type for each field), and a large language model16fills in one row per item. The table can be downloaded as a CSV file.17 18This is Application 1 of the paper "What Quality Engineers Need to Know About Generative19AI: Structured Data Extraction" by Fadel M. Megahed, Ying-Ju Chen, Yamin Dahwich, Arthur20Carvalho, L. Allison Jones-Farmer, Ibrahim Yousif, and Inez M. Zwetsloot. The paper is in21preparation for submission to *Quality Engineering*.22 23Hosted app: https://huggingface.co/spaces/fmegahed/structured_text_extraction24 25## What the app does26 271. **Input source.** Choose one of six tabs: demo data (two NHTSA recall summaries), pasted28 text, a text file (.txt, .csv, .md), a readable PDF, a scanned PDF, or images (PNG, JPEG,29 WebP, GIF). Only the tab that is open when you click Extract Data is used.302. **Field definitions.** Each field has a label (the column name), a description (the31 instruction the model receives for that field) and a type (text, number, yes/no, or list32 of text). You can start from a template for a common quality document: recall notice,33 8D / corrective action report, customer complaint record, or inspection / nonconformance34 note. Field definitions can be saved to a JSON file and loaded again.353. **Extract Data.** The app sends one request per item to the model. Each item is sent in a36 new conversation. If an item fails, the others are still processed.374. **Results.** One row per item, with a `source_id` column, one column per field, and the38 columns `status` (`ok` or `failed`) and `status_message`.395. **Reproduce in code.** The app shows R ([ellmer](https://ellmer.tidyverse.org/)) and40 Python ([chatlas](https://posit-dev.github.io/chatlas/)) code that runs the same41 extraction with your fields on your own data.42 43The "?" buttons in the app explain each part, including what is sent to the model.44 45## What to keep in mind46 47- **The results need checking.** The model can misread a document or return a value that the48 document does not contain. A status of `ok` means that the request completed. It does not49 mean that the values are correct.50- **Your input leaves your computer.** The hosted app sends the content of the active input51 source and your field definitions to an OpenAI model. Do not submit confidential or52 proprietary records to the hosted app. To keep data in house, run the app yourself or53 adapt the code from "Reproduce in code" to another provider or a local model.54- **The descriptions are the prompt.** The app adds no instructions of its own. A good55 description says what to return, in what format, and what to return when the document56 does not contain the information.57 58## Limits59 60The limits are constants in `R/config.R`, and the text in the app is generated from them.61 62| Limit | Constant |63| --- | --- |64| PDF pages (readable or scanned) or images per run | `MAX_PAGES_OR_IMAGES` |65| Size of one uploaded file, in MB | `MAX_UPLOAD_MB` |66| Number of fields | `MAX_FIELDS` |67 68A readable PDF is sent as one item (the text of its pages). A scanned PDF gives one item per69page, and each image is one item. Text is split into items at blank lines. A .csv file gives70one item per row of its first column, and its first row is treated as a header.71 72## Model73 74The model id is set in one place, `DEFAULT_EXTRACTION_MODEL` in `R/config.R`. To use another75OpenAI model without editing code, set the environment variable `EXTRACTION_MODEL`. The76call, the help text and the generated code all read this value. A test fails if a model id77is written anywhere else in the app code, apart from the record of the model that the78paper's version used.79 80## Versions81 82| Version | Date | Model | Notes |83|---|---|---|---|84| 3.0.1 | 9 October 2026 | `gpt-6-luna` | Released after the paper was drafted. Field templates, saved field definitions, per-item progress and failure handling, code that reproduces an extraction, and in-app help. |85| [2.0.1](https://huggingface.co/spaces/fmegahed/structured_text_extraction/tree/v2.0.1-paper) | 25 June 2026 | `gpt-5-mini` | The version the paper describes (tag `v2.0.1-paper`). |86 87The app has been updated since the paper was drafted, and this is intentional. Language88models are replaced often, and users find problems that only appear in use, so an app built89on a language model needs regular updating. The header of the app links to a short "What90changed" note that is generated from the values in `R/config.R`.91 92## Run it yourself93 94The app needs an OpenAI API key in the environment variable `OPENAI_API_KEY`. On Hugging95Face, set it as a Space secret. Never commit the key (`.env` is listed in `.gitignore` and96`.dockerignore`).97 98With R (4.5.2 is the version used in the Docker image) and the packages shiny, ellmer,99pdftools, magick, base64enc and jsonlite:100 101```r102Sys.setenv(OPENAI_API_KEY = "your key")103shiny::runApp(".", port = 7860)104```105 106With Docker:107 108```bash109docker build -t structured-text-extraction .110docker run --rm -p 7860:7860 -e OPENAI_API_KEY="your key" structured-text-extraction111```112 113Then open http://localhost:7860.114 115## Repository layout116 117| Path | Content |118| --- | --- |119| `app.R` | User interface and server |120| `R/config.R` | Model id and limits |121| `R/fields.R` | Field definitions: validation, JSON import and export, templates, ellmer types |122| `R/sources.R` | Reading the input sources and resolving the active tab into items |123| `R/extraction.R` | One model call per item, failure handling, and assembly of the results |124| `R/snippets.R` | Generated R and Python code |125| `R/help.R` | Text of the "?" help dialogs |126| `config/templates/` | Field templates, one JSON file each |127| `www/` | Theme, client-side script, logos |128| `tests/` | Unit and server tests, and a live smoke test |129 130To add a template, add a JSON file to `config/templates/` with the keys `id`, `name`,131`description`, `order` and `fields` (each field has a `label`, a `description` and a `type`132that is one of `text`, `number`, `yes_no` or `list`).133 134## Tests135 136```bash137Rscript tests/testthat.R138```139 140The tests need the packages testthat and withr. They replace the model with a mock, so they141need no API key and make no model calls. `.github/workflows/tests.yml` runs them on every142push when the repository is hosted on GitHub.143 144`tests/live_smoke.R` is a separate script that makes three short calls to the configured145model (the two demo texts and one generated image). It needs an API key:146 147```bash148Rscript tests/live_smoke.R # key from OPENAI_API_KEY149Rscript tests/live_smoke.R path/to/.env # key from a KEY=VALUE file150```151 152## License153 154MIT155 