ARslan-Ahamd/grounded-document-extraction
Where Did That Number Come From?
Field extraction from invoices as span selection, not generation. The model emits a pair of token indices, so the value it returns is a slice of the page — a string that does not appear in the document is not a low-probability output, it is not in the output space at all.
The provenance check in the demo is computed live on whatever document you generate, not quoted from the paper.
Measured in the repository: 1,075,037 opportunities, 0 ungrounded values, with verification switched off. A generative head on the same encoder, data, budget and seed hallucinates at 0.9986.
- Code: https://github.com/arslan-ahm/grounded-document-extraction
- Full results: https://grounded-extraction-arslan.surge.sh
- All seven projects: https://seven-ai-projects-arslan.surge.sh
Research artefact on synthetic documents. Not a production system.
Runs entirely in your browser through Pyodide — no server, nothing uploaded, nothing leaves your machine. The first load fetches roughly 12 MB and takes 10–40 seconds depending on the demo; after that it is cached.
