llamaindex/ExtractBench
ExtractBench Quick links: [🌐 Website] [📜 Paper] [💻 Code] Given a document and a schema, a system returns structured data with evidence. The input is a full document, born-digital or scanned, and a schema written by the user. The output is a schema-valid JSON object, with the source page and a bounding box for each value as evidence. It must return correct, exhaustive values (including repeated records), correctly use null for absent information, and ground each extracted… See the full description on the dataset page: https://huggingface.co/datasets/llamaindex/ExtractBench.
ExtractBench v1.1: tighter boxes, new box ground truth, value fixes (#2)
Sync repaired ExtractBench ground truth
Update README.md
Reground ishares printed values: subtotal amounts and inline issuer names cite their printed text, not the row region
Repair evidence bboxes: wrapped-run unions marked coarse, row grounding rebuilt from the printed ordinal column (45 docs)
Repair evidence bboxes: wrapped-run unions marked coarse, row grounding rebuilt from the printed ordinal column (45 docs)
Repair evidence bboxes: wrapped-run unions marked coarse, row grounding rebuilt from the printed ordinal column (45 docs)
Restore ExtractBench
initial commit
