Team Ai
Datasetpublic

llamaindex/ExtractBench

ExtractBench Quick links: [🌐 Website] [📜 Paper] [💻 Code] Given a document and a schema, a system returns structured data with evidence. The input is a full document, born-digital or scanned, and a schema written by the user. The output is a schema-valid JSON object, with the source page and a bounding box for each value as evidence. It must return correct, exhaustive values (including repeated records), correctly use null for absent information, and ground each extracted… See the full description on the dataset page: https://huggingface.co/datasets/llamaindex/ExtractBench.

sourceHugging Faceapache-2.0updated 11d agoView on Hugging Face
58likes16kdownloads
9 commits on main
51bf3a011d ago

ExtractBench v1.1: tighter boxes, new box ground truth, value fixes (#2)

boyang-runllama
f6180e92mo ago

Sync repaired ExtractBench ground truth

boyang-runllama
49d80b12mo ago

Update README.md

boyang-runllama
1fd0c652mo ago

Reground ishares printed values: subtotal amounts and inline issuer names cite their printed text, not the row region

boyang-runllama
42cecd42mo ago

Repair evidence bboxes: wrapped-run unions marked coarse, row grounding rebuilt from the printed ordinal column (45 docs)

boyang-runllama
16941352mo ago

Repair evidence bboxes: wrapped-run unions marked coarse, row grounding rebuilt from the printed ordinal column (45 docs)

boyang-runllama
bd30a2d2mo ago

Repair evidence bboxes: wrapped-run unions marked coarse, row grounding rebuilt from the printed ordinal column (45 docs)

boyang-runllama
c1338dd2mo ago

Restore ExtractBench

boyang-runllama
83e239f2mo ago

initial commit

boyang-runllama