small-models-for-glam/glam-extraction-benchmark
GLAM extraction benchmark Structured extraction from cultural-heritage documents. The first configuration is nls-index-cards: 98 manuscript catalogue cards from the National Library of Scotland. Source and credits Derived from NationalLibraryOfScotland/index-cards-eval, revision 2a81070549d8493c2c538744a9dbbc1dc72cb146 (CC0). Images and checked outputs are preserved. NLS cataloguers reviewed the model-drafted labels: 66 accepted as drafted, 32 corrected. Drafting… See the full description on the dataset page: https://huggingface.co/datasets/small-models-for-glam/glam-extraction-benchmark.
GLAM extraction benchmark
Structured extraction from cultural-heritage documents. The first configuration is nls-index-cards: 98 manuscript catalogue cards from the National Library of Scotland.
Source and credits
Derived from NationalLibraryOfScotland/index-cards-eval, revision 2a81070549d8493c2c538744a9dbbc1dc72cb146 (CC0). Images and checked outputs are preserved. NLS cataloguers reviewed the model-drafted labels: 66 accepted as drafted, 32 corrected. Drafting model and original review metadata are preserved in provenance. The original source remains maintained by the National Library of Scotland.
Format
id, image, target_schema, expected_output are the common benchmark columns. Schema and expected output are JSON strings. provenance is a JSON string containing source IDs, revision, image checksum and the source's review metadata. Only image, target schema and declared task instructions are model inputs.
nls-index-cards/manifest.json records the source-to-test split mapping and scoring policy. nls-index-cards/source-schema.json preserves the original schema. The target schema retains its nullable fields and adds x-match: exact annotations to manuscript numbers and folios. The matching rule normalizes text; it is not byte-for-byte equality. The notes field is preserved in gold but excluded from the extraction score, matching the original board.
Scope
NLS has 98 cards; Harvard has 30 scored cards. These are small evaluation sets, not representative of every collection. Keep results for different configs separate. Historical benchmark predictions used a simplified schema without nullable branches; rescoring those predictions does not constitute rerunning models with this schema. This dataset is not yet registered as a Hub benchmark. Future task IDs will be separate from dataset configs; result records must identify config, data revision and scorer version.
Harvard botany headers
harvard-botany-headers, split test: 30 curated cards from Harvard University Botany Libraries, compiled by Walter Deane in 1894–1895. Source: biglam/index-cards-harvard-botany-metropolitan-flora, pinned at 986938d535c7710d33f73219a918ecddf16b3483. The source card reports Harvard's NOT_IN_COPYRIGHT rights statement; retain Harvard attribution. This source has a public-domain rights statement; NLS retains its separate CC0 licence.
Fields are taxon_name, taxon_authority, and taxon_correction, each a string or null. Preserve the original printed heading and author abbreviation, even when corrected. Record an explicit handwritten replacement separately, including its authority when present. Do not expand abbreviations or modernise taxonomy. Body notes, initials, rulings and strike-throughs alone are not replacements. Null means absent, not unknown or illegible. Annotation scope and task instructions are in the config manifest; schemas and outputs are JSON strings as for NLS.
Printed headings on all 30 cards were visually drafted by assistants and reviewed by a human. Both non-null correction strings were also human-confirmed. Absence of corrections was visually audited by an assistant, with direct human confirmation for one ambiguous initial. Per-field review flags are retained in provenance; we do not label every added null as individually human-reviewed. Existing source OCR was not used as extraction ground truth.
The separate review split retains five cards found in a random 25-card schema audit: four handwritten-only headings and one with multiple heading interventions. They have schema_applicable: false and scoring_eligible: false in provenance, a reason, and no extraction gold (expected_output is null). Do not score this split. The harness manifest selects only test, so these cards are kept without requiring special filtering in inference code. Applicability is annotation metadata, not a prediction target. The other 20 audit cards were not fully annotated.
This is a curated POC spanning seven source items; the 25-card audit is separate from the curated scored sample. Score configs separately. Adding Harvard does not invalidate NLS runs pinned to their existing config and revision.
