hashmortar/multimodal-annual-reports
Multimodal Annual Reports A document question-answering benchmark built from 20 complete corporate annual and integrated reports. It contains 595 English questions with reference answers, source evidence, original PDFs, reviewed HTML and Markdown representations, and figure/table crops. Questions require interpreting narratives, tables, and non-tabular visuals, including Japanese and French sources. MIT covers original benchmark contributions only. Source reports and their… See the full description on the dataset page: https://huggingface.co/datasets/hashmortar/multimodal-annual-reports.
Multimodal Annual Reports
A document question-answering benchmark built from 20 complete corporate annual and integrated reports. It contains 595 English questions with reference answers, source evidence, original PDFs, reviewed HTML and Markdown representations, and figure/table crops. Questions require interpreting narratives, tables, and non-tabular visuals, including Japanese and French sources.
MIT covers original benchmark contributions only. Source reports and their derivatives retain third-party rights and restrictions; see [RIGHTS.md](RIGHTS.md). Publication does not grant additional rights in those materials.
Contents and scope
Nineteen reports have 30 questions each. Rémy Cointreau has 25; its lower count was justified during review to avoid repetitive additions. The retained partial Redeia conversion and its ten questions are excluded, as are reports outside the final twenty-report expansion inventory. The project's broader target of 165 issuers is a future corpus goal, not the size of this release.
The issuers cover all eleven GICS sectors. Seventeen are from Japan, one from France, one from New Zealand, and one is the UK/Australia Rio Tinto group. Sector classifications and their supporting sources are recorded in document metadata. This is a small, deliberately visually rich, geographically concentrated corpus rather than a representative sample of global corporate reporting.
Dataset configurations
All configurations use the single split evaluation. This denotes the complete evaluation collection; no train, development, or test allocation has been created.
The JSONL files are data/qa.jsonl, data/inputs.jsonl, data/documents.jsonl, and data/assets.jsonl. Explicit configuration paths keep unrelated canonical JSON files out of automatic dataset loading. The package preserves the original corpus directory layout under corpus/data/processed/ and corpus/data/raw_pdfs/.
The inputs configuration excludes answers, evidence locations, categories, family IDs, calculations, and gold/review content. It is a convenient inference view. Gold answers remain available elsewhere in the same package, so this view is not an access-control mechanism.
Load questions and document text
Use the standard Hugging Face loader; no custom loading script is required.
from datasets import load_dataset
repo_id = "hashmortar/multimodal-annual-reports"
qa = load_dataset(repo_id, "qa", split="evaluation")
inputs = load_dataset(repo_id, "inputs", split="evaluation")
documents = load_dataset(repo_id, "documents", split="evaluation")
assets = load_dataset(repo_id, "assets", split="evaluation")
question = inputs[0]["question"]
markdown = documents[0]["markdown"]
html = documents[0]["html"]Document strings can be consumed directly. For PDFs, local HTML/Markdown browsing, or crops, download the repository snapshot and resolve package-relative paths from its root:
from pathlib import Path
from huggingface_hub import snapshot_download
from PIL import Image
snapshot_root = Path(snapshot_download(repo_id, repo_type="dataset"))
image = Image.open(snapshot_root / assets[0]["image_path"])Image paths are strings, not automatically decoded Hugging Face Image features. The snapshot retains sibling crops and source PDFs needed by relative links in report HTML and Markdown. See the official guides for loading datasets, configuration metadata, and repository downloads.
Fields and canonical records
The qa rows expose these fields:
Each evidence record contains source_pdf_path, source_pdf_url, physical_pdf_page, report_page, printed_page, location, support, html_anchor, markdown_anchor, asset_ids, asset_paths, original_language_evidence, and raw_evidence_json. Physical and logical page numbers are integers; printed page labels are nullable strings. Asset ID/path fields are lists. Heterogeneous original evidence fields remain available through the serialized raw record and canonical QA file.
physical_pdf_page is one-based within the named original PDF component. report_page is the one-based logical page used by the combined extraction and report anchors. These differ for multipart reports. Terumo's Sustainability component begins at logical page 35; Chubu's Financial Section begins at logical page 107. For example, Chubu logical page 110 is physical page 4 of its financial PDF. Do not treat a combined-report page number as a page number in every source component.
The original qa.json, report.md, report.html, extraction.json, and assets.json files are retained alongside metadata, source PDFs, crops, and selected audit records. Canonical objects preserve source-specific fields; normalized JSONL records make heterogeneous data browsable and loadable. Calculation objects and other variable structures are serialized rather than forcing incompatible types into one column. The release includes selected acceptance/audit provenance rather than the entire working history.
Some canonical question verification dictionaries still contain historical pending flags or author-stage wording. These fields were preserved rather than rewritten. The normalized verification_status follows the completed final QA expansion review and its linked accepted per-report reviews. Consult those final reviews when determining acceptance; a preserved intermediate flag alone is not the final verdict.
Non-English reports have English analytical text and visual/table descriptions while retaining original labels, quotations, source transcripts, and original-language evidence where recorded. These representations should not be interpreted as certified literal translations. Original PDFs and crops remain the authority for checking labels, units, graphical relationships, and translation choices.
Construction and review
The corpus was acquired from publicly accessible issuer report sources and converted into document representations with separate QA files. Questions and reference answers were not inserted into report HTML or Markdown.
The extraction history used Astra and Sol models. The final expansion used GPT-6.1 Sol, with medium reasoning for coordination and high reasoning for authors and independent solvers/reviewers. Independent solvers worked from original source PDFs and neutral questions before comparing their frozen answers with authored gold. Material replacements received fresh uninformed solves; affected answer/precision repairs received targeted source rechecks as documented. These were separate model contexts, including models from the same family; this is not a claim of human adjudication or independent human validation.
The final expansion accepted all twenty reports and all 395 additions. It preserved the original 200 question objects and protected document inputs. Mechanical audits checked counts, identifiers, links, artifact hashes, and provenance consistency. Such checks do not by themselves prove semantic correctness, translation accuracy, visual necessity, extraction completeness, or reviewer independence. Source-based review records document that additional work and its qualifications.
Intended use and evaluation
Use this corpus to compare models or harnesses that read corporate reports as PDFs, extracted text, HTML, Markdown, and optionally crops. It supports investigation of visual reasoning, table interpretation, source discovery, calculation, and cross-language question answering. It does not provide established performance scores or a universal automatic answer metric.
For retrieval or inference inputs, allowlist original PDFs, report HTML/Markdown, source extraction content, and crops. Exclude qa data, canonical qa.json, reference answers, blind-solve records, and review/proof artifacts from retrieval indexes and model context. A snapshot containing both inputs and gold requires deliberate harness separation.
Report results by document, modality, and source-language tier as well as overall. A macro average across reports helps prevent reports with more questions from dominating comparisons. Inspect source support, calculations, units, and tolerances when judging free-text answers; answers may admit equivalent wording without matching the reference verbatim.
If creating training or validation partitions, keep shared evidence families together. Prefer company/report holdouts for evaluating transfer beyond a known document. Family grouping reduces leakage from shared figures, tables, claims, and calculation inputs; it does not establish statistical independence between all questions or reports.
Limitations
- The corpus is small and strongly concentrated in Japan, with intentionally rich visual content and uneven language coverage.
- Reference answers, translations, visual descriptions, and family assignments are model-authored or model-reviewed and may retain errors despite audits.
- Reviewed extraction is not a claim of perfect parsing. Source discrepancies, legibility constraints, multipart boundaries, and review qualifications remain in the source metadata and audit records.
- The original 200 questions and later additions have different historical authoring/review records. Final normalization preserves that history.
- Report publication dates, fiscal-year labels, and document IDs are separate metadata concepts; an ID year should not substitute for the recorded reporting period.
- This release has no benchmark scores, prescribed train/test split, or claim that it measures general financial-analysis competence.
Rights and licensing
Original QA annotations, metadata, and documentation are offered under MIT only to the extent the contributor holds the necessary rights. Publisher PDFs, images, reproduced text, translations and other document-derived content are not covered by that grant; their owners’ terms remain applicable.
The license above applies only to original contributions within its stated scope. Original issuer PDFs, report content, images, trademarks, auditor material, and derived excerpts retain their respective third-party rights. Public accessibility does not establish permission to redistribute those materials.
The corpus includes third-party source material; no blanket redistribution permission is claimed. Existing source terms include local/personal-use restrictions and other limitations. Particular records include Chubu's KPMG material on logical pages 160–166, Fisher & Paykel Healthcare upload restrictions, and SoftBank restrictions beyond private use. See RIGHTS.md and the per-document metadata for source-specific details. A private Hugging Face repository does not automatically satisfy those terms.
Users must assess the applicable source terms for their intended use and obtain permission where required. Do not infer a blanket license for third-party material from the license assigned to the dataset author's own contributions.
Attribution
Prepared by Harsh Tomar, 2026. Cite this dataset by its repository URL and the commit revision used for evaluation. Source issuers, report URLs, component hashes and retained notices are recorded in the document metadata and provenance files.
