Team Ai
Datasetpublic

COMHIS/eacl26-detect-latin

Dataset of the EACL 2026 paper Detecting Latin in Historical Books with Large Language Models: A Multimodal Benchmark (arXiv:2510.19585) The benchmark contains 724 annotated pages from eighteenth-century British books. 594 pages contain Latin and 130 pages do not. Each example pairs the original OCR text, LLM-corrected OCR text, page-level metadata, Latin text segments, image references, and manual annotations for Latin regions and text spans. The original dataset record is… See the full description on the dataset page: https://huggingface.co/datasets/COMHIS/eacl26-detect-latin.

sourceHugging Facecc-by-nc-4.0updated 1d agoView on Hugging Face
0likes50downloads
Dataset Card

Dataset of the EACL 2026 paper Detecting Latin in Historical Books with Large Language Models: A Multimodal Benchmark (arXiv:2510.19585)

The benchmark contains 724 annotated pages from eighteenth-century British books. 594 pages contain Latin and 130 pages do not. Each example pairs the original OCR text, LLM-corrected OCR text, page-level metadata, Latin text segments, image references, and manual annotations for Latin regions and text spans.

The original dataset record is available on Zenodo: https://doi.org/10.5281/zenodo.18377924. The code repository is available at https://github.com/COMHIS/EACL26-detect-latin.

Tasks

The dataset supports two core tasks:

  1. 1.Page-level Latin detection: predict whether a page contains at least one Latin segment.
  2. 2.Latin segment extraction: extract the Latin text spans from pages that contain Latin.

The annotations also support multimodal evaluation because each page has a corresponding page image and bounding-box annotations for Latin regions.

Files

  • —data/pages.jsonl: normalized page-level records for Hugging Face use.
  • —data/latin_annotation_final.json: original Label Studio-style annotation export from Zenodo.
  • —images.zip: page image archive from Zenodo.
  • —DATASET_SPEC.md: detailed schema, label taxonomy, and usage notes.

Data Fields

Each row in data/pages.jsonl contains:

  • —page_id: stable identifier in the form <ecco_id>_<page_number>.
  • —ecco_id: ECCO document identifier.
  • —page_number: page number in the source item.
  • —image_path: relative path to the page image inside images.zip or an extracted images/ directory.
  • —has_latin: whether the page contains Latin according to the provided annotation.
  • —page_text: original OCR text from the pageText field.
  • —page_text_clean: LLM-corrected OCR text from the pageTextClean field.
  • —page_latin: Latin content associated with the page.
  • —latin_regions: bounding boxes marked as Latin on the page, represented with percentage coordinates and original image dimensions.
  • —latin_spans: text-span annotations with start/end offsets, text, and label. Offsets originate from page_latin; due to formatting and normalization differences, text should be treated as the authoritative span content.

Label Notes

Text-span labels in the manual annotations include Direct Quote, Independent, Footnote, Code-switching, Side-note, Bilingual, Emblematic, Legal, Indices and Catalogs, Ecclesiastical, Tables and Charts, Dictionary, and Unlocatable.

Licensing and Terms

This dataset is released under CC BY-NC 4.0. The underlying historical materials are sourced from Eighteenth Century Collections Online (ECCO, Gale) and are used for non-commercial academic research.

Users are responsible for complying with the non-commercial license and any access conditions that apply to the underlying ECCO material.

Citation

If you use this dataset, please cite the accompanying EACL 2026 paper:

bibtex
@inproceedings{wu-etal-2026-detecting,
  title = "Detecting {L}atin in Historical Books with Large Language Models: A Multimodal Benchmark",
  author = {Wu, Yu and
    Shu, Ke and
    Fischer, Jonas and
    Pivovarova, Lidia and
    Rosson, David and
    M{\"a}kel{\"a}, Eetu and
    Tolonen, Mikko},
  editor = "Demberg, Vera and
    Inui, Kentaro and
    Marquez, Llu{\'i}s",
  booktitle = "Proceedings of the 19th Conference of the {E}uropean Chapter of the {A}ssociation for {C}omputational {L}inguistics (Volume 1: Long Papers)",
  month = mar,
  year = "2026",
  address = "Rabat, Morocco",
  publisher = "Association for Computational Linguistics",
  url = "https://aclanthology.org/2026.eacl-long.245/",
  doi = "10.18653/v1/2026.eacl-long.245",
  pages = "5305--5328",
  ISBN = "979-8-89176-380-7",
  abstract = "This paper presents a novel task of extracting low-resourced and noisy Latin fragments from mixed-language historical documents with varied layouts. We benchmark and evaluate the performance of large foundation models against a multimodal dataset of 724 annotated pages. The results demonstrate that reliable Latin detection with contemporary zero-shot models is achievable, yet these models lack a functional comprehension of Latin. This study establishes a comprehensive baseline for processing Latin within mixed-language corpora, supporting quantitative analysis in intellectual history and historical linguistics. Both the dataset and code are available at https://github.com/COMHIS/EACL26-detect-latin."
}