COMHIS/eacl26-detect-latin
Dataset of the EACL 2026 paper Detecting Latin in Historical Books with Large Language Models: A Multimodal Benchmark (arXiv:2510.19585) The benchmark contains 724 annotated pages from eighteenth-century British books. 594 pages contain Latin and 130 pages do not. Each example pairs the original OCR text, LLM-corrected OCR text, page-level metadata, Latin text segments, image references, and manual annotations for Latin regions and text spans. The original dataset record is… See the full description on the dataset page: https://huggingface.co/datasets/COMHIS/eacl26-detect-latin.
Dataset of the EACL 2026 paper Detecting Latin in Historical Books with Large Language Models: A Multimodal Benchmark (arXiv:2510.19585)
The benchmark contains 724 annotated pages from eighteenth-century British books. 594 pages contain Latin and 130 pages do not. Each example pairs the original OCR text, LLM-corrected OCR text, page-level metadata, Latin text segments, image references, and manual annotations for Latin regions and text spans.
The original dataset record is available on Zenodo: https://doi.org/10.5281/zenodo.18377924. The code repository is available at https://github.com/COMHIS/EACL26-detect-latin.
Tasks
The dataset supports two core tasks:
- Page-level Latin detection: predict whether a page contains at least one Latin segment.
- Latin segment extraction: extract the Latin text spans from pages that contain Latin.
The annotations also support multimodal evaluation because each page has a corresponding page image and bounding-box annotations for Latin regions.
Files
data/pages.jsonl: normalized page-level records for Hugging Face use.data/latin_annotation_final.json: original Label Studio-style annotation export from Zenodo.images.zip: page image archive from Zenodo.DATASET_SPEC.md: detailed schema, label taxonomy, and usage notes.
Data Fields
Each row in data/pages.jsonl contains:
page_id: stable identifier in the form<ecco_id>_<page_number>.ecco_id: ECCO document identifier.page_number: page number in the source item.image_path: relative path to the page image insideimages.zipor an extractedimages/directory.has_latin: whether the page contains Latin according to the provided annotation.page_text: original OCR text from thepageTextfield.page_text_clean: LLM-corrected OCR text from thepageTextCleanfield.page_latin: Latin content associated with the page.latin_regions: bounding boxes marked as Latin on the page, represented with percentage coordinates and original image dimensions.latin_spans: text-span annotations with start/end offsets, text, and label. Offsets originate frompage_latin; due to formatting and normalization differences,textshould be treated as the authoritative span content.
Label Notes
Text-span labels in the manual annotations include Direct Quote, Independent, Footnote, Code-switching, Side-note, Bilingual, Emblematic, Legal, Indices and Catalogs, Ecclesiastical, Tables and Charts, Dictionary, and Unlocatable.
Licensing and Terms
This dataset is released under CC BY-NC 4.0. The underlying historical materials are sourced from Eighteenth Century Collections Online (ECCO, Gale) and are used for non-commercial academic research.
Users are responsible for complying with the non-commercial license and any access conditions that apply to the underlying ECCO material.
Citation
If you use this dataset, please cite the accompanying EACL 2026 paper:
@inproceedings{wu-etal-2026-detecting,
title = "Detecting {L}atin in Historical Books with Large Language Models: A Multimodal Benchmark",
author = {Wu, Yu and
Shu, Ke and
Fischer, Jonas and
Pivovarova, Lidia and
Rosson, David and
M{\"a}kel{\"a}, Eetu and
Tolonen, Mikko},
editor = "Demberg, Vera and
Inui, Kentaro and
Marquez, Llu{\'i}s",
booktitle = "Proceedings of the 19th Conference of the {E}uropean Chapter of the {A}ssociation for {C}omputational {L}inguistics (Volume 1: Long Papers)",
month = mar,
year = "2026",
address = "Rabat, Morocco",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2026.eacl-long.245/",
doi = "10.18653/v1/2026.eacl-long.245",
pages = "5305--5328",
ISBN = "979-8-89176-380-7",
abstract = "This paper presents a novel task of extracting low-resourced and noisy Latin fragments from mixed-language historical documents with varied layouts. We benchmark and evaluate the performance of large foundation models against a multimodal dataset of 724 annotated pages. The results demonstrate that reliable Latin detection with contemporary zero-shot models is achievable, yet these models lack a functional comprehension of Latin. This study establishes a comprehensive baseline for processing Latin within mixed-language corpora, supporting quantitative analysis in intellectual history and historical linguistics. Both the dataset and code are available at https://github.com/COMHIS/EACL26-detect-latin."
}