datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ocr-pl-lines
ocr-pl-lines
Syntetyczny zbiór linii tekstu po polsku do fine-tuningu OCR (TrOCR).
Pary NNNNN.png (obraz linii) + NNNNN.txt (transkrypcja).
Struktura
train/ — 2000 par (seed 42)
val/ — 200 par (seed 123)
Generowanie
OCR_engine —
python -m training.generate_synthetic
Korpus: zdania potoczne i urzędowe, domeny (faktury, umowy, medyczne,
prawnicze), losowe daty/kwoty/adresy/NIP/PESEL, zdania z pl.wikipedia.org.
Augmentacje: pochylenie, blur, szum… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/ocr-pl-lines.lazy-lines-medialines-dataset
P&ID Line Detection Dataset
This dataset contains cropped images from P&ID (Piping and Instrumentation Diagrams)
with line segment annotations for line detection and segmentation tasks.
Dataset Splits
The dataset is split by source image to prevent data leakage:
Split
Source Images
Samples
Line Segments
Train
400 (80%)
8,000
69,483
Validation
50 (10%)
1,000
9,485
Test
50 (10%)
1,000
8,940
Dataset Structure
lines_dataset/
├── train/
│… See the full description on the dataset page: https://huggingface.co/datasets/prasatee/lines-dataset.pid_lines_dataset
P&ID Line Detection Dataset
This dataset contains cropped images from P&ID (Piping and Instrumentation Diagrams)
with line segment annotations for line detection and segmentation tasks.
Dataset Structure
Each sample contains:
file_name: Image filename
source_image_idx: Index of the original P&ID image
crop_idx: Index of this crop from the source image
width: Crop width in pixels
height: Crop height in pixels
lines: Dictionary with:
segments: List of line segments as [x1… See the full description on the dataset page: https://huggingface.co/datasets/prasatee/pid_lines_dataset.ehri-pl-lines
ehri-pl-lines
Line-level OCR/HTR dataset of Polish typewritten historical documents, derived
from the Polish sub-corpus of the EHRI dataset.
Each example is a single text-line crop plus its ground-truth transcription. Line
bounding boxes come from the original ALTO XML ground truth (no automatic
detection was used), so transcriptions are reliable and aligned.
Content and provenance
468 line crops from 15 pages across 6 documents (ZIH collection).
Source images… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/ehri-pl-lines.pid_lines_dataset
P&ID Line Detection Dataset
This dataset contains cropped images from P&ID (Piping and Instrumentation Diagrams)
with line segment annotations for line detection and segmentation tasks.
Dataset Structure
Each sample contains:
file_name: Image filename
source_image_idx: Index of the original P&ID image
crop_idx: Index of this crop from the source image
width: Crop width in pixels
height: Crop height in pixels
lines: Dictionary with:
segments: List of line… See the full description on the dataset page: https://huggingface.co/datasets/samiksha9874/pid_lines_dataset.Copiale_Lines
Copiale Lines
Copiale Lines is a line-level image-to-text dataset for historical cipher decipherment. It contains cropped line images from the Copiale manuscript paired with plaintext ground truth.
This dataset was presented in the paper Learning to Decipher from Pixels -- A Case Study of Copiale (HistoCrypt 2026).
Code: https://github.com/leitro/Decipher-from-Pixels-Copiale
Dataset Structure
The dataset is split into:
train: 1,269 samples
valid: 175 samples… See the full description on the dataset page: https://huggingface.co/datasets/leitro/Copiale_Lines.goteborgs_poliskammare_fore_1900_linesarams28k-htr-lineshebrew_synth_linesEn_IAM_lineseval_htr_out_of_domain_linesOpenITI-MAKHZAN-Ottoman-Lines
Dataset Card for OpenITI MAKHZAN Ottoman Lines
Dataset Summary
This dataset contains line-level image-text pairs of historical Ottoman Turkish manuscripts and printed documents. It is derived from the OpenITI MAKHZAN dataset, a large aggregation of Arabic-script ground truth and evaluation data developed by the Open Islamicate Texts Initiative (OpenITI).
The dataset specifically focuses on Ottoman Turkish texts and is highly valuable for training and evaluating… See the full description on the dataset page: https://huggingface.co/datasets/OttomanNLP/OpenITI-MAKHZAN-Ottoman-Lines.pid_lines_dataset
P&ID Line Detection Dataset
This dataset contains cropped images from P&ID (Piping and Instrumentation Diagrams)
with line segment annotations for line detection and segmentation tasks.
Dataset Structure
Each sample contains:
file_name: Image filename
source_image_idx: Index of the original P&ID image
crop_idx: Index of this crop from the source image
width: Crop width in pixels
height: Crop height in pixels
lines: Dictionary with:
segments: List of line… See the full description on the dataset page: https://huggingface.co/datasets/Sri1311/pid_lines_dataset.image-text_bullinger-autoren-lines
image-text_bullinger-autoren-lines
Line-level variant of dh-unibe/image-text_bullinger-autoren at revision 85ea246a5021a5f244c392f97f2c8ac9a42cc717. One row per transcribed line — the cropped line image and its transcription — so the material can be read without a PageXML parser or a line segmenter.
Lines
73502
Characters
3866274
Source
dh-unibe/image-text_bullinger-autoren
Source revision
85ea246a5021a5f244c392f97f2c8ac9a42cc717
What the source reported… See the full description on the dataset page: https://huggingface.co/datasets/dh-unibe/image-text_bullinger-autoren-lines.bergskollegium_relationer_och_skrivelser_linesalvsborgs_losen_linessvea_hovratt_linesinstance-synthetic-lineskurrent-hanse-xvi-test-lines-with-evaluation
Dataset Card for evaluation-hf-printed
This dataset is derived from danameyer/evaluation-hf-printed and has been enriched with inference results.
Dataset Summary
This dataset contains 164 samples across 1 split(s).
Projects Included
1505-02-10_Hanserezess,Lübeck_Dienstag_nach_Scholastice_1505(SAHST_Rep__2,_I_040-4)
Evaluation Results
This repository includes evaluation artifacts for the latest inference output.
Timestamp:… See the full description on the dataset page: https://huggingface.co/datasets/jwidmer/kurrent-hanse-xvi-test-lines-with-evaluation.pinkas-hebrew-lines
Pinkas — Hebrew cursive text lines (line crops)
Line-level crops of The Pinkas Dataset: 30 pages of early-modern European Jewish community
records (c. 1500–1800), written by non-professional scribes in Hebrew and Yiddish, with
line-level ground truth in PAGE XML.
Original dataset — all credit to its authors:
Irina Rabaev, Berat Kurar Barakat, Jihad El-Sana. The Pinkas Dataset. Zenodo, 2019.
https://doi.org/10.5281/zenodo.3569694 (record 3569694). License: CC-BY-4.0.
This is a… See the full description on the dataset page: https://huggingface.co/datasets/cyttic/pinkas-hebrew-lines.image-text_aaeb-xiv-xvii-part-2-lines
image-text_aaeb-xiv-xvii-part-2-lines
Line-level variant of dh-unibe/image-text_aaeb-xiv-xvii-part-2 at revision ce2739fce41c0b8f8673127525839842411020b0. One row per transcribed line — the cropped line image and its transcription — so the material can be read without a PageXML parser or a line segmenter.
Lines
3386
Characters
107134
Source
dh-unibe/image-text_aaeb-xiv-xvii-part-2
Source revision
ce2739fce41c0b8f8673127525839842411020b0
What the source… See the full description on the dataset page: https://huggingface.co/datasets/dh-unibe/image-text_aaeb-xiv-xvii-part-2-lines.jonkopings_radhusratt_och_magistrat_lineskrigshovrattens_dombocker_linestrolldomskommissionen_linesazerbaijani-ocr-lines
Azerbaijani OCR Lines
Line-level training data for Azerbaijani text recognition, in Latin and
Cyrillic script, extracted from scanned books.
Fields
field
description
image
cropped text line, grayscale, height 48 px
text
transcription
script
az_latin or az_cyrillic
book
anonymised source-book id
How it was built
Pages come from scanned PDFs that already carried an OCR text layer. Line
boxes were taken from that layer, rendered… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/azerbaijani-ocr-lines.gota_hovratt_linesimage-text_kurrent-xix-lines
image-text_kurrent-xix-lines
Line-level variant of dh-unibe/image-text_kurrent-xix at revision 20743e9f6982460b243e014091626da43e2b57f7. One row per transcribed line — the cropped line image and its transcription — so the material can be read without a PageXML parser or a line segmenter.
Lines
622298
Characters
26008983
Source
dh-unibe/image-text_kurrent-xix
Source revision
20743e9f6982460b243e014091626da43e2b57f7
What the source reported while being… See the full description on the dataset page: https://huggingface.co/datasets/dh-unibe/image-text_kurrent-xix-lines.ehri-danish-typewritten-lines
EHRI Danish Typewritten Lines
This dataset contains 1,007 cropped line images from 36 Danish typewritten pages in the EHRI Dataset. The source material consists of Danish diplomatic reports from the Second World War held by the Danish National Archives.
It is published as a ready-to-use line recognition evaluation dataset. It is typewritten material, not book or newspaper typesetting.
Dataset structure
The dataset has one test split because the upstream dataset… See the full description on the dataset page: https://huggingface.co/datasets/V4ldeLund/ehri-danish-typewritten-lines.hanse-kurrent-xvi-lines
Dataset Card for hanse-kurrent-xvi-lines
This dataset was created using pagexml-hf converter from Transkribus PageXML data.
The dataset contains transkriptions of the Minutes of the low German town assemblies (Niederdeutsche Städtetage) from the 16th Century from various archives. The Texts include middle low German and New High German Languages
Transcriptions were created according to the following guidelines: Forschungsstelle für die Geschichte der Hanse und des Ostseeraums… See the full description on the dataset page: https://huggingface.co/datasets/fgho/hanse-kurrent-xvi-lines.
