Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01PiotrSty /ocr-pl-lines ocr-pl-lines Syntetyczny zbiór linii tekstu po polsku do fine-tuningu OCR (TrOCR). Pary NNNNN.png (obraz linii) + NNNNN.txt (transkrypcja). Struktura train/ — 2000 par (seed 42) val/ — 200 par (seed 123) Generowanie OCR_engine — python -m training.generate_synthetic Korpus: zdania potoczne i urzędowe, domeny (faktury, umowy, medyczne, prawnicze), losowe daty/kwoty/adresy/NIP/PESEL, zdania z pl.wikipedia.org. Augmentacje: pochylenie, blur, szum… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/ocr-pl-lines.image1K<n<10K0 likes1.9k downloads27d agoHugging Face02nader39 /lazy-lines-mediaimagen<1K0 likes912 downloads6d agoHugging Face03prasatee /lines-dataset P&ID Line Detection Dataset This dataset contains cropped images from P&ID (Piping and Instrumentation Diagrams) with line segment annotations for line detection and segmentation tasks. Dataset Splits The dataset is split by source image to prevent data leakage: Split Source Images Samples Line Segments Train 400 (80%) 8,000 69,483 Validation 50 (10%) 1,000 9,485 Test 50 (10%) 1,000 8,940 Dataset Structure lines_dataset/ ├── train/ │… See the full description on the dataset page: https://huggingface.co/datasets/prasatee/lines-dataset.imageimage-segmentation10K<n<100K0 likes877 downloads10mo agoHugging Face04prasatee /pid_lines_dataset P&ID Line Detection Dataset This dataset contains cropped images from P&ID (Piping and Instrumentation Diagrams) with line segment annotations for line detection and segmentation tasks. Dataset Structure Each sample contains: file_name: Image filename source_image_idx: Index of the original P&ID image crop_idx: Index of this crop from the source image width: Crop width in pixels height: Crop height in pixels lines: Dictionary with: segments: List of line segments as [x1… See the full description on the dataset page: https://huggingface.co/datasets/prasatee/pid_lines_dataset.imageimage-segmentation1K<n<10K1 likes853 downloads10mo agoHugging Face05PiotrSty /ehri-pl-lines ehri-pl-lines Line-level OCR/HTR dataset of Polish typewritten historical documents, derived from the Polish sub-corpus of the EHRI dataset. Each example is a single text-line crop plus its ground-truth transcription. Line bounding boxes come from the original ALTO XML ground truth (no automatic detection was used), so transcriptions are reliable and aligned. Content and provenance 468 line crops from 15 pages across 6 documents (ZIH collection). Source images… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/ehri-pl-lines.imageimage-to-textn<1K0 likes479 downloads26d agoHugging Face06samiksha9874 /pid_lines_dataset P&ID Line Detection Dataset This dataset contains cropped images from P&ID (Piping and Instrumentation Diagrams) with line segment annotations for line detection and segmentation tasks. Dataset Structure Each sample contains: file_name: Image filename source_image_idx: Index of the original P&ID image crop_idx: Index of this crop from the source image width: Crop width in pixels height: Crop height in pixels lines: Dictionary with: segments: List of line… See the full description on the dataset page: https://huggingface.co/datasets/samiksha9874/pid_lines_dataset.imageimage-segmentation1K<n<10K0 likes411 downloads4mo agoHugging Face07leitro /Copiale_Lines Copiale Lines Copiale Lines is a line-level image-to-text dataset for historical cipher decipherment. It contains cropped line images from the Copiale manuscript paired with plaintext ground truth. This dataset was presented in the paper Learning to Decipher from Pixels -- A Case Study of Copiale (HistoCrypt 2026). Code: https://github.com/leitro/Decipher-from-Pixels-Copiale Dataset Structure The dataset is split into: train: 1,269 samples valid: 175 samples… See the full description on the dataset page: https://huggingface.co/datasets/leitro/Copiale_Lines.imageimage-to-text1K<n<10K0 likes237 downloads5mo agoHugging Face08Riksarkivet /goteborgs_poliskammare_fore_1900_linesimage100K<n<1M1 likes198 downloads2y agoHugging Face09SubSpring /arams28k-htr-linesimage10K<n<100K1 likes177 downloads1mo agoHugging Face10samaritan-ai /hebrew_synth_linesimagetext-generation100K<n<1M1 likes169 downloads1y agoHugging Face11AlhitawiMohammed22 /En_IAM_linesimage10K<n<100K1 likes142 downloads4y agoHugging Face12Riksarkivet /eval_htr_out_of_domain_linesimage1K<n<10K1 likes138 downloads2y agoHugging Face13OttomanNLP /OpenITI-MAKHZAN-Ottoman-Lines Dataset Card for OpenITI MAKHZAN Ottoman Lines Dataset Summary This dataset contains line-level image-text pairs of historical Ottoman Turkish manuscripts and printed documents. It is derived from the OpenITI MAKHZAN dataset, a large aggregation of Arabic-script ground truth and evaluation data developed by the Open Islamicate Texts Initiative (OpenITI). The dataset specifically focuses on Ottoman Turkish texts and is highly valuable for training and evaluating… See the full description on the dataset page: https://huggingface.co/datasets/OttomanNLP/OpenITI-MAKHZAN-Ottoman-Lines.imageimage-to-text1K<n<10K1 likes68 downloads3mo agoHugging Face14Sri1311 /pid_lines_dataset P&ID Line Detection Dataset This dataset contains cropped images from P&ID (Piping and Instrumentation Diagrams) with line segment annotations for line detection and segmentation tasks. Dataset Structure Each sample contains: file_name: Image filename source_image_idx: Index of the original P&ID image crop_idx: Index of this crop from the source image width: Crop width in pixels height: Crop height in pixels lines: Dictionary with: segments: List of line… See the full description on the dataset page: https://huggingface.co/datasets/Sri1311/pid_lines_dataset.imageimage-segmentation10K<n<100K1 likes60 downloads3mo agoHugging Face15dh-unibe /image-text_bullinger-autoren-lines image-text_bullinger-autoren-lines Line-level variant of dh-unibe/image-text_bullinger-autoren at revision 85ea246a5021a5f244c392f97f2c8ac9a42cc717. One row per transcribed line — the cropped line image and its transcription — so the material can be read without a PageXML parser or a line segmenter. Lines 73502 Characters 3866274 Source dh-unibe/image-text_bullinger-autoren Source revision 85ea246a5021a5f244c392f97f2c8ac9a42cc717 What the source reported… See the full description on the dataset page: https://huggingface.co/datasets/dh-unibe/image-text_bullinger-autoren-lines.image10K<n<100K0 likes53 downloads3d agoHugging Face16Riksarkivet /bergskollegium_relationer_och_skrivelser_linesimage10K<n<100K0 likes50 downloads2y agoHugging Face17Riksarkivet /alvsborgs_losen_linesimage1K<n<10K1 likes49 downloads2y agoHugging Face18Riksarkivet /svea_hovratt_linesimage10K<n<100K2 likes49 downloads2y agoHugging Face19LuzianU /instance-synthetic-linesimage10K<n<100K0 likes46 downloads2y agoHugging Face20jwidmer /kurrent-hanse-xvi-test-lines-with-evaluation Dataset Card for evaluation-hf-printed This dataset is derived from danameyer/evaluation-hf-printed and has been enriched with inference results. Dataset Summary This dataset contains 164 samples across 1 split(s). Projects Included 1505-02-10_Hanserezess,Lübeck_Dienstag_nach_Scholastice_1505(SAHST_Rep__2,_I_040-4) Evaluation Results This repository includes evaluation artifacts for the latest inference output. Timestamp:… See the full description on the dataset page: https://huggingface.co/datasets/jwidmer/kurrent-hanse-xvi-test-lines-with-evaluation.imagen<1K0 likes44 downloads4mo agoHugging Face21cyttic /pinkas-hebrew-lines Pinkas — Hebrew cursive text lines (line crops) Line-level crops of The Pinkas Dataset: 30 pages of early-modern European Jewish community records (c. 1500–1800), written by non-professional scribes in Hebrew and Yiddish, with line-level ground truth in PAGE XML. Original dataset — all credit to its authors: Irina Rabaev, Berat Kurar Barakat, Jihad El-Sana. The Pinkas Dataset. Zenodo, 2019. https://doi.org/10.5281/zenodo.3569694 (record 3569694). License: CC-BY-4.0. This is a… See the full description on the dataset page: https://huggingface.co/datasets/cyttic/pinkas-hebrew-lines.imageimage-to-textn<1K0 likes43 downloads11d agoHugging Face22dh-unibe /image-text_aaeb-xiv-xvii-part-2-lines image-text_aaeb-xiv-xvii-part-2-lines Line-level variant of dh-unibe/image-text_aaeb-xiv-xvii-part-2 at revision ce2739fce41c0b8f8673127525839842411020b0. One row per transcribed line — the cropped line image and its transcription — so the material can be read without a PageXML parser or a line segmenter. Lines 3386 Characters 107134 Source dh-unibe/image-text_aaeb-xiv-xvii-part-2 Source revision ce2739fce41c0b8f8673127525839842411020b0 What the source… See the full description on the dataset page: https://huggingface.co/datasets/dh-unibe/image-text_aaeb-xiv-xvii-part-2-lines.image1K<n<10K0 likes42 downloads5d agoHugging Face23Riksarkivet /jonkopings_radhusratt_och_magistrat_linesimage1K<n<10K0 likes41 downloads2y agoHugging Face24Riksarkivet /krigshovrattens_dombocker_linesimage10K<n<100K0 likes40 downloads2y agoHugging Face25Riksarkivet /trolldomskommissionen_linesimage10K<n<100K0 likes40 downloads2y agoHugging Face26LocalDoc /azerbaijani-ocr-lines Azerbaijani OCR Lines Line-level training data for Azerbaijani text recognition, in Latin and Cyrillic script, extracted from scanned books. Fields field description image cropped text line, grayscale, height 48 px text transcription script az_latin or az_cyrillic book anonymised source-book id How it was built Pages come from scanned PDFs that already carried an OCR text layer. Line boxes were taken from that layer, rendered… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/azerbaijani-ocr-lines.imageimage-to-text100K<n<1M1 likes40 downloads1mo agoHugging Face27Riksarkivet /gota_hovratt_linesimage1K<n<10K0 likes39 downloads2y agoHugging Face28dh-unibe /image-text_kurrent-xix-lines image-text_kurrent-xix-lines Line-level variant of dh-unibe/image-text_kurrent-xix at revision 20743e9f6982460b243e014091626da43e2b57f7. One row per transcribed line — the cropped line image and its transcription — so the material can be read without a PageXML parser or a line segmenter. Lines 622298 Characters 26008983 Source dh-unibe/image-text_kurrent-xix Source revision 20743e9f6982460b243e014091626da43e2b57f7 What the source reported while being… See the full description on the dataset page: https://huggingface.co/datasets/dh-unibe/image-text_kurrent-xix-lines.image100K<n<1M0 likes39 downloads2d agoHugging Face29V4ldeLund /ehri-danish-typewritten-lines EHRI Danish Typewritten Lines This dataset contains 1,007 cropped line images from 36 Danish typewritten pages in the EHRI Dataset. The source material consists of Danish diplomatic reports from the Second World War held by the Danish National Archives. It is published as a ready-to-use line recognition evaluation dataset. It is typewritten material, not book or newspaper typesetting. Dataset structure The dataset has one test split because the upstream dataset… See the full description on the dataset page: https://huggingface.co/datasets/V4ldeLund/ehri-danish-typewritten-lines.imageimage-to-text1K<n<10K0 likes38 downloads1mo agoHugging Face30fgho /hanse-kurrent-xvi-lines Dataset Card for hanse-kurrent-xvi-lines This dataset was created using pagexml-hf converter from Transkribus PageXML data. The dataset contains transkriptions of the Minutes of the low German town assemblies (Niederdeutsche Städtetage) from the 16th Century from various archives. The Texts include middle low German and New High German Languages Transcriptions were created according to the following guidelines: Forschungsstelle für die Geschichte der Hanse und des Ostseeraums… See the full description on the dataset page: https://huggingface.co/datasets/fgho/hanse-kurrent-xvi-lines.image10K<n<100K0 likes36 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.