lines
Datasets
All datasets matching “lines”c4-pt-randMore35M-part04-deduplicated-128000-no-digit-split-mask-train-15003771-lines
Dataset Card for "c4-pt-randMore35M-part04-deduplicated-128000-no-digit-split-mask-train-15003771-lines"
More Information needed
ocr-pl-lines
ocr-pl-lines
Syntetyczny zbiór linii tekstu po polsku do fine-tuningu OCR (TrOCR).
Pary NNNNN.png (obraz linii) + NNNNN.txt (transkrypcja).
Struktura
train/ — 2000 par (seed 42)
val/ — 200 par (seed 123)
Generowanie
OCR_engine —
python -m training.generate_synthetic
Korpus: zdania potoczne i urzędowe, domeny (faktury, umowy, medyczne,
prawnicze), losowe daty/kwoty/adresy/NIP/PESEL, zdania z pl.wikipedia.org.
Augmentacje: pochylenie, blur, szum… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/ocr-pl-lines.simpsons_script_lineslazy-lines-medialines-dataset
P&ID Line Detection Dataset
This dataset contains cropped images from P&ID (Piping and Instrumentation Diagrams)
with line segment annotations for line detection and segmentation tasks.
Dataset Splits
The dataset is split by source image to prevent data leakage:
Split
Source Images
Samples
Line Segments
Train
400 (80%)
8,000
69,483
Validation
50 (10%)
1,000
9,485
Test
50 (10%)
1,000
8,940
Dataset Structure
lines_dataset/
├── train/
│… See the full description on the dataset page: https://huggingface.co/datasets/prasatee/lines-dataset.jam-alt-lines
