Team Ai
Datasetpublic

prithivMLmods/OpenDoc-Pdf-Preview

OpenDoc-Pdf-Preview OpenDoc-Pdf-Preview is a compact visual preview dataset containing 6,000 high-resolution document images extracted from PDFs. This dataset is designed for Image-to-Text tasks such as document OCR pretraining, layout understanding, and multimodal document analysis. Dataset Summary Modality: Image-to-Text Content Type: PDF-based document previews Number of Samples: 6,000 Language: English Format: Parquet Split: train only Size: 606 MB License:… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/OpenDoc-Pdf-Preview.

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
2likes233downloads

prithivMLmods/OpenDoc-Pdf-Preview · main · files are served by the source, never re-hosted here