Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01docling-project /DocLayNet-v1.2 Dataset Card for DocLayNet v1.2 Dataset Summary This dataset is an extention of the original DocLayNet dataset which embeds the PDF files of the document images inside a binary column. DocLayNet provides page-by-page layout segmentation ground-truth using bounding-boxes for 11 distinct class labels on 80863 unique pages from 6 document categories. It provides several unique features compared to related work such as PubLayNet or DocBank: Human Annotation: DocLayNet is… See the full description on the dataset page: https://huggingface.co/datasets/docling-project/DocLayNet-v1.2.image10K<n<100K21 likes5.6k downloads2y agoHugging Face02docling-project /DocLayNet-v1.1 Dataset Card for DocLayNet v1.1 Dataset Summary DocLayNet provides page-by-page layout segmentation ground-truth using bounding-boxes for 11 distinct class labels on 80863 unique pages from 6 document categories. It provides several unique features compared to related work such as PubLayNet or DocBank: Human Annotation: DocLayNet is hand-annotated by well-trained experts, providing a gold-standard in layout segmentation through human recognition and interpretation of… See the full description on the dataset page: https://huggingface.co/datasets/docling-project/DocLayNet-v1.1.imageobject-detection10K<n<100K27 likes3.1k downloads3y agoHugging Face03vikp /doclaynet_processed Dataset Card for "doclaynet_processed" Clean version of DocLayNet ready for finetuning. image10K<n<100K6 likes1.2k downloads3y agoHugging Face04vikp /doclaynet_benchimage1K<n<10K3 likes871 downloads3y agoHugging Face05pierreguillou /DocLayNet-baseAccurate document layout analysis is a key requirement for high-quality PDF document conversion. With the recent availability of public, large ground-truth datasets such as PubLayNet and DocBank, deep-learning models have proven to be very effective at layout detection and segmentation. While these datasets are of adequate size to train such models, they severely lack in layout variability since they are sourced from scientific article repositories such as PubMed and arXiv only. Consequently, the accuracy of the layout segmentation drops significantly when these models are applied on more challenging and diverse layouts. In this paper, we present \textit{DocLayNet}, a new, publicly available, document-layout annotation dataset in COCO format. It contains 80863 manually annotated pages from diverse data sources to represent a wide variability in layouts. For each PDF page, the layout annotations provide labelled bounding-boxes with a choice of 11 distinct classes. DocLayNet also provides a subset of double- and triple-annotated pages to determine the inter-annotator agreement. In multiple experiments, we provide smallline accuracy scores (in mAP) for a set of popular object detection models. We also demonstrate that these models fall approximately 10\% behind the inter-annotator agreement. Furthermore, we provide evidence that DocLayNet is of sufficient size. Lastly, we compare models trained on PubLayNet, DocBank and DocLayNet, showing that layout predictions of the DocLayNet-trained models are more robust and thus the preferred choice for general-purpose document-layout analysis.imageobject-detection1K<n<10K19 likes745 downloads3y agoHugging Face06pierreguillou /DocLayNet-smallAccurate document layout analysis is a key requirement for high-quality PDF document conversion. With the recent availability of public, large ground-truth datasets such as PubLayNet and DocBank, deep-learning models have proven to be very effective at layout detection and segmentation. While these datasets are of adequate size to train such models, they severely lack in layout variability since they are sourced from scientific article repositories such as PubMed and arXiv only. Consequently, the accuracy of the layout segmentation drops significantly when these models are applied on more challenging and diverse layouts. In this paper, we present \textit{DocLayNet}, a new, publicly available, document-layout annotation dataset in COCO format. It contains 80863 manually annotated pages from diverse data sources to represent a wide variability in layouts. For each PDF page, the layout annotations provide labelled bounding-boxes with a choice of 11 distinct classes. DocLayNet also provides a subset of double- and triple-annotated pages to determine the inter-annotator agreement. In multiple experiments, we provide smallline accuracy scores (in mAP) for a set of popular object detection models. We also demonstrate that these models fall approximately 10\% behind the inter-annotator agreement. Furthermore, we provide evidence that DocLayNet is of sufficient size. Lastly, we compare models trained on PubLayNet, DocBank and DocLayNet, showing that layout predictions of the DocLayNet-trained models are more robust and thus the preferred choice for general-purpose document-layout analysis.imageobject-detectionn<1K13 likes435 downloads3y agoHugging Face07MingxuChai /DocLayNet_rankimage10K<n<100K0 likes267 downloads10mo agoHugging Face08docling-project /doclaynet-pt-enriched-formulaimage100K<n<1M2 likes193 downloads1y agoHugging Face09ahmedheakl /arocrbench_doclaynetPlease see paper & code for more information: https://github.com/mbzuai-oryx/KITAB-Bench https://arxiv.org/abs/2502.14949 imagen<1K1 likes56 downloads2y agoHugging Face10thewalnutaisg /Doclaynet-Full NOtice: The category Ids are not mapped btw 0-10 doclaynet classes, rather they are 2-12 Use the following Classes map. {'caption': 2, 'footnote': 3, 'formula': 4, 'list_item': 5, 'page_footer': 6, 'page_header': 7, 'picture': 8, 'section_header': 9, 'table': 10, 'text': 11, 'title': 12} dataset_info: config_name: all features: name: image, dtype: image name: category_ids, sequence: int32 name: image_id, dtype: int32 name: boxes, sequence:… See the full description on the dataset page: https://huggingface.co/datasets/thewalnutaisg/Doclaynet-Full.image10K<n<100K0 likes53 downloads2y agoHugging Face11katphlab /doclaynet-full Category IDs 1 - Caption 2 - Footnote 3 - Formula 4 - List Item 5 - Page Footer 6 - Page Header 7 - Picture 8 - Section Header 9 - Table 10 - Text 11 - Title image10K<n<100K0 likes35 downloads2y agoHugging Face12tasksource /doclaynet-region doclaynet-region Native DocLayNet v1.1 page regions, with 11 source-defined layout classes. CDLA-Permissive-1.0. Native train/val/test partitions are preserved (val is named validation). PDF text cells never enter the inputs. A uniform red box identifies the selected region. The original page/document identity and geometry remain in metadata; out-of-frame boxes are excluded, not clipped. Original data: docling-project/DocLayNet-v1.1, docling-project/DocLayNet. Repackaged as… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/doclaynet-region.image1K<n<10K0 likes34 downloads1d agoHugging Face13nevernever69 /small-DocLayNet-v1.1image1K<n<10K0 likes33 downloads2y agoHugging Face14nnul /DocLayNet-Instruct-v1-preprocessedimage10K<n<100K0 likes31 downloads1y agoHugging Face15hantian /doclaynet-for-yoloUsed by https://github.com/ppaanngggg/yolo-doclaynet imageobject-detection1 likes29 downloads1y agoHugging Face16nevernever69 /DocLayNet-Smallimage1K<n<10K0 likes26 downloads2y agoHugging Face17miikatoi /DocLayNet-tiny Dataset Card for "DocLayNet-tiny" Tiny set for unit tests based on https://huggingface.co/datasets/pierreguillou/DocLayNet-small. Total ~0.1% of DocLayNet. imagen<1K0 likes25 downloads3y agoHugging Face18Elliot-Data /doclaynet_train_cleanedgated doclaynet_train_cleaned The doclaynet_train family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning. images 54,199 QA turns 145,370 answers rewritten by the cleaning pass 0 QA created by the cleaning pass (new_qa) not measured for this family shards 43 How this was cleaned A vision-language model read each image together with its QA and judged the item. The pass is not a filter that only removes rows — it rewrites answers… See the full description on the dataset page: https://huggingface.co/datasets/Elliot-Data/doclaynet_train_cleaned.imagevisual-question-answering10K<n<100K0 likes25 downloads1mo agoHugging Face19alaperna /DocLayNet_ValidationJust the validation part of the DocLayNet dataset. The full dataset can be found at https://developer.ibm.com/exchanges/data/all/doclaynet/ image1K<n<10K0 likes23 downloads2y agoHugging Face20MathDG /DocLayNet-base-lawimage1K<n<10K0 likes19 downloads3y agoHugging Face21Epicguest97 /doclaynet10classesimage1K<n<10K0 likes19 downloads2y agoHugging Face22merve /doclaynet-smallimage1K<n<10K2 likes18 downloads1y agoHugging Face23ahmedheakl /arocrbench_doclaynetv2imagen<1K0 likes16 downloads2y agoHugging Face24nnul /DocLayNet-Instruct-v1image10K<n<100K0 likes14 downloads1y agoHugging Face25showgun01 /doclaynet-yoloimagen<1K0 likes13 downloads2y agoHugging Face26Epicguest97 /doclaynet3classesimagen<1K0 likes12 downloads2y agoHugging Face27vikp /doclaynet_mathimage1K<n<10K1 likes10 downloads3y agoHugging Face28huyhoangt2201 /table-extracted-yolo-data-doclaynet-zipimage0 likes9 downloads10mo agoHugging Face29elliot-mllm /doclaynet_train_cleanedgated doclaynet_train_cleaned The doclaynet_train family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning. images 54,199 QA turns 145,370 answers rewritten by the cleaning pass 0 QA created by the cleaning pass (new_qa) not measured for this family shards 43 How this was cleaned A vision-language model read each image together with its QA and judged the item. The pass is not a filter that only removes rows — it rewrites answers… See the full description on the dataset page: https://huggingface.co/datasets/elliot-mllm/doclaynet_train_cleaned.imagevisual-question-answering10K<n<100K0 likes9 downloads1mo agoHugging Face30Epicguest97 /Doclaynet3imagen<1K0 likes8 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.