Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01docling-project /DocLayNet-v1.2 Dataset Card for DocLayNet v1.2 Dataset Summary This dataset is an extention of the original DocLayNet dataset which embeds the PDF files of the document images inside a binary column. DocLayNet provides page-by-page layout segmentation ground-truth using bounding-boxes for 11 distinct class labels on 80863 unique pages from 6 document categories. It provides several unique features compared to related work such as PubLayNet or DocBank: Human Annotation: DocLayNet is… See the full description on the dataset page: https://huggingface.co/datasets/docling-project/DocLayNet-v1.2.image10K<n<100K21 likes5.8k downloads2y agoHugging Face02docling-project /DocLayNet-v1.1 Dataset Card for DocLayNet v1.1 Dataset Summary DocLayNet provides page-by-page layout segmentation ground-truth using bounding-boxes for 11 distinct class labels on 80863 unique pages from 6 document categories. It provides several unique features compared to related work such as PubLayNet or DocBank: Human Annotation: DocLayNet is hand-annotated by well-trained experts, providing a gold-standard in layout segmentation through human recognition and interpretation of… See the full description on the dataset page: https://huggingface.co/datasets/docling-project/DocLayNet-v1.1.imageobject-detection10K<n<100K27 likes3.2k downloads3y agoHugging Face03vikp /doclaynet_processed Dataset Card for "doclaynet_processed" Clean version of DocLayNet ready for finetuning. image10K<n<100K6 likes1.5k downloads3y agoHugging Face04vikp /doclaynet_benchimage1K<n<10K3 likes870 downloads3y agoHugging Face05pierreguillou /DocLayNet-baseAccurate document layout analysis is a key requirement for high-quality PDF document conversion. With the recent availability of public, large ground-truth datasets such as PubLayNet and DocBank, deep-learning models have proven to be very effective at layout detection and segmentation. While these datasets are of adequate size to train such models, they severely lack in layout variability since they are sourced from scientific article repositories such as PubMed and arXiv only. Consequently, the accuracy of the layout segmentation drops significantly when these models are applied on more challenging and diverse layouts. In this paper, we present \textit{DocLayNet}, a new, publicly available, document-layout annotation dataset in COCO format. It contains 80863 manually annotated pages from diverse data sources to represent a wide variability in layouts. For each PDF page, the layout annotations provide labelled bounding-boxes with a choice of 11 distinct classes. DocLayNet also provides a subset of double- and triple-annotated pages to determine the inter-annotator agreement. In multiple experiments, we provide smallline accuracy scores (in mAP) for a set of popular object detection models. We also demonstrate that these models fall approximately 10\% behind the inter-annotator agreement. Furthermore, we provide evidence that DocLayNet is of sufficient size. Lastly, we compare models trained on PubLayNet, DocBank and DocLayNet, showing that layout predictions of the DocLayNet-trained models are more robust and thus the preferred choice for general-purpose document-layout analysis.imageobject-detection1K<n<10K19 likes773 downloads3y agoHugging Face06docling-project /DocLayNetDocLayNet is a human-annotated document layout segmentation dataset from a broad variety of document sources.object-detection10K<n<100K150 likes584 downloads4y agoHugging Face07pierreguillou /DocLayNet-smallAccurate document layout analysis is a key requirement for high-quality PDF document conversion. With the recent availability of public, large ground-truth datasets such as PubLayNet and DocBank, deep-learning models have proven to be very effective at layout detection and segmentation. While these datasets are of adequate size to train such models, they severely lack in layout variability since they are sourced from scientific article repositories such as PubMed and arXiv only. Consequently, the accuracy of the layout segmentation drops significantly when these models are applied on more challenging and diverse layouts. In this paper, we present \textit{DocLayNet}, a new, publicly available, document-layout annotation dataset in COCO format. It contains 80863 manually annotated pages from diverse data sources to represent a wide variability in layouts. For each PDF page, the layout annotations provide labelled bounding-boxes with a choice of 11 distinct classes. DocLayNet also provides a subset of double- and triple-annotated pages to determine the inter-annotator agreement. In multiple experiments, we provide smallline accuracy scores (in mAP) for a set of popular object detection models. We also demonstrate that these models fall approximately 10\% behind the inter-annotator agreement. Furthermore, we provide evidence that DocLayNet is of sufficient size. Lastly, we compare models trained on PubLayNet, DocBank and DocLayNet, showing that layout predictions of the DocLayNet-trained models are more robust and thus the preferred choice for general-purpose document-layout analysis.imageobject-detectionn<1K13 likes443 downloads3y agoHugging Face08MingxuChai /DocLayNet_rankimage10K<n<100K0 likes274 downloads10mo agoHugging Face09pierreguillou /DocLayNet-largeAccurate document layout analysis is a key requirement for high-quality PDF document conversion. With the recent availability of public, large ground-truth datasets such as PubLayNet and DocBank, deep-learning models have proven to be very effective at layout detection and segmentation. While these datasets are of adequate size to train such models, they severely lack in layout variability since they are sourced from scientific article repositories such as PubMed and arXiv only. Consequently, the accuracy of the layout segmentation drops significantly when these models are applied on more challenging and diverse layouts. In this paper, we present \textit{DocLayNet}, a new, publicly available, document-layout annotation dataset in COCO format. It contains 80863 manually annotated pages from diverse data sources to represent a wide variability in layouts. For each PDF page, the layout annotations provide labelled bounding-boxes with a choice of 11 distinct classes. DocLayNet also provides a subset of double- and triple-annotated pages to determine the inter-annotator agreement. In multiple experiments, we provide largeline accuracy scores (in mAP) for a set of popular object detection models. We also demonstrate that these models fall approximately 10\% behind the inter-annotator agreement. Furthermore, we provide evidence that DocLayNet is of sufficient size. Lastly, we compare models trained on PubLayNet, DocBank and DocLayNet, showing that layout predictions of the DocLayNet-trained models are more robust and thus the preferred choice for general-purpose document-layout analysis.object-detection10K<n<100K14 likes249 downloads3y agoHugging Face10docling-project /doclaynet-pt-enriched-formulaimage100K<n<1M2 likes225 downloads11mo agoHugging Face11operant-ai /doclaynet-document-level DocLayNet Document-Level Reconstruction and 8K Expansion This dataset is a normalized, one-row-per-document view over the page-level DocLayNet v1.1 dataset. Pages are grouped using DocLayNet's source metadata and ordered by their original page number. Dataset summary 2,944 logical documents 80,863 observed pages 896 complete document groups 2,048 partial document groups Train: 2,355 documents / 60,810 pages Validation: 294 documents / 7,964 pages Test: 295… See the full description on the dataset page: https://huggingface.co/datasets/operant-ai/doclaynet-document-level.tabular10K<n<100K1 likes155 downloads27d agoHugging Face12PassionPrc /doclaynet-grpotext100K<n<1M0 likes125 downloads4mo agoHugging Face13thewalnutaisg /Doclaynet-Full NOtice: The category Ids are not mapped btw 0-10 doclaynet classes, rather they are 2-12 Use the following Classes map. {'caption': 2, 'footnote': 3, 'formula': 4, 'list_item': 5, 'page_footer': 6, 'page_header': 7, 'picture': 8, 'section_header': 9, 'table': 10, 'text': 11, 'title': 12} dataset_info: config_name: all features: name: image, dtype: image name: category_ids, sequence: int32 name: image_id, dtype: int32 name: boxes, sequence:… See the full description on the dataset page: https://huggingface.co/datasets/thewalnutaisg/Doclaynet-Full.image10K<n<100K0 likes61 downloads2y agoHugging Face14anonymous-iccv1968 /DocLayout_YOLO_DocLayNet_iccv1968text10K<n<100K0 likes55 downloads2y agoHugging Face15Jaki-73 /DocLayNetDocLayNet is a human-annotated document layout segmentation dataset from a broad variety of document sources.object-detection10K<n<100K0 likes53 downloads28d agoHugging Face16ahmedheakl /arocrbench_doclaynetPlease see paper & code for more information: https://github.com/mbzuai-oryx/KITAB-Bench https://arxiv.org/abs/2502.14949 imagen<1K1 likes52 downloads2y agoHugging Face17docling-project /icdar2023-doclaynetDataset for the ICDAR 2023 Competition on Robust Layout Segmentation in Corporate Documents.object-detectionn<1K4 likes40 downloads4y agoHugging Face18juliozhao /doclayout-yolo-DocLayNethttps://huggingface.co/papers/2410.12628 7 likes39 downloads2y agoHugging Face19nevernever69 /small-DocLayNet-v1.1image1K<n<10K0 likes39 downloads2y agoHugging Face20Elliot-Data /doclaynet_train_cleanedgated doclaynet_train_cleaned The doclaynet_train family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning. images 54,199 QA turns 145,370 answers rewritten by the cleaning pass 0 QA created by the cleaning pass (new_qa) not measured for this family shards 43 How this was cleaned A vision-language model read each image together with its QA and judged the item. The pass is not a filter that only removes rows — it rewrites answers… See the full description on the dataset page: https://huggingface.co/datasets/Elliot-Data/doclaynet_train_cleaned.imagevisual-question-answering10K<n<100K0 likes37 downloads29d agoHugging Face21agomberto /DoCLayNet-large-wt-image Dataset Card for DocLayNet large without image About this card (02/14/2024) Property and license All information from this page but the content of this paragraph "About this card (02/14/2025)" has been copied/pasted from Dataset Card for DocLayNet. DocLayNet is a dataset created by Deep Search (IBM Research) published under license CDLA-Permissive-1.0. I do not claim any rights to the data taken from this dataset and published on this page. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/agomberto/DoCLayNet-large-wt-image.tabularobject-detection10K<n<100K0 likes34 downloads3y agoHugging Face22katphlab /doclaynet-full Category IDs 1 - Caption 2 - Footnote 3 - Formula 4 - List Item 5 - Page Footer 6 - Page Header 7 - Picture 8 - Section Header 9 - Table 10 - Text 11 - Title image10K<n<100K0 likes34 downloads2y agoHugging Face23nnul /DocLayNet-Instruct-v1-preprocessedimage10K<n<100K0 likes33 downloads1y agoHugging Face24miikatoi /DocLayNet-tiny Dataset Card for "DocLayNet-tiny" Tiny set for unit tests based on https://huggingface.co/datasets/pierreguillou/DocLayNet-small. Total ~0.1% of DocLayNet. imagen<1K0 likes30 downloads3y agoHugging Face25nevernever69 /DocLayNet-Smallimage1K<n<10K0 likes29 downloads2y agoHugging Face26hantian /doclaynet-for-yoloUsed by https://github.com/ppaanngggg/yolo-doclaynet imageobject-detection1 likes29 downloads1y agoHugging Face27MathDG /DocLayNet-base-lawimage1K<n<10K0 likes25 downloads3y agoHugging Face28alaperna /DocLayNet_ValidationJust the validation part of the DocLayNet dataset. The full dataset can be found at https://developer.ibm.com/exchanges/data/all/doclaynet/ image1K<n<10K0 likes25 downloads1y agoHugging Face29kbang2021 /doclaynet-6class DocLayNet 6-Class Filtered Dataset Dataset Description This is a filtered version of the DocLayNet dataset containing only 6 most relevant layout element classes for document layout analysis tasks. Original Dataset DocLayNet is a human-annotated document layout segmentation dataset containing 80,863 pages from diverse sources with 11 distinct layout categories. Citation: @article{doclaynet2022, title = {DocLayNet: A Large Human-Annotated Dataset for… See the full description on the dataset page: https://huggingface.co/datasets/kbang2021/doclaynet-6class.object-detection10K<n<100K0 likes22 downloads7mo agoHugging Face30Epicguest97 /doclaynet10classesimage1K<n<10K0 likes20 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.