Team Ai
Datasetpublic

sitloboi2012/CMDS_Multimodal_Document

Dataset Card for Cyrillic Multimodel Document (CMDS) This is the dataset consists of 3789 pairs of images and text across 31 categories downloaded from the Bulgarian ministry of finance Dataset Summary This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Supported Tasks and Leaderboards Uses this dataset for downstream task like Document Classification, Image Classification or Text… See the full description on the dataset page: https://huggingface.co/datasets/sitloboi2012/CMDS_Multimodal_Document.

sourceHugging Faceapache-2.0updated 3y agoView on Hugging Face
0likes51downloads
Dataset Card

Dataset Card for Cyrillic Multimodel Document (CMDS)

This is the dataset consists of 3789 pairs of images and text across 31 categories downloaded from the Bulgarian ministry of finance

Dataset Summary

This dataset card aims to be a base template for new datasets. It has been generated using this raw template.

Supported Tasks and Leaderboards

Uses this dataset for downstream task like Document Classification, Image Classification or Text Classification (Sequences Classification). Suitable for multimodal Model like LayoutLm Family, Donut, etc.

Languages

Bulgarian

Data Fields

  • —_text_ (bytes): the text appear in the document
  • —_filename_ (str): the name of the file
  • —_image_ (PIL.Image): the image of the document
  • —_label_ (str): the label of the document. There are 31 differences labels