Team Ai
Datasetpublic

rishitdagli/cppe-5

Dataset Card for CPPE - 5 Dataset Summary CPPE - 5 (Medical Personal Protective Equipment) is a new challenging dataset with the goal to allow the study of subordinate categorization of medical personal protective equipments, which is not possible with other popular data sets that focus on broad level categories. Some features of this dataset are: high quality images and annotations (~4.6 bounding boxes per image) real-life images unlike any current such dataset… See the full description on the dataset page: https://huggingface.co/datasets/rishitdagli/cppe-5.

sourceHugging Faceunknownupdated 3y agoView on Hugging Face
23likes3.4kdownloads
README.md265 linesDownload Raw Back to root
1---2annotations_creators:3- crowdsourced4language_creators:5- found6language:7- en8license:9- unknown10multilinguality:11- monolingual12size_categories:13- 1K<n<10K14source_datasets:15- original16task_categories:17- object-detection18task_ids: []19paperswithcode_id: cppe-520pretty_name: CPPE - 521tags:22- medical-personal-protective-equipment-detection23dataset_info:24  features:25  - name: image_id26    dtype: int6427  - name: image28    dtype: image29  - name: width30    dtype: int3231  - name: height32    dtype: int3233  - name: objects34    sequence:35    - name: id36      dtype: int6437    - name: area38      dtype: int6439    - name: bbox40      sequence: float3241      length: 442    - name: category43      dtype:44        class_label:45          names:46            '0': Coverall47            '1': Face_Shield48            '2': Gloves49            '3': Goggles50            '4': Mask51  splits:52  - name: train53    num_bytes: 240463364.054    num_examples: 100055  - name: test56    num_bytes: 4172164.057    num_examples: 2958  download_size: 24115265359  dataset_size: 244635528.060configs:61- config_name: default62  data_files:63  - split: train64    path: data/train-*65  - split: test66    path: data/test-*67---68 69# Dataset Card for CPPE - 570 71## Table of Contents72- [Table of Contents](#table-of-contents)73- [Dataset Description](#dataset-description)74  - [Dataset Summary](#dataset-summary)75  - [Supported Tasks and Leaderboards](#supported-tasks-and-leaderboards)76  - [Languages](#languages)77- [Dataset Structure](#dataset-structure)78  - [Data Instances](#data-instances)79  - [Data Fields](#data-fields)80  - [Data Splits](#data-splits)81- [Dataset Creation](#dataset-creation)82  - [Curation Rationale](#curation-rationale)83  - [Source Data](#source-data)84  - [Annotations](#annotations)85  - [Personal and Sensitive Information](#personal-and-sensitive-information)86- [Considerations for Using the Data](#considerations-for-using-the-data)87  - [Social Impact of Dataset](#social-impact-of-dataset)88  - [Discussion of Biases](#discussion-of-biases)89  - [Other Known Limitations](#other-known-limitations)90- [Additional Information](#additional-information)91  - [Dataset Curators](#dataset-curators)92  - [Licensing Information](#licensing-information)93  - [Citation Information](#citation-information)94  - [Contributions](#contributions)95 96## Dataset Description97 98- **Homepage:**99- **Repository:** https://github.com/Rishit-dagli/CPPE-Dataset100- **Paper:** [CPPE-5: Medical Personal Protective Equipment Dataset](https://arxiv.org/abs/2112.09569)101- **Leaderboard:** https://paperswithcode.com/sota/object-detection-on-cppe-5102- **Point of Contact:** rishit.dagli@gmail.com103 104### Dataset Summary105 106CPPE - 5 (Medical Personal Protective Equipment) is a new challenging dataset with the goal to allow the study of subordinate categorization of medical personal protective equipments, which is not possible with other popular data sets that focus on broad level categories.107 108Some features of this dataset are:109 110* high quality images and annotations (~4.6 bounding boxes per image)111* real-life images unlike any current such dataset112* majority of non-iconic images (allowing easy deployment to real-world environments)113 114### Supported Tasks and Leaderboards115 116- `object-detection`: The dataset can be used to train a model for Object Detection. This task has an active leaderboard which can be found at https://paperswithcode.com/sota/object-detection-on-cppe-5. The metrics for this task are adopted from the COCO detection evaluation criteria, and include the mean Average Precision (AP) across IoU thresholds ranging from 0.50 to 0.95 at different scales.117 118### Languages119 120English121 122## Dataset Structure123 124### Data Instances125 126A data point comprises an image and its object annotations.127 128```129{130  'image_id': 15,131  'image': <PIL.JpegImagePlugin.JpegImageFile image mode=RGB size=943x663 at 0x2373B065C18>,132  'width': 943,133  'height': 663,134  'objects': {135    'id': [114, 115, 116, 117], 136    'area': [3796, 1596, 152768, 81002],137    'bbox': [138      [302.0, 109.0, 73.0, 52.0],139      [810.0, 100.0, 57.0, 28.0],140      [160.0, 31.0, 248.0, 616.0],141      [741.0, 68.0, 202.0, 401.0]142    ], 143    'category': [4, 4, 0, 0]144  }145}146```147 148### Data Fields149 150- `image`: the image id151- `image`: `PIL.Image.Image` object containing the image. Note that when accessing the image column: `dataset[0]["image"]` the image file is automatically decoded. Decoding of a large number of image files might take a significant amount of time. Thus it is important to first query the sample index before the `"image"` column, *i.e.* `dataset[0]["image"]` should **always** be preferred over `dataset["image"][0]`152- `width`: the image width153- `height`: the image height154- `objects`: a dictionary containing bounding box metadata for the objects present on the image155  - `id`: the annotation id156  - `area`: the area of the bounding box157  - `bbox`: the object's bounding box (in the [coco](https://albumentations.ai/docs/getting_started/bounding_boxes_augmentation/#coco) format)158  - `category`: the object's category, with possible values including `Coverall` (0),`Face_Shield` (1),`Gloves` (2),`Goggles` (3) and `Mask` (4)159 160### Data Splits161 162The data is split into training and testing set. The training set contains 1000 images and test set 29 images.163 164## Dataset Creation165 166### Curation Rationale167 168From the paper:169> With CPPE-5 dataset, we hope to facilitate research and use in applications at multiple public places to autonomously identify if a PPE (Personal Protective Equipment) kit has been worn and also which part of the PPE kit has been worn. One of the main aims with this dataset was to also capture a higher ratio of non-iconic images or non-canonical perspectives [5] of the objects in this dataset. We further hope to see high use of this dataset to aid in medical scenarios which would have a huge effect170worldwide.171 172### Source Data173 174#### Initial Data Collection and Normalization175 176The images in the CPPE-5 dataset were collected using the following process:177* Obtain Images from Flickr: Following the object categories we identified earlier, we first download images from Flickr and save them at the "Original" size. On Flickr, images are served at multiple different sizes (Square 75, Small 240, Large 1024, X-Large 4K etc.), the "Original" size is an exact copy of the image uploaded by author.178* Extract relevant metadata: Flickr contains images each with searchable metadata, we extract the following relevant179metadata:180  * A direct link to the original image on Flickr181  * Width and height of the image182  * Title given to the image by the author183  * Date and time the image was uploaded on184  * Flickr username of the author of the image185  * Flickr Name of the author of the image186  * Flickr profile of the author of the image187  * The License image is licensed under188  * MD5 hash of the original image189* Obtain Images from Google Images: Due to the reasons we mention earlier, we only collect a very small proportion190of images from Google Images. For these set of images we extract the following metadata:191  * A direct link to the original image192  * Width and height of the image193  * MD5 hash of the original image194* Filter inappropriate images: Though very rare in the collected images, we also remove images containing inappropriate content using the safety filters on Flickr and Google Safe Search.195* Filter near-similar images: We then remove near-duplicate images in the dataset using GIST descriptors196 197#### Who are the source language producers?198 199The images for this dataset were collected from Flickr and Google Images.200 201### Annotations202 203#### Annotation process204 205The dataset was labelled in two phases: the first phase included labelling 416 images and the second phase included labelling 613 images. For all the images in the dataset volunteers were provided the following table:206 207|Item        |Description                                                              |208|------------|---------------------------------------------------------------------    | 209|coveralls | Coveralls are hospital gowns worn by medical professionals as in order to provide a barrier between patient and professional, these usually cover most of the exposed skin surfaces of the professional medics.|210|mask | Mask prevents airborne transmission of infections between patients and/or treating personnel by blocking the movement of pathogens (primarily bacteria and viruses) shed in respiratory droplets and aerosols into and from the wearer’s mouth and nose.|211face shield | Face shield aims to protect the wearer’s entire face (or part of it) from hazards such as flying objects and road debris, chemical splashes (in laboratories or in industry), or potentially infectious materials (in medical and laboratory environments).|212gloves | Gloves are used during medical examinations and procedures to help prevent cross-contamination between caregivers and patients.|213|goggles | Goggles, or safety glasses, are forms of protective eye wear that usually enclose or protect the area surrounding the eye in order to prevent particulates, water or chemicals from striking the eyes.|214 215as well as examples of: correctly labelled images, incorrectly labelled images, and not applicable images. Before the labelling task, each volunteer was provided with an exercise to verify if the volunteer was able to correctly identify categories as well as identify if an annotated image is correctly labelled, incorrectly labelled, or not applicable. The labelling process first involved two volunteers independently labelling an image from the dataset. In any of the cases that: the number of bounding boxes are different, the labels for on or more of the bounding boxes are different or two volunteer annotations are sufficiently different; a third volunteer compiles the result from the two annotations to come up with a correctly labelled image. After this step, a volunteer verifies the bounding box annotations. Following this method of labelling the dataset we ensured that all images were labelled accurately and contained exhaustive216annotations. As a result of this, our dataset consists of 1029 high-quality, majorly non-iconic, and accurately annotated images.217 218#### Who are the annotators?219 220In both the phases crowd-sourcing techniques were used with multiple volunteers labelling the dataset using the open-source tool LabelImg.221 222### Personal and Sensitive Information223 224[More Information Needed]225 226## Considerations for Using the Data227 228### Social Impact of Dataset229 230[More Information Needed]231 232### Discussion of Biases233 234[More Information Needed]235 236### Other Known Limitations237 238[More Information Needed]239 240## Additional Information241 242### Dataset Curators243 244Dagli, Rishit, and Ali Mustufa Shaikh.245 246### Licensing Information247 248[More Information Needed]249 250### Citation Information251 252```253@misc{dagli2021cppe5,254      title={CPPE-5: Medical Personal Protective Equipment Dataset},255      author={Rishit Dagli and Ali Mustufa Shaikh},256      year={2021},257      eprint={2112.09569},258      archivePrefix={arXiv},259      primaryClass={cs.CV}260}261```262 263### Contributions264 265Thanks to [@mariosasko](https://github.com/mariosasko) for adding this dataset.