Team Ai
Apppublic

Arulkumar03/Wheat_HEAD_Detection_Counting_ComputerVision_Model

sourceHugging Facemitupdated 3y agoView on Hugging Face
0likes
datasets.md291 linesDownload Raw Back to tutorials
1# Use Custom Datasets2 3This document explains how the dataset APIs4([DatasetCatalog](../modules/data.html#detectron2.data.DatasetCatalog), [MetadataCatalog](../modules/data.html#detectron2.data.MetadataCatalog))5work, and how to use them to add custom datasets.6 7Datasets that have builtin support in detectron2 are listed in [builtin datasets](builtin_datasets.md).8If you want to use a custom dataset while also reusing detectron2's data loaders,9you will need to:10 111. __Register__ your dataset (i.e., tell detectron2 how to obtain your dataset).122. Optionally, __register metadata__ for your dataset.13 14Next, we explain the above two concepts in detail.15 16The [Colab tutorial](https://colab.research.google.com/drive/16jcaJoc6bCFAQ96jDe2HwtXj7BMD_-m5)17has a live example of how to register and train on a dataset of custom formats.18 19### Register a Dataset20 21To let detectron2 know how to obtain a dataset named "my_dataset", users need to implement22a function that returns the items in your dataset and then tell detectron2 about this23function:24```python25def my_dataset_function():26  ...27  return list[dict] in the following format28 29from detectron2.data import DatasetCatalog30DatasetCatalog.register("my_dataset", my_dataset_function)31# later, to access the data:32data: List[Dict] = DatasetCatalog.get("my_dataset")33```34 35Here, the snippet associates a dataset named "my_dataset" with a function that returns the data.36The function must return the same data (with same order) if called multiple times.37The registration stays effective until the process exits.38 39The function can do arbitrary things and should return the data in `list[dict]`, each dict in either40of the following formats:411. Detectron2's standard dataset dict, described below. This will make it work with many other builtin42   features in detectron2, so it's recommended to use it when it's sufficient.432. Any custom format. You can also return arbitrary dicts in your own format,44   such as adding extra keys for new tasks.45   Then you will need to handle them properly downstream as well.46   See below for more details.47 48#### Standard Dataset Dicts49 50For standard tasks51(instance detection, instance/semantic/panoptic segmentation, keypoint detection),52we load the original dataset into `list[dict]` with a specification similar to COCO's annotations.53This is our standard representation for a dataset.54 55Each dict contains information about one image.56The dict may have the following fields,57and the required fields vary based on what the dataloader or the task needs (see more below).58 59```eval_rst60.. list-table::61  :header-rows: 162 63  * - Task64    - Fields65  * - Common66    - file_name, height, width, image_id67 68  * - Instance detection/segmentation69    - annotations70 71  * - Semantic segmentation72    - sem_seg_file_name73 74  * - Panoptic segmentation75    - pan_seg_file_name, segments_info76```77 78+ `file_name`: the full path to the image file.79+ `height`, `width`: integer. The shape of the image.80+ `image_id` (str or int): a unique id that identifies this image. Required by many81  evaluators to identify the images, but a dataset may use it for different purposes.82+ `annotations` (list[dict]): Required by __instance detection/segmentation or keypoint detection__ tasks.83  Each dict corresponds to annotations of one instance in this image, and84  may contain the following keys:85  + `bbox` (list[float], required): list of 4 numbers representing the bounding box of the instance.86  + `bbox_mode` (int, required): the format of bbox.  It must be a member of87    [structures.BoxMode](../modules/structures.html#detectron2.structures.BoxMode).88    Currently supports: `BoxMode.XYXY_ABS`, `BoxMode.XYWH_ABS`.89  + `category_id` (int, required): an integer in the range [0, num_categories-1] representing the category label.90    The value num_categories is reserved to represent the "background" category, if applicable.91  + `segmentation` (list[list[float]] or dict): the segmentation mask of the instance.92    + If `list[list[float]]`, it represents a list of polygons, one for each connected component93      of the object. Each `list[float]` is one simple polygon in the format of `[x1, y1, ..., xn, yn]` (n≥3).94      The Xs and Ys are absolute coordinates in unit of pixels.95    + If `dict`, it represents the per-pixel segmentation mask in COCO's compressed RLE format.96      The dict should have keys "size" and "counts". You can convert a uint8 segmentation mask of 0s and97      1s into such dict by `pycocotools.mask.encode(np.asarray(mask, order="F"))`.98      `cfg.INPUT.MASK_FORMAT` must be set to `bitmask` if using the default data loader with such format.99  + `keypoints` (list[float]): in the format of [x1, y1, v1,..., xn, yn, vn].100    v[i] means the [visibility](http://cocodataset.org/#format-data) of this keypoint.101    `n` must be equal to the number of keypoint categories.102    The Xs and Ys are absolute real-value coordinates in range [0, W or H].103 104    (Note that the keypoint coordinates in COCO format are integers in range [0, W-1 or H-1], which is different105    from our standard format. Detectron2 adds 0.5 to COCO keypoint coordinates to convert them from discrete106    pixel indices to floating point coordinates.)107  + `iscrowd`: 0 (default) or 1. Whether this instance is labeled as COCO's "crowd108    region". Don't include this field if you don't know what it means.109 110  If `annotations` is an empty list, it means the image is labeled to have no objects.111  Such images will by default be removed from training,112  but can be included using `DATALOADER.FILTER_EMPTY_ANNOTATIONS`.113 114+ `sem_seg_file_name` (str):115  The full path to the semantic segmentation ground truth file.116  It should be a grayscale image whose pixel values are integer labels.117+ `pan_seg_file_name` (str):118  The full path to panoptic segmentation ground truth file.119  It should be an RGB image whose pixel values are integer ids encoded using the120  [panopticapi.utils.id2rgb](https://github.com/cocodataset/panopticapi/) function.121  The ids are defined by `segments_info`.122  If an id does not appear in `segments_info`, the pixel is considered unlabeled123  and is usually ignored in training & evaluation.124+ `segments_info` (list[dict]): defines the meaning of each id in panoptic segmentation ground truth.125  Each dict has the following keys:126  + `id` (int): integer that appears in the ground truth image.127  + `category_id` (int): an integer in the range [0, num_categories-1] representing the category label.128  + `iscrowd`: 0 (default) or 1. Whether this instance is labeled as COCO's "crowd region".129 130 131```eval_rst132 133.. note::134 135   The PanopticFPN model does not use the panoptic segmentation136   format defined here, but a combination of both instance segmentation and semantic segmentation data137   format. See :doc:`builtin_datasets` for instructions on COCO.138 139```140 141Fast R-CNN (with pre-computed proposals) models are rarely used today.142To train a Fast R-CNN, the following extra keys are needed:143 144+ `proposal_boxes` (array): 2D numpy array with shape (K, 4) representing K precomputed proposal boxes for this image.145+ `proposal_objectness_logits` (array): numpy array with shape (K, ), which corresponds to the objectness146 logits of proposals in 'proposal_boxes'.147+ `proposal_bbox_mode` (int): the format of the precomputed proposal bbox.148 It must be a member of149 [structures.BoxMode](../modules/structures.html#detectron2.structures.BoxMode).150 Default is `BoxMode.XYXY_ABS`.151 152 153 154#### Custom Dataset Dicts for New Tasks155 156In the `list[dict]` that your dataset function returns, the dictionary can also have __arbitrary custom data__.157This will be useful for a new task that needs extra information not covered158by the standard dataset dicts. In this case, you need to make sure the downstream code can handle your data159correctly. Usually this requires writing a new `mapper` for the dataloader (see [Use Custom Dataloaders](./data_loading.md)).160 161When designing a custom format, note that all dicts are stored in memory162(sometimes serialized and with multiple copies).163To save memory, each dict is meant to contain __small__ but sufficient information164about each sample, such as file names and annotations.165Loading full samples typically happens in the data loader.166 167For attributes shared among the entire dataset, use `Metadata` (see below).168To avoid extra memory, do not save such information inside each sample.169 170### "Metadata" for Datasets171 172Each dataset is associated with some metadata, accessible through173`MetadataCatalog.get(dataset_name).some_metadata`.174Metadata is a key-value mapping that contains information that's shared among175the entire dataset, and usually is used to interpret what's in the dataset, e.g.,176names of classes, colors of classes, root of files, etc.177This information will be useful for augmentation, evaluation, visualization, logging, etc.178The structure of metadata depends on what is needed from the corresponding downstream code.179 180If you register a new dataset through `DatasetCatalog.register`,181you may also want to add its corresponding metadata through182`MetadataCatalog.get(dataset_name).some_key = some_value`, to enable any features that need the metadata.183You can do it like this (using the metadata key "thing_classes" as an example):184 185```python186from detectron2.data import MetadataCatalog187MetadataCatalog.get("my_dataset").thing_classes = ["person", "dog"]188```189 190Here is a list of metadata keys that are used by builtin features in detectron2.191If you add your own dataset without these metadata, some features may be192unavailable to you:193 194* `thing_classes` (list[str]): Used by all instance detection/segmentation tasks.195  A list of names for each instance/thing category.196  If you load a COCO format dataset, it will be automatically set by the function `load_coco_json`.197 198* `thing_colors` (list[tuple(r, g, b)]): Pre-defined color (in [0, 255]) for each thing category.199  Used for visualization. If not given, random colors will be used.200 201* `stuff_classes` (list[str]): Used by semantic and panoptic segmentation tasks.202  A list of names for each stuff category.203 204* `stuff_colors` (list[tuple(r, g, b)]): Pre-defined color (in [0, 255]) for each stuff category.205  Used for visualization. If not given, random colors are used.206 207* `ignore_label` (int): Used by semantic and panoptic segmentation tasks. Pixels in ground-truth208  annotations with this category label should be ignored in evaluation. Typically these are "unlabeled"209  pixels.210 211* `keypoint_names` (list[str]): Used by keypoint detection. A list of names for each keypoint.212 213* `keypoint_flip_map` (list[tuple[str]]): Used by keypoint detection. A list of pairs of names,214  where each pair are the two keypoints that should be flipped if the image is215  flipped horizontally during augmentation.216* `keypoint_connection_rules`: list[tuple(str, str, (r, g, b))]. Each tuple specifies a pair of keypoints217  that are connected and the color (in [0, 255]) to use for the line between them when visualized.218 219Some additional metadata that are specific to the evaluation of certain datasets (e.g. COCO):220 221* `thing_dataset_id_to_contiguous_id` (dict[int->int]): Used by all instance detection/segmentation tasks in the COCO format.222  A mapping from instance class ids in the dataset to contiguous ids in range [0, #class).223  Will be automatically set by the function `load_coco_json`.224 225* `stuff_dataset_id_to_contiguous_id` (dict[int->int]): Used when generating prediction json files for226  semantic/panoptic segmentation.227  A mapping from semantic segmentation class ids in the dataset228  to contiguous ids in [0, num_categories). It is useful for evaluation only.229 230* `json_file`: The COCO annotation json file. Used by COCO evaluation for COCO-format datasets.231* `panoptic_root`, `panoptic_json`: Used by COCO-format panoptic evaluation.232* `evaluator_type`: Used by the builtin main training script to select233   evaluator. Don't use it in a new training script.234   You can just provide the [DatasetEvaluator](../modules/evaluation.html#detectron2.evaluation.DatasetEvaluator)235   for your dataset directly in your main script.236 237```eval_rst238.. note::239 240   In recognition, sometimes we use the term "thing" for instance-level tasks,241   and "stuff" for semantic segmentation tasks.242   Both are used in panoptic segmentation tasks.243   For background on the concept of "thing" and "stuff", see244   `On Seeing Stuff: The Perception of Materials by Humans and Machines245   <http://persci.mit.edu/pub_pdfs/adelson_spie_01.pdf>`_.246```247 248### Register a COCO Format Dataset249 250If your instance-level (detection, segmentation, keypoint) dataset is already a json file in the COCO format,251the dataset and its associated metadata can be registered easily with:252```python253from detectron2.data.datasets import register_coco_instances254register_coco_instances("my_dataset", {}, "json_annotation.json", "path/to/image/dir")255```256 257If your dataset is in COCO format but need to be further processed, or has extra custom per-instance annotations,258the [load_coco_json](../modules/data.html#detectron2.data.datasets.load_coco_json)259function might be useful.260 261### Update the Config for New Datasets262 263Once you've registered the dataset, you can use the name of the dataset (e.g., "my_dataset" in264example above) in `cfg.DATASETS.{TRAIN,TEST}`.265There are other configs you might want to change to train or evaluate on new datasets:266 267* `MODEL.ROI_HEADS.NUM_CLASSES` and `MODEL.RETINANET.NUM_CLASSES` are the number of thing classes268  for R-CNN and RetinaNet models, respectively.269* `MODEL.ROI_KEYPOINT_HEAD.NUM_KEYPOINTS` sets the number of keypoints for Keypoint R-CNN.270  You'll also need to set [Keypoint OKS](http://cocodataset.org/#keypoints-eval)271  with `TEST.KEYPOINT_OKS_SIGMAS` for evaluation.272* `MODEL.SEM_SEG_HEAD.NUM_CLASSES` sets the number of stuff classes for Semantic FPN & Panoptic FPN.273* `TEST.DETECTIONS_PER_IMAGE` controls the maximum number of objects to be detected.274  Set it to a larger number if test images may contain >100 objects.275* If you're training Fast R-CNN (with precomputed proposals), `DATASETS.PROPOSAL_FILES_{TRAIN,TEST}`276  need to match the datasets. The format of proposal files are documented277  [here](../modules/data.html#detectron2.data.load_proposals_into_dataset).278 279New models280(e.g. [TensorMask](../../projects/TensorMask),281[PointRend](../../projects/PointRend))282often have similar configs of their own that need to be changed as well.283 284```eval_rst285.. tip::286 287   After changing the number of classes, certain layers in a pre-trained model will become incompatible288   and therefore cannot be loaded to the new model.289   This is expected, and loading such pre-trained models will produce warnings about such layers.290```291