Arulkumar03/Wheat_HEAD_Detection_Counting_ComputerVision_Model
0
1 2# Data Augmentation3 4Augmentation is an important part of training.5Detectron2's data augmentation system aims at addressing the following goals:6 71. Allow augmenting multiple data types together8 (e.g., images together with their bounding boxes and masks)92. Allow applying a sequence of statically-declared augmentation103. Allow adding custom new data types to augment (rotated bounding boxes, video clips, etc.)114. Process and manipulate the __operations__ that are applied by augmentations12 13The first two features cover most of the common use cases, and is also14available in other libraries such as [albumentations](https://medium.com/pytorch/multi-target-in-albumentations-16a777e9006e).15Supporting other features adds some overhead to detectron2's augmentation API,16which we'll explain in this tutorial.17 18This tutorial focuses on how to use augmentations when writing new data loaders,19and how to write new augmentations.20If you use the default data loader in detectron2, it already supports taking a user-provided list of custom augmentations,21as explained in the [Dataloader tutorial](data_loading).22 23## Basic Usage24 25The basic usage of feature (1) and (2) is like the following:26```python27from detectron2.data import transforms as T28# Define a sequence of augmentations:29augs = T.AugmentationList([30 T.RandomBrightness(0.9, 1.1),31 T.RandomFlip(prob=0.5),32 T.RandomCrop("absolute", (640, 640))33]) # type: T.Augmentation34 35# Define the augmentation input ("image" required, others optional):36input = T.AugInput(image, boxes=boxes, sem_seg=sem_seg)37# Apply the augmentation:38transform = augs(input) # type: T.Transform39image_transformed = input.image # new image40sem_seg_transformed = input.sem_seg # new semantic segmentation41 42# For any extra data that needs to be augmented together, use transform, e.g.:43image2_transformed = transform.apply_image(image2)44polygons_transformed = transform.apply_polygons(polygons)45```46 47Three basic concepts are involved here. They are:48* [T.Augmentation](../modules/data_transforms.html#detectron2.data.transforms.Augmentation) defines the __"policy"__ to modify inputs.49 * its `__call__(AugInput) -> Transform` method augments the inputs in-place, and returns the operation that is applied50* [T.Transform](../modules/data_transforms.html#detectron2.data.transforms.Transform)51 implements the actual __operations__ to transform data52 * it has methods such as `apply_image`, `apply_coords` that define how to transform each data type53* [T.AugInput](../modules/data_transforms.html#detectron2.data.transforms.AugInput)54 stores inputs needed by `T.Augmentation` and how they should be transformed.55 This concept is needed for some advanced usage.56 Using this class directly should be sufficient for all common use cases,57 since extra data not in `T.AugInput` can be augmented using the returned58 `transform`, as shown in the above example.59 60## Write New Augmentations61 62Most 2D augmentations only need to know about the input image. Such augmentation can be implemented easily like this:63 64```python65class MyColorAugmentation(T.Augmentation):66 def get_transform(self, image):67 r = np.random.rand(2)68 return T.ColorTransform(lambda x: x * r[0] + r[1] * 10)69 70class MyCustomResize(T.Augmentation):71 def get_transform(self, image):72 old_h, old_w = image.shape[:2]73 new_h, new_w = int(old_h * np.random.rand()), int(old_w * 1.5)74 return T.ResizeTransform(old_h, old_w, new_h, new_w)75 76augs = MyCustomResize()77transform = augs(input)78```79 80In addition to image, any attributes of the given `AugInput` can be used as long81as they are part of the function signature, e.g.:82 83```python84class MyCustomCrop(T.Augmentation):85 def get_transform(self, image, sem_seg):86 # decide where to crop using both image and sem_seg87 return T.CropTransform(...)88 89augs = MyCustomCrop()90assert hasattr(input, "image") and hasattr(input, "sem_seg")91transform = augs(input)92```93 94New transform operation can also be added by subclassing95[T.Transform](../modules/data_transforms.html#detectron2.data.transforms.Transform).96 97## Advanced Usage98 99We give a few examples of advanced usages that100are enabled by our system.101These options can be interesting to new research,102although changing them is often not needed103for standard use cases.104 105### Custom transform strategy106 107Instead of only returning the augmented data, detectron2's `Augmentation` returns the __operations__ as `T.Transform`.108This allows users to apply custom transform strategy on their data.109We use keypoints data as an example.110 111Keypoints are (x, y) coordinates, but they are not so trivial to augment due to the semantic meaning they carry.112Such meaning is only known to the users, therefore users may want to augment them manually113by looking at the returned `transform`.114For example, when an image is horizontally flipped, we'd like to swap the keypoint annotations for "left eye" and "right eye".115This can be done like this (included by default in detectron2's default data loader):116```python117# augs, input are defined as in previous examples118transform = augs(input) # type: T.Transform119keypoints_xy = transform.apply_coords(keypoints_xy) # transform the coordinates120 121# get a list of all transforms that were applied122transforms = T.TransformList([transform]).transforms123# check if it is flipped for odd number of times124do_hflip = sum(isinstance(t, T.HFlipTransform) for t in transforms) % 2 == 1125if do_hflip:126 keypoints_xy = keypoints_xy[flip_indices_mapping]127```128 129As another example, keypoints annotations often have a "visibility" field.130A sequence of augmentations might augment a visible keypoint out of the image boundary (e.g. with cropping),131but then bring it back within the boundary afterwards (e.g. with image padding).132If users decide to label such keypoints "invisible",133then the visibility check has to happen after every transform step.134This can be achieved by:135 136```python137transform = augs(input) # type: T.TransformList138assert isinstance(transform, T.TransformList)139for t in transform.transforms:140 keypoints_xy = t.apply_coords(keypoints_xy)141 visibility &= (keypoints_xy >= [0, 0] & keypoints_xy <= [W, H]).all(axis=1)142 143# btw, detectron2's `transform_keypoint_annotations` function chooses to label such keypoints "visible":144# keypoints_xy = transform.apply_coords(keypoints_xy)145# visibility &= (keypoints_xy >= [0, 0] & keypoints_xy <= [W, H]).all(axis=1)146```147 148 149### Geometrically invert the transform150If images are pre-processed by augmentations before inference, the predicted results151such as segmentation masks are localized on the augmented image.152We'd like to invert the applied augmentation with the [inverse()](../modules/data_transforms.html#detectron2.data.transforms.Transform.inverse)153API, to obtain results on the original image:154```python155transform = augs(input)156pred_mask = make_prediction(input.image)157inv_transform = transform.inverse()158pred_mask_orig = inv_transform.apply_segmentation(pred_mask)159```160 161### Add new data types162 163[T.Transform](../modules/data_transforms.html#detectron2.data.transforms.Transform)164supports a few common data types to transform, including images, coordinates, masks, boxes, polygons.165It allows registering new data types, e.g.:166```python167@T.HFlipTransform.register_type("rotated_boxes")168def func(flip_transform: T.HFlipTransform, rotated_boxes: Any):169 # do the work170 return flipped_rotated_boxes171 172t = HFlipTransform(width=800)173transformed_rotated_boxes = t.apply_rotated_boxes(rotated_boxes) # func will be called174```175 176### Extend T.AugInput177 178An augmentation can only access attributes available in the given input.179[T.AugInput](../modules/data_transforms.html#detectron2.data.transforms.StandardAugInput) defines "image", "boxes", "sem_seg",180which are sufficient for common augmentation strategies to decide how to augment.181If not, a custom implementation is needed.182 183By re-implement the "transform()" method in AugInput, it is also possible to184augment different fields in ways that are dependent on each other.185Such use case is uncommon (e.g. post-process bounding box based on augmented masks), but allowed by the system.186 187 