Arulkumar03/Wheat_HEAD_Detection_Counting_ComputerVision_Model
0
1# Use Models2 3## Build Models from Yacs Config4From a yacs config object,5models (and their sub-models) can be built by6functions such as `build_model`, `build_backbone`, `build_roi_heads`:7```python8from detectron2.modeling import build_model9model = build_model(cfg) # returns a torch.nn.Module10```11 12`build_model` only builds the model structure and fills it with random parameters.13See below for how to load an existing checkpoint to the model and how to use the `model` object.14 15### Load/Save a Checkpoint16```python17from detectron2.checkpoint import DetectionCheckpointer18DetectionCheckpointer(model).load(file_path_or_url) # load a file, usually from cfg.MODEL.WEIGHTS19 20checkpointer = DetectionCheckpointer(model, save_dir="output")21checkpointer.save("model_999") # save to output/model_999.pth22```23 24Detectron2's checkpointer recognizes models in pytorch's `.pth` format, as well as the `.pkl` files25in our model zoo.26See [API doc](../modules/checkpoint.html#detectron2.checkpoint.DetectionCheckpointer)27for more details about its usage.28 29The model files can be arbitrarily manipulated using `torch.{load,save}` for `.pth` files or30`pickle.{dump,load}` for `.pkl` files.31 32### Use a Model33 34A model can be called by `outputs = model(inputs)`, where `inputs` is a `list[dict]`.35Each dict corresponds to one image and the required keys36depend on the type of model, and whether the model is in training or evaluation mode.37For example, in order to do inference,38all existing models expect the "image" key, and optionally "height" and "width".39The detailed format of inputs and outputs of existing models are explained below.40 41__Training__: When in training mode, all models are required to be used under an `EventStorage`.42The training statistics will be put into the storage:43```python44from detectron2.utils.events import EventStorage45with EventStorage() as storage:46 losses = model(inputs)47```48 49__Inference__: If you only want to do simple inference using an existing model,50[DefaultPredictor](../modules/engine.html#detectron2.engine.defaults.DefaultPredictor)51is a wrapper around model that provides such basic functionality.52It includes default behavior including model loading, preprocessing,53and operates on single image rather than batches. See its documentation for usage.54 55You can also run inference directly like this:56```python57model.eval()58with torch.no_grad():59 outputs = model(inputs)60```61 62### Model Input Format63 64Users can implement custom models that support any arbitrary input format.65Here we describe the standard input format that all builtin models support in detectron2.66They all take a `list[dict]` as the inputs. Each dict67corresponds to information about one image.68 69The dict may contain the following keys:70 71* "image": `Tensor` in (C, H, W) format. The meaning of channels are defined by `cfg.INPUT.FORMAT`.72 Image normalization, if any, will be performed inside the model using73 `cfg.MODEL.PIXEL_{MEAN,STD}`.74* "height", "width": the **desired** output height and width **in inference**, which is not necessarily the same75 as the height or width of the `image` field.76 For example, the `image` field contains the resized image, if resize is used as a preprocessing step.77 But you may want the outputs to be in **original** resolution.78 If provided, the model will produce output in this resolution,79 rather than in the resolution of the `image` as input into the model. This is more efficient and accurate.80* "instances": an [Instances](../modules/structures.html#detectron2.structures.Instances)81 object for training, with the following fields:82 + "gt_boxes": a [Boxes](../modules/structures.html#detectron2.structures.Boxes) object storing N boxes, one for each instance.83 + "gt_classes": `Tensor` of long type, a vector of N labels, in range [0, num_categories).84 + "gt_masks": a [PolygonMasks](../modules/structures.html#detectron2.structures.PolygonMasks)85 or [BitMasks](../modules/structures.html#detectron2.structures.BitMasks) object storing N masks, one for each instance.86 + "gt_keypoints": a [Keypoints](../modules/structures.html#detectron2.structures.Keypoints)87 object storing N keypoint sets, one for each instance.88* "sem_seg": `Tensor[int]` in (H, W) format. The semantic segmentation ground truth for training.89 Values represent category labels starting from 0.90* "proposals": an [Instances](../modules/structures.html#detectron2.structures.Instances)91 object used only in Fast R-CNN style models, with the following fields:92 + "proposal_boxes": a [Boxes](../modules/structures.html#detectron2.structures.Boxes) object storing P proposal boxes.93 + "objectness_logits": `Tensor`, a vector of P scores, one for each proposal.94 95For inference of builtin models, only "image" key is required, and "width/height" are optional.96 97We currently don't define standard input format for panoptic segmentation training,98because models now use custom formats produced by custom data loaders.99 100#### How it connects to data loader:101 102The output of the default [DatasetMapper]( ../modules/data.html#detectron2.data.DatasetMapper) is a dict103that follows the above format.104After the data loader performs batching, it becomes `list[dict]` which the builtin models support.105 106 107### Model Output Format108 109When in training mode, the builtin models output a `dict[str->ScalarTensor]` with all the losses.110 111When in inference mode, the builtin models output a `list[dict]`, one dict for each image.112Based on the tasks the model is doing, each dict may contain the following fields:113 114* "instances": [Instances](../modules/structures.html#detectron2.structures.Instances)115 object with the following fields:116 * "pred_boxes": [Boxes](../modules/structures.html#detectron2.structures.Boxes) object storing N boxes, one for each detected instance.117 * "scores": `Tensor`, a vector of N confidence scores.118 * "pred_classes": `Tensor`, a vector of N labels in range [0, num_categories).119 + "pred_masks": a `Tensor` of shape (N, H, W), masks for each detected instance.120 + "pred_keypoints": a `Tensor` of shape (N, num_keypoint, 3).121 Each row in the last dimension is (x, y, score). Confidence scores are larger than 0.122* "sem_seg": `Tensor` of (num_categories, H, W), the semantic segmentation prediction.123* "proposals": [Instances](../modules/structures.html#detectron2.structures.Instances)124 object with the following fields:125 * "proposal_boxes": [Boxes](../modules/structures.html#detectron2.structures.Boxes)126 object storing N boxes.127 * "objectness_logits": a torch vector of N confidence scores.128* "panoptic_seg": A tuple of `(pred: Tensor, segments_info: Optional[list[dict]])`.129 The `pred` tensor has shape (H, W), containing the segment id of each pixel.130 131 * If `segments_info` exists, each dict describes one segment id in `pred` and has the following fields:132 133 * "id": the segment id134 * "isthing": whether the segment is a thing or stuff135 * "category_id": the category id of this segment.136 137 If a pixel's id does not exist in `segments_info`, it is considered to be void label138 defined in [Panoptic Segmentation](https://arxiv.org/abs/1801.00868).139 140 * If `segments_info` is None, all pixel values in `pred` must be ≥ -1.141 Pixels with value -1 are assigned void labels.142 Otherwise, the category id of each pixel is obtained by143 `category_id = pixel // metadata.label_divisor`.144 145 146### Partially execute a model:147 148Sometimes you may want to obtain an intermediate tensor inside a model,149such as the input of certain layer, the output before post-processing.150Since there are typically hundreds of intermediate tensors, there isn't an API that provides you151the intermediate result you need.152You have the following options:153 1541. Write a (sub)model. Following the [tutorial](./write-models.md), you can155 rewrite a model component (e.g. a head of a model), such that it156 does the same thing as the existing component, but returns the output157 you need.1582. Partially execute a model. You can create the model as usual,159 but use custom code to execute it instead of its `forward()`. For example,160 the following code obtains mask features before mask head.161 162 ```python163 images = ImageList.from_tensors(...) # preprocessed input tensor164 model = build_model(cfg)165 model.eval()166 features = model.backbone(images.tensor)167 proposals, _ = model.proposal_generator(images, features)168 instances, _ = model.roi_heads(images, features, proposals)169 mask_features = [features[f] for f in model.roi_heads.in_features]170 mask_features = model.roi_heads.mask_pooler(mask_features, [x.pred_boxes for x in instances])171 ```172 1733. Use [forward hooks](https://pytorch.org/tutorials/beginner/former_torchies/nnft_tutorial.html#forward-and-backward-function-hooks).174 Forward hooks can help you obtain inputs or outputs of a certain module.175 If they are not exactly what you want, they can at least be used together with partial execution176 to obtain other tensors.177 178All options require you to read documentation and sometimes code179of the existing models to understand the internal logic,180in order to write code to obtain the internal tensors.181 