Team Ai
Apppublic

Arulkumar03/Wheat_HEAD_Detection_Counting_ComputerVision_Model

sourceHugging Facemitupdated 3y agoView on Hugging Face
0likes
benchmarks.md197 linesDownload Raw Back to notes
1 2# Benchmarks3 4Here we benchmark the training speed of a Mask R-CNN in detectron2,5with some other popular open source Mask R-CNN implementations.6 7 8### Settings9 10* Hardware: 8 NVIDIA V100s with NVLink.11* Software: Python 3.7, CUDA 10.1, cuDNN 7.6.5, PyTorch 1.5,12  TensorFlow 1.15.0rc2, Keras 2.2.5, MxNet 1.6.0b20190820.13* Model: an end-to-end R-50-FPN Mask-RCNN model, using the same hyperparameter as the14  [Detectron baseline config](https://github.com/facebookresearch/Detectron/blob/master/configs/12_2017_baselines/e2e_mask_rcnn_R-50-FPN_1x.yaml)15  (it does not have scale augmentation).16* Metrics: We use the average throughput in iterations 100-500 to skip GPU warmup time.17  Note that for R-CNN-style models, the throughput of a model typically changes during training, because18  it depends on the predictions of the model. Therefore this metric is not directly comparable with19  "train speed" in model zoo, which is the average speed of the entire training run.20 21 22### Main Results23 24```eval_rst25+-------------------------------+--------------------+26| Implementation                | Throughput (img/s) |27+===============================+====================+28| |D2| |PT|                     | 62                 |29+-------------------------------+--------------------+30| mmdetection_  |PT|            | 53                 |31+-------------------------------+--------------------+32| maskrcnn-benchmark_  |PT|     | 53                 |33+-------------------------------+--------------------+34| tensorpack_ |TF|              | 50                 |35+-------------------------------+--------------------+36| simpledet_ |mxnet|            | 39                 |37+-------------------------------+--------------------+38| Detectron_  |C2|              | 19                 |39+-------------------------------+--------------------+40| `matterport/Mask_RCNN`__ |TF| | 14                 |41+-------------------------------+--------------------+42 43.. _maskrcnn-benchmark: https://github.com/facebookresearch/maskrcnn-benchmark/44.. _tensorpack: https://github.com/tensorpack/tensorpack/tree/master/examples/FasterRCNN45.. _mmdetection: https://github.com/open-mmlab/mmdetection/46.. _simpledet: https://github.com/TuSimple/simpledet/47.. _Detectron: https://github.com/facebookresearch/Detectron48__ https://github.com/matterport/Mask_RCNN/49 50.. |D2| image:: https://github.com/facebookresearch/detectron2/raw/main/.github/Detectron2-Logo-Horz.svg?sanitize=true51   :height: 15pt52   :target: https://github.com/facebookresearch/detectron2/53.. |PT| image:: https://pytorch.org/assets/images/logo-icon.svg54   :width: 15pt55   :height: 15pt56   :target: https://pytorch.org57.. |TF| image:: https://static.nvidiagrid.net/ngc/containers/tensorflow.png58   :width: 15pt59   :height: 15pt60   :target: https://tensorflow.org61.. |mxnet| image:: https://github.com/dmlc/web-data/raw/master/mxnet/image/mxnet_favicon.png62   :width: 15pt63   :height: 15pt64   :target: https://mxnet.apache.org/65.. |C2| image:: https://caffe2.ai/static/logo.svg66   :width: 15pt67   :height: 15pt68   :target: https://caffe2.ai69```70 71 72Details for each implementation:73 74* __Detectron2__: with release v0.1.2, run:75  ```76  python tools/train_net.py  --config-file configs/Detectron1-Comparisons/mask_rcnn_R_50_FPN_noaug_1x.yaml --num-gpus 877  ```78 79* __mmdetection__: at commit `b0d845f`, run80  ```81  ./tools/dist_train.sh configs/mask_rcnn/mask_rcnn_r50_caffe_fpn_1x_coco.py 882  ```83 84* __maskrcnn-benchmark__: use commit `0ce8f6f` with `sed -i 's/torch.uint8/torch.bool/g' **/*.py; sed -i 's/AT_CHECK/TORCH_CHECK/g' **/*.cu`85  to make it compatible with PyTorch 1.5. Then, run training with86  ```87  python -m torch.distributed.launch --nproc_per_node=8 tools/train_net.py --config-file configs/e2e_mask_rcnn_R_50_FPN_1x.yaml88  ```89  The speed we observed is faster than its model zoo, likely due to different software versions.90 91* __tensorpack__: at commit `caafda`, `export TF_CUDNN_USE_AUTOTUNE=0`, then run92  ```93  mpirun -np 8 ./train.py --config DATA.BASEDIR=/data/coco TRAINER=horovod BACKBONE.STRIDE_1X1=True TRAIN.STEPS_PER_EPOCH=50 --load ImageNet-R50-AlignPadding.npz94  ```95 96* __SimpleDet__: at commit `9187a1`, run97  ```98  python detection_train.py --config config/mask_r50v1_fpn_1x.py99  ```100 101* __Detectron__: run102  ```103  python tools/train_net.py --cfg configs/12_2017_baselines/e2e_mask_rcnn_R-50-FPN_1x.yaml104  ```105  Note that many of its ops run on CPUs, therefore the performance is limited.106 107* __matterport/Mask_RCNN__: at commit `3deaec`, apply the following diff, `export TF_CUDNN_USE_AUTOTUNE=0`, then run108  ```109  python coco.py train --dataset=/data/coco/ --model=imagenet110  ```111  Note that many small details in this implementation might be different112  from Detectron's standards.113 114  <details>115  <summary>116  (diff to make it use the same hyperparameters - click to expand)117  </summary>118 119  ```diff120  diff --git i/mrcnn/model.py w/mrcnn/model.py121  index 62cb2b0..61d7779 100644122  --- i/mrcnn/model.py123  +++ w/mrcnn/model.py124  @@ -2367,8 +2367,8 @@ class MaskRCNN():125        epochs=epochs,126        steps_per_epoch=self.config.STEPS_PER_EPOCH,127        callbacks=callbacks,128  -            validation_data=val_generator,129  -            validation_steps=self.config.VALIDATION_STEPS,130  +            #validation_data=val_generator,131  +            #validation_steps=self.config.VALIDATION_STEPS,132        max_queue_size=100,133        workers=workers,134        use_multiprocessing=True,135  diff --git i/mrcnn/parallel_model.py w/mrcnn/parallel_model.py136  index d2bf53b..060172a 100644137  --- i/mrcnn/parallel_model.py138  +++ w/mrcnn/parallel_model.py139  @@ -32,6 +32,7 @@ class ParallelModel(KM.Model):140      keras_model: The Keras model to parallelize141      gpu_count: Number of GPUs. Must be > 1142      """143  +        super().__init__()144      self.inner_model = keras_model145      self.gpu_count = gpu_count146      merged_outputs = self.make_parallel()147  diff --git i/samples/coco/coco.py w/samples/coco/coco.py148  index 5d172b5..239ed75 100644149  --- i/samples/coco/coco.py150  +++ w/samples/coco/coco.py151  @@ -81,7 +81,10 @@ class CocoConfig(Config):152    IMAGES_PER_GPU = 2153 154    # Uncomment to train on 8 GPUs (default is 1)155  -    # GPU_COUNT = 8156  +    GPU_COUNT = 8157  +    BACKBONE = "resnet50"158  +    STEPS_PER_EPOCH = 50159  +    TRAIN_ROIS_PER_IMAGE = 512160 161    # Number of classes (including background)162    NUM_CLASSES = 1 + 80  # COCO has 80 classes163  @@ -496,29 +499,10 @@ if __name__ == '__main__':164      # *** This training schedule is an example. Update to your needs ***165 166      # Training - Stage 1167  -        print("Training network heads")168      model.train(dataset_train, dataset_val,169            learning_rate=config.LEARNING_RATE,170            epochs=40,171  -                    layers='heads',172  -                    augmentation=augmentation)173  -174  -        # Training - Stage 2175  -        # Finetune layers from ResNet stage 4 and up176  -        print("Fine tune Resnet stage 4 and up")177  -        model.train(dataset_train, dataset_val,178  -                    learning_rate=config.LEARNING_RATE,179  -                    epochs=120,180  -                    layers='4+',181  -                    augmentation=augmentation)182  -183  -        # Training - Stage 3184  -        # Fine tune all layers185  -        print("Fine tune all layers")186  -        model.train(dataset_train, dataset_val,187  -                    learning_rate=config.LEARNING_RATE / 10,188  -                    epochs=160,189  -                    layers='all',190  +                    layers='3+',191            augmentation=augmentation)192 193    elif args.command == "evaluate":194  ```195 196  </details>197