Team Ai
Apppublic

MLVLab/Human_Object_Interaction

sourceHugging Faceapache-2.0updated 3y agoView on Hugging Face
1likes
launch.py192 linesDownload Raw Back to tools
1# --------------------------------------------------------------------------------------------------------------------------2# Deformable DETR3# Copyright (c) 2020 SenseTime. All Rights Reserved.4# Licensed under the Apache License, Version 2.0 [see LICENSE for details]5# --------------------------------------------------------------------------------------------------------------------------6# Modified from https://github.com/pytorch/pytorch/blob/173f224570017b4b1a3a1a13d0bff280a54d9cd9/torch/distributed/launch.py7# --------------------------------------------------------------------------------------------------------------------------8 9r"""10`torch.distributed.launch` is a module that spawns up multiple distributed11training processes on each of the training nodes.12The utility can be used for single-node distributed training, in which one or13more processes per node will be spawned. The utility can be used for either14CPU training or GPU training. If the utility is used for GPU training,15each distributed process will be operating on a single GPU. This can achieve16well-improved single-node training performance. It can also be used in17multi-node distributed training, by spawning up multiple processes on each node18for well-improved multi-node distributed training performance as well.19This will especially be benefitial for systems with multiple Infiniband20interfaces that have direct-GPU support, since all of them can be utilized for21aggregated communication bandwidth.22In both cases of single-node distributed training or multi-node distributed23training, this utility will launch the given number of processes per node24(``--nproc_per_node``). If used for GPU training, this number needs to be less25or euqal to the number of GPUs on the current system (``nproc_per_node``),26and each process will be operating on a single GPU from *GPU 0 to27GPU (nproc_per_node - 1)*.28**How to use this module:**291. Single-Node multi-process distributed training30::31    >>> python -m torch.distributed.launch --nproc_per_node=NUM_GPUS_YOU_HAVE32               YOUR_TRAINING_SCRIPT.py (--arg1 --arg2 --arg3 and all other33               arguments of your training script)342. Multi-Node multi-process distributed training: (e.g. two nodes)35Node 1: *(IP: 192.168.1.1, and has a free port: 1234)*36::37    >>> python -m torch.distributed.launch --nproc_per_node=NUM_GPUS_YOU_HAVE38               --nnodes=2 --node_rank=0 --master_addr="192.168.1.1"39               --master_port=1234 YOUR_TRAINING_SCRIPT.py (--arg1 --arg2 --arg340               and all other arguments of your training script)41Node 2:42::43    >>> python -m torch.distributed.launch --nproc_per_node=NUM_GPUS_YOU_HAVE44               --nnodes=2 --node_rank=1 --master_addr="192.168.1.1"45               --master_port=1234 YOUR_TRAINING_SCRIPT.py (--arg1 --arg2 --arg346               and all other arguments of your training script)473. To look up what optional arguments this module offers:48::49    >>> python -m torch.distributed.launch --help50**Important Notices:**511. This utilty and multi-process distributed (single-node or52multi-node) GPU training currently only achieves the best performance using53the NCCL distributed backend. Thus NCCL backend is the recommended backend to54use for GPU training.552. In your training program, you must parse the command-line argument:56``--local_rank=LOCAL_PROCESS_RANK``, which will be provided by this module.57If your training program uses GPUs, you should ensure that your code only58runs on the GPU device of LOCAL_PROCESS_RANK. This can be done by:59Parsing the local_rank argument60::61    >>> import argparse62    >>> parser = argparse.ArgumentParser()63    >>> parser.add_argument("--local_rank", type=int)64    >>> args = parser.parse_args()65Set your device to local rank using either66::67    >>> torch.cuda.set_device(arg.local_rank)  # before your code runs68or69::70    >>> with torch.cuda.device(arg.local_rank):71    >>>    # your code to run723. In your training program, you are supposed to call the following function73at the beginning to start the distributed backend. You need to make sure that74the init_method uses ``env://``, which is the only supported ``init_method``75by this module.76::77    torch.distributed.init_process_group(backend='YOUR BACKEND',78                                         init_method='env://')794. In your training program, you can either use regular distributed functions80or use :func:`torch.nn.parallel.DistributedDataParallel` module. If your81training program uses GPUs for training and you would like to use82:func:`torch.nn.parallel.DistributedDataParallel` module,83here is how to configure it.84::85    model = torch.nn.parallel.DistributedDataParallel(model,86                                                      device_ids=[arg.local_rank],87                                                      output_device=arg.local_rank)88Please ensure that ``device_ids`` argument is set to be the only GPU device id89that your code will be operating on. This is generally the local rank of the90process. In other words, the ``device_ids`` needs to be ``[args.local_rank]``,91and ``output_device`` needs to be ``args.local_rank`` in order to use this92utility935. Another way to pass ``local_rank`` to the subprocesses via environment variable94``LOCAL_RANK``. This behavior is enabled when you launch the script with95``--use_env=True``. You must adjust the subprocess example above to replace96``args.local_rank`` with ``os.environ['LOCAL_RANK']``; the launcher97will not pass ``--local_rank`` when you specify this flag.98.. warning::99    ``local_rank`` is NOT globally unique: it is only unique per process100    on a machine.  Thus, don't use it to decide if you should, e.g.,101    write to a networked filesystem.  See102    https://github.com/pytorch/pytorch/issues/12042 for an example of103    how things can go wrong if you don't do this correctly.104"""105 106 107import sys108import subprocess109import os110import socket111from argparse import ArgumentParser, REMAINDER112 113import torch114 115 116def parse_args():117    """118    Helper function parsing the command line options119    @retval ArgumentParser120    """121    parser = ArgumentParser(description="PyTorch distributed training launch "122                                        "helper utilty that will spawn up "123                                        "multiple distributed processes")124 125    # Optional arguments for the launch helper126    parser.add_argument("--nnodes", type=int, default=1,127                        help="The number of nodes to use for distributed "128                             "training")129    parser.add_argument("--node_rank", type=int, default=0,130                        help="The rank of the node for multi-node distributed "131                             "training")132    parser.add_argument("--nproc_per_node", type=int, default=1,133                        help="The number of processes to launch on each node, "134                             "for GPU training, this is recommended to be set "135                             "to the number of GPUs in your system so that "136                             "each process can be bound to a single GPU.")137    parser.add_argument("--master_addr", default="127.0.0.1", type=str,138                        help="Master node (rank 0)'s address, should be either "139                             "the IP address or the hostname of node 0, for "140                             "single node multi-proc training, the "141                             "--master_addr can simply be 127.0.0.1")142    parser.add_argument("--master_port", default=29500, type=int,143                        help="Master node (rank 0)'s free port that needs to "144                             "be used for communciation during distributed "145                             "training")146 147    # positional148    parser.add_argument("training_script", type=str,149                        help="The full path to the single GPU training "150                             "program/script to be launched in parallel, "151                             "followed by all the arguments for the "152                             "training script")153 154    # rest from the training program155    parser.add_argument('training_script_args', nargs=REMAINDER)156    return parser.parse_args()157 158 159def main():160    args = parse_args()161 162    # world size in terms of number of processes163    dist_world_size = args.nproc_per_node * args.nnodes164 165    # set PyTorch distributed related environmental variables166    current_env = os.environ.copy()167    current_env["MASTER_ADDR"] = args.master_addr168    current_env["MASTER_PORT"] = str(args.master_port)169    current_env["WORLD_SIZE"] = str(dist_world_size)170 171    processes = []172 173    for local_rank in range(0, args.nproc_per_node):174        # each process's rank175        dist_rank = args.nproc_per_node * args.node_rank + local_rank176        current_env["RANK"] = str(dist_rank)177        current_env["LOCAL_RANK"] = str(local_rank)178 179        cmd = [args.training_script] + args.training_script_args180 181        process = subprocess.Popen(cmd, env=current_env)182        processes.append(process)183 184    for process in processes:185        process.wait()186        if process.returncode != 0:187            raise subprocess.CalledProcessError(returncode=process.returncode,188                                                cmd=process.args)189 190 191if __name__ == "__main__":192    main()