Team Ai
Apppublic

Arulkumar03/Wheat_HEAD_Detection_Counting_ComputerVision_Model

sourceHugging Facemitupdated 3y agoView on Hugging Face
0likes
data_loading.md96 linesDownload Raw Back to tutorials
1 2# Dataloader3 4Dataloader is the component that provides data to models.5A dataloader usually (but not necessarily) takes raw information from [datasets](./datasets.md),6and process them into a format needed by the model.7 8## How the Existing Dataloader Works9 10Detectron2 contains a builtin data loading pipeline.11It's good to understand how it works, in case you need to write a custom one.12 13Detectron2 provides two functions14[build_detection_{train,test}_loader](../modules/data.html#detectron2.data.build_detection_train_loader)15that create a default data loader from a given config.16Here is how `build_detection_{train,test}_loader` work:17 181. It takes the name of a registered dataset (e.g., "coco_2017_train") and loads a `list[dict]` representing the dataset items19   in a lightweight format. These dataset items are not yet ready to be used by the model (e.g., images are20   not loaded into memory, random augmentations have not been applied, etc.).21   Details about the dataset format and dataset registration can be found in22   [datasets](./datasets.md).232. Each dict in this list is mapped by a function ("mapper"):24   * Users can customize this mapping function by specifying the "mapper" argument in25        `build_detection_{train,test}_loader`. The default mapper is [DatasetMapper](../modules/data.html#detectron2.data.DatasetMapper).26   * The output format of the mapper can be arbitrary, as long as it is accepted by the consumer of this data loader (usually the model).27     The outputs of the default mapper, after batching, follow the default model input format documented in28     [Use Models](./models.html#model-input-format).29   * The role of the mapper is to transform the lightweight representation of a dataset item into a format30     that is ready for the model to consume (including, e.g., read images, perform random data augmentation and convert to torch Tensors).31     If you would like to perform custom transformations to data, you often want a custom mapper.323. The outputs of the mapper are batched (simply into a list).334. This batched data is the output of the data loader. Typically, it's also the input of34   `model.forward()`.35 36 37## Write a Custom Dataloader38 39Using a different "mapper" with `build_detection_{train,test}_loader(mapper=)` works for most use cases40of custom data loading.41For example, if you want to resize all images to a fixed size for training, use:42 43```python44import detectron2.data.transforms as T45from detectron2.data import DatasetMapper   # the default mapper46dataloader = build_detection_train_loader(cfg,47   mapper=DatasetMapper(cfg, is_train=True, augmentations=[48      T.Resize((800, 800))49   ]))50# use this dataloader instead of the default51```52If the arguments of the default [DatasetMapper](../modules/data.html#detectron2.data.DatasetMapper)53does not provide what you need, you may write a custom mapper function and use it instead, e.g.:54 55```python56from detectron2.data import detection_utils as utils57 # Show how to implement a minimal mapper, similar to the default DatasetMapper58def mapper(dataset_dict):59    dataset_dict = copy.deepcopy(dataset_dict)  # it will be modified by code below60    # can use other ways to read image61    image = utils.read_image(dataset_dict["file_name"], format="BGR")62    # See "Data Augmentation" tutorial for details usage63    auginput = T.AugInput(image)64    transform = T.Resize((800, 800))(auginput)65    image = torch.from_numpy(auginput.image.transpose(2, 0, 1))66    annos = [67        utils.transform_instance_annotations(annotation, [transform], image.shape[1:])68        for annotation in dataset_dict.pop("annotations")69    ]70    return {71       # create the format that the model expects72       "image": image,73       "instances": utils.annotations_to_instances(annos, image.shape[1:])74    }75dataloader = build_detection_train_loader(cfg, mapper=mapper)76```77 78If you want to change not only the mapper (e.g., in order to implement different sampling or batching logic),79`build_detection_train_loader` won't work and you will need to write a different data loader.80The data loader is simply a81python iterator that produces [the format](./models.md) that the model accepts.82You can implement it using any tools you like.83 84No matter what to implement, it's recommended to85check out [API documentation of detectron2.data](../modules/data) to learn more about the APIs of86these functions.87 88## Use a Custom Dataloader89 90If you use [DefaultTrainer](../modules/engine.html#detectron2.engine.defaults.DefaultTrainer),91you can overwrite its `build_{train,test}_loader` method to use your own dataloader.92See the [deeplab dataloader](../../projects/DeepLab/train_net.py)93for an example.94 95If you write your own training loop, you can plug in your data loader easily.96