{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.7.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":24800,"databundleVersionId":1831594,"sourceType":"competition"},{"sourceId":1800067,"sourceType":"datasetVersion","datasetId":1069810}],"dockerImageVersionId":30043,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"VinBigData detectron2 train","metadata":{}},{"cell_type":"markdown","source":"# Table of Contents\n\n** [Dataset preparation](#dataset)** <br/>\n** [Installation](#installation)** <br/>\n** [Training method implementations](#train_method)** <br/>\n** [Customizing detectron2 trainer](#custom_trainer) ** [Advanced topic, skip it first time] <br/>\n**   - [Mapper for augmentation](#mapper)** <br/>\n**   - [Evaluator](#evaluator)** <br/>\n**   - [Loss evaluation hook](#loss_hook)** <br/>\n** [Loading Data](#load_data)** <br/>\n** [Data Visualization](#data_vis)** <br/>\n** [Training](#training)** <br/>\n** [Visualize loss curve & competition metric AP40](#vis_loss)** <br/>\n** [Visualization of augmentation by Mapper](#vis_aug)** <br/>\n** [Next step](#next_step)** <br/>","metadata":{}},{"cell_type":"markdown","source":"<a id=\"dataset\"></a>\n# Dataset preparation\n","metadata":{}},{"cell_type":"code","source":"import gc\nimport os\nfrom pathlib import Path\nimport random\nimport sys\n\nfrom tqdm.notebook import tqdm\nimport numpy as np\nimport pandas as pd\nimport scipy as sp\n\n\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nfrom IPython.core.display import display, HTML\n\n# --- plotly ---\nfrom plotly import tools, subplots\nimport plotly.offline as py\npy.init_notebook_mode(connected=True)\nimport plotly.graph_objs as go\nimport plotly.express as px\nimport plotly.figure_factory as ff\nimport plotly.io as pio\npio.templates.default = \"plotly_dark\"\n\n# --- models ---\nfrom sklearn import preprocessing\nfrom sklearn.model_selection import KFold\nimport lightgbm as lgb\nimport xgboost as xgb\nimport catboost as cb\n\n# --- setup ---\npd.set_option('max_columns', 50)\n","metadata":{"_kg_hide-input":true,"trusted":true,"execution":{"iopub.status.busy":"2025-04-24T13:57:28.171961Z","iopub.execute_input":"2025-04-24T13:57:28.1723Z","iopub.status.idle":"2025-04-24T13:57:28.199379Z","shell.execute_reply.started":"2025-04-24T13:57:28.172273Z","shell.execute_reply":"2025-04-24T13:57:28.19866Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<a id=\"installation\"></a>\n# Installation\n\ndetectron2 is not pre-installed in this kaggle docker, so let's install it. \nWe can follow [installation instruction](https://github.com/facebookresearch/detectron2/blob/master/INSTALL.md), we need to know CUDA and pytorch version to install correct `detectron2`.","metadata":{}},{"cell_type":"code","source":"!nvidia-smi","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true,"execution":{"iopub.status.busy":"2025-04-24T13:57:28.201034Z","iopub.execute_input":"2025-04-24T13:57:28.201258Z","iopub.status.idle":"2025-04-24T13:57:29.260699Z","shell.execute_reply.started":"2025-04-24T13:57:28.201235Z","shell.execute_reply":"2025-04-24T13:57:29.259623Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"!nvcc --version","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true,"execution":{"iopub.status.busy":"2025-04-24T13:57:29.262363Z","iopub.execute_input":"2025-04-24T13:57:29.26274Z","iopub.status.idle":"2025-04-24T13:57:30.280786Z","shell.execute_reply.started":"2025-04-24T13:57:29.262699Z","shell.execute_reply":"2025-04-24T13:57:30.280023Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import torch\n\ntorch.__version__","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true,"execution":{"iopub.status.busy":"2025-04-24T13:57:30.282691Z","iopub.execute_input":"2025-04-24T13:57:30.283121Z","iopub.status.idle":"2025-04-24T13:57:30.289096Z","shell.execute_reply.started":"2025-04-24T13:57:30.283076Z","shell.execute_reply":"2025-04-24T13:57:30.288226Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"It seems CUDA=10.2 and torch==1.7.0 is used in this kaggle docker image.\n\nSee [installation](https://detectron2.readthedocs.io/tutorials/install.html) for details.","metadata":{}},{"cell_type":"code","source":"!pip install detectron2 -f \\\n  https://dl.fbaipublicfiles.com/detectron2/wheels/cu102/torch1.7/index.html","metadata":{"_kg_hide-output":true,"trusted":true,"execution":{"iopub.status.busy":"2025-04-24T13:57:30.291834Z","iopub.execute_input":"2025-04-24T13:57:30.292106Z","iopub.status.idle":"2025-04-24T13:57:53.686103Z","shell.execute_reply.started":"2025-04-24T13:57:30.292082Z","shell.execute_reply":"2025-04-24T13:57:53.685241Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<a id=\"train_method\"></a>\n# Training method implementations\n\nBasically we don't need to implement neural network part, `detectron2` already implements famous architectures and provides its pre-trained weights. We can finetune these pre-trained architectures.\n\nThese models are summarized in [MODEL_ZOO.md](https://github.com/facebookresearch/detectron2/blob/master/MODEL_ZOO.md).\n\nIn this competition, we need object detection model, I will choose [R50-FPN](https://github.com/facebookresearch/detectron2/blob/master/configs/COCO-Detection/faster_rcnn_R_50_FPN_3x.yaml) for this kernel.","metadata":{}},{"cell_type":"markdown","source":"## Data preparation\n\n`detectron2` provides high-level API for training custom dataset.\n\nTo define custom dataset, we need to create **list of dict** (`dataset_dicts`) where each dict contains following:\n\n - file_name: file name of the image.\n - image_id: id of the image, index is used here.\n - height: height of the image.\n - width: width of the image.\n - annotation: This is the ground truth annotation data for object detection, which contains following\n     - bbox: bounding box pixel location with shape (n_boxes, 4)\n     - bbox_mode: `BoxMode.XYXY_ABS` is used here, meaning that absolute value of (xmin, ymin, xmax, ymax) annotation is used in the `bbox`.\n     - category_id: class label id for each bounding box, with shape (n_boxes,)\n\n`get_vinbigdata_dicts` is for train dataset preparation and `get_vinbigdata_dicts_test` is for test dataset preparation.\n\nThis `dataset_dicts` contains the metadata for actual data fed into the neural network.<br/>\nIt is loaded beforehand of the training **on memory**, so it should contain all the metadata (image filepath etc) to construct training dataset, but **should not contain heavy data**.<br/>\n\nIn practice, loading all the taining image arrays are too heavy to be loaded on memory, so these are loaded inside `DataLoader` on-demand (This is done by mapper class in `detectron2`, as I will expain later).","metadata":{}},{"cell_type":"code","source":"import pickle\nfrom pathlib import Path\nfrom typing import Optional\n\nimport cv2\nimport numpy as np\nimport pandas as pd\nfrom detectron2.structures import BoxMode\nfrom tqdm import tqdm\n\n\ndef get_vinbigdata_dicts(\n    imgdir: Path,\n    train_df: pd.DataFrame,\n    train_data_type: str = \"original\",\n    use_cache: bool = True,\n    debug: bool = True,\n    target_indices: Optional[np.ndarray] = None,\n    use_class14: bool = False,\n):\n    debug_str = f\"_debug{int(debug)}\"\n    train_data_type_str = f\"_{train_data_type}\"\n    class14_str = f\"_14class{int(use_class14)}\"\n    cache_path = Path(\".\") / f\"dataset_dicts_cache{train_data_type_str}{class14_str}{debug_str}.pkl\"\n    if not use_cache or not cache_path.exists():\n        print(\"Creating data...\")\n        train = pd.read_csv(imgdir / \"train.csv\")\n        if debug:\n            train = train.iloc[:500]  # For debug....\n\n        # Load 1 image to get image size.\n        image_id = train.loc[0, \"image_id\"]\n        image_path = str(imgdir / \"train\" / f\"{image_id}.png\")\n        image = cv2.imread(image_path)\n        resized_height, resized_width, ch = image.shape\n        print(f\"image shape: {image.shape}\")\n\n        dataset_dicts = []\n        for index, train_row in tqdm(train.iterrows(), total=len(train)):\n            record = {}\n            image_id, height, width = train_row[[\"image_id\", \"height\", \"width\"]].values\n            filename = str(imgdir / \"train\" / f\"{image_id}.png\")\n            record[\"file_name\"] = filename\n            record[\"image_id\"] = image_id\n            record[\"height\"] = resized_height\n            record[\"width\"] = resized_width\n            objs = []\n            for index2, row in train_df.query(\"image_id == @image_id\").iterrows():\n                # print(row)\n                # print(row[\"class_name\"])\n                # class_name = row[\"class_name\"]\n                class_id = row[\"class_id\"]\n                if class_id == 14:\n                    # It is \"No finding\"\n                    if use_class14:\n                        # Use this No finding class with the bbox covering all image area.\n                        bbox_resized = [0, 0, resized_width, resized_height]\n                        obj = {\n                            \"bbox\": bbox_resized,\n                            \"bbox_mode\": BoxMode.XYXY_ABS,\n                            \"category_id\": class_id,\n                        }\n                        objs.append(obj)\n                    else:\n                        # This annotator does not find anything, skip.\n                        pass\n                else:\n                    # bbox_original = [int(row[\"x_min\"]), int(row[\"y_min\"]), int(row[\"x_max\"]), int(row[\"y_max\"])]\n                    h_ratio = resized_height / height\n                    w_ratio = resized_width / width\n                    bbox_resized = [\n                        float(row[\"x_min\"]) * w_ratio,\n                        float(row[\"y_min\"]) * h_ratio,\n                        float(row[\"x_max\"]) * w_ratio,\n                        float(row[\"y_max\"]) * h_ratio,\n                    ]\n                    obj = {\n                        \"bbox\": bbox_resized,\n                        \"bbox_mode\": BoxMode.XYXY_ABS,\n                        \"category_id\": class_id,\n                    }\n                    objs.append(obj)\n            record[\"annotations\"] = objs\n            dataset_dicts.append(record)\n        with open(cache_path, mode=\"wb\") as f:\n            pickle.dump(dataset_dicts, f)\n\n    print(f\"Load from cache {cache_path}\")\n    with open(cache_path, mode=\"rb\") as f:\n        dataset_dicts = pickle.load(f)\n    if target_indices is not None:\n        dataset_dicts = [dataset_dicts[i] for i in target_indices]\n    return dataset_dicts\n\n\ndef get_vinbigdata_dicts_test(\n    imgdir: Path, test_meta: pd.DataFrame, use_cache: bool = True, debug: bool = True,\n):\n    debug_str = f\"_debug{int(debug)}\"\n    cache_path = Path(\".\") / f\"dataset_dicts_cache_test{debug_str}.pkl\"\n    if not use_cache or not cache_path.exists():\n        print(\"Creating data...\")\n        # test_meta = pd.read_csv(imgdir / \"test_meta.csv\")\n        if debug:\n            test_meta = test_meta.iloc[:500]  # For debug....\n\n        # Load 1 image to get image size.\n        image_id = test_meta.loc[0, \"image_id\"]\n        image_path = str(imgdir / \"test\" / f\"{image_id}.png\")\n        image = cv2.imread(image_path)\n        resized_height, resized_width, ch = image.shape\n        print(f\"image shape: {image.shape}\")\n\n        dataset_dicts = []\n        for index, test_meta_row in tqdm(test_meta.iterrows(), total=len(test_meta)):\n            record = {}\n\n            image_id, height, width = test_meta_row.values\n            filename = str(imgdir / \"test\" / f\"{image_id}.png\")\n            record[\"file_name\"] = filename\n            # record[\"image_id\"] = index\n            record[\"image_id\"] = image_id\n            record[\"height\"] = resized_height\n            record[\"width\"] = resized_width\n            # objs = []\n            # record[\"annotations\"] = objs\n            dataset_dicts.append(record)\n        with open(cache_path, mode=\"wb\") as f:\n            pickle.dump(dataset_dicts, f)\n\n    print(f\"Load from cache {cache_path}\")\n    with open(cache_path, mode=\"rb\") as f:\n        dataset_dicts = pickle.load(f)\n    return dataset_dicts\n","metadata":{"_kg_hide-input":true,"trusted":true,"execution":{"iopub.status.busy":"2025-04-24T13:57:53.688768Z","iopub.execute_input":"2025-04-24T13:57:53.68909Z","iopub.status.idle":"2025-04-24T13:57:54.093987Z","shell.execute_reply.started":"2025-04-24T13:57:53.689052Z","shell.execute_reply":"2025-04-24T13:57:54.093224Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# --- utils ---\nfrom pathlib import Path\nfrom typing import Any, Union\n\nimport yaml\n\n\ndef save_yaml(filepath: Union[str, Path], content: Any, width: int = 120):\n    with open(filepath, \"w\") as f:\n        yaml.dump(content, f, width=width)\n\n\ndef load_yaml(filepath: Union[str, Path]) -> Any:\n    with open(filepath, \"r\") as f:\n        content = yaml.full_load(f)\n    return content\n","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true,"execution":{"iopub.status.busy":"2025-04-24T13:57:54.095328Z","iopub.execute_input":"2025-04-24T13:57:54.095668Z","iopub.status.idle":"2025-04-24T13:57:54.102631Z","shell.execute_reply.started":"2025-04-24T13:57:54.09563Z","shell.execute_reply":"2025-04-24T13:57:54.101839Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# --- configs ---\nthing_classes = [\n    \"Aortic enlargement\",\n    \"Atelectasis\",\n    \"Calcification\",\n    \"Cardiomegaly\",\n    \"Consolidation\",\n    \"ILD\",\n    \"Infiltration\",\n    \"Lung Opacity\",\n    \"Nodule/Mass\",\n    \"Other lesion\",\n    \"Pleural effusion\",\n    \"Pleural thickening\",\n    \"Pneumothorax\",\n    \"Pulmonary fibrosis\"\n]\ncategory_name_to_id = {class_name: index for index, class_name in enumerate(thing_classes)}\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-24T13:57:54.103988Z","iopub.execute_input":"2025-04-24T13:57:54.104292Z","iopub.status.idle":"2025-04-24T13:57:54.113659Z","shell.execute_reply.started":"2025-04-24T13:57:54.104267Z","shell.execute_reply":"2025-04-24T13:57:54.113039Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<a id=\"custom_trainer\"></a>\n# Customizing detectron2 trainer\n\n※ This section is advanced, I recommend to jump to **Training scripts** section for the first time of reading.\n\nYou can refer the [Detectron2 Beginner's Tutorial](https://colab.research.google.com/drive/16jcaJoc6bCFAQ96jDe2HwtXj7BMD_-m5#scrollTo=QHnVupBBn9eR) Colab Notebook (or [version 7 of this kernel](https://www.kaggle.com/corochann/vinbigdata-detectron2-train?scriptVersionId=51628272)) for the simple usage of detectron2 how to train custom dataset. `DefaultTrainer` is used in the example which provides the starting point to train your model with custom dataset.\n\nIt is nice to start with, however I want to customize the training behavior more to improve the model's performance.\nWe can make own Trainer class (`MyTrainer` here) for this purpose, and override methods to provide customized behavior.","metadata":{}},{"cell_type":"markdown","source":"<a id=\"mapper\"></a>\n## Mapper for augmentation\n\n`Mapper` class is used inside pytorch `DataLoader`. It is responsible for converting `dataset_dicts` into actual data fed into the neural network, and we can insert augmentation process in this Mapper class.\n\n - Ref: [detectron2 docs \"Dataloader\"](https://detectron2.readthedocs.io/en/latest/tutorials/data_loading.html)\n\nI implemented `MyMapper` which uses augmentations implemented in `detectron2`, and `AlbumentationsMapper` which uses albumentations library augmentations.<br/> \nI will demonstrate these augmentations later, so you can skip reading the code and please just jump to next.","metadata":{}},{"cell_type":"code","source":"\"\"\"\nReferenced:\n - https://detectron2.readthedocs.io/en/latest/tutorials/data_loading.html\n - https://www.kaggle.com/dhiiyaur/detectron-2-compare-models-augmentation/#data\n\"\"\"\nimport copy\nimport logging\n\nimport detectron2.data.transforms as T\nimport torch\nfrom detectron2.data import detection_utils as utils\n\n\nclass MyMapper:\n    \"\"\"Mapper which uses `detectron2.data.transforms` augmentations\"\"\"\n\n    def __init__(self, cfg, is_train: bool = True):\n        aug_kwargs = cfg.aug_kwargs\n        aug_list = [\n            # T.Resize((800, 800)),\n        ]\n        if is_train:\n            aug_list.extend([getattr(T, name)(**kwargs) for name, kwargs in aug_kwargs.items()])\n        self.augmentations = T.AugmentationList(aug_list)\n        self.is_train = is_train\n\n        mode = \"training\" if is_train else \"inference\"\n        print(f\"[MyDatasetMapper] Augmentations used in {mode}: {self.augmentations}\")\n\n    def __call__(self, dataset_dict):\n        dataset_dict = copy.deepcopy(dataset_dict)  # it will be modified by code below\n        image = utils.read_image(dataset_dict[\"file_name\"], format=\"BGR\")\n\n        aug_input = T.AugInput(image)\n        transforms = self.augmentations(aug_input)\n        image = aug_input.image\n\n        # if not self.is_train:\n        #     # USER: Modify this if you want to keep them for some reason.\n        #     dataset_dict.pop(\"annotations\", None)\n        #     dataset_dict.pop(\"sem_seg_file_name\", None)\n        #     return dataset_dict\n\n        image_shape = image.shape[:2]  # h, w\n        dataset_dict[\"image\"] = torch.as_tensor(image.transpose(2, 0, 1).astype(\"float32\"))\n        annos = [\n            utils.transform_instance_annotations(obj, transforms, image_shape)\n            for obj in dataset_dict.pop(\"annotations\")\n            if obj.get(\"iscrowd\", 0) == 0\n        ]\n        instances = utils.annotations_to_instances(annos, image_shape)\n        dataset_dict[\"instances\"] = utils.filter_empty_instances(instances)\n        return dataset_dict","metadata":{"_kg_hide-input":true,"trusted":true,"execution":{"iopub.status.busy":"2025-04-24T13:57:54.114801Z","iopub.execute_input":"2025-04-24T13:57:54.115065Z","iopub.status.idle":"2025-04-24T13:57:54.190772Z","shell.execute_reply.started":"2025-04-24T13:57:54.115042Z","shell.execute_reply":"2025-04-24T13:57:54.190189Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\"\"\"\nReferenced:\n - https://detectron2.readthedocs.io/en/latest/tutorials/data_loading.html\n - https://www.kaggle.com/dhiiyaur/detectron-2-compare-models-augmentation/#data\n\"\"\"\nimport albumentations as A\nimport copy\nimport numpy as np\n\nimport torch\nfrom detectron2.data import detection_utils as utils\n\n\nclass AlbumentationsMapper:\n    \"\"\"Mapper which uses `albumentations` augmentations\"\"\"\n    def __init__(self, cfg, is_train: bool = True):\n        aug_kwargs = cfg.aug_kwargs\n        aug_list = [\n        ]\n        if is_train:\n            aug_list.extend([getattr(A, name)(**kwargs) for name, kwargs in aug_kwargs.items()])\n        self.transform = A.Compose(\n            aug_list, bbox_params=A.BboxParams(format=\"pascal_voc\", label_fields=[\"category_ids\"])\n        )\n        self.is_train = is_train\n\n        mode = \"training\" if is_train else \"inference\"\n        print(f\"[AlbumentationsMapper] Augmentations used in {mode}: {self.transform}\")\n\n    def __call__(self, dataset_dict):\n        dataset_dict = copy.deepcopy(dataset_dict)  # it will be modified by code below\n        image = utils.read_image(dataset_dict[\"file_name\"], format=\"BGR\")\n\n        # aug_input = T.AugInput(image)\n        # transforms = self.augmentations(aug_input)\n        # image = aug_input.image\n\n        prev_anno = dataset_dict[\"annotations\"]\n        bboxes = np.array([obj[\"bbox\"] for obj in prev_anno], dtype=np.float32)\n        # category_id = np.array([obj[\"category_id\"] for obj in dataset_dict[\"annotations\"]], dtype=np.int64)\n        category_id = np.arange(len(dataset_dict[\"annotations\"]))\n\n        transformed = self.transform(image=image, bboxes=bboxes, category_ids=category_id)\n        image = transformed[\"image\"]\n        annos = []\n        for i, j in enumerate(transformed[\"category_ids\"]):\n            d = prev_anno[j]\n            d[\"bbox\"] = transformed[\"bboxes\"][i]\n            annos.append(d)\n        dataset_dict.pop(\"annotations\", None)  # Remove unnecessary field.\n\n        # if not self.is_train:\n        #     # USER: Modify this if you want to keep them for some reason.\n        #     dataset_dict.pop(\"annotations\", None)\n        #     dataset_dict.pop(\"sem_seg_file_name\", None)\n        #     return dataset_dict\n\n        image_shape = image.shape[:2]  # h, w\n        dataset_dict[\"image\"] = torch.as_tensor(image.transpose(2, 0, 1).astype(\"float32\"))\n        instances = utils.annotations_to_instances(annos, image_shape)\n        dataset_dict[\"instances\"] = utils.filter_empty_instances(instances)\n        return dataset_dict\n","metadata":{"_kg_hide-input":true,"trusted":true,"execution":{"iopub.status.busy":"2025-04-24T13:57:54.191763Z","iopub.execute_input":"2025-04-24T13:57:54.192038Z","iopub.status.idle":"2025-04-24T13:57:54.512104Z","shell.execute_reply.started":"2025-04-24T13:57:54.192013Z","shell.execute_reply":"2025-04-24T13:57:54.511214Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<a id=\"loss_hook\"></a>\n## Loss evaluation hook\n\nWe implemented Evaluator and now we can calculate competition metric, however, validation loss is not calculated inside Evaluator. This is because model's evaluation is done in `model.eval()` mode and it outputs bounding box prediction but does not output `loss`.\n\nTo calculate validation loss, we need to call `model` with the training mode. This can be done by adding `Hook` which calculates the loss to the trainer.<br/>\nTrainer has attribute `storage` and calculated metrics are summarized. Its content is saved to `metric.json` (jsonl format) during training.\n\nBelow `LossEvalHook` calculates validation loss in `_do_loss_eval` method, and `self.trainer.storage.put_scalars(validation_loss=mean_loss)` is called to put this validation loss to the `storage`, which will be saved to `metrics.json`.<br/>\nNote that current implementation is not efficient in the sense that Evaluator's evaluation and `LossEvalHook`'s loss calculation run separately, even if both need a model forward calculation for same validation data.\n\n - Ref: [Training on Detectron2 with a Validation set, and plot loss on it to avoid overfitting](https://medium.com/@apofeniaco/training-on-detectron2-with-a-validation-set-and-plot-loss-on-it-to-avoid-overfitting-6449418fbf4e)","metadata":{}},{"cell_type":"code","source":"\"\"\"\nTo calculate & record validation loss\n\nOriginal code from https://medium.com/@apofeniaco/training-on-detectron2-with-a-validation-set-and-plot-loss-on-it-to-avoid-overfitting-6449418fbf4e\nby @apofeniaco\n\"\"\"\nimport numpy as np\nimport logging\n\nfrom detectron2.engine.hooks import HookBase\nfrom detectron2.utils.logger import log_every_n_seconds\nimport detectron2.utils.comm as comm\nimport torch\nimport time\nimport datetime\n\n\nclass LossEvalHook(HookBase):\n    def __init__(self, eval_period, model, data_loader):\n        self._model = model\n        self._period = eval_period\n        self._data_loader = data_loader\n\n    def _do_loss_eval(self):\n        # Copying inference_on_dataset from evaluator.py\n        total = len(self._data_loader)\n        num_warmup = min(5, total - 1)\n\n        start_time = time.perf_counter()\n        total_compute_time = 0\n        losses = []\n        for idx, inputs in enumerate(self._data_loader):\n            if idx == num_warmup:\n                start_time = time.perf_counter()\n                total_compute_time = 0\n            start_compute_time = time.perf_counter()\n            if torch.cuda.is_available():\n                torch.cuda.synchronize()\n            total_compute_time += time.perf_counter() - start_compute_time\n            iters_after_start = idx + 1 - num_warmup * int(idx >= num_warmup)\n            seconds_per_img = total_compute_time / iters_after_start\n            if idx >= num_warmup * 2 or seconds_per_img > 5:\n                total_seconds_per_img = (time.perf_counter() - start_time) / iters_after_start\n                eta = datetime.timedelta(seconds=int(total_seconds_per_img * (total - idx - 1)))\n                log_every_n_seconds(\n                    logging.INFO,\n                    \"Loss on Validation  done {}/{}. {:.4f} s / img. ETA={}\".format(\n                        idx + 1, total, seconds_per_img, str(eta)\n                    ),\n                    n=5,\n                )\n            loss_batch = self._get_loss(inputs)\n            losses.append(loss_batch)\n        mean_loss = np.mean(losses)\n        # self.trainer.storage.put_scalar('validation_loss', mean_loss)\n        comm.synchronize()\n\n        # return losses\n        return mean_loss\n\n    def _get_loss(self, data):\n        # How loss is calculated on train_loop\n        metrics_dict = self._model(data)\n        metrics_dict = {\n            k: v.detach().cpu().item() if isinstance(v, torch.Tensor) else float(v)\n            for k, v in metrics_dict.items()\n        }\n        total_losses_reduced = sum(loss for loss in metrics_dict.values())\n        return total_losses_reduced\n\n    def after_step(self):\n        next_iter = int(self.trainer.iter) + 1\n        is_final = next_iter == self.trainer.max_iter\n        if is_final or (self._period > 0 and next_iter % self._period == 0):\n            mean_loss = self._do_loss_eval()\n            self.trainer.storage.put_scalars(validation_loss=mean_loss)\n            print(\"validation do loss eval\", mean_loss)\n        else:\n            pass\n            # self.trainer.storage.put_scalars(timetest=11)\n","metadata":{"_kg_hide-input":true,"trusted":true,"execution":{"iopub.status.busy":"2025-04-24T13:57:54.51333Z","iopub.execute_input":"2025-04-24T13:57:54.513584Z","iopub.status.idle":"2025-04-24T13:57:54.581826Z","shell.execute_reply.started":"2025-04-24T13:57:54.513557Z","shell.execute_reply":"2025-04-24T13:57:54.581307Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Now all the preparation has done!\n\n`MyTrainer` overwraps `build_evaluator` method of `DefaultTrainer` provided by `detectron2` to support validation dataset evaluation.\n\n\n1. `build_train_loader` & `build_test_loader`: \nThese class methods deine how to construct DataLoader for training data & validation data respectively.\nHere `AlbumentationMapper` is passed to construct DataLoader to insert customized augmentation process.\n\n2. `build_evaluator`: COCOEvaluator\n\n3. `build_hooks`:\nThis method defines how to construct hooks. I insert `LossEvalHook` before evalutor to work well.","metadata":{}},{"cell_type":"code","source":"import os\n\nfrom detectron2.data import build_detection_test_loader, build_detection_train_loader\nfrom detectron2.engine import DefaultPredictor, DefaultTrainer, launch\n\nfrom detectron2.evaluation import COCOEvaluator\n\n\nclass MyTrainer(DefaultTrainer):\n    @classmethod\n    def build_train_loader(cls, cfg, sampler=None):\n        return build_detection_train_loader(\n            cfg, mapper=AlbumentationsMapper(cfg, True), sampler=sampler\n        )\n\n    @classmethod\n    def build_test_loader(cls, cfg, dataset_name):\n        return build_detection_test_loader(\n            cfg, dataset_name, mapper=AlbumentationsMapper(cfg, False)\n        )\n\n    @classmethod\n    def build_evaluator(cls, cfg, dataset_name, output_folder=None):\n        if output_folder is None:\n            output_folder = os.path.join(cfg.OUTPUT_DIR, \"inference\")\n        return COCOEvaluator(dataset_name, (\"bbox\",), False, output_dir=output_folder)\n\n    def build_hooks(self):\n        hooks = super(MyTrainer, self).build_hooks()\n        cfg = self.cfg\n        if len(cfg.DATASETS.TEST) > 0:\n            loss_eval_hook = LossEvalHook(\n                cfg.TEST.EVAL_PERIOD,\n                self.model,\n                MyTrainer.build_test_loader(cfg, cfg.DATASETS.TEST[0]),\n            )\n            hooks.insert(-1, loss_eval_hook)\n\n        return hooks\n","metadata":{"_kg_hide-input":true,"trusted":true,"execution":{"iopub.status.busy":"2025-04-24T13:57:54.582739Z","iopub.execute_input":"2025-04-24T13:57:54.582973Z","iopub.status.idle":"2025-04-24T13:57:54.591274Z","shell.execute_reply.started":"2025-04-24T13:57:54.582951Z","shell.execute_reply":"2025-04-24T13:57:54.590391Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## LR scheduling\n\nTo further customize learning rate scheduling you may override `build_lr_scheduler` class method to construct any pytorch LRScheduler.\n\nDefault `build_lr_schduler` method ([docs](https://detectron2.readthedocs.io/en/latest/_modules/detectron2/solver/build.html#build_lr_scheduler)) supports only 2 types of LR scheduling, WarmupMultiStepLR (default) & WarmupCosineLR. You can change which one to use by setting `cfg.SOLVER.LR_SCHEDULER_NAME` as you can see from the docs.","metadata":{}},{"cell_type":"markdown","source":"Now the methods are ready. main scripts starts from here.","metadata":{}},{"cell_type":"markdown","source":"<a id=\"load_data\"></a>\n# Loading Data\n\nThis `Flags` class is to manage experiments. I will tune these parameters through the competition to improve model's performance.","metadata":{}},{"cell_type":"code","source":"import argparse\nimport dataclasses\nimport json\nimport os\nimport pickle\nimport random\nimport sys\nfrom dataclasses import dataclass\nfrom distutils.util import strtobool\nfrom pathlib import Path\n\nimport cv2\nimport detectron2\nimport numpy as np\nimport pandas as pd\nimport torch\nfrom detectron2 import model_zoo\nfrom detectron2.config import get_cfg\nfrom detectron2.data import DatasetCatalog, MetadataCatalog\nfrom detectron2.engine import DefaultPredictor, DefaultTrainer, launch\nfrom detectron2.evaluation import COCOEvaluator\nfrom detectron2.structures import BoxMode\nfrom detectron2.utils.logger import setup_logger\nfrom detectron2.utils.visualizer import Visualizer\nfrom tqdm import tqdm\n\nsetup_logger()","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true,"execution":{"iopub.status.busy":"2025-04-24T13:57:54.592485Z","iopub.execute_input":"2025-04-24T13:57:54.592728Z","iopub.status.idle":"2025-04-24T13:57:54.608639Z","shell.execute_reply.started":"2025-04-24T13:57:54.592705Z","shell.execute_reply":"2025-04-24T13:57:54.608018Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# --- flags ---\nfrom dataclasses import dataclass, field\nfrom typing import Dict\n\n\n@dataclass\nclass Flags:\n    # General\n    debug: bool = True\n    outdir: str = \"results/det\"\n\n    # Data config\n    imgdir_name: str = \"vinbigdata-512-image-dataset/vinbigdata\"\n    split_mode: str = \"all_train\"  # all_train or valid20\n    seed: int = 111\n    train_data_type: str = \"original\"  # original or wbf\n    use_class14: bool = False\n    # Training config\n    iter: int = 10000\n    ims_per_batch: int = 2  # images per batch, this corresponds to \"total batch size\"\n    num_workers: int = 4\n    lr_scheduler_name: str = \"WarmupMultiStepLR\"  # WarmupMultiStepLR (default) or WarmupCosineLR\n    base_lr: float = 0.00025\n    roi_batch_size_per_image: int = 512\n    eval_period: int = 20\n    aug_kwargs: Dict = field(default_factory=lambda: {})\n\n    def update(self, param_dict: Dict) -> \"Flags\":\n        # Overwrite by `param_dict`\n        for key, value in param_dict.items():\n            if not hasattr(self, key):\n                raise ValueError(f\"[ERROR] Unexpected key for flag = {key}\")\n            setattr(self, key, value)\n        return self","metadata":{"_kg_hide-input":true,"trusted":true,"execution":{"iopub.status.busy":"2025-04-24T13:57:54.609669Z","iopub.execute_input":"2025-04-24T13:57:54.609913Z","iopub.status.idle":"2025-04-24T13:57:54.61972Z","shell.execute_reply.started":"2025-04-24T13:57:54.60989Z","shell.execute_reply":"2025-04-24T13:57:54.619103Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"flags_dict = {\n    \"debug\": False,\n    \"outdir\": \"results/v9\", \n    \"imgdir_name\": \"vinbigdata-512-image-dataset/vinbigdata\",\n    \"split_mode\": \"valid20\",\n    \"iter\": 10000,\n    \"roi_batch_size_per_image\": 512,\n    \"eval_period\": 1000,\n    \"lr_scheduler_name\": \"WarmupCosineLR\",\n    \"base_lr\": 0.001,\n    \"num_workers\": 4,\n    \"aug_kwargs\": {\n        \"HorizontalFlip\": {\"p\": 0.5},\n        \"ShiftScaleRotate\": {\"scale_limit\": 0.15, \"rotate_limit\": 10, \"p\": 0.5},\n        \"RandomBrightnessContrast\": {\"p\": 0.5}\n    }\n}","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-24T13:57:54.620708Z","iopub.execute_input":"2025-04-24T13:57:54.621022Z","iopub.status.idle":"2025-04-24T13:57:54.63421Z","shell.execute_reply.started":"2025-04-24T13:57:54.620998Z","shell.execute_reply":"2025-04-24T13:57:54.633565Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# args = parse()\nprint(\"torch\", torch.__version__)\nflags = Flags().update(flags_dict)\nprint(\"flags\", flags)\ndebug = flags.debug\noutdir = Path(flags.outdir)\nos.makedirs(str(outdir), exist_ok=True)\nflags_dict = dataclasses.asdict(flags)\nsave_yaml(outdir / \"flags.yaml\", flags_dict)\n# --- Read data ---\ninputdir = Path(\"/kaggle/input\")\ndatadir = inputdir / \"vinbigdata-512-image-dataset\" / \"vinbigdata\"\nimgdir = inputdir / flags.imgdir_name\n\n# Read in the data CSV files\ntrain_df = pd.read_csv(datadir / \"train.csv\")\ntrain_df = train_df[train_df.class_id!=14].reset_index(drop = True)\n\ntrain = train_df  # alias\n# sample_submission = pd.read_csv(datadir / 'sample_submission.csv')","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true,"execution":{"iopub.status.busy":"2025-04-24T13:57:54.635215Z","iopub.execute_input":"2025-04-24T13:57:54.635444Z","iopub.status.idle":"2025-04-24T13:57:54.800301Z","shell.execute_reply.started":"2025-04-24T13:57:54.635421Z","shell.execute_reply":"2025-04-24T13:57:54.799452Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_data_type = flags.train_data_type\nif flags.use_class14:\n    thing_classes.append(\"No finding\")\n\nsplit_mode = flags.split_mode\nif split_mode == \"all_train\":\n    DatasetCatalog.register(\n        \"vinbigdata_train\",\n        lambda: get_vinbigdata_dicts(\n            imgdir, train_df, train_data_type, debug=debug, use_class14=flags.use_class14\n        ),\n    )\n    MetadataCatalog.get(\"vinbigdata_train\").set(thing_classes=thing_classes)\nelif split_mode == \"valid20\":\n    # To get number of data...\n    n_dataset = len(\n        get_vinbigdata_dicts(\n            imgdir, train_df, train_data_type, debug=debug, use_class14=flags.use_class14\n        )\n    )\n    n_train = int(n_dataset * 0.8)\n    print(\"n_dataset\", n_dataset, \"n_train\", n_train)\n    rs = np.random.RandomState(flags.seed)\n    inds = rs.permutation(n_dataset)\n    train_inds, valid_inds = inds[:n_train], inds[n_train:]\n    DatasetCatalog.register(\n        \"vinbigdata_train\",\n        lambda: get_vinbigdata_dicts(\n            imgdir,\n            train_df,\n            train_data_type,\n            debug=debug,\n            target_indices=train_inds,\n            use_class14=flags.use_class14,\n        ),\n    )\n    MetadataCatalog.get(\"vinbigdata_train\").set(thing_classes=thing_classes)\n    DatasetCatalog.register(\n        \"vinbigdata_valid\",\n        lambda: get_vinbigdata_dicts(\n            imgdir,\n            train_df,\n            train_data_type,\n            debug=debug,\n            target_indices=valid_inds,\n            use_class14=flags.use_class14,\n        ),\n    )\n    MetadataCatalog.get(\"vinbigdata_valid\").set(thing_classes=thing_classes)\nelse:\n    raise ValueError(f\"[ERROR] Unexpected value split_mode={split_mode}\")\n","metadata":{"_kg_hide-input":true,"trusted":true,"execution":{"iopub.status.busy":"2025-04-24T13:57:54.801291Z","iopub.execute_input":"2025-04-24T13:57:54.801509Z","iopub.status.idle":"2025-04-24T14:02:08.491694Z","shell.execute_reply.started":"2025-04-24T13:57:54.801487Z","shell.execute_reply":"2025-04-24T14:02:08.490781Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"dataset_dicts = get_vinbigdata_dicts(imgdir, train, debug=debug)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-24T14:02:08.49309Z","iopub.execute_input":"2025-04-24T14:02:08.493403Z","iopub.status.idle":"2025-04-24T14:02:09.637071Z","shell.execute_reply.started":"2025-04-24T14:02:08.493374Z","shell.execute_reply":"2025-04-24T14:02:09.636393Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<a id=\"data_vis\"></a>\n# Data Visualization\n\nIt's also very easy to visualize prepared training dataset with `detectron2`.<br/>\nIt provides `Visualizer` class, we can use it to draw an image with bounding box as following.","metadata":{}},{"cell_type":"code","source":"# Visualize data...\nanomaly_image_ids = train.query(\"class_id != 14\")[\"image_id\"].unique()\ntrain = pd.read_csv(imgdir/\"train.csv\")\nanomaly_inds = np.argwhere(train[\"image_id\"].isin(anomaly_image_ids).values)[:, 0]\n\nvinbigdata_metadata = MetadataCatalog.get(\"vinbigdata_train\")\n\ncols = 3\nrows = 3\nfig, axes = plt.subplots(rows, cols, figsize=(18, 18))\naxes = axes.flatten()\n\nfor index, anom_ind in enumerate(anomaly_inds[:cols * rows]):\n    ax = axes[index]\n    # print(anom_ind)\n    d = dataset_dicts[anom_ind]\n    img = cv2.imread(d[\"file_name\"])\n    visualizer = Visualizer(img[:, :, ::-1], metadata=vinbigdata_metadata, scale=0.5)\n    out = visualizer.draw_dataset_dict(d)\n    # cv2_imshow(out.get_image()[:, :, ::-1])\n    #cv2.imwrite(str(outdir / f\"vinbigdata{index}.jpg\"), out.get_image()[:, :, ::-1])\n    ax.imshow(out.get_image()[:, :, ::-1])\n    ax.set_title(f\"{anom_ind}: image_id {anomaly_image_ids[index]}\")","metadata":{"_kg_hide-input":true,"trusted":true,"execution":{"iopub.status.busy":"2025-04-24T14:02:09.638091Z","iopub.execute_input":"2025-04-24T14:02:09.638324Z","iopub.status.idle":"2025-04-24T14:02:11.657331Z","shell.execute_reply.started":"2025-04-24T14:02:09.638302Z","shell.execute_reply":"2025-04-24T14:02:11.656267Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<a id=\"training\"></a>\n# Training","metadata":{}},{"cell_type":"code","source":"from detectron2.config.config import CfgNode as CN\n\ncfg = get_cfg()\ncfg.aug_kwargs = CN(flags.aug_kwargs)  # pass aug_kwargs to cfg\n\noriginal_output_dir = cfg.OUTPUT_DIR\ncfg.OUTPUT_DIR = str(outdir)\nprint(f\"cfg.OUTPUT_DIR {original_output_dir} -> {cfg.OUTPUT_DIR}\")\n\nconfig_name = \"COCO-Detection/faster_rcnn_R_101_FPN_3x.yaml\"\ncfg.merge_from_file(model_zoo.get_config_file(config_name))\ncfg.DATASETS.TRAIN = (\"vinbigdata_train\",)\nif split_mode == \"all_train\":\n    cfg.DATASETS.TEST = ()\nelse:\n    cfg.DATASETS.TEST = (\"vinbigdata_valid\",)\n    cfg.TEST.EVAL_PERIOD = flags.eval_period\n\ncfg.DATALOADER.NUM_WORKERS = flags.num_workers\n# Let training initialize from model zoo\ncfg.MODEL.WEIGHTS = model_zoo.get_checkpoint_url(config_name)\ncfg.SOLVER.IMS_PER_BATCH = flags.ims_per_batch\ncfg.SOLVER.LR_SCHEDULER_NAME = flags.lr_scheduler_name\ncfg.SOLVER.BASE_LR = flags.base_lr  # pick a good LR\ncfg.SOLVER.MAX_ITER = flags.iter\ncfg.SOLVER.CHECKPOINT_PERIOD = 100000  # Small value=Frequent save need a lot of storage.\ncfg.MODEL.ROI_HEADS.BATCH_SIZE_PER_IMAGE = flags.roi_batch_size_per_image\ncfg.MODEL.ROI_HEADS.NUM_CLASSES = len(thing_classes)\n# NOTE: this config means the number of classes,\n# but a few popular unofficial tutorials incorrect uses num_classes+1 here.\n\nos.makedirs(cfg.OUTPUT_DIR, exist_ok=True)\n","metadata":{"trusted":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2025-04-24T14:02:11.658462Z","iopub.execute_input":"2025-04-24T14:02:11.658723Z","iopub.status.idle":"2025-04-24T14:02:11.682746Z","shell.execute_reply.started":"2025-04-24T14:02:11.658695Z","shell.execute_reply":"2025-04-24T14:02:11.681897Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"trainer = MyTrainer(cfg)\ntrainer.resume_or_load(resume=False)\ntrainer.train()","metadata":{"trusted":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2025-04-24T14:02:11.68396Z","iopub.execute_input":"2025-04-24T14:02:11.684281Z","iopub.status.idle":"2025-04-24T17:40:52.736111Z","shell.execute_reply.started":"2025-04-24T14:02:11.684246Z","shell.execute_reply":"2025-04-24T17:40:52.734765Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"It's actually very easy to use multiple gpus for training.\n\nYou just need to wrap above training scripts by `main` method and use `launch` method provided by `detectron2`.\n\nPlease refer official example [train_net.py](https://github.com/facebookresearch/detectron2/blob/master/tools/train_net.py#L161) for details.","metadata":{}},{"cell_type":"markdown","source":"<a id=\"vis_loss\"></a>\n# Visualize loss curve & competition metric AP50\n\nAs I explained, the calculated metrics are saved in `metrics.json`. We can analyze/plot them to check how the training proceeded.","metadata":{}},{"cell_type":"code","source":"metrics_df = pd.read_json(outdir / \"metrics.json\", orient=\"records\", lines=True)\nmdf = metrics_df.sort_values(\"iteration\")\nmdf","metadata":{"_kg_hide-input":true,"trusted":true,"execution":{"iopub.status.busy":"2025-04-24T17:40:52.737686Z","iopub.execute_input":"2025-04-24T17:40:52.737974Z","iopub.status.idle":"2025-04-24T17:40:52.920365Z","shell.execute_reply.started":"2025-04-24T17:40:52.737947Z","shell.execute_reply":"2025-04-24T17:40:52.919573Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 1. Loss curve\nfig, ax = plt.subplots()\n\nmdf1 = mdf[~mdf[\"total_loss\"].isna()]\nax.plot(mdf1[\"iteration\"], mdf1[\"total_loss\"], c=\"C0\", label=\"train\")\nif \"validation_loss\" in mdf.columns:\n    mdf2 = mdf[~mdf[\"validation_loss\"].isna()]\n    ax.plot(mdf2[\"iteration\"], mdf2[\"validation_loss\"], c=\"C1\", label=\"validation\")\n\n# ax.set_ylim([0, 0.5])\nax.legend()\nax.set_title(\"Loss curve\")\nplt.show()\nplt.savefig(outdir/\"loss.png\")","metadata":{"_kg_hide-input":true,"trusted":true,"execution":{"iopub.status.busy":"2025-04-24T17:40:52.921694Z","iopub.execute_input":"2025-04-24T17:40:52.921985Z","iopub.status.idle":"2025-04-24T17:40:53.105042Z","shell.execute_reply.started":"2025-04-24T17:40:52.921956Z","shell.execute_reply":"2025-04-24T17:40:53.104106Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"fig, ax = plt.subplots()\nmdf3 = mdf[~mdf[\"bbox/AP50\"].isna()]\nax.plot(mdf3[\"iteration\"], mdf3[\"bbox/AP50\"] / 100., c=\"C2\", label=\"validation\")\n\nax.legend()\nax.set_title(\"AP50\")\nplt.show()\nplt.savefig(outdir / \"AP50.png\")\n\nfig, ax = plt.subplots()\nmdf3 = mdf[~mdf[\"bbox/AP75\"].isna()]\nax.plot(mdf3[\"iteration\"], mdf3[\"bbox/AP75\"] / 100., c=\"C2\", label=\"validation\")\n\nax.legend()\nax.set_title(\"AP75\")\nplt.show()\nplt.savefig(outdir / \"AP75.png\")","metadata":{"_kg_hide-input":true,"trusted":true,"execution":{"iopub.status.busy":"2025-04-24T17:40:53.106456Z","iopub.execute_input":"2025-04-24T17:40:53.106892Z","iopub.status.idle":"2025-04-24T17:40:53.419864Z","shell.execute_reply.started":"2025-04-24T17:40:53.106826Z","shell.execute_reply":"2025-04-24T17:40:53.418921Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"fig, ax = plt.subplots()\nmdf_bbox_class = mdf3.iloc[-1][[f\"bbox/AP-{col}\" for col in thing_classes]]\nmdf_bbox_class.plot(kind=\"bar\", ax=ax)\n_ = ax.set_title(\"AP by class\")","metadata":{"_kg_hide-input":true,"trusted":true,"execution":{"iopub.status.busy":"2025-04-24T17:40:53.421284Z","iopub.execute_input":"2025-04-24T17:40:53.421688Z","iopub.status.idle":"2025-04-24T17:40:53.617678Z","shell.execute_reply.started":"2025-04-24T17:40:53.421646Z","shell.execute_reply":"2025-04-24T17:40:53.616693Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Our Evaluator calculaes AP by class, and it is easy to check which class is diffucult to train.\n\nIn my experiment, **\"Calcification\" seems to be the most difficult class to predict**.","metadata":{}},{"cell_type":"markdown","source":"<a id=\"vis_aug\"></a>\n# Visualization of augmentation by Mapper\n\nLet's check the behavior of Mapper method. Since mapper is used inside DataLoader, we can check its behavior by constucting DataLoader and visualize the data processed by the DataLoader.\n\nThe defined Trainer class has **class method** `build_train_loader`. We can construct train_loader purely from `cfg`, without instantiating `trainer` since it's class method.\n\nBelow code is to visualize the same data 4 times. You can check that augmentation is applied and every time the image looks different.\n\nNote that both `detectron2.data.transforms` & `albumentations` augmentations properly handles bounding box. Thus bounding box is adjusted when the image is scaled, rotated etc!\n\nAt first I was using `detectron2.data.transforms` with `MyMapper` class, it provides basic augmentations.<br/>\nThen I noticed that we can use many augmentations in `albumentations`, so I implemented `AlbumentationsMapper` to support it.<br/>\nHow many augmentations can be used in albumentations?<br/>\nYou can see official github page, all [Pixel-level transforms](https://github.com/albumentations-team/albumentations#pixel-level-transforms) and [Spatial-level transforms](https://github.com/albumentations-team/albumentations#spatial-level-transforms) with \"BBoxes\" checked can be used. There are really many!!!","metadata":{}},{"cell_type":"code","source":"# Visualize data...\n# import matplotlib.pyplot as plt\nfrom detectron2.data.samplers import TrainingSampler\n\nn_images = 2\nn_aug = 4\n\nfig, axes = plt.subplots(n_images, n_aug, figsize=(16, 8))\n\n# Ref https://github.com/facebookresearch/detectron2/blob/22b70a8078eb09da38d0fefa130d0f537562bebc/tools/visualize_data.py#L79-L88\nfor i in range(n_aug):\n    sampler = TrainingSampler(len(dataset_dicts), shuffle=False)\n    train_vis_loader = MyTrainer.build_train_loader(\n        cfg, sampler=sampler\n    )  # For visualization...\n    for batch in train_vis_loader:\n        for j, per_image in enumerate(batch):\n            ax = axes[j, i]\n\n            img_arr = per_image[\"image\"].cpu().numpy().transpose((1, 2, 0))\n            visualizer = Visualizer(\n                img_arr[:, :, ::-1], metadata=vinbigdata_metadata, scale=1.0\n            )\n            target_fields = per_image[\"instances\"].get_fields()\n            labels = [\n                vinbigdata_metadata.thing_classes[i] for i in target_fields[\"gt_classes\"]\n            ]\n            out = visualizer.overlay_instances(\n                labels=labels,\n                boxes=target_fields.get(\"gt_boxes\", None),\n                masks=target_fields.get(\"gt_masks\", None),\n                keypoints=target_fields.get(\"gt_keypoints\", None),\n            )\n            # out = visualizer.draw_dataset_dict(per_image)\n\n            img = out.get_image()[:, :, ::-1]\n            filepath = str(outdir / f\"vinbigdata_{j}_aug{i}.jpg\")\n            cv2.imwrite(filepath, img)\n            print(f\"Visualization img {img_arr.shape} saved in {filepath}\")\n            ax.imshow(img)\n            ax.set_title(f\"image{j}, {i}-th aug\")\n        break","metadata":{"_kg_hide-input":true,"trusted":true,"execution":{"iopub.status.busy":"2025-04-24T17:40:53.619415Z","iopub.execute_input":"2025-04-24T17:40:53.619896Z","iopub.status.idle":"2025-04-24T17:41:10.444233Z","shell.execute_reply.started":"2025-04-24T17:40:53.61983Z","shell.execute_reply":"2025-04-24T17:41:10.443347Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"That's all! \n\nI found that the competition data is not so many (15000 for all images, 4000 images after filtering \"No finding\" images).<br/>\nIt does not take long time to train (less than a day), so this competition may be a good choice for beginners who want to learn object detection!\n\n<h3 style=\"color:red\">If this kernel helps you, please upvote to keep me motivated 😁<br>Thanks!</h3>","metadata":{}},{"cell_type":"markdown","source":"<a id=\"next_step\"></a>\n# Next step\n\n[📸VinBigData detectron2 prediction](https://www.kaggle.com/corochann/vinbigdata-detectron2-prediction) kernel explains how to use trained model for the prediction and submisssion for this competition.\n\n[📸VinBigData 2-class classifier complete pipeline](https://www.kaggle.com/corochann/vinbigdata-2-class-classifier-complete-pipeline) kernel explains how to train 2 class classifier model for the prediction and submisssion for this competition.\n\n## Discussions\nThese discussions are useful to further utilize this training notebook to conduct deeper experiment.\n\n - [1-step training & prediction](https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/219672): The 1-step pipeline which does not use any 2-class classifier approach is proposed.\n - [What anchor size & aspect ratio should be used?](https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/220295): Suggests how to predict more smaller sized, high aspect ratio bonding boxes. It affects to the score a lot!!!\n - [Preferable radiologist's id in the test dataset?](https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/219221): Investigation of test dataset annotation distribution.\n","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}