{
  "id": 229687,
  "title": "76th place solution - EfficientDet",
  "url": "/competitions/vinbigdata-chest-xray-abnormalities-detection/discussion/229687",
  "author_name": "Sunghyun Jun",
  "post_date": "2021-03-31T08:38:12.284000",
  "votes": 15,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Congratulations to all the winners and thanks to the organizers of this competition. And also thanks to all of the competitors. I have learned a lot from this competition.</p>\n<p>I tried to get better results, but couldn't find a better way to solve the problem.</p>\n<h2>Summary</h2>\n<p>The ensembled results of the two-stage approach and the one-stage approach. The two-stage approach consists of 2-class classifier and 14-class detector. And the one-stage approach is 15-class detector.</p>\n<h2>Tools</h2>\n<ul>\n<li>Colab Pro, Tesla V100 16GB single GPU</li>\n<li>GCS</li>\n<li>Pytorch Lightning</li>\n<li>Neptune</li>\n<li>Kaggle API</li>\n</ul>\n<h2>Validation</h2>\n<p>StratifiedKFold was used and the data set was composed of 5 folds.</p>\n<p>The classifier of the two-stage approach was trained using all normal and abnormal images, and only abnormal images were used for the detector.</p>\n<p>The one-stage detector was trained using all images of normal and abnormal.</p>\n<h2>Fusing BBoxes</h2>\n<p>The overlapped bboxes in both train set and validation set were fused using nms. I tested it using batched_nms of torchvision, nms and wbf of ZFTurbo.</p>\n<p>14-class Efficientdet d4 896px 30 epochs without classifier, local cv on positive image only</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>cv(mAP@iou=0.4)</th>\n<th>public LB</th>\n<th>private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>torchvision batched_nms</td>\n<td>0.4317</td>\n<td>0.155</td>\n<td>0.168</td>\n</tr>\n<tr>\n<td>ZFTurbo nms</td>\n<td>0.4419</td>\n<td>0.164</td>\n<td>0.181</td>\n</tr>\n<tr>\n<td>ZFTurbo wbf</td>\n<td>0.4158</td>\n<td>0.157</td>\n<td>0.185</td>\n</tr>\n</tbody>\n</table>\n<p>I thought direct comparison of local cv was not possible, and LB was also somewhat difficult to use as a criterion for judgement.</p>\n<p>After considering the labeling method of the test set, I thought that the nms method was more similar to the labeling method, so I decided to use nms.</p>\n<p><a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/214906#1199596\" target=\"_blank\">[Discussion] nms &gt; weighted box fusion?</a></p>\n<h3>torchvision's nms vs ZFTurbo's nms</h3>\n<p>Torchvision'nms has some problems in the order of results when the bbox scores are the same.</p>\n<p><a href=\"https://stackoverflow.com/questions/56176439/pytorch-argsort-ordered-with-duplicate-elements-in-the-tensor\" target=\"_blank\">[Stackoverflow] Pytorch argsort ordered, with duplicate elements in the tensor</a></p>\n<p><a href=\"https://pytorch.org/vision/stable/ops.html#torchvision.ops.nms\" target=\"_blank\">torchvision.ops.nms</a></p>\n<blockquote>\n  <p>If multiple boxes have the exact same score and satisfy the IoU criterion with respect to a reference box, the selected box is not guaranteed to be the same between CPU and GPU. This is similar to the behavior of argsort in PyTorch when repeated values are present.</p>\n</blockquote>\n<p>It doesn't matter when you convert the label only once at first and then save and load it as csv, pickle, etc., but if you convert the label and use it in a new environment, the consistency of the bbox may not be maintained.</p>\n<p>I decided to use ZFTurbo's nms which uses numpy.argsort, as this part would interfere with experiment flexibility and consistency during training.</p>\n<p><a href=\"https://github.com/ZFTurbo/Weighted-Boxes-Fusion\" target=\"_blank\">[GitHub] ZFTurbo/Weighted boxes fusion</a></p>\n<h2>Model training</h2>\n<p>All models were trained on Colab Pro's V100 16GB single GPU.</p>\n<ul>\n<li>AdamW</li>\n<li>CosineAnnealingLR</li>\n<li>epochs = 50</li>\n<li>checkpoint selection : max mAP among top-3 min val_loss</li>\n</ul>\n<h3>Two-stage approach</h3>\n<ul>\n<li>2-class classifier : EfficientNet, Resnet200d, total 15 models</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>image size(px)</th>\n<th>folds</th>\n<th>batch size</th>\n<th>init lr</th>\n<th>weight decay</th>\n<th>val_acc</th>\n<th>auc</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>b5</td>\n<td>600</td>\n<td>5 of 5</td>\n<td>12</td>\n<td>7.5e-5</td>\n<td>1.0e-4</td>\n<td>0.9601</td>\n<td>0.9930</td>\n</tr>\n<tr>\n<td>b6</td>\n<td>528</td>\n<td>5 of 5</td>\n<td>12</td>\n<td>7.5e-5</td>\n<td>1.0e-3</td>\n<td>0.9553</td>\n<td>0.9927</td>\n</tr>\n<tr>\n<td>resnet200d</td>\n<td>600</td>\n<td>3 of 5</td>\n<td>12</td>\n<td>7.5e-5</td>\n<td>1.0e-4</td>\n<td>0.9541</td>\n<td>0.9934</td>\n</tr>\n<tr>\n<td>b5</td>\n<td>456</td>\n<td>single</td>\n<td>16</td>\n<td>1.0e-4</td>\n<td>1.0e-4</td>\n<td>0.9557</td>\n<td>0.9927</td>\n</tr>\n<tr>\n<td>b5</td>\n<td>1024</td>\n<td>single</td>\n<td>4</td>\n<td>2.5e-5</td>\n<td>1.0e-4</td>\n<td>0.9577</td>\n<td>0.9936</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>14-class detector : EfficientDet with 2-class classifier, total 18 models, local cv on positive image only</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>image size(px)</th>\n<th>folds</th>\n<th>batch size</th>\n<th>init lr</th>\n<th>weight decay</th>\n<th>cv(mAP@iou=0.4)</th>\n<th>public LB</th>\n<th>private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>d3</td>\n<td>1024</td>\n<td>single</td>\n<td>3</td>\n<td>3e-4</td>\n<td>1e-3</td>\n<td>0.4545</td>\n<td>0.209</td>\n<td>0.250</td>\n</tr>\n<tr>\n<td>d4</td>\n<td>896</td>\n<td>5 of 5</td>\n<td>4</td>\n<td>4e-4</td>\n<td>1e-4</td>\n<td>0.4541</td>\n<td>0.218</td>\n<td>0.250</td>\n</tr>\n<tr>\n<td>d4</td>\n<td>896</td>\n<td>single</td>\n<td>4</td>\n<td>4e-4</td>\n<td>1e-3</td>\n<td>0.4606</td>\n<td>0.257</td>\n<td>0.247</td>\n</tr>\n<tr>\n<td>d4</td>\n<td>1024</td>\n<td>single</td>\n<td>3</td>\n<td>3e-4</td>\n<td>1e-3</td>\n<td>0.4545</td>\n<td>0.228</td>\n<td>0.249</td>\n</tr>\n<tr>\n<td>d5</td>\n<td>768</td>\n<td>5 of 5</td>\n<td>4</td>\n<td>4e-4</td>\n<td>1e-3</td>\n<td>0.4472</td>\n<td>0.225</td>\n<td>0.253</td>\n</tr>\n<tr>\n<td>d5</td>\n<td>896</td>\n<td>4 of 5</td>\n<td>3</td>\n<td>3e-4</td>\n<td>1e-3</td>\n<td>0.4522</td>\n<td>0.214</td>\n<td>0.250</td>\n</tr>\n<tr>\n<td>d5</td>\n<td>1024</td>\n<td>single</td>\n<td>2</td>\n<td>2e-4</td>\n<td>1e-3</td>\n<td>0.4462</td>\n<td>0.214</td>\n<td>0.232</td>\n</tr>\n</tbody>\n</table>\n<h3>One-stage approach</h3>\n<ul>\n<li>15-class detector, total 2 models, local cv on positive image only</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>image size(px)</th>\n<th>folds</th>\n<th>batch size</th>\n<th>init lr</th>\n<th>weight decay</th>\n<th>cv(mAP@iou=0.4)</th>\n<th>public LB</th>\n<th>private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>d4</td>\n<td>896</td>\n<td>2 of 5</td>\n<td>4</td>\n<td>4e-4</td>\n<td>1e-3</td>\n<td>0.4546</td>\n<td>0.230</td>\n<td>0.246</td>\n</tr>\n</tbody>\n</table>\n<p>At batch size &lt; 4, the mAP result was poor. The larger the image size, the better the mAP, but no further training was possible.</p>\n<p>I tried Freeze BatchNorm, accumulate grad batches, and GroupNorm but I didn't get any better results.</p>\n<p>I also tested Downconv, but the AP of small objects like Calcification increased, but the AP of ILD and large objects decreased, so the overall mAP was not improved.</p>\n<p><a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/226557\" target=\"_blank\">[RANZCR CLiP] 11th Place Solution - Utilizing High resolution, Annotations, and Unlabeled data</a></p>\n<h2>Augmentation</h2>\n<p>Resize, scale, and crop were configured by referring to the method in the paper of EfficientDet.</p>\n<p>CLAHE, equalize, invertimg, huesaturationvalue, randomgamma, shiftscalerotate did not work.</p>\n<pre><code>A.Compose(\n[\n    A.Resize(height=self.resize_height, width=self.resize_width),\n    A.RandomScale(scale_limit=(-, ), p=),\n    A.PadIfNeeded(\n        min_height=self.resize_height,\n        min_width=self.resize_width,\n        border_mode=cv2.BORDER_CONSTANT,\n        value=,\n        p=,\n    ),\n    A.RandomCrop(height=self.resize_height, width=self.resize_width, p=),\n    A.RandomBrightnessContrast(p=),\n    A.ChannelDropout(p=),\n    A.OneOf(\n        [\n            A.MotionBlur(p=),\n            A.MedianBlur(p=),\n            A.GaussianBlur(p=),\n            A.GaussNoise(p=),\n        ],\n        p=,\n    ),\n    A.HorizontalFlip(p=),\n    A.Normalize(),\n    ToTensorV2(),\n],\n</code></pre>\n<h2>Post processing</h2>\n<p>In the case of the two-stage approach, it was difficult to find an appropriate threshold for the classifier. The threshold was determined to be a 60~70% normal case. And if it was normal, all detections of the detector were excluded.</p>\n<p>In the case of the one-stage approach, similarly, detections were excluded if it was normal based on the threshold.</p>\n<h2>Blending</h2>\n<p>The results of the two-stage approach and the one-stage approach were blended using nms, and the final result is as follows.</p>\n<h3>two-stage approach, 15 classifier + 18 detector</h3>\n<table>\n<thead>\n<tr>\n<th>normal thr</th>\n<th>public LB</th>\n<th>private LB</th>\n<th></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.70</td>\n<td>0.219</td>\n<td><strong>0.246</strong></td>\n<td><strong>final submission</strong></td>\n</tr>\n<tr>\n<td>0.65</td>\n<td>0.217</td>\n<td>0.255</td>\n<td></td>\n</tr>\n<tr>\n<td>0.60</td>\n<td>0.215</td>\n<td>0.256</td>\n<td></td>\n</tr>\n</tbody>\n</table>\n<h3>one-stage approach, 2 detector</h3>\n<table>\n<thead>\n<tr>\n<th>normal thr</th>\n<th>public LB</th>\n<th>private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.10</td>\n<td>0.198</td>\n<td>0.245</td>\n</tr>\n<tr>\n<td>0.00</td>\n<td>0.198</td>\n<td>0.245</td>\n</tr>\n</tbody>\n</table>\n<h3>two-stage + one-stage</h3>\n<table>\n<thead>\n<tr>\n<th>thr</th>\n<th>public LB</th>\n<th>private LB</th>\n<th></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.70-0.30</td>\n<td>0.224</td>\n<td><strong>0.253</strong></td>\n<td><strong>final submission</strong></td>\n</tr>\n<tr>\n<td>0.65-0.30</td>\n<td>0.222</td>\n<td>0.259</td>\n<td></td>\n</tr>\n<tr>\n<td>0.60-0.30</td>\n<td>0.220</td>\n<td>0.258</td>\n<td></td>\n</tr>\n</tbody>\n</table>\n<h2>What did not work</h2>\n<p>Freeze BatchNorm</p>\n<p>accumulate grad batches</p>\n<p>GroupNorm, group per channel = 8</p>\n<p>GroupNorm, pretrained backbone with GroupNorm</p>\n<p>Downconv</p>\n<p>NIH dataset concat classifier</p>\n<p>change aspect ratio</p>\n<h2>Source Code</h2>\n<p>Source code is available at <a href=\"https://github.com/sunghyunjun/kaggle-vinbigdata-chest-xray-abnormalities-detection\" target=\"_blank\">https://github.com/sunghyunjun/kaggle-vinbigdata-chest-xray-abnormalities-detection</a></p>\n<p>Submission notebook is <a href=\"https://www.kaggle.com/sunghyunjun/91th-place-vbd-inference-clf-det-and-det-all\" target=\"_blank\">91th place VBD inference CLF DET and DET ALL</a></p>",
  "messages": [
    {
      "id": 1257962,
      "postDate": "2021-03-31T08:38:12.283Z",
      "content": "<p>Congratulations to all the winners and thanks to the organizers of this competition. And also thanks to all of the competitors. I have learned a lot from this competition.</p>\n<p>I tried to get better results, but couldn't find a better way to solve the problem.</p>\n<h2>Summary</h2>\n<p>The ensembled results of the two-stage approach and the one-stage approach. The two-stage approach consists of 2-class classifier and 14-class detector. And the one-stage approach is 15-class detector.</p>\n<h2>Tools</h2>\n<ul>\n<li>Colab Pro, Tesla V100 16GB single GPU</li>\n<li>GCS</li>\n<li>Pytorch Lightning</li>\n<li>Neptune</li>\n<li>Kaggle API</li>\n</ul>\n<h2>Validation</h2>\n<p>StratifiedKFold was used and the data set was composed of 5 folds.</p>\n<p>The classifier of the two-stage approach was trained using all normal and abnormal images, and only abnormal images were used for the detector.</p>\n<p>The one-stage detector was trained using all images of normal and abnormal.</p>\n<h2>Fusing BBoxes</h2>\n<p>The overlapped bboxes in both train set and validation set were fused using nms. I tested it using batched_nms of torchvision, nms and wbf of ZFTurbo.</p>\n<p>14-class Efficientdet d4 896px 30 epochs without classifier, local cv on positive image only</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>cv(mAP@iou=0.4)</th>\n<th>public LB</th>\n<th>private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>torchvision batched_nms</td>\n<td>0.4317</td>\n<td>0.155</td>\n<td>0.168</td>\n</tr>\n<tr>\n<td>ZFTurbo nms</td>\n<td>0.4419</td>\n<td>0.164</td>\n<td>0.181</td>\n</tr>\n<tr>\n<td>ZFTurbo wbf</td>\n<td>0.4158</td>\n<td>0.157</td>\n<td>0.185</td>\n</tr>\n</tbody>\n</table>\n<p>I thought direct comparison of local cv was not possible, and LB was also somewhat difficult to use as a criterion for judgement.</p>\n<p>After considering the labeling method of the test set, I thought that the nms method was more similar to the labeling method, so I decided to use nms.</p>\n<p><a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/214906#1199596\" target=\"_blank\">[Discussion] nms &gt; weighted box fusion?</a></p>\n<h3>torchvision's nms vs ZFTurbo's nms</h3>\n<p>Torchvision'nms has some problems in the order of results when the bbox scores are the same.</p>\n<p><a href=\"https://stackoverflow.com/questions/56176439/pytorch-argsort-ordered-with-duplicate-elements-in-the-tensor\" target=\"_blank\">[Stackoverflow] Pytorch argsort ordered, with duplicate elements in the tensor</a></p>\n<p><a href=\"https://pytorch.org/vision/stable/ops.html#torchvision.ops.nms\" target=\"_blank\">torchvision.ops.nms</a></p>\n<blockquote>\n  <p>If multiple boxes have the exact same score and satisfy the IoU criterion with respect to a reference box, the selected box is not guaranteed to be the same between CPU and GPU. This is similar to the behavior of argsort in PyTorch when repeated values are present.</p>\n</blockquote>\n<p>It doesn't matter when you convert the label only once at first and then save and load it as csv, pickle, etc., but if you convert the label and use it in a new environment, the consistency of the bbox may not be maintained.</p>\n<p>I decided to use ZFTurbo's nms which uses numpy.argsort, as this part would interfere with experiment flexibility and consistency during training.</p>\n<p><a href=\"https://github.com/ZFTurbo/Weighted-Boxes-Fusion\" target=\"_blank\">[GitHub] ZFTurbo/Weighted boxes fusion</a></p>\n<h2>Model training</h2>\n<p>All models were trained on Colab Pro's V100 16GB single GPU.</p>\n<ul>\n<li>AdamW</li>\n<li>CosineAnnealingLR</li>\n<li>epochs = 50</li>\n<li>checkpoint selection : max mAP among top-3 min val_loss</li>\n</ul>\n<h3>Two-stage approach</h3>\n<ul>\n<li>2-class classifier : EfficientNet, Resnet200d, total 15 models</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>image size(px)</th>\n<th>folds</th>\n<th>batch size</th>\n<th>init lr</th>\n<th>weight decay</th>\n<th>val_acc</th>\n<th>auc</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>b5</td>\n<td>600</td>\n<td>5 of 5</td>\n<td>12</td>\n<td>7.5e-5</td>\n<td>1.0e-4</td>\n<td>0.9601</td>\n<td>0.9930</td>\n</tr>\n<tr>\n<td>b6</td>\n<td>528</td>\n<td>5 of 5</td>\n<td>12</td>\n<td>7.5e-5</td>\n<td>1.0e-3</td>\n<td>0.9553</td>\n<td>0.9927</td>\n</tr>\n<tr>\n<td>resnet200d</td>\n<td>600</td>\n<td>3 of 5</td>\n<td>12</td>\n<td>7.5e-5</td>\n<td>1.0e-4</td>\n<td>0.9541</td>\n<td>0.9934</td>\n</tr>\n<tr>\n<td>b5</td>\n<td>456</td>\n<td>single</td>\n<td>16</td>\n<td>1.0e-4</td>\n<td>1.0e-4</td>\n<td>0.9557</td>\n<td>0.9927</td>\n</tr>\n<tr>\n<td>b5</td>\n<td>1024</td>\n<td>single</td>\n<td>4</td>\n<td>2.5e-5</td>\n<td>1.0e-4</td>\n<td>0.9577</td>\n<td>0.9936</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>14-class detector : EfficientDet with 2-class classifier, total 18 models, local cv on positive image only</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>image size(px)</th>\n<th>folds</th>\n<th>batch size</th>\n<th>init lr</th>\n<th>weight decay</th>\n<th>cv(mAP@iou=0.4)</th>\n<th>public LB</th>\n<th>private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>d3</td>\n<td>1024</td>\n<td>single</td>\n<td>3</td>\n<td>3e-4</td>\n<td>1e-3</td>\n<td>0.4545</td>\n<td>0.209</td>\n<td>0.250</td>\n</tr>\n<tr>\n<td>d4</td>\n<td>896</td>\n<td>5 of 5</td>\n<td>4</td>\n<td>4e-4</td>\n<td>1e-4</td>\n<td>0.4541</td>\n<td>0.218</td>\n<td>0.250</td>\n</tr>\n<tr>\n<td>d4</td>\n<td>896</td>\n<td>single</td>\n<td>4</td>\n<td>4e-4</td>\n<td>1e-3</td>\n<td>0.4606</td>\n<td>0.257</td>\n<td>0.247</td>\n</tr>\n<tr>\n<td>d4</td>\n<td>1024</td>\n<td>single</td>\n<td>3</td>\n<td>3e-4</td>\n<td>1e-3</td>\n<td>0.4545</td>\n<td>0.228</td>\n<td>0.249</td>\n</tr>\n<tr>\n<td>d5</td>\n<td>768</td>\n<td>5 of 5</td>\n<td>4</td>\n<td>4e-4</td>\n<td>1e-3</td>\n<td>0.4472</td>\n<td>0.225</td>\n<td>0.253</td>\n</tr>\n<tr>\n<td>d5</td>\n<td>896</td>\n<td>4 of 5</td>\n<td>3</td>\n<td>3e-4</td>\n<td>1e-3</td>\n<td>0.4522</td>\n<td>0.214</td>\n<td>0.250</td>\n</tr>\n<tr>\n<td>d5</td>\n<td>1024</td>\n<td>single</td>\n<td>2</td>\n<td>2e-4</td>\n<td>1e-3</td>\n<td>0.4462</td>\n<td>0.214</td>\n<td>0.232</td>\n</tr>\n</tbody>\n</table>\n<h3>One-stage approach</h3>\n<ul>\n<li>15-class detector, total 2 models, local cv on positive image only</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>image size(px)</th>\n<th>folds</th>\n<th>batch size</th>\n<th>init lr</th>\n<th>weight decay</th>\n<th>cv(mAP@iou=0.4)</th>\n<th>public LB</th>\n<th>private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>d4</td>\n<td>896</td>\n<td>2 of 5</td>\n<td>4</td>\n<td>4e-4</td>\n<td>1e-3</td>\n<td>0.4546</td>\n<td>0.230</td>\n<td>0.246</td>\n</tr>\n</tbody>\n</table>\n<p>At batch size &lt; 4, the mAP result was poor. The larger the image size, the better the mAP, but no further training was possible.</p>\n<p>I tried Freeze BatchNorm, accumulate grad batches, and GroupNorm but I didn't get any better results.</p>\n<p>I also tested Downconv, but the AP of small objects like Calcification increased, but the AP of ILD and large objects decreased, so the overall mAP was not improved.</p>\n<p><a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/226557\" target=\"_blank\">[RANZCR CLiP] 11th Place Solution - Utilizing High resolution, Annotations, and Unlabeled data</a></p>\n<h2>Augmentation</h2>\n<p>Resize, scale, and crop were configured by referring to the method in the paper of EfficientDet.</p>\n<p>CLAHE, equalize, invertimg, huesaturationvalue, randomgamma, shiftscalerotate did not work.</p>\n<pre><code>A.Compose(\n[\n    A.Resize(height=self.resize_height, width=self.resize_width),\n    A.RandomScale(scale_limit=(-, ), p=),\n    A.PadIfNeeded(\n        min_height=self.resize_height,\n        min_width=self.resize_width,\n        border_mode=cv2.BORDER_CONSTANT,\n        value=,\n        p=,\n    ),\n    A.RandomCrop(height=self.resize_height, width=self.resize_width, p=),\n    A.RandomBrightnessContrast(p=),\n    A.ChannelDropout(p=),\n    A.OneOf(\n        [\n            A.MotionBlur(p=),\n            A.MedianBlur(p=),\n            A.GaussianBlur(p=),\n            A.GaussNoise(p=),\n        ],\n        p=,\n    ),\n    A.HorizontalFlip(p=),\n    A.Normalize(),\n    ToTensorV2(),\n],\n</code></pre>\n<h2>Post processing</h2>\n<p>In the case of the two-stage approach, it was difficult to find an appropriate threshold for the classifier. The threshold was determined to be a 60~70% normal case. And if it was normal, all detections of the detector were excluded.</p>\n<p>In the case of the one-stage approach, similarly, detections were excluded if it was normal based on the threshold.</p>\n<h2>Blending</h2>\n<p>The results of the two-stage approach and the one-stage approach were blended using nms, and the final result is as follows.</p>\n<h3>two-stage approach, 15 classifier + 18 detector</h3>\n<table>\n<thead>\n<tr>\n<th>normal thr</th>\n<th>public LB</th>\n<th>private LB</th>\n<th></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.70</td>\n<td>0.219</td>\n<td><strong>0.246</strong></td>\n<td><strong>final submission</strong></td>\n</tr>\n<tr>\n<td>0.65</td>\n<td>0.217</td>\n<td>0.255</td>\n<td></td>\n</tr>\n<tr>\n<td>0.60</td>\n<td>0.215</td>\n<td>0.256</td>\n<td></td>\n</tr>\n</tbody>\n</table>\n<h3>one-stage approach, 2 detector</h3>\n<table>\n<thead>\n<tr>\n<th>normal thr</th>\n<th>public LB</th>\n<th>private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.10</td>\n<td>0.198</td>\n<td>0.245</td>\n</tr>\n<tr>\n<td>0.00</td>\n<td>0.198</td>\n<td>0.245</td>\n</tr>\n</tbody>\n</table>\n<h3>two-stage + one-stage</h3>\n<table>\n<thead>\n<tr>\n<th>thr</th>\n<th>public LB</th>\n<th>private LB</th>\n<th></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.70-0.30</td>\n<td>0.224</td>\n<td><strong>0.253</strong></td>\n<td><strong>final submission</strong></td>\n</tr>\n<tr>\n<td>0.65-0.30</td>\n<td>0.222</td>\n<td>0.259</td>\n<td></td>\n</tr>\n<tr>\n<td>0.60-0.30</td>\n<td>0.220</td>\n<td>0.258</td>\n<td></td>\n</tr>\n</tbody>\n</table>\n<h2>What did not work</h2>\n<p>Freeze BatchNorm</p>\n<p>accumulate grad batches</p>\n<p>GroupNorm, group per channel = 8</p>\n<p>GroupNorm, pretrained backbone with GroupNorm</p>\n<p>Downconv</p>\n<p>NIH dataset concat classifier</p>\n<p>change aspect ratio</p>\n<h2>Source Code</h2>\n<p>Source code is available at <a href=\"https://github.com/sunghyunjun/kaggle-vinbigdata-chest-xray-abnormalities-detection\" target=\"_blank\">https://github.com/sunghyunjun/kaggle-vinbigdata-chest-xray-abnormalities-detection</a></p>\n<p>Submission notebook is <a href=\"https://www.kaggle.com/sunghyunjun/91th-place-vbd-inference-clf-det-and-det-all\" target=\"_blank\">91th place VBD inference CLF DET and DET ALL</a></p>",
      "rawMarkdown": "Congratulations to all the winners and thanks to the organizers of this competition. And also thanks to all of the competitors. I have learned a lot from this competition.\n\nI tried to get better results, but couldn't find a better way to solve the problem.\n\n## Summary\n\nThe ensembled results of the two-stage approach and the one-stage approach. The two-stage approach consists of 2-class classifier and 14-class detector. And the one-stage approach is 15-class detector.\n\n## Tools\n\n- Colab Pro, Tesla V100 16GB single GPU\n- GCS\n- Pytorch Lightning\n- Neptune\n- Kaggle API\n\n## Validation\n\nStratifiedKFold was used and the data set was composed of 5 folds.\n\nThe classifier of the two-stage approach was trained using all normal and abnormal images, and only abnormal images were used for the detector.\n\nThe one-stage detector was trained using all images of normal and abnormal.\n\n## Fusing BBoxes\n\nThe overlapped bboxes in both train set and validation set were fused using nms. I tested it using batched_nms of torchvision, nms and wbf of ZFTurbo.\n\n14-class Efficientdet d4 896px 30 epochs without classifier, local cv on positive image only\n||cv(mAP@iou=0.4)|public LB|private LB|\n|---|---|---|---|\n|torchvision batched_nms|0.4317|0.155|0.168|\n|ZFTurbo nms|0.4419|0.164|0.181|\n|ZFTurbo wbf|0.4158|0.157|0.185|\n\nI thought direct comparison of local cv was not possible, and LB was also somewhat difficult to use as a criterion for judgement.\n\nAfter considering the labeling method of the test set, I thought that the nms method was more similar to the labeling method, so I decided to use nms.\n\n[[Discussion] nms > weighted box fusion?](https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/214906#1199596)\n\n### torchvision's nms vs ZFTurbo's nms\n\nTorchvision'nms has some problems in the order of results when the bbox scores are the same.\n\n[[Stackoverflow] Pytorch argsort ordered, with duplicate elements in the tensor](https://stackoverflow.com/questions/56176439/pytorch-argsort-ordered-with-duplicate-elements-in-the-tensor)\n\n[torchvision.ops.nms](https://pytorch.org/vision/stable/ops.html#torchvision.ops.nms)\n\n> If multiple boxes have the exact same score and satisfy the IoU criterion with respect to a reference box, the selected box is not guaranteed to be the same between CPU and GPU. This is similar to the behavior of argsort in PyTorch when repeated values are present.\n\nIt doesn't matter when you convert the label only once at first and then save and load it as csv, pickle, etc., but if you convert the label and use it in a new environment, the consistency of the bbox may not be maintained.\n\nI decided to use ZFTurbo's nms which uses numpy.argsort, as this part would interfere with experiment flexibility and consistency during training.\n\n[[GitHub] ZFTurbo/Weighted boxes fusion](https://github.com/ZFTurbo/Weighted-Boxes-Fusion)\n\n## Model training\n\nAll models were trained on Colab Pro's V100 16GB single GPU.\n\n- AdamW\n- CosineAnnealingLR\n- epochs = 50\n- checkpoint selection : max mAP among top-3 min val_loss\n\n### Two-stage approach\n\n- 2-class classifier : EfficientNet, Resnet200d, total 15 models\n\n|model|image size(px)|folds|batch size|init lr|weight decay|val_acc|auc|\n|---|---|---|---|---|---|---|---|\n|b5|600|5 of 5|12|7.5e-5|1.0e-4|0.9601|0.9930|\n|b6|528|5 of 5|12|7.5e-5|1.0e-3|0.9553|0.9927|\n|resnet200d|600|3 of 5|12|7.5e-5|1.0e-4|0.9541|0.9934|\n|b5|456|single|16|1.0e-4|1.0e-4|0.9557|0.9927|\n|b5|1024|single|4|2.5e-5|1.0e-4|0.9577|0.9936|\n\n- 14-class detector : EfficientDet with 2-class classifier, total 18 models, local cv on positive image only\n\n|model|image size(px)|folds|batch size|init lr|weight decay|cv(mAP@iou=0.4)|public LB|private LB|\n|---|---|---|---|---|---|---|---|---|\n|d3|1024|single|3|3e-4|1e-3|0.4545|0.209|0.250|\n|d4|896|5 of 5|4|4e-4|1e-4|0.4541|0.218|0.250|\n|d4|896|single|4|4e-4|1e-3|0.4606|0.257|0.247|\n|d4|1024|single|3|3e-4|1e-3|0.4545|0.228|0.249|\n|d5|768|5 of 5|4|4e-4|1e-3|0.4472|0.225|0.253|\n|d5|896|4 of 5|3|3e-4|1e-3|0.4522|0.214|0.250|\n|d5|1024|single|2|2e-4|1e-3|0.4462|0.214|0.232|\n\n### One-stage approach\n\n- 15-class detector, total 2 models, local cv on positive image only\n\n|model|image size(px)|folds|batch size|init lr|weight decay|cv(mAP@iou=0.4)|public LB|private LB|\n|---|---|---|---|---|---|---|---|---|\n|d4|896|2 of 5|4|4e-4|1e-3|0.4546|0.230|0.246|\n\nAt batch size < 4, the mAP result was poor. The larger the image size, the better the mAP, but no further training was possible.\n\nI tried Freeze BatchNorm, accumulate grad batches, and GroupNorm but I didn't get any better results.\n\nI also tested Downconv, but the AP of small objects like Calcification increased, but the AP of ILD and large objects decreased, so the overall mAP was not improved.\n\n[[RANZCR CLiP] 11th Place Solution - Utilizing High resolution, Annotations, and Unlabeled data](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/226557)\n\n## Augmentation\n\nResize, scale, and crop were configured by referring to the method in the paper of EfficientDet.\n\nCLAHE, equalize, invertimg, huesaturationvalue, randomgamma, shiftscalerotate did not work.\n\n```python\nA.Compose(\n[\n    A.Resize(height=self.resize_height, width=self.resize_width),\n    A.RandomScale(scale_limit=(-0.9, 1.0), p=1.0),\n    A.PadIfNeeded(\n        min_height=self.resize_height,\n        min_width=self.resize_width,\n        border_mode=cv2.BORDER_CONSTANT,\n        value=0,\n        p=1.0,\n    ),\n    A.RandomCrop(height=self.resize_height, width=self.resize_width, p=1.0),\n    A.RandomBrightnessContrast(p=0.8),\n    A.ChannelDropout(p=0.5),\n    A.OneOf(\n        [\n            A.MotionBlur(p=0.5),\n            A.MedianBlur(p=0.5),\n            A.GaussianBlur(p=0.5),\n            A.GaussNoise(p=0.5),\n        ],\n        p=0.5,\n    ),\n    A.HorizontalFlip(p=0.5),\n    A.Normalize(),\n    ToTensorV2(),\n],\n```\n\n## Post processing\n\nIn the case of the two-stage approach, it was difficult to find an appropriate threshold for the classifier. The threshold was determined to be a 60~70% normal case. And if it was normal, all detections of the detector were excluded.\n\nIn the case of the one-stage approach, similarly, detections were excluded if it was normal based on the threshold.\n\n## Blending\n\nThe results of the two-stage approach and the one-stage approach were blended using nms, and the final result is as follows.\n\n### two-stage approach, 15 classifier + 18 detector\n\n|normal thr|public LB|private LB||\n|---|---|---|---|\n|0.70|0.219|**0.246**|**final submission**|\n|0.65|0.217|0.255|\n|0.60|0.215|0.256|\n\n### one-stage approach, 2 detector\n\n|normal thr|public LB|private LB|\n|---|---|---|\n|0.10|0.198|0.245|\n|0.00|0.198|0.245|\n\n### two-stage + one-stage\n\n|thr|public LB|private LB||\n|---|---|---|---|\n|0.70-0.30|0.224|**0.253**|**final submission**|\n|0.65-0.30|0.222|0.259|\n|0.60-0.30|0.220|0.258|\n\n## What did not work\n\nFreeze BatchNorm\n\naccumulate grad batches\n\nGroupNorm, group per channel = 8\n\nGroupNorm, pretrained backbone with GroupNorm\n\nDownconv\n\nNIH dataset concat classifier\n\nchange aspect ratio\n\n## Source Code\n\nSource code is available at [https://github.com/sunghyunjun/kaggle-vinbigdata-chest-xray-abnormalities-detection](https://github.com/sunghyunjun/kaggle-vinbigdata-chest-xray-abnormalities-detection)\n\nSubmission notebook is [91th place VBD inference CLF DET and DET ALL](https://www.kaggle.com/sunghyunjun/91th-place-vbd-inference-clf-det-and-det-all)",
      "votes": 15
    },
    {
      "id": 1257968,
      "postDate": "2021-03-31T08:45:04.840Z",
      "content": "<p><a href=\"https://www.kaggle.com/sunghyunjun\" target=\"_blank\">@sunghyunjun</a> Congratulations on Bronze medal and thanks for sharing your approach</p>",
      "rawMarkdown": "@sunghyunjun Congratulations on Bronze medal and thanks for sharing your approach",
      "votes": 1,
      "replies": [
        {
          "id": 1257979,
          "postDate": "2021-03-31T08:58:16.177Z",
          "content": "<p>Thank you!</p>",
          "rawMarkdown": "Thank you!"
        }
      ]
    },
    {
      "id": 1262462,
      "postDate": "2021-04-04T10:17:40.957Z",
      "content": "<p>Congratulations on Bronze medal<br>\nCan you share the EfficientDet CV and LB scoring please</p>",
      "rawMarkdown": "Congratulations on Bronze medal\nCan you share the EfficientDet CV and LB scoring please",
      "replies": [
        {
          "id": 1263167,
          "postDate": "2021-04-05T06:28:24.387Z",
          "content": "<p>Updated public, private LB. Local CV is mAP(@iou=0.4) and positive images only.</p>",
          "rawMarkdown": "Updated public, private LB. Local CV is mAP(@iou=0.4) and positive images only."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1257968,
      "author_name": "Tensor Girl",
      "author_url": "",
      "post_date": "2021-03-31T08:45:04.840000",
      "content": "<p><a href=\"https://www.kaggle.com/sunghyunjun\" target=\"_blank\">@sunghyunjun</a> Congratulations on Bronze medal and thanks for sharing your approach</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1257979,
          "author_name": "Sunghyun Jun",
          "author_url": "",
          "post_date": "2021-03-31T08:58:16.177000",
          "content": "<p>Thank you!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1262462,
      "author_name": "Mohammed Rizin V K",
      "author_url": "",
      "post_date": "2021-04-04T10:17:40.957000",
      "content": "<p>Congratulations on Bronze medal<br>\nCan you share the EfficientDet CV and LB scoring please</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1263167,
          "author_name": "Sunghyun Jun",
          "author_url": "",
          "post_date": "2021-04-05T06:28:24.387000",
          "content": "<p>Updated public, private LB. Local CV is mAP(@iou=0.4) and positive images only.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1257962": "Congratulations to all the winners and thanks to the organizers of this competition. And also thanks to all of the competitors. I have learned a lot from this competition.\n\nI tried to get better results, but couldn't find a better way to solve the problem.\n\n## Summary\n\nThe ensembled results of the two-stage approach and the one-stage approach. The two-stage approach consists of 2-class classifier and 14-class detector. And the one-stage approach is 15-class detector.\n\n## Tools\n\n- Colab Pro, Tesla V100 16GB single GPU\n- GCS\n- Pytorch Lightning\n- Neptune\n- Kaggle API\n\n## Validation\n\nStratifiedKFold was used and the data set was composed of 5 folds.\n\nThe classifier of the two-stage approach was trained using all normal and abnormal images, and only abnormal images were used for the detector.\n\nThe one-stage detector was trained using all images of normal and abnormal.\n\n## Fusing BBoxes\n\nThe overlapped bboxes in both train set and validation set were fused using nms. I tested it using batched_nms of torchvision, nms and wbf of ZFTurbo.\n\n14-class Efficientdet d4 896px 30 epochs without classifier, local cv on positive image only\n||cv(mAP@iou=0.4)|public LB|private LB|\n|---|---|---|---|\n|torchvision batched_nms|0.4317|0.155|0.168|\n|ZFTurbo nms|0.4419|0.164|0.181|\n|ZFTurbo wbf|0.4158|0.157|0.185|\n\nI thought direct comparison of local cv was not possible, and LB was also somewhat difficult to use as a criterion for judgement.\n\nAfter considering the labeling method of the test set, I thought that the nms method was more similar to the labeling method, so I decided to use nms.\n\n[[Discussion] nms > weighted box fusion?](https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/214906#1199596)\n\n### torchvision's nms vs ZFTurbo's nms\n\nTorchvision'nms has some problems in the order of results when the bbox scores are the same.\n\n[[Stackoverflow] Pytorch argsort ordered, with duplicate elements in the tensor](https://stackoverflow.com/questions/56176439/pytorch-argsort-ordered-with-duplicate-elements-in-the-tensor)\n\n[torchvision.ops.nms](https://pytorch.org/vision/stable/ops.html#torchvision.ops.nms)\n\n> If multiple boxes have the exact same score and satisfy the IoU criterion with respect to a reference box, the selected box is not guaranteed to be the same between CPU and GPU. This is similar to the behavior of argsort in PyTorch when repeated values are present.\n\nIt doesn't matter when you convert the label only once at first and then save and load it as csv, pickle, etc., but if you convert the label and use it in a new environment, the consistency of the bbox may not be maintained.\n\nI decided to use ZFTurbo's nms which uses numpy.argsort, as this part would interfere with experiment flexibility and consistency during training.\n\n[[GitHub] ZFTurbo/Weighted boxes fusion](https://github.com/ZFTurbo/Weighted-Boxes-Fusion)\n\n## Model training\n\nAll models were trained on Colab Pro's V100 16GB single GPU.\n\n- AdamW\n- CosineAnnealingLR\n- epochs = 50\n- checkpoint selection : max mAP among top-3 min val_loss\n\n### Two-stage approach\n\n- 2-class classifier : EfficientNet, Resnet200d, total 15 models\n\n|model|image size(px)|folds|batch size|init lr|weight decay|val_acc|auc|\n|---|---|---|---|---|---|---|---|\n|b5|600|5 of 5|12|7.5e-5|1.0e-4|0.9601|0.9930|\n|b6|528|5 of 5|12|7.5e-5|1.0e-3|0.9553|0.9927|\n|resnet200d|600|3 of 5|12|7.5e-5|1.0e-4|0.9541|0.9934|\n|b5|456|single|16|1.0e-4|1.0e-4|0.9557|0.9927|\n|b5|1024|single|4|2.5e-5|1.0e-4|0.9577|0.9936|\n\n- 14-class detector : EfficientDet with 2-class classifier, total 18 models, local cv on positive image only\n\n|model|image size(px)|folds|batch size|init lr|weight decay|cv(mAP@iou=0.4)|public LB|private LB|\n|---|---|---|---|---|---|---|---|---|\n|d3|1024|single|3|3e-4|1e-3|0.4545|0.209|0.250|\n|d4|896|5 of 5|4|4e-4|1e-4|0.4541|0.218|0.250|\n|d4|896|single|4|4e-4|1e-3|0.4606|0.257|0.247|\n|d4|1024|single|3|3e-4|1e-3|0.4545|0.228|0.249|\n|d5|768|5 of 5|4|4e-4|1e-3|0.4472|0.225|0.253|\n|d5|896|4 of 5|3|3e-4|1e-3|0.4522|0.214|0.250|\n|d5|1024|single|2|2e-4|1e-3|0.4462|0.214|0.232|\n\n### One-stage approach\n\n- 15-class detector, total 2 models, local cv on positive image only\n\n|model|image size(px)|folds|batch size|init lr|weight decay|cv(mAP@iou=0.4)|public LB|private LB|\n|---|---|---|---|---|---|---|---|---|\n|d4|896|2 of 5|4|4e-4|1e-3|0.4546|0.230|0.246|\n\nAt batch size < 4, the mAP result was poor. The larger the image size, the better the mAP, but no further training was possible.\n\nI tried Freeze BatchNorm, accumulate grad batches, and GroupNorm but I didn't get any better results.\n\nI also tested Downconv, but the AP of small objects like Calcification increased, but the AP of ILD and large objects decreased, so the overall mAP was not improved.\n\n[[RANZCR CLiP] 11th Place Solution - Utilizing High resolution, Annotations, and Unlabeled data](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/226557)\n\n## Augmentation\n\nResize, scale, and crop were configured by referring to the method in the paper of EfficientDet.\n\nCLAHE, equalize, invertimg, huesaturationvalue, randomgamma, shiftscalerotate did not work.\n\n```python\nA.Compose(\n[\n    A.Resize(height=self.resize_height, width=self.resize_width),\n    A.RandomScale(scale_limit=(-0.9, 1.0), p=1.0),\n    A.PadIfNeeded(\n        min_height=self.resize_height,\n        min_width=self.resize_width,\n        border_mode=cv2.BORDER_CONSTANT,\n        value=0,\n        p=1.0,\n    ),\n    A.RandomCrop(height=self.resize_height, width=self.resize_width, p=1.0),\n    A.RandomBrightnessContrast(p=0.8),\n    A.ChannelDropout(p=0.5),\n    A.OneOf(\n        [\n            A.MotionBlur(p=0.5),\n            A.MedianBlur(p=0.5),\n            A.GaussianBlur(p=0.5),\n            A.GaussNoise(p=0.5),\n        ],\n        p=0.5,\n    ),\n    A.HorizontalFlip(p=0.5),\n    A.Normalize(),\n    ToTensorV2(),\n],\n```\n\n## Post processing\n\nIn the case of the two-stage approach, it was difficult to find an appropriate threshold for the classifier. The threshold was determined to be a 60~70% normal case. And if it was normal, all detections of the detector were excluded.\n\nIn the case of the one-stage approach, similarly, detections were excluded if it was normal based on the threshold.\n\n## Blending\n\nThe results of the two-stage approach and the one-stage approach were blended using nms, and the final result is as follows.\n\n### two-stage approach, 15 classifier + 18 detector\n\n|normal thr|public LB|private LB||\n|---|---|---|---|\n|0.70|0.219|**0.246**|**final submission**|\n|0.65|0.217|0.255|\n|0.60|0.215|0.256|\n\n### one-stage approach, 2 detector\n\n|normal thr|public LB|private LB|\n|---|---|---|\n|0.10|0.198|0.245|\n|0.00|0.198|0.245|\n\n### two-stage + one-stage\n\n|thr|public LB|private LB||\n|---|---|---|---|\n|0.70-0.30|0.224|**0.253**|**final submission**|\n|0.65-0.30|0.222|0.259|\n|0.60-0.30|0.220|0.258|\n\n## What did not work\n\nFreeze BatchNorm\n\naccumulate grad batches\n\nGroupNorm, group per channel = 8\n\nGroupNorm, pretrained backbone with GroupNorm\n\nDownconv\n\nNIH dataset concat classifier\n\nchange aspect ratio\n\n## Source Code\n\nSource code is available at [https://github.com/sunghyunjun/kaggle-vinbigdata-chest-xray-abnormalities-detection](https://github.com/sunghyunjun/kaggle-vinbigdata-chest-xray-abnormalities-detection)\n\nSubmission notebook is [91th place VBD inference CLF DET and DET ALL](https://www.kaggle.com/sunghyunjun/91th-place-vbd-inference-clf-det-and-det-all)",
    "1257968": "@sunghyunjun Congratulations on Bronze medal and thanks for sharing your approach",
    "1262462": "Congratulations on Bronze medal\nCan you share the EfficientDet CV and LB scoring please"
  }
}