{
  "id": 231511,
  "title": "1st Place Solution",
  "url": "/competitions/vinbigdata-chest-xray-abnormalities-detection/discussion/231511",
  "author_name": "Mohammed Rizin V K",
  "post_date": "2021-04-08T18:14:14.225000",
  "votes": 62,
  "comment_count": 7,
  "views": 0,
  "content": "<p><strong>Thanks to Kaggle and hosts for this very interesting competition with a too challenging dataset. This has been a great collaborative effort. Please give your upvotes to <a href=\"https://www.kaggle.com/fatihozturk\" target=\"_blank\">@fatihozturk</a> <a href=\"https://www.kaggle.com/socom20\" target=\"_blank\">@socom20</a> <a href=\"https://www.kaggle.com/avsanjay\" target=\"_blank\">@avsanjay</a> Congrats to The Winners.</strong></p>\n<h3>TLDR</h3>\n<p><a href=\"https://www.kaggle.com/fatihozturk\" target=\"_blank\">@fatihozturk</a> explained in his <a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229724\" target=\"_blank\">post</a> how was his path through this competition, this post is intended as an overall view of our final solution which achieved 0.354/0..314 PublicLB/PrivateLB. It is worth mentioning that we also had a better solution with 0.330/0.321 PublicLB/PrivateLB (which unfortunately we didn't select). As you all already know the main magic we have used for this competition was our Ensembling procedure, also there were lots of other important findings that we mention in this post.</p>\n<p><strong>CV strategy</strong>:  This was our main challenge, we didn't have a unique validation scheme. Since we started the competition as different teams, we all were using different CV splits to train our models and to build our separated best ensembles. To overcome this difficulty we finally ended up using the public LB as an extra validation dataset. We felt that this strategy was a bit risky, so we also retrained some of our best models using a common validation split and then we ensemble them (this last approach didn't have the best score on the private LB by itself). We think that our last ensemble, using the public LB as a validation source, helped out our final solution to fit better the consensus method used to build the test dataset. </p>\n<h3>Our Validation Strategy</h3>\n<p>Our strategy for the final submission can be divided into 3 stages :</p>\n<ul>\n<li>Fully Validated Stage</li>\n<li>Partially Validated Stage</li>\n<li>Not validated Stage (Just validated using the Public LB)</li>\n</ul>\n<h3>Models, we used for Ensembling</h3>\n<ul>\n<li><p><strong>Fully Validated Stage:</strong></p>\n<ul>\n<li>Detectron2 Resnet101 (notebook by <a href=\"https://www.kaggle.com/corochann\" target=\"_blank\">@corochann</a>, Thanks)</li>\n<li>YoloV5</li>\n<li>EffDetD2</li></ul></li>\n<li><p><strong>Partially Validated Stage:</strong></p>\n<ul>\n<li>YoloV5 5Folds</li></ul></li>\n<li><p><strong>Not Validated Stage:</strong></p>\n<ul>\n<li>2 EffdetD2 model  </li>\n<li>3 YoloV5 (1x w/ TTA , 2x w/o TTA)</li>\n<li>16 Classes Yolo model</li>\n<li>5 Folds YoloV5 by <a href=\"https://www.kaggle.com/awsaf49\" target=\"_blank\">@awsaf49</a></li>\n<li>New Anchor Yolo</li>\n<li>Detectron2 Resnet50 (notebook by <a href=\"https://www.kaggle.com/corochann\" target=\"_blank\">@corochann</a>, Thanks)</li>\n<li>YoloV5 <a href=\"https://www.kaggle.com/nxhong93\" target=\"_blank\">@nxhong93</a> (<a href=\"https://www.kaggle.com/nxhong93/yolov5-chest-512\" target=\"_blank\">https://www.kaggle.com/nxhong93/yolov5-chest-512</a>)</li>\n<li>Yolov5 with image size (640)</li></ul></li>\n</ul>\n<h3>Ensembling Strategy</h3>\n<p>To ensemble all the partial models we mainly used <a href=\"https://www.kaggle.com/zfturbo\" target=\"_blank\">@zfturbo</a>'s ensembling repository <a href=\"https://github.com/ZFTurbo/Weighted-Boxes-Fusion\" target=\"_blank\">https://github.com/ZFTurbo/Weighted-Boxes-Fusion</a>. Our ensemble technic is shown in the following figure:</p>\n<p><img src=\"https://i.postimg.cc/NjfVC2Mv/2x-Page-1-6.png\" alt=\"enter image description here\"></p>\n<p>As you can see at the end of the Not Validated Stage, we use WBF+p_sum. This last blending method is a variant of WBF proposed by <a href=\"https://www.kaggle.com/socom20\" target=\"_blank\">@socom20</a> and was intended to simulate the consensus of radiologists used in this particular test dataset. By using WBF+p_sum improved our Not Validated stage from 0.319 to 0.331 in the PublicLB.</p>\n<p><strong>Modified WBF:</strong>  It has more flexibility than WBF, it has 4 bbox blending technics, we found that only 2 of them were useful in this competition:</p>\n<ul>\n<li>\"p_det_weight_pmean\", has a behaviour similar to WBF, it finds the set of bboxes with iou &gt; thr, and then weights them using their detection probabilities (detection confidences). The new p_det (detection confidence) is an <strong>average</strong> over the detection probabilities of detections with iou &gt; thr.</li>\n<li>\"p_det_weight_psum\": does the same with the bboxes. In this case, the new p_det is a <strong>sum</strong> over the detection probability of bboxes with iou &gt; thr.<br>\nBecause of the sum of the p_dets, \"p_det_weight_psum\" needs a final normalization step. The idea behind psum is to try to replicate \"the consensus of radiologists\" using the partial models. When adding the probabilities of detection, we weigh more those boxes in which more models/radiologists agree.</li>\n</ul>\n<p>In our final submission, the not validated stage had a high weighting, which lead to an overfitted problem in our final ensemble we got  0.354/0.314 Public/Private. If we had reduced the weighting for this not validated stage, the LB score would have become 0.330/0.321 Public/Private as we previously have mentioned in the post.</p>\n<h3>Some Interesting Findings</h3>\n<ul>\n<li>After spending more time understanding the competition metric, we realized that there is No Penalty For Adding More Bbox as <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> explains clearly here: <a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229637\" target=\"_blank\">https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229637</a>. Having more confident boxes is good but, at the same time, having many low confidence bboxes improves the Recall and can improve the mAP as well. We tested this approach by using two models, a  Model A with a good mAP, and a Model B with a good Recall for all the classes. The following figure shows on the left the Precision vs Recall curves of Model A. It can be seen that many class tails end with a maximum recall = 0.7. We managed to add to Model A all the bboxes predicted for Model B which doesn't overlap with any bbox of predicted for Model A, the result is in the plot of the right. It can be seen that the main shape of the PvsR curves is still the same, but the tails now reach over recall=0.8. We didn't use this approach in our final model because the final ensemble already had a very good recall for all the classes.</li>\n</ul>\n<p><img src=\"https://i.postimg.cc/BnL1z77F/photo-2021-03-31-18-37-48.jpg\" alt=\"enter image description here\"></p>\n<ul>\n<li>Last competition days, we spent analyzing the predictions made for the ensemble both by submitting and also visually inspecting predicted bboxes. We found that our improved submissions were indeed improving for almost all classes including the easy and the rare ones.</li>\n</ul>\n<p><img src=\"https://i.ibb.co/k0NRCXJ/photo-2021-04-01-00-30-35.jpg\" alt=\"enter image description here\"></p>\n<h3>Things didn't work</h3>\n<ul>\n<li>Multi-Label Classifier: We trained a multilabel classifier using as output the 14 classes. We used sigmoid as activation because we needed the probability of each disease and many images have more than one disease. The idea was to remove from the ensemble all the bboxes not predicted by this classifier model, This approach led to a reduction in performance.  We also tried to switch the pooling layer and to use PCAM pooling (it is an attention pooling layer used in the <a href=\"https://github.com/jfhealthcare/Chexpert\" target=\"_blank\">Top1 Solution of CheXpert</a>), but it didn't work either. </li>\n<li>Bbox Filter: We trained an EfficientNetB6 which classifies whether the detected bboxes belong to a selected class or not just by seeing the crop selected by the ensemble. The model couldn't understand the difference between each disease, perhaps it needed more information on the whole image.</li>\n<li>Usage of External Dataset: We tried to use the NIH dataset. We first trained a backbone as a Multi-Label Classifier using NIH classes, and then we transfer its weights to the detector's backbone. Finally, we trained only the detection heads using this competition's dataset. We couldn't detect any appreciable improvement in the particular detector score.</li>\n<li>ClassWise Model: we trained different models, one for each detection class. We used clsX + cls14 as each model's dataset. In this case, each model had its own set of anchors prepared for detecting the class of interest. Some classes improved their CV performance but others didn't. Joining all the models we got a very good model, unfortunately, the validation split used in this model was different from the others and it was unreliable to add it to the Validated Stage so we decided to include it in the Partially Validated Stage, but still, didn't give any improvements. We had a final ensemble using this model, which wasn't finally selected, but it had the same score as our winning submission.</li>\n</ul>\n<h3>Models we choose.</h3>\n<ul>\n<li>Best at local validation (around 0.47+ mAP on CV), Public: 0.300, private: 0.287</li>\n<li>Best at LB. Public: 0.354, private: 0.314</li>\n<li>Our best submission in private was 0.321 which was 0.330 in Public</li>\n</ul>\n<h3>Hardware we used</h3>\n<ul>\n<li>4x Quadro GV100</li>\n<li>4x Geforce GTX 1080</li>\n<li>4x Titan X</li>\n<li>Ryzen 9 3950x</li>\n<li>Nvidia RTX 3080</li>\n<li>Quadro RTX 6000<br>\nEDIT 1:  The solution is reproducible by one GPU itself.</li>\n</ul>",
  "messages": [
    {
      "id": 1267727,
      "postDate": "2021-04-08T18:14:14.227Z",
      "content": "<p><strong>Thanks to Kaggle and hosts for this very interesting competition with a too challenging dataset. This has been a great collaborative effort. Please give your upvotes to <a href=\"https://www.kaggle.com/fatihozturk\" target=\"_blank\">@fatihozturk</a> <a href=\"https://www.kaggle.com/socom20\" target=\"_blank\">@socom20</a> <a href=\"https://www.kaggle.com/avsanjay\" target=\"_blank\">@avsanjay</a> Congrats to The Winners.</strong></p>\n<h3>TLDR</h3>\n<p><a href=\"https://www.kaggle.com/fatihozturk\" target=\"_blank\">@fatihozturk</a> explained in his <a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229724\" target=\"_blank\">post</a> how was his path through this competition, this post is intended as an overall view of our final solution which achieved 0.354/0..314 PublicLB/PrivateLB. It is worth mentioning that we also had a better solution with 0.330/0.321 PublicLB/PrivateLB (which unfortunately we didn't select). As you all already know the main magic we have used for this competition was our Ensembling procedure, also there were lots of other important findings that we mention in this post.</p>\n<p><strong>CV strategy</strong>:  This was our main challenge, we didn't have a unique validation scheme. Since we started the competition as different teams, we all were using different CV splits to train our models and to build our separated best ensembles. To overcome this difficulty we finally ended up using the public LB as an extra validation dataset. We felt that this strategy was a bit risky, so we also retrained some of our best models using a common validation split and then we ensemble them (this last approach didn't have the best score on the private LB by itself). We think that our last ensemble, using the public LB as a validation source, helped out our final solution to fit better the consensus method used to build the test dataset. </p>\n<h3>Our Validation Strategy</h3>\n<p>Our strategy for the final submission can be divided into 3 stages :</p>\n<ul>\n<li>Fully Validated Stage</li>\n<li>Partially Validated Stage</li>\n<li>Not validated Stage (Just validated using the Public LB)</li>\n</ul>\n<h3>Models, we used for Ensembling</h3>\n<ul>\n<li><p><strong>Fully Validated Stage:</strong></p>\n<ul>\n<li>Detectron2 Resnet101 (notebook by <a href=\"https://www.kaggle.com/corochann\" target=\"_blank\">@corochann</a>, Thanks)</li>\n<li>YoloV5</li>\n<li>EffDetD2</li></ul></li>\n<li><p><strong>Partially Validated Stage:</strong></p>\n<ul>\n<li>YoloV5 5Folds</li></ul></li>\n<li><p><strong>Not Validated Stage:</strong></p>\n<ul>\n<li>2 EffdetD2 model  </li>\n<li>3 YoloV5 (1x w/ TTA , 2x w/o TTA)</li>\n<li>16 Classes Yolo model</li>\n<li>5 Folds YoloV5 by <a href=\"https://www.kaggle.com/awsaf49\" target=\"_blank\">@awsaf49</a></li>\n<li>New Anchor Yolo</li>\n<li>Detectron2 Resnet50 (notebook by <a href=\"https://www.kaggle.com/corochann\" target=\"_blank\">@corochann</a>, Thanks)</li>\n<li>YoloV5 <a href=\"https://www.kaggle.com/nxhong93\" target=\"_blank\">@nxhong93</a> (<a href=\"https://www.kaggle.com/nxhong93/yolov5-chest-512\" target=\"_blank\">https://www.kaggle.com/nxhong93/yolov5-chest-512</a>)</li>\n<li>Yolov5 with image size (640)</li></ul></li>\n</ul>\n<h3>Ensembling Strategy</h3>\n<p>To ensemble all the partial models we mainly used <a href=\"https://www.kaggle.com/zfturbo\" target=\"_blank\">@zfturbo</a>'s ensembling repository <a href=\"https://github.com/ZFTurbo/Weighted-Boxes-Fusion\" target=\"_blank\">https://github.com/ZFTurbo/Weighted-Boxes-Fusion</a>. Our ensemble technic is shown in the following figure:</p>\n<p><img src=\"https://i.postimg.cc/NjfVC2Mv/2x-Page-1-6.png\" alt=\"enter image description here\"></p>\n<p>As you can see at the end of the Not Validated Stage, we use WBF+p_sum. This last blending method is a variant of WBF proposed by <a href=\"https://www.kaggle.com/socom20\" target=\"_blank\">@socom20</a> and was intended to simulate the consensus of radiologists used in this particular test dataset. By using WBF+p_sum improved our Not Validated stage from 0.319 to 0.331 in the PublicLB.</p>\n<p><strong>Modified WBF:</strong>  It has more flexibility than WBF, it has 4 bbox blending technics, we found that only 2 of them were useful in this competition:</p>\n<ul>\n<li>\"p_det_weight_pmean\", has a behaviour similar to WBF, it finds the set of bboxes with iou &gt; thr, and then weights them using their detection probabilities (detection confidences). The new p_det (detection confidence) is an <strong>average</strong> over the detection probabilities of detections with iou &gt; thr.</li>\n<li>\"p_det_weight_psum\": does the same with the bboxes. In this case, the new p_det is a <strong>sum</strong> over the detection probability of bboxes with iou &gt; thr.<br>\nBecause of the sum of the p_dets, \"p_det_weight_psum\" needs a final normalization step. The idea behind psum is to try to replicate \"the consensus of radiologists\" using the partial models. When adding the probabilities of detection, we weigh more those boxes in which more models/radiologists agree.</li>\n</ul>\n<p>In our final submission, the not validated stage had a high weighting, which lead to an overfitted problem in our final ensemble we got  0.354/0.314 Public/Private. If we had reduced the weighting for this not validated stage, the LB score would have become 0.330/0.321 Public/Private as we previously have mentioned in the post.</p>\n<h3>Some Interesting Findings</h3>\n<ul>\n<li>After spending more time understanding the competition metric, we realized that there is No Penalty For Adding More Bbox as <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> explains clearly here: <a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229637\" target=\"_blank\">https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229637</a>. Having more confident boxes is good but, at the same time, having many low confidence bboxes improves the Recall and can improve the mAP as well. We tested this approach by using two models, a  Model A with a good mAP, and a Model B with a good Recall for all the classes. The following figure shows on the left the Precision vs Recall curves of Model A. It can be seen that many class tails end with a maximum recall = 0.7. We managed to add to Model A all the bboxes predicted for Model B which doesn't overlap with any bbox of predicted for Model A, the result is in the plot of the right. It can be seen that the main shape of the PvsR curves is still the same, but the tails now reach over recall=0.8. We didn't use this approach in our final model because the final ensemble already had a very good recall for all the classes.</li>\n</ul>\n<p><img src=\"https://i.postimg.cc/BnL1z77F/photo-2021-03-31-18-37-48.jpg\" alt=\"enter image description here\"></p>\n<ul>\n<li>Last competition days, we spent analyzing the predictions made for the ensemble both by submitting and also visually inspecting predicted bboxes. We found that our improved submissions were indeed improving for almost all classes including the easy and the rare ones.</li>\n</ul>\n<p><img src=\"https://i.ibb.co/k0NRCXJ/photo-2021-04-01-00-30-35.jpg\" alt=\"enter image description here\"></p>\n<h3>Things didn't work</h3>\n<ul>\n<li>Multi-Label Classifier: We trained a multilabel classifier using as output the 14 classes. We used sigmoid as activation because we needed the probability of each disease and many images have more than one disease. The idea was to remove from the ensemble all the bboxes not predicted by this classifier model, This approach led to a reduction in performance.  We also tried to switch the pooling layer and to use PCAM pooling (it is an attention pooling layer used in the <a href=\"https://github.com/jfhealthcare/Chexpert\" target=\"_blank\">Top1 Solution of CheXpert</a>), but it didn't work either. </li>\n<li>Bbox Filter: We trained an EfficientNetB6 which classifies whether the detected bboxes belong to a selected class or not just by seeing the crop selected by the ensemble. The model couldn't understand the difference between each disease, perhaps it needed more information on the whole image.</li>\n<li>Usage of External Dataset: We tried to use the NIH dataset. We first trained a backbone as a Multi-Label Classifier using NIH classes, and then we transfer its weights to the detector's backbone. Finally, we trained only the detection heads using this competition's dataset. We couldn't detect any appreciable improvement in the particular detector score.</li>\n<li>ClassWise Model: we trained different models, one for each detection class. We used clsX + cls14 as each model's dataset. In this case, each model had its own set of anchors prepared for detecting the class of interest. Some classes improved their CV performance but others didn't. Joining all the models we got a very good model, unfortunately, the validation split used in this model was different from the others and it was unreliable to add it to the Validated Stage so we decided to include it in the Partially Validated Stage, but still, didn't give any improvements. We had a final ensemble using this model, which wasn't finally selected, but it had the same score as our winning submission.</li>\n</ul>\n<h3>Models we choose.</h3>\n<ul>\n<li>Best at local validation (around 0.47+ mAP on CV), Public: 0.300, private: 0.287</li>\n<li>Best at LB. Public: 0.354, private: 0.314</li>\n<li>Our best submission in private was 0.321 which was 0.330 in Public</li>\n</ul>\n<h3>Hardware we used</h3>\n<ul>\n<li>4x Quadro GV100</li>\n<li>4x Geforce GTX 1080</li>\n<li>4x Titan X</li>\n<li>Ryzen 9 3950x</li>\n<li>Nvidia RTX 3080</li>\n<li>Quadro RTX 6000<br>\nEDIT 1:  The solution is reproducible by one GPU itself.</li>\n</ul>",
      "rawMarkdown": "**Thanks to Kaggle and hosts for this very interesting competition with a too challenging dataset. This has been a great collaborative effort. Please give your upvotes to @fatihozturk @socom20 @avsanjay Congrats to The Winners.**\n\n### TLDR\n@fatihozturk explained in his [post](https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229724) how was his path through this competition, this post is intended as an overall view of our final solution which achieved 0.354/0..314 PublicLB/PrivateLB. It is worth mentioning that we also had a better solution with 0.330/0.321 PublicLB/PrivateLB (which unfortunately we didn't select). As you all already know the main magic we have used for this competition was our Ensembling procedure, also there were lots of other important findings that we mention in this post.\n\n**CV strategy**:  This was our main challenge, we didn't have a unique validation scheme. Since we started the competition as different teams, we all were using different CV splits to train our models and to build our separated best ensembles. To overcome this difficulty we finally ended up using the public LB as an extra validation dataset. We felt that this strategy was a bit risky, so we also retrained some of our best models using a common validation split and then we ensemble them (this last approach didn't have the best score on the private LB by itself). We think that our last ensemble, using the public LB as a validation source, helped out our final solution to fit better the consensus method used to build the test dataset. \n\n### Our Validation Strategy\nOur strategy for the final submission can be divided into 3 stages :\n* Fully Validated Stage\n* Partially Validated Stage\n* Not validated Stage (Just validated using the Public LB)\n\n### Models, we used for Ensembling\n\n- **Fully Validated Stage:**\n  - Detectron2 Resnet101 (notebook by @corochann, Thanks)\n  - YoloV5\n  - EffDetD2\n\n- **Partially Validated Stage:**\n  - YoloV5 5Folds\n\n- **Not Validated Stage:**\n  - 2 EffdetD2 model  \n  - 3 YoloV5 (1x w/ TTA , 2x w/o TTA)\n  - 16 Classes Yolo model\n  - 5 Folds YoloV5 by @awsaf49\n  - New Anchor Yolo\n  - Detectron2 Resnet50 (notebook by @corochann, Thanks)\n  - YoloV5 @nxhong93 (https://www.kaggle.com/nxhong93/yolov5-chest-512)\n  - Yolov5 with image size (640)\n\n### Ensembling Strategy\n\nTo ensemble all the partial models we mainly used @zfturbo's ensembling repository https://github.com/ZFTurbo/Weighted-Boxes-Fusion. Our ensemble technic is shown in the following figure:\n\n![enter image description here](https://i.postimg.cc/NjfVC2Mv/2x-Page-1-6.png)\n\nAs you can see at the end of the Not Validated Stage, we use WBF+p_sum. This last blending method is a variant of WBF proposed by @socom20 and was intended to simulate the consensus of radiologists used in this particular test dataset. By using WBF+p_sum improved our Not Validated stage from 0.319 to 0.331 in the PublicLB.\n\n**Modified WBF:**  It has more flexibility than WBF, it has 4 bbox blending technics, we found that only 2 of them were useful in this competition:\n- \"p_det_weight_pmean\", has a behaviour similar to WBF, it finds the set of bboxes with iou > thr, and then weights them using their detection probabilities (detection confidences). The new p_det (detection confidence) is an **average** over the detection probabilities of detections with iou > thr.\n- \"p_det_weight_psum\": does the same with the bboxes. In this case, the new p_det is a **sum** over the detection probability of bboxes with iou > thr.\nBecause of the sum of the p_dets, \"p_det_weight_psum\" needs a final normalization step. The idea behind psum is to try to replicate \"the consensus of radiologists\" using the partial models. When adding the probabilities of detection, we weigh more those boxes in which more models/radiologists agree.\n\nIn our final submission, the not validated stage had a high weighting, which lead to an overfitted problem in our final ensemble we got  0.354/0.314 Public/Private. If we had reduced the weighting for this not validated stage, the LB score would have become 0.330/0.321 Public/Private as we previously have mentioned in the post.\n\n### Some Interesting Findings\n\n- After spending more time understanding the competition metric, we realized that there is No Penalty For Adding More Bbox as @cdeotte explains clearly here: https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229637. Having more confident boxes is good but, at the same time, having many low confidence bboxes improves the Recall and can improve the mAP as well. We tested this approach by using two models, a  Model A with a good mAP, and a Model B with a good Recall for all the classes. The following figure shows on the left the Precision vs Recall curves of Model A. It can be seen that many class tails end with a maximum recall = 0.7. We managed to add to Model A all the bboxes predicted for Model B which doesn't overlap with any bbox of predicted for Model A, the result is in the plot of the right. It can be seen that the main shape of the PvsR curves is still the same, but the tails now reach over recall=0.8. We didn't use this approach in our final model because the final ensemble already had a very good recall for all the classes.\n\n![enter image description here](https://i.postimg.cc/BnL1z77F/photo-2021-03-31-18-37-48.jpg)\n\n- Last competition days, we spent analyzing the predictions made for the ensemble both by submitting and also visually inspecting predicted bboxes. We found that our improved submissions were indeed improving for almost all classes including the easy and the rare ones.\n\n![enter image description here](https://i.ibb.co/k0NRCXJ/photo-2021-04-01-00-30-35.jpg)\n\n### Things didn't work \n- Multi-Label Classifier: We trained a multilabel classifier using as output the 14 classes. We used sigmoid as activation because we needed the probability of each disease and many images have more than one disease. The idea was to remove from the ensemble all the bboxes not predicted by this classifier model, This approach led to a reduction in performance.  We also tried to switch the pooling layer and to use PCAM pooling (it is an attention pooling layer used in the [Top1 Solution of CheXpert](https://github.com/jfhealthcare/Chexpert)), but it didn't work either. \n- Bbox Filter: We trained an EfficientNetB6 which classifies whether the detected bboxes belong to a selected class or not just by seeing the crop selected by the ensemble. The model couldn't understand the difference between each disease, perhaps it needed more information on the whole image.\n- Usage of External Dataset: We tried to use the NIH dataset. We first trained a backbone as a Multi-Label Classifier using NIH classes, and then we transfer its weights to the detector's backbone. Finally, we trained only the detection heads using this competition's dataset. We couldn't detect any appreciable improvement in the particular detector score.\n- ClassWise Model: we trained different models, one for each detection class. We used clsX + cls14 as each model's dataset. In this case, each model had its own set of anchors prepared for detecting the class of interest. Some classes improved their CV performance but others didn't. Joining all the models we got a very good model, unfortunately, the validation split used in this model was different from the others and it was unreliable to add it to the Validated Stage so we decided to include it in the Partially Validated Stage, but still, didn't give any improvements. We had a final ensemble using this model, which wasn't finally selected, but it had the same score as our winning submission.\n\n### Models we choose.\n- Best at local validation (around 0.47+ mAP on CV), Public: 0.300, private: 0.287\n- Best at LB. Public: 0.354, private: 0.314\n- Our best submission in private was 0.321 which was 0.330 in Public\n\n### Hardware we used\n- 4x Quadro GV100\n- 4x Geforce GTX 1080\n- 4x Titan X\n- Ryzen 9 3950x\n- Nvidia RTX 3080\n- Quadro RTX 6000\nEDIT 1:  The solution is reproducible by one GPU itself.",
      "votes": 62
    },
    {
      "id": 1269892,
      "postDate": "2021-04-11T03:19:04.300Z",
      "content": "<p>Thank you for sharing! And you deserve it :) Congrats!</p>",
      "rawMarkdown": "Thank you for sharing! And you deserve it :) Congrats!",
      "votes": 1
    },
    {
      "id": 1267740,
      "postDate": "2021-04-08T18:27:53.217Z",
      "content": "<p>Sorry for being so late. We were busy in preparing the codes.</p>",
      "rawMarkdown": "Sorry for being so late. We were busy in preparing the codes.",
      "votes": 1
    },
    {
      "id": 1324896,
      "postDate": "2021-05-27T10:22:01.450Z",
      "content": "<p>hi,<br>\nis it possible so you could share predictions of your final model made on 15k samples trainset? with my colleagues we would like to analyze your model deeper to learn more from the winner:)<br>\n e.g. how model performed on different classes, and so on.</p>",
      "rawMarkdown": "hi,\nis it possible so you could share predictions of your final model made on 15k samples trainset? with my colleagues we would like to analyze your model deeper to learn more from the winner:)\n e.g. how model performed on different classes, and so on.",
      "replies": [
        {
          "id": 1325426,
          "postDate": "2021-05-27T18:29:16.990Z",
          "content": "<p>Do you need the final prediction file or what else ? i am confused</p>",
          "rawMarkdown": "Do you need the final prediction file or what else ? i am confused"
        },
        {
          "id": 1349584,
          "postDate": "2021-06-14T23:29:36.797Z",
          "content": "<p>if possible, we would need file containing predictions performed on trainset (data which model saw during training), not final predicitons. </p>",
          "rawMarkdown": "if possible, we would need file containing predictions performed on trainset (data which model saw during training), not final predicitons. "
        }
      ]
    },
    {
      "id": 1272976,
      "postDate": "2021-04-14T00:53:15.830Z",
      "content": "<p>the pic in <strong>Ensembling Strategy</strong> is so small, can you please give me another more clear one</p>",
      "rawMarkdown": "the pic in **Ensembling Strategy** is so small, can you please give me another more clear one",
      "replies": [
        {
          "id": 1274399,
          "postDate": "2021-04-15T09:14:43.783Z",
          "content": "<p><img src=\"https://i.postimg.cc/zvD77bts/2x-Page-1-7.png\" alt=\"\"></p>",
          "rawMarkdown": "![](https://i.postimg.cc/zvD77bts/2x-Page-1-7.png)",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1269892,
      "author_name": "Chanran Kim",
      "author_url": "",
      "post_date": "2021-04-11T03:19:04.300000",
      "content": "<p>Thank you for sharing! And you deserve it :) Congrats!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1267740,
      "author_name": "Mohammed Rizin V K",
      "author_url": "",
      "post_date": "2021-04-08T18:27:53.217000",
      "content": "<p>Sorry for being so late. We were busy in preparing the codes.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1324896,
      "author_name": "Piotr Czarnecki",
      "author_url": "",
      "post_date": "2021-05-27T10:22:01.450000",
      "content": "<p>hi,<br>\nis it possible so you could share predictions of your final model made on 15k samples trainset? with my colleagues we would like to analyze your model deeper to learn more from the winner:)<br>\n e.g. how model performed on different classes, and so on.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1325426,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2021-05-27T18:29:16.990000",
          "content": "<p>Do you need the final prediction file or what else ? i am confused</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1349584,
          "author_name": "Piotr Czarnecki",
          "author_url": "",
          "post_date": "2021-06-14T23:29:36.797000",
          "content": "<p>if possible, we would need file containing predictions performed on trainset (data which model saw during training), not final predicitons. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1272976,
      "author_name": "xuzhaocheng",
      "author_url": "",
      "post_date": "2021-04-14T00:53:15.830000",
      "content": "<p>the pic in <strong>Ensembling Strategy</strong> is so small, can you please give me another more clear one</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1274399,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2021-04-15T09:14:43.783000",
          "content": "<p><img src=\"https://i.postimg.cc/zvD77bts/2x-Page-1-7.png\" alt=\"\"></p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1267727": "**Thanks to Kaggle and hosts for this very interesting competition with a too challenging dataset. This has been a great collaborative effort. Please give your upvotes to @fatihozturk @socom20 @avsanjay Congrats to The Winners.**\n\n### TLDR\n@fatihozturk explained in his [post](https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229724) how was his path through this competition, this post is intended as an overall view of our final solution which achieved 0.354/0..314 PublicLB/PrivateLB. It is worth mentioning that we also had a better solution with 0.330/0.321 PublicLB/PrivateLB (which unfortunately we didn't select). As you all already know the main magic we have used for this competition was our Ensembling procedure, also there were lots of other important findings that we mention in this post.\n\n**CV strategy**:  This was our main challenge, we didn't have a unique validation scheme. Since we started the competition as different teams, we all were using different CV splits to train our models and to build our separated best ensembles. To overcome this difficulty we finally ended up using the public LB as an extra validation dataset. We felt that this strategy was a bit risky, so we also retrained some of our best models using a common validation split and then we ensemble them (this last approach didn't have the best score on the private LB by itself). We think that our last ensemble, using the public LB as a validation source, helped out our final solution to fit better the consensus method used to build the test dataset. \n\n### Our Validation Strategy\nOur strategy for the final submission can be divided into 3 stages :\n* Fully Validated Stage\n* Partially Validated Stage\n* Not validated Stage (Just validated using the Public LB)\n\n### Models, we used for Ensembling\n\n- **Fully Validated Stage:**\n  - Detectron2 Resnet101 (notebook by @corochann, Thanks)\n  - YoloV5\n  - EffDetD2\n\n- **Partially Validated Stage:**\n  - YoloV5 5Folds\n\n- **Not Validated Stage:**\n  - 2 EffdetD2 model  \n  - 3 YoloV5 (1x w/ TTA , 2x w/o TTA)\n  - 16 Classes Yolo model\n  - 5 Folds YoloV5 by @awsaf49\n  - New Anchor Yolo\n  - Detectron2 Resnet50 (notebook by @corochann, Thanks)\n  - YoloV5 @nxhong93 (https://www.kaggle.com/nxhong93/yolov5-chest-512)\n  - Yolov5 with image size (640)\n\n### Ensembling Strategy\n\nTo ensemble all the partial models we mainly used @zfturbo's ensembling repository https://github.com/ZFTurbo/Weighted-Boxes-Fusion. Our ensemble technic is shown in the following figure:\n\n![enter image description here](https://i.postimg.cc/NjfVC2Mv/2x-Page-1-6.png)\n\nAs you can see at the end of the Not Validated Stage, we use WBF+p_sum. This last blending method is a variant of WBF proposed by @socom20 and was intended to simulate the consensus of radiologists used in this particular test dataset. By using WBF+p_sum improved our Not Validated stage from 0.319 to 0.331 in the PublicLB.\n\n**Modified WBF:**  It has more flexibility than WBF, it has 4 bbox blending technics, we found that only 2 of them were useful in this competition:\n- \"p_det_weight_pmean\", has a behaviour similar to WBF, it finds the set of bboxes with iou > thr, and then weights them using their detection probabilities (detection confidences). The new p_det (detection confidence) is an **average** over the detection probabilities of detections with iou > thr.\n- \"p_det_weight_psum\": does the same with the bboxes. In this case, the new p_det is a **sum** over the detection probability of bboxes with iou > thr.\nBecause of the sum of the p_dets, \"p_det_weight_psum\" needs a final normalization step. The idea behind psum is to try to replicate \"the consensus of radiologists\" using the partial models. When adding the probabilities of detection, we weigh more those boxes in which more models/radiologists agree.\n\nIn our final submission, the not validated stage had a high weighting, which lead to an overfitted problem in our final ensemble we got  0.354/0.314 Public/Private. If we had reduced the weighting for this not validated stage, the LB score would have become 0.330/0.321 Public/Private as we previously have mentioned in the post.\n\n### Some Interesting Findings\n\n- After spending more time understanding the competition metric, we realized that there is No Penalty For Adding More Bbox as @cdeotte explains clearly here: https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229637. Having more confident boxes is good but, at the same time, having many low confidence bboxes improves the Recall and can improve the mAP as well. We tested this approach by using two models, a  Model A with a good mAP, and a Model B with a good Recall for all the classes. The following figure shows on the left the Precision vs Recall curves of Model A. It can be seen that many class tails end with a maximum recall = 0.7. We managed to add to Model A all the bboxes predicted for Model B which doesn't overlap with any bbox of predicted for Model A, the result is in the plot of the right. It can be seen that the main shape of the PvsR curves is still the same, but the tails now reach over recall=0.8. We didn't use this approach in our final model because the final ensemble already had a very good recall for all the classes.\n\n![enter image description here](https://i.postimg.cc/BnL1z77F/photo-2021-03-31-18-37-48.jpg)\n\n- Last competition days, we spent analyzing the predictions made for the ensemble both by submitting and also visually inspecting predicted bboxes. We found that our improved submissions were indeed improving for almost all classes including the easy and the rare ones.\n\n![enter image description here](https://i.ibb.co/k0NRCXJ/photo-2021-04-01-00-30-35.jpg)\n\n### Things didn't work \n- Multi-Label Classifier: We trained a multilabel classifier using as output the 14 classes. We used sigmoid as activation because we needed the probability of each disease and many images have more than one disease. The idea was to remove from the ensemble all the bboxes not predicted by this classifier model, This approach led to a reduction in performance.  We also tried to switch the pooling layer and to use PCAM pooling (it is an attention pooling layer used in the [Top1 Solution of CheXpert](https://github.com/jfhealthcare/Chexpert)), but it didn't work either. \n- Bbox Filter: We trained an EfficientNetB6 which classifies whether the detected bboxes belong to a selected class or not just by seeing the crop selected by the ensemble. The model couldn't understand the difference between each disease, perhaps it needed more information on the whole image.\n- Usage of External Dataset: We tried to use the NIH dataset. We first trained a backbone as a Multi-Label Classifier using NIH classes, and then we transfer its weights to the detector's backbone. Finally, we trained only the detection heads using this competition's dataset. We couldn't detect any appreciable improvement in the particular detector score.\n- ClassWise Model: we trained different models, one for each detection class. We used clsX + cls14 as each model's dataset. In this case, each model had its own set of anchors prepared for detecting the class of interest. Some classes improved their CV performance but others didn't. Joining all the models we got a very good model, unfortunately, the validation split used in this model was different from the others and it was unreliable to add it to the Validated Stage so we decided to include it in the Partially Validated Stage, but still, didn't give any improvements. We had a final ensemble using this model, which wasn't finally selected, but it had the same score as our winning submission.\n\n### Models we choose.\n- Best at local validation (around 0.47+ mAP on CV), Public: 0.300, private: 0.287\n- Best at LB. Public: 0.354, private: 0.314\n- Our best submission in private was 0.321 which was 0.330 in Public\n\n### Hardware we used\n- 4x Quadro GV100\n- 4x Geforce GTX 1080\n- 4x Titan X\n- Ryzen 9 3950x\n- Nvidia RTX 3080\n- Quadro RTX 6000\nEDIT 1:  The solution is reproducible by one GPU itself.",
    "1269892": "Thank you for sharing! And you deserve it :) Congrats!",
    "1267740": "Sorry for being so late. We were busy in preparing the codes.",
    "1324896": "hi,\nis it possible so you could share predictions of your final model made on 15k samples trainset? with my colleagues we would like to analyze your model deeper to learn more from the winner:)\n e.g. how model performed on different classes, and so on.",
    "1272976": "the pic in **Ensembling Strategy** is so small, can you please give me another more clear one"
  }
}