{
  "id": 229696,
  "title": "2nd place solution (quick overview)",
  "url": "/competitions/vinbigdata-chest-xray-abnormalities-detection/discussion/229696",
  "author_name": "ZFTurbo",
  "post_date": "2021-03-31T09:23:29.521000",
  "votes": 93,
  "comment_count": 31,
  "views": 0,
  "content": "<h2>Part 1: Initial models</h2>\n<p>At first I started to try different OD models for this dataset. I tried: EffDet (B5), CenterNet (HourGlass), Yolo_v5, Retinanet (RsNet50, ResNet101, ResNet152), Faster RCNN</p>\n<p>Main initial tricks:</p>\n<ol>\n<li>To train models I used <a href=\"https://github.com/ZFTurbo/Weighted-Boxes-Fusion\" target=\"_blank\">WBF</a> on train.csv with IoU: 0.4. I usually used this file with most of models for training and for validation.</li>\n<li>Using validation I find out that it’s critical to use lower values for IoU in NMS which filters boxes on output of models. So I used NMS IoU: 0.25 for Yolo_v5 and NMS IoU: 0.3 for RetinaNet. Looks like, it was better for high confident boxes to consume more boxes around it.</li>\n<li>I didn’t want to change anchors for RetinaNet so I created version of train.csv with boxes which is close to square, because default Retina doesn’t like elongated boxes.</li>\n</ol>\n<p>Best models were the following:</p>\n<ul>\n<li>Yolo_v5: ~0.251 Public LB (Std + Mirror images 5KFold, 640px)</li>\n<li>RetinaNet (ResNet101): ~0.246 Public LB (Std + Mirror images 5KFold, 1024px)</li>\n<li>RetinaNet (ResNet152): ~0.222 Public LB (Std + Mirror images 5KFold, 800px)</li>\n<li>CenterNet (HourGlass): ~0.196 Public LB(Std + Mirror images 5KFold, 512px)</li>\n</ul>\n<p>EffDet – same code as Wheat competition, but works pretty poor on this dataset. With Faster-RCNN I think I have some problem with code, because it has the very bad score.</p>\n<p>Using WBF ensemble of these models + 2 public models (Detectron2 and Yolo_v5):</p>\n<ul>\n<li>0.221: <a href=\"https://www.kaggle.com/corochann/vinbigdata-detectron2-prediction\" target=\"_blank\">https://www.kaggle.com/corochann/vinbigdata-detectron2-prediction</a></li>\n<li>0.204: <a href=\"https://www.kaggle.com/awsaf49/vinbigdata-2-class-filter?scriptVersionId=51562634\" target=\"_blank\">https://www.kaggle.com/awsaf49/vinbigdata-2-class-filter?scriptVersionId=51562634</a></li>\n</ul>\n<p>I was able to achieve 0.300 on Public LB (0.289 Private LB). Then I merge with my teammates. They had more mmdetection models as well as Yolo_v5, but together we were able to got only 0.301.</p>\n<h2>Part 2: Investigation</h2>\n<p>While we were stuck we decided to get individual score for each class and compare with our OOF validation. For 3 days we got the following table:</p>\n<pre><code>Class 0: 0.014 (mAP LB: 0.210 Valid: 0.900)\nClass 1: 0.009 (mAP LB: 0.135 Valid: 0.335)\nClass 2: 0.017 (mAP LB: 0.255 Valid: 0.278)\nClass 3: 0.046 (mAP LB: 0.690 Valid: 0.915)\nClass 4: 0.019 (mAP LB: 0.285 Valid: 0.438)\nClass 5: 0.013 (mAP LB: 0.195 Valid: 0.376)\nClass 6: 0.010 (mAP LB: 0.150 Valid: 0.465)\nClass 7: 0.003 (mAP LB: 0.045 Valid: 0.400)\nClass 8: 0.011 (mAP LB: 0.165 Valid: 0.372)\nClass 9: 0.004 (mAP LB: 0.060 Valid: 0.162)\nClass 10: 0.027 (mAP LB: 0.405 Valid: 0.548)\nClass 11: 0.012 (mAP LB: 0.180 Valid: 0.343)\nClass 12: 0.037 (mAP LB: 0.555 Valid: 0.388)\nClass 13: 0.009 (mAP LB: 0.135 Valid: 0.432)\nClass 14: 0.063 (mAP LB: 0.945 Valid: 0.772)\nSum: 0.294\n</code></pre>\n<p>As it can be seen LB mAP the very different from local mAP. We started from investigation of Class 0 difference (it was most suspicious). We have near perfect prediction on validation while LB score is so small. The first thing we tried is to calculate validation mAP based on Radiologist ID.</p>\n<pre><code>Class ID: 0 Validation file: ensemble_iou_0.4._mAP_0.454_extracted_classes_0.csv\nRadiologist: R8 mAP: 0.991058\nRadiologist: R9 mAP: 0.954593\nRadiologist: R10 mAP: 0.971534\nRadiologist: R11 mAP: 0.379283\nRadiologist: R12 mAP: 0.019819\nRadiologist: R13 mAP: 0.511284\nRadiologist: R14 mAP: 0.527909\nRadiologist: R15 mAP: 0.279870\nRadiologist: R16 mAP: 0.506167\nRadiologist: R17 mAP: 0.447092\n</code></pre>\n<p>Interesting, we have perfect match for R8, R9 and R10, but very poor on other radiologists. Bad thing is that there are too small entries by other radiologists in train:</p>\n<pre><code>R11 - Class:  0 Entries:    30 AP: 0.379283\nR12 - Class:  0 Entries:    12 AP: 0.019819\nR13 - Class:  0 Entries:    32 AP: 0.511284\nR14 - Class:  0 Entries:    74 AP: 0.527909\nR15 - Class:  0 Entries:    25 AP: 0.279870\nR16 - Class:  0 Entries:    23 AP: 0.506167\nR17 - Class:  0 Entries:    10 AP: 0.447092\n</code></pre>\n<p>One more thing that R8, R9 and R10 never markup with any other radiologists. And what if test set marked up by some other radiologists (may be even by R13, R14 etc) and making boxes too similar to these 3 we overfit to train? And what if with better models we went further from test mark up?</p>\n<p>I checked by eyes images of class 0 in train marked up by R8, R9 and R10 and by rare radiologists to check the difference. I was lucky to see that other radiologists often use much larger boxes comparing to R8, R9 and R10. So I made a small trick by hand:<br>\n<code>x1 -= 100 y2 += 100 for all boxes with class 0</code><br>\nAnd checked local validation, which improved a lot for other radiologists, while R8, R9 and R10 became worse:</p>\n<pre><code>Radiologist: R10 Class:  0 Entries:  2349 AP: 0.559458\nRadiologist: R11 Class:  0 Entries:    30 AP: 0.559355\nRadiologist: R12 Class:  0 Entries:    12 AP: 0.351739\nRadiologist: R13 Class:  0 Entries:    32 AP: 0.483730\nRadiologist: R14 Class:  0 Entries:    74 AP: 0.623019\nRadiologist: R15 Class:  0 Entries:    25 AP: 0.341545\nRadiologist: R16 Class:  0 Entries:    23 AP: 0.506167\nRadiologist: R17 Class:  0 Entries:    10 AP: 0.452887\n</code></pre>\n<p>I made the same trick on our 0.301 submission and got 0.309 on LB. This showed us the right direction. We finetuned our models using only rare radiologists. And it improved LB score for all of models:</p>\n<p>RetinaNet (ResNet101): 0.246 -&gt; 0.264<br>\nRetinaNet (ResNet152): 0.222 -&gt; 0.237<br>\nYolo_v5: 0.251 -&gt; 0.271<br>\nEnsemble for all models including original moved us to Public LB &gt; 0.330</p>\n<h2>Part 3 (Single classes):</h2>\n<p>At this point it became hard to add new models to ensemble without proper validation. So we decided to switch to single class improvements. It was much easier to find proper stop epochs for this case. Also training procedure became more complicated. We added NIH dataset boxes for classes where it was possible. We prepared <a href=\"https://www.kaggle.com/zfturbo/nih-chest-xray-dataset-bbox-for-vinbigdata\" target=\"_blank\">NIH dataset in format of VinBigData Chest X-ray competition</a>. We added it as separate radiologist R99. We also prepeared <a href=\"https://www.kaggle.com/zfturbo/siimacr-pneumothorax-for-vinbigdata-chest-xray\" target=\"_blank\">dataset for Pneumothorax, based on SIIM competition</a>, but it didn't help. </p>\n<p>We used only one RetinaNet-R101 model for finetuning. We started from best previous weights obtained for multiclass case. So finetuning was pretty fast. We changed sampling in following way:</p>\n<ul>\n<li>We first select random radiologist, then select random class (target_class ot empty image), then select random image which were marked by selected radiologist and selected class.</li>\n<li>Validation was also slightly changed. We excluded all non-empty images which were marked by R8, R9 and R10 radiologists.</li>\n</ul>\n<p>In parallel we fine-tuned some models on very high resolution (2048px). Some classes from these models catch small boxes better than previous models. It improved our local validation.<br>\nThis approaches together allowed us to improve even more: &gt;0.36 on public</p>\n<h2>Part 4. Final steps</h2>\n<p>Actually we decided to create 2 submissions: best on validation and best on LB. We made first submission pretty early. So at the late stage of competition we mostly work on LB-overfitting. That was a little bit risky, but after private LB was revealed we found out it was right thing to do. )</p>\n<h3>Models we choose.</h3>\n<ul>\n<li>Best at local validation (around 0.51 mAP locally), Public: 0.304, private: 0.304</li>\n<li>Best at LB. Public: 0.365, private: 0.307</li>\n<li>Our best model at private has 0.319 score, but we had no chance to choose it because it had too low score on public - only 0.292. And it wasn't best at local validation.</li>\n</ul>\n<h2>Code</h2>\n<p>You can find code for out solution on github:<br>\n<a href=\"https://github.com/ZFTurbo/2nd-place-solution-for-VinBigData-Chest-X-ray-Abnormalities-Detection\" target=\"_blank\">https://github.com/ZFTurbo/2nd-place-solution-for-VinBigData-Chest-X-ray-Abnormalities-Detection</a></p>",
  "messages": [
    {
      "id": 1258009,
      "postDate": "2021-03-31T09:23:29.520Z",
      "content": "<h2>Part 1: Initial models</h2>\n<p>At first I started to try different OD models for this dataset. I tried: EffDet (B5), CenterNet (HourGlass), Yolo_v5, Retinanet (RsNet50, ResNet101, ResNet152), Faster RCNN</p>\n<p>Main initial tricks:</p>\n<ol>\n<li>To train models I used <a href=\"https://github.com/ZFTurbo/Weighted-Boxes-Fusion\" target=\"_blank\">WBF</a> on train.csv with IoU: 0.4. I usually used this file with most of models for training and for validation.</li>\n<li>Using validation I find out that it’s critical to use lower values for IoU in NMS which filters boxes on output of models. So I used NMS IoU: 0.25 for Yolo_v5 and NMS IoU: 0.3 for RetinaNet. Looks like, it was better for high confident boxes to consume more boxes around it.</li>\n<li>I didn’t want to change anchors for RetinaNet so I created version of train.csv with boxes which is close to square, because default Retina doesn’t like elongated boxes.</li>\n</ol>\n<p>Best models were the following:</p>\n<ul>\n<li>Yolo_v5: ~0.251 Public LB (Std + Mirror images 5KFold, 640px)</li>\n<li>RetinaNet (ResNet101): ~0.246 Public LB (Std + Mirror images 5KFold, 1024px)</li>\n<li>RetinaNet (ResNet152): ~0.222 Public LB (Std + Mirror images 5KFold, 800px)</li>\n<li>CenterNet (HourGlass): ~0.196 Public LB(Std + Mirror images 5KFold, 512px)</li>\n</ul>\n<p>EffDet – same code as Wheat competition, but works pretty poor on this dataset. With Faster-RCNN I think I have some problem with code, because it has the very bad score.</p>\n<p>Using WBF ensemble of these models + 2 public models (Detectron2 and Yolo_v5):</p>\n<ul>\n<li>0.221: <a href=\"https://www.kaggle.com/corochann/vinbigdata-detectron2-prediction\" target=\"_blank\">https://www.kaggle.com/corochann/vinbigdata-detectron2-prediction</a></li>\n<li>0.204: <a href=\"https://www.kaggle.com/awsaf49/vinbigdata-2-class-filter?scriptVersionId=51562634\" target=\"_blank\">https://www.kaggle.com/awsaf49/vinbigdata-2-class-filter?scriptVersionId=51562634</a></li>\n</ul>\n<p>I was able to achieve 0.300 on Public LB (0.289 Private LB). Then I merge with my teammates. They had more mmdetection models as well as Yolo_v5, but together we were able to got only 0.301.</p>\n<h2>Part 2: Investigation</h2>\n<p>While we were stuck we decided to get individual score for each class and compare with our OOF validation. For 3 days we got the following table:</p>\n<pre><code>Class 0: 0.014 (mAP LB: 0.210 Valid: 0.900)\nClass 1: 0.009 (mAP LB: 0.135 Valid: 0.335)\nClass 2: 0.017 (mAP LB: 0.255 Valid: 0.278)\nClass 3: 0.046 (mAP LB: 0.690 Valid: 0.915)\nClass 4: 0.019 (mAP LB: 0.285 Valid: 0.438)\nClass 5: 0.013 (mAP LB: 0.195 Valid: 0.376)\nClass 6: 0.010 (mAP LB: 0.150 Valid: 0.465)\nClass 7: 0.003 (mAP LB: 0.045 Valid: 0.400)\nClass 8: 0.011 (mAP LB: 0.165 Valid: 0.372)\nClass 9: 0.004 (mAP LB: 0.060 Valid: 0.162)\nClass 10: 0.027 (mAP LB: 0.405 Valid: 0.548)\nClass 11: 0.012 (mAP LB: 0.180 Valid: 0.343)\nClass 12: 0.037 (mAP LB: 0.555 Valid: 0.388)\nClass 13: 0.009 (mAP LB: 0.135 Valid: 0.432)\nClass 14: 0.063 (mAP LB: 0.945 Valid: 0.772)\nSum: 0.294\n</code></pre>\n<p>As it can be seen LB mAP the very different from local mAP. We started from investigation of Class 0 difference (it was most suspicious). We have near perfect prediction on validation while LB score is so small. The first thing we tried is to calculate validation mAP based on Radiologist ID.</p>\n<pre><code>Class ID: 0 Validation file: ensemble_iou_0.4._mAP_0.454_extracted_classes_0.csv\nRadiologist: R8 mAP: 0.991058\nRadiologist: R9 mAP: 0.954593\nRadiologist: R10 mAP: 0.971534\nRadiologist: R11 mAP: 0.379283\nRadiologist: R12 mAP: 0.019819\nRadiologist: R13 mAP: 0.511284\nRadiologist: R14 mAP: 0.527909\nRadiologist: R15 mAP: 0.279870\nRadiologist: R16 mAP: 0.506167\nRadiologist: R17 mAP: 0.447092\n</code></pre>\n<p>Interesting, we have perfect match for R8, R9 and R10, but very poor on other radiologists. Bad thing is that there are too small entries by other radiologists in train:</p>\n<pre><code>R11 - Class:  0 Entries:    30 AP: 0.379283\nR12 - Class:  0 Entries:    12 AP: 0.019819\nR13 - Class:  0 Entries:    32 AP: 0.511284\nR14 - Class:  0 Entries:    74 AP: 0.527909\nR15 - Class:  0 Entries:    25 AP: 0.279870\nR16 - Class:  0 Entries:    23 AP: 0.506167\nR17 - Class:  0 Entries:    10 AP: 0.447092\n</code></pre>\n<p>One more thing that R8, R9 and R10 never markup with any other radiologists. And what if test set marked up by some other radiologists (may be even by R13, R14 etc) and making boxes too similar to these 3 we overfit to train? And what if with better models we went further from test mark up?</p>\n<p>I checked by eyes images of class 0 in train marked up by R8, R9 and R10 and by rare radiologists to check the difference. I was lucky to see that other radiologists often use much larger boxes comparing to R8, R9 and R10. So I made a small trick by hand:<br>\n<code>x1 -= 100 y2 += 100 for all boxes with class 0</code><br>\nAnd checked local validation, which improved a lot for other radiologists, while R8, R9 and R10 became worse:</p>\n<pre><code>Radiologist: R10 Class:  0 Entries:  2349 AP: 0.559458\nRadiologist: R11 Class:  0 Entries:    30 AP: 0.559355\nRadiologist: R12 Class:  0 Entries:    12 AP: 0.351739\nRadiologist: R13 Class:  0 Entries:    32 AP: 0.483730\nRadiologist: R14 Class:  0 Entries:    74 AP: 0.623019\nRadiologist: R15 Class:  0 Entries:    25 AP: 0.341545\nRadiologist: R16 Class:  0 Entries:    23 AP: 0.506167\nRadiologist: R17 Class:  0 Entries:    10 AP: 0.452887\n</code></pre>\n<p>I made the same trick on our 0.301 submission and got 0.309 on LB. This showed us the right direction. We finetuned our models using only rare radiologists. And it improved LB score for all of models:</p>\n<p>RetinaNet (ResNet101): 0.246 -&gt; 0.264<br>\nRetinaNet (ResNet152): 0.222 -&gt; 0.237<br>\nYolo_v5: 0.251 -&gt; 0.271<br>\nEnsemble for all models including original moved us to Public LB &gt; 0.330</p>\n<h2>Part 3 (Single classes):</h2>\n<p>At this point it became hard to add new models to ensemble without proper validation. So we decided to switch to single class improvements. It was much easier to find proper stop epochs for this case. Also training procedure became more complicated. We added NIH dataset boxes for classes where it was possible. We prepared <a href=\"https://www.kaggle.com/zfturbo/nih-chest-xray-dataset-bbox-for-vinbigdata\" target=\"_blank\">NIH dataset in format of VinBigData Chest X-ray competition</a>. We added it as separate radiologist R99. We also prepeared <a href=\"https://www.kaggle.com/zfturbo/siimacr-pneumothorax-for-vinbigdata-chest-xray\" target=\"_blank\">dataset for Pneumothorax, based on SIIM competition</a>, but it didn't help. </p>\n<p>We used only one RetinaNet-R101 model for finetuning. We started from best previous weights obtained for multiclass case. So finetuning was pretty fast. We changed sampling in following way:</p>\n<ul>\n<li>We first select random radiologist, then select random class (target_class ot empty image), then select random image which were marked by selected radiologist and selected class.</li>\n<li>Validation was also slightly changed. We excluded all non-empty images which were marked by R8, R9 and R10 radiologists.</li>\n</ul>\n<p>In parallel we fine-tuned some models on very high resolution (2048px). Some classes from these models catch small boxes better than previous models. It improved our local validation.<br>\nThis approaches together allowed us to improve even more: &gt;0.36 on public</p>\n<h2>Part 4. Final steps</h2>\n<p>Actually we decided to create 2 submissions: best on validation and best on LB. We made first submission pretty early. So at the late stage of competition we mostly work on LB-overfitting. That was a little bit risky, but after private LB was revealed we found out it was right thing to do. )</p>\n<h3>Models we choose.</h3>\n<ul>\n<li>Best at local validation (around 0.51 mAP locally), Public: 0.304, private: 0.304</li>\n<li>Best at LB. Public: 0.365, private: 0.307</li>\n<li>Our best model at private has 0.319 score, but we had no chance to choose it because it had too low score on public - only 0.292. And it wasn't best at local validation.</li>\n</ul>\n<h2>Code</h2>\n<p>You can find code for out solution on github:<br>\n<a href=\"https://github.com/ZFTurbo/2nd-place-solution-for-VinBigData-Chest-X-ray-Abnormalities-Detection\" target=\"_blank\">https://github.com/ZFTurbo/2nd-place-solution-for-VinBigData-Chest-X-ray-Abnormalities-Detection</a></p>",
      "rawMarkdown": "## Part 1: Initial models\n\nAt first I started to try different OD models for this dataset. I tried: EffDet (B5), CenterNet (HourGlass), Yolo_v5, Retinanet (RsNet50, ResNet101, ResNet152), Faster RCNN\n\nMain initial tricks:\n1. To train models I used [WBF](https://github.com/ZFTurbo/Weighted-Boxes-Fusion) on train.csv with IoU: 0.4. I usually used this file with most of models for training and for validation.\n2. Using validation I find out that it’s critical to use lower values for IoU in NMS which filters boxes on output of models. So I used NMS IoU: 0.25 for Yolo_v5 and NMS IoU: 0.3 for RetinaNet. Looks like, it was better for high confident boxes to consume more boxes around it.\n3. I didn’t want to change anchors for RetinaNet so I created version of train.csv with boxes which is close to square, because default Retina doesn’t like elongated boxes.\n\nBest models were the following:\n- Yolo_v5: ~0.251 Public LB (Std + Mirror images 5KFold, 640px)\n- RetinaNet (ResNet101): ~0.246 Public LB (Std + Mirror images 5KFold, 1024px)\n- RetinaNet (ResNet152): ~0.222 Public LB (Std + Mirror images 5KFold, 800px)\n- CenterNet (HourGlass): ~0.196 Public LB(Std + Mirror images 5KFold, 512px)\n\nEffDet – same code as Wheat competition, but works pretty poor on this dataset. With Faster-RCNN I think I have some problem with code, because it has the very bad score.\n\nUsing WBF ensemble of these models + 2 public models (Detectron2 and Yolo_v5):\n- 0.221: https://www.kaggle.com/corochann/vinbigdata-detectron2-prediction\n- 0.204: https://www.kaggle.com/awsaf49/vinbigdata-2-class-filter?scriptVersionId=51562634\n\nI was able to achieve 0.300 on Public LB (0.289 Private LB). Then I merge with my teammates. They had more mmdetection models as well as Yolo_v5, but together we were able to got only 0.301.\n\n## Part 2: Investigation\nWhile we were stuck we decided to get individual score for each class and compare with our OOF validation. For 3 days we got the following table:\n\n```\nClass 0: 0.014 (mAP LB: 0.210 Valid: 0.900)\nClass 1: 0.009 (mAP LB: 0.135 Valid: 0.335)\nClass 2: 0.017 (mAP LB: 0.255 Valid: 0.278)\nClass 3: 0.046 (mAP LB: 0.690 Valid: 0.915)\nClass 4: 0.019 (mAP LB: 0.285 Valid: 0.438)\nClass 5: 0.013 (mAP LB: 0.195 Valid: 0.376)\nClass 6: 0.010 (mAP LB: 0.150 Valid: 0.465)\nClass 7: 0.003 (mAP LB: 0.045 Valid: 0.400)\nClass 8: 0.011 (mAP LB: 0.165 Valid: 0.372)\nClass 9: 0.004 (mAP LB: 0.060 Valid: 0.162)\nClass 10: 0.027 (mAP LB: 0.405 Valid: 0.548)\nClass 11: 0.012 (mAP LB: 0.180 Valid: 0.343)\nClass 12: 0.037 (mAP LB: 0.555 Valid: 0.388)\nClass 13: 0.009 (mAP LB: 0.135 Valid: 0.432)\nClass 14: 0.063 (mAP LB: 0.945 Valid: 0.772)\nSum: 0.294\n```\nAs it can be seen LB mAP the very different from local mAP. We started from investigation of Class 0 difference (it was most suspicious). We have near perfect prediction on validation while LB score is so small. The first thing we tried is to calculate validation mAP based on Radiologist ID.\n```\nClass ID: 0 Validation file: ensemble_iou_0.4._mAP_0.454_extracted_classes_0.csv\nRadiologist: R8 mAP: 0.991058\nRadiologist: R9 mAP: 0.954593\nRadiologist: R10 mAP: 0.971534\nRadiologist: R11 mAP: 0.379283\nRadiologist: R12 mAP: 0.019819\nRadiologist: R13 mAP: 0.511284\nRadiologist: R14 mAP: 0.527909\nRadiologist: R15 mAP: 0.279870\nRadiologist: R16 mAP: 0.506167\nRadiologist: R17 mAP: 0.447092\n```\n\nInteresting, we have perfect match for R8, R9 and R10, but very poor on other radiologists. Bad thing is that there are too small entries by other radiologists in train:\n```\nR11 - Class:  0 Entries:    30 AP: 0.379283\nR12 - Class:  0 Entries:    12 AP: 0.019819\nR13 - Class:  0 Entries:    32 AP: 0.511284\nR14 - Class:  0 Entries:    74 AP: 0.527909\nR15 - Class:  0 Entries:    25 AP: 0.279870\nR16 - Class:  0 Entries:    23 AP: 0.506167\nR17 - Class:  0 Entries:    10 AP: 0.447092\n```\nOne more thing that R8, R9 and R10 never markup with any other radiologists. And what if test set marked up by some other radiologists (may be even by R13, R14 etc) and making boxes too similar to these 3 we overfit to train? And what if with better models we went further from test mark up?\n\nI checked by eyes images of class 0 in train marked up by R8, R9 and R10 and by rare radiologists to check the difference. I was lucky to see that other radiologists often use much larger boxes comparing to R8, R9 and R10. So I made a small trick by hand:\n`x1 -= 100 y2 += 100 for all boxes with class 0 `\nAnd checked local validation, which improved a lot for other radiologists, while R8, R9 and R10 became worse:\n```\nRadiologist: R10 Class:  0 Entries:  2349 AP: 0.559458\nRadiologist: R11 Class:  0 Entries:    30 AP: 0.559355\nRadiologist: R12 Class:  0 Entries:    12 AP: 0.351739\nRadiologist: R13 Class:  0 Entries:    32 AP: 0.483730\nRadiologist: R14 Class:  0 Entries:    74 AP: 0.623019\nRadiologist: R15 Class:  0 Entries:    25 AP: 0.341545\nRadiologist: R16 Class:  0 Entries:    23 AP: 0.506167\nRadiologist: R17 Class:  0 Entries:    10 AP: 0.452887\n```\nI made the same trick on our 0.301 submission and got 0.309 on LB. This showed us the right direction. We finetuned our models using only rare radiologists. And it improved LB score for all of models:\n\nRetinaNet (ResNet101): 0.246 -> 0.264\nRetinaNet (ResNet152): 0.222 -> 0.237\nYolo_v5: 0.251 -> 0.271\nEnsemble for all models including original moved us to Public LB > 0.330\n\n## Part 3 (Single classes):\n\nAt this point it became hard to add new models to ensemble without proper validation. So we decided to switch to single class improvements. It was much easier to find proper stop epochs for this case. Also training procedure became more complicated. We added NIH dataset boxes for classes where it was possible. We prepared [NIH dataset in format of VinBigData Chest X-ray competition](https://www.kaggle.com/zfturbo/nih-chest-xray-dataset-bbox-for-vinbigdata). We added it as separate radiologist R99. We also prepeared [dataset for Pneumothorax, based on SIIM competition](https://www.kaggle.com/zfturbo/siimacr-pneumothorax-for-vinbigdata-chest-xray), but it didn't help. \n\nWe used only one RetinaNet-R101 model for finetuning. We started from best previous weights obtained for multiclass case. So finetuning was pretty fast. We changed sampling in following way:\n- We first select random radiologist, then select random class (target_class ot empty image), then select random image which were marked by selected radiologist and selected class.\n- Validation was also slightly changed. We excluded all non-empty images which were marked by R8, R9 and R10 radiologists.\n\nIn parallel we fine-tuned some models on very high resolution (2048px). Some classes from these models catch small boxes better than previous models. It improved our local validation.\nThis approaches together allowed us to improve even more: >0.36 on public\n\n## Part 4. Final steps\n\nActually we decided to create 2 submissions: best on validation and best on LB. We made first submission pretty early. So at the late stage of competition we mostly work on LB-overfitting. That was a little bit risky, but after private LB was revealed we found out it was right thing to do. )\n\n### Models we choose.\n- Best at local validation (around 0.51 mAP locally), Public: 0.304, private: 0.304\n- Best at LB. Public: 0.365, private: 0.307\n- Our best model at private has 0.319 score, but we had no chance to choose it because it had too low score on public - only 0.292. And it wasn't best at local validation.\n\n## Code\n\nYou can find code for out solution on github:\nhttps://github.com/ZFTurbo/2nd-place-solution-for-VinBigData-Chest-X-ray-Abnormalities-Detection\n",
      "votes": 93
    },
    {
      "id": 1258660,
      "postDate": "2021-03-31T19:37:33.237Z",
      "content": "<p>A small addition.</p>\n<p>At first, I trained using pictures with boxes only, and used postprocessing with a binary classifier using threshold to remove boxes. But then, following <a href=\"https://www.kaggle.com/zfturbo\" target=\"_blank\">@zfturbo</a>'s example, I trained with empty images and just adding 14 class with the probability of being healthy.</p>\n<p>It is tricky to write a special dataset loader in Yolo v5, so I decided to triple every image with boxes. To accelerate training, I started with weights from previous training on WBF markup (the first stage). The second stage was training on rare radiologists only. I have also tried the following techniques:<br>\n•    Choosing checkpoints for each class by the local validation. I see now it was profitable on private LB, but not so obvious on public LB.<br>\n•    Using pseudo-labels from the test set. I used ~700 sure unhealthy images. I don’t know the impact of that alone since it was along with tripling together. <br>\n•    Adding NIH dataset.<br>\n•    Training with higher resolution (didn’t give an evident boost)<br>\nThe best Yolo model for public LB is 0.292, but it is not so good on private LB. The best Yolo for private LB is 0.288.</p>",
      "rawMarkdown": "A small addition.\n\nAt first, I trained using pictures with boxes only, and used postprocessing with a binary classifier using threshold to remove boxes. But then, following @zfturbo's example, I trained with empty images and just adding 14 class with the probability of being healthy.\n\nIt is tricky to write a special dataset loader in Yolo v5, so I decided to triple every image with boxes. To accelerate training, I started with weights from previous training on WBF markup (the first stage). The second stage was training on rare radiologists only. I have also tried the following techniques:\n•\tChoosing checkpoints for each class by the local validation. I see now it was profitable on private LB, but not so obvious on public LB.\n•\tUsing pseudo-labels from the test set. I used ~700 sure unhealthy images. I don’t know the impact of that alone since it was along with tripling together. \n•\tAdding NIH dataset.\n•\tTraining with higher resolution (didn’t give an evident boost)\nThe best Yolo model for public LB is 0.292, but it is not so good on private LB. The best Yolo for private LB is 0.288.\n\n",
      "votes": 6
    },
    {
      "id": 1258027,
      "postDate": "2021-03-31T09:32:15.503Z",
      "content": "<p>Congrats and thanks a lot for this detailed elaboration. Do you know the boost you got by blending in public solutions? It seems to help quite a bit as also shown by others.</p>\n<p>We did very similar investigation on class0 and came to the same conclusion as you.</p>\n<p>We mostly based our whole validation on rare annotators as those are way more prominent in test.</p>",
      "rawMarkdown": "Congrats and thanks a lot for this detailed elaboration. Do you know the boost you got by blending in public solutions? It seems to help quite a bit as also shown by others.\n\nWe did very similar investigation on class0 and came to the same conclusion as you.\n\nWe mostly based our whole validation on rare annotators as those are way more prominent in test.",
      "votes": 6,
      "replies": [
        {
          "id": 1258040,
          "postDate": "2021-03-31T09:42:12.987Z",
          "content": "<p>The score boost from public solutions is due to class 12 (it's the rarest class actually). It is ~0.03 on public LB, but not so good on private LB</p>",
          "rawMarkdown": "The score boost from public solutions is due to class 12 (it's the rarest class actually). It is ~0.03 on public LB, but not so good on private LB",
          "votes": 7
        },
        {
          "id": 1258043,
          "postDate": "2021-03-31T09:43:21.677Z",
          "content": "<p>Thanks! As I remember public submissions catch rare 12 class [Pneumothorax] very good. I think there are only few (2-5?) boxes of Pneumothorax class on public. Looks like it can be the main source of shake up. )</p>\n<p>You can see it in my table (it was even better than our validation):<br>\n<code>Class 12: 0.037 (mAP LB: 0.555 Valid: 0.388)</code></p>",
          "rawMarkdown": "Thanks! As I remember public submissions catch rare 12 class [Pneumothorax] very good. I think there are only few (2-5?) boxes of Pneumothorax class on public. Looks like it can be the main source of shake up. )\n\nYou can see it in my table (it was even better than our validation):\n`Class 12: 0.037 (mAP LB: 0.555 Valid: 0.388)`",
          "votes": 3
        },
        {
          "id": 1258048,
          "postDate": "2021-03-31T09:45:35.967Z",
          "content": "<p>I see, we lacked subs to check all classes. What's the private score of that class 12 submission now?</p>",
          "rawMarkdown": "I see, we lacked subs to check all classes. What's the private score of that class 12 submission now?",
          "votes": 1
        },
        {
          "id": 1258054,
          "postDate": "2021-03-31T09:52:12.810Z",
          "content": "<p>Class 12: 0.037 -&gt; 0.013 (Note: as I remember it contains not only public solution, but ensembled with our models as well)</p>",
          "rawMarkdown": "Class 12: 0.037 -> 0.013 (Note: as I remember it contains not only public solution, but ensembled with our models as well)",
          "votes": 1
        },
        {
          "id": 1258071,
          "postDate": "2021-03-31T10:11:20.677Z",
          "content": "<p>So most likely the public score difference came mostly from class 12 and the gap got smaller after class 12 did worse on private. Thanks!</p>",
          "rawMarkdown": "So most likely the public score difference came mostly from class 12 and the gap got smaller after class 12 did worse on private. Thanks!",
          "votes": 3
        }
      ]
    },
    {
      "id": 1258081,
      "postDate": "2021-03-31T10:20:43.280Z",
      "content": "<p>Congrats for the second place and thanks for the solution<br>\nI have some questions <a href=\"https://www.kaggle.com/zfturbo\" target=\"_blank\">@zfturbo</a> :</p>\n<ol>\n<li>How do you use wbf on train.csv, do you consider each Radiologist as a \"model\" ?</li>\n<li>What is \"Std + Mirror images\", is this an data augmentation technique ?</li>\n<li>How do you \"finetuned our models using only rare radiologists\" ? </li>\n<li>Do you train \"no fingding\" images with your OD models ? Since i note that you have class 14 mAP on Local Valid ?</li>\n</ol>",
      "rawMarkdown": "Congrats for the second place and thanks for the solution\nI have some questions @zfturbo :\n1. How do you use wbf on train.csv, do you consider each Radiologist as a \"model\" ?\n2. What is \"Std + Mirror images\", is this an data augmentation technique ?\n3. How do you \"finetuned our models using only rare radiologists\" ? \n4. Do you train \"no fingding\" images with your OD models ? Since i note that you have class 14 mAP on Local Valid ?",
      "votes": 3,
      "replies": [
        {
          "id": 1258100,
          "postDate": "2021-03-31T10:51:07.843Z",
          "content": "<blockquote>\n  <p>How do you use wbf on train.csv, do you consider each Radiologist as a \"model\" ?</p>\n</blockquote>\n<p>I just apply WBF to all boxes for given image with same confidence = 1. If boxes intersect with large IoU (&gt;0.4) they fused in single box, otherwise stays \"as is\".</p>\n<blockquote>\n  <p>What is \"Std + Mirror images\", is this an data augmentation technique ?</p>\n</blockquote>\n<p>It's standard Test Time Augmentation. I used original image and mirrored image - then ensemble obtained boxes for these 2 images using WBF. </p>\n<blockquote>\n  <p>How do you \"finetuned our models using only rare radiologists\" ? </p>\n</blockquote>\n<p>I removed images which were marked by R8, R9 and R10 radiologists from training and from validation. Finetuning started from best weights of these models trained on full data.</p>\n<blockquote>\n  <p>Do you train \"no fingding\" images with your OD models ? Since i note that you have class 14 mAP on Local Valid?</p>\n</blockquote>\n<p>From the begining for class 14 I used predictions from binary classifier (EffNetB5). But I don't remember what exactly we used in final submission. I think we changed it to some other method involving object detection models, because it was better on LB. Need to ask Sergey.</p>",
          "rawMarkdown": "> How do you use wbf on train.csv, do you consider each Radiologist as a \"model\" ?\n\nI just apply WBF to all boxes for given image with same confidence = 1. If boxes intersect with large IoU (>0.4) they fused in single box, otherwise stays \"as is\".\n\n> What is \"Std + Mirror images\", is this an data augmentation technique ?\n\nIt's standard Test Time Augmentation. I used original image and mirrored image - then ensemble obtained boxes for these 2 images using WBF. \n\n> How do you \"finetuned our models using only rare radiologists\" ? \n\nI removed images which were marked by R8, R9 and R10 radiologists from training and from validation. Finetuning started from best weights of these models trained on full data.\n\n> Do you train \"no fingding\" images with your OD models ? Since i note that you have class 14 mAP on Local Valid?\n\nFrom the begining for class 14 I used predictions from binary classifier (EffNetB5). But I don't remember what exactly we used in final submission. I think we changed it to some other method involving object detection models, because it was better on LB. Need to ask Sergey.",
          "votes": 3
        },
        {
          "id": 1258120,
          "postDate": "2021-03-31T11:15:16.053Z",
          "content": "<p>Once (accidentally) we made WBF for class 14, there was the binary classifier and the result of an OD model. And we saw that this is better (on public LB) than a pure classifier. Then we used this variant. We need to check if there is an effect on private LB from this.</p>",
          "rawMarkdown": "Once (accidentally) we made WBF for class 14, there was the binary classifier and the result of an OD model. And we saw that this is better (on public LB) than a pure classifier. Then we used this variant. We need to check if there is an effect on private LB from this.",
          "votes": 3
        },
        {
          "id": 1258195,
          "postDate": "2021-03-31T12:38:21.280Z",
          "content": "<p>Interesting. What was your best 14 class AP score on the private LB?</p>",
          "rawMarkdown": "Interesting. What was your best 14 class AP score on the private LB?"
        },
        {
          "id": 1258254,
          "postDate": "2021-03-31T13:39:36.197Z",
          "content": "<p>I've checked all our variants on private LB. And we used the right one :)<br>\n0.064 both private and public LB<br>\nIt is the variant with object detection injected. Although the boost from OD is quite small.</p>",
          "rawMarkdown": "I've checked all our variants on private LB. And we used the right one :)\n0.064 both private and public LB\nIt is the variant with object detection injected. Although the boost from OD is quite small.\n",
          "votes": 2
        },
        {
          "id": 1258883,
          "postDate": "2021-04-01T01:14:32.417Z",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/zfturbo\" target=\"_blank\">@zfturbo</a> <a href=\"https://www.kaggle.com/sergeyzlobin\" target=\"_blank\">@sergeyzlobin</a> </p>",
          "rawMarkdown": "Thank you @zfturbo @sergeyzlobin ",
          "votes": 1
        },
        {
          "id": 1258887,
          "postDate": "2021-04-01T01:31:15.047Z",
          "content": "<p>When you use the OD model to detect the 14th class, do you set the bbox to 0 0 1 1 or something else?</p>",
          "rawMarkdown": "When you use the OD model to detect the 14th class, do you set the bbox to 0 0 1 1 or something else?"
        }
      ]
    },
    {
      "id": 1258117,
      "postDate": "2021-03-31T11:13:31.777Z",
      "content": "<p>Congratz on #2 Win. You were a worthy competitor for us.</p>",
      "rawMarkdown": "Congratz on #2 Win. You were a worthy competitor for us.",
      "votes": 3
    },
    {
      "id": 1258809,
      "postDate": "2021-03-31T22:47:22.013Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/zfturbo\" target=\"_blank\">@zfturbo</a> and team! Great detective work analyzing the classes. And great discovery that IOU &lt; 0.4 worked well. We didn't try low IOU threshold values.</p>\n<p>Thanks for sharing your <a href=\"https://github.com/ZFTurbo/Weighted-Boxes-Fusion\" target=\"_blank\">Ensemble Boxes GitHub repo</a>, we used it and it is awesome!</p>",
      "rawMarkdown": "Congratulations @zfturbo and team! Great detective work analyzing the classes. And great discovery that IOU < 0.4 worked well. We didn't try low IOU threshold values.\n\nThanks for sharing your [Ensemble Boxes GitHub repo][1], we used it and it is awesome!\n\n[1]: https://github.com/ZFTurbo/Weighted-Boxes-Fusion",
      "votes": 1
    },
    {
      "id": 1258129,
      "postDate": "2021-03-31T11:23:05.523Z",
      "content": "<p><a href=\"https://www.kaggle.com/zfturbo\" target=\"_blank\">@zfturbo</a> amazing. Congrats. I had spent time on CenterNet. Wish I had seen that one all the way through!</p>",
      "rawMarkdown": "@zfturbo amazing. Congrats. I had spent time on CenterNet. Wish I had seen that one all the way through!",
      "votes": 1
    },
    {
      "id": 1258070,
      "postDate": "2021-03-31T10:11:19.870Z",
      "content": "<p>Congratulations on the second place and thanks for the detailed explanation!<br>\nI got a lot from <a href=\"https://www.kaggle.com/zfturbo\" target=\"_blank\">@zfturbo</a> team's great work.</p>",
      "rawMarkdown": "Congratulations on the second place and thanks for the detailed explanation!\nI got a lot from @zfturbo team's great work.",
      "votes": 1
    },
    {
      "id": 1258058,
      "postDate": "2021-03-31T09:56:51.427Z",
      "content": "<p>Congratulations and thanks a lot for the detailed solution! The investigation and your approach is so amazing!</p>",
      "rawMarkdown": "Congratulations and thanks a lot for the detailed solution! The investigation and your approach is so amazing!",
      "votes": 1
    },
    {
      "id": 1258051,
      "postDate": "2021-03-31T09:48:37.587Z",
      "content": "<p>Congrats ! Your investigation on individual classes and radiologists is very strong</p>",
      "rawMarkdown": "Congrats ! Your investigation on individual classes and radiologists is very strong",
      "votes": 1
    },
    {
      "id": 1258055,
      "postDate": "2021-03-31T09:52:32.623Z",
      "content": "<p>Congrats for the second place and thanks for the detailed writeup! </p>\n<p>We also noticed the systematic differences between annotators  (<a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229642\" target=\"_blank\">https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229642</a>)<br>\nI wonder if applying the bigger boxes finetune trick only for images not annotated by R8-R10 would further boost the score.</p>",
      "rawMarkdown": "Congrats for the second place and thanks for the detailed writeup! \n\n\nWe also noticed the systematic differences between annotators  (https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229642)\nI wonder if applying the bigger boxes finetune trick only for images not annotated by R8-R10 would further boost the score.",
      "votes": 2
    },
    {
      "id": 1258038,
      "postDate": "2021-03-31T09:40:09.613Z",
      "content": "<p>Congratulations for 2nd place and thanks for sharing detailed solution <a href=\"https://www.kaggle.com/zfturbo\" target=\"_blank\">@zfturbo</a> and team</p>",
      "rawMarkdown": "Congratulations for 2nd place and thanks for sharing detailed solution @zfturbo and team",
      "votes": 2
    },
    {
      "id": 1276625,
      "postDate": "2021-04-17T18:46:44.387Z",
      "content": "<p>Congrats! Great work!</p>",
      "rawMarkdown": "Congrats! Great work!"
    },
    {
      "id": 1260921,
      "postDate": "2021-04-02T14:34:18.100Z",
      "content": "<p>Congratz, which repo did you use for *<em>RetinaNet</em> ? <a href=\"https://www.kaggle.com/zfturbo\" target=\"_blank\">@zfturbo</a> </p>",
      "rawMarkdown": "Congratz, which repo did you use for **RetinaNet* ? @zfturbo ",
      "replies": [
        {
          "id": 1260965,
          "postDate": "2021-04-02T15:03:34.087Z",
          "content": "<p>I still use this code:<br>\n<a href=\"https://github.com/ZFTurbo/Keras-RetinaNet-for-Open-Images-Challenge-2018\" target=\"_blank\">https://github.com/ZFTurbo/Keras-RetinaNet-for-Open-Images-Challenge-2018</a></p>\n<p>It's outdated a little bit )</p>",
          "rawMarkdown": "I still use this code:\nhttps://github.com/ZFTurbo/Keras-RetinaNet-for-Open-Images-Challenge-2018\n\nIt's outdated a little bit )",
          "votes": 1
        }
      ]
    },
    {
      "id": 1259469,
      "postDate": "2021-04-01T12:34:12.717Z",
      "content": "<p>Congrats, <a href=\"https://www.kaggle.com/zfturbo\" target=\"_blank\">@zfturbo</a> ! Amazing work</p>",
      "rawMarkdown": "Congrats, @zfturbo ! Amazing work"
    },
    {
      "id": 1259370,
      "postDate": "2021-04-01T10:43:26.160Z",
      "content": "<p>\"The first thing we tried is to calculate validation mAP based on Radiologist ID.\"<br>\n=&gt; How did you calculate the mAP for each Radiologist ? Did you just remove all the other radiologists' bbox in ground truth and calculate just the remained one ? </p>",
      "rawMarkdown": "\"The first thing we tried is to calculate validation mAP based on Radiologist ID.\"\n=> How did you calculate the mAP for each Radiologist ? Did you just remove all the other radiologists' bbox in ground truth and calculate just the remained one ? ",
      "replies": [
        {
          "id": 1259679,
          "postDate": "2021-04-01T15:21:41.427Z",
          "content": "<p>Yes, just use images with bboxes marked by given radiologist discarding all other radiologist for this image. I don't remember exactly, but probably I also used all empty images.</p>",
          "rawMarkdown": "Yes, just use images with bboxes marked by given radiologist discarding all other radiologist for this image. I don't remember exactly, but probably I also used all empty images.",
          "votes": 1
        },
        {
          "id": 1260258,
          "postDate": "2021-04-02T01:12:44.233Z",
          "content": "<p>Thank you, and congrats on the Prize !</p>",
          "rawMarkdown": "Thank you, and congrats on the Prize !"
        },
        {
          "id": 1260535,
          "postDate": "2021-04-02T07:52:31.503Z",
          "content": "<blockquote>\n  <p>I don't remember exactly, but probably I also used all empty images.</p>\n</blockquote>\n<p>Yes, empty images were included. That's why the validation was not fast. :(</p>",
          "rawMarkdown": "> I don't remember exactly, but probably I also used all empty images.\n\nYes, empty images were included. That's why the validation was not fast. :(",
          "votes": 1
        }
      ]
    },
    {
      "id": 1258035,
      "postDate": "2021-03-31T09:38:59.577Z",
      "content": "<p><a href=\"https://www.kaggle.com/zfturbo\" target=\"_blank\">@zfturbo</a> Congratulations on 2 nd Prize and Thanks for sharing the approach</p>",
      "rawMarkdown": "@zfturbo Congratulations on 2 nd Prize and Thanks for sharing the approach"
    }
  ],
  "comments": [
    {
      "id": 1258660,
      "author_name": "Sergey Zlobin",
      "author_url": "",
      "post_date": "2021-03-31T19:37:33.237000",
      "content": "<p>A small addition.</p>\n<p>At first, I trained using pictures with boxes only, and used postprocessing with a binary classifier using threshold to remove boxes. But then, following <a href=\"https://www.kaggle.com/zfturbo\" target=\"_blank\">@zfturbo</a>'s example, I trained with empty images and just adding 14 class with the probability of being healthy.</p>\n<p>It is tricky to write a special dataset loader in Yolo v5, so I decided to triple every image with boxes. To accelerate training, I started with weights from previous training on WBF markup (the first stage). The second stage was training on rare radiologists only. I have also tried the following techniques:<br>\n•    Choosing checkpoints for each class by the local validation. I see now it was profitable on private LB, but not so obvious on public LB.<br>\n•    Using pseudo-labels from the test set. I used ~700 sure unhealthy images. I don’t know the impact of that alone since it was along with tripling together. <br>\n•    Adding NIH dataset.<br>\n•    Training with higher resolution (didn’t give an evident boost)<br>\nThe best Yolo model for public LB is 0.292, but it is not so good on private LB. The best Yolo for private LB is 0.288.</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 1258027,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2021-03-31T09:32:15.503000",
      "content": "<p>Congrats and thanks a lot for this detailed elaboration. Do you know the boost you got by blending in public solutions? It seems to help quite a bit as also shown by others.</p>\n<p>We did very similar investigation on class0 and came to the same conclusion as you.</p>\n<p>We mostly based our whole validation on rare annotators as those are way more prominent in test.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 1258040,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2021-03-31T09:42:12.987000",
          "content": "<p>The score boost from public solutions is due to class 12 (it's the rarest class actually). It is ~0.03 on public LB, but not so good on private LB</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 1258043,
          "author_name": "ZFTurbo",
          "author_url": "",
          "post_date": "2021-03-31T09:43:21.677000",
          "content": "<p>Thanks! As I remember public submissions catch rare 12 class [Pneumothorax] very good. I think there are only few (2-5?) boxes of Pneumothorax class on public. Looks like it can be the main source of shake up. )</p>\n<p>You can see it in my table (it was even better than our validation):<br>\n<code>Class 12: 0.037 (mAP LB: 0.555 Valid: 0.388)</code></p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1258048,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-03-31T09:45:35.967000",
          "content": "<p>I see, we lacked subs to check all classes. What's the private score of that class 12 submission now?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1258054,
          "author_name": "ZFTurbo",
          "author_url": "",
          "post_date": "2021-03-31T09:52:12.810000",
          "content": "<p>Class 12: 0.037 -&gt; 0.013 (Note: as I remember it contains not only public solution, but ensembled with our models as well)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1258071,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-03-31T10:11:20.677000",
          "content": "<p>So most likely the public score difference came mostly from class 12 and the gap got smaller after class 12 did worse on private. Thanks!</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1258081,
      "author_name": "Bùi Nhật Trường",
      "author_url": "",
      "post_date": "2021-03-31T10:20:43.280000",
      "content": "<p>Congrats for the second place and thanks for the solution<br>\nI have some questions <a href=\"https://www.kaggle.com/zfturbo\" target=\"_blank\">@zfturbo</a> :</p>\n<ol>\n<li>How do you use wbf on train.csv, do you consider each Radiologist as a \"model\" ?</li>\n<li>What is \"Std + Mirror images\", is this an data augmentation technique ?</li>\n<li>How do you \"finetuned our models using only rare radiologists\" ? </li>\n<li>Do you train \"no fingding\" images with your OD models ? Since i note that you have class 14 mAP on Local Valid ?</li>\n</ol>",
      "votes": 3,
      "replies": [
        {
          "id": 1258100,
          "author_name": "ZFTurbo",
          "author_url": "",
          "post_date": "2021-03-31T10:51:07.843000",
          "content": "<blockquote>\n  <p>How do you use wbf on train.csv, do you consider each Radiologist as a \"model\" ?</p>\n</blockquote>\n<p>I just apply WBF to all boxes for given image with same confidence = 1. If boxes intersect with large IoU (&gt;0.4) they fused in single box, otherwise stays \"as is\".</p>\n<blockquote>\n  <p>What is \"Std + Mirror images\", is this an data augmentation technique ?</p>\n</blockquote>\n<p>It's standard Test Time Augmentation. I used original image and mirrored image - then ensemble obtained boxes for these 2 images using WBF. </p>\n<blockquote>\n  <p>How do you \"finetuned our models using only rare radiologists\" ? </p>\n</blockquote>\n<p>I removed images which were marked by R8, R9 and R10 radiologists from training and from validation. Finetuning started from best weights of these models trained on full data.</p>\n<blockquote>\n  <p>Do you train \"no fingding\" images with your OD models ? Since i note that you have class 14 mAP on Local Valid?</p>\n</blockquote>\n<p>From the begining for class 14 I used predictions from binary classifier (EffNetB5). But I don't remember what exactly we used in final submission. I think we changed it to some other method involving object detection models, because it was better on LB. Need to ask Sergey.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1258120,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2021-03-31T11:15:16.053000",
          "content": "<p>Once (accidentally) we made WBF for class 14, there was the binary classifier and the result of an OD model. And we saw that this is better (on public LB) than a pure classifier. Then we used this variant. We need to check if there is an effect on private LB from this.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1258195,
          "author_name": "beluga",
          "author_url": "",
          "post_date": "2021-03-31T12:38:21.280000",
          "content": "<p>Interesting. What was your best 14 class AP score on the private LB?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1258254,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2021-03-31T13:39:36.197000",
          "content": "<p>I've checked all our variants on private LB. And we used the right one :)<br>\n0.064 both private and public LB<br>\nIt is the variant with object detection injected. Although the boost from OD is quite small.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1258883,
          "author_name": "Bùi Nhật Trường",
          "author_url": "",
          "post_date": "2021-04-01T01:14:32.417000",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/zfturbo\" target=\"_blank\">@zfturbo</a> <a href=\"https://www.kaggle.com/sergeyzlobin\" target=\"_blank\">@sergeyzlobin</a> </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1258887,
          "author_name": "tik_boa",
          "author_url": "",
          "post_date": "2021-04-01T01:31:15.047000",
          "content": "<p>When you use the OD model to detect the 14th class, do you set the bbox to 0 0 1 1 or something else?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1258117,
      "author_name": "Mohammed Rizin V K",
      "author_url": "",
      "post_date": "2021-03-31T11:13:31.777000",
      "content": "<p>Congratz on #2 Win. You were a worthy competitor for us.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1258809,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2021-03-31T22:47:22.013000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/zfturbo\" target=\"_blank\">@zfturbo</a> and team! Great detective work analyzing the classes. And great discovery that IOU &lt; 0.4 worked well. We didn't try low IOU threshold values.</p>\n<p>Thanks for sharing your <a href=\"https://github.com/ZFTurbo/Weighted-Boxes-Fusion\" target=\"_blank\">Ensemble Boxes GitHub repo</a>, we used it and it is awesome!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1258129,
      "author_name": "Charlie Craine",
      "author_url": "",
      "post_date": "2021-03-31T11:23:05.523000",
      "content": "<p><a href=\"https://www.kaggle.com/zfturbo\" target=\"_blank\">@zfturbo</a> amazing. Congrats. I had spent time on CenterNet. Wish I had seen that one all the way through!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1258070,
      "author_name": "Wang Xing",
      "author_url": "",
      "post_date": "2021-03-31T10:11:19.870000",
      "content": "<p>Congratulations on the second place and thanks for the detailed explanation!<br>\nI got a lot from <a href=\"https://www.kaggle.com/zfturbo\" target=\"_blank\">@zfturbo</a> team's great work.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1258058,
      "author_name": "Sunghyun Jun",
      "author_url": "",
      "post_date": "2021-03-31T09:56:51.427000",
      "content": "<p>Congratulations and thanks a lot for the detailed solution! The investigation and your approach is so amazing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1258051,
      "author_name": "Matthieu Planté",
      "author_url": "",
      "post_date": "2021-03-31T09:48:37.587000",
      "content": "<p>Congrats ! Your investigation on individual classes and radiologists is very strong</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1258055,
      "author_name": "beluga",
      "author_url": "",
      "post_date": "2021-03-31T09:52:32.623000",
      "content": "<p>Congrats for the second place and thanks for the detailed writeup! </p>\n<p>We also noticed the systematic differences between annotators  (<a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229642\" target=\"_blank\">https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229642</a>)<br>\nI wonder if applying the bigger boxes finetune trick only for images not annotated by R8-R10 would further boost the score.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1258038,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2021-03-31T09:40:09.613000",
      "content": "<p>Congratulations for 2nd place and thanks for sharing detailed solution <a href=\"https://www.kaggle.com/zfturbo\" target=\"_blank\">@zfturbo</a> and team</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1276625,
      "author_name": "Sai Prakash",
      "author_url": "",
      "post_date": "2021-04-17T18:46:44.387000",
      "content": "<p>Congrats! Great work!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1260921,
      "author_name": "Awsaf",
      "author_url": "",
      "post_date": "2021-04-02T14:34:18.100000",
      "content": "<p>Congratz, which repo did you use for *<em>RetinaNet</em> ? <a href=\"https://www.kaggle.com/zfturbo\" target=\"_blank\">@zfturbo</a> </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1260965,
          "author_name": "ZFTurbo",
          "author_url": "",
          "post_date": "2021-04-02T15:03:34.087000",
          "content": "<p>I still use this code:<br>\n<a href=\"https://github.com/ZFTurbo/Keras-RetinaNet-for-Open-Images-Challenge-2018\" target=\"_blank\">https://github.com/ZFTurbo/Keras-RetinaNet-for-Open-Images-Challenge-2018</a></p>\n<p>It's outdated a little bit )</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1259469,
      "author_name": "Sidney Ng",
      "author_url": "",
      "post_date": "2021-04-01T12:34:12.717000",
      "content": "<p>Congrats, <a href=\"https://www.kaggle.com/zfturbo\" target=\"_blank\">@zfturbo</a> ! Amazing work</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1259370,
      "author_name": "Tung Vu",
      "author_url": "",
      "post_date": "2021-04-01T10:43:26.160000",
      "content": "<p>\"The first thing we tried is to calculate validation mAP based on Radiologist ID.\"<br>\n=&gt; How did you calculate the mAP for each Radiologist ? Did you just remove all the other radiologists' bbox in ground truth and calculate just the remained one ? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1259679,
          "author_name": "ZFTurbo",
          "author_url": "",
          "post_date": "2021-04-01T15:21:41.427000",
          "content": "<p>Yes, just use images with bboxes marked by given radiologist discarding all other radiologist for this image. I don't remember exactly, but probably I also used all empty images.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1260258,
          "author_name": "Tung Vu",
          "author_url": "",
          "post_date": "2021-04-02T01:12:44.233000",
          "content": "<p>Thank you, and congrats on the Prize !</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1260535,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2021-04-02T07:52:31.503000",
          "content": "<blockquote>\n  <p>I don't remember exactly, but probably I also used all empty images.</p>\n</blockquote>\n<p>Yes, empty images were included. That's why the validation was not fast. :(</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1258035,
      "author_name": "Tensor Girl",
      "author_url": "",
      "post_date": "2021-03-31T09:38:59.577000",
      "content": "<p><a href=\"https://www.kaggle.com/zfturbo\" target=\"_blank\">@zfturbo</a> Congratulations on 2 nd Prize and Thanks for sharing the approach</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1258009": "## Part 1: Initial models\n\nAt first I started to try different OD models for this dataset. I tried: EffDet (B5), CenterNet (HourGlass), Yolo_v5, Retinanet (RsNet50, ResNet101, ResNet152), Faster RCNN\n\nMain initial tricks:\n1. To train models I used [WBF](https://github.com/ZFTurbo/Weighted-Boxes-Fusion) on train.csv with IoU: 0.4. I usually used this file with most of models for training and for validation.\n2. Using validation I find out that it’s critical to use lower values for IoU in NMS which filters boxes on output of models. So I used NMS IoU: 0.25 for Yolo_v5 and NMS IoU: 0.3 for RetinaNet. Looks like, it was better for high confident boxes to consume more boxes around it.\n3. I didn’t want to change anchors for RetinaNet so I created version of train.csv with boxes which is close to square, because default Retina doesn’t like elongated boxes.\n\nBest models were the following:\n- Yolo_v5: ~0.251 Public LB (Std + Mirror images 5KFold, 640px)\n- RetinaNet (ResNet101): ~0.246 Public LB (Std + Mirror images 5KFold, 1024px)\n- RetinaNet (ResNet152): ~0.222 Public LB (Std + Mirror images 5KFold, 800px)\n- CenterNet (HourGlass): ~0.196 Public LB(Std + Mirror images 5KFold, 512px)\n\nEffDet – same code as Wheat competition, but works pretty poor on this dataset. With Faster-RCNN I think I have some problem with code, because it has the very bad score.\n\nUsing WBF ensemble of these models + 2 public models (Detectron2 and Yolo_v5):\n- 0.221: https://www.kaggle.com/corochann/vinbigdata-detectron2-prediction\n- 0.204: https://www.kaggle.com/awsaf49/vinbigdata-2-class-filter?scriptVersionId=51562634\n\nI was able to achieve 0.300 on Public LB (0.289 Private LB). Then I merge with my teammates. They had more mmdetection models as well as Yolo_v5, but together we were able to got only 0.301.\n\n## Part 2: Investigation\nWhile we were stuck we decided to get individual score for each class and compare with our OOF validation. For 3 days we got the following table:\n\n```\nClass 0: 0.014 (mAP LB: 0.210 Valid: 0.900)\nClass 1: 0.009 (mAP LB: 0.135 Valid: 0.335)\nClass 2: 0.017 (mAP LB: 0.255 Valid: 0.278)\nClass 3: 0.046 (mAP LB: 0.690 Valid: 0.915)\nClass 4: 0.019 (mAP LB: 0.285 Valid: 0.438)\nClass 5: 0.013 (mAP LB: 0.195 Valid: 0.376)\nClass 6: 0.010 (mAP LB: 0.150 Valid: 0.465)\nClass 7: 0.003 (mAP LB: 0.045 Valid: 0.400)\nClass 8: 0.011 (mAP LB: 0.165 Valid: 0.372)\nClass 9: 0.004 (mAP LB: 0.060 Valid: 0.162)\nClass 10: 0.027 (mAP LB: 0.405 Valid: 0.548)\nClass 11: 0.012 (mAP LB: 0.180 Valid: 0.343)\nClass 12: 0.037 (mAP LB: 0.555 Valid: 0.388)\nClass 13: 0.009 (mAP LB: 0.135 Valid: 0.432)\nClass 14: 0.063 (mAP LB: 0.945 Valid: 0.772)\nSum: 0.294\n```\nAs it can be seen LB mAP the very different from local mAP. We started from investigation of Class 0 difference (it was most suspicious). We have near perfect prediction on validation while LB score is so small. The first thing we tried is to calculate validation mAP based on Radiologist ID.\n```\nClass ID: 0 Validation file: ensemble_iou_0.4._mAP_0.454_extracted_classes_0.csv\nRadiologist: R8 mAP: 0.991058\nRadiologist: R9 mAP: 0.954593\nRadiologist: R10 mAP: 0.971534\nRadiologist: R11 mAP: 0.379283\nRadiologist: R12 mAP: 0.019819\nRadiologist: R13 mAP: 0.511284\nRadiologist: R14 mAP: 0.527909\nRadiologist: R15 mAP: 0.279870\nRadiologist: R16 mAP: 0.506167\nRadiologist: R17 mAP: 0.447092\n```\n\nInteresting, we have perfect match for R8, R9 and R10, but very poor on other radiologists. Bad thing is that there are too small entries by other radiologists in train:\n```\nR11 - Class:  0 Entries:    30 AP: 0.379283\nR12 - Class:  0 Entries:    12 AP: 0.019819\nR13 - Class:  0 Entries:    32 AP: 0.511284\nR14 - Class:  0 Entries:    74 AP: 0.527909\nR15 - Class:  0 Entries:    25 AP: 0.279870\nR16 - Class:  0 Entries:    23 AP: 0.506167\nR17 - Class:  0 Entries:    10 AP: 0.447092\n```\nOne more thing that R8, R9 and R10 never markup with any other radiologists. And what if test set marked up by some other radiologists (may be even by R13, R14 etc) and making boxes too similar to these 3 we overfit to train? And what if with better models we went further from test mark up?\n\nI checked by eyes images of class 0 in train marked up by R8, R9 and R10 and by rare radiologists to check the difference. I was lucky to see that other radiologists often use much larger boxes comparing to R8, R9 and R10. So I made a small trick by hand:\n`x1 -= 100 y2 += 100 for all boxes with class 0 `\nAnd checked local validation, which improved a lot for other radiologists, while R8, R9 and R10 became worse:\n```\nRadiologist: R10 Class:  0 Entries:  2349 AP: 0.559458\nRadiologist: R11 Class:  0 Entries:    30 AP: 0.559355\nRadiologist: R12 Class:  0 Entries:    12 AP: 0.351739\nRadiologist: R13 Class:  0 Entries:    32 AP: 0.483730\nRadiologist: R14 Class:  0 Entries:    74 AP: 0.623019\nRadiologist: R15 Class:  0 Entries:    25 AP: 0.341545\nRadiologist: R16 Class:  0 Entries:    23 AP: 0.506167\nRadiologist: R17 Class:  0 Entries:    10 AP: 0.452887\n```\nI made the same trick on our 0.301 submission and got 0.309 on LB. This showed us the right direction. We finetuned our models using only rare radiologists. And it improved LB score for all of models:\n\nRetinaNet (ResNet101): 0.246 -> 0.264\nRetinaNet (ResNet152): 0.222 -> 0.237\nYolo_v5: 0.251 -> 0.271\nEnsemble for all models including original moved us to Public LB > 0.330\n\n## Part 3 (Single classes):\n\nAt this point it became hard to add new models to ensemble without proper validation. So we decided to switch to single class improvements. It was much easier to find proper stop epochs for this case. Also training procedure became more complicated. We added NIH dataset boxes for classes where it was possible. We prepared [NIH dataset in format of VinBigData Chest X-ray competition](https://www.kaggle.com/zfturbo/nih-chest-xray-dataset-bbox-for-vinbigdata). We added it as separate radiologist R99. We also prepeared [dataset for Pneumothorax, based on SIIM competition](https://www.kaggle.com/zfturbo/siimacr-pneumothorax-for-vinbigdata-chest-xray), but it didn't help. \n\nWe used only one RetinaNet-R101 model for finetuning. We started from best previous weights obtained for multiclass case. So finetuning was pretty fast. We changed sampling in following way:\n- We first select random radiologist, then select random class (target_class ot empty image), then select random image which were marked by selected radiologist and selected class.\n- Validation was also slightly changed. We excluded all non-empty images which were marked by R8, R9 and R10 radiologists.\n\nIn parallel we fine-tuned some models on very high resolution (2048px). Some classes from these models catch small boxes better than previous models. It improved our local validation.\nThis approaches together allowed us to improve even more: >0.36 on public\n\n## Part 4. Final steps\n\nActually we decided to create 2 submissions: best on validation and best on LB. We made first submission pretty early. So at the late stage of competition we mostly work on LB-overfitting. That was a little bit risky, but after private LB was revealed we found out it was right thing to do. )\n\n### Models we choose.\n- Best at local validation (around 0.51 mAP locally), Public: 0.304, private: 0.304\n- Best at LB. Public: 0.365, private: 0.307\n- Our best model at private has 0.319 score, but we had no chance to choose it because it had too low score on public - only 0.292. And it wasn't best at local validation.\n\n## Code\n\nYou can find code for out solution on github:\nhttps://github.com/ZFTurbo/2nd-place-solution-for-VinBigData-Chest-X-ray-Abnormalities-Detection\n",
    "1258660": "A small addition.\n\nAt first, I trained using pictures with boxes only, and used postprocessing with a binary classifier using threshold to remove boxes. But then, following @zfturbo's example, I trained with empty images and just adding 14 class with the probability of being healthy.\n\nIt is tricky to write a special dataset loader in Yolo v5, so I decided to triple every image with boxes. To accelerate training, I started with weights from previous training on WBF markup (the first stage). The second stage was training on rare radiologists only. I have also tried the following techniques:\n•\tChoosing checkpoints for each class by the local validation. I see now it was profitable on private LB, but not so obvious on public LB.\n•\tUsing pseudo-labels from the test set. I used ~700 sure unhealthy images. I don’t know the impact of that alone since it was along with tripling together. \n•\tAdding NIH dataset.\n•\tTraining with higher resolution (didn’t give an evident boost)\nThe best Yolo model for public LB is 0.292, but it is not so good on private LB. The best Yolo for private LB is 0.288.\n\n",
    "1258027": "Congrats and thanks a lot for this detailed elaboration. Do you know the boost you got by blending in public solutions? It seems to help quite a bit as also shown by others.\n\nWe did very similar investigation on class0 and came to the same conclusion as you.\n\nWe mostly based our whole validation on rare annotators as those are way more prominent in test.",
    "1258081": "Congrats for the second place and thanks for the solution\nI have some questions @zfturbo :\n1. How do you use wbf on train.csv, do you consider each Radiologist as a \"model\" ?\n2. What is \"Std + Mirror images\", is this an data augmentation technique ?\n3. How do you \"finetuned our models using only rare radiologists\" ? \n4. Do you train \"no fingding\" images with your OD models ? Since i note that you have class 14 mAP on Local Valid ?",
    "1258117": "Congratz on #2 Win. You were a worthy competitor for us.",
    "1258809": "Congratulations @zfturbo and team! Great detective work analyzing the classes. And great discovery that IOU < 0.4 worked well. We didn't try low IOU threshold values.\n\nThanks for sharing your [Ensemble Boxes GitHub repo][1], we used it and it is awesome!\n\n[1]: https://github.com/ZFTurbo/Weighted-Boxes-Fusion",
    "1258129": "@zfturbo amazing. Congrats. I had spent time on CenterNet. Wish I had seen that one all the way through!",
    "1258070": "Congratulations on the second place and thanks for the detailed explanation!\nI got a lot from @zfturbo team's great work.",
    "1258058": "Congratulations and thanks a lot for the detailed solution! The investigation and your approach is so amazing!",
    "1258051": "Congrats ! Your investigation on individual classes and radiologists is very strong",
    "1258055": "Congrats for the second place and thanks for the detailed writeup! \n\n\nWe also noticed the systematic differences between annotators  (https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229642)\nI wonder if applying the bigger boxes finetune trick only for images not annotated by R8-R10 would further boost the score.",
    "1258038": "Congratulations for 2nd place and thanks for sharing detailed solution @zfturbo and team",
    "1276625": "Congrats! Great work!",
    "1260921": "Congratz, which repo did you use for **RetinaNet* ? @zfturbo ",
    "1259469": "Congrats, @zfturbo ! Amazing work",
    "1259370": "\"The first thing we tried is to calculate validation mAP based on Radiologist ID.\"\n=> How did you calculate the mAP for each Radiologist ? Did you just remove all the other radiologists' bbox in ground truth and calculate just the remained one ? ",
    "1258035": "@zfturbo Congratulations on 2 nd Prize and Thanks for sharing the approach"
  }
}