{
  "id": 229740,
  "title": "2nd place solution (addition about MMDetection)",
  "url": "/competitions/vinbigdata-chest-xray-abnormalities-detection/discussion/229740",
  "author_name": "Ivan Panshin",
  "post_date": "2021-03-31T13:46:50.311000",
  "votes": 41,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I would like to add a little bit of information to the great <a href=\"url\" target=\"_blank\">topic</a> by <a href=\"https://www.kaggle.com/zfturbo\" target=\"_blank\">@zfturbo</a>. In this one, I want to  talk about MMDetection pipelines that we used. </p>\n<h2>Part 1: Grid Search</h2>\n<p>MMDetection is great in terms of the amount of models/tricks. However, that also brings some difficulties, like what approaches to use. I started running some tests with ATSS, Cascade_RFP, GFL, RetinaNet, UniverseNet. Based on initial validation scores + LB, I decided to settle on the following approaches and tune them more carefully: </p>\n<ul>\n<li>Cascade_RFP_R50 </li>\n<li>GFL_R101</li>\n<li>GFL_X101</li>\n<li>RetinaNet_X101 </li>\n</ul>\n<h2>Part 2: Training tricks</h2>\n<ul>\n<li>Add albumentations for validation: ShiftScaleRotate, IAAAffine, Blur/GaussianBlur/MedianBlur, RandomBrightnessContrast, IAAAdditiveGaussianNoise/GaussNoise, HorizontalFlip.</li>\n<li>Use 1024x1024 for all models, and tune a couple of them (GFL_R101, GFL_X101) with 2048x2048 in order to catch small bboxes. </li>\n<li>Apply FP16 to all models except Cascade_RFP in order to increase batch size and speed-up the training. In my experience, FP16 currently doesn't work with Cascade_RFP.</li>\n<li>Train with empty images in order to get rid of classifier. But not all empty images (there are too many of them), treat empty just like any other class, and add N random empty images, where N = average amount of samples for all classes</li>\n<li>CosineAnnealing with linear warm-up instead of classic StepLP. </li>\n<li>Tried to use sampler so that the amount of bboxes for each class is more or less the same. Had to patch ClassBalancedDataset (with oversample_thr) in order to do that, but that didn't boost the performance (perhaps due to the fact that it's not about class imbalance, but radiologists imbalance). </li>\n</ul>\n<h2>Part 3: 2-step training</h2>\n<p>Like <a href=\"https://www.kaggle.com/zfturbo\" target=\"_blank\">@zfturbo</a> noticed, it's profitable to train with rare radiologists (all but R8, R9, R10). Btw, my greatest thanks and respects for this idea. I'm not sure I would have been able to come up with that on my own. </p>\n<ul>\n<li>1-step: pre-train with all data for 30 epochs. Select the best checkpoints based on AP@0.4 (wanted to try selecting checkpoints for each class, or using SWA on different checkpoints, but eventually didn't have time for this).</li>\n<li>2-step: take only rare radiologists (again, with empty images. In this case, I added about 50 images. Moreover, I didn't apply WBF to bboxes. Instead, I created 3 versions of each image with the corresponding set of bboxes. This idea belongs to <a href=\"https://www.kaggle.com/sergeyzlobin\" target=\"_blank\">@sergeyzlobin</a>, I appreciate it!). It's also important to handle folds correctly. What I mean by that: data for 2-step is a subset of data for the 1-step. So take rare + empty images from the same folds. In other words, do not introduce leakage. Start training from the best weights from 1-step. I trained with the same augs and optimizer. Moreover, taking the same LR as in the beginning of the 1-step worked the best.</li>\n</ul>\n<h2>Part 4: Models scores (public / private)</h2>\n<ul>\n<li><p>GFL_R101_rare_with_empty_2step_1024 - 0.239 / 0.248</p></li>\n<li><p>GFL_X101_rare_with_empty_2step_1024 - 0.264 / 0.255</p></li>\n<li><p>Cascade_RFP_R50_rare_with_empty_2step_1024 - 0.251 / 0.243</p></li>\n<li><p>Retina_rare_2step_1024 (no empty) - 0.228 / 0.239</p></li>\n<li><p>GFL_R101_rare_with_empty_2step_2048 - 0.248 / 0.235</p></li>\n<li><p>GFL_X101_rare_with_empty_2step_2048 - 0.256 / 0.238</p></li>\n</ul>\n<p>In the last days of the competition I also decided to add VFNet to the mix, so:</p>\n<ul>\n<li>VFNet_X101_rare_with_empty_2step_1024 - 0.259 / 0.252</li>\n</ul>\n<p>All models are 5 folds. </p>\n<h2>Conclusion</h2>\n<p>This competition is one of the most difficult things that I've done. We've worked for a couple of months pretty much non-stop, and it was awesome. Thanks so much to the organizers, and my team <a href=\"https://www.kaggle.com/zfturbo\" target=\"_blank\">@zfturbo</a>, <a href=\"https://www.kaggle.com/sergeyzlobin\" target=\"_blank\">@sergeyzlobin</a>, you're insanely good and I hope to work again on something else. </p>",
  "messages": [
    {
      "id": 1258264,
      "postDate": "2021-03-31T13:46:50.313Z",
      "content": "<p>I would like to add a little bit of information to the great <a href=\"url\" target=\"_blank\">topic</a> by <a href=\"https://www.kaggle.com/zfturbo\" target=\"_blank\">@zfturbo</a>. In this one, I want to  talk about MMDetection pipelines that we used. </p>\n<h2>Part 1: Grid Search</h2>\n<p>MMDetection is great in terms of the amount of models/tricks. However, that also brings some difficulties, like what approaches to use. I started running some tests with ATSS, Cascade_RFP, GFL, RetinaNet, UniverseNet. Based on initial validation scores + LB, I decided to settle on the following approaches and tune them more carefully: </p>\n<ul>\n<li>Cascade_RFP_R50 </li>\n<li>GFL_R101</li>\n<li>GFL_X101</li>\n<li>RetinaNet_X101 </li>\n</ul>\n<h2>Part 2: Training tricks</h2>\n<ul>\n<li>Add albumentations for validation: ShiftScaleRotate, IAAAffine, Blur/GaussianBlur/MedianBlur, RandomBrightnessContrast, IAAAdditiveGaussianNoise/GaussNoise, HorizontalFlip.</li>\n<li>Use 1024x1024 for all models, and tune a couple of them (GFL_R101, GFL_X101) with 2048x2048 in order to catch small bboxes. </li>\n<li>Apply FP16 to all models except Cascade_RFP in order to increase batch size and speed-up the training. In my experience, FP16 currently doesn't work with Cascade_RFP.</li>\n<li>Train with empty images in order to get rid of classifier. But not all empty images (there are too many of them), treat empty just like any other class, and add N random empty images, where N = average amount of samples for all classes</li>\n<li>CosineAnnealing with linear warm-up instead of classic StepLP. </li>\n<li>Tried to use sampler so that the amount of bboxes for each class is more or less the same. Had to patch ClassBalancedDataset (with oversample_thr) in order to do that, but that didn't boost the performance (perhaps due to the fact that it's not about class imbalance, but radiologists imbalance). </li>\n</ul>\n<h2>Part 3: 2-step training</h2>\n<p>Like <a href=\"https://www.kaggle.com/zfturbo\" target=\"_blank\">@zfturbo</a> noticed, it's profitable to train with rare radiologists (all but R8, R9, R10). Btw, my greatest thanks and respects for this idea. I'm not sure I would have been able to come up with that on my own. </p>\n<ul>\n<li>1-step: pre-train with all data for 30 epochs. Select the best checkpoints based on AP@0.4 (wanted to try selecting checkpoints for each class, or using SWA on different checkpoints, but eventually didn't have time for this).</li>\n<li>2-step: take only rare radiologists (again, with empty images. In this case, I added about 50 images. Moreover, I didn't apply WBF to bboxes. Instead, I created 3 versions of each image with the corresponding set of bboxes. This idea belongs to <a href=\"https://www.kaggle.com/sergeyzlobin\" target=\"_blank\">@sergeyzlobin</a>, I appreciate it!). It's also important to handle folds correctly. What I mean by that: data for 2-step is a subset of data for the 1-step. So take rare + empty images from the same folds. In other words, do not introduce leakage. Start training from the best weights from 1-step. I trained with the same augs and optimizer. Moreover, taking the same LR as in the beginning of the 1-step worked the best.</li>\n</ul>\n<h2>Part 4: Models scores (public / private)</h2>\n<ul>\n<li><p>GFL_R101_rare_with_empty_2step_1024 - 0.239 / 0.248</p></li>\n<li><p>GFL_X101_rare_with_empty_2step_1024 - 0.264 / 0.255</p></li>\n<li><p>Cascade_RFP_R50_rare_with_empty_2step_1024 - 0.251 / 0.243</p></li>\n<li><p>Retina_rare_2step_1024 (no empty) - 0.228 / 0.239</p></li>\n<li><p>GFL_R101_rare_with_empty_2step_2048 - 0.248 / 0.235</p></li>\n<li><p>GFL_X101_rare_with_empty_2step_2048 - 0.256 / 0.238</p></li>\n</ul>\n<p>In the last days of the competition I also decided to add VFNet to the mix, so:</p>\n<ul>\n<li>VFNet_X101_rare_with_empty_2step_1024 - 0.259 / 0.252</li>\n</ul>\n<p>All models are 5 folds. </p>\n<h2>Conclusion</h2>\n<p>This competition is one of the most difficult things that I've done. We've worked for a couple of months pretty much non-stop, and it was awesome. Thanks so much to the organizers, and my team <a href=\"https://www.kaggle.com/zfturbo\" target=\"_blank\">@zfturbo</a>, <a href=\"https://www.kaggle.com/sergeyzlobin\" target=\"_blank\">@sergeyzlobin</a>, you're insanely good and I hope to work again on something else. </p>",
      "rawMarkdown": "I would like to add a little bit of information to the great [topic](url) by @zfturbo. In this one, I want to  talk about MMDetection pipelines that we used. \n\n## Part 1: Grid Search\n\nMMDetection is great in terms of the amount of models/tricks. However, that also brings some difficulties, like what approaches to use. I started running some tests with ATSS, Cascade_RFP, GFL, RetinaNet, UniverseNet. Based on initial validation scores + LB, I decided to settle on the following approaches and tune them more carefully: \n\n- Cascade_RFP_R50 \n- GFL_R101\n- GFL_X101\n- RetinaNet_X101 \n\n## Part 2: Training tricks\n\n- Add albumentations for validation: ShiftScaleRotate, IAAAffine, Blur/GaussianBlur/MedianBlur, RandomBrightnessContrast, IAAAdditiveGaussianNoise/GaussNoise, HorizontalFlip.\n- Use 1024x1024 for all models, and tune a couple of them (GFL_R101, GFL_X101) with 2048x2048 in order to catch small bboxes. \n- Apply FP16 to all models except Cascade_RFP in order to increase batch size and speed-up the training. In my experience, FP16 currently doesn't work with Cascade_RFP.\n- Train with empty images in order to get rid of classifier. But not all empty images (there are too many of them), treat empty just like any other class, and add N random empty images, where N = average amount of samples for all classes\n- CosineAnnealing with linear warm-up instead of classic StepLP. \n- Tried to use sampler so that the amount of bboxes for each class is more or less the same. Had to patch ClassBalancedDataset (with oversample_thr) in order to do that, but that didn't boost the performance (perhaps due to the fact that it's not about class imbalance, but radiologists imbalance). \n\n## Part 3: 2-step training\n\nLike @zfturbo noticed, it's profitable to train with rare radiologists (all but R8, R9, R10). Btw, my greatest thanks and respects for this idea. I'm not sure I would have been able to come up with that on my own. \n\n- 1-step: pre-train with all data for 30 epochs. Select the best checkpoints based on AP@0.4 (wanted to try selecting checkpoints for each class, or using SWA on different checkpoints, but eventually didn't have time for this).\n- 2-step: take only rare radiologists (again, with empty images. In this case, I added about 50 images. Moreover, I didn't apply WBF to bboxes. Instead, I created 3 versions of each image with the corresponding set of bboxes. This idea belongs to @sergeyzlobin, I appreciate it!). It's also important to handle folds correctly. What I mean by that: data for 2-step is a subset of data for the 1-step. So take rare + empty images from the same folds. In other words, do not introduce leakage. Start training from the best weights from 1-step. I trained with the same augs and optimizer. Moreover, taking the same LR as in the beginning of the 1-step worked the best.\n\n## Part 4: Models scores (public / private)\n\n- GFL_R101_rare_with_empty_2step_1024 - 0.239 / 0.248\n- GFL_X101_rare_with_empty_2step_1024 - 0.264 / 0.255\n- Cascade_RFP_R50_rare_with_empty_2step_1024 - 0.251 / 0.243\n- Retina_rare_2step_1024 (no empty) - 0.228 / 0.239\n\n- GFL_R101_rare_with_empty_2step_2048 - 0.248 / 0.235\n- GFL_X101_rare_with_empty_2step_2048 - 0.256 / 0.238\n\nIn the last days of the competition I also decided to add VFNet to the mix, so:\n- VFNet_X101_rare_with_empty_2step_1024 - 0.259 / 0.252\n\nAll models are 5 folds. \n\n## Conclusion  \n\nThis competition is one of the most difficult things that I've done. We've worked for a couple of months pretty much non-stop, and it was awesome. Thanks so much to the organizers, and my team @zfturbo, @sergeyzlobin, you're insanely good and I hope to work again on something else. \n\n\n",
      "votes": 41
    },
    {
      "id": 1270209,
      "postDate": "2021-04-11T12:29:30.913Z",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!",
      "votes": 1
    },
    {
      "id": 1265590,
      "postDate": "2021-04-07T03:53:24.310Z",
      "content": "<p>hi, congrats, and thanks for your sharing!! I'm a little confuse with your part 3: the sentence \"I didn't apply WBF to bboxes. Instead, I created 3 versions of each image with the corresponding set of bboxes. \"If don't mind, could you explain more about how you created 3 version of each image, such as in the image, there are class2 with 3 boxes and class7 with one box, How did you separate these boxes into 3 version. Congrats again</p>",
      "rawMarkdown": "hi, congrats, and thanks for your sharing!! I'm a little confuse with your part 3: the sentence \"I didn't apply WBF to bboxes. Instead, I created 3 versions of each image with the corresponding set of bboxes. \"If don't mind, could you explain more about how you created 3 version of each image, such as in the image, there are class2 with 3 boxes and class7 with one box, How did you separate these boxes into 3 version. Congrats again",
      "votes": 1,
      "replies": [
        {
          "id": 1265877,
          "postDate": "2021-04-07T09:14:45.233Z",
          "content": "<p>Thanks! It actually depends on the radiologists. Each image was annotated by 3 radiologists (in train). So I created the 3 versions of the same image as following: image i + annotations only from 1st RAD; image i + annotations only from 2nd RAD; image i + annotations only form 3rd RAD. </p>",
          "rawMarkdown": "Thanks! It actually depends on the radiologists. Each image was annotated by 3 radiologists (in train). So I created the 3 versions of the same image as following: image i + annotations only from 1st RAD; image i + annotations only from 2nd RAD; image i + annotations only form 3rd RAD. ",
          "votes": 1
        },
        {
          "id": 1266724,
          "postDate": "2021-04-08T02:57:28.893Z",
          "content": "<p>thanks for your replay!</p>",
          "rawMarkdown": "thanks for your replay!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1260382,
      "postDate": "2021-04-02T04:24:28.967Z",
      "content": "<p>Congrats on the second place, the new competitions master <a href=\"https://www.kaggle.com/ivanpan\" target=\"_blank\">@ivanpan</a> </p>",
      "rawMarkdown": "Congrats on the second place, the new competitions master @ivanpan ",
      "votes": 1
    },
    {
      "id": 1258688,
      "postDate": "2021-03-31T20:05:57.717Z",
      "content": "<p><a href=\"https://www.kaggle.com/ivanpan\" target=\"_blank\">@ivanpan</a> Thank you for sharing. Really great insight. Congrats!</p>",
      "rawMarkdown": "@ivanpan Thank you for sharing. Really great insight. Congrats!",
      "votes": 1
    },
    {
      "id": 1258342,
      "postDate": "2021-03-31T14:44:28.530Z",
      "content": "<p>Thanks for sharing some information from your team solution, congrats on 2nd place <a href=\"https://www.kaggle.com/ivanpan\" target=\"_blank\">@ivanpan</a> </p>",
      "rawMarkdown": "Thanks for sharing some information from your team solution, congrats on 2nd place @ivanpan ",
      "votes": 2
    },
    {
      "id": 1258485,
      "postDate": "2021-03-31T16:52:37.700Z",
      "content": "<p><a href=\"https://www.kaggle.com/ivanpan\" target=\"_blank\">@ivanpan</a> Congratulations on second place Finish and Thanks for sharing the approach</p>",
      "rawMarkdown": "@ivanpan Congratulations on second place Finish and Thanks for sharing the approach"
    }
  ],
  "comments": [
    {
      "id": 1270209,
      "author_name": "Wonjun Park",
      "author_url": "",
      "post_date": "2021-04-11T12:29:30.913000",
      "content": "<p>Congratulations!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1265590,
      "author_name": "Betty",
      "author_url": "",
      "post_date": "2021-04-07T03:53:24.310000",
      "content": "<p>hi, congrats, and thanks for your sharing!! I'm a little confuse with your part 3: the sentence \"I didn't apply WBF to bboxes. Instead, I created 3 versions of each image with the corresponding set of bboxes. \"If don't mind, could you explain more about how you created 3 version of each image, such as in the image, there are class2 with 3 boxes and class7 with one box, How did you separate these boxes into 3 version. Congrats again</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1265877,
          "author_name": "Ivan Panshin",
          "author_url": "",
          "post_date": "2021-04-07T09:14:45.233000",
          "content": "<p>Thanks! It actually depends on the radiologists. Each image was annotated by 3 radiologists (in train). So I created the 3 versions of the same image as following: image i + annotations only from 1st RAD; image i + annotations only from 2nd RAD; image i + annotations only form 3rd RAD. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1266724,
          "author_name": "Betty",
          "author_url": "",
          "post_date": "2021-04-08T02:57:28.893000",
          "content": "<p>thanks for your replay!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1260382,
      "author_name": "( ͡° ͜ʖ ͡°)",
      "author_url": "",
      "post_date": "2021-04-02T04:24:28.967000",
      "content": "<p>Congrats on the second place, the new competitions master <a href=\"https://www.kaggle.com/ivanpan\" target=\"_blank\">@ivanpan</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1258688,
      "author_name": "Charlie Craine",
      "author_url": "",
      "post_date": "2021-03-31T20:05:57.717000",
      "content": "<p><a href=\"https://www.kaggle.com/ivanpan\" target=\"_blank\">@ivanpan</a> Thank you for sharing. Really great insight. Congrats!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1258342,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2021-03-31T14:44:28.530000",
      "content": "<p>Thanks for sharing some information from your team solution, congrats on 2nd place <a href=\"https://www.kaggle.com/ivanpan\" target=\"_blank\">@ivanpan</a> </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1258485,
      "author_name": "Tensor Girl",
      "author_url": "",
      "post_date": "2021-03-31T16:52:37.700000",
      "content": "<p><a href=\"https://www.kaggle.com/ivanpan\" target=\"_blank\">@ivanpan</a> Congratulations on second place Finish and Thanks for sharing the approach</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1258264": "I would like to add a little bit of information to the great [topic](url) by @zfturbo. In this one, I want to  talk about MMDetection pipelines that we used. \n\n## Part 1: Grid Search\n\nMMDetection is great in terms of the amount of models/tricks. However, that also brings some difficulties, like what approaches to use. I started running some tests with ATSS, Cascade_RFP, GFL, RetinaNet, UniverseNet. Based on initial validation scores + LB, I decided to settle on the following approaches and tune them more carefully: \n\n- Cascade_RFP_R50 \n- GFL_R101\n- GFL_X101\n- RetinaNet_X101 \n\n## Part 2: Training tricks\n\n- Add albumentations for validation: ShiftScaleRotate, IAAAffine, Blur/GaussianBlur/MedianBlur, RandomBrightnessContrast, IAAAdditiveGaussianNoise/GaussNoise, HorizontalFlip.\n- Use 1024x1024 for all models, and tune a couple of them (GFL_R101, GFL_X101) with 2048x2048 in order to catch small bboxes. \n- Apply FP16 to all models except Cascade_RFP in order to increase batch size and speed-up the training. In my experience, FP16 currently doesn't work with Cascade_RFP.\n- Train with empty images in order to get rid of classifier. But not all empty images (there are too many of them), treat empty just like any other class, and add N random empty images, where N = average amount of samples for all classes\n- CosineAnnealing with linear warm-up instead of classic StepLP. \n- Tried to use sampler so that the amount of bboxes for each class is more or less the same. Had to patch ClassBalancedDataset (with oversample_thr) in order to do that, but that didn't boost the performance (perhaps due to the fact that it's not about class imbalance, but radiologists imbalance). \n\n## Part 3: 2-step training\n\nLike @zfturbo noticed, it's profitable to train with rare radiologists (all but R8, R9, R10). Btw, my greatest thanks and respects for this idea. I'm not sure I would have been able to come up with that on my own. \n\n- 1-step: pre-train with all data for 30 epochs. Select the best checkpoints based on AP@0.4 (wanted to try selecting checkpoints for each class, or using SWA on different checkpoints, but eventually didn't have time for this).\n- 2-step: take only rare radiologists (again, with empty images. In this case, I added about 50 images. Moreover, I didn't apply WBF to bboxes. Instead, I created 3 versions of each image with the corresponding set of bboxes. This idea belongs to @sergeyzlobin, I appreciate it!). It's also important to handle folds correctly. What I mean by that: data for 2-step is a subset of data for the 1-step. So take rare + empty images from the same folds. In other words, do not introduce leakage. Start training from the best weights from 1-step. I trained with the same augs and optimizer. Moreover, taking the same LR as in the beginning of the 1-step worked the best.\n\n## Part 4: Models scores (public / private)\n\n- GFL_R101_rare_with_empty_2step_1024 - 0.239 / 0.248\n- GFL_X101_rare_with_empty_2step_1024 - 0.264 / 0.255\n- Cascade_RFP_R50_rare_with_empty_2step_1024 - 0.251 / 0.243\n- Retina_rare_2step_1024 (no empty) - 0.228 / 0.239\n\n- GFL_R101_rare_with_empty_2step_2048 - 0.248 / 0.235\n- GFL_X101_rare_with_empty_2step_2048 - 0.256 / 0.238\n\nIn the last days of the competition I also decided to add VFNet to the mix, so:\n- VFNet_X101_rare_with_empty_2step_1024 - 0.259 / 0.252\n\nAll models are 5 folds. \n\n## Conclusion  \n\nThis competition is one of the most difficult things that I've done. We've worked for a couple of months pretty much non-stop, and it was awesome. Thanks so much to the organizers, and my team @zfturbo, @sergeyzlobin, you're insanely good and I hope to work again on something else. \n\n\n",
    "1270209": "Congratulations!",
    "1265590": "hi, congrats, and thanks for your sharing!! I'm a little confuse with your part 3: the sentence \"I didn't apply WBF to bboxes. Instead, I created 3 versions of each image with the corresponding set of bboxes. \"If don't mind, could you explain more about how you created 3 version of each image, such as in the image, there are class2 with 3 boxes and class7 with one box, How did you separate these boxes into 3 version. Congrats again",
    "1260382": "Congrats on the second place, the new competitions master @ivanpan ",
    "1258688": "@ivanpan Thank you for sharing. Really great insight. Congrats!",
    "1258342": "Thanks for sharing some information from your team solution, congrats on 2nd place @ivanpan ",
    "1258485": "@ivanpan Congratulations on second place Finish and Thanks for sharing the approach"
  }
}