{
  "id": 229711,
  "title": "5th Place Solution",
  "url": "/competitions/vinbigdata-chest-xray-abnormalities-detection/discussion/229711",
  "author_name": "Ahmet Erdem",
  "post_date": "2021-03-31T11:03:47.617000",
  "votes": 50,
  "comment_count": 33,
  "views": 0,
  "content": "<p>In this competition, I am happy that I have teamed up with one of the most dangerous gangs of Kaggle Street, namely <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a>, <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a>, <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a>. </p>\n<p><strong>Validation and Data Bias</strong><br>\nIt was very difficult to make reliable Cross-Validation since the test set was generated by different labelling process and Public LB was not trustworthy due to its small size. We first noticed that our LB score is much lower than our CV score. It makes sense to some degree but the gap was really big, therefore we have investigated the data. We have trained an adversarial validation model on train vs test images. It has 0.61 AUC. The model was almost certainly able to distinguish some train set images from the test set but the rest looks like random split. These distinguishable images were all labelled by R8-R9-R10 with findings. So the different data sources were not evenly distributed among radiologists. This finding didn't really help us since we don't have the luxury to remove data with labellings. Therefore we have tried to remove the bias. They all had black corners covering some information. It made sense to add the same black boxes to the corners of the each image but it didn't remove the bias. Then we have tried gradCAM and according to it the bias was everywhere in the image, making it hard to remove.</p>\n<p><img src=\"https://i.imgur.com/mY34STX.png\" alt=\"gradCAM Result\"></p>\n<p>We have trained another model to detect which radiologists labelled the images. This model was very accurate on R1-7 vs the rest. Since R1-7 didn't have any labels, we have also removed them from our training set hoping that it can solve some of the bias issue. Knowing that bias is everywhere in the image, and not on a specific location, we thought that style transfer may solve the issue and ended up in this art.</p>\n<p><img src=\"https://i.imgur.com/Gqk1RBq.jpg\" alt=\"Chest Art\"></p>\n<p>This didn't work (obviously :D). </p>\n<p><strong>Rare Annotators</strong><br>\nLosing hope on removing more bias, we have focused on improving our CV as it is with more focus on R11-17. We have trained several effdets with different backbones and standard augmentations. Data is sampled for each radiologist if the image has a finding, otherwise the image is sampled once during each epoch. We also had a YOLO model performing worse than our effdet. Ensembling them all with WBF@0.4 gave us our best score. Some experiments  with pretraining on R8-10, finetuning on R11-17 and applying WBF before training by weighting R11-17 more worked not good on Public LB, therefore we didn't use them. But they seem to work much better on Private LB like 0.224 -&gt; 0.275.</p>\n<p><strong>Aortic Enlargement</strong><br>\nWe noticed that something was wrong with this class because our LB score for it was way less than our CV score. Considering that it may be the way boxes are labelled in the test set, we have made a post processing for them which gave us an extra 0.01 on Public, 0.005 on Private.</p>\n<pre><code>sub.loc[sub.class_id==0, \"y_max\"] += 110\nsub.loc[sub.class_id==0, \"x_min\"] += 40\n</code></pre>",
  "messages": [
    {
      "id": 1258107,
      "postDate": "2021-03-31T11:03:47.617Z",
      "content": "<p>In this competition, I am happy that I have teamed up with one of the most dangerous gangs of Kaggle Street, namely <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a>, <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a>, <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a>. </p>\n<p><strong>Validation and Data Bias</strong><br>\nIt was very difficult to make reliable Cross-Validation since the test set was generated by different labelling process and Public LB was not trustworthy due to its small size. We first noticed that our LB score is much lower than our CV score. It makes sense to some degree but the gap was really big, therefore we have investigated the data. We have trained an adversarial validation model on train vs test images. It has 0.61 AUC. The model was almost certainly able to distinguish some train set images from the test set but the rest looks like random split. These distinguishable images were all labelled by R8-R9-R10 with findings. So the different data sources were not evenly distributed among radiologists. This finding didn't really help us since we don't have the luxury to remove data with labellings. Therefore we have tried to remove the bias. They all had black corners covering some information. It made sense to add the same black boxes to the corners of the each image but it didn't remove the bias. Then we have tried gradCAM and according to it the bias was everywhere in the image, making it hard to remove.</p>\n<p><img src=\"https://i.imgur.com/mY34STX.png\" alt=\"gradCAM Result\"></p>\n<p>We have trained another model to detect which radiologists labelled the images. This model was very accurate on R1-7 vs the rest. Since R1-7 didn't have any labels, we have also removed them from our training set hoping that it can solve some of the bias issue. Knowing that bias is everywhere in the image, and not on a specific location, we thought that style transfer may solve the issue and ended up in this art.</p>\n<p><img src=\"https://i.imgur.com/Gqk1RBq.jpg\" alt=\"Chest Art\"></p>\n<p>This didn't work (obviously :D). </p>\n<p><strong>Rare Annotators</strong><br>\nLosing hope on removing more bias, we have focused on improving our CV as it is with more focus on R11-17. We have trained several effdets with different backbones and standard augmentations. Data is sampled for each radiologist if the image has a finding, otherwise the image is sampled once during each epoch. We also had a YOLO model performing worse than our effdet. Ensembling them all with WBF@0.4 gave us our best score. Some experiments  with pretraining on R8-10, finetuning on R11-17 and applying WBF before training by weighting R11-17 more worked not good on Public LB, therefore we didn't use them. But they seem to work much better on Private LB like 0.224 -&gt; 0.275.</p>\n<p><strong>Aortic Enlargement</strong><br>\nWe noticed that something was wrong with this class because our LB score for it was way less than our CV score. Considering that it may be the way boxes are labelled in the test set, we have made a post processing for them which gave us an extra 0.01 on Public, 0.005 on Private.</p>\n<pre><code>sub.loc[sub.class_id==0, \"y_max\"] += 110\nsub.loc[sub.class_id==0, \"x_min\"] += 40\n</code></pre>",
      "rawMarkdown": "In this competition, I am happy that I have teamed up with one of the most dangerous gangs of Kaggle Street, namely @christofhenkel, @ilu000, @philippsinger. \n\n**Validation and Data Bias**\nIt was very difficult to make reliable Cross-Validation since the test set was generated by different labelling process and Public LB was not trustworthy due to its small size. We first noticed that our LB score is much lower than our CV score. It makes sense to some degree but the gap was really big, therefore we have investigated the data. We have trained an adversarial validation model on train vs test images. It has 0.61 AUC. The model was almost certainly able to distinguish some train set images from the test set but the rest looks like random split. These distinguishable images were all labelled by R8-R9-R10 with findings. So the different data sources were not evenly distributed among radiologists. This finding didn't really help us since we don't have the luxury to remove data with labellings. Therefore we have tried to remove the bias. They all had black corners covering some information. It made sense to add the same black boxes to the corners of the each image but it didn't remove the bias. Then we have tried gradCAM and according to it the bias was everywhere in the image, making it hard to remove.\n\n![gradCAM Result](https://i.imgur.com/mY34STX.png)\n\nWe have trained another model to detect which radiologists labelled the images. This model was very accurate on R1-7 vs the rest. Since R1-7 didn't have any labels, we have also removed them from our training set hoping that it can solve some of the bias issue. Knowing that bias is everywhere in the image, and not on a specific location, we thought that style transfer may solve the issue and ended up in this art.\n\n![Chest Art](https://i.imgur.com/Gqk1RBq.jpg)\n\nThis didn't work (obviously :D). \n\n**Rare Annotators**\nLosing hope on removing more bias, we have focused on improving our CV as it is with more focus on R11-17. We have trained several effdets with different backbones and standard augmentations. Data is sampled for each radiologist if the image has a finding, otherwise the image is sampled once during each epoch. We also had a YOLO model performing worse than our effdet. Ensembling them all with WBF@0.4 gave us our best score. Some experiments  with pretraining on R8-10, finetuning on R11-17 and applying WBF before training by weighting R11-17 more worked not good on Public LB, therefore we didn't use them. But they seem to work much better on Private LB like 0.224 -> 0.275.\n\n**Aortic Enlargement**\nWe noticed that something was wrong with this class because our LB score for it was way less than our CV score. Considering that it may be the way boxes are labelled in the test set, we have made a post processing for them which gave us an extra 0.01 on Public, 0.005 on Private.\n```\nsub.loc[sub.class_id==0, \"y_max\"] += 110\nsub.loc[sub.class_id==0, \"x_min\"] += 40\n```",
      "votes": 50
    },
    {
      "id": 1258142,
      "postDate": "2021-03-31T11:37:11.220Z",
      "content": "<p>Thanks for the awesome and professional team-up again <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> <a href=\"https://www.kaggle.com/aerdem4\" target=\"_blank\">@aerdem4</a> <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> !<br>\nThat style transfer experiment was definetly a highlight.</p>",
      "rawMarkdown": "Thanks for the awesome and professional team-up again @christofhenkel @aerdem4 @philippsinger !\nThat style transfer experiment was definetly a highlight.",
      "votes": 5
    },
    {
      "id": 1258118,
      "postDate": "2021-03-31T11:14:37.537Z",
      "content": "<p>Cool style transfer art!</p>",
      "rawMarkdown": "Cool style transfer art!",
      "votes": 3
    },
    {
      "id": 1259169,
      "postDate": "2021-04-01T07:30:20.697Z",
      "content": "<p>Congratulations. Winning in a great jump!</p>",
      "rawMarkdown": "Congratulations. Winning in a great jump!",
      "votes": 1
    },
    {
      "id": 1258531,
      "postDate": "2021-03-31T17:35:10.920Z",
      "content": "<p>Very interesting write-up (and approach but it goes without saying), varies a lot from the classical blending description, thanks a lot for that !<br>\nAnd congratz on the impressive result but I guess that's no big surprise after all!</p>",
      "rawMarkdown": "Very interesting write-up (and approach but it goes without saying), varies a lot from the classical blending description, thanks a lot for that !\nAnd congratz on the impressive result but I guess that's no big surprise after all!",
      "votes": 1
    },
    {
      "id": 1258418,
      "postDate": "2021-03-31T15:35:09.100Z",
      "content": "<p>Thank you for the write up and congrats on your strong finish! I am seeing that working with Radialogist bias potentially had a lot of value…<br>\nEveryone is talking about Aortic Enlargement, but I am curious, how was your \"Other Lesion\" class score compare to others? It was one of the most difficult classes to deal with in my case</p>",
      "rawMarkdown": "Thank you for the write up and congrats on your strong finish! I am seeing that working with Radialogist bias potentially had a lot of value...\nEveryone is talking about Aortic Enlargement, but I am curious, how was your \"Other Lesion\" class score compare to others? It was one of the most difficult classes to deal with in my case",
      "votes": 1,
      "replies": [
        {
          "id": 1258437,
          "postDate": "2021-03-31T15:56:22.487Z",
          "content": "<p>We only submit the common classes separately to see if we can postprocess. Postprocessing rare classes based on small LB feedback was dangerous. Therefore I don't know its LB score but maybe my  teammates have an idea.</p>",
          "rawMarkdown": "We only submit the common classes separately to see if we can postprocess. Postprocessing rare classes based on small LB feedback was dangerous. Therefore I don't know its LB score but maybe my  teammates have an idea.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1258331,
      "postDate": "2021-03-31T14:37:58.123Z",
      "content": "<p>Congratulations on an awesome finish and great jump upward from public to private LB! The style gan art is great.</p>\n<p>We also noticed the huge difference between Aortic Enlargement CV and LB. On CV, Aortic Enlargement is the best class with local validation <code>mAP = 0.867</code> (on Yolo single model), but on private LB it only has <code>mAP = 0.210</code>. So test data 5 doctor consensus label is very different than train label. We didn't discover why. </p>\n<p>Great detective work making the bbox larger. That explains some of the difference. Using your trick boost our private Aortic Enlargment to <code>mAP = 0.255</code> (on same Yolo single model).</p>",
      "rawMarkdown": "Congratulations on an awesome finish and great jump upward from public to private LB! The style gan art is great.\n\nWe also noticed the huge difference between Aortic Enlargement CV and LB. On CV, Aortic Enlargement is the best class with local validation `mAP = 0.867` (on Yolo single model), but on private LB it only has `mAP = 0.210`. So test data 5 doctor consensus label is very different than train label. We didn't discover why. \n\nGreat detective work making the bbox larger. That explains some of the difference. Using your trick boost our private Aortic Enlargment to `mAP = 0.255` (on same Yolo single model).",
      "votes": 1,
      "replies": [
        {
          "id": 1258337,
          "postDate": "2021-03-31T14:40:22.320Z",
          "content": "<p>Perhaps it isn't different bbox, but maybe the consensus decided many of the Aortic Enlargement are not Aortic Enlargement. So maybe finding some rule to decrease certain image confidence scores for Aortic Enlargement would boost that class.</p>\n<p>Now that we can make unlimited submissions, i'm tempted to solve this mystery </p>",
          "rawMarkdown": "Perhaps it isn't different bbox, but maybe the consensus decided many of the Aortic Enlargement are not Aortic Enlargement. So maybe finding some rule to decrease certain image confidence scores for Aortic Enlargement would boost that class.\n\nNow that we can make unlimited submissions, i'm tempted to solve this mystery "
        },
        {
          "id": 1258346,
          "postDate": "2021-03-31T14:47:32.117Z",
          "content": "<p>2nd place posted lots of Aortic Enlargment investigation information here:<br>\n<a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229696\" target=\"_blank\">https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229696</a></p>",
          "rawMarkdown": "2nd place posted lots of Aortic Enlargment investigation information here:\nhttps://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229696"
        },
        {
          "id": 1258361,
          "postDate": "2021-03-31T14:51:14.433Z",
          "content": "<blockquote>\n  <p>Now that we can make unlimited submissions, i'm tempted to solve this mystery</p>\n</blockquote>\n<p>Now that the competition is finished, I will delete everything including Python and won't look back:)</p>",
          "rawMarkdown": "> Now that we can make unlimited submissions, i'm tempted to solve this mystery\n\nNow that the competition is finished, I will delete everything including Python and won't look back:)",
          "votes": 7
        },
        {
          "id": 1258364,
          "postDate": "2021-03-31T14:53:40.620Z",
          "content": "<p>We noticed that not ony the boxes are smaller but the highlighted aortas too. Half of the R8-R10 annotations had 20-40mm box width (<code>box_width_in_pixel x pixel_spacing</code>) while most definitions we found started aortic enlargement from 45-55mm. We tried to filter the smaller annotations but that experiment did not improve our LB score…</p>\n<p>We tried to use external data too but CRX14 and Chexpert does not report this class.</p>",
          "rawMarkdown": "We noticed that not ony the boxes are smaller but the highlighted aortas too. Half of the R8-R10 annotations had 20-40mm box width (`box_width_in_pixel x pixel_spacing`) while most definitions we found started aortic enlargement from 45-55mm. We tried to filter the smaller annotations but that experiment did not improve our LB score...\n\nWe tried to use external data too but CRX14 and Chexpert does not report this class."
        },
        {
          "id": 1258373,
          "postDate": "2021-03-31T14:58:15.883Z",
          "content": "<p>From my investigations it seems that annotator 12 seems most similar to at least parts of the test class0 labels. Maybe he is a specialist? Maybe there are even different specialists for different diseases?</p>",
          "rawMarkdown": "From my investigations it seems that annotator 12 seems most similar to at least parts of the test class0 labels. Maybe he is a specialist? Maybe there are even different specialists for different diseases?",
          "votes": 1
        },
        {
          "id": 1258422,
          "postDate": "2021-03-31T15:38:32.583Z",
          "content": "<p><a href=\"https://www.kaggle.com/aerdem4\" target=\"_blank\">@aerdem4</a> now that you potentially deleted python and everything associated, you can do the next competition in C++…or just create beautiful art work ;)</p>",
          "rawMarkdown": "@aerdem4 now that you potentially deleted python and everything associated, you can do the next competition in C++...or just create beautiful art work ;)",
          "votes": 2
        },
        {
          "id": 1258796,
          "postDate": "2021-03-31T22:24:49.670Z",
          "content": "<p>I would like to see <a href=\"https://www.kaggle.com/aerdem4\" target=\"_blank\">@aerdem4</a> competes in raw C++.<br>\nThat would be the next level ;)</p>\n<p>(But I'm afraid after he finally wins 1 competition that way, he will not only delete c++, but will also throw out his whole computer too :D and then go to the mountains and live there as a monk xD)</p>",
          "rawMarkdown": "I would like to see @aerdem4 competes in raw C++.\nThat would be the next level ;)\n\n(But I'm afraid after he finally wins 1 competition that way, he will not only delete c++, but will also throw out his whole computer too :D and then go to the mountains and live there as a monk xD)",
          "votes": 2
        }
      ]
    },
    {
      "id": 1258308,
      "postDate": "2021-03-31T14:19:20.600Z",
      "content": "<p>One of the most dangerous gangs of Kaggle Street :D ^^</p>\n<p>Bonus results:<br>\nAfter this comp we have a lung poem + a style transfer art picture xD</p>",
      "rawMarkdown": "One of the most dangerous gangs of Kaggle Street :D ^^\n\nBonus results:\nAfter this comp we have a lung poem + a style transfer art picture xD",
      "votes": 1,
      "replies": [
        {
          "id": 1258356,
          "postDate": "2021-03-31T14:49:56.383Z",
          "content": "<p>I would never expect a medical image competition to produce those great art pieces. Maybe Kaggle needs Artist ranking too:)</p>",
          "rawMarkdown": "I would never expect a medical image competition to produce those great art pieces. Maybe Kaggle needs Artist ranking too:)",
          "votes": 1
        },
        {
          "id": 1258499,
          "postDate": "2021-03-31T17:03:01.527Z",
          "content": "<p>yeah, totally ^^</p>",
          "rawMarkdown": "yeah, totally ^^"
        }
      ]
    },
    {
      "id": 1258158,
      "postDate": "2021-03-31T11:54:13Z",
      "content": "<p>Congrats Ahmet and to the rest of the gangs! I also tried some postprocessing with box sizes but failed…</p>",
      "rawMarkdown": "Congrats Ahmet and to the rest of the gangs! I also tried some postprocessing with box sizes but failed...",
      "votes": 1
    },
    {
      "id": 1258147,
      "postDate": "2021-03-31T11:44:17.660Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/aerdem4\" target=\"_blank\">@aerdem4</a> <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> and <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> on 5th place. Nice jumping!</p>",
      "rawMarkdown": "Congrats @aerdem4 @ilu000 @christofhenkel and @philippsinger on 5th place. Nice jumping!",
      "votes": 2
    },
    {
      "id": 1320952,
      "postDate": "2021-05-24T12:45:29.053Z",
      "content": "<p>Can you share with us the GitHub link from where you have implemented the Effdet code?</p>",
      "rawMarkdown": "Can you share with us the GitHub link from where you have implemented the Effdet code?",
      "replies": [
        {
          "id": 1321468,
          "postDate": "2021-05-24T17:39:53.353Z",
          "content": "<p>sure: <a href=\"https://github.com/rwightman/pytorch-image-models\" target=\"_blank\">https://github.com/rwightman/pytorch-image-models</a><br>\n(Edit: sorry, this didn't contain the effdet, he had it in an extra repo, see my comment below)</p>",
          "rawMarkdown": "sure: https://github.com/rwightman/pytorch-image-models\n(Edit: sorry, this didn't contain the effdet, he had it in an extra repo, see my comment below)"
        },
        {
          "id": 1321640,
          "postDate": "2021-05-24T20:16:52.173Z",
          "content": "<p>wow, rwightman had detection models too! didn't know that, thanks</p>",
          "rawMarkdown": "wow, rwightman had detection models too! didn't know that, thanks"
        },
        {
          "id": 1321671,
          "postDate": "2021-05-24T20:54:11.720Z",
          "content": "<p>Yeah, sorry above link wasn't exactly correct.</p>\n<p><a href=\"https://github.com/rwightman/efficientdet-pytorch\" target=\"_blank\">https://github.com/rwightman/efficientdet-pytorch</a></p>\n<p>Here is his effdet. It's a standalone repo.</p>",
          "rawMarkdown": "Yeah, sorry above link wasn't exactly correct.\n\nhttps://github.com/rwightman/efficientdet-pytorch\n\nHere is his effdet. It's a standalone repo."
        }
      ]
    },
    {
      "id": 1261883,
      "postDate": "2021-04-03T14:41:23.650Z",
      "content": "<p>Very cool.<br>\nI have just made it a reference for me</p>",
      "rawMarkdown": "Very cool.\nI have just made it a reference for me"
    },
    {
      "id": 1261881,
      "postDate": "2021-04-03T14:40:45.747Z",
      "content": "<p>Very cool.<br>\nI have just made it a reference for me</p>",
      "rawMarkdown": "Very cool.\nI have just made it a reference for me"
    },
    {
      "id": 1259712,
      "postDate": "2021-04-01T15:47:54.977Z",
      "content": "<p>From another discussion, i read that you trained your object detection on all images. (And you did not train a classifier model). Did you submit your object detection model's bbox confidence scores as is (for classes 0 thru 13) or did you post process them?</p>",
      "rawMarkdown": "From another discussion, i read that you trained your object detection on all images. (And you did not train a classifier model). Did you submit your object detection model's bbox confidence scores as is (for classes 0 thru 13) or did you post process them?",
      "replies": [
        {
          "id": 1259720,
          "postDate": "2021-04-01T15:55:03.287Z",
          "content": "<p>I'm curious because Ivan asked <a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229637#1259605\" target=\"_blank\">here</a>, if our team's post process is the same as training object detection on all images. If you read my answer, i say it is not. </p>\n<p>Many other teams including 6th place Guanshuo Xu found the formula <code>score = base * prob(finding)**0.2</code> to be optimal for the <code>mAP</code> metric in this comp. Where <code>base</code> are the confidence scores from an object detection model trained on <strong>only</strong> \"finding\" images. And <code>prob(finding)</code> is a classifier.</p>\n<p>I would think that training object detection on all images would be similar to <code>score = base * prob(finding)**1.0</code>.</p>\n<p>So the first formula is equivalent to <code>score = base**0.84 * prob(finding)**0.16</code>. And the second formula is equivalent to <code>score = base**0.5 * prob(finding)**0.5</code>.</p>\n<p>So if you didn't post process, i'm wondering if your confidence scores are similar to <code>score = base**0.5 * prob(finding)**0.5</code> and if so, i'm wondering whether you can increase your CV LB by making it more like <code>score = base**0.84 * prob(finding)**0.16</code></p>",
          "rawMarkdown": "I'm curious because Ivan asked [here][1], if our team's post process is the same as training object detection on all images. If you read my answer, i say it is not. \n\nMany other teams including 6th place Guanshuo Xu found the formula `score = base * prob(finding)**0.2` to be optimal for the `mAP` metric in this comp. Where `base` are the confidence scores from an object detection model trained on **only** \"finding\" images. And `prob(finding)` is a classifier.\n\nI would think that training object detection on all images would be similar to `score = base * prob(finding)**1.0`.\n\nSo the first formula is equivalent to `score = base**0.84 * prob(finding)**0.16`. And the second formula is equivalent to `score = base**0.5 * prob(finding)**0.5`.\n\nSo if you didn't post process, i'm wondering if your confidence scores are similar to `score = base**0.5 * prob(finding)**0.5` and if so, i'm wondering whether you can increase your CV LB by making it more like `score = base**0.84 * prob(finding)**0.16`\n\n[1]: https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229637#1259605"
        },
        {
          "id": 1259839,
          "postDate": "2021-04-01T17:36:35.647Z",
          "content": "<p>We didn't postprocess box confidences. We also didn't use any classifier model. Class 14 predictions are multiplication of (1 - box_conf[i]).</p>",
          "rawMarkdown": "We didn't postprocess box confidences. We also didn't use any classifier model. Class 14 predictions are multiplication of (1 - box_conf[i]).",
          "votes": 2
        },
        {
          "id": 1259865,
          "postDate": "2021-04-01T17:53:21.037Z",
          "content": "<p>Wow great job! </p>\n<p>I'm curious if there is a way to post process your confidence scores to increase <code>mAP</code> CV LB. Did you try any post processing experiments on your bbox confidence scores?</p>\n<p>Like for each image with <code>prob(finding)&gt;0.5</code> then update all bbox (class 0 thru 13) confidence scores for that same image with <code>new_score = old_score ** 1.1</code>. And if that doesn't work, try <code>new score = old_score ** 0.9</code>.</p>\n<p>That formula won't affect scores with <code>prob=1</code> or <code>prob=0</code> but will affect middle probs. And i'm wondering if middle probs of images with finding should be treated differently than middle probs of images without findings.</p>",
          "rawMarkdown": "Wow great job! \n\nI'm curious if there is a way to post process your confidence scores to increase `mAP` CV LB. Did you try any post processing experiments on your bbox confidence scores?\n\nLike for each image with `prob(finding)>0.5` then update all bbox (class 0 thru 13) confidence scores for that same image with `new_score = old_score ** 1.1`. And if that doesn't work, try `new score = old_score ** 0.9`.\n\nThat formula won't affect scores with `prob=1` or `prob=0` but will affect middle probs. And i'm wondering if middle probs of images with finding should be treated differently than middle probs of images without findings.",
          "votes": 1
        },
        {
          "id": 1261984,
          "postDate": "2021-04-03T16:43:25.057Z",
          "content": "<p>Another question: did you try using 1 - mult(box_conf[i]) instead of mult(1-box_conf[i]) as prediction for class 14? </p>",
          "rawMarkdown": "Another question: did you try using 1 - mult(box_conf[i]) instead of mult(1-box_conf[i]) as prediction for class 14? "
        },
        {
          "id": 1261999,
          "postDate": "2021-04-03T17:03:30.470Z",
          "content": "<p>Never mind, my math is wrong </p>",
          "rawMarkdown": "Never mind, my math is wrong "
        }
      ]
    },
    {
      "id": 1258764,
      "postDate": "2021-03-31T21:22:55.043Z",
      "content": "<p>If it was a Kernel only competition; because of very less test images i think u couldn't have done adversial validation; then what have been your approach?</p>",
      "rawMarkdown": "If it was a Kernel only competition; because of very less test images i think u couldn't have done adversial validation; then what have been your approach?",
      "replies": [
        {
          "id": 1259138,
          "postDate": "2021-04-01T07:04:56.187Z",
          "content": "<p>Exactly the same as we did not use this information in the end as also written in the post.</p>",
          "rawMarkdown": "Exactly the same as we did not use this information in the end as also written in the post."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1258142,
      "author_name": "Pascal Pfeiffer",
      "author_url": "",
      "post_date": "2021-03-31T11:37:11.220000",
      "content": "<p>Thanks for the awesome and professional team-up again <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> <a href=\"https://www.kaggle.com/aerdem4\" target=\"_blank\">@aerdem4</a> <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> !<br>\nThat style transfer experiment was definetly a highlight.</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 1258118,
      "author_name": "Nick Sergievskiy",
      "author_url": "",
      "post_date": "2021-03-31T11:14:37.537000",
      "content": "<p>Cool style transfer art!</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1259169,
      "author_name": "Mohammed Rizin V K",
      "author_url": "",
      "post_date": "2021-04-01T07:30:20.697000",
      "content": "<p>Congratulations. Winning in a great jump!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1258531,
      "author_name": "Theo Viel",
      "author_url": "",
      "post_date": "2021-03-31T17:35:10.920000",
      "content": "<p>Very interesting write-up (and approach but it goes without saying), varies a lot from the classical blending description, thanks a lot for that !<br>\nAnd congratz on the impressive result but I guess that's no big surprise after all!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1258418,
      "author_name": "InDSweTrust",
      "author_url": "",
      "post_date": "2021-03-31T15:35:09.100000",
      "content": "<p>Thank you for the write up and congrats on your strong finish! I am seeing that working with Radialogist bias potentially had a lot of value…<br>\nEveryone is talking about Aortic Enlargement, but I am curious, how was your \"Other Lesion\" class score compare to others? It was one of the most difficult classes to deal with in my case</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1258437,
          "author_name": "Ahmet Erdem",
          "author_url": "",
          "post_date": "2021-03-31T15:56:22.487000",
          "content": "<p>We only submit the common classes separately to see if we can postprocess. Postprocessing rare classes based on small LB feedback was dangerous. Therefore I don't know its LB score but maybe my  teammates have an idea.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1258331,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2021-03-31T14:37:58.123000",
      "content": "<p>Congratulations on an awesome finish and great jump upward from public to private LB! The style gan art is great.</p>\n<p>We also noticed the huge difference between Aortic Enlargement CV and LB. On CV, Aortic Enlargement is the best class with local validation <code>mAP = 0.867</code> (on Yolo single model), but on private LB it only has <code>mAP = 0.210</code>. So test data 5 doctor consensus label is very different than train label. We didn't discover why. </p>\n<p>Great detective work making the bbox larger. That explains some of the difference. Using your trick boost our private Aortic Enlargment to <code>mAP = 0.255</code> (on same Yolo single model).</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1258337,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-03-31T14:40:22.320000",
          "content": "<p>Perhaps it isn't different bbox, but maybe the consensus decided many of the Aortic Enlargement are not Aortic Enlargement. So maybe finding some rule to decrease certain image confidence scores for Aortic Enlargement would boost that class.</p>\n<p>Now that we can make unlimited submissions, i'm tempted to solve this mystery </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1258346,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-03-31T14:47:32.117000",
          "content": "<p>2nd place posted lots of Aortic Enlargment investigation information here:<br>\n<a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229696\" target=\"_blank\">https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229696</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1258361,
          "author_name": "Ahmet Erdem",
          "author_url": "",
          "post_date": "2021-03-31T14:51:14.433000",
          "content": "<blockquote>\n  <p>Now that we can make unlimited submissions, i'm tempted to solve this mystery</p>\n</blockquote>\n<p>Now that the competition is finished, I will delete everything including Python and won't look back:)</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 1258364,
          "author_name": "beluga",
          "author_url": "",
          "post_date": "2021-03-31T14:53:40.620000",
          "content": "<p>We noticed that not ony the boxes are smaller but the highlighted aortas too. Half of the R8-R10 annotations had 20-40mm box width (<code>box_width_in_pixel x pixel_spacing</code>) while most definitions we found started aortic enlargement from 45-55mm. We tried to filter the smaller annotations but that experiment did not improve our LB score…</p>\n<p>We tried to use external data too but CRX14 and Chexpert does not report this class.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1258373,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-03-31T14:58:15.883000",
          "content": "<p>From my investigations it seems that annotator 12 seems most similar to at least parts of the test class0 labels. Maybe he is a specialist? Maybe there are even different specialists for different diseases?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1258422,
          "author_name": "InDSweTrust",
          "author_url": "",
          "post_date": "2021-03-31T15:38:32.583000",
          "content": "<p><a href=\"https://www.kaggle.com/aerdem4\" target=\"_blank\">@aerdem4</a> now that you potentially deleted python and everything associated, you can do the next competition in C++…or just create beautiful art work ;)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1258796,
          "author_name": "AmorfEvo",
          "author_url": "",
          "post_date": "2021-03-31T22:24:49.670000",
          "content": "<p>I would like to see <a href=\"https://www.kaggle.com/aerdem4\" target=\"_blank\">@aerdem4</a> competes in raw C++.<br>\nThat would be the next level ;)</p>\n<p>(But I'm afraid after he finally wins 1 competition that way, he will not only delete c++, but will also throw out his whole computer too :D and then go to the mountains and live there as a monk xD)</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1258308,
      "author_name": "AmorfEvo",
      "author_url": "",
      "post_date": "2021-03-31T14:19:20.600000",
      "content": "<p>One of the most dangerous gangs of Kaggle Street :D ^^</p>\n<p>Bonus results:<br>\nAfter this comp we have a lung poem + a style transfer art picture xD</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1258356,
          "author_name": "Ahmet Erdem",
          "author_url": "",
          "post_date": "2021-03-31T14:49:56.383000",
          "content": "<p>I would never expect a medical image competition to produce those great art pieces. Maybe Kaggle needs Artist ranking too:)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1258499,
          "author_name": "AmorfEvo",
          "author_url": "",
          "post_date": "2021-03-31T17:03:01.527000",
          "content": "<p>yeah, totally ^^</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1258158,
      "author_name": "Fatih Öztürk",
      "author_url": "",
      "post_date": "2021-03-31T11:54:13",
      "content": "<p>Congrats Ahmet and to the rest of the gangs! I also tried some postprocessing with box sizes but failed…</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1258147,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2021-03-31T11:44:17.660000",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/aerdem4\" target=\"_blank\">@aerdem4</a> <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> and <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> on 5th place. Nice jumping!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1320952,
      "author_name": "DeepUnderstanding",
      "author_url": "",
      "post_date": "2021-05-24T12:45:29.053000",
      "content": "<p>Can you share with us the GitHub link from where you have implemented the Effdet code?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1321468,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2021-05-24T17:39:53.353000",
          "content": "<p>sure: <a href=\"https://github.com/rwightman/pytorch-image-models\" target=\"_blank\">https://github.com/rwightman/pytorch-image-models</a><br>\n(Edit: sorry, this didn't contain the effdet, he had it in an extra repo, see my comment below)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1321640,
          "author_name": "DeepUnderstanding",
          "author_url": "",
          "post_date": "2021-05-24T20:16:52.173000",
          "content": "<p>wow, rwightman had detection models too! didn't know that, thanks</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1321671,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2021-05-24T20:54:11.720000",
          "content": "<p>Yeah, sorry above link wasn't exactly correct.</p>\n<p><a href=\"https://github.com/rwightman/efficientdet-pytorch\" target=\"_blank\">https://github.com/rwightman/efficientdet-pytorch</a></p>\n<p>Here is his effdet. It's a standalone repo.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1261883,
      "author_name": "Mohamed Bakrey Mahmoud",
      "author_url": "",
      "post_date": "2021-04-03T14:41:23.650000",
      "content": "<p>Very cool.<br>\nI have just made it a reference for me</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1261881,
      "author_name": "Mohamed Bakrey Mahmoud",
      "author_url": "",
      "post_date": "2021-04-03T14:40:45.747000",
      "content": "<p>Very cool.<br>\nI have just made it a reference for me</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1259712,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2021-04-01T15:47:54.977000",
      "content": "<p>From another discussion, i read that you trained your object detection on all images. (And you did not train a classifier model). Did you submit your object detection model's bbox confidence scores as is (for classes 0 thru 13) or did you post process them?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1259720,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-04-01T15:55:03.287000",
          "content": "<p>I'm curious because Ivan asked <a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229637#1259605\" target=\"_blank\">here</a>, if our team's post process is the same as training object detection on all images. If you read my answer, i say it is not. </p>\n<p>Many other teams including 6th place Guanshuo Xu found the formula <code>score = base * prob(finding)**0.2</code> to be optimal for the <code>mAP</code> metric in this comp. Where <code>base</code> are the confidence scores from an object detection model trained on <strong>only</strong> \"finding\" images. And <code>prob(finding)</code> is a classifier.</p>\n<p>I would think that training object detection on all images would be similar to <code>score = base * prob(finding)**1.0</code>.</p>\n<p>So the first formula is equivalent to <code>score = base**0.84 * prob(finding)**0.16</code>. And the second formula is equivalent to <code>score = base**0.5 * prob(finding)**0.5</code>.</p>\n<p>So if you didn't post process, i'm wondering if your confidence scores are similar to <code>score = base**0.5 * prob(finding)**0.5</code> and if so, i'm wondering whether you can increase your CV LB by making it more like <code>score = base**0.84 * prob(finding)**0.16</code></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1259839,
          "author_name": "Ahmet Erdem",
          "author_url": "",
          "post_date": "2021-04-01T17:36:35.647000",
          "content": "<p>We didn't postprocess box confidences. We also didn't use any classifier model. Class 14 predictions are multiplication of (1 - box_conf[i]).</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1259865,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-04-01T17:53:21.037000",
          "content": "<p>Wow great job! </p>\n<p>I'm curious if there is a way to post process your confidence scores to increase <code>mAP</code> CV LB. Did you try any post processing experiments on your bbox confidence scores?</p>\n<p>Like for each image with <code>prob(finding)&gt;0.5</code> then update all bbox (class 0 thru 13) confidence scores for that same image with <code>new_score = old_score ** 1.1</code>. And if that doesn't work, try <code>new score = old_score ** 0.9</code>.</p>\n<p>That formula won't affect scores with <code>prob=1</code> or <code>prob=0</code> but will affect middle probs. And i'm wondering if middle probs of images with finding should be treated differently than middle probs of images without findings.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1261984,
          "author_name": "Ivan Panshin",
          "author_url": "",
          "post_date": "2021-04-03T16:43:25.057000",
          "content": "<p>Another question: did you try using 1 - mult(box_conf[i]) instead of mult(1-box_conf[i]) as prediction for class 14? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1261999,
          "author_name": "Ivan Panshin",
          "author_url": "",
          "post_date": "2021-04-03T17:03:30.470000",
          "content": "<p>Never mind, my math is wrong </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1258764,
      "author_name": "DeepUnderstanding",
      "author_url": "",
      "post_date": "2021-03-31T21:22:55.043000",
      "content": "<p>If it was a Kernel only competition; because of very less test images i think u couldn't have done adversial validation; then what have been your approach?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1259138,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-04-01T07:04:56.187000",
          "content": "<p>Exactly the same as we did not use this information in the end as also written in the post.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1258107": "In this competition, I am happy that I have teamed up with one of the most dangerous gangs of Kaggle Street, namely @christofhenkel, @ilu000, @philippsinger. \n\n**Validation and Data Bias**\nIt was very difficult to make reliable Cross-Validation since the test set was generated by different labelling process and Public LB was not trustworthy due to its small size. We first noticed that our LB score is much lower than our CV score. It makes sense to some degree but the gap was really big, therefore we have investigated the data. We have trained an adversarial validation model on train vs test images. It has 0.61 AUC. The model was almost certainly able to distinguish some train set images from the test set but the rest looks like random split. These distinguishable images were all labelled by R8-R9-R10 with findings. So the different data sources were not evenly distributed among radiologists. This finding didn't really help us since we don't have the luxury to remove data with labellings. Therefore we have tried to remove the bias. They all had black corners covering some information. It made sense to add the same black boxes to the corners of the each image but it didn't remove the bias. Then we have tried gradCAM and according to it the bias was everywhere in the image, making it hard to remove.\n\n![gradCAM Result](https://i.imgur.com/mY34STX.png)\n\nWe have trained another model to detect which radiologists labelled the images. This model was very accurate on R1-7 vs the rest. Since R1-7 didn't have any labels, we have also removed them from our training set hoping that it can solve some of the bias issue. Knowing that bias is everywhere in the image, and not on a specific location, we thought that style transfer may solve the issue and ended up in this art.\n\n![Chest Art](https://i.imgur.com/Gqk1RBq.jpg)\n\nThis didn't work (obviously :D). \n\n**Rare Annotators**\nLosing hope on removing more bias, we have focused on improving our CV as it is with more focus on R11-17. We have trained several effdets with different backbones and standard augmentations. Data is sampled for each radiologist if the image has a finding, otherwise the image is sampled once during each epoch. We also had a YOLO model performing worse than our effdet. Ensembling them all with WBF@0.4 gave us our best score. Some experiments  with pretraining on R8-10, finetuning on R11-17 and applying WBF before training by weighting R11-17 more worked not good on Public LB, therefore we didn't use them. But they seem to work much better on Private LB like 0.224 -> 0.275.\n\n**Aortic Enlargement**\nWe noticed that something was wrong with this class because our LB score for it was way less than our CV score. Considering that it may be the way boxes are labelled in the test set, we have made a post processing for them which gave us an extra 0.01 on Public, 0.005 on Private.\n```\nsub.loc[sub.class_id==0, \"y_max\"] += 110\nsub.loc[sub.class_id==0, \"x_min\"] += 40\n```",
    "1258142": "Thanks for the awesome and professional team-up again @christofhenkel @aerdem4 @philippsinger !\nThat style transfer experiment was definetly a highlight.",
    "1258118": "Cool style transfer art!",
    "1259169": "Congratulations. Winning in a great jump!",
    "1258531": "Very interesting write-up (and approach but it goes without saying), varies a lot from the classical blending description, thanks a lot for that !\nAnd congratz on the impressive result but I guess that's no big surprise after all!",
    "1258418": "Thank you for the write up and congrats on your strong finish! I am seeing that working with Radialogist bias potentially had a lot of value...\nEveryone is talking about Aortic Enlargement, but I am curious, how was your \"Other Lesion\" class score compare to others? It was one of the most difficult classes to deal with in my case",
    "1258331": "Congratulations on an awesome finish and great jump upward from public to private LB! The style gan art is great.\n\nWe also noticed the huge difference between Aortic Enlargement CV and LB. On CV, Aortic Enlargement is the best class with local validation `mAP = 0.867` (on Yolo single model), but on private LB it only has `mAP = 0.210`. So test data 5 doctor consensus label is very different than train label. We didn't discover why. \n\nGreat detective work making the bbox larger. That explains some of the difference. Using your trick boost our private Aortic Enlargment to `mAP = 0.255` (on same Yolo single model).",
    "1258308": "One of the most dangerous gangs of Kaggle Street :D ^^\n\nBonus results:\nAfter this comp we have a lung poem + a style transfer art picture xD",
    "1258158": "Congrats Ahmet and to the rest of the gangs! I also tried some postprocessing with box sizes but failed...",
    "1258147": "Congrats @aerdem4 @ilu000 @christofhenkel and @philippsinger on 5th place. Nice jumping!",
    "1320952": "Can you share with us the GitHub link from where you have implemented the Effdet code?",
    "1261883": "Very cool.\nI have just made it a reference for me",
    "1261881": "Very cool.\nI have just made it a reference for me",
    "1259712": "From another discussion, i read that you trained your object detection on all images. (And you did not train a classifier model). Did you submit your object detection model's bbox confidence scores as is (for classes 0 thru 13) or did you post process them?",
    "1258764": "If it was a Kernel only competition; because of very less test images i think u couldn't have done adversial validation; then what have been your approach?"
  }
}