{
  "id": 229637,
  "title": "7th Place - Metric is like AUC",
  "url": "/competitions/vinbigdata-chest-xray-abnormalities-detection/discussion/229637",
  "author_name": "Chris Deotte",
  "post_date": "2021-03-31T03:57:51.310000",
  "votes": 107,
  "comment_count": 48,
  "views": 0,
  "content": "<p>I would like to thank my teammates <a href=\"https://www.kaggle.com/matthieuplante\" target=\"_blank\">@matthieuplante</a> <a href=\"https://www.kaggle.com/seb6084\" target=\"_blank\">@seb6084</a> for inviting me onto their team. I have been wanting to learn more about object detection and my teammates were the perfect teachers. </p>\n<p>When i joined them 1 week ago, they already had full pipelines to train VFNet, YoloV5, EfficientDet4 and more! Together we tuned the models, built some classifier models, and ensembled everything using weighted box fusion and Bayesian statistics!</p>\n<h1>Trust Your CV - mAP 0.488</h1>\n<p>Our final solution was an ensemble of three object detection models:</p>\n<ul>\n<li>VFNet on 1024x1024, infer 800x800 - CV mAP 0.463, Public LB 0.287, <strong>Private LB 0.289</strong></li>\n<li>Yolov5 on 1024x1024, infer 768x768 - CV mAP 0.464, Public LB 0.268, <strong>Private LB 0.290</strong></li>\n<li>EfficientDet4 on 1024x1024, infer 1024x1024 - CV 0.460, Public LB 0.237, <strong>Private LB 0.282</strong></li>\n</ul>\n<p>and three 15 class classifiers, i.e. 15 binary outs, with ensemble \"no finding\" <strong>CV AUC 0.9934</strong>:</p>\n<ul>\n<li>EfficientNet B4 on 512x512 - CV on finding/no finding output, AUC 0.990</li>\n<li>EfficientNet B4 on 768x768 - CV on finding/no finding, AUC 0.991</li>\n<li>DenseNet201 on 768x768 - CV on finding/no finding, AUC 0.991</li>\n</ul>\n<p>All decisions on how to ensemble folds and models. And how to combine object detection models with classifier models were determined via local validation.</p>\n<p>With ensemble, we achieved a validation score of <strong>0.488 mAP, public LB 0.272, and private LB 0.296</strong>. To simulate the test set with our validation set, we randomly only kept one doctor's predictions per class per image otherwise validation would not be like LB scoring.</p>\n<h1>Competition Metric is like AUC</h1>\n<p>For a given class, every bbox confidence score you predict needs to be ordered from most likely to least likely. Just like AUC metric. This is not per image, this is per CSV file for each class. If you predict a total of 30,000 bbox for class 0, then all 30,000 bbox confidence scores need to be correctly ordered from most likely to be class 0 to least likely to be class 0. </p>\n<h1>Bbox Confidence Score Order is Important</h1>\n<p>By simply adjusting the confidence scores of your bbox and not changing their xmin, xmax, ymin, ymax you can increase your CV and LB score! In the following example, the entire plot is 1 class, each dot is 1 bbox and the numbers are confidence scores. If the confidence score for bbox D and G are changed then the mAP for this class increases by +0.12 !! Therefore it is important to calibrate bbox probabilities.</p>\n<h2>Order 1</h2>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Feb-2022/map1.png\" alt=\"\"></p>\n<h2>Order 2</h2>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Feb-2022/map2.png\" alt=\"\"></p>\n<h1>No Penalty For Adding More Bbox</h1>\n<p>If you add new bbox with confidence score less than the lowest confidence score for all of that class in your entire CSV file, then your CV LB can only increase, it cannot decrease! Therefore instead of removing bbox it is better to calibrate probabilities, i.e. lower probability for unlikely bbox.<br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Feb-2022/map3.png\" alt=\"\"></p>\n<h1>Calibrated Confidence Scores</h1>\n<p>The most effective way to train an object detection model is to remove all the images without any bbox. After doing this, your object detection model will predict a confidence score equal to <br>\n$$ \\text{confidence score} = \\text{prob}(class \\text{ given } finding = \\text{True})$$<br>\nTherefore to get the calibrated probability, i.e correct confidence score for classes 0 thru 13, you must use conditional probability rule and multiply by <code>prob(finding) = 1 - prob(no finding)</code>. So we use our classifier and the formula<br>\n$$ \\text{confidence score} = \\text{prob}(class \\text{ given }finding = \\text{True}) * \\text{prob}(finding =\\text{True}) $$<br>\nThe first prob comes from object detection model and the second prob comes from classifier model ensemble. Using CV we discovered you can improve mAP metric with a 15 class classifier and the following formula:<br>\n$$ \\text{score} = base * \\text{prob}(class)^{0.1} * \\text{prob}(finding)^{0.1} $$<br>\nwhere <code>base = prob(class given finding = True)</code> that comes from your object detection model. And <code>prob(class)</code> is the probability that specific class is present on image from classifier model. And <code>prob(finding)</code> is the probability that image has a finding from classifier model.</p>\n<h1>Add Class 14, finding/ no finding</h1>\n<p>There is nothing fancy about adding predictions for class 14. For each image just add the string <code>14 PR 0 0 1 1</code> where <code>PR = prob(no finding)</code> from your classifier. If your classifier is perfect, then you will score full <code>mAP = 1.0</code> for class 14. Most classifiers have CV AUC 0.99+, so most teams will maximize this class mAP</p>",
  "messages": [
    {
      "id": 1257715,
      "postDate": "2021-03-31T03:57:51.310Z",
      "content": "<p>I would like to thank my teammates <a href=\"https://www.kaggle.com/matthieuplante\" target=\"_blank\">@matthieuplante</a> <a href=\"https://www.kaggle.com/seb6084\" target=\"_blank\">@seb6084</a> for inviting me onto their team. I have been wanting to learn more about object detection and my teammates were the perfect teachers. </p>\n<p>When i joined them 1 week ago, they already had full pipelines to train VFNet, YoloV5, EfficientDet4 and more! Together we tuned the models, built some classifier models, and ensembled everything using weighted box fusion and Bayesian statistics!</p>\n<h1>Trust Your CV - mAP 0.488</h1>\n<p>Our final solution was an ensemble of three object detection models:</p>\n<ul>\n<li>VFNet on 1024x1024, infer 800x800 - CV mAP 0.463, Public LB 0.287, <strong>Private LB 0.289</strong></li>\n<li>Yolov5 on 1024x1024, infer 768x768 - CV mAP 0.464, Public LB 0.268, <strong>Private LB 0.290</strong></li>\n<li>EfficientDet4 on 1024x1024, infer 1024x1024 - CV 0.460, Public LB 0.237, <strong>Private LB 0.282</strong></li>\n</ul>\n<p>and three 15 class classifiers, i.e. 15 binary outs, with ensemble \"no finding\" <strong>CV AUC 0.9934</strong>:</p>\n<ul>\n<li>EfficientNet B4 on 512x512 - CV on finding/no finding output, AUC 0.990</li>\n<li>EfficientNet B4 on 768x768 - CV on finding/no finding, AUC 0.991</li>\n<li>DenseNet201 on 768x768 - CV on finding/no finding, AUC 0.991</li>\n</ul>\n<p>All decisions on how to ensemble folds and models. And how to combine object detection models with classifier models were determined via local validation.</p>\n<p>With ensemble, we achieved a validation score of <strong>0.488 mAP, public LB 0.272, and private LB 0.296</strong>. To simulate the test set with our validation set, we randomly only kept one doctor's predictions per class per image otherwise validation would not be like LB scoring.</p>\n<h1>Competition Metric is like AUC</h1>\n<p>For a given class, every bbox confidence score you predict needs to be ordered from most likely to least likely. Just like AUC metric. This is not per image, this is per CSV file for each class. If you predict a total of 30,000 bbox for class 0, then all 30,000 bbox confidence scores need to be correctly ordered from most likely to be class 0 to least likely to be class 0. </p>\n<h1>Bbox Confidence Score Order is Important</h1>\n<p>By simply adjusting the confidence scores of your bbox and not changing their xmin, xmax, ymin, ymax you can increase your CV and LB score! In the following example, the entire plot is 1 class, each dot is 1 bbox and the numbers are confidence scores. If the confidence score for bbox D and G are changed then the mAP for this class increases by +0.12 !! Therefore it is important to calibrate bbox probabilities.</p>\n<h2>Order 1</h2>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Feb-2022/map1.png\" alt=\"\"></p>\n<h2>Order 2</h2>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Feb-2022/map2.png\" alt=\"\"></p>\n<h1>No Penalty For Adding More Bbox</h1>\n<p>If you add new bbox with confidence score less than the lowest confidence score for all of that class in your entire CSV file, then your CV LB can only increase, it cannot decrease! Therefore instead of removing bbox it is better to calibrate probabilities, i.e. lower probability for unlikely bbox.<br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Feb-2022/map3.png\" alt=\"\"></p>\n<h1>Calibrated Confidence Scores</h1>\n<p>The most effective way to train an object detection model is to remove all the images without any bbox. After doing this, your object detection model will predict a confidence score equal to <br>\n$$ \\text{confidence score} = \\text{prob}(class \\text{ given } finding = \\text{True})$$<br>\nTherefore to get the calibrated probability, i.e correct confidence score for classes 0 thru 13, you must use conditional probability rule and multiply by <code>prob(finding) = 1 - prob(no finding)</code>. So we use our classifier and the formula<br>\n$$ \\text{confidence score} = \\text{prob}(class \\text{ given }finding = \\text{True}) * \\text{prob}(finding =\\text{True}) $$<br>\nThe first prob comes from object detection model and the second prob comes from classifier model ensemble. Using CV we discovered you can improve mAP metric with a 15 class classifier and the following formula:<br>\n$$ \\text{score} = base * \\text{prob}(class)^{0.1} * \\text{prob}(finding)^{0.1} $$<br>\nwhere <code>base = prob(class given finding = True)</code> that comes from your object detection model. And <code>prob(class)</code> is the probability that specific class is present on image from classifier model. And <code>prob(finding)</code> is the probability that image has a finding from classifier model.</p>\n<h1>Add Class 14, finding/ no finding</h1>\n<p>There is nothing fancy about adding predictions for class 14. For each image just add the string <code>14 PR 0 0 1 1</code> where <code>PR = prob(no finding)</code> from your classifier. If your classifier is perfect, then you will score full <code>mAP = 1.0</code> for class 14. Most classifiers have CV AUC 0.99+, so most teams will maximize this class mAP</p>",
      "rawMarkdown": "I would like to thank my teammates @matthieuplante @seb6084 for inviting me onto their team. I have been wanting to learn more about object detection and my teammates were the perfect teachers. \n\nWhen i joined them 1 week ago, they already had full pipelines to train VFNet, YoloV5, EfficientDet4 and more! Together we tuned the models, built some classifier models, and ensembled everything using weighted box fusion and Bayesian statistics!\n\n# Trust Your CV - mAP 0.488\nOur final solution was an ensemble of three object detection models:\n* VFNet on 1024x1024, infer 800x800 - CV mAP 0.463, Public LB 0.287, **Private LB 0.289**\n* Yolov5 on 1024x1024, infer 768x768 - CV mAP 0.464, Public LB 0.268, **Private LB 0.290**\n* EfficientDet4 on 1024x1024, infer 1024x1024 - CV 0.460, Public LB 0.237, **Private LB 0.282**\n\nand three 15 class classifiers, i.e. 15 binary outs, with ensemble \"no finding\" **CV AUC 0.9934**:\n* EfficientNet B4 on 512x512 - CV on finding/no finding output, AUC 0.990\n* EfficientNet B4 on 768x768 - CV on finding/no finding, AUC 0.991\n* DenseNet201 on 768x768 - CV on finding/no finding, AUC 0.991\n\nAll decisions on how to ensemble folds and models. And how to combine object detection models with classifier models were determined via local validation.\n\nWith ensemble, we achieved a validation score of **0.488 mAP, public LB 0.272, and private LB 0.296**. To simulate the test set with our validation set, we randomly only kept one doctor's predictions per class per image otherwise validation would not be like LB scoring.\n\n# Competition Metric is like AUC\nFor a given class, every bbox confidence score you predict needs to be ordered from most likely to least likely. Just like AUC metric. This is not per image, this is per CSV file for each class. If you predict a total of 30,000 bbox for class 0, then all 30,000 bbox confidence scores need to be correctly ordered from most likely to be class 0 to least likely to be class 0. \n\n# Bbox Confidence Score Order is Important\nBy simply adjusting the confidence scores of your bbox and not changing their xmin, xmax, ymin, ymax you can increase your CV and LB score! In the following example, the entire plot is 1 class, each dot is 1 bbox and the numbers are confidence scores. If the confidence score for bbox D and G are changed then the mAP for this class increases by +0.12 !! Therefore it is important to calibrate bbox probabilities.\n## Order 1\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Feb-2022/map1.png)\n## Order 2\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Feb-2022/map2.png)\n\n# No Penalty For Adding More Bbox\nIf you add new bbox with confidence score less than the lowest confidence score for all of that class in your entire CSV file, then your CV LB can only increase, it cannot decrease! Therefore instead of removing bbox it is better to calibrate probabilities, i.e. lower probability for unlikely bbox.\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Feb-2022/map3.png)\n\n# Calibrated Confidence Scores\nThe most effective way to train an object detection model is to remove all the images without any bbox. After doing this, your object detection model will predict a confidence score equal to \n$$ \\text{confidence score} = \\text{prob}(class \\text{ given } finding = \\text{True})$$\nTherefore to get the calibrated probability, i.e correct confidence score for classes 0 thru 13, you must use conditional probability rule and multiply by `prob(finding) = 1 - prob(no finding)`. So we use our classifier and the formula\n$$ \\text{confidence score} = \\text{prob}(class \\text{ given }finding = \\text{True}) * \\text{prob}(finding =\\text{True}) $$\nThe first prob comes from object detection model and the second prob comes from classifier model ensemble. Using CV we discovered you can improve mAP metric with a 15 class classifier and the following formula:\n$$ \\text{score} = base * \\text{prob}(class)^{0.1} * \\text{prob}(finding)^{0.1} $$\nwhere `base = prob(class given finding = True)` that comes from your object detection model. And `prob(class)` is the probability that specific class is present on image from classifier model. And `prob(finding)` is the probability that image has a finding from classifier model.\n\n# Add Class 14, finding/ no finding\nThere is nothing fancy about adding predictions for class 14. For each image just add the string `14 PR 0 0 1 1` where `PR = prob(no finding)` from your classifier. If your classifier is perfect, then you will score full `mAP = 1.0` for class 14. Most classifiers have CV AUC 0.99+, so most teams will maximize this class mAP\n",
      "votes": 107
    },
    {
      "id": 1262621,
      "postDate": "2021-04-04T14:21:19.057Z",
      "content": "<p>Brilliant! From the melonoma classification, I have learned a lot from your discussion. thank you</p>\n<p>I have a question why your team trained the models with 1024 but inferred with lower resolutions?</p>",
      "rawMarkdown": "Brilliant! From the melonoma classification, I have learned a lot from your discussion. thank you\n\nI have a question why your team trained the models with 1024 but inferred with lower resolutions?",
      "votes": 5
    },
    {
      "id": 1263240,
      "postDate": "2021-04-05T08:30:04.340Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , thank for sharing your post-processing approach. I have a question. You said the accuracy will increase if the boxes confidence score increase. In the Figure 2, how about if all confidence score all increases 0.1, so all the position from A-G still remains same as Figure 1 right, so is it make any change different from Figure 1?</p>",
      "rawMarkdown": "Hi @cdeotte , thank for sharing your post-processing approach. I have a question. You said the accuracy will increase if the boxes confidence score increase. In the Figure 2, how about if all confidence score all increases 0.1, so all the position from A-G still remains same as Figure 1 right, so is it make any change different from Figure 1?\n",
      "votes": 1,
      "replies": [
        {
          "id": 1263638,
          "postDate": "2021-04-05T14:57:29.507Z",
          "content": "<p>No, then <code>mAP</code> will not increase. We increase <code>mAP</code> by rearranging the ranking order into a more accurate order just like the metric AUC.</p>",
          "rawMarkdown": "No, then `mAP` will not increase. We increase `mAP` by rearranging the ranking order into a more accurate order just like the metric AUC.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1259605,
      "postDate": "2021-04-01T14:38:30.853Z",
      "content": "<p>Awesome post-processing approach… as always :D </p>\n<p>Just one question: do you think this approach is more robust and just generally works better than simply training with some empty images, and not using classifiers to calibrate confidence scores for each box? </p>",
      "rawMarkdown": "Awesome post-processing approach... as always :D \n\nJust one question: do you think this approach is more robust and just generally works better than simply training with some empty images, and not using classifiers to calibrate confidence scores for each box? ",
      "votes": 1,
      "replies": [
        {
          "id": 1259671,
          "postDate": "2021-04-01T15:15:21.803Z",
          "content": "<p>Great question <a href=\"https://www.kaggle.com/ivanpan\" target=\"_blank\">@ivanpan</a> . Note that this post process is not the same as training object detection with all images.</p>\n<p>Bayesian statistics say we should use <code>score = base * prob(finding)</code> where base is the probability from object detection training on only \"finding\" images. I believe this formula would produce the same bbox confidence scores as training an object detection on all images.</p>\n<p>The above formula is equivilent (i.e. same rank order, same <code>mAP</code>) as using <code>score = base**0.5 * prob(finding)**0.5</code>. This is the geometric mean between the two. Now, <code>base</code> is over zealous because it was trained on only positive images, and multiplying by <code>prob(finding)</code> makes the result realistic.</p>\n<p>Our final formula is equivilent to <code>score = base**0.84 * prob(finding)**0.16</code>. This is the geometric average where we give more weight to over zealous bbox confidence scores which optimizes the loop holes of the <code>mAP</code> metric.</p>",
          "rawMarkdown": "Great question @ivanpan . Note that this post process is not the same as training object detection with all images.\n\nBayesian statistics say we should use `score = base * prob(finding)` where base is the probability from object detection training on only \"finding\" images. I believe this formula would produce the same bbox confidence scores as training an object detection on all images.\n\nThe above formula is equivilent (i.e. same rank order, same `mAP`) as using `score = base**0.5 * prob(finding)**0.5`. This is the geometric mean between the two. Now, `base` is over zealous because it was trained on only positive images, and multiplying by `prob(finding)` makes the result realistic.\n\nOur final formula is equivilent to `score = base**0.84 * prob(finding)**0.16`. This is the geometric average where we give more weight to over zealous bbox confidence scores which optimizes the loop holes of the `mAP` metric.",
          "votes": 1
        },
        {
          "id": 1259676,
          "postDate": "2021-04-01T15:17:40.260Z",
          "content": "<p>To clarify further, our formula is <code>score = base * prob(finding)**0.1 * prob(class)**0.1</code> which is more or less <code>score = base * prob(finding)**0.2</code> which is equivalent (same ranking ordering, same <code>mAP</code>) to <code>score = base**0.84 * prob(finding)**0.16</code> if you raise the right side to the exponent <code>5/6</code>. </p>",
          "rawMarkdown": "To clarify further, our formula is `score = base * prob(finding)**0.1 * prob(class)**0.1` which is more or less `score = base * prob(finding)**0.2` which is equivalent (same ranking ordering, same `mAP`) to `score = base**0.84 * prob(finding)**0.16` if you raise the right side to the exponent `5/6`. ",
          "votes": 1
        },
        {
          "id": 1259793,
          "postDate": "2021-04-01T17:03:25.073Z",
          "content": "<p>Hm, interesting. I actually understand the math perfectly (+ saw your response about Xu's solution), but it doesn't seem intuitive at all :) </p>",
          "rawMarkdown": "Hm, interesting. I actually understand the math perfectly (+ saw your response about Xu's solution), but it doesn't seem intuitive at all :) ",
          "votes": 1
        },
        {
          "id": 1259801,
          "postDate": "2021-04-01T17:08:19.173Z",
          "content": "<p>Now that I think about it, it's actually more or less okay. Since your ending formula is pretty much the same as <code>base**0.84 * prob(finding)**0.16</code>, which is just geometric mean. But how to get 0.84/0.16? Grid search on <code>[0, 1]</code> with step 0.01 + check validation scores? </p>",
          "rawMarkdown": "Now that I think about it, it's actually more or less okay. Since your ending formula is pretty much the same as `base**0.84 * prob(finding)**0.16`, which is just geometric mean. But how to get 0.84/0.16? Grid search on `[0, 1]` with step 0.01 + check validation scores? ",
          "votes": 1
        },
        {
          "id": 1259852,
          "postDate": "2021-04-01T17:43:29.680Z",
          "content": "<p>Yes, we found the formula using grid search on local validation. </p>\n<p>Your question has me wondering how similar or different our post process is to training object detection on all images. As i said above, i think they are different and i think if i had an object detection model trained on all images, i could post process the confidence scores with a new formula not published anywhere that could increase it's mAP CV LB. </p>\n<p>Because the <code>mAP</code> metric rewards both precision and recall. And the best way to increase recall is to make more bbox on the images that we believe have findings (i.e. <code>prob(finding)&gt;0.5</code>) and <code>mAP</code> does not penalize you for trying to find more bbox and being wrong.</p>",
          "rawMarkdown": "Yes, we found the formula using grid search on local validation. \n\nYour question has me wondering how similar or different our post process is to training object detection on all images. As i said above, i think they are different and i think if i had an object detection model trained on all images, i could post process the confidence scores with a new formula not published anywhere that could increase it's mAP CV LB. \n\nBecause the `mAP` metric rewards both precision and recall. And the best way to increase recall is to make more bbox on the images that we believe have findings (i.e. `prob(finding)>0.5`) and `mAP` does not penalize you for trying to find more bbox and being wrong.",
          "votes": 1
        },
        {
          "id": 1259863,
          "postDate": "2021-04-01T17:50:38.973Z",
          "content": "<p>Okay, it's official: my goal for 2021-2022 is to get into a team with you. I take models, you take post-processing 💪</p>",
          "rawMarkdown": "Okay, it's official: my goal for 2021-2022 is to get into a team with you. I take models, you take post-processing 💪"
        }
      ]
    },
    {
      "id": 1258818,
      "postDate": "2021-03-31T23:08:24.803Z",
      "content": "<p>Very well explained <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> !  Congrats once again !</p>",
      "rawMarkdown": "Very well explained @cdeotte !  Congrats once again !",
      "votes": 1
    },
    {
      "id": 1258729,
      "postDate": "2021-03-31T20:55:16.780Z",
      "content": "<p>Congrats! thanks for sharing insights about competition metric in detailed explanation.</p>\n<p>I could not understand how to calibrate probability in \"Order2\" way, isn't it possible to calibrate only D &amp; G class without knowing which is the true class??</p>",
      "rawMarkdown": "Congrats! thanks for sharing insights about competition metric in detailed explanation.\n\nI could not understand how to calibrate probability in \"Order2\" way, isn't it possible to calibrate only D & G class without knowing which is the true class??",
      "votes": 1,
      "replies": [
        {
          "id": 1258795,
          "postDate": "2021-03-31T22:23:05.167Z",
          "content": "<p>Thanks. The letters are not classes. The entire picture is a single class. Each letter is one bbox. The numbers are the confidence scores assigned to a bbox. And the bboxes are all bboxes in your entire <code>submission.csv</code> file for the 1 pictured class. (This is a made up example with few bbox).</p>\n<p>So given 1 class with its many bboxes in our submitted CSV file, if we can modify all the confidence scores in a smart way, then we can increase that 1 class's <code>mAP</code>.</p>",
          "rawMarkdown": "Thanks. The letters are not classes. The entire picture is a single class. Each letter is one bbox. The numbers are the confidence scores assigned to a bbox. And the bboxes are all bboxes in your entire `submission.csv` file for the 1 pictured class. (This is a made up example with few bbox).\n\nSo given 1 class with its many bboxes in our submitted CSV file, if we can modify all the confidence scores in a smart way, then we can increase that 1 class's `mAP`.",
          "votes": 1
        },
        {
          "id": 1260257,
          "postDate": "2021-04-02T01:08:49.767Z",
          "content": "<blockquote>\n  <p>if we can modify all the confidence scores in a smart way</p>\n</blockquote>\n<p>Thanks for reply, but I'm sorry that I still could not understand how to achieve this. Order does not affect to the metric, so just multiplying some value to all confidence does not change the metric.<br>\nThen, how you can change the order of confidence scores without knowing ground truth \"truth=0 or 1\"?</p>",
          "rawMarkdown": "> if we can modify all the confidence scores in a smart way\n\nThanks for reply, but I'm sorry that I still could not understand how to achieve this. Order does not affect to the metric, so just multiplying some value to all confidence does not change the metric.\nThen, how you can change the order of confidence scores without knowing ground truth \"truth=0 or 1\"?"
        }
      ]
    },
    {
      "id": 1258659,
      "postDate": "2021-03-31T19:37:12.467Z",
      "content": "<p>Congrats! I admire your post-processing very much.👍</p>",
      "rawMarkdown": "Congrats! I admire your post-processing very much.👍",
      "votes": 1
    },
    {
      "id": 1258221,
      "postDate": "2021-03-31T13:06:52.853Z",
      "content": "<p>Congrats and thanks for sharing your details! I have learned a lot from you. Thanks you.</p>",
      "rawMarkdown": "Congrats and thanks for sharing your details! I have learned a lot from you. Thanks you.",
      "votes": 1
    },
    {
      "id": 1257893,
      "postDate": "2021-03-31T07:30:21.377Z",
      "content": "<p>Congrats!! And thanks for the sharing in details, you always did a very fascinating job in post-process </p>",
      "rawMarkdown": "Congrats!! And thanks for the sharing in details, you always did a very fascinating job in post-process ",
      "votes": 1,
      "replies": [
        {
          "id": 1258367,
          "postDate": "2021-03-31T14:54:37.763Z",
          "content": "<p>Thanks Tuan</p>",
          "rawMarkdown": "Thanks Tuan",
          "votes": 1
        }
      ]
    },
    {
      "id": 1257822,
      "postDate": "2021-03-31T05:59:54.567Z",
      "content": "<p>Congratulations on the standing! It is people like you and your team whom really inspire beginner like myself to improve and become better. To have that intuition and experience to apply theories i never thought i could apply out of school ( Bayesian statistics )  👍</p>",
      "rawMarkdown": "Congratulations on the standing! It is people like you and your team whom really inspire beginner like myself to improve and become better. To have that intuition and experience to apply theories i never thought i could apply out of school ( Bayesian statistics )  👍",
      "votes": 1,
      "replies": [
        {
          "id": 1257850,
          "postDate": "2021-03-31T06:31:29.913Z",
          "content": "<p>Thanks. Congratulations on earning a bronze medal, top 10%</p>",
          "rawMarkdown": "Thanks. Congratulations on earning a bronze medal, top 10%",
          "votes": 1
        }
      ]
    },
    {
      "id": 1257801,
      "postDate": "2021-03-31T05:35:54.403Z",
      "content": "<p>Congrats for the good models and stable validation results! I like the postprocessing ideas!</p>\n<p>I wonder if you tested your Class 14 predictions separately by just submitting the 1 pixel boxes what would be the LB AP score?</p>",
      "rawMarkdown": "Congrats for the good models and stable validation results! I like the postprocessing ideas!\n\nI wonder if you tested your Class 14 predictions separately by just submitting the 1 pixel boxes what would be the LB AP score?",
      "votes": 1,
      "replies": [
        {
          "id": 1257817,
          "postDate": "2021-03-31T05:53:28.913Z",
          "content": "<p>Thanks. We did not during the competition, but i just submitted now. Our ensembled classifier output for finding / no finding achieves CV AUC 0.9934. And when submitted by itself achieves LB 0.064</p>",
          "rawMarkdown": "Thanks. We did not during the competition, but i just submitted now. Our ensembled classifier output for finding / no finding achieves CV AUC 0.9934. And when submitted by itself achieves LB 0.064",
          "votes": 1
        }
      ]
    },
    {
      "id": 1257785,
      "postDate": "2021-03-31T05:11:18.987Z",
      "content": "<p>Thanks for sharing ur great kernel! BTW, I have some questions here.</p>\n<ol>\n<li>What`s the purpose of 15 class classifiers? </li>\n<li>What`s the meaning of cv on \"class14\"? Is it a model that distinguishes between find/nofind? </li>\n</ol>\n<p>Also, the calibrated confidence is genius. For score ordering, I used sum-based nms method that sums up scores of overlapping boxes that predicted by ensemble models.</p>\n<p>Thanks again !</p>",
      "rawMarkdown": "Thanks for sharing ur great kernel! BTW, I have some questions here.\n\n1.  What`s the purpose of 15 class classifiers? \n2. What`s the meaning of cv on \"class14\"? Is it a model that distinguishes between find/nofind? \n\nAlso, the calibrated confidence is genius. For score ordering, I used sum-based nms method that sums up scores of overlapping boxes that predicted by ensemble models.\n\nThanks again !",
      "votes": 1,
      "replies": [
        {
          "id": 1257792,
          "postDate": "2021-03-31T05:27:13.510Z",
          "content": "<p>Thanks. Statistics say we only need to multiply the object detection confidence scores by the classifier probabililty \"no finding / finding\" scores to correct them. However through experimentation, we found we could increase CV mAP score more by also multiplying by the probability of class too.</p>\n<p>So for example if object detection says confidence for bbox of \"class 0\" is 0.78. Then we multiply 0.78 by probability that image has \"class 0\" and we multiply by probability that image has a \"finding\". We add exponents of 0.1 for each with formula</p>\n<p>$$ score = base * prob(class)^{0.1} * prob(finding)^{0.1}$$</p>\n<p>Our \"15 class classifier\" has 15 binary outputs and predicts the probability that a specific class is present. It was trained using the number of doctors who predict a class divided by 3 as ground truth.</p>\n<p>When i say CV on \"class 14\", yes i mean the OOF AUC of binary classifier for \"no finding / finding\". The ensemble of 3 classifiers achieves CV OOF of AUC 0.9934</p>\n<p>I like your idea your sum based NMS idea.</p>",
          "rawMarkdown": "Thanks. Statistics say we only need to multiply the object detection confidence scores by the classifier probabililty \"no finding / finding\" scores to correct them. However through experimentation, we found we could increase CV mAP score more by also multiplying by the probability of class too.\n\nSo for example if object detection says confidence for bbox of \"class 0\" is 0.78. Then we multiply 0.78 by probability that image has \"class 0\" and we multiply by probability that image has a \"finding\". We add exponents of 0.1 for each with formula\n\n$$ score = base * prob(class)^{0.1} * prob(finding)^{0.1}$$\n\nOur \"15 class classifier\" has 15 binary outputs and predicts the probability that a specific class is present. It was trained using the number of doctors who predict a class divided by 3 as ground truth.\n\nWhen i say CV on \"class 14\", yes i mean the OOF AUC of binary classifier for \"no finding / finding\". The ensemble of 3 classifiers achieves CV OOF of AUC 0.9934\n\nI like your idea your sum based NMS idea.",
          "votes": 2
        },
        {
          "id": 1257802,
          "postDate": "2021-03-31T05:36:04.047Z",
          "content": "<p>Thanks for your details !</p>",
          "rawMarkdown": "Thanks for your details !",
          "votes": 1
        }
      ]
    },
    {
      "id": 1257780,
      "postDate": "2021-03-31T05:02:12.750Z",
      "content": "<p>Congratulation on your final result. May I ask how many folds did you use for CV and did you stratify that or just randomly split the folds</p>",
      "rawMarkdown": "Congratulation on your final result. May I ask how many folds did you use for CV and did you stratify that or just randomly split the folds",
      "votes": 1,
      "replies": [
        {
          "id": 1257784,
          "postDate": "2021-03-31T05:11:03.827Z",
          "content": "<p>We had two validation schemes. The first was 10 stratified folds. We used this to make predictions on the test dataset. And we computed a CV mAP score on this.</p>\n<p>Next, to compare different fold ensembling techniques, we had a second validation scheme. In our second validation scheme, we put fold 9 and 10 aside as a holdout validation set. The we trained an 8 fold model on folds 1-8 without adding fold 9 and 10 to train data. Lastly we used the 8 folds to predict folds 9 and 10. We compared different ways to ensemble the 8 folds like NMS versus WBF etc. We computed a validation mAP score on this.</p>",
          "rawMarkdown": "We had two validation schemes. The first was 10 stratified folds. We used this to make predictions on the test dataset. And we computed a CV mAP score on this.\n\nNext, to compare different fold ensembling techniques, we had a second validation scheme. In our second validation scheme, we put fold 9 and 10 aside as a holdout validation set. The we trained an 8 fold model on folds 1-8 without adding fold 9 and 10 to train data. Lastly we used the 8 folds to predict folds 9 and 10. We compared different ways to ensemble the 8 folds like NMS versus WBF etc. We computed a validation mAP score on this.",
          "votes": 1
        },
        {
          "id": 1257959,
          "postDate": "2021-03-31T08:35:42.620Z",
          "content": "<p>The WBF has weights to denote the importances of different models when fusing boxes. Did you tune those, or use all weights=1 by default ?</p>",
          "rawMarkdown": "The WBF has weights to denote the importances of different models when fusing boxes. Did you tune those, or use all weights=1 by default ?",
          "votes": 1
        },
        {
          "id": 1258228,
          "postDate": "2021-03-31T13:14:58.137Z",
          "content": "<p>We tried tuning the weights, but using weights=1 gave best CV score. First we applied <code>NMS 0.4</code> to each fold separately. Next per model, we used <code>WBF iou_thr=0.40, conf_type='max'</code> to combine 10 model folds into 1 set of predictions. Next we used <code>WBF iou_thr=0.40, conf_type='max'</code> again to combine the 3 models into <code>submission.csv</code></p>",
          "rawMarkdown": "We tried tuning the weights, but using weights=1 gave best CV score. First we applied `NMS 0.4` to each fold separately. Next per model, we used `WBF iou_thr=0.40, conf_type='max'` to combine 10 model folds into 1 set of predictions. Next we used `WBF iou_thr=0.40, conf_type='max'` again to combine the 3 models into `submission.csv`"
        }
      ]
    },
    {
      "id": 1257750,
      "postDate": "2021-03-31T04:42:25.947Z",
      "content": "<p>Congrats on 7th place 👍</p>",
      "rawMarkdown": "Congrats on 7th place 👍",
      "votes": 1
    },
    {
      "id": 1257719,
      "postDate": "2021-03-31T04:01:55.977Z",
      "content": "<p>Congrats on 7th place <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> and team</p>",
      "rawMarkdown": "Congrats on 7th place @cdeotte and team",
      "votes": 1
    },
    {
      "id": 1267101,
      "postDate": "2021-04-08T10:21:00.433Z",
      "content": "<p><code>already had full pipelines to train VFNet, YoloV5, EfficientDet4 and more!</code><br>\nIt would be really beneficial for beginners if you shred your pipelines.<br>\nTo be honest, I tried to code a whole pipeline on my own but was never successful, hope to learn from your codes<br>\nit is better if its in PyTorch ; )</p>",
      "rawMarkdown": "`already had full pipelines to train VFNet, YoloV5, EfficientDet4 and more! `\nIt would be really beneficial for beginners if you shred your pipelines.\nTo be honest, I tried to code a whole pipeline on my own but was never successful, hope to learn from your codes\nit is better if its in PyTorch ; )",
      "votes": 2
    },
    {
      "id": 1258726,
      "postDate": "2021-03-31T20:45:57.153Z",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> any reason for <code>0.1</code> as power?</p>",
      "rawMarkdown": "@cdeotte any reason for `0.1` as power?",
      "votes": 2,
      "replies": [
        {
          "id": 1258793,
          "postDate": "2021-03-31T22:19:47.980Z",
          "content": "<p>Yes. It's a geometric mean. The formula <code>p = a * b**0.1 * c **0.1</code> produces the same ranking ordering (i.e. same <code>mAP</code>) as <code>a**0.84 * b**0.08 * c**0.08 = p**(5/6)</code>. This second expression is the geometric mean of numbers <code>a, b, c</code> with weights <code>0.84, 0.08, 0.08</code>. </p>\n<p>So we use 84% of the prediction from object detection model, 8% prediction from \"finding / no finding\" prediction, and 8% from \"specific class prediction\". These weights were found using local validation.</p>\n<p>You will also notice that Guanshuo Xu in his 6th place <a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229770\" target=\"_blank\">write-up</a> used a geometric mean of <code>a**0.84 * b**0.16</code>. He used 84% object detection and 16% \"finding / no finding\" prediction. Most teams found this geometric mean to work best.</p>",
          "rawMarkdown": "Yes. It's a geometric mean. The formula `p = a * b**0.1 * c **0.1` produces the same ranking ordering (i.e. same `mAP`) as `a**0.84 * b**0.08 * c**0.08 = p**(5/6)`. This second expression is the geometric mean of numbers `a, b, c` with weights `0.84, 0.08, 0.08`. \n\nSo we use 84% of the prediction from object detection model, 8% prediction from \"finding / no finding\" prediction, and 8% from \"specific class prediction\". These weights were found using local validation.\n\nYou will also notice that Guanshuo Xu in his 6th place [write-up][1] used a geometric mean of `a**0.84 * b**0.16`. He used 84% object detection and 16% \"finding / no finding\" prediction. Most teams found this geometric mean to work best.\n\n[1]: https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229770",
          "votes": 5
        },
        {
          "id": 1260990,
          "postDate": "2021-04-02T15:33:08.483Z",
          "content": "<blockquote>\n  <p>8% from \"specific class prediction\"</p>\n</blockquote>\n<p>But why did you need this class prediction? I suggest object detection models have a classifier inside. </p>\n<p>(As far as I understand you had special 15-class classifier.)</p>\n<blockquote>\n  <p>8% prediction from \"finding / no finding\" </p>\n</blockquote>\n<p>This part is clear because you trained without empty images.</p>\n<p>Or is it just empirical evidence that works better here?</p>",
          "rawMarkdown": ">  8% from \"specific class prediction\"\n\nBut why did you need this class prediction? I suggest object detection models have a classifier inside. \n\n(As far as I understand you had special 15-class classifier.)\n\n> 8% prediction from \"finding / no finding\" \n\nThis part is clear because you trained without empty images.\n\nOr is it just empirical evidence that works better here?\n",
          "votes": 1
        },
        {
          "id": 1261086,
          "postDate": "2021-04-02T17:19:54.883Z",
          "content": "<p>Good point. I agree it shouldn't be needed. Bayesian statistics says only <code>prediction from \"finding / no finding\"</code> is needed when we train object detection on just positive images.</p>\n<p>I think it helped for three reasons. </p>\n<ul>\n<li>First, having the extra 14 outputs on our classifier probably made the output for \"finding/ no finding\" more accurate due to multi-task learning. </li>\n<li>Second using 8% \"class\" and 8% \"finding\", is like ensembling two different diverse classifiers to achieve the goal of predicting \"finding\" better. </li>\n<li>Third, the <code>mAP</code> metric doesn't want realistic probabilities. After predicting bbox, you can increase <code>mAP</code> by adding more bbox and using a \"class\" classifier helps find which additional bbox to add.</li>\n</ul>\n<p>But ultimately the decision was based on local validation. Using our local validation we tried both </p>\n<ul>\n<li>16% \"finding\" classifier</li>\n<li>8% \"class\" ensembled with 8% \"finding\"</li>\n</ul>\n<p>The latter increased CV and private LB by +0.002</p>",
          "rawMarkdown": "Good point. I agree it shouldn't be needed. Bayesian statistics says only `prediction from \"finding / no finding\"` is needed when we train object detection on just positive images.\n\nI think it helped for three reasons. \n* First, having the extra 14 outputs on our classifier probably made the output for \"finding/ no finding\" more accurate due to multi-task learning. \n* Second using 8% \"class\" and 8% \"finding\", is like ensembling two different diverse classifiers to achieve the goal of predicting \"finding\" better. \n* Third, the `mAP` metric doesn't want realistic probabilities. After predicting bbox, you can increase `mAP` by adding more bbox and using a \"class\" classifier helps find which additional bbox to add.\n\nBut ultimately the decision was based on local validation. Using our local validation we tried both \n* 16% \"finding\" classifier\n* 8% \"class\" ensembled with 8% \"finding\"\n\nThe latter increased CV and private LB by +0.002",
          "votes": 2
        },
        {
          "id": 1261190,
          "postDate": "2021-04-02T19:06:33.543Z",
          "content": "<p>Wow, thank you for your detailed answer!<br>\nI think all your points are valid!</p>",
          "rawMarkdown": "Wow, thank you for your detailed answer!\nI think all your points are valid!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1258566,
      "postDate": "2021-03-31T18:00:27.537Z",
      "content": "<p>Congratz ! </p>\n<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> They should give you the Post-processing GM rank, I'm pretty sure that's the 5th gold you get with such ideas :)</p>",
      "rawMarkdown": "Congratz ! \n\n@cdeotte They should give you the Post-processing GM rank, I'm pretty sure that's the 5th gold you get with such ideas :)",
      "votes": 2,
      "replies": [
        {
          "id": 1258618,
          "postDate": "2021-03-31T18:53:21.743Z",
          "content": "<p>haha, true <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> <br>\nKaggle Post Process Grandmaster 😄</p>",
          "rawMarkdown": "haha, true @theoviel \nKaggle Post Process Grandmaster 😄",
          "votes": 1
        }
      ]
    },
    {
      "id": 1258018,
      "postDate": "2021-03-31T09:26:57.553Z",
      "content": "<p>Congrats on the result. We also used the probability calibration technique in post-processing . Our calibrated confidence was calculated as: confidence score=(prob(class given finding=True)^2)∗prob(finding=True)<br>\nWe missed the 15 class classifier part!! That was actually a very good idea !!</p>",
      "rawMarkdown": "Congrats on the result. We also used the probability calibration technique in post-processing . Our calibrated confidence was calculated as: confidence score=(prob(class given finding=True)^2)∗prob(finding=True)\nWe missed the 15 class classifier part!! That was actually a very good idea !!",
      "votes": 2,
      "replies": [
        {
          "id": 1258049,
          "postDate": "2021-03-31T09:46:40.197Z",
          "content": "<p>Congrats for your score too ! Yes the 15 class classifier had no big impact on LB but was very strong on CV so we trusted it </p>",
          "rawMarkdown": "Congrats for your score too ! Yes the 15 class classifier had no big impact on LB but was very strong on CV so we trusted it ",
          "votes": 1
        },
        {
          "id": 1258560,
          "postDate": "2021-03-31T17:55:50.197Z",
          "content": "<p>Congratulations KS and team, well done!</p>\n<p>Just to confirm, is your formula the \"output score from object detection squared\" multiplied by \"output score from classifier\"?</p>",
          "rawMarkdown": "Congratulations KS and team, well done!\n\nJust to confirm, is your formula the \"output score from object detection squared\" multiplied by \"output score from classifier\"?",
          "votes": 1
        },
        {
          "id": 1258568,
          "postDate": "2021-03-31T18:01:02.517Z",
          "content": "<p>That should work good. That is the same as <code>confidence score= prob(class given finding=True) ∗ prob(finding=True)**0.5</code> because both formulas keep scores in the same order. The formula here is just the square root of yours above.</p>\n<p>For us, using <code>confidence score= prob(class given finding=True) ∗ prob(finding=True)**0.2</code> worked nearly as well as our formula posted above which uses our 15 class classifier. </p>\n<p>Did you try the formula <br>\n<code>confidence score=(prob(class given finding=True)^5)∗prob(finding=True)</code> <br>\nor the equivilent<br>\n <code>confidence score= prob(class given finding=True) ∗ prob(finding=True)**0.2</code> ?</p>",
          "rawMarkdown": "That should work good. That is the same as `confidence score= prob(class given finding=True) ∗ prob(finding=True)**0.5` because both formulas keep scores in the same order. The formula here is just the square root of yours above.\n\nFor us, using `confidence score= prob(class given finding=True) ∗ prob(finding=True)**0.2` worked nearly as well as our formula posted above which uses our 15 class classifier. \n\nDid you try the formula \n`confidence score=(prob(class given finding=True)^5)∗prob(finding=True)` \nor the equivilent\n `confidence score= prob(class given finding=True) ∗ prob(finding=True)**0.2` ?",
          "votes": 1
        },
        {
          "id": 1259058,
          "postDate": "2021-04-01T05:44:10.213Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Yes, the formula is  \"output score from object detection squared\" multiplied by \"output score from classifier\". I tried powers from 2 to 10 for prob ( class given finding=True ) on CV. I got the best results with power=2!! Results with power=2 and 3 were similar and it started dropping after power =4!!</p>",
          "rawMarkdown": "@cdeotte Yes, the formula is  \"output score from object detection squared\" multiplied by \"output score from classifier\". I tried powers from 2 to 10 for prob ( class given finding=True ) on CV. I got the best results with power=2!! Results with power=2 and 3 were similar and it started dropping after power =4!!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1257957,
      "postDate": "2021-03-31T08:34:40.187Z",
      "content": "<p>Congrats!    </p>",
      "rawMarkdown": "Congrats!    ",
      "votes": 2
    },
    {
      "id": 1257902,
      "postDate": "2021-03-31T07:40:18.623Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , for detection, did you use 14 or 15 class training?</p>",
      "rawMarkdown": "Congratulations @cdeotte , for detection, did you use 14 or 15 class training?",
      "votes": 2,
      "replies": [
        {
          "id": 1258028,
          "postDate": "2021-03-31T09:35:00.590Z",
          "content": "<p>All our models are trained with 14 classes </p>",
          "rawMarkdown": "All our models are trained with 14 classes ",
          "votes": 2
        }
      ]
    },
    {
      "id": 1322642,
      "postDate": "2021-05-25T15:10:34.697Z",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> we are all fan of your post-processing techniques. Most of them are based on Bayesian methods. Can you share some good resources from where you have studied this.<br>\nThanks</p>",
      "rawMarkdown": "Hi, @cdeotte we are all fan of your post-processing techniques. Most of them are based on Bayesian methods. Can you share some good resources from where you have studied this.\nThanks"
    }
  ],
  "comments": [
    {
      "id": 1262621,
      "author_name": "kaggler",
      "author_url": "",
      "post_date": "2021-04-04T14:21:19.057000",
      "content": "<p>Brilliant! From the melonoma classification, I have learned a lot from your discussion. thank you</p>\n<p>I have a question why your team trained the models with 1024 but inferred with lower resolutions?</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 1263240,
      "author_name": "David",
      "author_url": "",
      "post_date": "2021-04-05T08:30:04.340000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , thank for sharing your post-processing approach. I have a question. You said the accuracy will increase if the boxes confidence score increase. In the Figure 2, how about if all confidence score all increases 0.1, so all the position from A-G still remains same as Figure 1 right, so is it make any change different from Figure 1?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1263638,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-04-05T14:57:29.507000",
          "content": "<p>No, then <code>mAP</code> will not increase. We increase <code>mAP</code> by rearranging the ranking order into a more accurate order just like the metric AUC.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1259605,
      "author_name": "Ivan Panshin",
      "author_url": "",
      "post_date": "2021-04-01T14:38:30.853000",
      "content": "<p>Awesome post-processing approach… as always :D </p>\n<p>Just one question: do you think this approach is more robust and just generally works better than simply training with some empty images, and not using classifiers to calibrate confidence scores for each box? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1259671,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-04-01T15:15:21.803000",
          "content": "<p>Great question <a href=\"https://www.kaggle.com/ivanpan\" target=\"_blank\">@ivanpan</a> . Note that this post process is not the same as training object detection with all images.</p>\n<p>Bayesian statistics say we should use <code>score = base * prob(finding)</code> where base is the probability from object detection training on only \"finding\" images. I believe this formula would produce the same bbox confidence scores as training an object detection on all images.</p>\n<p>The above formula is equivilent (i.e. same rank order, same <code>mAP</code>) as using <code>score = base**0.5 * prob(finding)**0.5</code>. This is the geometric mean between the two. Now, <code>base</code> is over zealous because it was trained on only positive images, and multiplying by <code>prob(finding)</code> makes the result realistic.</p>\n<p>Our final formula is equivilent to <code>score = base**0.84 * prob(finding)**0.16</code>. This is the geometric average where we give more weight to over zealous bbox confidence scores which optimizes the loop holes of the <code>mAP</code> metric.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1259676,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-04-01T15:17:40.260000",
          "content": "<p>To clarify further, our formula is <code>score = base * prob(finding)**0.1 * prob(class)**0.1</code> which is more or less <code>score = base * prob(finding)**0.2</code> which is equivalent (same ranking ordering, same <code>mAP</code>) to <code>score = base**0.84 * prob(finding)**0.16</code> if you raise the right side to the exponent <code>5/6</code>. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1259793,
          "author_name": "Ivan Panshin",
          "author_url": "",
          "post_date": "2021-04-01T17:03:25.073000",
          "content": "<p>Hm, interesting. I actually understand the math perfectly (+ saw your response about Xu's solution), but it doesn't seem intuitive at all :) </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1259801,
          "author_name": "Ivan Panshin",
          "author_url": "",
          "post_date": "2021-04-01T17:08:19.173000",
          "content": "<p>Now that I think about it, it's actually more or less okay. Since your ending formula is pretty much the same as <code>base**0.84 * prob(finding)**0.16</code>, which is just geometric mean. But how to get 0.84/0.16? Grid search on <code>[0, 1]</code> with step 0.01 + check validation scores? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1259852,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-04-01T17:43:29.680000",
          "content": "<p>Yes, we found the formula using grid search on local validation. </p>\n<p>Your question has me wondering how similar or different our post process is to training object detection on all images. As i said above, i think they are different and i think if i had an object detection model trained on all images, i could post process the confidence scores with a new formula not published anywhere that could increase it's mAP CV LB. </p>\n<p>Because the <code>mAP</code> metric rewards both precision and recall. And the best way to increase recall is to make more bbox on the images that we believe have findings (i.e. <code>prob(finding)&gt;0.5</code>) and <code>mAP</code> does not penalize you for trying to find more bbox and being wrong.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1259863,
          "author_name": "Ivan Panshin",
          "author_url": "",
          "post_date": "2021-04-01T17:50:38.973000",
          "content": "<p>Okay, it's official: my goal for 2021-2022 is to get into a team with you. I take models, you take post-processing 💪</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1258818,
      "author_name": "Sidney Ng",
      "author_url": "",
      "post_date": "2021-03-31T23:08:24.803000",
      "content": "<p>Very well explained <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> !  Congrats once again !</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1258729,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "2021-03-31T20:55:16.780000",
      "content": "<p>Congrats! thanks for sharing insights about competition metric in detailed explanation.</p>\n<p>I could not understand how to calibrate probability in \"Order2\" way, isn't it possible to calibrate only D &amp; G class without knowing which is the true class??</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1258795,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-03-31T22:23:05.167000",
          "content": "<p>Thanks. The letters are not classes. The entire picture is a single class. Each letter is one bbox. The numbers are the confidence scores assigned to a bbox. And the bboxes are all bboxes in your entire <code>submission.csv</code> file for the 1 pictured class. (This is a made up example with few bbox).</p>\n<p>So given 1 class with its many bboxes in our submitted CSV file, if we can modify all the confidence scores in a smart way, then we can increase that 1 class's <code>mAP</code>.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1260257,
          "author_name": "corochann",
          "author_url": "",
          "post_date": "2021-04-02T01:08:49.767000",
          "content": "<blockquote>\n  <p>if we can modify all the confidence scores in a smart way</p>\n</blockquote>\n<p>Thanks for reply, but I'm sorry that I still could not understand how to achieve this. Order does not affect to the metric, so just multiplying some value to all confidence does not change the metric.<br>\nThen, how you can change the order of confidence scores without knowing ground truth \"truth=0 or 1\"?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1258659,
      "author_name": "Alien",
      "author_url": "",
      "post_date": "2021-03-31T19:37:12.467000",
      "content": "<p>Congrats! I admire your post-processing very much.👍</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1258221,
      "author_name": "Sunghyun Jun",
      "author_url": "",
      "post_date": "2021-03-31T13:06:52.853000",
      "content": "<p>Congrats and thanks for sharing your details! I have learned a lot from you. Thanks you.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1257893,
      "author_name": "Tuan Ho Lai",
      "author_url": "",
      "post_date": "2021-03-31T07:30:21.377000",
      "content": "<p>Congrats!! And thanks for the sharing in details, you always did a very fascinating job in post-process </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1258367,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-03-31T14:54:37.763000",
          "content": "<p>Thanks Tuan</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1257822,
      "author_name": "aqx",
      "author_url": "",
      "post_date": "2021-03-31T05:59:54.567000",
      "content": "<p>Congratulations on the standing! It is people like you and your team whom really inspire beginner like myself to improve and become better. To have that intuition and experience to apply theories i never thought i could apply out of school ( Bayesian statistics )  👍</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1257850,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-03-31T06:31:29.913000",
          "content": "<p>Thanks. Congratulations on earning a bronze medal, top 10%</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1257801,
      "author_name": "beluga",
      "author_url": "",
      "post_date": "2021-03-31T05:35:54.403000",
      "content": "<p>Congrats for the good models and stable validation results! I like the postprocessing ideas!</p>\n<p>I wonder if you tested your Class 14 predictions separately by just submitting the 1 pixel boxes what would be the LB AP score?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1257817,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-03-31T05:53:28.913000",
          "content": "<p>Thanks. We did not during the competition, but i just submitted now. Our ensembled classifier output for finding / no finding achieves CV AUC 0.9934. And when submitted by itself achieves LB 0.064</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1257785,
      "author_name": "Wonho Song",
      "author_url": "",
      "post_date": "2021-03-31T05:11:18.987000",
      "content": "<p>Thanks for sharing ur great kernel! BTW, I have some questions here.</p>\n<ol>\n<li>What`s the purpose of 15 class classifiers? </li>\n<li>What`s the meaning of cv on \"class14\"? Is it a model that distinguishes between find/nofind? </li>\n</ol>\n<p>Also, the calibrated confidence is genius. For score ordering, I used sum-based nms method that sums up scores of overlapping boxes that predicted by ensemble models.</p>\n<p>Thanks again !</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1257792,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-03-31T05:27:13.510000",
          "content": "<p>Thanks. Statistics say we only need to multiply the object detection confidence scores by the classifier probabililty \"no finding / finding\" scores to correct them. However through experimentation, we found we could increase CV mAP score more by also multiplying by the probability of class too.</p>\n<p>So for example if object detection says confidence for bbox of \"class 0\" is 0.78. Then we multiply 0.78 by probability that image has \"class 0\" and we multiply by probability that image has a \"finding\". We add exponents of 0.1 for each with formula</p>\n<p>$$ score = base * prob(class)^{0.1} * prob(finding)^{0.1}$$</p>\n<p>Our \"15 class classifier\" has 15 binary outputs and predicts the probability that a specific class is present. It was trained using the number of doctors who predict a class divided by 3 as ground truth.</p>\n<p>When i say CV on \"class 14\", yes i mean the OOF AUC of binary classifier for \"no finding / finding\". The ensemble of 3 classifiers achieves CV OOF of AUC 0.9934</p>\n<p>I like your idea your sum based NMS idea.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1257802,
          "author_name": "Wonho Song",
          "author_url": "",
          "post_date": "2021-03-31T05:36:04.047000",
          "content": "<p>Thanks for your details !</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1257780,
      "author_name": "Liam Nguyen",
      "author_url": "",
      "post_date": "2021-03-31T05:02:12.750000",
      "content": "<p>Congratulation on your final result. May I ask how many folds did you use for CV and did you stratify that or just randomly split the folds</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1257784,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-03-31T05:11:03.827000",
          "content": "<p>We had two validation schemes. The first was 10 stratified folds. We used this to make predictions on the test dataset. And we computed a CV mAP score on this.</p>\n<p>Next, to compare different fold ensembling techniques, we had a second validation scheme. In our second validation scheme, we put fold 9 and 10 aside as a holdout validation set. The we trained an 8 fold model on folds 1-8 without adding fold 9 and 10 to train data. Lastly we used the 8 folds to predict folds 9 and 10. We compared different ways to ensemble the 8 folds like NMS versus WBF etc. We computed a validation mAP score on this.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1257959,
          "author_name": "Liam Nguyen",
          "author_url": "",
          "post_date": "2021-03-31T08:35:42.620000",
          "content": "<p>The WBF has weights to denote the importances of different models when fusing boxes. Did you tune those, or use all weights=1 by default ?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1258228,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-03-31T13:14:58.137000",
          "content": "<p>We tried tuning the weights, but using weights=1 gave best CV score. First we applied <code>NMS 0.4</code> to each fold separately. Next per model, we used <code>WBF iou_thr=0.40, conf_type='max'</code> to combine 10 model folds into 1 set of predictions. Next we used <code>WBF iou_thr=0.40, conf_type='max'</code> again to combine the 3 models into <code>submission.csv</code></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1257750,
      "author_name": "Jagadish Sivakumar",
      "author_url": "",
      "post_date": "2021-03-31T04:42:25.947000",
      "content": "<p>Congrats on 7th place 👍</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1257719,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2021-03-31T04:01:55.977000",
      "content": "<p>Congrats on 7th place <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> and team</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1267101,
      "author_name": "DeepUnderstanding",
      "author_url": "",
      "post_date": "2021-04-08T10:21:00.433000",
      "content": "<p><code>already had full pipelines to train VFNet, YoloV5, EfficientDet4 and more!</code><br>\nIt would be really beneficial for beginners if you shred your pipelines.<br>\nTo be honest, I tried to code a whole pipeline on my own but was never successful, hope to learn from your codes<br>\nit is better if its in PyTorch ; )</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1258726,
      "author_name": "Awsaf",
      "author_url": "",
      "post_date": "2021-03-31T20:45:57.153000",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> any reason for <code>0.1</code> as power?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1258793,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-03-31T22:19:47.980000",
          "content": "<p>Yes. It's a geometric mean. The formula <code>p = a * b**0.1 * c **0.1</code> produces the same ranking ordering (i.e. same <code>mAP</code>) as <code>a**0.84 * b**0.08 * c**0.08 = p**(5/6)</code>. This second expression is the geometric mean of numbers <code>a, b, c</code> with weights <code>0.84, 0.08, 0.08</code>. </p>\n<p>So we use 84% of the prediction from object detection model, 8% prediction from \"finding / no finding\" prediction, and 8% from \"specific class prediction\". These weights were found using local validation.</p>\n<p>You will also notice that Guanshuo Xu in his 6th place <a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229770\" target=\"_blank\">write-up</a> used a geometric mean of <code>a**0.84 * b**0.16</code>. He used 84% object detection and 16% \"finding / no finding\" prediction. Most teams found this geometric mean to work best.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1260990,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2021-04-02T15:33:08.483000",
          "content": "<blockquote>\n  <p>8% from \"specific class prediction\"</p>\n</blockquote>\n<p>But why did you need this class prediction? I suggest object detection models have a classifier inside. </p>\n<p>(As far as I understand you had special 15-class classifier.)</p>\n<blockquote>\n  <p>8% prediction from \"finding / no finding\" </p>\n</blockquote>\n<p>This part is clear because you trained without empty images.</p>\n<p>Or is it just empirical evidence that works better here?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1261086,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-04-02T17:19:54.883000",
          "content": "<p>Good point. I agree it shouldn't be needed. Bayesian statistics says only <code>prediction from \"finding / no finding\"</code> is needed when we train object detection on just positive images.</p>\n<p>I think it helped for three reasons. </p>\n<ul>\n<li>First, having the extra 14 outputs on our classifier probably made the output for \"finding/ no finding\" more accurate due to multi-task learning. </li>\n<li>Second using 8% \"class\" and 8% \"finding\", is like ensembling two different diverse classifiers to achieve the goal of predicting \"finding\" better. </li>\n<li>Third, the <code>mAP</code> metric doesn't want realistic probabilities. After predicting bbox, you can increase <code>mAP</code> by adding more bbox and using a \"class\" classifier helps find which additional bbox to add.</li>\n</ul>\n<p>But ultimately the decision was based on local validation. Using our local validation we tried both </p>\n<ul>\n<li>16% \"finding\" classifier</li>\n<li>8% \"class\" ensembled with 8% \"finding\"</li>\n</ul>\n<p>The latter increased CV and private LB by +0.002</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1261190,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2021-04-02T19:06:33.543000",
          "content": "<p>Wow, thank you for your detailed answer!<br>\nI think all your points are valid!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1258566,
      "author_name": "Theo Viel",
      "author_url": "",
      "post_date": "2021-03-31T18:00:27.537000",
      "content": "<p>Congratz ! </p>\n<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> They should give you the Post-processing GM rank, I'm pretty sure that's the 5th gold you get with such ideas :)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1258618,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-03-31T18:53:21.743000",
          "content": "<p>haha, true <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> <br>\nKaggle Post Process Grandmaster 😄</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1258018,
      "author_name": "Kumar Shubham",
      "author_url": "",
      "post_date": "2021-03-31T09:26:57.553000",
      "content": "<p>Congrats on the result. We also used the probability calibration technique in post-processing . Our calibrated confidence was calculated as: confidence score=(prob(class given finding=True)^2)∗prob(finding=True)<br>\nWe missed the 15 class classifier part!! That was actually a very good idea !!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1258049,
          "author_name": "Matthieu Planté",
          "author_url": "",
          "post_date": "2021-03-31T09:46:40.197000",
          "content": "<p>Congrats for your score too ! Yes the 15 class classifier had no big impact on LB but was very strong on CV so we trusted it </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1258560,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-03-31T17:55:50.197000",
          "content": "<p>Congratulations KS and team, well done!</p>\n<p>Just to confirm, is your formula the \"output score from object detection squared\" multiplied by \"output score from classifier\"?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1258568,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-03-31T18:01:02.517000",
          "content": "<p>That should work good. That is the same as <code>confidence score= prob(class given finding=True) ∗ prob(finding=True)**0.5</code> because both formulas keep scores in the same order. The formula here is just the square root of yours above.</p>\n<p>For us, using <code>confidence score= prob(class given finding=True) ∗ prob(finding=True)**0.2</code> worked nearly as well as our formula posted above which uses our 15 class classifier. </p>\n<p>Did you try the formula <br>\n<code>confidence score=(prob(class given finding=True)^5)∗prob(finding=True)</code> <br>\nor the equivilent<br>\n <code>confidence score= prob(class given finding=True) ∗ prob(finding=True)**0.2</code> ?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1259058,
          "author_name": "Kumar Shubham",
          "author_url": "",
          "post_date": "2021-04-01T05:44:10.213000",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Yes, the formula is  \"output score from object detection squared\" multiplied by \"output score from classifier\". I tried powers from 2 to 10 for prob ( class given finding=True ) on CV. I got the best results with power=2!! Results with power=2 and 3 were similar and it started dropping after power =4!!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1257957,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2021-03-31T08:34:40.187000",
      "content": "<p>Congrats!    </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1257902,
      "author_name": "Awsaf",
      "author_url": "",
      "post_date": "2021-03-31T07:40:18.623000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , for detection, did you use 14 or 15 class training?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1258028,
          "author_name": "Matthieu Planté",
          "author_url": "",
          "post_date": "2021-03-31T09:35:00.590000",
          "content": "<p>All our models are trained with 14 classes </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1322642,
      "author_name": "DeepUnderstanding",
      "author_url": "",
      "post_date": "2021-05-25T15:10:34.697000",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> we are all fan of your post-processing techniques. Most of them are based on Bayesian methods. Can you share some good resources from where you have studied this.<br>\nThanks</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1257715": "I would like to thank my teammates @matthieuplante @seb6084 for inviting me onto their team. I have been wanting to learn more about object detection and my teammates were the perfect teachers. \n\nWhen i joined them 1 week ago, they already had full pipelines to train VFNet, YoloV5, EfficientDet4 and more! Together we tuned the models, built some classifier models, and ensembled everything using weighted box fusion and Bayesian statistics!\n\n# Trust Your CV - mAP 0.488\nOur final solution was an ensemble of three object detection models:\n* VFNet on 1024x1024, infer 800x800 - CV mAP 0.463, Public LB 0.287, **Private LB 0.289**\n* Yolov5 on 1024x1024, infer 768x768 - CV mAP 0.464, Public LB 0.268, **Private LB 0.290**\n* EfficientDet4 on 1024x1024, infer 1024x1024 - CV 0.460, Public LB 0.237, **Private LB 0.282**\n\nand three 15 class classifiers, i.e. 15 binary outs, with ensemble \"no finding\" **CV AUC 0.9934**:\n* EfficientNet B4 on 512x512 - CV on finding/no finding output, AUC 0.990\n* EfficientNet B4 on 768x768 - CV on finding/no finding, AUC 0.991\n* DenseNet201 on 768x768 - CV on finding/no finding, AUC 0.991\n\nAll decisions on how to ensemble folds and models. And how to combine object detection models with classifier models were determined via local validation.\n\nWith ensemble, we achieved a validation score of **0.488 mAP, public LB 0.272, and private LB 0.296**. To simulate the test set with our validation set, we randomly only kept one doctor's predictions per class per image otherwise validation would not be like LB scoring.\n\n# Competition Metric is like AUC\nFor a given class, every bbox confidence score you predict needs to be ordered from most likely to least likely. Just like AUC metric. This is not per image, this is per CSV file for each class. If you predict a total of 30,000 bbox for class 0, then all 30,000 bbox confidence scores need to be correctly ordered from most likely to be class 0 to least likely to be class 0. \n\n# Bbox Confidence Score Order is Important\nBy simply adjusting the confidence scores of your bbox and not changing their xmin, xmax, ymin, ymax you can increase your CV and LB score! In the following example, the entire plot is 1 class, each dot is 1 bbox and the numbers are confidence scores. If the confidence score for bbox D and G are changed then the mAP for this class increases by +0.12 !! Therefore it is important to calibrate bbox probabilities.\n## Order 1\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Feb-2022/map1.png)\n## Order 2\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Feb-2022/map2.png)\n\n# No Penalty For Adding More Bbox\nIf you add new bbox with confidence score less than the lowest confidence score for all of that class in your entire CSV file, then your CV LB can only increase, it cannot decrease! Therefore instead of removing bbox it is better to calibrate probabilities, i.e. lower probability for unlikely bbox.\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Feb-2022/map3.png)\n\n# Calibrated Confidence Scores\nThe most effective way to train an object detection model is to remove all the images without any bbox. After doing this, your object detection model will predict a confidence score equal to \n$$ \\text{confidence score} = \\text{prob}(class \\text{ given } finding = \\text{True})$$\nTherefore to get the calibrated probability, i.e correct confidence score for classes 0 thru 13, you must use conditional probability rule and multiply by `prob(finding) = 1 - prob(no finding)`. So we use our classifier and the formula\n$$ \\text{confidence score} = \\text{prob}(class \\text{ given }finding = \\text{True}) * \\text{prob}(finding =\\text{True}) $$\nThe first prob comes from object detection model and the second prob comes from classifier model ensemble. Using CV we discovered you can improve mAP metric with a 15 class classifier and the following formula:\n$$ \\text{score} = base * \\text{prob}(class)^{0.1} * \\text{prob}(finding)^{0.1} $$\nwhere `base = prob(class given finding = True)` that comes from your object detection model. And `prob(class)` is the probability that specific class is present on image from classifier model. And `prob(finding)` is the probability that image has a finding from classifier model.\n\n# Add Class 14, finding/ no finding\nThere is nothing fancy about adding predictions for class 14. For each image just add the string `14 PR 0 0 1 1` where `PR = prob(no finding)` from your classifier. If your classifier is perfect, then you will score full `mAP = 1.0` for class 14. Most classifiers have CV AUC 0.99+, so most teams will maximize this class mAP\n",
    "1262621": "Brilliant! From the melonoma classification, I have learned a lot from your discussion. thank you\n\nI have a question why your team trained the models with 1024 but inferred with lower resolutions?",
    "1263240": "Hi @cdeotte , thank for sharing your post-processing approach. I have a question. You said the accuracy will increase if the boxes confidence score increase. In the Figure 2, how about if all confidence score all increases 0.1, so all the position from A-G still remains same as Figure 1 right, so is it make any change different from Figure 1?\n",
    "1259605": "Awesome post-processing approach... as always :D \n\nJust one question: do you think this approach is more robust and just generally works better than simply training with some empty images, and not using classifiers to calibrate confidence scores for each box? ",
    "1258818": "Very well explained @cdeotte !  Congrats once again !",
    "1258729": "Congrats! thanks for sharing insights about competition metric in detailed explanation.\n\nI could not understand how to calibrate probability in \"Order2\" way, isn't it possible to calibrate only D & G class without knowing which is the true class??",
    "1258659": "Congrats! I admire your post-processing very much.👍",
    "1258221": "Congrats and thanks for sharing your details! I have learned a lot from you. Thanks you.",
    "1257893": "Congrats!! And thanks for the sharing in details, you always did a very fascinating job in post-process ",
    "1257822": "Congratulations on the standing! It is people like you and your team whom really inspire beginner like myself to improve and become better. To have that intuition and experience to apply theories i never thought i could apply out of school ( Bayesian statistics )  👍",
    "1257801": "Congrats for the good models and stable validation results! I like the postprocessing ideas!\n\nI wonder if you tested your Class 14 predictions separately by just submitting the 1 pixel boxes what would be the LB AP score?",
    "1257785": "Thanks for sharing ur great kernel! BTW, I have some questions here.\n\n1.  What`s the purpose of 15 class classifiers? \n2. What`s the meaning of cv on \"class14\"? Is it a model that distinguishes between find/nofind? \n\nAlso, the calibrated confidence is genius. For score ordering, I used sum-based nms method that sums up scores of overlapping boxes that predicted by ensemble models.\n\nThanks again !",
    "1257780": "Congratulation on your final result. May I ask how many folds did you use for CV and did you stratify that or just randomly split the folds",
    "1257750": "Congrats on 7th place 👍",
    "1257719": "Congrats on 7th place @cdeotte and team",
    "1267101": "`already had full pipelines to train VFNet, YoloV5, EfficientDet4 and more! `\nIt would be really beneficial for beginners if you shred your pipelines.\nTo be honest, I tried to code a whole pipeline on my own but was never successful, hope to learn from your codes\nit is better if its in PyTorch ; )",
    "1258726": "@cdeotte any reason for `0.1` as power?",
    "1258566": "Congratz ! \n\n@cdeotte They should give you the Post-processing GM rank, I'm pretty sure that's the 5th gold you get with such ideas :)",
    "1258018": "Congrats on the result. We also used the probability calibration technique in post-processing . Our calibrated confidence was calculated as: confidence score=(prob(class given finding=True)^2)∗prob(finding=True)\nWe missed the 15 class classifier part!! That was actually a very good idea !!",
    "1257957": "Congrats!    ",
    "1257902": "Congratulations @cdeotte , for detection, did you use 14 or 15 class training?",
    "1322642": "Hi, @cdeotte we are all fan of your post-processing techniques. Most of them are based on Bayesian methods. Can you share some good resources from where you have studied this.\nThanks"
  }
}