{
  "id": 212287,
  "title": "Explanation of scoring metric (mAP@0.4)",
  "url": "/competitions/vinbigdata-chest-xray-abnormalities-detection/discussion/212287",
  "author_name": "Peter",
  "post_date": "2021-01-18T11:06:18.825000",
  "votes": 99,
  "comment_count": 25,
  "views": 0,
  "content": "<h1>Mean Average Precision</h1>\n<blockquote>\n  <p>The challenge uses the standard PASCAL VOC 2010 mean Average Precision (mAP) at IoU &gt; 0.4.</p>\n</blockquote>\n<h2>Basics</h2>\n<h3>Precision and Recall</h3>\n<p><strong>Precision</strong>: What proportion of the positive identifications (predicted boxes) was actually correct?<br>\nPrecision = TP/(TP + FP)</p>\n<p><strong>Recall</strong> What proportion of actual positives was identified correctly?<br>\nRecall = TP / (TP + FN)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F864684%2Fb3a58f56a1bfdbc4841dfd47747193ef%2Fprecision_recall_tp_fp.png?generation=1610967621000417&amp;alt=media\" alt=\"\"></p>\n<p>TP: True positives. You predict, and it is correct.<br>\nFP: False positives. You predict, but it is wrong.<br>\nFN: False negatives. You did not predict, and it is wrong. You should have predicted.</p>\n<p>There are no true negatives (TN) in object detection. </p>\n<h3>Intersection over union (IoU)</h3>\n<p>IoU measures the overlap between 2 boxes. <br>\nIoU = (area of overlap) / (area of union)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F864684%2F3a82ceb91fbb96d5154f6b15062804df%2Fiou.jpg?generation=1610967851297996&amp;alt=media\" alt=\"\"></p>\n<h4>IoU @0.4</h4>\n<p>IoU at 0.4 means that we consider a predicted box as a true positive (TP) if the IoU of this box and one of the GT box is greater than 0.4.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F864684%2F50b91585a79d9e0ecb24d874e78be99e%2Fiou_examples.jpg?generation=1610967875651682&amp;alt=media\" alt=\"\"></p>\n<h3>Score</h3>\n<p>For this metric, we have to add a confidence score for our predictions. It acts as a tie-breaker. For example, if there are two or more predicted boxes with IoU&gt;0.4 with the same GT, the metric will consider our most confident (highest score) prediction as TP for that GT-Pred pair.</p>\n<h2>How to determine TP, FP, FN</h2>\n<ul>\n<li>Sort the predicted boxes in descendent (by score)</li>\n<li>Pick the most confident prediction</li>\n<li>Calculate the IoU between the selected prediction and <strong>all</strong> of the GT boxes</li>\n<li>Select the highest IoU</li>\n<li>If the IoU is greater than our threshold (0.4), we have a true positive (TP) match. In this case, we remove the matched boxes from both the predicted and the GT list of boxes.</li>\n<li>Repeat 2-4</li>\n<li>After we iterate through in our predicted boxes, what is left in the predicted list are the false positives (FP) and what is left in the GT list are the false negatives (FN).</li>\n</ul>\n<p><em>All of these are per class!!</em></p>\n<h2>Average Precision</h2>\n<table>\n<thead>\n<tr>\n<th>Rank</th>\n<th>Correct (TP)?</th>\n<th>Precision</th>\n<th>Recall</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>True</td>\n<td>1.0</td>\n<td>0.2</td>\n</tr>\n<tr>\n<td>2</td>\n<td>True</td>\n<td>1.0</td>\n<td>0.4</td>\n</tr>\n<tr>\n<td>3</td>\n<td>False</td>\n<td>0.67</td>\n<td>0.4</td>\n</tr>\n<tr>\n<td>4</td>\n<td>False</td>\n<td>0.5</td>\n<td>0.4</td>\n</tr>\n<tr>\n<td>5</td>\n<td>False</td>\n<td>0.4</td>\n<td>0.4</td>\n</tr>\n<tr>\n<td>6</td>\n<td>True</td>\n<td>0.5</td>\n<td>0.6</td>\n</tr>\n<tr>\n<td>7</td>\n<td>True</td>\n<td>0.57</td>\n<td>0.8</td>\n</tr>\n<tr>\n<td>8</td>\n<td>False</td>\n<td>0.5</td>\n<td>0.8</td>\n</tr>\n<tr>\n<td>9</td>\n<td>False</td>\n<td>0.44</td>\n<td>0.8</td>\n</tr>\n<tr>\n<td>10</td>\n<td>True</td>\n<td>0.5</td>\n<td>1.0</td>\n</tr>\n</tbody>\n</table>\n<p>In the table above the rows are our predictions. Let's assume we have 5 GT boxes in this example. So you have to iterate over your prediction and calculate Precisions and recalls after every row.</p>\n<p>Row #1, you were correct. So far, we have: TP=1, FP=0, FN=4; Precision = 1/(1+0)=1, Recall = 1/(1+4) = 0.2<br>\nRow #2, you were correct. So far, we have: TP=2, FP=0, FN=3; Precision = 2/(2+0)=1, Recall = 2/(2+3) = 0.4<br>\nRow #3, you were incorrect. So far, we have: TP=2, FP=1, FN=3; Precision = 2/(2+1)=0.67, Recall = 2/(2+3)=0.4<br>\nand so on…</p>\n<p>Note: The recall value is always increasing (or stays the same). You can plot this: precision-recall curve, where the Recall is the x-axis, Precision is y.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F864684%2F667d2e7cf09649a352aaeeee7ad27fee%2Fprecision_recall_curve.png?generation=1610967690204686&amp;alt=media\" alt=\"\"></p>\n<p>The Average Precision (AP) is the area under the precision-recall curve. We could integrate this, but fortunately, we don't have to. COCOEval approximates this by divided the recall-axis into 100 intervals and take the highest Precision at every interval. Kind of Riemann sum.</p>\n<h2>Healthy samples</h2>\n<p>By default, object detection can not deal with true negatives. You can not predict a box for an object that is not there. The host's solution is this 1x1 pixel at the corner. From the AP point of view, it is just another box, and the metric handles these boxes as the same.</p>\n<p>You can experiment with this fact. For example, if your model predicts one uncertain box with 0.51 confidence score, that means there is a chance that the sample is healthy. It might improve your overall score if you add \"14 1 0 0 1 1\" to your prediction.<br>\nMore details in this <a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/211971\" target=\"_blank\">discussion</a></p>\n<h2>Calculation</h2>\n<p>See my <a href=\"https://www.kaggle.com/pestipeti/competition-metric-map-0-4\" target=\"_blank\">Competition Metric Calculator</a> notebook.</p>\n<h2>References</h2>\n<ul>\n<li><a href=\"https://jonathan-hui.medium.com/map-mean-average-precision-for-object-detection-45c121a31173\" target=\"_blank\">mAP (mean Average Precision) for Object Detection</a> by Jonathan Hui</li>\n<li><a href=\"https://blog.paperspace.com/mean-average-precision/\" target=\"_blank\">Evaluating Object Detection Models Using Mean Average Precision (mAP)</a> by Ahmed Fawzy Gad - Paperspace</li>\n<li><a href=\"https://en.wikipedia.org/wiki/Precision_and_recall\" target=\"_blank\">Precision and Recall</a> - Wikipedia</li>\n</ul>",
  "messages": [
    {
      "id": 1158063,
      "postDate": "2021-01-18T11:06:18.827Z",
      "content": "<h1>Mean Average Precision</h1>\n<blockquote>\n  <p>The challenge uses the standard PASCAL VOC 2010 mean Average Precision (mAP) at IoU &gt; 0.4.</p>\n</blockquote>\n<h2>Basics</h2>\n<h3>Precision and Recall</h3>\n<p><strong>Precision</strong>: What proportion of the positive identifications (predicted boxes) was actually correct?<br>\nPrecision = TP/(TP + FP)</p>\n<p><strong>Recall</strong> What proportion of actual positives was identified correctly?<br>\nRecall = TP / (TP + FN)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F864684%2Fb3a58f56a1bfdbc4841dfd47747193ef%2Fprecision_recall_tp_fp.png?generation=1610967621000417&amp;alt=media\" alt=\"\"></p>\n<p>TP: True positives. You predict, and it is correct.<br>\nFP: False positives. You predict, but it is wrong.<br>\nFN: False negatives. You did not predict, and it is wrong. You should have predicted.</p>\n<p>There are no true negatives (TN) in object detection. </p>\n<h3>Intersection over union (IoU)</h3>\n<p>IoU measures the overlap between 2 boxes. <br>\nIoU = (area of overlap) / (area of union)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F864684%2F3a82ceb91fbb96d5154f6b15062804df%2Fiou.jpg?generation=1610967851297996&amp;alt=media\" alt=\"\"></p>\n<h4>IoU @0.4</h4>\n<p>IoU at 0.4 means that we consider a predicted box as a true positive (TP) if the IoU of this box and one of the GT box is greater than 0.4.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F864684%2F50b91585a79d9e0ecb24d874e78be99e%2Fiou_examples.jpg?generation=1610967875651682&amp;alt=media\" alt=\"\"></p>\n<h3>Score</h3>\n<p>For this metric, we have to add a confidence score for our predictions. It acts as a tie-breaker. For example, if there are two or more predicted boxes with IoU&gt;0.4 with the same GT, the metric will consider our most confident (highest score) prediction as TP for that GT-Pred pair.</p>\n<h2>How to determine TP, FP, FN</h2>\n<ul>\n<li>Sort the predicted boxes in descendent (by score)</li>\n<li>Pick the most confident prediction</li>\n<li>Calculate the IoU between the selected prediction and <strong>all</strong> of the GT boxes</li>\n<li>Select the highest IoU</li>\n<li>If the IoU is greater than our threshold (0.4), we have a true positive (TP) match. In this case, we remove the matched boxes from both the predicted and the GT list of boxes.</li>\n<li>Repeat 2-4</li>\n<li>After we iterate through in our predicted boxes, what is left in the predicted list are the false positives (FP) and what is left in the GT list are the false negatives (FN).</li>\n</ul>\n<p><em>All of these are per class!!</em></p>\n<h2>Average Precision</h2>\n<table>\n<thead>\n<tr>\n<th>Rank</th>\n<th>Correct (TP)?</th>\n<th>Precision</th>\n<th>Recall</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>True</td>\n<td>1.0</td>\n<td>0.2</td>\n</tr>\n<tr>\n<td>2</td>\n<td>True</td>\n<td>1.0</td>\n<td>0.4</td>\n</tr>\n<tr>\n<td>3</td>\n<td>False</td>\n<td>0.67</td>\n<td>0.4</td>\n</tr>\n<tr>\n<td>4</td>\n<td>False</td>\n<td>0.5</td>\n<td>0.4</td>\n</tr>\n<tr>\n<td>5</td>\n<td>False</td>\n<td>0.4</td>\n<td>0.4</td>\n</tr>\n<tr>\n<td>6</td>\n<td>True</td>\n<td>0.5</td>\n<td>0.6</td>\n</tr>\n<tr>\n<td>7</td>\n<td>True</td>\n<td>0.57</td>\n<td>0.8</td>\n</tr>\n<tr>\n<td>8</td>\n<td>False</td>\n<td>0.5</td>\n<td>0.8</td>\n</tr>\n<tr>\n<td>9</td>\n<td>False</td>\n<td>0.44</td>\n<td>0.8</td>\n</tr>\n<tr>\n<td>10</td>\n<td>True</td>\n<td>0.5</td>\n<td>1.0</td>\n</tr>\n</tbody>\n</table>\n<p>In the table above the rows are our predictions. Let's assume we have 5 GT boxes in this example. So you have to iterate over your prediction and calculate Precisions and recalls after every row.</p>\n<p>Row #1, you were correct. So far, we have: TP=1, FP=0, FN=4; Precision = 1/(1+0)=1, Recall = 1/(1+4) = 0.2<br>\nRow #2, you were correct. So far, we have: TP=2, FP=0, FN=3; Precision = 2/(2+0)=1, Recall = 2/(2+3) = 0.4<br>\nRow #3, you were incorrect. So far, we have: TP=2, FP=1, FN=3; Precision = 2/(2+1)=0.67, Recall = 2/(2+3)=0.4<br>\nand so on…</p>\n<p>Note: The recall value is always increasing (or stays the same). You can plot this: precision-recall curve, where the Recall is the x-axis, Precision is y.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F864684%2F667d2e7cf09649a352aaeeee7ad27fee%2Fprecision_recall_curve.png?generation=1610967690204686&amp;alt=media\" alt=\"\"></p>\n<p>The Average Precision (AP) is the area under the precision-recall curve. We could integrate this, but fortunately, we don't have to. COCOEval approximates this by divided the recall-axis into 100 intervals and take the highest Precision at every interval. Kind of Riemann sum.</p>\n<h2>Healthy samples</h2>\n<p>By default, object detection can not deal with true negatives. You can not predict a box for an object that is not there. The host's solution is this 1x1 pixel at the corner. From the AP point of view, it is just another box, and the metric handles these boxes as the same.</p>\n<p>You can experiment with this fact. For example, if your model predicts one uncertain box with 0.51 confidence score, that means there is a chance that the sample is healthy. It might improve your overall score if you add \"14 1 0 0 1 1\" to your prediction.<br>\nMore details in this <a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/211971\" target=\"_blank\">discussion</a></p>\n<h2>Calculation</h2>\n<p>See my <a href=\"https://www.kaggle.com/pestipeti/competition-metric-map-0-4\" target=\"_blank\">Competition Metric Calculator</a> notebook.</p>\n<h2>References</h2>\n<ul>\n<li><a href=\"https://jonathan-hui.medium.com/map-mean-average-precision-for-object-detection-45c121a31173\" target=\"_blank\">mAP (mean Average Precision) for Object Detection</a> by Jonathan Hui</li>\n<li><a href=\"https://blog.paperspace.com/mean-average-precision/\" target=\"_blank\">Evaluating Object Detection Models Using Mean Average Precision (mAP)</a> by Ahmed Fawzy Gad - Paperspace</li>\n<li><a href=\"https://en.wikipedia.org/wiki/Precision_and_recall\" target=\"_blank\">Precision and Recall</a> - Wikipedia</li>\n</ul>",
      "rawMarkdown": "# Mean Average Precision\n> The challenge uses the standard PASCAL VOC 2010 mean Average Precision (mAP) at IoU > 0.4.\n\n\n## Basics\n\n### Precision and Recall\n**Precision**: What proportion of the positive identifications (predicted boxes) was actually correct?\nPrecision = TP/(TP + FP)\n\n**Recall** What proportion of actual positives was identified correctly?\nRecall = TP / (TP + FN)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F864684%2Fb3a58f56a1bfdbc4841dfd47747193ef%2Fprecision_recall_tp_fp.png?generation=1610967621000417&alt=media)\n\nTP: True positives. You predict, and it is correct.\nFP: False positives. You predict, but it is wrong.\nFN: False negatives. You did not predict, and it is wrong. You should have predicted.\n\nThere are no true negatives (TN) in object detection. \n\n### Intersection over union (IoU)\nIoU measures the overlap between 2 boxes. \nIoU = (area of overlap) / (area of union)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F864684%2F3a82ceb91fbb96d5154f6b15062804df%2Fiou.jpg?generation=1610967851297996&alt=media)\n\n#### IoU \\@0.4\nIoU at 0.4 means that we consider a predicted box as a true positive (TP) if the IoU of this box and one of the GT box is greater than 0.4.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F864684%2F50b91585a79d9e0ecb24d874e78be99e%2Fiou_examples.jpg?generation=1610967875651682&alt=media)\n\n### Score\nFor this metric, we have to add a confidence score for our predictions. It acts as a tie-breaker. For example, if there are two or more predicted boxes with IoU>0.4 with the same GT, the metric will consider our most confident (highest score) prediction as TP for that GT-Pred pair.\n\n## How to determine TP, FP, FN\n- Sort the predicted boxes in descendent (by score)\n- Pick the most confident prediction\n- Calculate the IoU between the selected prediction and **all** of the GT boxes\n- Select the highest IoU\n- If the IoU is greater than our threshold (0.4), we have a true positive (TP) match. In this case, we remove the matched boxes from both the predicted and the GT list of boxes.\n- Repeat 2-4\n- After we iterate through in our predicted boxes, what is left in the predicted list are the false positives (FP) and what is left in the GT list are the false negatives (FN).\n\n*All of these are per class!!*\n\n## Average Precision\n| Rank | Correct (TP)? | Precision | Recall |\n| ---- | ------------- | --------- | ------ |\n| 1    | True          | 1.0       | 0.2    |\n| 2    | True          | 1.0       | 0.4    |\n| 3    | False         | 0.67      | 0.4    |\n| 4    | False         | 0.5       | 0.4    |\n| 5    | False         | 0.4       | 0.4    |\n| 6    | True          | 0.5       | 0.6    |\n| 7    | True          | 0.57      | 0.8    |\n| 8    | False         | 0.5       | 0.8    |\n| 9    | False         | 0.44      | 0.8    |\n| 10   | True          | 0.5       | 1.0    |\n\nIn the table above the rows are our predictions. Let's assume we have 5 GT boxes in this example. So you have to iterate over your prediction and calculate Precisions and recalls after every row.\n\nRow #1, you were correct. So far, we have: TP=1, FP=0, FN=4; Precision = 1/(1+0)=1, Recall = 1/(1+4) = 0.2\nRow #2, you were correct. So far, we have: TP=2, FP=0, FN=3; Precision = 2/(2+0)=1, Recall = 2/(2+3) = 0.4\nRow #3, you were incorrect. So far, we have: TP=2, FP=1, FN=3; Precision = 2/(2+1)=0.67, Recall = 2/(2+3)=0.4\nand so on...\n\nNote: The recall value is always increasing (or stays the same). You can plot this: precision-recall curve, where the Recall is the x-axis, Precision is y.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F864684%2F667d2e7cf09649a352aaeeee7ad27fee%2Fprecision_recall_curve.png?generation=1610967690204686&alt=media)\n\nThe Average Precision (AP) is the area under the precision-recall curve. We could integrate this, but fortunately, we don't have to. COCOEval approximates this by divided the recall-axis into 100 intervals and take the highest Precision at every interval. Kind of Riemann sum.\n\n\n## Healthy samples\nBy default, object detection can not deal with true negatives. You can not predict a box for an object that is not there. The host's solution is this 1x1 pixel at the corner. From the AP point of view, it is just another box, and the metric handles these boxes as the same.\n\nYou can experiment with this fact. For example, if your model predicts one uncertain box with 0.51 confidence score, that means there is a chance that the sample is healthy. It might improve your overall score if you add \"14 1 0 0 1 1\" to your prediction.\nMore details in this [discussion](https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/211971)\n\n## Calculation\nSee my [Competition Metric Calculator](https://www.kaggle.com/pestipeti/competition-metric-map-0-4) notebook.\n\n## References\n- [mAP (mean Average Precision) for Object Detection](https://jonathan-hui.medium.com/map-mean-average-precision-for-object-detection-45c121a31173) by Jonathan Hui\n- [Evaluating Object Detection Models Using Mean Average Precision (mAP)](https://blog.paperspace.com/mean-average-precision/) by Ahmed Fawzy Gad - Paperspace\n- [Precision and Recall](https://en.wikipedia.org/wiki/Precision_and_recall) - Wikipedia\n\n\n\n\n\n\n\n\n",
      "votes": 99
    },
    {
      "id": 1187577,
      "postDate": "2021-02-05T14:44:00.710Z",
      "content": "<p>Nice wrapup!</p>\n<p>I'd like to add some important mAP features.</p>\n<p>1, The lower the confidence threshold, the higher the mAP we will get.<br>\nSo we should use as low a confidence threshold as possible.</p>\n<p>2, Class dependence is linear.</p>\n<pre>mAP(all class) = mAP(class_id == n) + mAP(class_id != n)\n</pre>\n<p>So you can check the scores for each class individually.</p>\n<p>3, Small sample class is equaly important as bit class.</p>\n<pre>max(mAP(small class)) == max(mAP(big class))\n</pre>\n<p>If mAP is mean \"waighted\" average precision, this would be not true.</p>\n<p>For more detail, I made a notebook about this.<br>\n<a href=\"https://www.kaggle.com/its7171/a-few-tips-on-map?scriptVersionId=53590041\" target=\"_blank\">https://www.kaggle.com/its7171/a-few-tips-on-map?scriptVersionId=53590041</a></p>\n<p>I hope this information is useful especially for beginners.</p>",
      "rawMarkdown": "Nice wrapup!\n\nI'd like to add some important mAP features.\n\n1, The lower the confidence threshold, the higher the mAP we will get.\nSo we should use as low a confidence threshold as possible.\n\n2, Class dependence is linear.\n<pre>\nmAP(all class) = mAP(class_id == n) + mAP(class_id != n)\n</pre>\nSo you can check the scores for each class individually.\n\n3, Small sample class is equaly important as bit class.\n<pre>\nmax(mAP(small class)) == max(mAP(big class))\n</pre>\nIf mAP is mean \"waighted\" average precision, this would be not true.\n\nFor more detail, I made a notebook about this.\nhttps://www.kaggle.com/its7171/a-few-tips-on-map?scriptVersionId=53590041\n\nI hope this information is useful especially for beginners.",
      "votes": 7,
      "replies": [
        {
          "id": 1187652,
          "postDate": "2021-02-05T15:55:20.783Z",
          "content": "<p>Very useful, thanks for sharing!</p>\n<p>With respect to 1.: Is this a weakness of the metric? A prediction of something like 100 of objects per image (which you get with a very low confidence threshold) doesn't seem very useful.. </p>\n<p>PS. The link is not working.</p>",
          "rawMarkdown": "Very useful, thanks for sharing!\n\nWith respect to 1.: Is this a weakness of the metric? A prediction of something like 100 of objects per image (which you get with a very low confidence threshold) doesn't seem very useful.. \n\nPS. The link is not working.",
          "votes": 1
        },
        {
          "id": 1187682,
          "postDate": "2021-02-05T16:19:40.163Z",
          "content": "<p>I think so.<br>\nIt might be a good idea to add some restrictions such as reasonable submission file size limitation.</p>",
          "rawMarkdown": "I think so.\nIt might be a good idea to add some restrictions such as reasonable submission file size limitation.",
          "votes": 1
        },
        {
          "id": 1187713,
          "postDate": "2021-02-05T16:47:16.397Z",
          "content": "<p>However, it is an advantage that we dont have to worry about tuning the confidence threshold because we can simply set the smallest possible threshold.</p>\n<p>On the other hand, this threshold setting can be very critical for other metrics .<br>\nIn some cases, this threshold setting can make the competition a lottery.</p>",
          "rawMarkdown": "However, it is an advantage that we dont have to worry about tuning the confidence threshold because we can simply set the smallest possible threshold.\n\nOn the other hand, this threshold setting can be very critical for other metrics .\nIn some cases, this threshold setting can make the competition a lottery.",
          "votes": 1
        },
        {
          "id": 1188110,
          "postDate": "2021-02-06T01:03:27.173Z",
          "content": "<p>Agreed. Or maybe a maximum number per image. </p>\n<p>It's a bit a downer, because you would like to produce realistic, useful predictions, but the metric counteracts this objective.</p>",
          "rawMarkdown": "Agreed. Or maybe a maximum number per image. \n\nIt's a bit a downer, because you would like to produce realistic, useful predictions, but the metric counteracts this objective.",
          "votes": 1
        },
        {
          "id": 1230981,
          "postDate": "2021-03-08T15:33:58.397Z",
          "content": "<blockquote>\n  <p>1, The lower the confidence threshold, the higher the mAP we will get.<br>\n  So we should use as low a confidence threshold as possible.</p>\n</blockquote>\n<p>I did a test, submission with conf_tresh: 0.01 and other submission with conf_tresh: 0.001, but my lb score drop with conf_tresh: 0.001 , what u think about ? </p>",
          "rawMarkdown": "> 1, The lower the confidence threshold, the higher the mAP we will get.\n> So we should use as low a confidence threshold as possible.\n\n\nI did a test, submission with conf_tresh: 0.01 and other submission with conf_tresh: 0.001, but my lb score drop with conf_tresh: 0.001 , what u think about ? "
        },
        {
          "id": 1231042,
          "postDate": "2021-03-08T16:15:24.380Z",
          "content": "<p>That is very strange. For me it improved till 0.001, further reduction didn't lead to an improvement, but also didn't hurt. That is expected because we are very much at the right of the precision-recall curve so there is not much room for improvement.</p>",
          "rawMarkdown": "That is very strange. For me it improved till 0.001, further reduction didn't lead to an improvement, but also didn't hurt. That is expected because we are very much at the right of the precision-recall curve so there is not much room for improvement."
        },
        {
          "id": 1360126,
          "postDate": "2021-06-21T20:36:31.583Z",
          "content": "<p><a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> </p>\n<p>Thanks for your added points.</p>\n<p>I have a question regarding the low threshold. Why does lowering the threshold generally increase mAP? Wouldn't lowering the threshold generally introduce a lot more false positives?</p>",
          "rawMarkdown": "@its7171 \n\nThanks for your added points.\n\nI have a question regarding the low threshold. Why does lowering the threshold generally increase mAP? Wouldn't lowering the threshold generally introduce a lot more false positives?"
        }
      ]
    },
    {
      "id": 1160070,
      "postDate": "2021-01-19T16:45:05.300Z",
      "content": "<p>Very clear and helpful, thank you !</p>",
      "rawMarkdown": "Very clear and helpful, thank you !",
      "votes": 1,
      "replies": [
        {
          "id": 1206727,
          "postDate": "2021-02-17T14:16:17.390Z",
          "content": "<p>You are very welcome my friend. We are all set now, let's do it ! 👌</p>",
          "rawMarkdown": "You are very welcome my friend. We are all set now, let's do it ! 👌"
        }
      ]
    },
    {
      "id": 1159505,
      "postDate": "2021-01-19T10:12:50.270Z",
      "content": "<p>Thank you very much for the clarification !</p>",
      "rawMarkdown": "Thank you very much for the clarification !",
      "votes": 1
    },
    {
      "id": 1158812,
      "postDate": "2021-01-18T19:42:46.343Z",
      "content": "<p><a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> thanks for summarizing this. I was reading a few other resources regarding mAP but when I got to calculating the precision-recall curve I got a bit turned around. Your example showing the calculation along with the table helped clarify that immensely. Thanks for posting!</p>",
      "rawMarkdown": "@pestipeti thanks for summarizing this. I was reading a few other resources regarding mAP but when I got to calculating the precision-recall curve I got a bit turned around. Your example showing the calculation along with the table helped clarify that immensely. Thanks for posting!",
      "votes": 1
    },
    {
      "id": 1251211,
      "postDate": "2021-03-24T15:22:53.993Z",
      "content": "<p>Thank you for the explanation! I have a clarification with regards to the following scenario: supposed an image have 2 GT box (class 2 and 3). There is a predicted box (class 3), but IOU &gt; defined threshold only with the class 2 GT box. Im not sure if i fully comprehend the mAP. Supposed you loop through the GT classes, it will be a FN on class 2, but it also appears as a FP on class 3?</p>",
      "rawMarkdown": "Thank you for the explanation! I have a clarification with regards to the following scenario: supposed an image have 2 GT box (class 2 and 3). There is a predicted box (class 3), but IOU > defined threshold only with the class 2 GT box. Im not sure if i fully comprehend the mAP. Supposed you loop through the GT classes, it will be a FN on class 2, but it also appears as a FP on class 3?"
    },
    {
      "id": 1197310,
      "postDate": "2021-02-12T05:19:46.137Z",
      "content": "<p>Thanks for sharing the metric details <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> !</p>",
      "rawMarkdown": "Thanks for sharing the metric details @pestipeti !"
    },
    {
      "id": 1161853,
      "postDate": "2021-01-20T19:26:04.583Z",
      "content": "<p>Thank you for the write-up and notebook.<br>\nI am not sure if your notebook returns the correct result, as** I'd expect a precision of 1** given prediction == annotation.<br>\nCould you provide a working sample?<br>\nAlso, why do you need to remove duplicates?</p>\n<p>Thank you!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3115179%2F567f0397b1ac4651a0d7ae86f50d4ee9%2FDeepinScreenshot_select-area_20210120202345.png?generation=1611170649011001&amp;alt=media\" alt=\"\"></p>\n<p>class_names = [\"image_id\", \"class_name\", \"class_id\", \"rad_id\", \"x_min\", \"y_min\",  \"x_max\",  \"y_max\"]<br>\ndf_sample = pd.DataFrame(dict(zip(class_names,<br>\n                                  [[\"0\", \"1\"], <br>\n                                   [\"No-Finding\", \"No-Finding\"], <br>\n                                   [14, 14], [7, 7], [0., 1.], [0., 1.], [1., 1.], [1., 1.]])))<br>\ndisplay(df_sample)<br>\nvineval_sample = VinBigDataEval(df_sample)<br>\npred_df = pd.DataFrame({\"image_id\": [\"0\", \"1\"],<br>\n                        \"PredictionString\": [\"14 1.0 0 0 1 1\",<br>\n                                             \"14 1.0 1 1 1 1\",<br>\n                                             ]})<br>\ndisplay(pred_df)<br>\ncocoEvalRes = vineval_sample.evaluate(pred_df, n_imgs=-1)</p>",
      "rawMarkdown": "Thank you for the write-up and notebook.\nI am not sure if your notebook returns the correct result, as** I'd expect a precision of 1** given prediction == annotation.\nCould you provide a working sample?\nAlso, why do you need to remove duplicates?\n\nThank you!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3115179%2F567f0397b1ac4651a0d7ae86f50d4ee9%2FDeepinScreenshot_select-area_20210120202345.png?generation=1611170649011001&alt=media)\n\nclass_names = [\"image_id\", \"class_name\", \"class_id\", \"rad_id\", \"x_min\", \"y_min\",  \"x_max\",  \"y_max\"]\ndf_sample = pd.DataFrame(dict(zip(class_names,\n                                  [[\"0\", \"1\"], \n                                   [\"No-Finding\", \"No-Finding\"], \n                                   [14, 14], [7, 7], [0., 1.], [0., 1.], [1., 1.], [1., 1.]])))\ndisplay(df_sample)\nvineval_sample = VinBigDataEval(df_sample)\npred_df = pd.DataFrame({\"image_id\": [\"0\", \"1\"],\n                        \"PredictionString\": [\"14 1.0 0 0 1 1\",\n                                             \"14 1.0 1 1 1 1\",\n                                             ]})\ndisplay(pred_df)\ncocoEvalRes = vineval_sample.evaluate(pred_df, n_imgs=-1)\n\n\n",
      "replies": [
        {
          "id": 1161881,
          "postDate": "2021-01-20T19:41:40.843Z",
          "content": "<p>The script is only a wrapper for COCOEval. It just converts our data to their format. It does not calculate the score. I've noticed the same result and it seems that COCOEval has a small bug. It does not count the first true positive sample. If you add more samples (TPs) to the calculation it will work. The one wrongly handled TP will be there, but its weight in the overall score would be less as you increase the number of samples.</p>",
          "rawMarkdown": "The script is only a wrapper for COCOEval. It just converts our data to their format. It does not calculate the score. I've noticed the same result and it seems that COCOEval has a small bug. It does not count the first true positive sample. If you add more samples (TPs) to the calculation it will work. The one wrongly handled TP will be there, but its weight in the overall score would be less as you increase the number of samples.",
          "votes": 1
        },
        {
          "id": 1162611,
          "postDate": "2021-01-21T08:31:44.660Z",
          "content": "<p>Speaking of duplicates, they have to be removed because the image id in your predictions file has to match the exact id in the ground truth. In the gt there are multiple rows with the same id, so one pred image id  will match to multiple gt rows and will therefore output an array with multiple bboxes. This format can cause  issues specifically with COCOEval, so duplicates have to be removed.</p>\n<p>However, <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> I am concerned about removing duplicates - this way the predicted box is evaluated only against gt from one radiologist without considering labelled images from other radiologists…</p>",
          "rawMarkdown": "Speaking of duplicates, they have to be removed because the image id in your predictions file has to match the exact id in the ground truth. In the gt there are multiple rows with the same id, so one pred image id  will match to multiple gt rows and will therefore output an array with multiple bboxes. This format can cause  issues specifically with COCOEval, so duplicates have to be removed.\n\nHowever, @pestipeti I am concerned about removing duplicates - this way the predicted box is evaluated only against gt from one radiologist without considering labelled images from other radiologists...\n",
          "votes": 2
        },
        {
          "id": 1162723,
          "postDate": "2021-01-21T09:49:03.033Z",
          "content": "<p><a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> thanks for your explanation.<br>\nI repeated the experiment for up to 200 samples and the error seems to converge to zero.<br>\nStill, it's a strange behavior. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3115179%2F6a711c0daf5e05cde69c293d0f092098%2FDeepinScreenshot_select-area_20210121104302.png?generation=1611222240444020&amp;alt=media\" alt=\"\"></p>\n<p>Regarding the duplicates:<br>\nIf I understood correctly, the mAP looks for the GT bounding box with the highest IoU. <br>\nHence, there should be only problems for <em>exact</em> duplicates, i.e. duplicate bounding boxes in an image.</p>",
          "rawMarkdown": "@pestipeti thanks for your explanation.\nI repeated the experiment for up to 200 samples and the error seems to converge to zero.\nStill, it's a strange behavior. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3115179%2F6a711c0daf5e05cde69c293d0f092098%2FDeepinScreenshot_select-area_20210121104302.png?generation=1611222240444020&alt=media)\n\nRegarding the duplicates:\nIf I understood correctly, the mAP looks for the GT bounding box with the highest IoU. \nHence, there should be only problems for *exact* duplicates, i.e. duplicate bounding boxes in an image.",
          "votes": 2
        },
        {
          "id": 1162740,
          "postDate": "2021-01-21T10:04:37.983Z",
          "content": "<p>I looked into this issue in further detail, and yes this is what I found as well. Removing duplicates is OK only in the case of identical bounding boxes for the same image id. In other cases, iou needs to be calculated for all of them and the highest selected.</p>",
          "rawMarkdown": "I looked into this issue in further detail, and yes this is what I found as well. Removing duplicates is OK only in the case of identical bounding boxes for the same image id. In other cases, iou needs to be calculated for all of them and the highest selected.",
          "votes": 1
        },
        {
          "id": 1162760,
          "postDate": "2021-01-21T10:27:36.253Z",
          "content": "<p>You don't necessarily have to remove duplicates; COCOEval handles them. The key is the order of your prediction confidence (score).<br>\nIf you have one GT and your model predicts two boxes (same xywh), then it will match (assuming IoU&gt;0.4) only with the highest score. If it is a match, the metric will remove the boxes (one GT and <strong>one</strong> predicted) and the algorithm repeats. You will have one TP and one FP (the other predicted without a matching pair).</p>\n<p>I haven't checked the details of this competition, I am working on the Rainforest. I assumed that in the test set we don't have multiple predictions from different radiologists.</p>\n<p><a href=\"https://www.kaggle.com/indswetrust\" target=\"_blank\">@indswetrust</a><br>\nYou are almost right. You can safely remove duplicates if the boxes are for the same image, <strong>the same class</strong>, and they are \"close enough.\" </p>\n<p><a href=\"https://www.kaggle.com/fold10\" target=\"_blank\">@fold10</a></p>\n<blockquote>\n  <p>If I understood correctly, the mAP looks for the GT bounding box with the highest IoU.</p>\n</blockquote>\n<p>Not exactly. It looks for IoU &gt; 0.4 and takes the highest confidence level/score.</p>",
          "rawMarkdown": "You don't necessarily have to remove duplicates; COCOEval handles them. The key is the order of your prediction confidence (score).\nIf you have one GT and your model predicts two boxes (same xywh), then it will match (assuming IoU>0.4) only with the highest score. If it is a match, the metric will remove the boxes (one GT and **one** predicted) and the algorithm repeats. You will have one TP and one FP (the other predicted without a matching pair).\n\nI haven't checked the details of this competition, I am working on the Rainforest. I assumed that in the test set we don't have multiple predictions from different radiologists.\n\n@indswetrust\nYou are almost right. You can safely remove duplicates if the boxes are for the same image, **the same class**, and they are \"close enough.\" \n\n@fold10\n> If I understood correctly, the mAP looks for the GT bounding box with the highest IoU.\n\nNot exactly. It looks for IoU > 0.4 and takes the highest confidence level/score.",
          "votes": 2
        },
        {
          "id": 1162834,
          "postDate": "2021-01-21T11:21:50.613Z",
          "content": "<p><a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a>  thank you, this makes sense now. So ideally, if a specific image id in the GT has two rows (labels/boxes), the prediction string for that ID should look like this for example:</p>\n<p><code>9306d40fe3b77bd159c6ab92c0a306c8,3 0.9352 713 1465 1644 1784 8 0.6065 471 1407 532 1475</code><br>\nAnd if IoU &gt; 0.4, those boxes will be TP (predicted in order of confidence score).</p>\n<p>Also, I think you're right about the test set, from <a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/207969\" target=\"_blank\">this discussion</a>:</p>\n<blockquote>\n  <p>Each image in the test set of 3,000 images was labeled by the consensus of 5 radiologists. That means there are no overlapping boxes of multiple radiologists in each image.</p>\n</blockquote>\n<p>I guess that means the test set doesn't have predictions from different radiologists. Which is not what I thought at first, that simplifies the issue a bit.</p>",
          "rawMarkdown": "@pestipeti  thank you, this makes sense now. So ideally, if a specific image id in the GT has two rows (labels/boxes), the prediction string for that ID should look like this for example:\n\n`9306d40fe3b77bd159c6ab92c0a306c8,3 0.9352 713 1465 1644 1784 8 0.6065 471 1407 532 1475`\nAnd if IoU > 0.4, those boxes will be TP (predicted in order of confidence score).\n\nAlso, I think you're right about the test set, from [this discussion](https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/207969):\n> Each image in the test set of 3,000 images was labeled by the consensus of 5 radiologists. That means there are no overlapping boxes of multiple radiologists in each image.\n\nI guess that means the test set doesn't have predictions from different radiologists. Which is not what I thought at first, that simplifies the issue a bit.",
          "votes": 1
        },
        {
          "id": 1162989,
          "postDate": "2021-01-21T12:43:50.867Z",
          "content": "<p><a href=\"https://www.kaggle.com/indswetrust\" target=\"_blank\">@indswetrust</a> </p>\n<blockquote>\n  <p><code>9306d40fe3b77bd159c6ab92c0a306c8,3 0.9352 713 1465 1644 1784 8 0.6065 471 1407 532 1475</code><br>\n  And if IoU &gt; 0.4, those boxes will be TP (predicted in order of confidence score).</p>\n</blockquote>\n<p>Yes. One note: in your example, there is one class #3 and one class #8, the order of confidence is only relevant in the same classes. So in this example, it does not matter.</p>",
          "rawMarkdown": "@indswetrust \n\n> `9306d40fe3b77bd159c6ab92c0a306c8,3 0.9352 713 1465 1644 1784 8 0.6065 471 1407 532 1475`\nAnd if IoU > 0.4, those boxes will be TP (predicted in order of confidence score).\n\nYes. One note: in your example, there is one class \\#3 and one class \\#8, the order of confidence is only relevant in the same classes. So in this example, it does not matter.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1346736,
      "postDate": "2021-06-12T15:22:35.007Z",
      "content": "<p>Thanks for explaination</p>",
      "rawMarkdown": "Thanks for explaination"
    },
    {
      "id": 1256130,
      "postDate": "2021-03-29T15:07:29.667Z",
      "content": "<p>Well written. Thanks for sharing! :)</p>",
      "rawMarkdown": "Well written. Thanks for sharing! :)"
    },
    {
      "id": 1241706,
      "postDate": "2021-03-17T07:02:26.813Z",
      "content": "<p>Very helpful. Thanks for sharing!</p>",
      "rawMarkdown": "Very helpful. Thanks for sharing!"
    }
  ],
  "comments": [
    {
      "id": 1187577,
      "author_name": "tito",
      "author_url": "",
      "post_date": "2021-02-05T14:44:00.710000",
      "content": "<p>Nice wrapup!</p>\n<p>I'd like to add some important mAP features.</p>\n<p>1, The lower the confidence threshold, the higher the mAP we will get.<br>\nSo we should use as low a confidence threshold as possible.</p>\n<p>2, Class dependence is linear.</p>\n<pre>mAP(all class) = mAP(class_id == n) + mAP(class_id != n)\n</pre>\n<p>So you can check the scores for each class individually.</p>\n<p>3, Small sample class is equaly important as bit class.</p>\n<pre>max(mAP(small class)) == max(mAP(big class))\n</pre>\n<p>If mAP is mean \"waighted\" average precision, this would be not true.</p>\n<p>For more detail, I made a notebook about this.<br>\n<a href=\"https://www.kaggle.com/its7171/a-few-tips-on-map?scriptVersionId=53590041\" target=\"_blank\">https://www.kaggle.com/its7171/a-few-tips-on-map?scriptVersionId=53590041</a></p>\n<p>I hope this information is useful especially for beginners.</p>",
      "votes": 7,
      "replies": [
        {
          "id": 1187652,
          "author_name": "Hannes Öhler",
          "author_url": "",
          "post_date": "2021-02-05T15:55:20.783000",
          "content": "<p>Very useful, thanks for sharing!</p>\n<p>With respect to 1.: Is this a weakness of the metric? A prediction of something like 100 of objects per image (which you get with a very low confidence threshold) doesn't seem very useful.. </p>\n<p>PS. The link is not working.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1187682,
          "author_name": "tito",
          "author_url": "",
          "post_date": "2021-02-05T16:19:40.163000",
          "content": "<p>I think so.<br>\nIt might be a good idea to add some restrictions such as reasonable submission file size limitation.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1187713,
          "author_name": "tito",
          "author_url": "",
          "post_date": "2021-02-05T16:47:16.397000",
          "content": "<p>However, it is an advantage that we dont have to worry about tuning the confidence threshold because we can simply set the smallest possible threshold.</p>\n<p>On the other hand, this threshold setting can be very critical for other metrics .<br>\nIn some cases, this threshold setting can make the competition a lottery.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1188110,
          "author_name": "Hannes Öhler",
          "author_url": "",
          "post_date": "2021-02-06T01:03:27.173000",
          "content": "<p>Agreed. Or maybe a maximum number per image. </p>\n<p>It's a bit a downer, because you would like to produce realistic, useful predictions, but the metric counteracts this objective.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1230981,
          "author_name": "adriel cabral",
          "author_url": "",
          "post_date": "2021-03-08T15:33:58.397000",
          "content": "<blockquote>\n  <p>1, The lower the confidence threshold, the higher the mAP we will get.<br>\n  So we should use as low a confidence threshold as possible.</p>\n</blockquote>\n<p>I did a test, submission with conf_tresh: 0.01 and other submission with conf_tresh: 0.001, but my lb score drop with conf_tresh: 0.001 , what u think about ? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1231042,
          "author_name": "Hannes Öhler",
          "author_url": "",
          "post_date": "2021-03-08T16:15:24.380000",
          "content": "<p>That is very strange. For me it improved till 0.001, further reduction didn't lead to an improvement, but also didn't hurt. That is expected because we are very much at the right of the precision-recall curve so there is not much room for improvement.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1360126,
          "author_name": "Yousef Rabi",
          "author_url": "",
          "post_date": "2021-06-21T20:36:31.583000",
          "content": "<p><a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> </p>\n<p>Thanks for your added points.</p>\n<p>I have a question regarding the low threshold. Why does lowering the threshold generally increase mAP? Wouldn't lowering the threshold generally introduce a lot more false positives?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1160070,
      "author_name": "Théo Dufort",
      "author_url": "",
      "post_date": "2021-01-19T16:45:05.300000",
      "content": "<p>Very clear and helpful, thank you !</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1206727,
          "author_name": "Jordan Laforet",
          "author_url": "",
          "post_date": "2021-02-17T14:16:17.390000",
          "content": "<p>You are very welcome my friend. We are all set now, let's do it ! 👌</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1159505,
      "author_name": "nbrosse",
      "author_url": "",
      "post_date": "2021-01-19T10:12:50.270000",
      "content": "<p>Thank you very much for the clarification !</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1158812,
      "author_name": "Craig Thomas",
      "author_url": "",
      "post_date": "2021-01-18T19:42:46.343000",
      "content": "<p><a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> thanks for summarizing this. I was reading a few other resources regarding mAP but when I got to calculating the precision-recall curve I got a bit turned around. Your example showing the calculation along with the table helped clarify that immensely. Thanks for posting!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1251211,
      "author_name": "aqx",
      "author_url": "",
      "post_date": "2021-03-24T15:22:53.993000",
      "content": "<p>Thank you for the explanation! I have a clarification with regards to the following scenario: supposed an image have 2 GT box (class 2 and 3). There is a predicted box (class 3), but IOU &gt; defined threshold only with the class 2 GT box. Im not sure if i fully comprehend the mAP. Supposed you loop through the GT classes, it will be a FN on class 2, but it also appears as a FP on class 3?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1197310,
      "author_name": "Old Monk",
      "author_url": "",
      "post_date": "2021-02-12T05:19:46.137000",
      "content": "<p>Thanks for sharing the metric details <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> !</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1161853,
      "author_name": "Alessandro",
      "author_url": "",
      "post_date": "2021-01-20T19:26:04.583000",
      "content": "<p>Thank you for the write-up and notebook.<br>\nI am not sure if your notebook returns the correct result, as** I'd expect a precision of 1** given prediction == annotation.<br>\nCould you provide a working sample?<br>\nAlso, why do you need to remove duplicates?</p>\n<p>Thank you!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3115179%2F567f0397b1ac4651a0d7ae86f50d4ee9%2FDeepinScreenshot_select-area_20210120202345.png?generation=1611170649011001&amp;alt=media\" alt=\"\"></p>\n<p>class_names = [\"image_id\", \"class_name\", \"class_id\", \"rad_id\", \"x_min\", \"y_min\",  \"x_max\",  \"y_max\"]<br>\ndf_sample = pd.DataFrame(dict(zip(class_names,<br>\n                                  [[\"0\", \"1\"], <br>\n                                   [\"No-Finding\", \"No-Finding\"], <br>\n                                   [14, 14], [7, 7], [0., 1.], [0., 1.], [1., 1.], [1., 1.]])))<br>\ndisplay(df_sample)<br>\nvineval_sample = VinBigDataEval(df_sample)<br>\npred_df = pd.DataFrame({\"image_id\": [\"0\", \"1\"],<br>\n                        \"PredictionString\": [\"14 1.0 0 0 1 1\",<br>\n                                             \"14 1.0 1 1 1 1\",<br>\n                                             ]})<br>\ndisplay(pred_df)<br>\ncocoEvalRes = vineval_sample.evaluate(pred_df, n_imgs=-1)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1161881,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2021-01-20T19:41:40.843000",
          "content": "<p>The script is only a wrapper for COCOEval. It just converts our data to their format. It does not calculate the score. I've noticed the same result and it seems that COCOEval has a small bug. It does not count the first true positive sample. If you add more samples (TPs) to the calculation it will work. The one wrongly handled TP will be there, but its weight in the overall score would be less as you increase the number of samples.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1162611,
          "author_name": "InDSweTrust",
          "author_url": "",
          "post_date": "2021-01-21T08:31:44.660000",
          "content": "<p>Speaking of duplicates, they have to be removed because the image id in your predictions file has to match the exact id in the ground truth. In the gt there are multiple rows with the same id, so one pred image id  will match to multiple gt rows and will therefore output an array with multiple bboxes. This format can cause  issues specifically with COCOEval, so duplicates have to be removed.</p>\n<p>However, <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> I am concerned about removing duplicates - this way the predicted box is evaluated only against gt from one radiologist without considering labelled images from other radiologists…</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1162723,
          "author_name": "Alessandro",
          "author_url": "",
          "post_date": "2021-01-21T09:49:03.033000",
          "content": "<p><a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> thanks for your explanation.<br>\nI repeated the experiment for up to 200 samples and the error seems to converge to zero.<br>\nStill, it's a strange behavior. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3115179%2F6a711c0daf5e05cde69c293d0f092098%2FDeepinScreenshot_select-area_20210121104302.png?generation=1611222240444020&amp;alt=media\" alt=\"\"></p>\n<p>Regarding the duplicates:<br>\nIf I understood correctly, the mAP looks for the GT bounding box with the highest IoU. <br>\nHence, there should be only problems for <em>exact</em> duplicates, i.e. duplicate bounding boxes in an image.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1162740,
          "author_name": "InDSweTrust",
          "author_url": "",
          "post_date": "2021-01-21T10:04:37.983000",
          "content": "<p>I looked into this issue in further detail, and yes this is what I found as well. Removing duplicates is OK only in the case of identical bounding boxes for the same image id. In other cases, iou needs to be calculated for all of them and the highest selected.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1162760,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2021-01-21T10:27:36.253000",
          "content": "<p>You don't necessarily have to remove duplicates; COCOEval handles them. The key is the order of your prediction confidence (score).<br>\nIf you have one GT and your model predicts two boxes (same xywh), then it will match (assuming IoU&gt;0.4) only with the highest score. If it is a match, the metric will remove the boxes (one GT and <strong>one</strong> predicted) and the algorithm repeats. You will have one TP and one FP (the other predicted without a matching pair).</p>\n<p>I haven't checked the details of this competition, I am working on the Rainforest. I assumed that in the test set we don't have multiple predictions from different radiologists.</p>\n<p><a href=\"https://www.kaggle.com/indswetrust\" target=\"_blank\">@indswetrust</a><br>\nYou are almost right. You can safely remove duplicates if the boxes are for the same image, <strong>the same class</strong>, and they are \"close enough.\" </p>\n<p><a href=\"https://www.kaggle.com/fold10\" target=\"_blank\">@fold10</a></p>\n<blockquote>\n  <p>If I understood correctly, the mAP looks for the GT bounding box with the highest IoU.</p>\n</blockquote>\n<p>Not exactly. It looks for IoU &gt; 0.4 and takes the highest confidence level/score.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1162834,
          "author_name": "InDSweTrust",
          "author_url": "",
          "post_date": "2021-01-21T11:21:50.613000",
          "content": "<p><a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a>  thank you, this makes sense now. So ideally, if a specific image id in the GT has two rows (labels/boxes), the prediction string for that ID should look like this for example:</p>\n<p><code>9306d40fe3b77bd159c6ab92c0a306c8,3 0.9352 713 1465 1644 1784 8 0.6065 471 1407 532 1475</code><br>\nAnd if IoU &gt; 0.4, those boxes will be TP (predicted in order of confidence score).</p>\n<p>Also, I think you're right about the test set, from <a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/207969\" target=\"_blank\">this discussion</a>:</p>\n<blockquote>\n  <p>Each image in the test set of 3,000 images was labeled by the consensus of 5 radiologists. That means there are no overlapping boxes of multiple radiologists in each image.</p>\n</blockquote>\n<p>I guess that means the test set doesn't have predictions from different radiologists. Which is not what I thought at first, that simplifies the issue a bit.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1162989,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2021-01-21T12:43:50.867000",
          "content": "<p><a href=\"https://www.kaggle.com/indswetrust\" target=\"_blank\">@indswetrust</a> </p>\n<blockquote>\n  <p><code>9306d40fe3b77bd159c6ab92c0a306c8,3 0.9352 713 1465 1644 1784 8 0.6065 471 1407 532 1475</code><br>\n  And if IoU &gt; 0.4, those boxes will be TP (predicted in order of confidence score).</p>\n</blockquote>\n<p>Yes. One note: in your example, there is one class #3 and one class #8, the order of confidence is only relevant in the same classes. So in this example, it does not matter.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1346736,
      "author_name": "ajmasih",
      "author_url": "",
      "post_date": "2021-06-12T15:22:35.007000",
      "content": "<p>Thanks for explaination</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1256130,
      "author_name": "antoreepjana",
      "author_url": "",
      "post_date": "2021-03-29T15:07:29.667000",
      "content": "<p>Well written. Thanks for sharing! :)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1241706,
      "author_name": "Ivan Stebakov",
      "author_url": "",
      "post_date": "2021-03-17T07:02:26.813000",
      "content": "<p>Very helpful. Thanks for sharing!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1158063": "# Mean Average Precision\n> The challenge uses the standard PASCAL VOC 2010 mean Average Precision (mAP) at IoU > 0.4.\n\n\n## Basics\n\n### Precision and Recall\n**Precision**: What proportion of the positive identifications (predicted boxes) was actually correct?\nPrecision = TP/(TP + FP)\n\n**Recall** What proportion of actual positives was identified correctly?\nRecall = TP / (TP + FN)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F864684%2Fb3a58f56a1bfdbc4841dfd47747193ef%2Fprecision_recall_tp_fp.png?generation=1610967621000417&alt=media)\n\nTP: True positives. You predict, and it is correct.\nFP: False positives. You predict, but it is wrong.\nFN: False negatives. You did not predict, and it is wrong. You should have predicted.\n\nThere are no true negatives (TN) in object detection. \n\n### Intersection over union (IoU)\nIoU measures the overlap between 2 boxes. \nIoU = (area of overlap) / (area of union)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F864684%2F3a82ceb91fbb96d5154f6b15062804df%2Fiou.jpg?generation=1610967851297996&alt=media)\n\n#### IoU \\@0.4\nIoU at 0.4 means that we consider a predicted box as a true positive (TP) if the IoU of this box and one of the GT box is greater than 0.4.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F864684%2F50b91585a79d9e0ecb24d874e78be99e%2Fiou_examples.jpg?generation=1610967875651682&alt=media)\n\n### Score\nFor this metric, we have to add a confidence score for our predictions. It acts as a tie-breaker. For example, if there are two or more predicted boxes with IoU>0.4 with the same GT, the metric will consider our most confident (highest score) prediction as TP for that GT-Pred pair.\n\n## How to determine TP, FP, FN\n- Sort the predicted boxes in descendent (by score)\n- Pick the most confident prediction\n- Calculate the IoU between the selected prediction and **all** of the GT boxes\n- Select the highest IoU\n- If the IoU is greater than our threshold (0.4), we have a true positive (TP) match. In this case, we remove the matched boxes from both the predicted and the GT list of boxes.\n- Repeat 2-4\n- After we iterate through in our predicted boxes, what is left in the predicted list are the false positives (FP) and what is left in the GT list are the false negatives (FN).\n\n*All of these are per class!!*\n\n## Average Precision\n| Rank | Correct (TP)? | Precision | Recall |\n| ---- | ------------- | --------- | ------ |\n| 1    | True          | 1.0       | 0.2    |\n| 2    | True          | 1.0       | 0.4    |\n| 3    | False         | 0.67      | 0.4    |\n| 4    | False         | 0.5       | 0.4    |\n| 5    | False         | 0.4       | 0.4    |\n| 6    | True          | 0.5       | 0.6    |\n| 7    | True          | 0.57      | 0.8    |\n| 8    | False         | 0.5       | 0.8    |\n| 9    | False         | 0.44      | 0.8    |\n| 10   | True          | 0.5       | 1.0    |\n\nIn the table above the rows are our predictions. Let's assume we have 5 GT boxes in this example. So you have to iterate over your prediction and calculate Precisions and recalls after every row.\n\nRow #1, you were correct. So far, we have: TP=1, FP=0, FN=4; Precision = 1/(1+0)=1, Recall = 1/(1+4) = 0.2\nRow #2, you were correct. So far, we have: TP=2, FP=0, FN=3; Precision = 2/(2+0)=1, Recall = 2/(2+3) = 0.4\nRow #3, you were incorrect. So far, we have: TP=2, FP=1, FN=3; Precision = 2/(2+1)=0.67, Recall = 2/(2+3)=0.4\nand so on...\n\nNote: The recall value is always increasing (or stays the same). You can plot this: precision-recall curve, where the Recall is the x-axis, Precision is y.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F864684%2F667d2e7cf09649a352aaeeee7ad27fee%2Fprecision_recall_curve.png?generation=1610967690204686&alt=media)\n\nThe Average Precision (AP) is the area under the precision-recall curve. We could integrate this, but fortunately, we don't have to. COCOEval approximates this by divided the recall-axis into 100 intervals and take the highest Precision at every interval. Kind of Riemann sum.\n\n\n## Healthy samples\nBy default, object detection can not deal with true negatives. You can not predict a box for an object that is not there. The host's solution is this 1x1 pixel at the corner. From the AP point of view, it is just another box, and the metric handles these boxes as the same.\n\nYou can experiment with this fact. For example, if your model predicts one uncertain box with 0.51 confidence score, that means there is a chance that the sample is healthy. It might improve your overall score if you add \"14 1 0 0 1 1\" to your prediction.\nMore details in this [discussion](https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/211971)\n\n## Calculation\nSee my [Competition Metric Calculator](https://www.kaggle.com/pestipeti/competition-metric-map-0-4) notebook.\n\n## References\n- [mAP (mean Average Precision) for Object Detection](https://jonathan-hui.medium.com/map-mean-average-precision-for-object-detection-45c121a31173) by Jonathan Hui\n- [Evaluating Object Detection Models Using Mean Average Precision (mAP)](https://blog.paperspace.com/mean-average-precision/) by Ahmed Fawzy Gad - Paperspace\n- [Precision and Recall](https://en.wikipedia.org/wiki/Precision_and_recall) - Wikipedia\n\n\n\n\n\n\n\n\n",
    "1187577": "Nice wrapup!\n\nI'd like to add some important mAP features.\n\n1, The lower the confidence threshold, the higher the mAP we will get.\nSo we should use as low a confidence threshold as possible.\n\n2, Class dependence is linear.\n<pre>\nmAP(all class) = mAP(class_id == n) + mAP(class_id != n)\n</pre>\nSo you can check the scores for each class individually.\n\n3, Small sample class is equaly important as bit class.\n<pre>\nmax(mAP(small class)) == max(mAP(big class))\n</pre>\nIf mAP is mean \"waighted\" average precision, this would be not true.\n\nFor more detail, I made a notebook about this.\nhttps://www.kaggle.com/its7171/a-few-tips-on-map?scriptVersionId=53590041\n\nI hope this information is useful especially for beginners.",
    "1160070": "Very clear and helpful, thank you !",
    "1159505": "Thank you very much for the clarification !",
    "1158812": "@pestipeti thanks for summarizing this. I was reading a few other resources regarding mAP but when I got to calculating the precision-recall curve I got a bit turned around. Your example showing the calculation along with the table helped clarify that immensely. Thanks for posting!",
    "1251211": "Thank you for the explanation! I have a clarification with regards to the following scenario: supposed an image have 2 GT box (class 2 and 3). There is a predicted box (class 3), but IOU > defined threshold only with the class 2 GT box. Im not sure if i fully comprehend the mAP. Supposed you loop through the GT classes, it will be a FN on class 2, but it also appears as a FP on class 3?",
    "1197310": "Thanks for sharing the metric details @pestipeti !",
    "1161853": "Thank you for the write-up and notebook.\nI am not sure if your notebook returns the correct result, as** I'd expect a precision of 1** given prediction == annotation.\nCould you provide a working sample?\nAlso, why do you need to remove duplicates?\n\nThank you!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3115179%2F567f0397b1ac4651a0d7ae86f50d4ee9%2FDeepinScreenshot_select-area_20210120202345.png?generation=1611170649011001&alt=media)\n\nclass_names = [\"image_id\", \"class_name\", \"class_id\", \"rad_id\", \"x_min\", \"y_min\",  \"x_max\",  \"y_max\"]\ndf_sample = pd.DataFrame(dict(zip(class_names,\n                                  [[\"0\", \"1\"], \n                                   [\"No-Finding\", \"No-Finding\"], \n                                   [14, 14], [7, 7], [0., 1.], [0., 1.], [1., 1.], [1., 1.]])))\ndisplay(df_sample)\nvineval_sample = VinBigDataEval(df_sample)\npred_df = pd.DataFrame({\"image_id\": [\"0\", \"1\"],\n                        \"PredictionString\": [\"14 1.0 0 0 1 1\",\n                                             \"14 1.0 1 1 1 1\",\n                                             ]})\ndisplay(pred_df)\ncocoEvalRes = vineval_sample.evaluate(pred_df, n_imgs=-1)\n\n\n",
    "1346736": "Thanks for explaination",
    "1256130": "Well written. Thanks for sharing! :)",
    "1241706": "Very helpful. Thanks for sharing!"
  }
}