{
  "id": 229817,
  "title": "Does this metric even work?",
  "url": "/competitions/vinbigdata-chest-xray-abnormalities-detection/discussion/229817",
  "author_name": "Awsaf",
  "post_date": "2021-03-31T21:12:31.536000",
  "votes": 15,
  "comment_count": 6,
  "views": 0,
  "content": "<p>The metric for this competition is well-known, <strong>mean average precision(mAP)</strong>.  One of the features of this metric is <a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229637\" target=\"_blank\">No Penalty For Adding More Bbox </a><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. So, if we submit with lots of <strong>False Positives</strong> still we'll be able to get a good score. Check out the images below which are beyond comprehension. It was still able to get 0.26+ score on lb. So it really makes me think, <strong>does this metric even work?</strong><br>\n<img src=\"https://i.ibb.co/16cJfqh/165867751-1777430435773207-7505470057566631699-n.png\" alt=\"\"></p>",
  "messages": [
    {
      "id": 1258750,
      "postDate": "2021-03-31T21:12:31.537Z",
      "content": "<p>The metric for this competition is well-known, <strong>mean average precision(mAP)</strong>.  One of the features of this metric is <a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229637\" target=\"_blank\">No Penalty For Adding More Bbox </a><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. So, if we submit with lots of <strong>False Positives</strong> still we'll be able to get a good score. Check out the images below which are beyond comprehension. It was still able to get 0.26+ score on lb. So it really makes me think, <strong>does this metric even work?</strong><br>\n<img src=\"https://i.ibb.co/16cJfqh/165867751-1777430435773207-7505470057566631699-n.png\" alt=\"\"></p>",
      "rawMarkdown": "The metric for this competition is well-known, **mean average precision(mAP)**.  One of the features of this metric is [No Penalty For Adding More Bbox @cdeotte](https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229637). So, if we submit with lots of **False Positives** still we'll be able to get a good score. Check out the images below which are beyond comprehension. It was still able to get 0.26+ score on lb. So it really makes me think, **does this metric even work?**\n![](https://i.ibb.co/16cJfqh/165867751-1777430435773207-7505470057566631699-n.png)",
      "votes": 14
    },
    {
      "id": 1260886,
      "postDate": "2021-04-02T13:52:10.373Z",
      "content": "<p>I think the potential problem with the metric is the following scenario. The purpose of a metric is to say \"one thing is better than another\", or \"two things are equal\". The following example have the same <code>mAP</code>. Do they have equal value to the host?</p>\n<p>Both make 12 predictions and both have <code>mAP = 0.50</code>. (Note that mAP does take the area, but uses horizontal lines like pictured below). If our models are in production, we do not know which bbox are TP or FP, we just have predictions that are ordered by confidence scores.</p>\n<p>If both of the below models are in production, and we select the 3 more confident bbox, then the model on the left is all TP and the model on the right is all FP. Are these models really equal? One could argue that the model on the left is better.</p>\n<p>Both have 50% area, but the one on the left has more area to the left. So one way to change the metric is to weight area on the left as more important than area on the right. (This would reward a model who top k bbox have more TP).</p>\n<p><img src=\"https://www.ccom.ucsd.edu/~cdeotte/Kaggle/compare-4-2.png\" alt=\"\"></p>",
      "rawMarkdown": "I think the potential problem with the metric is the following scenario. The purpose of a metric is to say \"one thing is better than another\", or \"two things are equal\". The following example have the same `mAP`. Do they have equal value to the host?\n\nBoth make 12 predictions and both have `mAP = 0.50`. (Note that mAP does take the area, but uses horizontal lines like pictured below). If our models are in production, we do not know which bbox are TP or FP, we just have predictions that are ordered by confidence scores.\n\nIf both of the below models are in production, and we select the 3 more confident bbox, then the model on the left is all TP and the model on the right is all FP. Are these models really equal? One could argue that the model on the left is better.\n\nBoth have 50% area, but the one on the left has more area to the left. So one way to change the metric is to weight area on the left as more important than area on the right. (This would reward a model who top k bbox have more TP).\n\n![](https://www.ccom.ucsd.edu/~cdeotte/Kaggle/compare-4-2.png)",
      "votes": 4,
      "replies": [
        {
          "id": 1260901,
          "postDate": "2021-04-02T14:14:25.383Z",
          "content": "<p>A little bit different metric was used in <a href=\"https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/overview/evaluation\" target=\"_blank\"><strong>Pneumonia Detection</strong></a> Competition with same name <strong>mAP</strong> but there instead of <strong>Area Under the Curve</strong>,  <strong>mAP</strong> is calculated using this formula, <br>\n<img src=\"https://latex.codecogs.com/gif.download?mAP%20%3D%20%5Cfrac%7BTP%28t%29%7D%7BTP%28t%29%20+%20FP%28t%29%20+%20FN%28t%29%7D\" alt=\"\"></p>",
          "rawMarkdown": "A little bit different metric was used in [**Pneumonia Detection**](https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/overview/evaluation) Competition with same name **mAP** but there instead of **Area Under the Curve**,  **mAP** is calculated using this formula, \n![](https://latex.codecogs.com/gif.download?mAP%20%3D%20%5Cfrac%7BTP%28t%29%7D%7BTP%28t%29%20+%20FP%28t%29%20+%20FN%28t%29%7D)"
        }
      ]
    },
    {
      "id": 1260655,
      "postDate": "2021-04-02T10:02:56.880Z",
      "content": "<p>I was also curious why this metric was selected. Same thoughts - huge number of bboxes with conf scores &lt; 0.001 which make the predictions incomprehensible, yet increase the mAP score.</p>\n<p>I was thinking maybe its because during medical diagnostics you would rather have a lot of predictions which include correct abnormality (TP) + lots of FPs, rather than a limited number of predictions where the TP might be missed out.<br>\nBut this is just a guess - I am new to object detection and haven't worked with other metrics yet, so maybe other metrics have different limitations.</p>",
      "rawMarkdown": "I was also curious why this metric was selected. Same thoughts - huge number of bboxes with conf scores < 0.001 which make the predictions incomprehensible, yet increase the mAP score.\n\nI was thinking maybe its because during medical diagnostics you would rather have a lot of predictions which include correct abnormality (TP) + lots of FPs, rather than a limited number of predictions where the TP might be missed out.\nBut this is just a guess - I am new to object detection and haven't worked with other metrics yet, so maybe other metrics have different limitations.",
      "replies": [
        {
          "id": 1260716,
          "postDate": "2021-04-02T11:18:08.663Z",
          "content": "<p>Yes you're right sometimes <code>TP</code> is more important than <code>FP</code> but with that many <code>FP</code> I'm not sure how health workers will treat the patient. </p>",
          "rawMarkdown": "Yes you're right sometimes `TP` is more important than `FP` but with that many `FP` I'm not sure how health workers will treat the patient. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1259230,
      "postDate": "2021-04-01T08:33:24.423Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, I think your comment somehow got deleted, I'm talking about this scenario. If we take <strong>top k</strong> then those low-confident boxes will be removed and it'll reduce the score. <img src=\"https://www.ccom.ucsd.edu/~cdeotte/Kaggle/map3.png\" alt=\"\"></p>",
      "rawMarkdown": "Hi @cdeotte, I think your comment somehow got deleted, I'm talking about this scenario. If we take **top k** then those low-confident boxes will be removed and it'll reduce the score. ![](https://www.ccom.ucsd.edu/~cdeotte/Kaggle/map3.png)"
    },
    {
      "id": 1258886,
      "postDate": "2021-04-01T01:27:59.280Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1260886,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2021-04-02T13:52:10.373000",
      "content": "<p>I think the potential problem with the metric is the following scenario. The purpose of a metric is to say \"one thing is better than another\", or \"two things are equal\". The following example have the same <code>mAP</code>. Do they have equal value to the host?</p>\n<p>Both make 12 predictions and both have <code>mAP = 0.50</code>. (Note that mAP does take the area, but uses horizontal lines like pictured below). If our models are in production, we do not know which bbox are TP or FP, we just have predictions that are ordered by confidence scores.</p>\n<p>If both of the below models are in production, and we select the 3 more confident bbox, then the model on the left is all TP and the model on the right is all FP. Are these models really equal? One could argue that the model on the left is better.</p>\n<p>Both have 50% area, but the one on the left has more area to the left. So one way to change the metric is to weight area on the left as more important than area on the right. (This would reward a model who top k bbox have more TP).</p>\n<p><img src=\"https://www.ccom.ucsd.edu/~cdeotte/Kaggle/compare-4-2.png\" alt=\"\"></p>",
      "votes": 4,
      "replies": [
        {
          "id": 1260901,
          "author_name": "Awsaf",
          "author_url": "",
          "post_date": "2021-04-02T14:14:25.383000",
          "content": "<p>A little bit different metric was used in <a href=\"https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/overview/evaluation\" target=\"_blank\"><strong>Pneumonia Detection</strong></a> Competition with same name <strong>mAP</strong> but there instead of <strong>Area Under the Curve</strong>,  <strong>mAP</strong> is calculated using this formula, <br>\n<img src=\"https://latex.codecogs.com/gif.download?mAP%20%3D%20%5Cfrac%7BTP%28t%29%7D%7BTP%28t%29%20+%20FP%28t%29%20+%20FN%28t%29%7D\" alt=\"\"></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1260655,
      "author_name": "InDSweTrust",
      "author_url": "",
      "post_date": "2021-04-02T10:02:56.880000",
      "content": "<p>I was also curious why this metric was selected. Same thoughts - huge number of bboxes with conf scores &lt; 0.001 which make the predictions incomprehensible, yet increase the mAP score.</p>\n<p>I was thinking maybe its because during medical diagnostics you would rather have a lot of predictions which include correct abnormality (TP) + lots of FPs, rather than a limited number of predictions where the TP might be missed out.<br>\nBut this is just a guess - I am new to object detection and haven't worked with other metrics yet, so maybe other metrics have different limitations.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1260716,
          "author_name": "Awsaf",
          "author_url": "",
          "post_date": "2021-04-02T11:18:08.663000",
          "content": "<p>Yes you're right sometimes <code>TP</code> is more important than <code>FP</code> but with that many <code>FP</code> I'm not sure how health workers will treat the patient. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1259230,
      "author_name": "Awsaf",
      "author_url": "",
      "post_date": "2021-04-01T08:33:24.423000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, I think your comment somehow got deleted, I'm talking about this scenario. If we take <strong>top k</strong> then those low-confident boxes will be removed and it'll reduce the score. <img src=\"https://www.ccom.ucsd.edu/~cdeotte/Kaggle/map3.png\" alt=\"\"></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1258886,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-04-01T01:27:59.280000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1258750": "The metric for this competition is well-known, **mean average precision(mAP)**.  One of the features of this metric is [No Penalty For Adding More Bbox @cdeotte](https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229637). So, if we submit with lots of **False Positives** still we'll be able to get a good score. Check out the images below which are beyond comprehension. It was still able to get 0.26+ score on lb. So it really makes me think, **does this metric even work?**\n![](https://i.ibb.co/16cJfqh/165867751-1777430435773207-7505470057566631699-n.png)",
    "1260886": "I think the potential problem with the metric is the following scenario. The purpose of a metric is to say \"one thing is better than another\", or \"two things are equal\". The following example have the same `mAP`. Do they have equal value to the host?\n\nBoth make 12 predictions and both have `mAP = 0.50`. (Note that mAP does take the area, but uses horizontal lines like pictured below). If our models are in production, we do not know which bbox are TP or FP, we just have predictions that are ordered by confidence scores.\n\nIf both of the below models are in production, and we select the 3 more confident bbox, then the model on the left is all TP and the model on the right is all FP. Are these models really equal? One could argue that the model on the left is better.\n\nBoth have 50% area, but the one on the left has more area to the left. So one way to change the metric is to weight area on the left as more important than area on the right. (This would reward a model who top k bbox have more TP).\n\n![](https://www.ccom.ucsd.edu/~cdeotte/Kaggle/compare-4-2.png)",
    "1260655": "I was also curious why this metric was selected. Same thoughts - huge number of bboxes with conf scores < 0.001 which make the predictions incomprehensible, yet increase the mAP score.\n\nI was thinking maybe its because during medical diagnostics you would rather have a lot of predictions which include correct abnormality (TP) + lots of FPs, rather than a limited number of predictions where the TP might be missed out.\nBut this is just a guess - I am new to object detection and haven't worked with other metrics yet, so maybe other metrics have different limitations.",
    "1259230": "Hi @cdeotte, I think your comment somehow got deleted, I'm talking about this scenario. If we take **top k** then those low-confident boxes will be removed and it'll reduce the score. ![](https://www.ccom.ucsd.edu/~cdeotte/Kaggle/map3.png)",
    "1258886": ""
  }
}