{
  "id": 369886,
  "title": "exploit the metric \"bug\"",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/369886",
  "author_name": "hengck23",
  "post_date": "2022-12-01T18:20:20.451000",
  "votes": 95,
  "comment_count": 10,
  "views": 0,
  "content": "<p>no metric is perfect. <br>\nhere you can \"improve\" your score by exploiting the metric \"bug\"<br>\n(the accuracy of the model was not improved but the lb score value was artifically  improved) </p>\n<ol>\n<li>plot the distribution of your model prediction for the pos cases and neg cases </li>\n<li>you note that the neg distribution is exponential-like, with a very sharp peak at p=0</li>\n<li>now the distribution is highly imbalance, with 99% being negative</li>\n<li>Hence you can get a better lb score if you do some hard thresholding, e.g. set predict[predict&lt;0.1]=0</li>\n</ol>\n<p>you can search for \"best parameters\": predict[??&lt;predict&lt;??]=??</p>\n<p>you can verify this with your validation set. </p>\n<hr>\n<p>final note : beware of shakeup using this method</p>",
  "messages": [
    {
      "id": 2051932,
      "postDate": "2022-12-01T18:20:20.450Z",
      "content": "<p>no metric is perfect. <br>\nhere you can \"improve\" your score by exploiting the metric \"bug\"<br>\n(the accuracy of the model was not improved but the lb score value was artifically  improved) </p>\n<ol>\n<li>plot the distribution of your model prediction for the pos cases and neg cases </li>\n<li>you note that the neg distribution is exponential-like, with a very sharp peak at p=0</li>\n<li>now the distribution is highly imbalance, with 99% being negative</li>\n<li>Hence you can get a better lb score if you do some hard thresholding, e.g. set predict[predict&lt;0.1]=0</li>\n</ol>\n<p>you can search for \"best parameters\": predict[??&lt;predict&lt;??]=??</p>\n<p>you can verify this with your validation set. </p>\n<hr>\n<p>final note : beware of shakeup using this method</p>",
      "rawMarkdown": "no metric is perfect. \nhere you can \"improve\" your score by exploiting the metric \"bug\"\n(the accuracy of the model was not improved but the lb score value was artifically  improved) \n\n1. plot the distribution of your model prediction for the pos cases and neg cases \n2. you note that the neg distribution is exponential-like, with a very sharp peak at p=0\n3. now the distribution is highly imbalance, with 99% being negative\n4. Hence you can get a better lb score if you do some hard thresholding, e.g. set predict[predict<0.1]=0\n\nyou can search for \"best parameters\": predict[??<predict<??]=??\n\nyou can verify this with your validation set. \n\n---\n\nfinal note : beware of shakeup using this method",
      "votes": 94
    },
    {
      "id": 2052489,
      "postDate": "2022-12-02T08:41:17.167Z",
      "content": "<p>Thanks for sharing Heng. You've saved a lot of trouble to many of us.</p>\n<p>As you can see from my current LB the boost is quite big aha</p>",
      "rawMarkdown": "Thanks for sharing Heng. You've saved a lot of trouble to many of us.\n\nAs you can see from my current LB the boost is quite big aha",
      "votes": 5,
      "replies": [
        {
          "id": 2053782,
          "postDate": "2022-12-03T16:00:11.697Z",
          "content": "<p>Hey Theo,</p>\n<p>Was your LB boost based on your resized 512x512 ?</p>",
          "rawMarkdown": "Hey Theo,\n\nWas your LB boost based on your resized 512x512 ?"
        },
        {
          "id": 2053823,
          "postDate": "2022-12-03T16:30:55.073Z",
          "content": "<p>yup                    </p>",
          "rawMarkdown": "yup                    ",
          "votes": 1
        }
      ]
    },
    {
      "id": 2053115,
      "postDate": "2022-12-02T21:04:28.947Z",
      "content": "<p>here are the details:<br>\n(results per image, not per patient)</p>\n<pre><code>one fold only.\ntrain images :\n    num_patient = 9530\n    num_image = 43771\n        cancer0 = 42854 (0.979)\n        cancer1 =   917 (0.021)\nvalidation images :\n    num_patient = 2383\n    num_image = 10935\n        cancer0 = 10694 (0.978)\n        cancer1 =   241 (0.022)\n</code></pre>\n<p><a href=\"https://ibb.co/TP6QPyP\"><img src=\"https://i.ibb.co/P1K21k1/Selection-052.png\" alt=\"Selection-052\"></a></p>",
      "rawMarkdown": "here are the details:\n(results per image, not per patient)\n\n```\none fold only.\ntrain images :\n\tnum_patient = 9530\n\tnum_image = 43771\n\t\tcancer0 = 42854 (0.979)\n\t\tcancer1 =   917 (0.021)\nvalidation images :\n\tnum_patient = 2383\n\tnum_image = 10935\n\t\tcancer0 = 10694 (0.978)\n\t\tcancer1 =   241 (0.022)\n```\n\n<a href=\"https://ibb.co/TP6QPyP\"><img src=\"https://i.ibb.co/P1K21k1/Selection-052.png\" alt=\"Selection-052\" border=\"0\"></a>\n",
      "votes": 6,
      "replies": [
        {
          "id": 2053241,
          "postDate": "2022-12-03T02:17:57.797Z",
          "content": "<p>yet another magic</p>\n<p><img src=\"https://i.ibb.co/R4B7RjS/Selection-059.png\" alt=\"https://i.ibb.co/R4B7RjS/Selection-059.png\"></p>\n<p>obviously if you train a model to predict using multiview inputs, the results will be better than averaging (and/or max)</p>",
          "rawMarkdown": "yet another magic\n\n![https://i.ibb.co/R4B7RjS/Selection-059.png](https://i.ibb.co/R4B7RjS/Selection-059.png)\n\nobviously if you train a model to predict using multiview inputs, the results will be better than averaging (and/or max)",
          "votes": 3
        }
      ]
    },
    {
      "id": 2053340,
      "postDate": "2022-12-03T06:14:54.637Z",
      "content": "<p>is this a leak?</p>\n<pre><code>HINT: you can do a probe\n\ntrain_df[['site_id','machine_id','cancer']].groupby(['site_id','machine_id']).mean()\nOut[6]: \n                      cancer\nsite_id machine_id          \n1       49          0.026095\n        93          0.007311\n        170         0.024919\n        190         0.034483\n        197         0.000000\n        210         0.000000\n        216         0.004193\n2       21          0.018733\n        29          0.019475\n        48          0.020577\n\n\ntrain_df[['site_id','machine_id','cancer']].groupby(['site_id','machine_id']).count()\nOut[7]: \n                    cancer\nsite_id machine_id        \n1       49           23529\n        93            1915\n        170            923\n        190            145\n        197             29\n        210           1070\n        216           1908\n2       21            8221\n        29            8267\n        48            8699\n</code></pre>",
      "rawMarkdown": "is this a leak?\n\n```\nHINT: you can do a probe\n\ntrain_df[['site_id','machine_id','cancer']].groupby(['site_id','machine_id']).mean()\nOut[6]: \n                      cancer\nsite_id machine_id          \n1       49          0.026095\n        93          0.007311\n        170         0.024919\n        190         0.034483\n        197         0.000000\n        210         0.000000\n        216         0.004193\n2       21          0.018733\n        29          0.019475\n        48          0.020577\n\n\ntrain_df[['site_id','machine_id','cancer']].groupby(['site_id','machine_id']).count()\nOut[7]: \n                    cancer\nsite_id machine_id        \n1       49           23529\n        93            1915\n        170            923\n        190            145\n        197             29\n        210           1070\n        216           1908\n2       21            8221\n        29            8267\n        48            8699\n```",
      "votes": 4
    },
    {
      "id": 2053121,
      "postDate": "2022-12-02T21:10:52.127Z",
      "content": "<p>now we have a serious problem:</p>\n<p>why use pfbeta (probabilistic) if everyone is submitting fbeta (fixed threshold)?<br>\nI think AUC(or raw class-weighted log(p)) is a better metric for lb score?</p>",
      "rawMarkdown": "now we have a serious problem:\n\nwhy use pfbeta (probabilistic) if everyone is submitting fbeta (fixed threshold)?\nI think AUC(or raw class-weighted log(p)) is a better metric for lb score?",
      "votes": 4,
      "replies": [
        {
          "id": 2053142,
          "postDate": "2022-12-02T21:30:52.997Z",
          "content": "<p>Isn't beta=1, hence fixed?</p>",
          "rawMarkdown": "Isn't beta=1, hence fixed?",
          "replies": [
            {
              "id": 2095647,
              "postDate": "2023-01-11T14:41:07.157Z",
              "content": "<p>I think what he means, is that people are not submitting the probabilites given by the model.<br>\nInstead they binarize they're output based on an optimized threshold to maximize the score</p>",
              "rawMarkdown": "I think what he means, is that people are not submitting the probabilites given by the model.\nInstead they binarize they're output based on an optimized threshold to maximize the score",
              "votes": 3
            },
            {
              "id": 2134348,
              "postDate": "2023-02-07T23:14:33.277Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2052489,
      "author_name": "Theo Viel",
      "author_url": "",
      "post_date": "2022-12-02T08:41:17.167000",
      "content": "<p>Thanks for sharing Heng. You've saved a lot of trouble to many of us.</p>\n<p>As you can see from my current LB the boost is quite big aha</p>",
      "votes": 5,
      "replies": [
        {
          "id": 2053782,
          "author_name": "Alin Cijov",
          "author_url": "",
          "post_date": "2022-12-03T16:00:11.697000",
          "content": "<p>Hey Theo,</p>\n<p>Was your LB boost based on your resized 512x512 ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2053823,
          "author_name": "Theo Viel",
          "author_url": "",
          "post_date": "2022-12-03T16:30:55.073000",
          "content": "<p>yup                    </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2053115,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-12-02T21:04:28.947000",
      "content": "<p>here are the details:<br>\n(results per image, not per patient)</p>\n<pre><code>one fold only.\ntrain images :\n    num_patient = 9530\n    num_image = 43771\n        cancer0 = 42854 (0.979)\n        cancer1 =   917 (0.021)\nvalidation images :\n    num_patient = 2383\n    num_image = 10935\n        cancer0 = 10694 (0.978)\n        cancer1 =   241 (0.022)\n</code></pre>\n<p><a href=\"https://ibb.co/TP6QPyP\"><img src=\"https://i.ibb.co/P1K21k1/Selection-052.png\" alt=\"Selection-052\"></a></p>",
      "votes": 6,
      "replies": [
        {
          "id": 2053241,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2022-12-03T02:17:57.797000",
          "content": "<p>yet another magic</p>\n<p><img src=\"https://i.ibb.co/R4B7RjS/Selection-059.png\" alt=\"https://i.ibb.co/R4B7RjS/Selection-059.png\"></p>\n<p>obviously if you train a model to predict using multiview inputs, the results will be better than averaging (and/or max)</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2053340,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-12-03T06:14:54.637000",
      "content": "<p>is this a leak?</p>\n<pre><code>HINT: you can do a probe\n\ntrain_df[['site_id','machine_id','cancer']].groupby(['site_id','machine_id']).mean()\nOut[6]: \n                      cancer\nsite_id machine_id          \n1       49          0.026095\n        93          0.007311\n        170         0.024919\n        190         0.034483\n        197         0.000000\n        210         0.000000\n        216         0.004193\n2       21          0.018733\n        29          0.019475\n        48          0.020577\n\n\ntrain_df[['site_id','machine_id','cancer']].groupby(['site_id','machine_id']).count()\nOut[7]: \n                    cancer\nsite_id machine_id        \n1       49           23529\n        93            1915\n        170            923\n        190            145\n        197             29\n        210           1070\n        216           1908\n2       21            8221\n        29            8267\n        48            8699\n</code></pre>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 2053121,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-12-02T21:10:52.127000",
      "content": "<p>now we have a serious problem:</p>\n<p>why use pfbeta (probabilistic) if everyone is submitting fbeta (fixed threshold)?<br>\nI think AUC(or raw class-weighted log(p)) is a better metric for lb score?</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2053142,
          "author_name": "moth",
          "author_url": "",
          "post_date": "2022-12-02T21:30:52.997000",
          "content": "<p>Isn't beta=1, hence fixed?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2095647,
              "author_name": "KSMCG90",
              "author_url": "",
              "post_date": "2023-01-11T14:41:07.157000",
              "content": "<p>I think what he means, is that people are not submitting the probabilites given by the model.<br>\nInstead they binarize they're output based on an optimized threshold to maximize the score</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2134348,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-02-07T23:14:33.277000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2051932": "no metric is perfect. \nhere you can \"improve\" your score by exploiting the metric \"bug\"\n(the accuracy of the model was not improved but the lb score value was artifically  improved) \n\n1. plot the distribution of your model prediction for the pos cases and neg cases \n2. you note that the neg distribution is exponential-like, with a very sharp peak at p=0\n3. now the distribution is highly imbalance, with 99% being negative\n4. Hence you can get a better lb score if you do some hard thresholding, e.g. set predict[predict<0.1]=0\n\nyou can search for \"best parameters\": predict[??<predict<??]=??\n\nyou can verify this with your validation set. \n\n---\n\nfinal note : beware of shakeup using this method",
    "2052489": "Thanks for sharing Heng. You've saved a lot of trouble to many of us.\n\nAs you can see from my current LB the boost is quite big aha",
    "2053115": "here are the details:\n(results per image, not per patient)\n\n```\none fold only.\ntrain images :\n\tnum_patient = 9530\n\tnum_image = 43771\n\t\tcancer0 = 42854 (0.979)\n\t\tcancer1 =   917 (0.021)\nvalidation images :\n\tnum_patient = 2383\n\tnum_image = 10935\n\t\tcancer0 = 10694 (0.978)\n\t\tcancer1 =   241 (0.022)\n```\n\n<a href=\"https://ibb.co/TP6QPyP\"><img src=\"https://i.ibb.co/P1K21k1/Selection-052.png\" alt=\"Selection-052\" border=\"0\"></a>\n",
    "2053340": "is this a leak?\n\n```\nHINT: you can do a probe\n\ntrain_df[['site_id','machine_id','cancer']].groupby(['site_id','machine_id']).mean()\nOut[6]: \n                      cancer\nsite_id machine_id          \n1       49          0.026095\n        93          0.007311\n        170         0.024919\n        190         0.034483\n        197         0.000000\n        210         0.000000\n        216         0.004193\n2       21          0.018733\n        29          0.019475\n        48          0.020577\n\n\ntrain_df[['site_id','machine_id','cancer']].groupby(['site_id','machine_id']).count()\nOut[7]: \n                    cancer\nsite_id machine_id        \n1       49           23529\n        93            1915\n        170            923\n        190            145\n        197             29\n        210           1070\n        216           1908\n2       21            8221\n        29            8267\n        48            8699\n```",
    "2053121": "now we have a serious problem:\n\nwhy use pfbeta (probabilistic) if everyone is submitting fbeta (fixed threshold)?\nI think AUC(or raw class-weighted log(p)) is a better metric for lb score?"
  }
}