{
  "id": 372175,
  "title": "What is the rationale for this metric above AUC?",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/372175",
  "author_name": "James Howard",
  "post_date": "2022-12-14T15:17:50.932000",
  "votes": 32,
  "comment_count": 32,
  "views": 0,
  "content": "<p>The stated rationale for this metric is it is probabilistic. Well, it's not - everyone's hacking the metric by dichotomising their predictions. It serves no purpose above an F1 score.</p>\n<p>Why not just use AUC (like the SETI challenge), or even BCEloss (like DFDC)?</p>\n<p>Also, for what it's worth, I dislike the F1 score, too. Precision and recall, on which it's based, don't use the count of 'true negatives'. This isn't a problem when you have a 50:50 mix of the two classes, but in medicine you almost never do.</p>\n<p>For example, imagine you deploy your classification model in the real world and then find the proportion of negative (vs positive) cases increases. Then imagine you find your model ALSO performs better in terms of specificity (i.e. identifying a negative case as a true negative). In this situation, your F1 score might go up, might go down, or might even stay the same (despite an improvement in model!) This is because your True positives, False positives, and False negatives may remain balanced, with only the True negatives varying. This is immensely relevant in a screening test. Cohen's Kappa, although it has it's faults, at least is able to adjust for class imbalance.</p>\n<p>I personally cannot see the point behind this metric even outside of the competition space; it punishes probabilistic models and favours dichotomising tests, despite the former obviously being advantageous.</p>",
  "messages": [
    {
      "id": 2065354,
      "postDate": "2022-12-14T15:17:50.933Z",
      "content": "<p>The stated rationale for this metric is it is probabilistic. Well, it's not - everyone's hacking the metric by dichotomising their predictions. It serves no purpose above an F1 score.</p>\n<p>Why not just use AUC (like the SETI challenge), or even BCEloss (like DFDC)?</p>\n<p>Also, for what it's worth, I dislike the F1 score, too. Precision and recall, on which it's based, don't use the count of 'true negatives'. This isn't a problem when you have a 50:50 mix of the two classes, but in medicine you almost never do.</p>\n<p>For example, imagine you deploy your classification model in the real world and then find the proportion of negative (vs positive) cases increases. Then imagine you find your model ALSO performs better in terms of specificity (i.e. identifying a negative case as a true negative). In this situation, your F1 score might go up, might go down, or might even stay the same (despite an improvement in model!) This is because your True positives, False positives, and False negatives may remain balanced, with only the True negatives varying. This is immensely relevant in a screening test. Cohen's Kappa, although it has it's faults, at least is able to adjust for class imbalance.</p>\n<p>I personally cannot see the point behind this metric even outside of the competition space; it punishes probabilistic models and favours dichotomising tests, despite the former obviously being advantageous.</p>",
      "rawMarkdown": "The stated rationale for this metric is it is probabilistic. Well, it's not - everyone's hacking the metric by dichotomising their predictions. It serves no purpose above an F1 score.\n\nWhy not just use AUC (like the SETI challenge), or even BCEloss (like DFDC)?\n\nAlso, for what it's worth, I dislike the F1 score, too. Precision and recall, on which it's based, don't use the count of 'true negatives'. This isn't a problem when you have a 50:50 mix of the two classes, but in medicine you almost never do.\n\nFor example, imagine you deploy your classification model in the real world and then find the proportion of negative (vs positive) cases increases. Then imagine you find your model ALSO performs better in terms of specificity (i.e. identifying a negative case as a true negative). In this situation, your F1 score might go up, might go down, or might even stay the same (despite an improvement in model!) This is because your True positives, False positives, and False negatives may remain balanced, with only the True negatives varying. This is immensely relevant in a screening test. Cohen's Kappa, although it has it's faults, at least is able to adjust for class imbalance.\n\nI personally cannot see the point behind this metric even outside of the competition space; it punishes probabilistic models and favours dichotomising tests, despite the former obviously being advantageous.",
      "votes": 32
    },
    {
      "id": 2080720,
      "postDate": "2022-12-30T12:20:02.533Z",
      "content": "<p>Ok I'm going to come out and say it - this metric is a bad choice for this competition, and it makes training and validating models very annoying.</p>\n<p>I have just trained a new model that has a worse:</p>\n<ul>\n<li>Validation loss (unweighted BCE)</li>\n<li>Validation AUC (both with and without cutoffs)</li>\n<li>Validation accuracy at 0.5 cutoff</li>\n</ul>\n<p>But has a MUCH better PF1 (with optimal thresholding for each).<br>\nOld model: 0.418 @ 0.45 cutoff<br>\nNew model: 0.460 @ 0.60 cutoff</p>\n<p>I have delved into why this is (below is the data for a single illustrative validation fold, so 20% of the training dataset).</p>\n<pre><code>ROC_AUC_SCORE CONTINUOUS\n    OLD \n    NEW \n\nPFBETA CONTINUOUS\n    OLD \n    NEW \n\n\nOLD\n    pf1_cont \n    pf1_best \n    pf1_thresh \n\nNEW\n    pf1_cont \n    pf1_best \n    pf1_thresh \n\nROC_AUC_SCORE DISCREET (USING OPTIMAL PF1 CUTOFF)\n    OLD \n    NEW \n\nPFBETA DISCREET (USING OPTIMAL PF1 CUTOFF)\n    OLD \n    NEW \n\n\nCM OLD\n[[   ]\n [     ]]\n\nCM NEW\n[[   ]\n [     ]]\n\nNormalised CM OLD\n[[ ]\n [ ]]\n\nNormalised CM NEW\n[[ ]\n [ ]]\n</code></pre>\n<p>The last bit explains why.</p>\n<p>The new model is less sensitive but more specific; it does a worse job in identifying cancers, resulting in more false negatives and less true positives.</p>\n<p>So is this a problem? YES - this is a cancer screening test - and yet the metric is favouring a model with a lower <em>sensitivity</em>! This is is the opposite of what you want.</p>\n<p>Some of this, of course, is caused by the thresholding issue, but that doesn't make this any easier.</p>\n<p>Until now I've been using AUC to CV my models, but now I've realised I can't even do that. 0.418 -&gt; 0.60 is a HUGE difference in performance. And yet, the PF1 metric is soooo noisy. I guess I just have to check LB score for every model and overfit to that instead?</p>",
      "rawMarkdown": "Ok I'm going to come out and say it - this metric is a bad choice for this competition, and it makes training and validating models very annoying.\n\nI have just trained a new model that has a worse:\n- Validation loss (unweighted BCE)\n- Validation AUC (both with and without cutoffs)\n- Validation accuracy at 0.5 cutoff\n\nBut has a MUCH better PF1 (with optimal thresholding for each).\nOld model: 0.418 @ 0.45 cutoff\nNew model: 0.460 @ 0.60 cutoff\n\nI have delved into why this is (below is the data for a single illustrative validation fold, so 20% of the training dataset).\n\n\n```python\nROC_AUC_SCORE CONTINUOUS\n\tOLD\t0.905\n\tNEW\t0.876\n\nPFBETA CONTINUOUS\n\tOLD\t0.167\n\tNEW\t0.170\n\n\nOLD\n\tpf1_cont 0.1668043726025578\n\tpf1_best 0.4180327868852459\n\tpf1_thresh 0.44999999999999996\n\nNEW\n\tpf1_cont 0.17028606548001943\n\tpf1_best 0.45989304812834225\n\tpf1_thresh 0.6\n\nROC_AUC_SCORE DISCREET (USING OPTIMAL PF1 CUTOFF)\n\tOLD\t0.725\n\tNEW\t0.693\n\nPFBETA DISCREET (USING OPTIMAL PF1 CUTOFF)\n\tOLD\t0.418\n\tNEW\t0.460\n\n\nCM OLD\n[[4577   84]\n [  58   51]]\n\nCM NEW\n[[4626   35]\n [  66   43]]\n\nNormalised CM OLD\n[[0.98197812 0.01802188]\n [0.53211009 0.46788991]]\n\nNormalised CM NEW\n[[0.99249088 0.00750912]\n [0.60550459 0.39449541]]\n```\n\nThe last bit explains why.\n\nThe new model is less sensitive but more specific; it does a worse job in identifying cancers, resulting in more false negatives and less true positives.\n\nSo is this a problem? YES - this is a cancer screening test - and yet the metric is favouring a model with a lower *sensitivity*! This is is the opposite of what you want.\n\nSome of this, of course, is caused by the thresholding issue, but that doesn't make this any easier.\n\nUntil now I've been using AUC to CV my models, but now I've realised I can't even do that. 0.418 -> 0.60 is a HUGE difference in performance. And yet, the PF1 metric is soooo noisy. I guess I just have to check LB score for every model and overfit to that instead?",
      "votes": 26,
      "replies": [
        {
          "id": 2080828,
          "postDate": "2022-12-30T14:26:44.607Z",
          "rawMarkdown": "",
          "votes": -1,
          "isDeleted": true
        },
        {
          "id": 2080943,
          "postDate": "2022-12-30T16:27:52.457Z",
          "content": "<p>I am tracking precision, sensitivity, specificity, f1, pf1, ap, auc after validation and I noticed that raw predictions are more sensitive by itself. I agree that auc makes the most sense for this problem since there are different negatives as well. Organizers probably thought we could somehow come up with generating confident/not confident outputs so the metric could be fully utilized but it didn't happen.  </p>",
          "rawMarkdown": "I am tracking precision, sensitivity, specificity, f1, pf1, ap, auc after validation and I noticed that raw predictions are more sensitive by itself. I agree that auc makes the most sense for this problem since there are different negatives as well. Organizers probably thought we could somehow come up with generating confident/not confident outputs so the metric could be fully utilized but it didn't happen.  ",
          "votes": 4
        },
        {
          "id": 2081062,
          "postDate": "2022-12-30T19:23:07.857Z",
          "content": "<p>Yeah, the metric is really really problematic. It should have been changed to AUC after it became imminent that binarization works much better than raw probabilities which basically nullifies the original idea of the metric. </p>",
          "rawMarkdown": "Yeah, the metric is really really problematic. It should have been changed to AUC after it became imminent that binarization works much better than raw probabilities which basically nullifies the original idea of the metric. ",
          "votes": 12
        }
      ]
    },
    {
      "id": 2158050,
      "postDate": "2023-02-24T15:29:08.327Z",
      "content": "<p>I wish the metric would have been changed 2 months ago when these discussions started. There was still plenty of time left to not disrupt any early work.</p>\n<p>The metric obviously does not serve its original intended purpose, and a pure F1 metric with a target ratio of 2% is very problematic.</p>",
      "rawMarkdown": "I wish the metric would have been changed 2 months ago when these discussions started. There was still plenty of time left to not disrupt any early work.\n\nThe metric obviously does not serve its original intended purpose, and a pure F1 metric with a target ratio of 2% is very problematic.",
      "votes": 14,
      "replies": [
        {
          "id": 2160405,
          "postDate": "2023-02-26T16:40:17.653Z",
          "content": "<p>I stopped submitting scores a month ago (when I was in gold zone) as I lost faith in the competition. I'm sure I could have continued to improve, but it was become too difficult to gauge progress and I figured even if I did have the best model, it would be complete luck of the draw if I ended up with gold or not.</p>",
          "rawMarkdown": "I stopped submitting scores a month ago (when I was in gold zone) as I lost faith in the competition. I'm sure I could have continued to improve, but it was become too difficult to gauge progress and I figured even if I did have the best model, it would be complete luck of the draw if I ended up with gold or not.",
          "votes": 13
        }
      ]
    },
    {
      "id": 2082042,
      "postDate": "2023-01-01T03:22:53.610Z",
      "content": "<p>With pf1, it seems to me that a slight difference in the tuning of the thresholds can cause a large fluctuation in scores. With AUC, I don't think this problem occurs. I think that if you happen to set a threshold that fits private test data, your rankings can improve significantly, and vice versa. In my experiments, I feel that this is more dependent on luck than on the performance of the model.</p>",
      "rawMarkdown": "With pf1, it seems to me that a slight difference in the tuning of the thresholds can cause a large fluctuation in scores. With AUC, I don't think this problem occurs. I think that if you happen to set a threshold that fits private test data, your rankings can improve significantly, and vice versa. In my experiments, I feel that this is more dependent on luck than on the performance of the model.",
      "votes": 12,
      "replies": [
        {
          "id": 2082447,
          "postDate": "2023-01-01T15:35:06.397Z",
          "rawMarkdown": "",
          "votes": -7,
          "isDeleted": true,
          "replies": [
            {
              "id": 2082479,
              "postDate": "2023-01-01T16:02:04.417Z",
              "content": "<p>I'm afraid I don't agree with you. If the training dataset is 8000 patients, 16000 breasts, then with a 2% cancer rate we have a ~320 cancers. This probably means ~90 cancers in the public leaderboard, and 230 in the final.</p>\n<p>Because the F1 is the harmonic mean of precision and recall (or sensitivity, as doctors prefer to call it), then its use is limited by the variance of the sensitivity. We can work out the 95% confidence interval of the sensitivity using the binomial distribution.</p>\n<p>If our model's true sensitivity is 80%, then it will identify 184 of the 230 cases on average, but with a 95% confidence interval of 74% to 85%. This is an absolutely huge interval, and will almost certainly dwarf the between model differences in performance at the top of the LB.</p>",
              "rawMarkdown": "I'm afraid I don't agree with you. If the training dataset is 8000 patients, 16000 breasts, then with a 2% cancer rate we have a ~320 cancers. This probably means ~90 cancers in the public leaderboard, and 230 in the final.\n\nBecause the F1 is the harmonic mean of precision and recall (or sensitivity, as doctors prefer to call it), then its use is limited by the variance of the sensitivity. We can work out the 95% confidence interval of the sensitivity using the binomial distribution.\n\nIf our model's true sensitivity is 80%, then it will identify 184 of the 230 cases on average, but with a 95% confidence interval of 74% to 85%. This is an absolutely huge interval, and will almost certainly dwarf the between model differences in performance at the top of the LB.",
              "votes": 5
            },
            {
              "id": 2082546,
              "postDate": "2023-01-01T17:03:52.843Z",
              "rawMarkdown": "",
              "votes": -4,
              "isDeleted": true
            },
            {
              "id": 2082730,
              "postDate": "2023-01-01T21:12:53.253Z",
              "content": "<p>You need to look at it from the other side, there is a significantly high chance that the best model will not win due to the randomness of the thresholded metric.</p>",
              "rawMarkdown": "You need to look at it from the other side, there is a significantly high chance that the best model will not win due to the randomness of the thresholded metric.",
              "votes": 6
            },
            {
              "id": 2082795,
              "postDate": "2023-01-02T00:38:23.243Z",
              "content": "<p>As models perform better, pF1 scores will increase and fluctuate less. It is natural that this is the case. I don't think anyone probably has an objection there.</p>\n<p>However, usually in kaggle competitions, the top tier will be competing against very small differences in performance, and with the current metrics, it is possible that the randomness could change whether you get a gold or silver medal. </p>",
              "rawMarkdown": "As models perform better, pF1 scores will increase and fluctuate less. It is natural that this is the case. I don't think anyone probably has an objection there.\n\nHowever, usually in kaggle competitions, the top tier will be competing against very small differences in performance, and with the current metrics, it is possible that the randomness could change whether you get a gold or silver medal. ",
              "votes": 4
            }
          ]
        }
      ]
    },
    {
      "id": 2081542,
      "postDate": "2022-12-31T10:26:14.023Z",
      "content": "<p>I was interested in seeing if I could demonstrate at what point it becomes 'worth it' to submit probabilities rather than thresholded scores.</p>\n<p>I have just created a public notebook where I create a series of synthetic logistic regression datasets, then fit a simple model, and calculate:</p>\n<ul>\n<li>A traditional PF1 using the data</li>\n<li>A thresholded PF1 (F1)</li>\n</ul>\n<p>Generating logistic regression datasets is non trivial (I haven't modelled for non-normal prediction distributions, e.g. bimodal confidences) - and so maybe my assumptions don't hold - apologies if I've made a mistake.</p>\n<p><strong>However, across 3610 simulations (10 random seeds, 19 strengths of correlation between x and y_true, 19 levels of class imbalance), I cannot find a single run for which submitting probabilities is better than thresholded scores.</strong></p>\n<p><a href=\"https://www.kaggle.com/code/jamesphoward/pf1-testing-is-it-ever-better-not-to-threshold?scriptVersionId=115154528\" target=\"_blank\">https://www.kaggle.com/code/jamesphoward/pf1-testing-is-it-ever-better-not-to-threshold?scriptVersionId=115154528</a></p>\n<p>Again, huge apologies if my approach is naive and I have made an error.</p>",
      "rawMarkdown": "I was interested in seeing if I could demonstrate at what point it becomes 'worth it' to submit probabilities rather than thresholded scores.\n\nI have just created a public notebook where I create a series of synthetic logistic regression datasets, then fit a simple model, and calculate:\n- A traditional PF1 using the data\n- A thresholded PF1 (F1)\n\nGenerating logistic regression datasets is non trivial (I haven't modelled for non-normal prediction distributions, e.g. bimodal confidences) - and so maybe my assumptions don't hold - apologies if I've made a mistake.\n\n**However, across 3610 simulations (10 random seeds, 19 strengths of correlation between x and y_true, 19 levels of class imbalance), I cannot find a single run for which submitting probabilities is better than thresholded scores.**\n\nhttps://www.kaggle.com/code/jamesphoward/pf1-testing-is-it-ever-better-not-to-threshold?scriptVersionId=115154528\n\nAgain, huge apologies if my approach is naive and I have made an error.",
      "votes": 8,
      "replies": [
        {
          "id": 2081566,
          "postDate": "2022-12-31T11:16:26.087Z",
          "content": "<p>Feels a bit like folks are piling on, and no one has yet said what the issue is.  Granted, it might just be an f1 score, but that's not the end of the world, is it?  When there is a close tie, it'll be a little unfair probably, but I think any serious improvement in terms of fn/tp should prove out the winner.</p>",
          "rawMarkdown": "Feels a bit like folks are piling on, and no one has yet said what the issue is.  Granted, it might just be an f1 score, but that's not the end of the world, is it?  When there is a close tie, it'll be a little unfair probably, but I think any serious improvement in terms of fn/tp should prove out the winner.",
          "replies": [
            {
              "id": 2081571,
              "postDate": "2022-12-31T11:25:11.800Z",
              "content": "<p>No but it makes the competition difficult to CV, because thresholding with such a small number of positive cases makes things very brittle.</p>",
              "rawMarkdown": "No but it makes the competition difficult to CV, because thresholding with such a small number of positive cases makes things very brittle.",
              "votes": 3
            },
            {
              "id": 2081697,
              "postDate": "2022-12-31T14:21:04.997Z",
              "content": "<p>In general yes, but F1 score on a 2% positive rate can be a lottery in the end. And this would be a shame for such a competition.</p>",
              "rawMarkdown": "In general yes, but F1 score on a 2% positive rate can be a lottery in the end. And this would be a shame for such a competition.",
              "votes": 3
            },
            {
              "id": 2082214,
              "postDate": "2023-01-01T09:22:26.593Z",
              "rawMarkdown": "",
              "votes": -3,
              "isDeleted": true
            },
            {
              "id": 2082239,
              "postDate": "2023-01-01T10:17:28.203Z",
              "content": "<p>Most threshold-based metrics on very imbalanced data are problematic - and F1 score is one of them.</p>",
              "rawMarkdown": "Most threshold-based metrics on very imbalanced data are problematic - and F1 score is one of them.",
              "votes": 4
            },
            {
              "id": 2082350,
              "postDate": "2023-01-01T13:11:47.433Z",
              "content": "<p>\"I'd prefer labeled ROI boxes and Intersection over Union\"</p>\n<p>this is even more problematic … even for radiologist. e.g. just based on visual inspections, the true positive rate (PPV) for calcification patterns ranges from 10 to 70% </p>",
              "rawMarkdown": "\"I'd prefer labeled ROI boxes and Intersection over Union\"\n\nthis is even more problematic ... even for radiologist. e.g. just based on visual inspections, the true positive rate (PPV) for calcification patterns ranges from 10 to 70% "
            },
            {
              "id": 2082395,
              "postDate": "2023-01-01T14:24:10.723Z",
              "content": "<p>The problem with the comp is that it's difficult to know why our models are rating cancer.  It just sort of seems like augment and train and hope for the best.</p>\n<p>Perhaps some folks are quietly training ROI/patch level prediction based on the different presentations in BC, but given the doubt behind vindr (I can't even <a href=\"https://www.kaggle.com/discussions/product-feedback/375132\" target=\"_blank\">upload </a>the dataset for some reason), not sure what they'd use to train with.   Using grad cam was a suggested idea, but that seems at risk for FP patches.</p>\n<p>As for the downvotes, I plotted some pfbeta results and they seemed informative to me on our imbalanced dataset (assuming binarization =).  eg:</p>\n<pre><code> matplotlib.pyplot  plt\n\ngrounds= []*+[]*\nlabels = []*\nplts = []\n i  (,):\n    ngr = labels+ []* + []*+[]*(-i)+[]*i\n    plts.append(pfbeta(grounds, ngr, ))\nplt.title()\nplt.plot(plts)\nplt.show()\n\nplts = []\n i  (,):\n    ngr = labels +[]*(-i)+[]*i +  []*\n    plts.append(pfbeta(grounds, ngr, ))\nplt.title()\nplt.plot(plts)\nplt.show()\n\nplts = []\n i  (,):\n    ngr = labels +[]*(-i)+[]*i +  []*+[]*(-i)+[]*i\n    plts.append(pfbeta(grounds, ngr, ))\nplt.title()\nplt.plot(plts)\nplt.show()\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9052057%2F433bf23ad67361b20ab411d9c8053d92%2Fpfbetares.png?generation=1672582678182032&amp;alt=media\" alt=\"\"></p>\n<p>However, I can certainly see how if the models are having a tough time with detecting needles in the haystack, than these results would be very sensitive to thresholds.  </p>\n<p>I also appreciate that folks who have closely similar models aren't guaranteed a win by scoring marginally better on the leaderboard.  But, frankly, I think that is a good thing.  Flame me if you want, but we need more exploration over exploitation on Kaggle.  Getting slightly more than the next person in score shouldn't be what this is all about.  It's also really not clear to me that scoring marginally better is anything more than just overfitting slightly better anyways.</p>\n<p>The original issue above I still claim was caused by pushing specificity to the limit.  Yes, it will increase the score by sacrificing sensitivity as it did, but he's maxed out there now and has to fix the rest of the confusion matrix to gain further in score.   Downvote if you wish, but please do explain how I'm wrong there.</p>\n<p>I can't argue that AUC wouldn't be better, but I don't see how it would change things significantly enough to disrupt the comp.  Using f1 score here seems reasonable enough.</p>",
              "rawMarkdown": "The problem with the comp is that it's difficult to know why our models are rating cancer.  It just sort of seems like augment and train and hope for the best.\n\nPerhaps some folks are quietly training ROI/patch level prediction based on the different presentations in BC, but given the doubt behind vindr (I can't even [upload ](https://www.kaggle.com/discussions/product-feedback/375132)the dataset for some reason), not sure what they'd use to train with.   Using grad cam was a suggested idea, but that seems at risk for FP patches.\n\nAs for the downvotes, I plotted some pfbeta results and they seemed informative to me on our imbalanced dataset (assuming binarization =).  eg:\n\n```python\nimport matplotlib.pyplot as plt\n\ngrounds= [0]*15680+[1]*320\nlabels = [0]*15630\nplts = []\nfor i in range(0,50):\n    ngr = labels+ [0]*50 + [1]*270+[1]*(50-i)+[0]*i\n    plts.append(pfbeta(grounds, ngr, 1))\nplt.title(\"increasing fn, 0 to 50\")\nplt.plot(plts)\nplt.show()\n\nplts = []\nfor i in range(0,50):\n    ngr = labels +[0]*(50-i)+[1]*i +  [1]*320\n    plts.append(pfbeta(grounds, ngr, 1))\nplt.title(\"increasing fp, 0 to 50\")\nplt.plot(plts)\nplt.show()\n\nplts = []\nfor i in range(0,50):\n    ngr = labels +[0]*(50-i)+[1]*i +  [1]*270+[1]*(50-i)+[0]*i\n    plts.append(pfbeta(grounds, ngr, 1))\nplt.title(\"increasing fp and fn, 0 to 50\")\nplt.plot(plts)\nplt.show()\n\n\n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9052057%2F433bf23ad67361b20ab411d9c8053d92%2Fpfbetares.png?generation=1672582678182032&alt=media)\n\nHowever, I can certainly see how if the models are having a tough time with detecting needles in the haystack, than these results would be very sensitive to thresholds.  \n\nI also appreciate that folks who have closely similar models aren't guaranteed a win by scoring marginally better on the leaderboard.  But, frankly, I think that is a good thing.  Flame me if you want, but we need more exploration over exploitation on Kaggle.  Getting slightly more than the next person in score shouldn't be what this is all about.  It's also really not clear to me that scoring marginally better is anything more than just overfitting slightly better anyways.\n\nThe original issue above I still claim was caused by pushing specificity to the limit.  Yes, it will increase the score by sacrificing sensitivity as it did, but he's maxed out there now and has to fix the rest of the confusion matrix to gain further in score.   Downvote if you wish, but please do explain how I'm wrong there.\n\nI can't argue that AUC wouldn't be better, but I don't see how it would change things significantly enough to disrupt the comp.  Using f1 score here seems reasonable enough."
            }
          ]
        },
        {
          "id": 2081568,
          "postDate": "2022-12-31T11:19:35.203Z",
          "content": "<p>if you use bce loss, you predicted probability distribution is exponential.<br>\nfor exponential distribution, thresholded pf1 will give better results ( which can be proven mathematically i think)</p>\n<p>hence if we used thresholded probability, we end up the usually f1 score</p>",
          "rawMarkdown": "if you use bce loss, you predicted probability distribution is exponential.\nfor exponential distribution, thresholded pf1 will give better results ( which can be proven mathematically i think)\n\nhence if we used thresholded probability, we end up the usually f1 score",
          "votes": 1
        }
      ]
    },
    {
      "id": 2066847,
      "postDate": "2022-12-16T05:45:54.533Z",
      "content": "<p>There's a pretty simple answer: I didn't expect dichotomizing to be such a popular approach! Some degree of thresholding is to be expected, but it's happening to a greater extent than I had anticipated. I am still optimistic that people will submit more meaningful probabilities as their models get stronger and generate better calibrated predictions. However, that's pure speculation and quite possibly wishful thinking on my part.</p>\n<p>Generally speaking there are a lot of metrics that would have been reasonable choices and if switching between them were a costless option I would consider doing so. Given that switching metrics midstream would disrupt the work competitors have already done I don't think it would be fair to do so unless there were quite severe issues with the metric. As even a completely traditional F1 score provides a reasonable method of ranking submissions, I don't think we're at that point.</p>",
      "rawMarkdown": "There's a pretty simple answer: I didn't expect dichotomizing to be such a popular approach! Some degree of thresholding is to be expected, but it's happening to a greater extent than I had anticipated. I am still optimistic that people will submit more meaningful probabilities as their models get stronger and generate better calibrated predictions. However, that's pure speculation and quite possibly wishful thinking on my part.\n\nGenerally speaking there are a lot of metrics that would have been reasonable choices and if switching between them were a costless option I would consider doing so. Given that switching metrics midstream would disrupt the work competitors have already done I don't think it would be fair to do so unless there were quite severe issues with the metric. As even a completely traditional F1 score provides a reasonable method of ranking submissions, I don't think we're at that point.\n",
      "votes": 2,
      "replies": [
        {
          "id": 2066946,
          "postDate": "2022-12-16T08:16:53.813Z",
          "content": "<p>I have a suggestion maybe for future kaggle competition. </p>\n<p>In some CVPR or computer science conference competitions/benchmarking, etc , they use the main metric (for ranking solution) and other metrics  (for analysis, etc, and not for ranking).</p>\n<p>I am thinking LB can have multiple metric columns for ranking and non-ranking models.<br>\nBy studying the metrics and performances, we can make and suggest better ones in the future.</p>",
          "rawMarkdown": "I have a suggestion maybe for future kaggle competition. \n\nIn some CVPR or computer science conference competitions/benchmarking, etc , they use the main metric (for ranking solution) and other metrics  (for analysis, etc, and not for ranking).\n\nI am thinking LB can have multiple metric columns for ranking and non-ranking models.\nBy studying the metrics and performances, we can make and suggest better ones in the future.",
          "votes": 5,
          "replies": [
            {
              "id": 2067368,
              "postDate": "2022-12-16T16:33:19.347Z",
              "content": "<p>To reduce probing risks, these can be provided at the end of the competition.</p>",
              "rawMarkdown": "To reduce probing risks, these can be provided at the end of the competition."
            },
            {
              "id": 2067412,
              "postDate": "2022-12-16T17:17:34.247Z",
              "content": "<p>It would definitely be nice to provide additional metrics, at least after the close of a competition. It's just not something we've been able to prioritize over other engineering work in the past.</p>",
              "rawMarkdown": "It would definitely be nice to provide additional metrics, at least after the close of a competition. It's just not something we've been able to prioritize over other engineering work in the past.",
              "votes": 1
            },
            {
              "id": 2067639,
              "postDate": "2022-12-16T23:45:27.180Z",
              "content": "<p>Wouldn't it be a fairly basic / straightforward notebook?  At the very least just scoring the selected submissions.   Something like this would probably be sufficient</p>\n<p><a href=\"https://www.kaggle.com/code/ryanholbrook/feedback-prize-3-efficiency-leaderboard\" target=\"_blank\">https://www.kaggle.com/code/ryanholbrook/feedback-prize-3-efficiency-leaderboard</a></p>\n<p>In theory, you could probably do this on a weekly basis.  In this way probing risk would be reduced.</p>\n<p>eg:</p>\n<pre><code> sub  selected:\n    scorefunc  scorefuncs:\n     scorename, score = scorefunc\n        res.append((sub[], scorename, score(lb.loc[pubinds, ], \n            sub[].loc[pubinds, ])))\n</code></pre>\n<p>The advantage to this, btw, is that you could get folks to keep their 'selected' updated during the comp.  It may also encourage people to submit more frequently and select more intelligently.</p>\n<p>I added a feature suggestion based on this.   <a href=\"https://www.kaggle.com/discussions/product-feedback/372619\" target=\"_blank\">https://www.kaggle.com/discussions/product-feedback/372619</a>   </p>\n<p>That said, the weekly auxiliary scoring might not fly as great as that would be.  If I had to guess, a lot of the constraints on what scoring methods can be used/shown is due to probing risk.</p>",
              "rawMarkdown": "Wouldn't it be a fairly basic / straightforward notebook?  At the very least just scoring the selected submissions.   Something like this would probably be sufficient\n\nhttps://www.kaggle.com/code/ryanholbrook/feedback-prize-3-efficiency-leaderboard\n\nIn theory, you could probably do this on a weekly basis.  In this way probing risk would be reduced.\n\n\neg:\n\n```python\nfor sub in selected:\n   for scorefunc in scorefuncs:\n     scorename, score = scorefunc\n        res.append((sub['username'], scorename, score(lb.loc[pubinds, 'cancer'], \n            sub['data'].loc[pubinds, 'cancer'])))\n```\n\nThe advantage to this, btw, is that you could get folks to keep their 'selected' updated during the comp.  It may also encourage people to submit more frequently and select more intelligently.\n\nI added a feature suggestion based on this.   https://www.kaggle.com/discussions/product-feedback/372619   \n\nThat said, the weekly auxiliary scoring might not fly as great as that would be.  If I had to guess, a lot of the constraints on what scoring methods can be used/shown is due to probing risk."
            },
            {
              "id": 2070240,
              "postDate": "2022-12-19T18:25:00.503Z",
              "content": "<p>Calculating new scores is the easy part. The time consuming steps would involve dealing with access to the submissions (most of which aren't public), refactoring the database to store multiple metrics per competition, adding the concepts of primary and secondary metrics to the codebase, building a page to display the results, and so on and so on.</p>",
              "rawMarkdown": "Calculating new scores is the easy part. The time consuming steps would involve dealing with access to the submissions (most of which aren't public), refactoring the database to store multiple metrics per competition, adding the concepts of primary and secondary metrics to the codebase, building a page to display the results, and so on and so on.",
              "votes": 1
            },
            {
              "id": 2070246,
              "postDate": "2022-12-19T18:29:41.643Z",
              "content": "<p>Right, though I think we wouldn't really need the scores stored.  Just showing them from a notebook run similar to what they did for the efficiency metric in the feedback comp would be enough. </p>\n<p>That said, I can imagine getting access to the submissions could be tricky if those have been locked down extra tight, which is reasonable.</p>",
              "rawMarkdown": "Right, though I think we wouldn't really need the scores stored.  Just showing them from a notebook run similar to what they did for the efficiency metric in the feedback comp would be enough. \n\nThat said, I can imagine getting access to the submissions could be tricky if those have been locked down extra tight, which is reasonable."
            }
          ]
        },
        {
          "id": 2068290,
          "postDate": "2022-12-17T17:59:13.393Z",
          "content": "<p>I think in the future rather than a single metric, metric can be weweighte average of some of them.</p>",
          "rawMarkdown": "I think in the future rather than a single metric, metric can be weweighte average of some of them.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2162734,
      "postDate": "2023-02-28T11:42:27.467Z",
      "rawMarkdown": "",
      "votes": -3,
      "isDeleted": true,
      "replies": [
        {
          "id": 2162804,
          "postDate": "2023-02-28T12:17:27.953Z",
          "content": "<blockquote>\n  <p>As I suspected, there wasn't much shakeup. Relative to other Kaggle contests, the results largely reflected the leaderboard.</p>\n</blockquote>\n<p>Much more shakeup than it should have been.</p>\n<p>We (and I suspect others as well) had unselected submissions that could have ended in gold zone just by chance (lucky threshold/seed/approach) with roughly same CV score.  </p>\n<p>For us the actual ranking it did not hurt much but I imagine it was much more annoying for the top teams especially close to the prize winning positions.</p>\n<p>During the competition everyone was affected. Finding any significant local CV improvement was also much more difficult than it should have been…</p>",
          "rawMarkdown": ">As I suspected, there wasn't much shakeup. Relative to other Kaggle contests, the results largely reflected the leaderboard.\n\nMuch more shakeup than it should have been.\n\nWe (and I suspect others as well) had unselected submissions that could have ended in gold zone just by chance (lucky threshold/seed/approach) with roughly same CV score.  \n\nFor us the actual ranking it did not hurt much but I imagine it was much more annoying for the top teams especially close to the prize winning positions.\n\nDuring the competition everyone was affected. Finding any significant local CV improvement was also much more difficult than it should have been...",
          "votes": 9
        }
      ]
    },
    {
      "id": 2066884,
      "postDate": "2022-12-16T06:48:30.563Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2080720,
      "author_name": "James Howard",
      "author_url": "",
      "post_date": "2022-12-30T12:20:02.533000",
      "content": "<p>Ok I'm going to come out and say it - this metric is a bad choice for this competition, and it makes training and validating models very annoying.</p>\n<p>I have just trained a new model that has a worse:</p>\n<ul>\n<li>Validation loss (unweighted BCE)</li>\n<li>Validation AUC (both with and without cutoffs)</li>\n<li>Validation accuracy at 0.5 cutoff</li>\n</ul>\n<p>But has a MUCH better PF1 (with optimal thresholding for each).<br>\nOld model: 0.418 @ 0.45 cutoff<br>\nNew model: 0.460 @ 0.60 cutoff</p>\n<p>I have delved into why this is (below is the data for a single illustrative validation fold, so 20% of the training dataset).</p>\n<pre><code>ROC_AUC_SCORE CONTINUOUS\n    OLD \n    NEW \n\nPFBETA CONTINUOUS\n    OLD \n    NEW \n\n\nOLD\n    pf1_cont \n    pf1_best \n    pf1_thresh \n\nNEW\n    pf1_cont \n    pf1_best \n    pf1_thresh \n\nROC_AUC_SCORE DISCREET (USING OPTIMAL PF1 CUTOFF)\n    OLD \n    NEW \n\nPFBETA DISCREET (USING OPTIMAL PF1 CUTOFF)\n    OLD \n    NEW \n\n\nCM OLD\n[[   ]\n [     ]]\n\nCM NEW\n[[   ]\n [     ]]\n\nNormalised CM OLD\n[[ ]\n [ ]]\n\nNormalised CM NEW\n[[ ]\n [ ]]\n</code></pre>\n<p>The last bit explains why.</p>\n<p>The new model is less sensitive but more specific; it does a worse job in identifying cancers, resulting in more false negatives and less true positives.</p>\n<p>So is this a problem? YES - this is a cancer screening test - and yet the metric is favouring a model with a lower <em>sensitivity</em>! This is is the opposite of what you want.</p>\n<p>Some of this, of course, is caused by the thresholding issue, but that doesn't make this any easier.</p>\n<p>Until now I've been using AUC to CV my models, but now I've realised I can't even do that. 0.418 -&gt; 0.60 is a HUGE difference in performance. And yet, the PF1 metric is soooo noisy. I guess I just have to check LB score for every model and overfit to that instead?</p>",
      "votes": 26,
      "replies": [
        {
          "id": 2080828,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-12-30T14:26:44.607000",
          "content": "",
          "votes": -1,
          "replies": []
        },
        {
          "id": 2080943,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2022-12-30T16:27:52.457000",
          "content": "<p>I am tracking precision, sensitivity, specificity, f1, pf1, ap, auc after validation and I noticed that raw predictions are more sensitive by itself. I agree that auc makes the most sense for this problem since there are different negatives as well. Organizers probably thought we could somehow come up with generating confident/not confident outputs so the metric could be fully utilized but it didn't happen.  </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 2081062,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2022-12-30T19:23:07.857000",
          "content": "<p>Yeah, the metric is really really problematic. It should have been changed to AUC after it became imminent that binarization works much better than raw probabilities which basically nullifies the original idea of the metric. </p>",
          "votes": 12,
          "replies": []
        }
      ]
    },
    {
      "id": 2158050,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2023-02-24T15:29:08.327000",
      "content": "<p>I wish the metric would have been changed 2 months ago when these discussions started. There was still plenty of time left to not disrupt any early work.</p>\n<p>The metric obviously does not serve its original intended purpose, and a pure F1 metric with a target ratio of 2% is very problematic.</p>",
      "votes": 14,
      "replies": [
        {
          "id": 2160405,
          "author_name": "James Howard",
          "author_url": "",
          "post_date": "2023-02-26T16:40:17.653000",
          "content": "<p>I stopped submitting scores a month ago (when I was in gold zone) as I lost faith in the competition. I'm sure I could have continued to improve, but it was become too difficult to gauge progress and I figured even if I did have the best model, it would be complete luck of the draw if I ended up with gold or not.</p>",
          "votes": 13,
          "replies": []
        }
      ]
    },
    {
      "id": 2082042,
      "author_name": "YYama",
      "author_url": "",
      "post_date": "2023-01-01T03:22:53.610000",
      "content": "<p>With pf1, it seems to me that a slight difference in the tuning of the thresholds can cause a large fluctuation in scores. With AUC, I don't think this problem occurs. I think that if you happen to set a threshold that fits private test data, your rankings can improve significantly, and vice versa. In my experiments, I feel that this is more dependent on luck than on the performance of the model.</p>",
      "votes": 12,
      "replies": [
        {
          "id": 2082447,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-01-01T15:35:06.397000",
          "content": "",
          "votes": -7,
          "replies": [
            {
              "id": 2082479,
              "author_name": "James Howard",
              "author_url": "",
              "post_date": "2023-01-01T16:02:04.417000",
              "content": "<p>I'm afraid I don't agree with you. If the training dataset is 8000 patients, 16000 breasts, then with a 2% cancer rate we have a ~320 cancers. This probably means ~90 cancers in the public leaderboard, and 230 in the final.</p>\n<p>Because the F1 is the harmonic mean of precision and recall (or sensitivity, as doctors prefer to call it), then its use is limited by the variance of the sensitivity. We can work out the 95% confidence interval of the sensitivity using the binomial distribution.</p>\n<p>If our model's true sensitivity is 80%, then it will identify 184 of the 230 cases on average, but with a 95% confidence interval of 74% to 85%. This is an absolutely huge interval, and will almost certainly dwarf the between model differences in performance at the top of the LB.</p>",
              "votes": 5,
              "replies": []
            },
            {
              "id": 2082546,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-01-01T17:03:52.843000",
              "content": "",
              "votes": -4,
              "replies": []
            },
            {
              "id": 2082730,
              "author_name": "Psi",
              "author_url": "",
              "post_date": "2023-01-01T21:12:53.253000",
              "content": "<p>You need to look at it from the other side, there is a significantly high chance that the best model will not win due to the randomness of the thresholded metric.</p>",
              "votes": 6,
              "replies": []
            },
            {
              "id": 2082795,
              "author_name": "YYama",
              "author_url": "",
              "post_date": "2023-01-02T00:38:23.243000",
              "content": "<p>As models perform better, pF1 scores will increase and fluctuate less. It is natural that this is the case. I don't think anyone probably has an objection there.</p>\n<p>However, usually in kaggle competitions, the top tier will be competing against very small differences in performance, and with the current metrics, it is possible that the randomness could change whether you get a gold or silver medal. </p>",
              "votes": 4,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2081542,
      "author_name": "James Howard",
      "author_url": "",
      "post_date": "2022-12-31T10:26:14.023000",
      "content": "<p>I was interested in seeing if I could demonstrate at what point it becomes 'worth it' to submit probabilities rather than thresholded scores.</p>\n<p>I have just created a public notebook where I create a series of synthetic logistic regression datasets, then fit a simple model, and calculate:</p>\n<ul>\n<li>A traditional PF1 using the data</li>\n<li>A thresholded PF1 (F1)</li>\n</ul>\n<p>Generating logistic regression datasets is non trivial (I haven't modelled for non-normal prediction distributions, e.g. bimodal confidences) - and so maybe my assumptions don't hold - apologies if I've made a mistake.</p>\n<p><strong>However, across 3610 simulations (10 random seeds, 19 strengths of correlation between x and y_true, 19 levels of class imbalance), I cannot find a single run for which submitting probabilities is better than thresholded scores.</strong></p>\n<p><a href=\"https://www.kaggle.com/code/jamesphoward/pf1-testing-is-it-ever-better-not-to-threshold?scriptVersionId=115154528\" target=\"_blank\">https://www.kaggle.com/code/jamesphoward/pf1-testing-is-it-ever-better-not-to-threshold?scriptVersionId=115154528</a></p>\n<p>Again, huge apologies if my approach is naive and I have made an error.</p>",
      "votes": 8,
      "replies": [
        {
          "id": 2081566,
          "author_name": "@kaggleqrdl",
          "author_url": "",
          "post_date": "2022-12-31T11:16:26.087000",
          "content": "<p>Feels a bit like folks are piling on, and no one has yet said what the issue is.  Granted, it might just be an f1 score, but that's not the end of the world, is it?  When there is a close tie, it'll be a little unfair probably, but I think any serious improvement in terms of fn/tp should prove out the winner.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2081571,
              "author_name": "James Howard",
              "author_url": "",
              "post_date": "2022-12-31T11:25:11.800000",
              "content": "<p>No but it makes the competition difficult to CV, because thresholding with such a small number of positive cases makes things very brittle.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2081697,
              "author_name": "Psi",
              "author_url": "",
              "post_date": "2022-12-31T14:21:04.997000",
              "content": "<p>In general yes, but F1 score on a 2% positive rate can be a lottery in the end. And this would be a shame for such a competition.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2082214,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-01-01T09:22:26.593000",
              "content": "",
              "votes": -3,
              "replies": []
            },
            {
              "id": 2082239,
              "author_name": "Psi",
              "author_url": "",
              "post_date": "2023-01-01T10:17:28.203000",
              "content": "<p>Most threshold-based metrics on very imbalanced data are problematic - and F1 score is one of them.</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2082350,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-01-01T13:11:47.433000",
              "content": "<p>\"I'd prefer labeled ROI boxes and Intersection over Union\"</p>\n<p>this is even more problematic … even for radiologist. e.g. just based on visual inspections, the true positive rate (PPV) for calcification patterns ranges from 10 to 70% </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2082395,
              "author_name": "@kaggleqrdl",
              "author_url": "",
              "post_date": "2023-01-01T14:24:10.723000",
              "content": "<p>The problem with the comp is that it's difficult to know why our models are rating cancer.  It just sort of seems like augment and train and hope for the best.</p>\n<p>Perhaps some folks are quietly training ROI/patch level prediction based on the different presentations in BC, but given the doubt behind vindr (I can't even <a href=\"https://www.kaggle.com/discussions/product-feedback/375132\" target=\"_blank\">upload </a>the dataset for some reason), not sure what they'd use to train with.   Using grad cam was a suggested idea, but that seems at risk for FP patches.</p>\n<p>As for the downvotes, I plotted some pfbeta results and they seemed informative to me on our imbalanced dataset (assuming binarization =).  eg:</p>\n<pre><code> matplotlib.pyplot  plt\n\ngrounds= []*+[]*\nlabels = []*\nplts = []\n i  (,):\n    ngr = labels+ []* + []*+[]*(-i)+[]*i\n    plts.append(pfbeta(grounds, ngr, ))\nplt.title()\nplt.plot(plts)\nplt.show()\n\nplts = []\n i  (,):\n    ngr = labels +[]*(-i)+[]*i +  []*\n    plts.append(pfbeta(grounds, ngr, ))\nplt.title()\nplt.plot(plts)\nplt.show()\n\nplts = []\n i  (,):\n    ngr = labels +[]*(-i)+[]*i +  []*+[]*(-i)+[]*i\n    plts.append(pfbeta(grounds, ngr, ))\nplt.title()\nplt.plot(plts)\nplt.show()\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9052057%2F433bf23ad67361b20ab411d9c8053d92%2Fpfbetares.png?generation=1672582678182032&amp;alt=media\" alt=\"\"></p>\n<p>However, I can certainly see how if the models are having a tough time with detecting needles in the haystack, than these results would be very sensitive to thresholds.  </p>\n<p>I also appreciate that folks who have closely similar models aren't guaranteed a win by scoring marginally better on the leaderboard.  But, frankly, I think that is a good thing.  Flame me if you want, but we need more exploration over exploitation on Kaggle.  Getting slightly more than the next person in score shouldn't be what this is all about.  It's also really not clear to me that scoring marginally better is anything more than just overfitting slightly better anyways.</p>\n<p>The original issue above I still claim was caused by pushing specificity to the limit.  Yes, it will increase the score by sacrificing sensitivity as it did, but he's maxed out there now and has to fix the rest of the confusion matrix to gain further in score.   Downvote if you wish, but please do explain how I'm wrong there.</p>\n<p>I can't argue that AUC wouldn't be better, but I don't see how it would change things significantly enough to disrupt the comp.  Using f1 score here seems reasonable enough.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2081568,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2022-12-31T11:19:35.203000",
          "content": "<p>if you use bce loss, you predicted probability distribution is exponential.<br>\nfor exponential distribution, thresholded pf1 will give better results ( which can be proven mathematically i think)</p>\n<p>hence if we used thresholded probability, we end up the usually f1 score</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2066847,
      "author_name": "Sohier Dane",
      "author_url": "",
      "post_date": "2022-12-16T05:45:54.533000",
      "content": "<p>There's a pretty simple answer: I didn't expect dichotomizing to be such a popular approach! Some degree of thresholding is to be expected, but it's happening to a greater extent than I had anticipated. I am still optimistic that people will submit more meaningful probabilities as their models get stronger and generate better calibrated predictions. However, that's pure speculation and quite possibly wishful thinking on my part.</p>\n<p>Generally speaking there are a lot of metrics that would have been reasonable choices and if switching between them were a costless option I would consider doing so. Given that switching metrics midstream would disrupt the work competitors have already done I don't think it would be fair to do so unless there were quite severe issues with the metric. As even a completely traditional F1 score provides a reasonable method of ranking submissions, I don't think we're at that point.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2066946,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2022-12-16T08:16:53.813000",
          "content": "<p>I have a suggestion maybe for future kaggle competition. </p>\n<p>In some CVPR or computer science conference competitions/benchmarking, etc , they use the main metric (for ranking solution) and other metrics  (for analysis, etc, and not for ranking).</p>\n<p>I am thinking LB can have multiple metric columns for ranking and non-ranking models.<br>\nBy studying the metrics and performances, we can make and suggest better ones in the future.</p>",
          "votes": 5,
          "replies": [
            {
              "id": 2067368,
              "author_name": "@kaggleqrdl",
              "author_url": "",
              "post_date": "2022-12-16T16:33:19.347000",
              "content": "<p>To reduce probing risks, these can be provided at the end of the competition.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2067412,
              "author_name": "Sohier Dane",
              "author_url": "",
              "post_date": "2022-12-16T17:17:34.247000",
              "content": "<p>It would definitely be nice to provide additional metrics, at least after the close of a competition. It's just not something we've been able to prioritize over other engineering work in the past.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2067639,
              "author_name": "@kaggleqrdl",
              "author_url": "",
              "post_date": "2022-12-16T23:45:27.180000",
              "content": "<p>Wouldn't it be a fairly basic / straightforward notebook?  At the very least just scoring the selected submissions.   Something like this would probably be sufficient</p>\n<p><a href=\"https://www.kaggle.com/code/ryanholbrook/feedback-prize-3-efficiency-leaderboard\" target=\"_blank\">https://www.kaggle.com/code/ryanholbrook/feedback-prize-3-efficiency-leaderboard</a></p>\n<p>In theory, you could probably do this on a weekly basis.  In this way probing risk would be reduced.</p>\n<p>eg:</p>\n<pre><code> sub  selected:\n    scorefunc  scorefuncs:\n     scorename, score = scorefunc\n        res.append((sub[], scorename, score(lb.loc[pubinds, ], \n            sub[].loc[pubinds, ])))\n</code></pre>\n<p>The advantage to this, btw, is that you could get folks to keep their 'selected' updated during the comp.  It may also encourage people to submit more frequently and select more intelligently.</p>\n<p>I added a feature suggestion based on this.   <a href=\"https://www.kaggle.com/discussions/product-feedback/372619\" target=\"_blank\">https://www.kaggle.com/discussions/product-feedback/372619</a>   </p>\n<p>That said, the weekly auxiliary scoring might not fly as great as that would be.  If I had to guess, a lot of the constraints on what scoring methods can be used/shown is due to probing risk.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2070240,
              "author_name": "Sohier Dane",
              "author_url": "",
              "post_date": "2022-12-19T18:25:00.503000",
              "content": "<p>Calculating new scores is the easy part. The time consuming steps would involve dealing with access to the submissions (most of which aren't public), refactoring the database to store multiple metrics per competition, adding the concepts of primary and secondary metrics to the codebase, building a page to display the results, and so on and so on.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2070246,
              "author_name": "@kaggleqrdl",
              "author_url": "",
              "post_date": "2022-12-19T18:29:41.643000",
              "content": "<p>Right, though I think we wouldn't really need the scores stored.  Just showing them from a notebook run similar to what they did for the efficiency metric in the feedback comp would be enough. </p>\n<p>That said, I can imagine getting access to the submissions could be tricky if those have been locked down extra tight, which is reasonable.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2068290,
          "author_name": "DeepUnderstanding",
          "author_url": "",
          "post_date": "2022-12-17T17:59:13.393000",
          "content": "<p>I think in the future rather than a single metric, metric can be weweighte average of some of them.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2162734,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-02-28T11:42:27.467000",
      "content": "",
      "votes": -3,
      "replies": [
        {
          "id": 2162804,
          "author_name": "beluga",
          "author_url": "",
          "post_date": "2023-02-28T12:17:27.953000",
          "content": "<blockquote>\n  <p>As I suspected, there wasn't much shakeup. Relative to other Kaggle contests, the results largely reflected the leaderboard.</p>\n</blockquote>\n<p>Much more shakeup than it should have been.</p>\n<p>We (and I suspect others as well) had unselected submissions that could have ended in gold zone just by chance (lucky threshold/seed/approach) with roughly same CV score.  </p>\n<p>For us the actual ranking it did not hurt much but I imagine it was much more annoying for the top teams especially close to the prize winning positions.</p>\n<p>During the competition everyone was affected. Finding any significant local CV improvement was also much more difficult than it should have been…</p>",
          "votes": 9,
          "replies": []
        }
      ]
    },
    {
      "id": 2066884,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-12-16T06:48:30.563000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2065354": "The stated rationale for this metric is it is probabilistic. Well, it's not - everyone's hacking the metric by dichotomising their predictions. It serves no purpose above an F1 score.\n\nWhy not just use AUC (like the SETI challenge), or even BCEloss (like DFDC)?\n\nAlso, for what it's worth, I dislike the F1 score, too. Precision and recall, on which it's based, don't use the count of 'true negatives'. This isn't a problem when you have a 50:50 mix of the two classes, but in medicine you almost never do.\n\nFor example, imagine you deploy your classification model in the real world and then find the proportion of negative (vs positive) cases increases. Then imagine you find your model ALSO performs better in terms of specificity (i.e. identifying a negative case as a true negative). In this situation, your F1 score might go up, might go down, or might even stay the same (despite an improvement in model!) This is because your True positives, False positives, and False negatives may remain balanced, with only the True negatives varying. This is immensely relevant in a screening test. Cohen's Kappa, although it has it's faults, at least is able to adjust for class imbalance.\n\nI personally cannot see the point behind this metric even outside of the competition space; it punishes probabilistic models and favours dichotomising tests, despite the former obviously being advantageous.",
    "2080720": "Ok I'm going to come out and say it - this metric is a bad choice for this competition, and it makes training and validating models very annoying.\n\nI have just trained a new model that has a worse:\n- Validation loss (unweighted BCE)\n- Validation AUC (both with and without cutoffs)\n- Validation accuracy at 0.5 cutoff\n\nBut has a MUCH better PF1 (with optimal thresholding for each).\nOld model: 0.418 @ 0.45 cutoff\nNew model: 0.460 @ 0.60 cutoff\n\nI have delved into why this is (below is the data for a single illustrative validation fold, so 20% of the training dataset).\n\n\n```python\nROC_AUC_SCORE CONTINUOUS\n\tOLD\t0.905\n\tNEW\t0.876\n\nPFBETA CONTINUOUS\n\tOLD\t0.167\n\tNEW\t0.170\n\n\nOLD\n\tpf1_cont 0.1668043726025578\n\tpf1_best 0.4180327868852459\n\tpf1_thresh 0.44999999999999996\n\nNEW\n\tpf1_cont 0.17028606548001943\n\tpf1_best 0.45989304812834225\n\tpf1_thresh 0.6\n\nROC_AUC_SCORE DISCREET (USING OPTIMAL PF1 CUTOFF)\n\tOLD\t0.725\n\tNEW\t0.693\n\nPFBETA DISCREET (USING OPTIMAL PF1 CUTOFF)\n\tOLD\t0.418\n\tNEW\t0.460\n\n\nCM OLD\n[[4577   84]\n [  58   51]]\n\nCM NEW\n[[4626   35]\n [  66   43]]\n\nNormalised CM OLD\n[[0.98197812 0.01802188]\n [0.53211009 0.46788991]]\n\nNormalised CM NEW\n[[0.99249088 0.00750912]\n [0.60550459 0.39449541]]\n```\n\nThe last bit explains why.\n\nThe new model is less sensitive but more specific; it does a worse job in identifying cancers, resulting in more false negatives and less true positives.\n\nSo is this a problem? YES - this is a cancer screening test - and yet the metric is favouring a model with a lower *sensitivity*! This is is the opposite of what you want.\n\nSome of this, of course, is caused by the thresholding issue, but that doesn't make this any easier.\n\nUntil now I've been using AUC to CV my models, but now I've realised I can't even do that. 0.418 -> 0.60 is a HUGE difference in performance. And yet, the PF1 metric is soooo noisy. I guess I just have to check LB score for every model and overfit to that instead?",
    "2158050": "I wish the metric would have been changed 2 months ago when these discussions started. There was still plenty of time left to not disrupt any early work.\n\nThe metric obviously does not serve its original intended purpose, and a pure F1 metric with a target ratio of 2% is very problematic.",
    "2082042": "With pf1, it seems to me that a slight difference in the tuning of the thresholds can cause a large fluctuation in scores. With AUC, I don't think this problem occurs. I think that if you happen to set a threshold that fits private test data, your rankings can improve significantly, and vice versa. In my experiments, I feel that this is more dependent on luck than on the performance of the model.",
    "2081542": "I was interested in seeing if I could demonstrate at what point it becomes 'worth it' to submit probabilities rather than thresholded scores.\n\nI have just created a public notebook where I create a series of synthetic logistic regression datasets, then fit a simple model, and calculate:\n- A traditional PF1 using the data\n- A thresholded PF1 (F1)\n\nGenerating logistic regression datasets is non trivial (I haven't modelled for non-normal prediction distributions, e.g. bimodal confidences) - and so maybe my assumptions don't hold - apologies if I've made a mistake.\n\n**However, across 3610 simulations (10 random seeds, 19 strengths of correlation between x and y_true, 19 levels of class imbalance), I cannot find a single run for which submitting probabilities is better than thresholded scores.**\n\nhttps://www.kaggle.com/code/jamesphoward/pf1-testing-is-it-ever-better-not-to-threshold?scriptVersionId=115154528\n\nAgain, huge apologies if my approach is naive and I have made an error.",
    "2066847": "There's a pretty simple answer: I didn't expect dichotomizing to be such a popular approach! Some degree of thresholding is to be expected, but it's happening to a greater extent than I had anticipated. I am still optimistic that people will submit more meaningful probabilities as their models get stronger and generate better calibrated predictions. However, that's pure speculation and quite possibly wishful thinking on my part.\n\nGenerally speaking there are a lot of metrics that would have been reasonable choices and if switching between them were a costless option I would consider doing so. Given that switching metrics midstream would disrupt the work competitors have already done I don't think it would be fair to do so unless there were quite severe issues with the metric. As even a completely traditional F1 score provides a reasonable method of ranking submissions, I don't think we're at that point.\n",
    "2162734": "",
    "2066884": ""
  }
}