{
  "id": 390958,
  "title": "[A funny story]The shake is not that big",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/390958",
  "author_name": "ForcewithMe",
  "post_date": "2023-02-28T00:23:15.106000",
  "votes": 19,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Thanks to the organizers and congratulations to the winners!</p>\n<p>We're not going to write the solution right now. Just share a funny story. Up until a week before the end of the competition, we were mystified by the high scores on public leaderboard. Because our single model only has a score of about 0.5 on CV. Until the other day, we ran simulations on the training set. This simulation is based on the fact that the score of site2 is much higher than that of site1. We guess that the high score of public leaderboard is related to the ratio of site1 and site2, as well as the proportion of positive.</p>\n<ol>\n<li><p>Divide public and private in the same proportion as kaggle.</p></li>\n<li><p>Only 40% of positive samples and 65% of negative site1 samples in public are retained.</p></li>\n<li><p>Use site2 in private to fill in the missing samples in public</p></li>\n<li><p>Add the samples removed from public to private</p></li>\n</ol>\n<p>Through the above operation, our public score finally exceeds 0.6. After that, we decided to forget kaggle public leaderboard. Our final submission was based largely on the CV and private scores in our simulation. We expect our final one submission to be about 0.524 and ultimately it's 0.53 on kaggle's private leaderboard.</p>\n<p>But!!!!!!!!!!! Our highest public submission, although CV was not that high and the simulated pvt was not that good, kaggle private leaderboard was actually higher than our carefully selected ones. I don't understand.. </p>",
  "messages": [
    {
      "id": 2162041,
      "postDate": "2023-02-28T00:23:15.107Z",
      "content": "<p>Thanks to the organizers and congratulations to the winners!</p>\n<p>We're not going to write the solution right now. Just share a funny story. Up until a week before the end of the competition, we were mystified by the high scores on public leaderboard. Because our single model only has a score of about 0.5 on CV. Until the other day, we ran simulations on the training set. This simulation is based on the fact that the score of site2 is much higher than that of site1. We guess that the high score of public leaderboard is related to the ratio of site1 and site2, as well as the proportion of positive.</p>\n<ol>\n<li><p>Divide public and private in the same proportion as kaggle.</p></li>\n<li><p>Only 40% of positive samples and 65% of negative site1 samples in public are retained.</p></li>\n<li><p>Use site2 in private to fill in the missing samples in public</p></li>\n<li><p>Add the samples removed from public to private</p></li>\n</ol>\n<p>Through the above operation, our public score finally exceeds 0.6. After that, we decided to forget kaggle public leaderboard. Our final submission was based largely on the CV and private scores in our simulation. We expect our final one submission to be about 0.524 and ultimately it's 0.53 on kaggle's private leaderboard.</p>\n<p>But!!!!!!!!!!! Our highest public submission, although CV was not that high and the simulated pvt was not that good, kaggle private leaderboard was actually higher than our carefully selected ones. I don't understand.. </p>",
      "rawMarkdown": "Thanks to the organizers and congratulations to the winners!\n\nWe're not going to write the solution right now. Just share a funny story. Up until a week before the end of the competition, we were mystified by the high scores on public leaderboard. Because our single model only has a score of about 0.5 on CV. Until the other day, we ran simulations on the training set. This simulation is based on the fact that the score of site2 is much higher than that of site1. We guess that the high score of public leaderboard is related to the ratio of site1 and site2, as well as the proportion of positive.\n\n1. Divide public and private in the same proportion as kaggle.\n\n2. Only 40% of positive samples and 65% of negative site1 samples in public are retained.\n\n3. Use site2 in private to fill in the missing samples in public\n\n4. Add the samples removed from public to private\n\nThrough the above operation, our public score finally exceeds 0.6. After that, we decided to forget kaggle public leaderboard. Our final submission was based largely on the CV and private scores in our simulation. We expect our final one submission to be about 0.524 and ultimately it's 0.53 on kaggle's private leaderboard.\n\nBut!!!!!!!!!!! Our highest public submission, although CV was not that high and the simulated pvt was not that good, kaggle private leaderboard was actually higher than our carefully selected ones. I don't understand.. ",
      "votes": 19
    },
    {
      "id": 2162045,
      "postDate": "2023-02-28T00:28:25.930Z",
      "content": "<p>Private LB still has a large random range. While the scores are now much closer to actual CV numbers, we would still have subs with 0.53 private LB that have much lower CV score than our selection. And I am sure many teams have unselected high scoring private subs, that did not make too much sense to select based on CV.</p>\n<p>You need to have a strong model/ensemble, and then also be lucky with the threshold to score high here.</p>",
      "rawMarkdown": "Private LB still has a large random range. While the scores are now much closer to actual CV numbers, we would still have subs with 0.53 private LB that have much lower CV score than our selection. And I am sure many teams have unselected high scoring private subs, that did not make too much sense to select based on CV.\n\nYou need to have a strong model/ensemble, and then also be lucky with the threshold to score high here.",
      "votes": 7,
      "replies": [
        {
          "id": 2162051,
          "postDate": "2023-02-28T00:37:22.960Z",
          "content": "<p>Yeah, one lucky thing for us is that the threshold is stable </p>",
          "rawMarkdown": "Yeah, one lucky thing for us is that the threshold is stable ",
          "replies": [
            {
              "id": 2162220,
              "postDate": "2023-02-28T04:13:36.303Z",
              "content": "<p>Congrats <a href=\"https://www.kaggle.com/forcewithme\" target=\"_blank\">@forcewithme</a> Many unselected good scoring submissions for me, I took the bet on cv scores. Are thresholds stable for you or in general? I can still see +- 0.05 private lb over slight changes in thresholds. </p>",
              "rawMarkdown": "Congrats @forcewithme Many unselected good scoring submissions for me, I took the bet on cv scores. Are thresholds stable for you or in general? I can still see +- 0.05 private lb over slight changes in thresholds. ",
              "votes": 1
            },
            {
              "id": 2162228,
              "postDate": "2023-02-28T04:18:10.080Z",
              "content": "<p>Thank you. We always select the best threshold on CV. I haven't tested how it performed on private LB. If someone probs the public LB to get the threshold in this competition, I think he will suffer a huge shake.</p>",
              "rawMarkdown": "Thank you. We always select the best threshold on CV. I haven't tested how it performed on private LB. If someone probs the public LB to get the threshold in this competition, I think he will suffer a huge shake."
            }
          ]
        },
        {
          "id": 2162055,
          "postDate": "2023-02-28T00:42:11.290Z",
          "content": "<p>Private and public lb scores are almost uncorrelated, but our best public sub. is also best private (0.69 public, 0.54 private). However, I didn't dare to select it since it only contained models' weights + optimized threshold from our first 2 folds.<br>\nP/s: I completely disliked the pf1 metric 😓</p>",
          "rawMarkdown": "Private and public lb scores are almost uncorrelated, but our best public sub. is also best private (0.69 public, 0.54 private). However, I didn't dare to select it since it only contained models' weights + optimized threshold from our first 2 folds.\nP/s: I completely disliked the pf1 metric 😓",
          "votes": 6,
          "replies": [
            {
              "id": 2162058,
              "postDate": "2023-02-28T00:47:52.997Z",
              "content": "<p>Yes, I don’t think it’s a good metric.</p>",
              "rawMarkdown": "Yes, I don’t think it’s a good metric.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2162050,
      "postDate": "2023-02-28T00:35:35.823Z",
      "content": "<p>We got more than 10 subs ranging from .46 cv to cv worse than that scoring .5 or above (gold medal range), we focused on improving cv, improved oof to .523, but because of stabality chose our .508 as the final sub which public lb scores in range from .65-.64 mostly, private did not follow our cv at all</p>",
      "rawMarkdown": "We got more than 10 subs ranging from .46 cv to cv worse than that scoring .5 or above (gold medal range), we focused on improving cv, improved oof to .523, but because of stabality chose our .508 as the final sub which public lb scores in range from .65-.64 mostly, private did not follow our cv at all",
      "votes": 6
    },
    {
      "id": 2162062,
      "postDate": "2023-02-28T00:56:53.557Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/forcewithme\" target=\"_blank\">@forcewithme</a> and team. I only had a couple of weekends to spend on this competition but jumped in anyways 13 days ago. I am happy with where my model ended up although I have one submission that would have put me a few points higher but still in the bronze range. I agree the metric is bonkers.</p>",
      "rawMarkdown": "Congrats @forcewithme and team. I only had a couple of weekends to spend on this competition but jumped in anyways 13 days ago. I am happy with where my model ended up although I have one submission that would have put me a few points higher but still in the bronze range. I agree the metric is bonkers.",
      "votes": 1,
      "replies": [
        {
          "id": 2162109,
          "postDate": "2023-02-28T02:11:43.777Z",
          "content": "<p>thank youuu</p>",
          "rawMarkdown": "thank youuu",
          "votes": 1
        }
      ]
    },
    {
      "id": 2162865,
      "postDate": "2023-02-28T13:02:10.753Z",
      "content": "<p><a href=\"https://www.kaggle.com/forcewithme\" target=\"_blank\">@forcewithme</a> did you select your thresholds based on an OOF?</p>",
      "rawMarkdown": "@forcewithme did you select your thresholds based on an OOF?",
      "replies": [
        {
          "id": 2162876,
          "postDate": "2023-02-28T13:10:22.887Z",
          "content": "<p>Yes, it’s the best way to suffer the shake in this dangerous competition.</p>",
          "rawMarkdown": "Yes, it’s the best way to suffer the shake in this dangerous competition.",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2162045,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2023-02-28T00:28:25.930000",
      "content": "<p>Private LB still has a large random range. While the scores are now much closer to actual CV numbers, we would still have subs with 0.53 private LB that have much lower CV score than our selection. And I am sure many teams have unselected high scoring private subs, that did not make too much sense to select based on CV.</p>\n<p>You need to have a strong model/ensemble, and then also be lucky with the threshold to score high here.</p>",
      "votes": 7,
      "replies": [
        {
          "id": 2162051,
          "author_name": "ForcewithMe",
          "author_url": "",
          "post_date": "2023-02-28T00:37:22.960000",
          "content": "<p>Yeah, one lucky thing for us is that the threshold is stable </p>",
          "votes": 0,
          "replies": [
            {
              "id": 2162220,
              "author_name": "Nischay Dhankhar",
              "author_url": "",
              "post_date": "2023-02-28T04:13:36.303000",
              "content": "<p>Congrats <a href=\"https://www.kaggle.com/forcewithme\" target=\"_blank\">@forcewithme</a> Many unselected good scoring submissions for me, I took the bet on cv scores. Are thresholds stable for you or in general? I can still see +- 0.05 private lb over slight changes in thresholds. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2162228,
              "author_name": "ForcewithMe",
              "author_url": "",
              "post_date": "2023-02-28T04:18:10.080000",
              "content": "<p>Thank you. We always select the best threshold on CV. I haven't tested how it performed on private LB. If someone probs the public LB to get the threshold in this competition, I think he will suffer a huge shake.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2162055,
          "author_name": "NguyenThanhNhan",
          "author_url": "",
          "post_date": "2023-02-28T00:42:11.290000",
          "content": "<p>Private and public lb scores are almost uncorrelated, but our best public sub. is also best private (0.69 public, 0.54 private). However, I didn't dare to select it since it only contained models' weights + optimized threshold from our first 2 folds.<br>\nP/s: I completely disliked the pf1 metric 😓</p>",
          "votes": 6,
          "replies": [
            {
              "id": 2162058,
              "author_name": "ForcewithMe",
              "author_url": "",
              "post_date": "2023-02-28T00:47:52.997000",
              "content": "<p>Yes, I don’t think it’s a good metric.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2162050,
      "author_name": "Harshit Sheoran",
      "author_url": "",
      "post_date": "2023-02-28T00:35:35.823000",
      "content": "<p>We got more than 10 subs ranging from .46 cv to cv worse than that scoring .5 or above (gold medal range), we focused on improving cv, improved oof to .523, but because of stabality chose our .508 as the final sub which public lb scores in range from .65-.64 mostly, private did not follow our cv at all</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 2162062,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2023-02-28T00:56:53.557000",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/forcewithme\" target=\"_blank\">@forcewithme</a> and team. I only had a couple of weekends to spend on this competition but jumped in anyways 13 days ago. I am happy with where my model ended up although I have one submission that would have put me a few points higher but still in the bronze range. I agree the metric is bonkers.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2162109,
          "author_name": "ForcewithMe",
          "author_url": "",
          "post_date": "2023-02-28T02:11:43.777000",
          "content": "<p>thank youuu</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2162865,
      "author_name": "moth",
      "author_url": "",
      "post_date": "2023-02-28T13:02:10.753000",
      "content": "<p><a href=\"https://www.kaggle.com/forcewithme\" target=\"_blank\">@forcewithme</a> did you select your thresholds based on an OOF?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2162876,
          "author_name": "ForcewithMe",
          "author_url": "",
          "post_date": "2023-02-28T13:10:22.887000",
          "content": "<p>Yes, it’s the best way to suffer the shake in this dangerous competition.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2162041": "Thanks to the organizers and congratulations to the winners!\n\nWe're not going to write the solution right now. Just share a funny story. Up until a week before the end of the competition, we were mystified by the high scores on public leaderboard. Because our single model only has a score of about 0.5 on CV. Until the other day, we ran simulations on the training set. This simulation is based on the fact that the score of site2 is much higher than that of site1. We guess that the high score of public leaderboard is related to the ratio of site1 and site2, as well as the proportion of positive.\n\n1. Divide public and private in the same proportion as kaggle.\n\n2. Only 40% of positive samples and 65% of negative site1 samples in public are retained.\n\n3. Use site2 in private to fill in the missing samples in public\n\n4. Add the samples removed from public to private\n\nThrough the above operation, our public score finally exceeds 0.6. After that, we decided to forget kaggle public leaderboard. Our final submission was based largely on the CV and private scores in our simulation. We expect our final one submission to be about 0.524 and ultimately it's 0.53 on kaggle's private leaderboard.\n\nBut!!!!!!!!!!! Our highest public submission, although CV was not that high and the simulated pvt was not that good, kaggle private leaderboard was actually higher than our carefully selected ones. I don't understand.. ",
    "2162045": "Private LB still has a large random range. While the scores are now much closer to actual CV numbers, we would still have subs with 0.53 private LB that have much lower CV score than our selection. And I am sure many teams have unselected high scoring private subs, that did not make too much sense to select based on CV.\n\nYou need to have a strong model/ensemble, and then also be lucky with the threshold to score high here.",
    "2162050": "We got more than 10 subs ranging from .46 cv to cv worse than that scoring .5 or above (gold medal range), we focused on improving cv, improved oof to .523, but because of stabality chose our .508 as the final sub which public lb scores in range from .65-.64 mostly, private did not follow our cv at all",
    "2162062": "Congrats @forcewithme and team. I only had a couple of weekends to spend on this competition but jumped in anyways 13 days ago. I am happy with where my model ended up although I have one submission that would have put me a few points higher but still in the bronze range. I agree the metric is bonkers.",
    "2162865": "@forcewithme did you select your thresholds based on an OOF?"
  }
}