{
  "id": 420135,
  "title": "Trust ur CV even though LB drops",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/420135",
  "author_name": "LUPIN11",
  "post_date": "2023-06-29T11:18:15.512000",
  "votes": 15,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Owing to the limited sample amount of public LB, I would like to come to this conclusion: <strong>Trust ur CV even through LB drops</strong>. </p>\n<p>What I want to express here is that if your idea is explanatory and shows an apparent improvement in CV, then there is no need to feel discouraged even if LB drops.</p>\n<p>Here is the basis on which I reached this conclusion:</p>\n<p>Recently, I managed to improve the global dice of my baseline model on my validation set from 0.6124 to 0.6204. But the better model got a lower LB(0.547) than previous LB(0.550). BTW, though my baseline model has a low LB,  it's easy for me to modify implementation details in order to test my ideas.<br>\nThen I tried to fix this weird drop of LB and explain it. By slightly adjusting the threshold, the model of CV 0.6204 got LB 0.551. But this LB improvement is trivial in contrast with the improvement on CV. I attribute this to the limited amount of samples of public LB (only about 300 samples).  With some experiments, I'd like to conclude that:</p>\n<ul>\n<li><p><strong>an apparent improvement on a dataset(e.g. 2000 samples) maybe a trivial improvement on its small subset (e.g. 300 samples)</strong></p></li>\n<li><p><strong>the optimal threshold may fluctuate on different datasets or different subsets and thereby the trivial improvement may becomes a drop on this subset.</strong></p></li>\n</ul>\n<p>Details:</p>\n<blockquote>\n  <p>The hidden test set is approximately the same size (± 5%) as the validation set<br>\n  This leaderboard is calculated with approximately 15% of the test data.</p>\n</blockquote>\n<p>So the public LB is calculated on just about 300 samples. My validation set has 2000 samples sampled from the train folder.</p>\n<p>Experiment A: <br>\nCalculate the global dice on subsets of my validation set. Each subset has only 300 samples and there is no overlap. These lables (0.6124 and so on) means the corresponding global dice on my validation set. As you can see, the improvement on the 4th subset is very small. For me, this explain my trivial LB improvement (0.550-&gt;0.551) compared with the CV improvement (0.6124-&gt;0.6204) <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9893459%2Fd13d006167e732cdf1775e96615b4de3%2Fk1.png?generation=1688035106333192&amp;alt=media\" alt=\"\"></p>\n<p>Experiments B:<br>\nIn my submission of the improved model (CV 0.6204), I use the optimal threshold (0.26) on my validation set. This results in a drop from 0.550 to 0.547 on LB. But when changing the threshold to 0.24, the LB is 0.551, which seems like a more reasonable result. Then I tried to search the optimal threshold on a different and small subset of my validation set. I found that the optimal threshold may fluctuate on a small subset. In that case, the comparison of models on this small subset becomes unfair. Since when considering the entire dataset, this comparison would be biased, and when we only consider this subset, the model may fall behind due to the fluctuation of the optimal threshold, even though it has great potential to surpass another model by just adjusting the threshold. Figures below show the optimal threshold of 2 models on a dataset of 2000 samples and its small subset of 300 samples. As you can see, the optimal threshold of the 2nd model changes from 0.26 to 0.24.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9893459%2F905f3901b3ae27db473d2ef8a2f2496d%2Fk2.png?generation=1688036297046790&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9893459%2F5aeee50bb9af65a1621f2d73983cf84f%2Fk3.png?generation=1688036309876562&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 2322599,
      "postDate": "2023-06-29T11:18:15.513Z",
      "content": "<p>Owing to the limited sample amount of public LB, I would like to come to this conclusion: <strong>Trust ur CV even through LB drops</strong>. </p>\n<p>What I want to express here is that if your idea is explanatory and shows an apparent improvement in CV, then there is no need to feel discouraged even if LB drops.</p>\n<p>Here is the basis on which I reached this conclusion:</p>\n<p>Recently, I managed to improve the global dice of my baseline model on my validation set from 0.6124 to 0.6204. But the better model got a lower LB(0.547) than previous LB(0.550). BTW, though my baseline model has a low LB,  it's easy for me to modify implementation details in order to test my ideas.<br>\nThen I tried to fix this weird drop of LB and explain it. By slightly adjusting the threshold, the model of CV 0.6204 got LB 0.551. But this LB improvement is trivial in contrast with the improvement on CV. I attribute this to the limited amount of samples of public LB (only about 300 samples).  With some experiments, I'd like to conclude that:</p>\n<ul>\n<li><p><strong>an apparent improvement on a dataset(e.g. 2000 samples) maybe a trivial improvement on its small subset (e.g. 300 samples)</strong></p></li>\n<li><p><strong>the optimal threshold may fluctuate on different datasets or different subsets and thereby the trivial improvement may becomes a drop on this subset.</strong></p></li>\n</ul>\n<p>Details:</p>\n<blockquote>\n  <p>The hidden test set is approximately the same size (± 5%) as the validation set<br>\n  This leaderboard is calculated with approximately 15% of the test data.</p>\n</blockquote>\n<p>So the public LB is calculated on just about 300 samples. My validation set has 2000 samples sampled from the train folder.</p>\n<p>Experiment A: <br>\nCalculate the global dice on subsets of my validation set. Each subset has only 300 samples and there is no overlap. These lables (0.6124 and so on) means the corresponding global dice on my validation set. As you can see, the improvement on the 4th subset is very small. For me, this explain my trivial LB improvement (0.550-&gt;0.551) compared with the CV improvement (0.6124-&gt;0.6204) <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9893459%2Fd13d006167e732cdf1775e96615b4de3%2Fk1.png?generation=1688035106333192&amp;alt=media\" alt=\"\"></p>\n<p>Experiments B:<br>\nIn my submission of the improved model (CV 0.6204), I use the optimal threshold (0.26) on my validation set. This results in a drop from 0.550 to 0.547 on LB. But when changing the threshold to 0.24, the LB is 0.551, which seems like a more reasonable result. Then I tried to search the optimal threshold on a different and small subset of my validation set. I found that the optimal threshold may fluctuate on a small subset. In that case, the comparison of models on this small subset becomes unfair. Since when considering the entire dataset, this comparison would be biased, and when we only consider this subset, the model may fall behind due to the fluctuation of the optimal threshold, even though it has great potential to surpass another model by just adjusting the threshold. Figures below show the optimal threshold of 2 models on a dataset of 2000 samples and its small subset of 300 samples. As you can see, the optimal threshold of the 2nd model changes from 0.26 to 0.24.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9893459%2F905f3901b3ae27db473d2ef8a2f2496d%2Fk2.png?generation=1688036297046790&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9893459%2F5aeee50bb9af65a1621f2d73983cf84f%2Fk3.png?generation=1688036309876562&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Owing to the limited sample amount of public LB, I would like to come to this conclusion: **Trust ur CV even through LB drops**. \n\nWhat I want to express here is that if your idea is explanatory and shows an apparent improvement in CV, then there is no need to feel discouraged even if LB drops.\n\nHere is the basis on which I reached this conclusion:\n\nRecently, I managed to improve the global dice of my baseline model on my validation set from 0.6124 to 0.6204. But the better model got a lower LB(0.547) than previous LB(0.550). BTW, though my baseline model has a low LB,  it's easy for me to modify implementation details in order to test my ideas.\nThen I tried to fix this weird drop of LB and explain it. By slightly adjusting the threshold, the model of CV 0.6204 got LB 0.551. But this LB improvement is trivial in contrast with the improvement on CV. I attribute this to the limited amount of samples of public LB (only about 300 samples).  With some experiments, I'd like to conclude that:\n+ **an apparent improvement on a dataset(e.g. 2000 samples) maybe a trivial improvement on its small subset (e.g. 300 samples)**\n\n+ **the optimal threshold may fluctuate on different datasets or different subsets and thereby the trivial improvement may becomes a drop on this subset.**\n\nDetails:\n> The hidden test set is approximately the same size (± 5%) as the validation set\n> This leaderboard is calculated with approximately 15% of the test data.\n\nSo the public LB is calculated on just about 300 samples. My validation set has 2000 samples sampled from the train folder.\n\nExperiment A: \nCalculate the global dice on subsets of my validation set. Each subset has only 300 samples and there is no overlap. These lables (0.6124 and so on) means the corresponding global dice on my validation set. As you can see, the improvement on the 4th subset is very small. For me, this explain my trivial LB improvement (0.550->0.551) compared with the CV improvement (0.6124->0.6204) \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9893459%2Fd13d006167e732cdf1775e96615b4de3%2Fk1.png?generation=1688035106333192&alt=media)\n\nExperiments B:\nIn my submission of the improved model (CV 0.6204), I use the optimal threshold (0.26) on my validation set. This results in a drop from 0.550 to 0.547 on LB. But when changing the threshold to 0.24, the LB is 0.551, which seems like a more reasonable result. Then I tried to search the optimal threshold on a different and small subset of my validation set. I found that the optimal threshold may fluctuate on a small subset. In that case, the comparison of models on this small subset becomes unfair. Since when considering the entire dataset, this comparison would be biased, and when we only consider this subset, the model may fall behind due to the fluctuation of the optimal threshold, even though it has great potential to surpass another model by just adjusting the threshold. Figures below show the optimal threshold of 2 models on a dataset of 2000 samples and its small subset of 300 samples. As you can see, the optimal threshold of the 2nd model changes from 0.26 to 0.24.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9893459%2F905f3901b3ae27db473d2ef8a2f2496d%2Fk2.png?generation=1688036297046790&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9893459%2F5aeee50bb9af65a1621f2d73983cf84f%2Fk3.png?generation=1688036309876562&alt=media)",
      "votes": 14
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2322599": "Owing to the limited sample amount of public LB, I would like to come to this conclusion: **Trust ur CV even through LB drops**. \n\nWhat I want to express here is that if your idea is explanatory and shows an apparent improvement in CV, then there is no need to feel discouraged even if LB drops.\n\nHere is the basis on which I reached this conclusion:\n\nRecently, I managed to improve the global dice of my baseline model on my validation set from 0.6124 to 0.6204. But the better model got a lower LB(0.547) than previous LB(0.550). BTW, though my baseline model has a low LB,  it's easy for me to modify implementation details in order to test my ideas.\nThen I tried to fix this weird drop of LB and explain it. By slightly adjusting the threshold, the model of CV 0.6204 got LB 0.551. But this LB improvement is trivial in contrast with the improvement on CV. I attribute this to the limited amount of samples of public LB (only about 300 samples).  With some experiments, I'd like to conclude that:\n+ **an apparent improvement on a dataset(e.g. 2000 samples) maybe a trivial improvement on its small subset (e.g. 300 samples)**\n\n+ **the optimal threshold may fluctuate on different datasets or different subsets and thereby the trivial improvement may becomes a drop on this subset.**\n\nDetails:\n> The hidden test set is approximately the same size (± 5%) as the validation set\n> This leaderboard is calculated with approximately 15% of the test data.\n\nSo the public LB is calculated on just about 300 samples. My validation set has 2000 samples sampled from the train folder.\n\nExperiment A: \nCalculate the global dice on subsets of my validation set. Each subset has only 300 samples and there is no overlap. These lables (0.6124 and so on) means the corresponding global dice on my validation set. As you can see, the improvement on the 4th subset is very small. For me, this explain my trivial LB improvement (0.550->0.551) compared with the CV improvement (0.6124->0.6204) \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9893459%2Fd13d006167e732cdf1775e96615b4de3%2Fk1.png?generation=1688035106333192&alt=media)\n\nExperiments B:\nIn my submission of the improved model (CV 0.6204), I use the optimal threshold (0.26) on my validation set. This results in a drop from 0.550 to 0.547 on LB. But when changing the threshold to 0.24, the LB is 0.551, which seems like a more reasonable result. Then I tried to search the optimal threshold on a different and small subset of my validation set. I found that the optimal threshold may fluctuate on a small subset. In that case, the comparison of models on this small subset becomes unfair. Since when considering the entire dataset, this comparison would be biased, and when we only consider this subset, the model may fall behind due to the fluctuation of the optimal threshold, even though it has great potential to surpass another model by just adjusting the threshold. Figures below show the optimal threshold of 2 models on a dataset of 2000 samples and its small subset of 300 samples. As you can see, the optimal threshold of the 2nd model changes from 0.26 to 0.24.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9893459%2F905f3901b3ae27db473d2ef8a2f2496d%2Fk2.png?generation=1688036297046790&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9893459%2F5aeee50bb9af65a1621f2d73983cf84f%2Fk3.png?generation=1688036309876562&alt=media)"
  }
}