{
  "id": 428853,
  "title": "Will there be a shake at the end?",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/428853",
  "author_name": "william.wu",
  "post_date": "2023-08-03T05:14:10.211000",
  "votes": 7,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I'm using 5 folds cross-validation, and evaluating the performance on oof. I came across several cases to make be believe there will be a shake on LB.</p>\n<ul>\n<li>Applying HFlip, almost no change on the CV but LB +0.01</li>\n<li>ResNet200e has a similar performance as EfficientNet-b7 on CV but worse than EfficientNet-b4 on LB( also &gt;0.01 gap).</li>\n</ul>",
  "messages": [
    {
      "id": 2371381,
      "postDate": "2023-08-03T05:14:10.210Z",
      "content": "<p>I'm using 5 folds cross-validation, and evaluating the performance on oof. I came across several cases to make be believe there will be a shake on LB.</p>\n<ul>\n<li>Applying HFlip, almost no change on the CV but LB +0.01</li>\n<li>ResNet200e has a similar performance as EfficientNet-b7 on CV but worse than EfficientNet-b4 on LB( also &gt;0.01 gap).</li>\n</ul>",
      "rawMarkdown": "I'm using 5 folds cross-validation, and evaluating the performance on oof. I came across several cases to make be believe there will be a shake on LB.\n- Applying HFlip, almost no change on the CV but LB +0.01\n- ResNet200e has a similar performance as EfficientNet-b7 on CV but worse than EfficientNet-b4 on LB( also >0.01 gap).",
      "votes": 6
    },
    {
      "id": 2371934,
      "postDate": "2023-08-03T11:06:38.527Z",
      "content": "<p>I guess the shake-up also depends on the threshold selected, I've seen Kagglers submissions with 0.01 thr, just because improves the LB score. I'm not an expert, but doesn't look good to me.</p>",
      "rawMarkdown": "I guess the shake-up also depends on the threshold selected, I've seen Kagglers submissions with 0.01 thr, just because improves the LB score. I'm not an expert, but doesn't look good to me.",
      "votes": 4,
      "replies": [
        {
          "id": 2372140,
          "postDate": "2023-08-03T14:00:08.223Z",
          "content": "<p>But you can see whats the optimal threshold on CV tho, it's the same type of 'postprocessing' as TTA or anything that increases CV</p>",
          "rawMarkdown": "But you can see whats the optimal threshold on CV tho, it's the same type of 'postprocessing' as TTA or anything that increases CV",
          "votes": 1,
          "replies": [
            {
              "id": 2372191,
              "postDate": "2023-08-03T14:38:40.093Z",
              "content": "<p>Hi J€ANMPIA, my point is the following, looping through a list of thresholds and selecting who's giving you the best validation dice or LB boost, it's risky at some point. TTA falls into another category in my opinion.</p>",
              "rawMarkdown": "Hi J€ANMPIA, my point is the following, looping through a list of thresholds and selecting who's giving you the best validation dice or LB boost, it's risky at some point. TTA falls into another category in my opinion."
            }
          ]
        }
      ]
    },
    {
      "id": 2374956,
      "postDate": "2023-08-05T09:40:41.210Z",
      "content": "<p>Hi! I have trained a model on 4 + 1 + 1 validation scheme (4 parts on train, 1 on val, 1 on test): </p>\n<ul>\n<li>train_dataset: 15509 samples, val_dataset: 3110 samples, test_dataset: 3766 samples</li>\n<li>optimal threshold by validation is 0.3, by train it is 0.25</li>\n<li>no model selection is performed by val, just EWMA of weights</li>\n<li>dice scores on val and test are 0.6804 and 0.6685 respectively</li>\n</ul>\n<p>I have predicted the test dataset and bootstrapped the F1 score with threshold 0.3 on sub-sample of 279 length with 1000 tries (basically, sample 279 records with replacement and calculate the score on it for 1000 times), the histogram of scores are on the figure below:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2F345e661cca104efd86898e6123296a14%2Fsimulate_public_lb_hist-preds-hearty-thunder-754-test-0.30-279-1000.png?generation=1691227988794649&amp;alt=media\" alt=\"\"></p>\n<p>According to this experiment, it seems that the shakeup could be quiet serious (at least for this model). If you have some proposal for better methodology or a question on this method please share it!</p>",
      "rawMarkdown": "Hi! I have trained a model on 4 + 1 + 1 validation scheme (4 parts on train, 1 on val, 1 on test): \n- train_dataset: 15509 samples, val_dataset: 3110 samples, test_dataset: 3766 samples\n- optimal threshold by validation is 0.3, by train it is 0.25\n- no model selection is performed by val, just EWMA of weights\n- dice scores on val and test are 0.6804 and 0.6685 respectively\n\nI have predicted the test dataset and bootstrapped the F1 score with threshold 0.3 on sub-sample of 279 length with 1000 tries (basically, sample 279 records with replacement and calculate the score on it for 1000 times), the histogram of scores are on the figure below:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2F345e661cca104efd86898e6123296a14%2Fsimulate_public_lb_hist-preds-hearty-thunder-754-test-0.30-279-1000.png?generation=1691227988794649&alt=media)\n\nAccording to this experiment, it seems that the shakeup could be quiet serious (at least for this model). If you have some proposal for better methodology or a question on this method please share it!",
      "votes": 4,
      "replies": [
        {
          "id": 2375160,
          "postDate": "2023-08-05T13:06:54.847Z",
          "content": "<p>I believe in my 5 folds CV and use the public LB as another validation set</p>",
          "rawMarkdown": "I believe in my 5 folds CV and use the public LB as another validation set",
          "votes": 1
        },
        {
          "id": 2375388,
          "postDate": "2023-08-05T15:41:13.440Z",
          "content": "<p><a href=\"https://www.kaggle.com/mkotyushev\" target=\"_blank\">@mkotyushev</a> I believe if you want to predict shakeup, you should add a second model here (or more). If you compare two similar models (~same CV but somehow different), then you can create a scatter plot of model_1_score vs model_2_score on the 1000 subsamples.<br>\nIf you see a straight line it means that even though some subsamples are easier, they are easier for all models and 279 examples are enough to compare models: so there will probably be a minimal shakeup.<br>\nIf you see a poor linear correlation: then comparing two similar models on 279 examples is not enough to judge the quality of a model -&gt; there will probably be a large shakeup.</p>\n<p>This won't take into account LB overfiting, people tweaking their models to optimize LB without looking at their CV much.</p>",
          "rawMarkdown": "@mkotyushev I believe if you want to predict shakeup, you should add a second model here (or more). If you compare two similar models (~same CV but somehow different), then you can create a scatter plot of model_1_score vs model_2_score on the 1000 subsamples.\nIf you see a straight line it means that even though some subsamples are easier, they are easier for all models and 279 examples are enough to compare models: so there will probably be a minimal shakeup.\nIf you see a poor linear correlation: then comparing two similar models on 279 examples is not enough to judge the quality of a model -> there will probably be a large shakeup.\n\nThis won't take into account LB overfiting, people tweaking their models to optimize LB without looking at their CV much.",
          "votes": 8,
          "replies": [
            {
              "id": 2375447,
              "postDate": "2023-08-05T16:25:07.493Z",
              "content": "<p>Hi, thanks for reply</p>\n<p>Agreed, I came to similar thought. I compared the same model as above with weaker (~0.65-0.66 on val / test sets) model of different architecture and in addition to the individual histograms (first is stronger model, second is weaker) added a shared scatter-plot with stronger model on x- and weaker on y-axis</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2F3962cd1199ed7f4fc41fbe36873e241b%2Fsimulate_public_lb_hist-preds-hearty-thunder-754-test-preds-hearty-thunder-754-val-preds-crimson-sweep-1-test-preds-crimson-sweep-1-val-0.30-0.85-279-1000-False.png?generation=1691252190694053&amp;alt=media\" alt=\"\"></p>\n<p>Here we see some variation (e. g. on single y-axis \"slice\") but less dramatic that one on a individual histogram. Unfortunately I do not have full CV results on same folds for many models, so I do not think this graph should be considered seriously, but it is definitely a way to estimate possible shakeup if you have more data</p>",
              "rawMarkdown": "Hi, thanks for reply\n\nAgreed, I came to similar thought. I compared the same model as above with weaker (~0.65-0.66 on val / test sets) model of different architecture and in addition to the individual histograms (first is stronger model, second is weaker) added a shared scatter-plot with stronger model on x- and weaker on y-axis\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2F3962cd1199ed7f4fc41fbe36873e241b%2Fsimulate_public_lb_hist-preds-hearty-thunder-754-test-preds-hearty-thunder-754-val-preds-crimson-sweep-1-test-preds-crimson-sweep-1-val-0.30-0.85-279-1000-False.png?generation=1691252190694053&alt=media)\n\nHere we see some variation (e. g. on single y-axis \"slice\") but less dramatic that one on a individual histogram. Unfortunately I do not have full CV results on same folds for many models, so I do not think this graph should be considered seriously, but it is definitely a way to estimate possible shakeup if you have more data",
              "votes": 5
            }
          ]
        }
      ]
    },
    {
      "id": 2372119,
      "postDate": "2023-08-03T13:39:49.463Z",
      "content": "<p>taking into account that the current LB is based on ~290 images aprox. and the validation set contains 1856… it's pretty sure that yes, prepare for the shake-up 🤠</p>",
      "rawMarkdown": "taking into account that the current LB is based on ~290 images aprox. and the validation set contains 1856... it's pretty sure that yes, prepare for the shake-up 🤠",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 2371934,
      "author_name": "Maximiliano Diaz Battan",
      "author_url": "",
      "post_date": "2023-08-03T11:06:38.527000",
      "content": "<p>I guess the shake-up also depends on the threshold selected, I've seen Kagglers submissions with 0.01 thr, just because improves the LB score. I'm not an expert, but doesn't look good to me.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2372140,
          "author_name": "JEANMPIA",
          "author_url": "",
          "post_date": "2023-08-03T14:00:08.223000",
          "content": "<p>But you can see whats the optimal threshold on CV tho, it's the same type of 'postprocessing' as TTA or anything that increases CV</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2372191,
              "author_name": "Maximiliano Diaz Battan",
              "author_url": "",
              "post_date": "2023-08-03T14:38:40.093000",
              "content": "<p>Hi J€ANMPIA, my point is the following, looping through a list of thresholds and selecting who's giving you the best validation dice or LB boost, it's risky at some point. TTA falls into another category in my opinion.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2374956,
      "author_name": "Mikhail Kotyushev",
      "author_url": "",
      "post_date": "2023-08-05T09:40:41.210000",
      "content": "<p>Hi! I have trained a model on 4 + 1 + 1 validation scheme (4 parts on train, 1 on val, 1 on test): </p>\n<ul>\n<li>train_dataset: 15509 samples, val_dataset: 3110 samples, test_dataset: 3766 samples</li>\n<li>optimal threshold by validation is 0.3, by train it is 0.25</li>\n<li>no model selection is performed by val, just EWMA of weights</li>\n<li>dice scores on val and test are 0.6804 and 0.6685 respectively</li>\n</ul>\n<p>I have predicted the test dataset and bootstrapped the F1 score with threshold 0.3 on sub-sample of 279 length with 1000 tries (basically, sample 279 records with replacement and calculate the score on it for 1000 times), the histogram of scores are on the figure below:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2F345e661cca104efd86898e6123296a14%2Fsimulate_public_lb_hist-preds-hearty-thunder-754-test-0.30-279-1000.png?generation=1691227988794649&amp;alt=media\" alt=\"\"></p>\n<p>According to this experiment, it seems that the shakeup could be quiet serious (at least for this model). If you have some proposal for better methodology or a question on this method please share it!</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2375160,
          "author_name": "william.wu",
          "author_url": "",
          "post_date": "2023-08-05T13:06:54.847000",
          "content": "<p>I believe in my 5 folds CV and use the public LB as another validation set</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2375388,
          "author_name": "Optimo",
          "author_url": "",
          "post_date": "2023-08-05T15:41:13.440000",
          "content": "<p><a href=\"https://www.kaggle.com/mkotyushev\" target=\"_blank\">@mkotyushev</a> I believe if you want to predict shakeup, you should add a second model here (or more). If you compare two similar models (~same CV but somehow different), then you can create a scatter plot of model_1_score vs model_2_score on the 1000 subsamples.<br>\nIf you see a straight line it means that even though some subsamples are easier, they are easier for all models and 279 examples are enough to compare models: so there will probably be a minimal shakeup.<br>\nIf you see a poor linear correlation: then comparing two similar models on 279 examples is not enough to judge the quality of a model -&gt; there will probably be a large shakeup.</p>\n<p>This won't take into account LB overfiting, people tweaking their models to optimize LB without looking at their CV much.</p>",
          "votes": 8,
          "replies": [
            {
              "id": 2375447,
              "author_name": "Mikhail Kotyushev",
              "author_url": "",
              "post_date": "2023-08-05T16:25:07.493000",
              "content": "<p>Hi, thanks for reply</p>\n<p>Agreed, I came to similar thought. I compared the same model as above with weaker (~0.65-0.66 on val / test sets) model of different architecture and in addition to the individual histograms (first is stronger model, second is weaker) added a shared scatter-plot with stronger model on x- and weaker on y-axis</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2F3962cd1199ed7f4fc41fbe36873e241b%2Fsimulate_public_lb_hist-preds-hearty-thunder-754-test-preds-hearty-thunder-754-val-preds-crimson-sweep-1-test-preds-crimson-sweep-1-val-0.30-0.85-279-1000-False.png?generation=1691252190694053&amp;alt=media\" alt=\"\"></p>\n<p>Here we see some variation (e. g. on single y-axis \"slice\") but less dramatic that one on a individual histogram. Unfortunately I do not have full CV results on same folds for many models, so I do not think this graph should be considered seriously, but it is definitely a way to estimate possible shakeup if you have more data</p>",
              "votes": 5,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2372119,
      "author_name": "Enric Domingo",
      "author_url": "",
      "post_date": "2023-08-03T13:39:49.463000",
      "content": "<p>taking into account that the current LB is based on ~290 images aprox. and the validation set contains 1856… it's pretty sure that yes, prepare for the shake-up 🤠</p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2371381": "I'm using 5 folds cross-validation, and evaluating the performance on oof. I came across several cases to make be believe there will be a shake on LB.\n- Applying HFlip, almost no change on the CV but LB +0.01\n- ResNet200e has a similar performance as EfficientNet-b7 on CV but worse than EfficientNet-b4 on LB( also >0.01 gap).",
    "2371934": "I guess the shake-up also depends on the threshold selected, I've seen Kagglers submissions with 0.01 thr, just because improves the LB score. I'm not an expert, but doesn't look good to me.",
    "2374956": "Hi! I have trained a model on 4 + 1 + 1 validation scheme (4 parts on train, 1 on val, 1 on test): \n- train_dataset: 15509 samples, val_dataset: 3110 samples, test_dataset: 3766 samples\n- optimal threshold by validation is 0.3, by train it is 0.25\n- no model selection is performed by val, just EWMA of weights\n- dice scores on val and test are 0.6804 and 0.6685 respectively\n\nI have predicted the test dataset and bootstrapped the F1 score with threshold 0.3 on sub-sample of 279 length with 1000 tries (basically, sample 279 records with replacement and calculate the score on it for 1000 times), the histogram of scores are on the figure below:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2F345e661cca104efd86898e6123296a14%2Fsimulate_public_lb_hist-preds-hearty-thunder-754-test-0.30-279-1000.png?generation=1691227988794649&alt=media)\n\nAccording to this experiment, it seems that the shakeup could be quiet serious (at least for this model). If you have some proposal for better methodology or a question on this method please share it!",
    "2372119": "taking into account that the current LB is based on ~290 images aprox. and the validation set contains 1856... it's pretty sure that yes, prepare for the shake-up 🤠"
  }
}