{
  "id": 168917,
  "title": "Estimate of Shakeup?",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/168917",
  "author_name": "Nicholas Lyu",
  "post_date": "2020-07-22T10:12:23.817000",
  "votes": 13,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Starting a customary near-end-of-the-competition thread: <strong>what is your estimate of the shakeup?</strong></p>\n<p>For me, one voice looked at the report and thinks that given public / private test set are similarly labeled via consensus, there should be small shakeup.<br>\nAnother voice looked at the disastrous (yes, disastrous) CV-LB correlation of my submissions and think that shakeup will be very high.</p>\n<p>The LB seems very very unstable. Several experiments I did:</p>\n<ol>\n<li>Dropping the worst single-fold of the 5-fold which constitutes my .911 submission: .911 &gt; .889</li>\n<li>Best single-fold (with minimal TTA and inference tricks): .901. With tricks added this fold could've accounted for my best score on its own</li>\n</ol>\n<p>The inconsistency of LB across folds and the noisy nature of data should be considered when estimating shakeup. What do you think?</p>\n<h1>Update</h1>\n<p>Made a quick local estimate of sampling noise of the LB (not even considering the label shift noise). Here is the sampling distribution of bootstrapped QWK of CV .91 model. CV .91 is calculated with ~2000 samples, boostrapping with 600 samples (~sample size of private leaderboard) and 2000 simulations.</p>\n<p><img src=\"https://www.googleapis.com/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1943421%2F8d50e923060c1782806644e268a80592%2Fsampling%20distribution.png\" alt=\"\"></p>\n<p>Really not much which can be done, it seems……Good luck everyone</p>",
  "messages": [
    {
      "id": 939587,
      "postDate": "2020-07-22T10:12:23.817Z",
      "content": "<p>Starting a customary near-end-of-the-competition thread: <strong>what is your estimate of the shakeup?</strong></p>\n<p>For me, one voice looked at the report and thinks that given public / private test set are similarly labeled via consensus, there should be small shakeup.<br>\nAnother voice looked at the disastrous (yes, disastrous) CV-LB correlation of my submissions and think that shakeup will be very high.</p>\n<p>The LB seems very very unstable. Several experiments I did:</p>\n<ol>\n<li>Dropping the worst single-fold of the 5-fold which constitutes my .911 submission: .911 &gt; .889</li>\n<li>Best single-fold (with minimal TTA and inference tricks): .901. With tricks added this fold could've accounted for my best score on its own</li>\n</ol>\n<p>The inconsistency of LB across folds and the noisy nature of data should be considered when estimating shakeup. What do you think?</p>\n<h1>Update</h1>\n<p>Made a quick local estimate of sampling noise of the LB (not even considering the label shift noise). Here is the sampling distribution of bootstrapped QWK of CV .91 model. CV .91 is calculated with ~2000 samples, boostrapping with 600 samples (~sample size of private leaderboard) and 2000 simulations.</p>\n<p><img src=\"https://www.googleapis.com/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1943421%2F8d50e923060c1782806644e268a80592%2Fsampling%20distribution.png\" alt=\"\"></p>\n<p>Really not much which can be done, it seems……Good luck everyone</p>",
      "rawMarkdown": "Starting a customary near-end-of-the-competition thread: **what is your estimate of the shakeup?**\n\nFor me, one voice looked at the report and thinks that given public / private test set are similarly labeled via consensus, there should be small shakeup.\nAnother voice looked at the disastrous (yes, disastrous) CV-LB correlation of my submissions and think that shakeup will be very high.\n\nThe LB seems very very unstable. Several experiments I did:\n1. Dropping the worst single-fold of the 5-fold which constitutes my .911 submission: .911 &gt; .889\n2. Best single-fold (with minimal TTA and inference tricks): .901. With tricks added this fold could've accounted for my best score on its own\n\nThe inconsistency of LB across folds and the noisy nature of data should be considered when estimating shakeup. What do you think?\n\n# Update\nMade a quick local estimate of sampling noise of the LB (not even considering the label shift noise). Here is the sampling distribution of bootstrapped QWK of CV .91 model. CV .91 is calculated with ~2000 samples, boostrapping with 600 samples (~sample size of private leaderboard) and 2000 simulations.\n\n![](https://www.googleapis.com/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1943421%2F8d50e923060c1782806644e268a80592%2Fsampling%20distribution.png)\n\nReally not much which can be done, it seems......Good luck everyone",
      "votes": 13
    },
    {
      "id": 940381,
      "postDate": "2020-07-22T22:25:34.477Z",
      "content": "<blockquote>\n  <p>Here is the sampling distribution of bootstrapped QWK of CV .91 model. CV .91 is calculated with ~2000 samples, boostrapping with 600 samples (~sample size of private leaderboard) and 2000 simulations.</p>\n</blockquote>\n\n<p>Test labels were not produced the same way as train labels.  Anything based on CV is questionable.</p>\n\n<p>Test data is small, hence qwk is very noisy.</p>\n\n<p>TL;DR  selecting sub is like throwing dice.</p>",
      "rawMarkdown": "&gt;  Here is the sampling distribution of bootstrapped QWK of CV .91 model. CV .91 is calculated with ~2000 samples, boostrapping with 600 samples (~sample size of private leaderboard) and 2000 simulations.\n\nTest labels were not produced the same way as train labels.  Anything based on CV is questionable.\n\nTest data is small, hence qwk is very noisy.\n\nTL;DR  selecting sub is like throwing dice.\n\n",
      "votes": 7
    },
    {
      "id": 940367,
      "postDate": "2020-07-22T21:49:16.617Z",
      "content": "<p>Happy shake-shake time 😲, big wave is coming. The LB is quite small, just ~500+500 samples, and uses QWK metric. My expectation that shake up may be ~0.005-0.01 based on my subs, and not wisely selected final submissions could be affected even more. Be careful and the best luck.</p>",
      "rawMarkdown": "Happy shake-shake time 😲, big wave is coming. The LB is quite small, just ~500+500 samples, and uses QWK metric. My expectation that shake up may be ~0.005-0.01 based on my subs, and not wisely selected final submissions could be affected even more. Be careful and the best luck.",
      "votes": 5
    },
    {
      "id": 939845,
      "postDate": "2020-07-22T14:14:22.787Z",
      "content": "<p>Yep local simulation shows that shake up can be really big… But we should not worry … There is famous saying: <strong>Panda works in mysterious ways.</strong> Book of Panda (23:7).</p>",
      "rawMarkdown": "Yep local simulation shows that shake up can be really big... But we should not worry ... There is famous saying: **Panda works in mysterious ways.** Book of Panda (23:7).\n",
      "votes": 3
    },
    {
      "id": 939738,
      "postDate": "2020-07-22T12:29:34.760Z",
      "content": "<p>Big shakeup, data is way too small for stable QWK.</p>",
      "rawMarkdown": "Big shakeup, data is way too small for stable QWK.",
      "votes": 3,
      "replies": [
        {
          "id": 939778,
          "postDate": "2020-07-22T13:03:58.827Z",
          "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> Agreed…bootstrapping CV with the sample size of the LB gives me a QWK distribution with SD of about 1…this indicates that quite scary shakeup might be on the way…</p>",
          "rawMarkdown": "@philippsinger Agreed...bootstrapping CV with the sample size of the LB gives me a QWK distribution with SD of about 1...this indicates that quite scary shakeup might be on the way...",
          "votes": 2
        }
      ]
    },
    {
      "id": 939716,
      "postDate": "2020-07-22T12:03:21.837Z",
      "content": "<p>In my experience, submitting the same inference kernel with a different random seed for TTA results in public LB differences of up to .006, which could indicate a big shakeup. Although public and private test should be similar, misclassifying just 1-2 additional observations out of 520 could cost places on the private LB.</p>",
      "rawMarkdown": "In my experience, submitting the same inference kernel with a different random seed for TTA results in public LB differences of up to .006, which could indicate a big shakeup. Although public and private test should be similar, misclassifying just 1-2 additional observations out of 520 could cost places on the private LB.",
      "votes": 3
    },
    {
      "id": 940080,
      "postDate": "2020-07-22T17:05:12.357Z",
      "content": "<p>The data size of private set is too small for qwk. 😿 </p>",
      "rawMarkdown": "The data size of private set is too small for qwk. 😿 ",
      "votes": 4
    },
    {
      "id": 940075,
      "postDate": "2020-07-22T17:01:31.757Z",
      "content": "<p>Big Shake up is comming...\nGood luck </p>",
      "rawMarkdown": "Big Shake up is comming...\nGood luck ",
      "votes": 1
    },
    {
      "id": 939754,
      "postDate": "2020-07-22T12:43:59.723Z",
      "content": "<p>I would rather rely on my CV which is based on about 2000 images (carefully balanced split) rather then Public LB which is based on about 420 images.</p>\n\n<p>In other words, I expect big shakeup )</p>",
      "rawMarkdown": "I would rather rely on my CV which is based on about 2000 images (carefully balanced split) rather then Public LB which is based on about 420 images.\n\nIn other words, I expect big shakeup )",
      "votes": 1,
      "replies": [
        {
          "id": 940170,
          "postDate": "2020-07-22T17:56:01.643Z",
          "content": "<p>I would generally agree with you, but in this case the train and test data have been labelled with two different processes, so saying which one is more reliable is really hard</p>",
          "rawMarkdown": "I would generally agree with you, but in this case the train and test data have been labelled with two different processes, so saying which one is more reliable is really hard",
          "votes": 1
        }
      ]
    },
    {
      "id": 939622,
      "postDate": "2020-07-22T10:40:44.380Z",
      "content": "<p>It has the potential to be significant. Private test has around 500 cases I think so QWK could be rather unstable.</p>",
      "rawMarkdown": "It has the potential to be significant. Private test has around 500 cases I think so QWK could be rather unstable.",
      "votes": 1
    },
    {
      "id": 941276,
      "postDate": "2020-07-23T06:43:34.863Z",
      "content": "<p>Since the competition is end. I want to ask you what's the inference tricks did you use if you don't mind disclosure? I remeber the day you get 0.889 and after sleeping, the next morning I find you jump to 0.911, it's very impressive. 👍 </p>",
      "rawMarkdown": "Since the competition is end. I want to ask you what's the inference tricks did you use if you don't mind disclosure? I remeber the day you get 0.889 and after sleeping, the next morning I find you jump to 0.911, it's very impressive. 👍 "
    }
  ],
  "comments": [
    {
      "id": 940381,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2020-07-22T22:25:34.477000",
      "content": "<blockquote>\n  <p>Here is the sampling distribution of bootstrapped QWK of CV .91 model. CV .91 is calculated with ~2000 samples, boostrapping with 600 samples (~sample size of private leaderboard) and 2000 simulations.</p>\n</blockquote>\n\n<p>Test labels were not produced the same way as train labels.  Anything based on CV is questionable.</p>\n\n<p>Test data is small, hence qwk is very noisy.</p>\n\n<p>TL;DR  selecting sub is like throwing dice.</p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 940367,
      "author_name": "Iafoss",
      "author_url": "",
      "post_date": "2020-07-22T21:49:16.617000",
      "content": "<p>Happy shake-shake time 😲, big wave is coming. The LB is quite small, just ~500+500 samples, and uses QWK metric. My expectation that shake up may be ~0.005-0.01 based on my subs, and not wisely selected final submissions could be affected even more. Be careful and the best luck.</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 939845,
      "author_name": "DrHB",
      "author_url": "",
      "post_date": "2020-07-22T14:14:22.787000",
      "content": "<p>Yep local simulation shows that shake up can be really big… But we should not worry … There is famous saying: <strong>Panda works in mysterious ways.</strong> Book of Panda (23:7).</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 939738,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2020-07-22T12:29:34.760000",
      "content": "<p>Big shakeup, data is way too small for stable QWK.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 939778,
          "author_name": "Nicholas Lyu",
          "author_url": "",
          "post_date": "2020-07-22T13:03:58.827000",
          "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> Agreed…bootstrapping CV with the sample size of the LB gives me a QWK distribution with SD of about 1…this indicates that quite scary shakeup might be on the way…</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 939716,
      "author_name": "Nikita Kozodoi",
      "author_url": "",
      "post_date": "2020-07-22T12:03:21.837000",
      "content": "<p>In my experience, submitting the same inference kernel with a different random seed for TTA results in public LB differences of up to .006, which could indicate a big shakeup. Although public and private test should be similar, misclassifying just 1-2 additional observations out of 520 could cost places on the private LB.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 940080,
      "author_name": "Sugawarya",
      "author_url": "",
      "post_date": "2020-07-22T17:05:12.357000",
      "content": "<p>The data size of private set is too small for qwk. 😿 </p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 940075,
      "author_name": "seefun",
      "author_url": "",
      "post_date": "2020-07-22T17:01:31.757000",
      "content": "<p>Big Shake up is comming...\nGood luck </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 939754,
      "author_name": "Dmitry A. Grechka",
      "author_url": "",
      "post_date": "2020-07-22T12:43:59.723000",
      "content": "<p>I would rather rely on my CV which is based on about 2000 images (carefully balanced split) rather then Public LB which is based on about 420 images.</p>\n\n<p>In other words, I expect big shakeup )</p>",
      "votes": 1,
      "replies": [
        {
          "id": 940170,
          "author_name": "Pasquale",
          "author_url": "",
          "post_date": "2020-07-22T17:56:01.643000",
          "content": "<p>I would generally agree with you, but in this case the train and test data have been labelled with two different processes, so saying which one is more reliable is really hard</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 939622,
      "author_name": "Ian Pan",
      "author_url": "",
      "post_date": "2020-07-22T10:40:44.380000",
      "content": "<p>It has the potential to be significant. Private test has around 500 cases I think so QWK could be rather unstable.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 941276,
      "author_name": "Shiyuan Zeng",
      "author_url": "",
      "post_date": "2020-07-23T06:43:34.863000",
      "content": "<p>Since the competition is end. I want to ask you what's the inference tricks did you use if you don't mind disclosure? I remeber the day you get 0.889 and after sleeping, the next morning I find you jump to 0.911, it's very impressive. 👍 </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "939587": "Starting a customary near-end-of-the-competition thread: **what is your estimate of the shakeup?**\n\nFor me, one voice looked at the report and thinks that given public / private test set are similarly labeled via consensus, there should be small shakeup.\nAnother voice looked at the disastrous (yes, disastrous) CV-LB correlation of my submissions and think that shakeup will be very high.\n\nThe LB seems very very unstable. Several experiments I did:\n1. Dropping the worst single-fold of the 5-fold which constitutes my .911 submission: .911 &gt; .889\n2. Best single-fold (with minimal TTA and inference tricks): .901. With tricks added this fold could've accounted for my best score on its own\n\nThe inconsistency of LB across folds and the noisy nature of data should be considered when estimating shakeup. What do you think?\n\n# Update\nMade a quick local estimate of sampling noise of the LB (not even considering the label shift noise). Here is the sampling distribution of bootstrapped QWK of CV .91 model. CV .91 is calculated with ~2000 samples, boostrapping with 600 samples (~sample size of private leaderboard) and 2000 simulations.\n\n![](https://www.googleapis.com/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1943421%2F8d50e923060c1782806644e268a80592%2Fsampling%20distribution.png)\n\nReally not much which can be done, it seems......Good luck everyone",
    "940381": "&gt;  Here is the sampling distribution of bootstrapped QWK of CV .91 model. CV .91 is calculated with ~2000 samples, boostrapping with 600 samples (~sample size of private leaderboard) and 2000 simulations.\n\nTest labels were not produced the same way as train labels.  Anything based on CV is questionable.\n\nTest data is small, hence qwk is very noisy.\n\nTL;DR  selecting sub is like throwing dice.\n\n",
    "940367": "Happy shake-shake time 😲, big wave is coming. The LB is quite small, just ~500+500 samples, and uses QWK metric. My expectation that shake up may be ~0.005-0.01 based on my subs, and not wisely selected final submissions could be affected even more. Be careful and the best luck.",
    "939845": "Yep local simulation shows that shake up can be really big... But we should not worry ... There is famous saying: **Panda works in mysterious ways.** Book of Panda (23:7).\n",
    "939738": "Big shakeup, data is way too small for stable QWK.",
    "939716": "In my experience, submitting the same inference kernel with a different random seed for TTA results in public LB differences of up to .006, which could indicate a big shakeup. Although public and private test should be similar, misclassifying just 1-2 additional observations out of 520 could cost places on the private LB.",
    "940080": "The data size of private set is too small for qwk. 😿 ",
    "940075": "Big Shake up is comming...\nGood luck ",
    "939754": "I would rather rely on my CV which is based on about 2000 images (carefully balanced split) rather then Public LB which is based on about 420 images.\n\nIn other words, I expect big shakeup )",
    "939622": "It has the potential to be significant. Private test has around 500 cases I think so QWK could be rather unstable.",
    "941276": "Since the competition is end. I want to ask you what's the inference tricks did you use if you don't mind disclosure? I remeber the day you get 0.889 and after sleeping, the next morning I find you jump to 0.911, it's very impressive. 👍 "
  }
}