{
  "id": 219090,
  "title": "possibility of public/private shake at this competition",
  "url": "/competitions/vinbigdata-chest-xray-abnormalities-detection/discussion/219090",
  "author_name": "Ryunosuke Ishizaki",
  "post_date": "2021-02-13T08:41:40.196000",
  "votes": 16,
  "comment_count": 18,
  "views": 0,
  "content": "<p>Hi there,<br>\nnow we know 10% of total 3000 test data are public, rest 90% are private.<br>\nwe can estimate host will prepare relatively more abnormal images for test, but if test distribution of each abnormalities are same as train/public test/private test, N of each abnormalities in test data<br>\ncan be estimated as<br>\npublic : [61,  3,  9, 46,  7,  7, 12, 26, 16, 22, 20, 39,  1, 32],<br>\nprivate : [552,  33,  81, 413,  63,  69, 110, 237, 148, 204, 185, 356,  17, 291]<br>\nthis may cause huge shakeup for last leaderboard. how do you think?</p>",
  "messages": [
    {
      "id": 1198730,
      "postDate": "2021-02-13T08:41:40.197Z",
      "content": "<p>Hi there,<br>\nnow we know 10% of total 3000 test data are public, rest 90% are private.<br>\nwe can estimate host will prepare relatively more abnormal images for test, but if test distribution of each abnormalities are same as train/public test/private test, N of each abnormalities in test data<br>\ncan be estimated as<br>\npublic : [61,  3,  9, 46,  7,  7, 12, 26, 16, 22, 20, 39,  1, 32],<br>\nprivate : [552,  33,  81, 413,  63,  69, 110, 237, 148, 204, 185, 356,  17, 291]<br>\nthis may cause huge shakeup for last leaderboard. how do you think?</p>",
      "rawMarkdown": "Hi there,\nnow we know 10% of total 3000 test data are public, rest 90% are private.\nwe can estimate host will prepare relatively more abnormal images for test, but if test distribution of each abnormalities are same as train/public test/private test, N of each abnormalities in test data\ncan be estimated as\npublic : [61,  3,  9, 46,  7,  7, 12, 26, 16, 22, 20, 39,  1, 32],\nprivate : [552,  33,  81, 413,  63,  69, 110, 237, 148, 204, 185, 356,  17, 291]\nthis may cause huge shakeup for last leaderboard. how do you think?",
      "votes": 16
    },
    {
      "id": 1199493,
      "postDate": "2021-02-13T20:54:24.503Z",
      "content": "<p>I think you're right. Even if you re-run your model for several times, the LB score will be quite different with little init randomization. And considering the annotations from train/test set are also different, I'm still having a hard time trying to find a reliable metric to select model.</p>",
      "rawMarkdown": "I think you're right. Even if you re-run your model for several times, the LB score will be quite different with little init randomization. And considering the annotations from train/test set are also different, I'm still having a hard time trying to find a reliable metric to select model.",
      "votes": 4,
      "replies": [
        {
          "id": 1210227,
          "postDate": "2021-02-19T09:24:36.240Z",
          "content": "<p><a href=\"https://www.kaggle.com/issacc\" target=\"_blank\">@issacc</a> I got lb range 12.0-18.0 map when re-running same model…  I am frustrated, what am i doing wrong lol?</p>",
          "rawMarkdown": "@issacc I got lb range 12.0-18.0 map when re-running same model...  I am frustrated, what am i doing wrong lol?",
          "votes": 2
        },
        {
          "id": 1211677,
          "postDate": "2021-02-20T12:44:00.320Z",
          "content": "<p>Thanks for sharing ideas, we wanna get higher public/private lb, but when we seek upper public lb score, we have more risk for shake at private lb for espessially in this competition. dilemma🤔</p>",
          "rawMarkdown": "Thanks for sharing ideas, we wanna get higher public/private lb, but when we seek upper public lb score, we have more risk for shake at private lb for espessially in this competition. dilemma🤔",
          "votes": 1
        }
      ]
    },
    {
      "id": 1215290,
      "postDate": "2021-02-23T14:04:01.527Z",
      "content": "<p>Agree!!<br>\nwhich makes is more important to trust your CV and models and not get swayed by the Public LB</p>",
      "rawMarkdown": "Agree!!\nwhich makes is more important to trust your CV and models and not get swayed by the Public LB",
      "votes": 1
    },
    {
      "id": 1212210,
      "postDate": "2021-02-21T02:27:22.610Z",
      "content": "<p>Hello , is this game only need to submit a table?</p>",
      "rawMarkdown": "Hello , is this game only need to submit a table?",
      "votes": 1
    },
    {
      "id": 1207011,
      "postDate": "2021-02-17T17:18:17.643Z",
      "content": "<p>I was just thinking: Looking at your frequency estimates, model selection based on public LB results restricted to the more frequent classes may be an option? These might be a bit more reliable than the very rare classes? On the other hand, different models may behaviour differently on rare vs. frequent classes.. But this I guess can only be assessed looking at the respective model and setup individualistically..</p>",
      "rawMarkdown": "I was just thinking: Looking at your frequency estimates, model selection based on public LB results restricted to the more frequent classes may be an option? These might be a bit more reliable than the very rare classes? On the other hand, different models may behaviour differently on rare vs. frequent classes.. But this I guess can only be assessed looking at the respective model and setup individualistically..",
      "votes": 1,
      "replies": [
        {
          "id": 1211668,
          "postDate": "2021-02-20T12:29:10.643Z",
          "content": "<p>I guess we can roughly refar to public LB results, for ex, if each models are single, public LB 0.25 vs 0.1, will be the same results as Private LB, but the more we start to add new process/system to test data,<br>\n(for ex, re-thinking bbox(nms, nmw, wbf, etc…)), you have more possibility of overfitting to Public LB.<br>\n(for my experience, Hold-out cv 0.25 vs Public LB 0.16 for simple model outputs,  hold-out cv 0.35 vs Public LB 0.12 for additional process.)</p>",
          "rawMarkdown": "I guess we can roughly refar to public LB results, for ex, if each models are single, public LB 0.25 vs 0.1, will be the same results as Private LB, but the more we start to add new process/system to test data,\n(for ex, re-thinking bbox(nms, nmw, wbf, etc...)), you have more possibility of overfitting to Public LB.\n(for my experience, Hold-out cv 0.25 vs Public LB 0.16 for simple model outputs,  hold-out cv 0.35 vs Public LB 0.12 for additional process.)",
          "votes": 2
        },
        {
          "id": 1211678,
          "postDate": "2021-02-20T12:44:52.837Z",
          "content": "<p>Yes, this aggrevates the issue.. In my experiment I included the 2-class predictions and NMS so that might be an issue.. on the other hand, I still think that a single model may accidentally fit very well let's say the few cases of Atelectasis in the public LB set while an other model might not. At the same time the latter model could show a better performance on the private LB set where more cases of Atelectasis are included.</p>",
          "rawMarkdown": "Yes, this aggrevates the issue.. In my experiment I included the 2-class predictions and NMS so that might be an issue.. on the other hand, I still think that a single model may accidentally fit very well let's say the few cases of Atelectasis in the public LB set while an other model might not. At the same time the latter model could show a better performance on the private LB set where more cases of Atelectasis are included.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1199078,
      "postDate": "2021-02-13T14:05:17.437Z",
      "content": "<p>All 3000 test images are in public. isn't it? </p>\n<p>am new to the competition so can you please tell me whether the leaderboard test data will be different from the available test data.</p>",
      "rawMarkdown": "All 3000 test images are in public. isn't it? \n\nam new to the competition so can you please tell me whether the leaderboard test data will be different from the available test data.",
      "votes": 1,
      "replies": [
        {
          "id": 1199169,
          "postDate": "2021-02-13T15:02:18.163Z",
          "content": "<p>They're all \"public\" as in you can view them, but only 300 are scored on the public LB. That's what he means.</p>",
          "rawMarkdown": "They're all \"public\" as in you can view them, but only 300 are scored on the public LB. That's what he means.",
          "votes": 3
        },
        {
          "id": 1200851,
          "postDate": "2021-02-15T03:47:38.020Z",
          "content": "<p>If I understand correctly, 3000 test images provided by the host include 300 for LB and 2700 remains for private LB?</p>",
          "rawMarkdown": "If I understand correctly, 3000 test images provided by the host include 300 for LB and 2700 remains for private LB?",
          "votes": 2
        },
        {
          "id": 1200853,
          "postDate": "2021-02-15T03:50:42.050Z",
          "content": "<p>Sure. right</p>",
          "rawMarkdown": "Sure. right",
          "votes": 2
        }
      ]
    },
    {
      "id": 1248038,
      "postDate": "2021-03-22T09:32:41.150Z",
      "content": "<p>I think there would be shake up.<br>\nAll previous X-ray competition had a great shake up. So I expect the same here.<br>\nBut ….<br>\nto trust on CV is quite hard in this competition. Since <br>\nCV doesn't correlate nor proportional to LB Score </p>",
      "rawMarkdown": "I think there would be shake up.\nAll previous X-ray competition had a great shake up. So I expect the same here.\nBut ....\nto trust on CV is quite hard in this competition. Since \nCV doesn't correlate nor proportional to LB Score ",
      "votes": 2,
      "replies": [
        {
          "id": 1248051,
          "postDate": "2021-03-22T09:46:08.700Z",
          "content": "<p>Yes, since my hold-out cv/lb score became mess, I stopped submitting,<br>\nthere are some factors of shake, public/private test label distribution difference, stage2 manual<br>\nlabeling from some radiologists, noisy training dataset of training set…etc<br>\nI guess last private lb top will be 0.2〜0.3, how do you think?</p>",
          "rawMarkdown": "Yes, since my hold-out cv/lb score became mess, I stopped submitting,\nthere are some factors of shake, public/private test label distribution difference, stage2 manual\nlabeling from some radiologists, noisy training dataset of training set...etc\nI guess last private lb top will be 0.2〜0.3, how do you think?"
        },
        {
          "id": 1248135,
          "postDate": "2021-03-22T11:29:17.403Z",
          "content": "<p>I think the same. but still i could imagine sometime our CV score might be the final private LB score</p>",
          "rawMarkdown": "I think the same. but still i could imagine sometime our CV score might be the final private LB score"
        },
        {
          "id": 1248243,
          "postDate": "2021-03-22T13:13:17.363Z",
          "content": "<p><a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> what is your best cv so far? </p>",
          "rawMarkdown": "@morizin what is your best cv so far? "
        },
        {
          "id": 1248246,
          "postDate": "2021-03-22T13:16:29.590Z",
          "content": "<p>The final CV is being calculating now</p>",
          "rawMarkdown": "The final CV is being calculating now"
        }
      ]
    },
    {
      "id": 1249225,
      "postDate": "2021-03-23T07:23:17.903Z",
      "content": "<p>Shake-up is obvious in this competition base on the antecedent of previous similar ones.<br>\nThe LB will bring forth this trend. Based on discussion trends a lot of participants expect the shake-up but do not know what structure it will be introduced.</p>",
      "rawMarkdown": "Shake-up is obvious in this competition base on the antecedent of previous similar ones.\nThe LB will bring forth this trend. Based on discussion trends a lot of participants expect the shake-up but do not know what structure it will be introduced."
    }
  ],
  "comments": [
    {
      "id": 1199493,
      "author_name": "Issac",
      "author_url": "",
      "post_date": "2021-02-13T20:54:24.503000",
      "content": "<p>I think you're right. Even if you re-run your model for several times, the LB score will be quite different with little init randomization. And considering the annotations from train/test set are also different, I'm still having a hard time trying to find a reliable metric to select model.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1210227,
          "author_name": "Gleb Anferov",
          "author_url": "",
          "post_date": "2021-02-19T09:24:36.240000",
          "content": "<p><a href=\"https://www.kaggle.com/issacc\" target=\"_blank\">@issacc</a> I got lb range 12.0-18.0 map when re-running same model…  I am frustrated, what am i doing wrong lol?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1211677,
          "author_name": "Ryunosuke Ishizaki",
          "author_url": "",
          "post_date": "2021-02-20T12:44:00.320000",
          "content": "<p>Thanks for sharing ideas, we wanna get higher public/private lb, but when we seek upper public lb score, we have more risk for shake at private lb for espessially in this competition. dilemma🤔</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1215290,
      "author_name": "Kamal Das",
      "author_url": "",
      "post_date": "2021-02-23T14:04:01.527000",
      "content": "<p>Agree!!<br>\nwhich makes is more important to trust your CV and models and not get swayed by the Public LB</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1212210,
      "author_name": "Zekun",
      "author_url": "",
      "post_date": "2021-02-21T02:27:22.610000",
      "content": "<p>Hello , is this game only need to submit a table?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1207011,
      "author_name": "Hannes Öhler",
      "author_url": "",
      "post_date": "2021-02-17T17:18:17.643000",
      "content": "<p>I was just thinking: Looking at your frequency estimates, model selection based on public LB results restricted to the more frequent classes may be an option? These might be a bit more reliable than the very rare classes? On the other hand, different models may behaviour differently on rare vs. frequent classes.. But this I guess can only be assessed looking at the respective model and setup individualistically..</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1211668,
          "author_name": "Ryunosuke Ishizaki",
          "author_url": "",
          "post_date": "2021-02-20T12:29:10.643000",
          "content": "<p>I guess we can roughly refar to public LB results, for ex, if each models are single, public LB 0.25 vs 0.1, will be the same results as Private LB, but the more we start to add new process/system to test data,<br>\n(for ex, re-thinking bbox(nms, nmw, wbf, etc…)), you have more possibility of overfitting to Public LB.<br>\n(for my experience, Hold-out cv 0.25 vs Public LB 0.16 for simple model outputs,  hold-out cv 0.35 vs Public LB 0.12 for additional process.)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1211678,
          "author_name": "Hannes Öhler",
          "author_url": "",
          "post_date": "2021-02-20T12:44:52.837000",
          "content": "<p>Yes, this aggrevates the issue.. In my experiment I included the 2-class predictions and NMS so that might be an issue.. on the other hand, I still think that a single model may accidentally fit very well let's say the few cases of Atelectasis in the public LB set while an other model might not. At the same time the latter model could show a better performance on the private LB set where more cases of Atelectasis are included.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1199078,
      "author_name": " Sachin",
      "author_url": "",
      "post_date": "2021-02-13T14:05:17.437000",
      "content": "<p>All 3000 test images are in public. isn't it? </p>\n<p>am new to the competition so can you please tell me whether the leaderboard test data will be different from the available test data.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1199169,
          "author_name": "Stanley Zheng",
          "author_url": "",
          "post_date": "2021-02-13T15:02:18.163000",
          "content": "<p>They're all \"public\" as in you can view them, but only 300 are scored on the public LB. That's what he means.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1200851,
          "author_name": "Phat Tran",
          "author_url": "",
          "post_date": "2021-02-15T03:47:38.020000",
          "content": "<p>If I understand correctly, 3000 test images provided by the host include 300 for LB and 2700 remains for private LB?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1200853,
          "author_name": "Ryunosuke Ishizaki",
          "author_url": "",
          "post_date": "2021-02-15T03:50:42.050000",
          "content": "<p>Sure. right</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1248038,
      "author_name": "Mohammed Rizin V K",
      "author_url": "",
      "post_date": "2021-03-22T09:32:41.150000",
      "content": "<p>I think there would be shake up.<br>\nAll previous X-ray competition had a great shake up. So I expect the same here.<br>\nBut ….<br>\nto trust on CV is quite hard in this competition. Since <br>\nCV doesn't correlate nor proportional to LB Score </p>",
      "votes": 2,
      "replies": [
        {
          "id": 1248051,
          "author_name": "Ryunosuke Ishizaki",
          "author_url": "",
          "post_date": "2021-03-22T09:46:08.700000",
          "content": "<p>Yes, since my hold-out cv/lb score became mess, I stopped submitting,<br>\nthere are some factors of shake, public/private test label distribution difference, stage2 manual<br>\nlabeling from some radiologists, noisy training dataset of training set…etc<br>\nI guess last private lb top will be 0.2〜0.3, how do you think?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1248135,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2021-03-22T11:29:17.403000",
          "content": "<p>I think the same. but still i could imagine sometime our CV score might be the final private LB score</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1248243,
          "author_name": "Awsaf",
          "author_url": "",
          "post_date": "2021-03-22T13:13:17.363000",
          "content": "<p><a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> what is your best cv so far? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1248246,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2021-03-22T13:16:29.590000",
          "content": "<p>The final CV is being calculating now</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1249225,
      "author_name": "Olusesi Adebisi",
      "author_url": "",
      "post_date": "2021-03-23T07:23:17.903000",
      "content": "<p>Shake-up is obvious in this competition base on the antecedent of previous similar ones.<br>\nThe LB will bring forth this trend. Based on discussion trends a lot of participants expect the shake-up but do not know what structure it will be introduced.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1198730": "Hi there,\nnow we know 10% of total 3000 test data are public, rest 90% are private.\nwe can estimate host will prepare relatively more abnormal images for test, but if test distribution of each abnormalities are same as train/public test/private test, N of each abnormalities in test data\ncan be estimated as\npublic : [61,  3,  9, 46,  7,  7, 12, 26, 16, 22, 20, 39,  1, 32],\nprivate : [552,  33,  81, 413,  63,  69, 110, 237, 148, 204, 185, 356,  17, 291]\nthis may cause huge shakeup for last leaderboard. how do you think?",
    "1199493": "I think you're right. Even if you re-run your model for several times, the LB score will be quite different with little init randomization. And considering the annotations from train/test set are also different, I'm still having a hard time trying to find a reliable metric to select model.",
    "1215290": "Agree!!\nwhich makes is more important to trust your CV and models and not get swayed by the Public LB",
    "1212210": "Hello , is this game only need to submit a table?",
    "1207011": "I was just thinking: Looking at your frequency estimates, model selection based on public LB results restricted to the more frequent classes may be an option? These might be a bit more reliable than the very rare classes? On the other hand, different models may behaviour differently on rare vs. frequent classes.. But this I guess can only be assessed looking at the respective model and setup individualistically..",
    "1199078": "All 3000 test images are in public. isn't it? \n\nam new to the competition so can you please tell me whether the leaderboard test data will be different from the available test data.",
    "1248038": "I think there would be shake up.\nAll previous X-ray competition had a great shake up. So I expect the same here.\nBut ....\nto trust on CV is quite hard in this competition. Since \nCV doesn't correlate nor proportional to LB Score ",
    "1249225": "Shake-up is obvious in this competition base on the antecedent of previous similar ones.\nThe LB will bring forth this trend. Based on discussion trends a lot of participants expect the shake-up but do not know what structure it will be introduced."
  }
}