{
  "id": 219201,
  "title": "LB score as an unreliable indicator",
  "url": "/competitions/vinbigdata-chest-xray-abnormalities-detection/discussion/219201",
  "author_name": "Issac",
  "post_date": "2021-02-13T21:08:54.277000",
  "votes": 8,
  "comment_count": 6,
  "views": 0,
  "content": "<p>If I understand correctly, LB score is only evaluated on 300 images, and the distribution of diseases in the 300 images might differ a lot from the training set. So the score only reflect how your model performs on the most diseases included in the 300 images.</p>\n<p>I'm pointing this out because I found that the change of the LB score doesn't match my val set (20% images with the similar disease distribution with training set) at all. </p>\n<p>E.g., the best LB score model is A and the best val map@0.5 model is B. B gets a lower LB score and A gets a lower val MAP. However, A does perform better on several diseases than B on the val set (e.g.,  long opacity). Based on these observation I'm assuming the 300 images are most with some specific diseases and are quite unbalanced. This could be a problem for stage 2 eval if you only look at the LB score.</p>\n<p>I'm not aware of any posts about evaluation set info, but if the 3000 evaluation set has a very different disease distribution, then I guess it'd be a problem. Think about an extreme case where there's only 5 diseases exist in the evaluation set.</p>\n<p>Update: in the original paper it says<br>\n<code>A set of 18,000\nCXRs were randomly chosen from the filtered data, of which 15,000 scans serve as the training set and the rest 3,000 form the test set</code></p>\n<p>I think this random selection might make the rare diseases in the training set even rarer in the test set?</p>",
  "messages": [
    {
      "id": 1199505,
      "postDate": "2021-02-13T21:08:54.277Z",
      "content": "<p>If I understand correctly, LB score is only evaluated on 300 images, and the distribution of diseases in the 300 images might differ a lot from the training set. So the score only reflect how your model performs on the most diseases included in the 300 images.</p>\n<p>I'm pointing this out because I found that the change of the LB score doesn't match my val set (20% images with the similar disease distribution with training set) at all. </p>\n<p>E.g., the best LB score model is A and the best val map@0.5 model is B. B gets a lower LB score and A gets a lower val MAP. However, A does perform better on several diseases than B on the val set (e.g.,  long opacity). Based on these observation I'm assuming the 300 images are most with some specific diseases and are quite unbalanced. This could be a problem for stage 2 eval if you only look at the LB score.</p>\n<p>I'm not aware of any posts about evaluation set info, but if the 3000 evaluation set has a very different disease distribution, then I guess it'd be a problem. Think about an extreme case where there's only 5 diseases exist in the evaluation set.</p>\n<p>Update: in the original paper it says<br>\n<code>A set of 18,000\nCXRs were randomly chosen from the filtered data, of which 15,000 scans serve as the training set and the rest 3,000 form the test set</code></p>\n<p>I think this random selection might make the rare diseases in the training set even rarer in the test set?</p>",
      "rawMarkdown": "If I understand correctly, LB score is only evaluated on 300 images, and the distribution of diseases in the 300 images might differ a lot from the training set. So the score only reflect how your model performs on the most diseases included in the 300 images.\n\nI'm pointing this out because I found that the change of the LB score doesn't match my val set (20% images with the similar disease distribution with training set) at all. \n\nE.g., the best LB score model is A and the best val map@0.5 model is B. B gets a lower LB score and A gets a lower val MAP. However, A does perform better on several diseases than B on the val set (e.g.,  long opacity). Based on these observation I'm assuming the 300 images are most with some specific diseases and are quite unbalanced. This could be a problem for stage 2 eval if you only look at the LB score.\n\nI'm not aware of any posts about evaluation set info, but if the 3000 evaluation set has a very different disease distribution, then I guess it'd be a problem. Think about an extreme case where there's only 5 diseases exist in the evaluation set.\n\nUpdate: in the original paper it says\n`A set of 18,000\nCXRs were randomly chosen from the filtered data, of which 15,000 scans serve as the training set and the rest 3,000 form the test set`\n\nI think this random selection might make the rare diseases in the training set even rarer in the test set?",
      "votes": 8
    },
    {
      "id": 1211656,
      "postDate": "2021-02-20T12:10:35.733Z",
      "content": "<p>Also the public test set do have lots of no finding images. I think when we suppress more into no finding by <a href=\"https://www.kaggle.com/corochann\" target=\"_blank\">@corochann</a> method we might overfit at private where may some images as no finding</p>",
      "rawMarkdown": "Also the public test set do have lots of no finding images. I think when we suppress more into no finding by @corochann method we might overfit at private where may some images as no finding"
    },
    {
      "id": 1200404,
      "postDate": "2021-02-14T16:18:48.827Z",
      "content": "<p>Oh, I think it's 3000 public test images, and 27000 private test images before.<br>\nSo did I run inference for 3000 images and only 300 are used for scoring?.</p>",
      "rawMarkdown": "Oh, I think it's 3000 public test images, and 27000 private test images before.\nSo did I run inference for 3000 images and only 300 are used for scoring?.",
      "replies": [
        {
          "id": 1211215,
          "postDate": "2021-02-20T04:26:50.960Z",
          "content": "<p>The test set contains 3000 images. The public score will be calculated on approx. 10%(300) of the test set and will be visible to you during the competition. Private LB will be calculated on the remaining 90%(2700) of the data.</p>",
          "rawMarkdown": "The test set contains 3000 images. The public score will be calculated on approx. 10%(300) of the test set and will be visible to you during the competition. Private LB will be calculated on the remaining 90%(2700) of the data."
        },
        {
          "id": 1212306,
          "postDate": "2021-02-21T05:13:41.743Z",
          "content": "<blockquote>\n  <ol>\n  <li>The test set contains 3000 images. The public score will be calculated on approx. 10% of the test set and will be visible to you during the competition. Final rankings will be determined on the remaining 90% i.e. private LB</li>\n  </ol>\n</blockquote>\n<p>This was said by the competition host. Can someone please confirm if I understood this correctly. </p>",
          "rawMarkdown": "> 1. The test set contains 3000 images. The public score will be calculated on approx. 10% of the test set and will be visible to you during the competition. Final rankings will be determined on the remaining 90% i.e. private LB\n> \n\nThis was said by the competition host. Can someone please confirm if I understood this correctly. ",
          "votes": 1
        },
        {
          "id": 1212311,
          "postDate": "2021-02-21T05:21:19.540Z",
          "content": "<p>out of the 3000 images 300 is public and other 2700 is private. You can see and visualize private set dicom files from the given whole test set. from the 300 public test set (which is selected randomly or in any other way) the public LB will show score only of that 300 images the rest of 2700 is being evaluated but it will be only visible after the deadline of the competition finishes</p>\n<p>I hope you got my point </p>",
          "rawMarkdown": "out of the 3000 images 300 is public and other 2700 is private. You can see and visualize private set dicom files from the given whole test set. from the 300 public test set (which is selected randomly or in any other way) the public LB will show score only of that 300 images the rest of 2700 is being evaluated but it will be only visible after the deadline of the competition finishes\n\nI hope you got my point ",
          "votes": 2
        },
        {
          "id": 1212317,
          "postDate": "2021-02-21T05:33:49.823Z",
          "content": "<p>Yeah, understood. Thanks <a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> </p>",
          "rawMarkdown": "Yeah, understood. Thanks @morizin ",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1211656,
      "author_name": "Mohammed Rizin V K",
      "author_url": "",
      "post_date": "2021-02-20T12:10:35.733000",
      "content": "<p>Also the public test set do have lots of no finding images. I think when we suppress more into no finding by <a href=\"https://www.kaggle.com/corochann\" target=\"_blank\">@corochann</a> method we might overfit at private where may some images as no finding</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1200404,
      "author_name": "Phat Tran",
      "author_url": "",
      "post_date": "2021-02-14T16:18:48.827000",
      "content": "<p>Oh, I think it's 3000 public test images, and 27000 private test images before.<br>\nSo did I run inference for 3000 images and only 300 are used for scoring?.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1211215,
          "author_name": "Mohit Gidwani",
          "author_url": "",
          "post_date": "2021-02-20T04:26:50.960000",
          "content": "<p>The test set contains 3000 images. The public score will be calculated on approx. 10%(300) of the test set and will be visible to you during the competition. Private LB will be calculated on the remaining 90%(2700) of the data.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1212306,
          "author_name": "Mohit Gidwani",
          "author_url": "",
          "post_date": "2021-02-21T05:13:41.743000",
          "content": "<blockquote>\n  <ol>\n  <li>The test set contains 3000 images. The public score will be calculated on approx. 10% of the test set and will be visible to you during the competition. Final rankings will be determined on the remaining 90% i.e. private LB</li>\n  </ol>\n</blockquote>\n<p>This was said by the competition host. Can someone please confirm if I understood this correctly. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1212311,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2021-02-21T05:21:19.540000",
          "content": "<p>out of the 3000 images 300 is public and other 2700 is private. You can see and visualize private set dicom files from the given whole test set. from the 300 public test set (which is selected randomly or in any other way) the public LB will show score only of that 300 images the rest of 2700 is being evaluated but it will be only visible after the deadline of the competition finishes</p>\n<p>I hope you got my point </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1212317,
          "author_name": "Mohit Gidwani",
          "author_url": "",
          "post_date": "2021-02-21T05:33:49.823000",
          "content": "<p>Yeah, understood. Thanks <a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> </p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1199505": "If I understand correctly, LB score is only evaluated on 300 images, and the distribution of diseases in the 300 images might differ a lot from the training set. So the score only reflect how your model performs on the most diseases included in the 300 images.\n\nI'm pointing this out because I found that the change of the LB score doesn't match my val set (20% images with the similar disease distribution with training set) at all. \n\nE.g., the best LB score model is A and the best val map@0.5 model is B. B gets a lower LB score and A gets a lower val MAP. However, A does perform better on several diseases than B on the val set (e.g.,  long opacity). Based on these observation I'm assuming the 300 images are most with some specific diseases and are quite unbalanced. This could be a problem for stage 2 eval if you only look at the LB score.\n\nI'm not aware of any posts about evaluation set info, but if the 3000 evaluation set has a very different disease distribution, then I guess it'd be a problem. Think about an extreme case where there's only 5 diseases exist in the evaluation set.\n\nUpdate: in the original paper it says\n`A set of 18,000\nCXRs were randomly chosen from the filtered data, of which 15,000 scans serve as the training set and the rest 3,000 form the test set`\n\nI think this random selection might make the rare diseases in the training set even rarer in the test set?",
    "1211656": "Also the public test set do have lots of no finding images. I think when we suppress more into no finding by @corochann method we might overfit at private where may some images as no finding",
    "1200404": "Oh, I think it's 3000 public test images, and 27000 private test images before.\nSo did I run inference for 3000 images and only 300 are used for scoring?."
  }
}