{
  "id": 214284,
  "title": "The  correlation between CV/Public score",
  "url": "/competitions/vinbigdata-chest-xray-abnormalities-detection/discussion/214284",
  "author_name": "ZehuiGong",
  "post_date": "2021-01-26T03:14:30.661000",
  "votes": 18,
  "comment_count": 28,
  "views": 0,
  "content": "<p>Hi, everyone. I found an issue here. My scores on CV vs. LB can not correspond to each other.  Below is an example:<br>\nCV: 0.332 --&gt; LB: 0.217<br>\nCV: 0.350 --&gt; LB: 0.173<br>\nCV: 0.331 --&gt;LB: 0.179<br>\nI am confused about the correlation between CV and LB. How are your scores on CV vs. LB? Posting mine below.</p>",
  "messages": [
    {
      "id": 1170154,
      "postDate": "2021-01-26T03:14:30.663Z",
      "content": "<p>Hi, everyone. I found an issue here. My scores on CV vs. LB can not correspond to each other.  Below is an example:<br>\nCV: 0.332 --&gt; LB: 0.217<br>\nCV: 0.350 --&gt; LB: 0.173<br>\nCV: 0.331 --&gt;LB: 0.179<br>\nI am confused about the correlation between CV and LB. How are your scores on CV vs. LB? Posting mine below.</p>",
      "rawMarkdown": "Hi, everyone. I found an issue here. My scores on CV vs. LB can not correspond to each other.  Below is an example:\nCV: 0.332 --> LB: 0.217\nCV: 0.350 --> LB: 0.173\nCV: 0.331 -->LB: 0.179\nI am confused about the correlation between CV and LB. How are your scores on CV vs. LB? Posting mine below.",
      "votes": 18
    },
    {
      "id": 1192229,
      "postDate": "2021-02-09T03:49:20.980Z",
      "content": "<p>Hi, Are your CV include \"No findings\" class and have you apply \"2 class filtering\" on your LB?</p>",
      "rawMarkdown": "Hi, Are your CV include \"No findings\" class and have you apply \"2 class filtering\" on your LB?",
      "replies": [
        {
          "id": 1193069,
          "postDate": "2021-02-09T12:45:48.143Z",
          "content": "<p>I found that there is no difference on CV, with or without \"No findings\". Yes, I have applied '2-class filtering'.</p>",
          "rawMarkdown": "I found that there is no difference on CV, with or without \"No findings\". Yes, I have applied '2-class filtering'.",
          "votes": 1
        },
        {
          "id": 1193141,
          "postDate": "2021-02-09T13:38:23.263Z",
          "content": "<p>I think it's because of noise in the annotation. As discussed <a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/212859\" target=\"_blank\">here</a>.</p>\n<p>My CV is ~0.27 and got only 0.102 PB (without 2-class filtering)</p>",
          "rawMarkdown": "I think it's because of noise in the annotation. As discussed [here](https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/212859).\n\nMy CV is ~0.27 and got only 0.102 PB (without 2-class filtering)"
        }
      ]
    },
    {
      "id": 1188203,
      "postDate": "2021-02-06T04:15:41.630Z",
      "content": "<blockquote>\n  <p>CV: 0.332 --&gt; LB: 0.217<br>\n  CV: 0.350 --&gt; LB: 0.173<br>\n  CV: 0.331 --&gt;LB: 0.179</p>\n</blockquote>\n<p>Are these single models? <a href=\"https://www.kaggle.com/zehuigong\" target=\"_blank\">@zehuigong</a> </p>",
      "rawMarkdown": "> CV: 0.332 --> LB: 0.217\nCV: 0.350 --> LB: 0.173\nCV: 0.331 -->LB: 0.179\n\nAre these single models? @zehuigong ",
      "replies": [
        {
          "id": 1190141,
          "postDate": "2021-02-07T14:09:17.947Z",
          "content": "<p>Yes, these are the result of a single model and single fold.</p>",
          "rawMarkdown": "Yes, these are the result of a single model and single fold."
        }
      ]
    },
    {
      "id": 1183045,
      "postDate": "2021-02-02T17:58:05.633Z",
      "content": "<p>The host said that , the test set is also annotated by 3 radiologists but these boxes were unified to one single box by 2 professional radiologists.<br>\nEither the 3 radiologists or 2 radiologists would be different. Maybe they are different in their annotations.</p>",
      "rawMarkdown": "The host said that , the test set is also annotated by 3 radiologists but these boxes were unified to one single box by 2 professional radiologists.\nEither the 3 radiologists or 2 radiologists would be different. Maybe they are different in their annotations.",
      "replies": [
        {
          "id": 1190154,
          "postDate": "2021-02-07T14:15:43.017Z",
          "content": "<p>Sorry for the late reply. As far as I know, the disagreement will be further judged by the radiologists and finally made a consensus.  In my opinion, the difference in annotation distribution between train/test sets makes this competition a more challenging one.</p>",
          "rawMarkdown": "Sorry for the late reply. As far as I know, the disagreement will be further judged by the radiologists and finally made a consensus.  In my opinion, the difference in annotation distribution between train/test sets makes this competition a more challenging one."
        }
      ]
    },
    {
      "id": 1181146,
      "postDate": "2021-02-01T17:20:28.553Z",
      "content": "<p>I guess we're about to see a huge shake up</p>",
      "rawMarkdown": "I guess we're about to see a huge shake up",
      "replies": [
        {
          "id": 1181643,
          "postDate": "2021-02-02T03:10:00.500Z",
          "content": "<p>Given the results that we have got now. I feel the same too. </p>",
          "rawMarkdown": "Given the results that we have got now. I feel the same too. "
        }
      ]
    },
    {
      "id": 1174174,
      "postDate": "2021-01-28T10:18:08.937Z",
      "content": "<p>i have the same problem. CV got 0.36 but PL so bad. </p>",
      "rawMarkdown": "i have the same problem. CV got 0.36 but PL so bad. ",
      "replies": [
        {
          "id": 1174200,
          "postDate": "2021-01-28T10:36:38.760Z",
          "content": "<p>Maybe we can trust our CV score and pay less attention to the PL.</p>",
          "rawMarkdown": "Maybe we can trust our CV score and pay less attention to the PL."
        }
      ]
    },
    {
      "id": 1170491,
      "postDate": "2021-01-26T08:47:09.157Z",
      "content": "<p>One of the thing we can understand from this is test set is almost different from the training set . we cannot find a reasonable score on LB as we get in CV</p>",
      "rawMarkdown": "One of the thing we can understand from this is test set is almost different from the training set . we cannot find a reasonable score on LB as we get in CV",
      "replies": [
        {
          "id": 1170506,
          "postDate": "2021-01-26T08:56:14.710Z",
          "content": "<p>Yes, if this is true, it is hard to improve our baseline model, because we do not know which score we could trust, CV or LB. May I ask does your improvement on CV also results in the improvement on LB?</p>",
          "rawMarkdown": "Yes, if this is true, it is hard to improve our baseline model, because we do not know which score we could trust, CV or LB. May I ask does your improvement on CV also results in the improvement on LB?"
        },
        {
          "id": 1170591,
          "postDate": "2021-01-26T09:56:36.630Z",
          "content": "<p>Yes My CV is almost 0.5 for the score of 0.318. We cant trust PublicLB because it is doing only in 300 images</p>",
          "rawMarkdown": "Yes My CV is almost 0.5 for the score of 0.318. We cant trust PublicLB because it is doing only in 300 images",
          "votes": 1
        },
        {
          "id": 1171040,
          "postDate": "2021-01-26T15:42:57.837Z",
          "content": "<p><a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a>, may I ask how you split your whole training set into the cross-validation set?</p>",
          "rawMarkdown": "@morizin, may I ask how you split your whole training set into the cross-validation set?"
        },
        {
          "id": 1173981,
          "postDate": "2021-01-28T08:02:13.953Z",
          "content": "<p>I have CV-LB disparencies too. <br>\nGuess it's a trust-CV competition given the small size of the public set.</p>",
          "rawMarkdown": "I have CV-LB disparencies too. \nGuess it's a trust-CV competition given the small size of the public set.",
          "votes": 2
        },
        {
          "id": 1174092,
          "postDate": "2021-01-28T09:05:58.310Z",
          "content": "<p><a href=\"https://www.kaggle.com/zehuigong\" target=\"_blank\">@zehuigong</a> I have done StratifiedKFold with labels</p>",
          "rawMarkdown": "@zehuigong I have done StratifiedKFold with labels"
        },
        {
          "id": 1174204,
          "postDate": "2021-01-28T10:40:12.593Z",
          "content": "<p>Got it, thanks for your sharing.</p>",
          "rawMarkdown": "Got it, thanks for your sharing."
        },
        {
          "id": 1174231,
          "postDate": "2021-01-28T11:00:40.483Z",
          "content": "<p>but why do you need mine it doesn't correlate with LB</p>",
          "rawMarkdown": "but why do you need mine it doesn't correlate with LB"
        },
        {
          "id": 1174268,
          "postDate": "2021-01-28T11:18:39.610Z",
          "content": "<p>Now, I think that many of us have encountered this issue, the CV-LB discrepancies. </p>",
          "rawMarkdown": "Now, I think that many of us have encountered this issue, the CV-LB discrepancies. "
        },
        {
          "id": 1181075,
          "postDate": "2021-02-01T16:26:44.103Z",
          "content": "<p>Are you measuring CV mAP against the raw ground truth? Or against NMS/WBF-transformed boxes?</p>",
          "rawMarkdown": "Are you measuring CV mAP against the raw ground truth? Or against NMS/WBF-transformed boxes?"
        },
        {
          "id": 1181085,
          "postDate": "2021-02-01T16:32:11.173Z",
          "content": "<p>I checked both with and without never correlates</p>",
          "rawMarkdown": "I checked both with and without never correlates"
        }
      ]
    },
    {
      "id": 1174651,
      "postDate": "2021-01-28T16:22:11.443Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 1175298,
          "postDate": "2021-01-29T04:32:06.910Z",
          "content": "<p>Have you tried different cross-validation set divisions, i.e., generating different 5 fold cross-validation?</p>",
          "rawMarkdown": "Have you tried different cross-validation set divisions, i.e., generating different 5 fold cross-validation?"
        },
        {
          "id": 1181525,
          "postDate": "2021-02-02T00:38:42.497Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1181646,
          "postDate": "2021-02-02T03:12:23.610Z",
          "content": "<p>Of all the divisions you make, none of them show a correlation between CV and LB score?</p>",
          "rawMarkdown": "Of all the divisions you make, none of them show a correlation between CV and LB score?"
        },
        {
          "id": 1182975,
          "postDate": "2021-02-02T17:04:53.377Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1182987,
          "postDate": "2021-02-02T17:11:27.850Z",
          "content": "<p>May I ask what kind of detection architecture(algorithm) are you experimenting with?</p>",
          "rawMarkdown": "May I ask what kind of detection architecture(algorithm) are you experimenting with?"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1192229,
      "author_name": "Phat Tran",
      "author_url": "",
      "post_date": "2021-02-09T03:49:20.980000",
      "content": "<p>Hi, Are your CV include \"No findings\" class and have you apply \"2 class filtering\" on your LB?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1193069,
          "author_name": "ZehuiGong",
          "author_url": "",
          "post_date": "2021-02-09T12:45:48.143000",
          "content": "<p>I found that there is no difference on CV, with or without \"No findings\". Yes, I have applied '2-class filtering'.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1193141,
          "author_name": "Phat Tran",
          "author_url": "",
          "post_date": "2021-02-09T13:38:23.263000",
          "content": "<p>I think it's because of noise in the annotation. As discussed <a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/212859\" target=\"_blank\">here</a>.</p>\n<p>My CV is ~0.27 and got only 0.102 PB (without 2-class filtering)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1188203,
      "author_name": "Mohammed Rizin V K",
      "author_url": "",
      "post_date": "2021-02-06T04:15:41.630000",
      "content": "<blockquote>\n  <p>CV: 0.332 --&gt; LB: 0.217<br>\n  CV: 0.350 --&gt; LB: 0.173<br>\n  CV: 0.331 --&gt;LB: 0.179</p>\n</blockquote>\n<p>Are these single models? <a href=\"https://www.kaggle.com/zehuigong\" target=\"_blank\">@zehuigong</a> </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1190141,
          "author_name": "ZehuiGong",
          "author_url": "",
          "post_date": "2021-02-07T14:09:17.947000",
          "content": "<p>Yes, these are the result of a single model and single fold.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1183045,
      "author_name": "Mohammed Rizin V K",
      "author_url": "",
      "post_date": "2021-02-02T17:58:05.633000",
      "content": "<p>The host said that , the test set is also annotated by 3 radiologists but these boxes were unified to one single box by 2 professional radiologists.<br>\nEither the 3 radiologists or 2 radiologists would be different. Maybe they are different in their annotations.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1190154,
          "author_name": "ZehuiGong",
          "author_url": "",
          "post_date": "2021-02-07T14:15:43.017000",
          "content": "<p>Sorry for the late reply. As far as I know, the disagreement will be further judged by the radiologists and finally made a consensus.  In my opinion, the difference in annotation distribution between train/test sets makes this competition a more challenging one.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1181146,
      "author_name": "Awsaf",
      "author_url": "",
      "post_date": "2021-02-01T17:20:28.553000",
      "content": "<p>I guess we're about to see a huge shake up</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1181643,
          "author_name": "ZehuiGong",
          "author_url": "",
          "post_date": "2021-02-02T03:10:00.500000",
          "content": "<p>Given the results that we have got now. I feel the same too. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1174174,
      "author_name": "Manh Lab",
      "author_url": "",
      "post_date": "2021-01-28T10:18:08.937000",
      "content": "<p>i have the same problem. CV got 0.36 but PL so bad. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1174200,
          "author_name": "ZehuiGong",
          "author_url": "",
          "post_date": "2021-01-28T10:36:38.760000",
          "content": "<p>Maybe we can trust our CV score and pay less attention to the PL.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1170491,
      "author_name": "Mohammed Rizin V K",
      "author_url": "",
      "post_date": "2021-01-26T08:47:09.157000",
      "content": "<p>One of the thing we can understand from this is test set is almost different from the training set . we cannot find a reasonable score on LB as we get in CV</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1170506,
          "author_name": "ZehuiGong",
          "author_url": "",
          "post_date": "2021-01-26T08:56:14.710000",
          "content": "<p>Yes, if this is true, it is hard to improve our baseline model, because we do not know which score we could trust, CV or LB. May I ask does your improvement on CV also results in the improvement on LB?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1170591,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2021-01-26T09:56:36.630000",
          "content": "<p>Yes My CV is almost 0.5 for the score of 0.318. We cant trust PublicLB because it is doing only in 300 images</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1171040,
          "author_name": "ZehuiGong",
          "author_url": "",
          "post_date": "2021-01-26T15:42:57.837000",
          "content": "<p><a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a>, may I ask how you split your whole training set into the cross-validation set?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1173981,
          "author_name": "arutema47",
          "author_url": "",
          "post_date": "2021-01-28T08:02:13.953000",
          "content": "<p>I have CV-LB disparencies too. <br>\nGuess it's a trust-CV competition given the small size of the public set.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1174092,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2021-01-28T09:05:58.310000",
          "content": "<p><a href=\"https://www.kaggle.com/zehuigong\" target=\"_blank\">@zehuigong</a> I have done StratifiedKFold with labels</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1174204,
          "author_name": "ZehuiGong",
          "author_url": "",
          "post_date": "2021-01-28T10:40:12.593000",
          "content": "<p>Got it, thanks for your sharing.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1174231,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2021-01-28T11:00:40.483000",
          "content": "<p>but why do you need mine it doesn't correlate with LB</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1174268,
          "author_name": "ZehuiGong",
          "author_url": "",
          "post_date": "2021-01-28T11:18:39.610000",
          "content": "<p>Now, I think that many of us have encountered this issue, the CV-LB discrepancies. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1181075,
          "author_name": "adamson",
          "author_url": "",
          "post_date": "2021-02-01T16:26:44.103000",
          "content": "<p>Are you measuring CV mAP against the raw ground truth? Or against NMS/WBF-transformed boxes?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1181085,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2021-02-01T16:32:11.173000",
          "content": "<p>I checked both with and without never correlates</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1174651,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-01-28T16:22:11.443000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1175298,
          "author_name": "ZehuiGong",
          "author_url": "",
          "post_date": "2021-01-29T04:32:06.910000",
          "content": "<p>Have you tried different cross-validation set divisions, i.e., generating different 5 fold cross-validation?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1181525,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-02-02T00:38:42.497000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1181646,
          "author_name": "ZehuiGong",
          "author_url": "",
          "post_date": "2021-02-02T03:12:23.610000",
          "content": "<p>Of all the divisions you make, none of them show a correlation between CV and LB score?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1182975,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-02-02T17:04:53.377000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1182987,
          "author_name": "ZehuiGong",
          "author_url": "",
          "post_date": "2021-02-02T17:11:27.850000",
          "content": "<p>May I ask what kind of detection architecture(algorithm) are you experimenting with?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1170154": "Hi, everyone. I found an issue here. My scores on CV vs. LB can not correspond to each other.  Below is an example:\nCV: 0.332 --> LB: 0.217\nCV: 0.350 --> LB: 0.173\nCV: 0.331 -->LB: 0.179\nI am confused about the correlation between CV and LB. How are your scores on CV vs. LB? Posting mine below.",
    "1192229": "Hi, Are your CV include \"No findings\" class and have you apply \"2 class filtering\" on your LB?",
    "1188203": "> CV: 0.332 --> LB: 0.217\nCV: 0.350 --> LB: 0.173\nCV: 0.331 -->LB: 0.179\n\nAre these single models? @zehuigong ",
    "1183045": "The host said that , the test set is also annotated by 3 radiologists but these boxes were unified to one single box by 2 professional radiologists.\nEither the 3 radiologists or 2 radiologists would be different. Maybe they are different in their annotations.",
    "1181146": "I guess we're about to see a huge shake up",
    "1174174": "i have the same problem. CV got 0.36 but PL so bad. ",
    "1170491": "One of the thing we can understand from this is test set is almost different from the training set . we cannot find a reasonable score on LB as we get in CV",
    "1174651": ""
  }
}