{
  "id": 191694,
  "title": "Getting good cv but no lb ",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/191694",
  "author_name": "Jaideep",
  "post_date": "2020-10-18T02:43:16.590000",
  "votes": 5,
  "comment_count": 7,
  "views": 0,
  "content": "<p>can any one help… i get around approximated cv of .238 but lb is double…where could lie the mistake</p>",
  "messages": [
    {
      "id": 1052614,
      "postDate": "2020-10-18T02:43:16.590Z",
      "content": "<p>can any one help… i get around approximated cv of .238 but lb is double…where could lie the mistake</p>",
      "rawMarkdown": "can any one help... i get around approximated cv of .238 but lb is double...where could lie the mistake",
      "votes": 5
    },
    {
      "id": 1053001,
      "postDate": "2020-10-18T13:50:44.540Z",
      "content": "<p>Do you use the same metrics as the one of the competition? Be aware of the fact that you have to take into account specific weights for each label.</p>",
      "rawMarkdown": "Do you use the same metrics as the one of the competition? Be aware of the fact that you have to take into account specific weights for each label.",
      "votes": 1,
      "replies": [
        {
          "id": 1053050,
          "postDate": "2020-10-18T14:49:10.327Z",
          "content": "<p>no i dont use that.. simple bce </p>",
          "rawMarkdown": "no i dont use that.. simple bce ",
          "votes": 1
        },
        {
          "id": 1053086,
          "postDate": "2020-10-18T15:26:13.920Z",
          "content": "<p>Then that might be the reason, among others. </p>",
          "rawMarkdown": "Then that might be the reason, among others. "
        }
      ]
    },
    {
      "id": 1052780,
      "postDate": "2020-10-18T08:42:36.460Z",
      "content": "<p>I think you should trust your cv score,  as the public  distribution of test set maybe have some difference with  all test set or train set.</p>",
      "rawMarkdown": "I think you should trust your cv score,  as the public  distribution of test set maybe have some difference with  all test set or train set.",
      "votes": 1
    },
    {
      "id": 1053495,
      "postDate": "2020-10-19T04:04:55.150Z",
      "content": "<p>BTW, they also have a label consistency check: <a href=\"url\" target=\"_blank\">https://www.kaggle.com/anthracene/host-confirmed-label-consistency-check</a>, which will only be applied to the top 10 though. </p>",
      "rawMarkdown": "BTW, they also have a label consistency check: [https://www.kaggle.com/anthracene/host-confirmed-label-consistency-check](url), which will only be applied to the top 10 though. "
    },
    {
      "id": 1052685,
      "postDate": "2020-10-18T06:22:38.963Z",
      "content": "<p>It is hard to say with limited info, but here is one guess:  if you used defaults that randomly select your validation set from the rows of the training data (by image), you have information leakage from your training data to your CV. This is because adjacent images are practically identical, and all scans in a series will have some anatomy in common, so you don’t want images from the same study to be found in both your training and validation sets.</p>",
      "rawMarkdown": "It is hard to say with limited info, but here is one guess:  if you used defaults that randomly select your validation set from the rows of the training data (by image), you have information leakage from your training data to your CV. This is because adjacent images are practically identical, and all scans in a series will have some anatomy in common, so you don’t want images from the same study to be found in both your training and validation sets.\n\n\n\n\n\n",
      "replies": [
        {
          "id": 1052707,
          "postDate": "2020-10-18T06:59:34.317Z",
          "content": "<p>i segregate based on patient ids selected for val/train set based on a criteria ..<br>\nso at the end val will have different std ids for validation and train will have diff std ids for training. <br>\nFor training i simple put img loss and std loss separately as shown in Data description and evaluation section</p>\n<p>Should the order of  ids in sub be same as in Sample Sub ?</p>",
          "rawMarkdown": "i segregate based on patient ids selected for val/train set based on a criteria ..\nso at the end val will have different std ids for validation and train will have diff std ids for training. \nFor training i simple put img loss and std loss separately as shown in Data description and evaluation section\n\nShould the order of  ids in sub be same as in Sample Sub ?\n\n",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1053001,
      "author_name": "Catadanna",
      "author_url": "",
      "post_date": "2020-10-18T13:50:44.540000",
      "content": "<p>Do you use the same metrics as the one of the competition? Be aware of the fact that you have to take into account specific weights for each label.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1053050,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-10-18T14:49:10.327000",
          "content": "<p>no i dont use that.. simple bce </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1053086,
          "author_name": "Catadanna",
          "author_url": "",
          "post_date": "2020-10-18T15:26:13.920000",
          "content": "<p>Then that might be the reason, among others. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1052780,
      "author_name": "insight001",
      "author_url": "",
      "post_date": "2020-10-18T08:42:36.460000",
      "content": "<p>I think you should trust your cv score,  as the public  distribution of test set maybe have some difference with  all test set or train set.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1053495,
      "author_name": "Yang Taorui",
      "author_url": "",
      "post_date": "2020-10-19T04:04:55.150000",
      "content": "<p>BTW, they also have a label consistency check: <a href=\"url\" target=\"_blank\">https://www.kaggle.com/anthracene/host-confirmed-label-consistency-check</a>, which will only be applied to the top 10 though. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1052685,
      "author_name": "RS Turley",
      "author_url": "",
      "post_date": "2020-10-18T06:22:38.963000",
      "content": "<p>It is hard to say with limited info, but here is one guess:  if you used defaults that randomly select your validation set from the rows of the training data (by image), you have information leakage from your training data to your CV. This is because adjacent images are practically identical, and all scans in a series will have some anatomy in common, so you don’t want images from the same study to be found in both your training and validation sets.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1052707,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-10-18T06:59:34.317000",
          "content": "<p>i segregate based on patient ids selected for val/train set based on a criteria ..<br>\nso at the end val will have different std ids for validation and train will have diff std ids for training. <br>\nFor training i simple put img loss and std loss separately as shown in Data description and evaluation section</p>\n<p>Should the order of  ids in sub be same as in Sample Sub ?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1052614": "can any one help... i get around approximated cv of .238 but lb is double...where could lie the mistake",
    "1053001": "Do you use the same metrics as the one of the competition? Be aware of the fact that you have to take into account specific weights for each label.",
    "1052780": "I think you should trust your cv score,  as the public  distribution of test set maybe have some difference with  all test set or train set.",
    "1053495": "BTW, they also have a label consistency check: [https://www.kaggle.com/anthracene/host-confirmed-label-consistency-check](url), which will only be applied to the top 10 though. ",
    "1052685": "It is hard to say with limited info, but here is one guess:  if you used defaults that randomly select your validation set from the rows of the training data (by image), you have information leakage from your training data to your CV. This is because adjacent images are practically identical, and all scans in a series will have some anatomy in common, so you don’t want images from the same study to be found in both your training and validation sets.\n\n\n\n\n\n"
  }
}