{
  "id": 189061,
  "title": "any hint to beat baseline",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/189061",
  "author_name": "Manh Lab",
  "post_date": "2020-10-06T14:10:15.459000",
  "votes": 1,
  "comment_count": 17,
  "views": 0,
  "content": "<p>It's a very unbalanced competition and difficult to beat baseline. Any experience is nice to share all !!!</p>",
  "messages": [
    {
      "id": 1040349,
      "postDate": "2020-10-07T05:08:39.953Z",
      "content": "<p>you can get 0.32 on LB without any image.</p>",
      "rawMarkdown": "you can get 0.32 on LB without any image.",
      "votes": 4,
      "replies": [
        {
          "id": 1046083,
          "postDate": "2020-10-11T10:28:27.347Z",
          "content": "<p>If you aren't using the image data, don't you suspect you would be mostly overfitting on noise in the dicom metadata?</p>",
          "rawMarkdown": "If you aren't using the image data, don't you suspect you would be mostly overfitting on noise in the dicom metadata?"
        },
        {
          "id": 1046109,
          "postDate": "2020-10-11T11:01:19.763Z",
          "content": "<p>I don't think the average can overfit. <a href=\"https://www.kaggle.com/osciiart/baseline-with-no-image\" target=\"_blank\">https://www.kaggle.com/osciiart/baseline-with-no-image</a></p>",
          "rawMarkdown": "I don't think the average can overfit. https://www.kaggle.com/osciiart/baseline-with-no-image",
          "votes": 4
        },
        {
          "id": 1046140,
          "postDate": "2020-10-11T11:34:51.723Z",
          "content": "<p>It's overfitting I think. you can see all labels: 0.6746805906295776. just the default</p>",
          "rawMarkdown": "It's overfitting I think. you can see all labels: 0.6746805906295776. just the default",
          "votes": -2
        },
        {
          "id": 1048042,
          "postDate": "2020-10-13T06:34:32.957Z",
          "content": "<p>Please study the kernel of <a href=\"https://www.kaggle.com/osciiart\" target=\"_blank\">@osciiart</a> again <a href=\"https://www.kaggle.com/doanquanvietnamca\" target=\"_blank\">@doanquanvietnamca</a><br>\nYou are right, that this is fitting the training data, but the test data seems to have a similar distribution, thus taking the means is a valid approach for a baseline. In the last part, <a href=\"https://www.kaggle.com/osciiart\" target=\"_blank\">@osciiart</a> cleverly used the image level evaluation metric (<strong>weighted</strong> logloss) to maximize the score (or better minimize). </p>",
          "rawMarkdown": "Please study the kernel of @osciiart again @doanquanvietnamca\nYou are right, that this is fitting the training data, but the test data seems to have a similar distribution, thus taking the means is a valid approach for a baseline. In the last part, @osciiart cleverly used the image level evaluation metric (**weighted** logloss) to maximize the score (or better minimize). ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1039347,
      "postDate": "2020-10-06T14:10:15.460Z",
      "content": "<p>It's a very unbalanced competition and difficult to beat baseline. Any experience is nice to share all !!!</p>",
      "rawMarkdown": "It's a very unbalanced competition and difficult to beat baseline. Any experience is nice to share all !!!",
      "votes": 1
    },
    {
      "id": 1039857,
      "postDate": "2020-10-06T20:45:01.600Z",
      "content": "<p>yeah </p>",
      "rawMarkdown": "yeah ",
      "votes": -2
    },
    {
      "id": 1039691,
      "postDate": "2020-10-06T18:10:59.273Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 1040990,
          "postDate": "2020-10-07T13:37:25.913Z",
          "content": "<p>check if you are exceeding the allowed compute. if this is the case, try to reduce the number of workers on the pytorch dataloader. Also check if your RAM is constantly increasing. this could be due to memory leak</p>",
          "rawMarkdown": "check if you are exceeding the allowed compute. if this is the case, try to reduce the number of workers on the pytorch dataloader. Also check if your RAM is constantly increasing. this could be due to memory leak"
        },
        {
          "id": 1041705,
          "postDate": "2020-10-07T22:03:51.283Z",
          "content": "<p>I'm using tensorflow, and even doing inference on the entire training set works (though it does take 8 hours). I'm absolutely stumped as to what's holding my inference up…. I've resorted to just doing public LB in a kernel then submitting the CSV with a few lines of code</p>",
          "rawMarkdown": "I'm using tensorflow, and even doing inference on the entire training set works (though it does take 8 hours). I'm absolutely stumped as to what's holding my inference up.... I've resorted to just doing public LB in a kernel then submitting the CSV with a few lines of code",
          "votes": 1
        },
        {
          "id": 1041757,
          "postDate": "2020-10-07T22:44:45.533Z",
          "content": "<p>check if you are getting submission error . this could be due to pydicom failing to read few image files and your submission csv might not include those entried . this will raise submission error.(assuming you have taken care of the ID_label and other formating requirements).</p>\n<p>see if your final submission csv is having an index column. remove it.</p>\n<p>its its exceed computer error , then its most likely due to CPU RAM memory leak. but, if you could run through train data without any issue then i guess this should not be your problem.</p>\n<p>these were some of the issues i faced before getting my inference submission working.<br>\ntook me 2 weeks to figure it out. </p>",
          "rawMarkdown": "check if you are getting submission error . this could be due to pydicom failing to read few image files and your submission csv might not include those entried . this will raise submission error.(assuming you have taken care of the ID_label and other formating requirements).\n\nsee if your final submission csv is having an index column. remove it.\n\nits its exceed computer error , then its most likely due to CPU RAM memory leak. but, if you could run through train data without any issue then i guess this should not be your problem.\n\nthese were some of the issues i faced before getting my inference submission working.\ntook me 2 weeks to figure it out. \n"
        },
        {
          "id": 1047879,
          "postDate": "2020-10-13T02:57:13.097Z",
          "content": "<p>Sorry, can you tell me what's the TTA meas ?</p>",
          "rawMarkdown": "Sorry, can you tell me what's the TTA meas ?"
        },
        {
          "id": 1047884,
          "postDate": "2020-10-13T03:04:52.367Z",
          "content": "<p>Test-time augmentation</p>",
          "rawMarkdown": "Test-time augmentation"
        },
        {
          "id": 1048151,
          "postDate": "2020-10-13T08:35:34.750Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1048570,
          "postDate": "2020-10-13T15:55:02.920Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1048574,
          "postDate": "2020-10-13T16:00:30.930Z",
          "content": "<p>Hi. can you elaborate on this </p>\n<blockquote>\n  <p>I also migrated to Colab using the GCS bucket path given by Kaggle to save hours.</p>\n</blockquote>",
          "rawMarkdown": "Hi. can you elaborate on this \n>  I also migrated to Colab using the GCS bucket path given by Kaggle to save hours."
        },
        {
          "id": 1048582,
          "postDate": "2020-10-13T16:08:09.633Z",
          "content": "<p>Use GCS path given by kaggle (format gcs://xxxxxxxxxxxxxxxxxxxxxxxx) to use the dataset with Colab. I guess you could also download the dataset using Kaggle API, but my dataset is private and I'm not sure if it works</p>",
          "rawMarkdown": "Use GCS path given by kaggle (format gcs://xxxxxxxxxxxxxxxxxxxxxxxx) to use the dataset with Colab. I guess you could also download the dataset using Kaggle API, but my dataset is private and I'm not sure if it works",
          "votes": 1
        },
        {
          "id": 1048633,
          "postDate": "2020-10-13T17:05:40.203Z",
          "content": "<p>Thank you!</p>",
          "rawMarkdown": "Thank you!"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1040349,
      "author_name": "OsciiArt",
      "author_url": "",
      "post_date": "2020-10-07T05:08:39.953000",
      "content": "<p>you can get 0.32 on LB without any image.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1046083,
          "author_name": "aksg87",
          "author_url": "",
          "post_date": "2020-10-11T10:28:27.347000",
          "content": "<p>If you aren't using the image data, don't you suspect you would be mostly overfitting on noise in the dicom metadata?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1046109,
          "author_name": "OsciiArt",
          "author_url": "",
          "post_date": "2020-10-11T11:01:19.763000",
          "content": "<p>I don't think the average can overfit. <a href=\"https://www.kaggle.com/osciiart/baseline-with-no-image\" target=\"_blank\">https://www.kaggle.com/osciiart/baseline-with-no-image</a></p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1046140,
          "author_name": "Manh Lab",
          "author_url": "",
          "post_date": "2020-10-11T11:34:51.723000",
          "content": "<p>It's overfitting I think. you can see all labels: 0.6746805906295776. just the default</p>",
          "votes": -2,
          "replies": []
        },
        {
          "id": 1048042,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2020-10-13T06:34:32.957000",
          "content": "<p>Please study the kernel of <a href=\"https://www.kaggle.com/osciiart\" target=\"_blank\">@osciiart</a> again <a href=\"https://www.kaggle.com/doanquanvietnamca\" target=\"_blank\">@doanquanvietnamca</a><br>\nYou are right, that this is fitting the training data, but the test data seems to have a similar distribution, thus taking the means is a valid approach for a baseline. In the last part, <a href=\"https://www.kaggle.com/osciiart\" target=\"_blank\">@osciiart</a> cleverly used the image level evaluation metric (<strong>weighted</strong> logloss) to maximize the score (or better minimize). </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1039857,
      "author_name": "Deependra Yadav",
      "author_url": "",
      "post_date": "2020-10-06T20:45:01.600000",
      "content": "<p>yeah </p>",
      "votes": -2,
      "replies": []
    },
    {
      "id": 1039691,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-06T18:10:59.273000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1040990,
          "author_name": "yuvaramsingh",
          "author_url": "",
          "post_date": "2020-10-07T13:37:25.913000",
          "content": "<p>check if you are exceeding the allowed compute. if this is the case, try to reduce the number of workers on the pytorch dataloader. Also check if your RAM is constantly increasing. this could be due to memory leak</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1041705,
          "author_name": "Stanley Zheng",
          "author_url": "",
          "post_date": "2020-10-07T22:03:51.283000",
          "content": "<p>I'm using tensorflow, and even doing inference on the entire training set works (though it does take 8 hours). I'm absolutely stumped as to what's holding my inference up…. I've resorted to just doing public LB in a kernel then submitting the CSV with a few lines of code</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1041757,
          "author_name": "yuvaramsingh",
          "author_url": "",
          "post_date": "2020-10-07T22:44:45.533000",
          "content": "<p>check if you are getting submission error . this could be due to pydicom failing to read few image files and your submission csv might not include those entried . this will raise submission error.(assuming you have taken care of the ID_label and other formating requirements).</p>\n<p>see if your final submission csv is having an index column. remove it.</p>\n<p>its its exceed computer error , then its most likely due to CPU RAM memory leak. but, if you could run through train data without any issue then i guess this should not be your problem.</p>\n<p>these were some of the issues i faced before getting my inference submission working.<br>\ntook me 2 weeks to figure it out. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1047879,
          "author_name": "saatuo",
          "author_url": "",
          "post_date": "2020-10-13T02:57:13.097000",
          "content": "<p>Sorry, can you tell me what's the TTA meas ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1047884,
          "author_name": "Johnny Lee",
          "author_url": "",
          "post_date": "2020-10-13T03:04:52.367000",
          "content": "<p>Test-time augmentation</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1048151,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-10-13T08:35:34.750000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1048570,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-10-13T15:55:02.920000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1048574,
          "author_name": "A.Demyanchuk",
          "author_url": "",
          "post_date": "2020-10-13T16:00:30.930000",
          "content": "<p>Hi. can you elaborate on this </p>\n<blockquote>\n  <p>I also migrated to Colab using the GCS bucket path given by Kaggle to save hours.</p>\n</blockquote>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1048582,
          "author_name": "Stanley Zheng",
          "author_url": "",
          "post_date": "2020-10-13T16:08:09.633000",
          "content": "<p>Use GCS path given by kaggle (format gcs://xxxxxxxxxxxxxxxxxxxxxxxx) to use the dataset with Colab. I guess you could also download the dataset using Kaggle API, but my dataset is private and I'm not sure if it works</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1048633,
          "author_name": "A.Demyanchuk",
          "author_url": "",
          "post_date": "2020-10-13T17:05:40.203000",
          "content": "<p>Thank you!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1040349": "you can get 0.32 on LB without any image.",
    "1039347": "It's a very unbalanced competition and difficult to beat baseline. Any experience is nice to share all !!!",
    "1039857": "yeah ",
    "1039691": ""
  }
}