{
  "id": 190278,
  "title": "How approach is better?",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/190278",
  "author_name": "imori",
  "post_date": "2020-10-11T02:03:57.927000",
  "votes": 0,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I somehow know output, and tried to use CNN.<br>\nBut this dataset contains too large image, so using all images take too much time.<br>\nI think do we need to know efficient sampling method for training?</p>\n<p>And, Some notebooks uses gbdt as tabluar data , but I think image is important when diagnosis…<br>\nIs this idea wrong?   Could you tell me your opinion..</p>",
  "messages": [
    {
      "id": 1045762,
      "postDate": "2020-10-11T02:25:26.233Z",
      "content": "<p>crop or resize images to make them smaller. work with 256x256 for example</p>\n<p>I don't expect tabular data by itself will be enough.</p>\n<p>You don't need to train on all the images. Make an enriched dataset with all the positive studies but not all the negative studies.</p>\n<p>There are posted datasets with jpegs and tfrecords. You can run a typical CNN with TPU for training with over half the training data in less than 3 hours.</p>\n<p>But, this is a challenging competition. No question about it!</p>",
      "rawMarkdown": "crop or resize images to make them smaller. work with 256x256 for example\n\nI don't expect tabular data by itself will be enough.\n\nYou don't need to train on all the images. Make an enriched dataset with all the positive studies but not all the negative studies.\n\nThere are posted datasets with jpegs and tfrecords. You can run a typical CNN with TPU for training with over half the training data in less than 3 hours.\n\nBut, this is a challenging competition. No question about it!",
      "votes": 7,
      "replies": [
        {
          "id": 1045998,
          "postDate": "2020-10-11T08:25:54.947Z",
          "content": "<p>I trained half of the tf-records but the score was less than the mean baseline even though I used only images to train the model. Do you think mean baseline (no images) will perform better than image only models with less public lb but stable cv score ?</p>",
          "rawMarkdown": "I trained half of the tf-records but the score was less than the mean baseline even though I used only images to train the model. Do you think mean baseline (no images) will perform better than image only models with less public lb but stable cv score ?"
        },
        {
          "id": 1046863,
          "postDate": "2020-10-12T04:05:00.937Z",
          "content": "<p><a href=\"https://www.kaggle.com/thakurudit\" target=\"_blank\">@thakurudit</a> Difficult to tell at this point. If your model isn't beating the baseline, I'd try new ideas and hparam tuning, but you have two final submissions, why not pick one mean baseline and one model?</p>",
          "rawMarkdown": "@thakurudit Difficult to tell at this point. If your model isn't beating the baseline, I'd try new ideas and hparam tuning, but you have two final submissions, why not pick one mean baseline and one model?",
          "votes": 1
        }
      ]
    },
    {
      "id": 1046750,
      "postDate": "2020-10-12T01:02:25.760Z",
      "content": "<p>You may also sample slices for a given study, e.g. if a study has ~1000 slice we can take every nth slice for downsampling it to ~200 slices which seems to be the average across all studies. But it looks like you already figured something out ;)</p>",
      "rawMarkdown": "You may also sample slices for a given study, e.g. if a study has ~1000 slice we can take every nth slice for downsampling it to ~200 slices which seems to be the average across all studies. But it looks like you already figured something out ;)",
      "votes": 2
    },
    {
      "id": 1045748,
      "postDate": "2020-10-11T02:03:57.927Z",
      "content": "<p>I somehow know output, and tried to use CNN.<br>\nBut this dataset contains too large image, so using all images take too much time.<br>\nI think do we need to know efficient sampling method for training?</p>\n<p>And, Some notebooks uses gbdt as tabluar data , but I think image is important when diagnosis…<br>\nIs this idea wrong?   Could you tell me your opinion..</p>",
      "rawMarkdown": "I somehow know output, and tried to use CNN.\nBut this dataset contains too large image, so using all images take too much time.\nI think do we need to know efficient sampling method for training?\n\nAnd, Some notebooks uses gbdt as tabluar data , but I think image is important when diagnosis...\nIs this idea wrong?   Could you tell me your opinion.."
    }
  ],
  "comments": [
    {
      "id": 1045762,
      "author_name": "quadcore/Richard Epstein",
      "author_url": "",
      "post_date": "2020-10-11T02:25:26.233000",
      "content": "<p>crop or resize images to make them smaller. work with 256x256 for example</p>\n<p>I don't expect tabular data by itself will be enough.</p>\n<p>You don't need to train on all the images. Make an enriched dataset with all the positive studies but not all the negative studies.</p>\n<p>There are posted datasets with jpegs and tfrecords. You can run a typical CNN with TPU for training with over half the training data in less than 3 hours.</p>\n<p>But, this is a challenging competition. No question about it!</p>",
      "votes": 7,
      "replies": [
        {
          "id": 1045998,
          "author_name": "ilovepotatoes",
          "author_url": "",
          "post_date": "2020-10-11T08:25:54.947000",
          "content": "<p>I trained half of the tf-records but the score was less than the mean baseline even though I used only images to train the model. Do you think mean baseline (no images) will perform better than image only models with less public lb but stable cv score ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1046863,
          "author_name": "Stanley Zheng",
          "author_url": "",
          "post_date": "2020-10-12T04:05:00.937000",
          "content": "<p><a href=\"https://www.kaggle.com/thakurudit\" target=\"_blank\">@thakurudit</a> Difficult to tell at this point. If your model isn't beating the baseline, I'd try new ideas and hparam tuning, but you have two final submissions, why not pick one mean baseline and one model?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1046750,
      "author_name": "Kerem Turgutlu",
      "author_url": "",
      "post_date": "2020-10-12T01:02:25.760000",
      "content": "<p>You may also sample slices for a given study, e.g. if a study has ~1000 slice we can take every nth slice for downsampling it to ~200 slices which seems to be the average across all studies. But it looks like you already figured something out ;)</p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1045762": "crop or resize images to make them smaller. work with 256x256 for example\n\nI don't expect tabular data by itself will be enough.\n\nYou don't need to train on all the images. Make an enriched dataset with all the positive studies but not all the negative studies.\n\nThere are posted datasets with jpegs and tfrecords. You can run a typical CNN with TPU for training with over half the training data in less than 3 hours.\n\nBut, this is a challenging competition. No question about it!",
    "1046750": "You may also sample slices for a given study, e.g. if a study has ~1000 slice we can take every nth slice for downsampling it to ~200 slices which seems to be the average across all studies. But it looks like you already figured something out ;)",
    "1045748": "I somehow know output, and tried to use CNN.\nBut this dataset contains too large image, so using all images take too much time.\nI think do we need to know efficient sampling method for training?\n\nAnd, Some notebooks uses gbdt as tabluar data , but I think image is important when diagnosis...\nIs this idea wrong?   Could you tell me your opinion.."
  }
}