{
  "id": 185991,
  "title": "Stratified Validation Strategy Starter and Ideas to Deal with Huge Dataset",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/185991",
  "author_name": "khyeh",
  "post_date": "2020-09-22T19:30:03.956000",
  "votes": 29,
  "comment_count": 8,
  "views": 0,
  "content": "<h4>Stratified Validation Strategy Starter</h4>\n<p>Doing proper validation is important, so <strong>we could trust our cv scores without bombarding the public leaderboard too often and get overfitted to it</strong>. The spirit of doing validation is to simulate how the train-test is split. In this competition, train and test sets are split by different patients, so it might be better to split cv set by patients as well (same patients should be kept in the same validation set)</p>\n<p>There are actually more possible ways to do the validation splits. Considering the number of images per patient, it could be further considered as a way to do the stratified cv splits. I published a kernel showing how to do 20 folds patient-level + stratified validation splits based on the number of image-per-patient. (The motivation is further described in the bottom)</p>\n<blockquote>\n  <p><a href=\"https://www.kaggle.com/khyeh0719/stratified-validation-strategy\" target=\"_blank\">https://www.kaggle.com/khyeh0719/stratified-validation-strategy</a></p>\n</blockquote>\n<p>Here is how each fold looks like in the distribution of numbers of images per patient</p>\n<p><img src=\"https://i.imgur.com/SfV6Riu.jpg\" alt=\"\"></p>\n<h4>Further usage of this kernel:</h4>\n<ul>\n<li>You could use the kernel output csv file directly to do patient-level subsampling (ex. select fold=1-5 to do 5 fold cross-validation)</li>\n<li>You could modify FOLD_NUM in the kernel to create a different number of stratified folds yourself</li>\n<li>You could modify bin_counts+digitize_cols in the kernel to digitize columns with designated bin counts, which will be further incorporated into the new \"key\" to do the stratified validation splits</li>\n</ul>\n<h4>Ideas to Deal with Huge Dataset</h4>\n<p><strong>1. Patient-Level Subsampling:</strong> In my kernel, I do stratified splits of patients into 20 groups, each group is consists of ~360 patients. To do patient-level subsampling, we could pick patients with fold=1-5 to do stratified 5-fold cross-validation. In each fold, ~360*4=1440 patients are in the train set, ~360 patients are in the validation set, there would be no leakage between train and validation set still (patients will be in the either train or valid, not in both)</p>\n<p><strong>2. Image-Level Subsampling:</strong> We could train our model faster by loading uniformly downsampled images from a single patient. For example, for a patient, we always sample 100 images to feed into the NN. <em>In this way, the model would always see 100 images from a single patient, without knowing how many images there were in the raw data</em>. In order to make sure NN could generalize well on validation/test set when we do image-level subsampling, it is better to make sure the distribution we are sampling from is similar between train, validation, and test set. As shown in the kernel, train-test distributions are similar already, we only need to make sure distributions of train-validation are consistent then. This is also the motivation that I would like to make stratified validation splits based on the number of images per patient.</p>\n<p>Hope you find the kernel and this post helpful as a starter :)</p>",
  "messages": [
    {
      "id": 1022876,
      "postDate": "2020-09-22T19:30:03.957Z",
      "content": "<h4>Stratified Validation Strategy Starter</h4>\n<p>Doing proper validation is important, so <strong>we could trust our cv scores without bombarding the public leaderboard too often and get overfitted to it</strong>. The spirit of doing validation is to simulate how the train-test is split. In this competition, train and test sets are split by different patients, so it might be better to split cv set by patients as well (same patients should be kept in the same validation set)</p>\n<p>There are actually more possible ways to do the validation splits. Considering the number of images per patient, it could be further considered as a way to do the stratified cv splits. I published a kernel showing how to do 20 folds patient-level + stratified validation splits based on the number of image-per-patient. (The motivation is further described in the bottom)</p>\n<blockquote>\n  <p><a href=\"https://www.kaggle.com/khyeh0719/stratified-validation-strategy\" target=\"_blank\">https://www.kaggle.com/khyeh0719/stratified-validation-strategy</a></p>\n</blockquote>\n<p>Here is how each fold looks like in the distribution of numbers of images per patient</p>\n<p><img src=\"https://i.imgur.com/SfV6Riu.jpg\" alt=\"\"></p>\n<h4>Further usage of this kernel:</h4>\n<ul>\n<li>You could use the kernel output csv file directly to do patient-level subsampling (ex. select fold=1-5 to do 5 fold cross-validation)</li>\n<li>You could modify FOLD_NUM in the kernel to create a different number of stratified folds yourself</li>\n<li>You could modify bin_counts+digitize_cols in the kernel to digitize columns with designated bin counts, which will be further incorporated into the new \"key\" to do the stratified validation splits</li>\n</ul>\n<h4>Ideas to Deal with Huge Dataset</h4>\n<p><strong>1. Patient-Level Subsampling:</strong> In my kernel, I do stratified splits of patients into 20 groups, each group is consists of ~360 patients. To do patient-level subsampling, we could pick patients with fold=1-5 to do stratified 5-fold cross-validation. In each fold, ~360*4=1440 patients are in the train set, ~360 patients are in the validation set, there would be no leakage between train and validation set still (patients will be in the either train or valid, not in both)</p>\n<p><strong>2. Image-Level Subsampling:</strong> We could train our model faster by loading uniformly downsampled images from a single patient. For example, for a patient, we always sample 100 images to feed into the NN. <em>In this way, the model would always see 100 images from a single patient, without knowing how many images there were in the raw data</em>. In order to make sure NN could generalize well on validation/test set when we do image-level subsampling, it is better to make sure the distribution we are sampling from is similar between train, validation, and test set. As shown in the kernel, train-test distributions are similar already, we only need to make sure distributions of train-validation are consistent then. This is also the motivation that I would like to make stratified validation splits based on the number of images per patient.</p>\n<p>Hope you find the kernel and this post helpful as a starter :)</p>",
      "rawMarkdown": "#### Stratified Validation Strategy Starter\n\nDoing proper validation is important, so **we could trust our cv scores without bombarding the public leaderboard too often and get overfitted to it**. The spirit of doing validation is to simulate how the train-test is split. In this competition, train and test sets are split by different patients, so it might be better to split cv set by patients as well (same patients should be kept in the same validation set)\n\nThere are actually more possible ways to do the validation splits. Considering the number of images per patient, it could be further considered as a way to do the stratified cv splits. I published a kernel showing how to do 20 folds patient-level + stratified validation splits based on the number of image-per-patient. (The motivation is further described in the bottom)\n\n> https://www.kaggle.com/khyeh0719/stratified-validation-strategy\n\nHere is how each fold looks like in the distribution of numbers of images per patient\n\n![](https://i.imgur.com/SfV6Riu.jpg)\n\n#### Further usage of this kernel:\n\n- You could use the kernel output csv file directly to do patient-level subsampling (ex. select fold=1-5 to do 5 fold cross-validation)\n- You could modify FOLD_NUM in the kernel to create a different number of stratified folds yourself\n- You could modify bin_counts+digitize_cols in the kernel to digitize columns with designated bin counts, which will be further incorporated into the new \"key\" to do the stratified validation splits\n\n\n#### Ideas to Deal with Huge Dataset\n\n**1. Patient-Level Subsampling:** In my kernel, I do stratified splits of patients into 20 groups, each group is consists of ~360 patients. To do patient-level subsampling, we could pick patients with fold=1-5 to do stratified 5-fold cross-validation. In each fold, ~360*4=1440 patients are in the train set, ~360 patients are in the validation set, there would be no leakage between train and validation set still (patients will be in the either train or valid, not in both)\n\n**2. Image-Level Subsampling:** We could train our model faster by loading uniformly downsampled images from a single patient. For example, for a patient, we always sample 100 images to feed into the NN. *In this way, the model would always see 100 images from a single patient, without knowing how many images there were in the raw data*. In order to make sure NN could generalize well on validation/test set when we do image-level subsampling, it is better to make sure the distribution we are sampling from is similar between train, validation, and test set. As shown in the kernel, train-test distributions are similar already, we only need to make sure distributions of train-validation are consistent then. This is also the motivation that I would like to make stratified validation splits based on the number of images per patient.\n\nHope you find the kernel and this post helpful as a starter :)",
      "votes": 29
    },
    {
      "id": 1027992,
      "postDate": "2020-09-26T14:07:00.273Z",
      "content": "<p><a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a>     Can you please explain me the \"key\" column . I didnt quite get it :(</p>",
      "rawMarkdown": "@khyeh0719     Can you please explain me the \"key\" column . I didnt quite get it :(",
      "replies": [
        {
          "id": 1027997,
          "postDate": "2020-09-26T14:13:40.687Z",
          "content": "<p>In the kernel, I did some kind of numerical quantization for the \"image_num_per_patient\" and treat the values as a categorical variable (which column name is \"key\" in my kernel). Then do stratified kfold based on that column. Hope it helps you understand my sharing in both post and kernel better :)</p>",
          "rawMarkdown": "In the kernel, I did some kind of numerical quantization for the \"image_num_per_patient\" and treat the values as a categorical variable (which column name is \"key\" in my kernel). Then do stratified kfold based on that column. Hope it helps you understand my sharing in both post and kernel better :)",
          "votes": 1
        },
        {
          "id": 1028001,
          "postDate": "2020-09-26T14:18:16.847Z",
          "content": "<p>Perfect . Now I am able to understand it .</p>",
          "rawMarkdown": "Perfect . Now I am able to understand it .",
          "votes": 1
        },
        {
          "id": 1028008,
          "postDate": "2020-09-26T14:23:58.723Z",
          "content": "<p>good to know :)</p>",
          "rawMarkdown": "good to know :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 1024074,
      "postDate": "2020-09-23T16:21:28.323Z",
      "content": "<p>Hello Kun Hao Yeh, thank for the validation idea and for sharing the kernel.</p>\n<p>One clarification: When you say patient, you mean the <code>SeriesInstanceUID</code> right?</p>\n<p>Given the <code>StudyInstanceUID</code> and <code>SeriesInstanceUID</code> are unique and the 3D images are inside the SeriesInstanceUID folder I thought this was the patient and the uniqueness of the Study - Seriesimplied no patient participated in more than one Study. Is that right?</p>\n<p>Thank you!</p>",
      "rawMarkdown": "Hello Kun Hao Yeh, thank for the validation idea and for sharing the kernel.\n\nOne clarification: When you say patient, you mean the `SeriesInstanceUID` right?\n\nGiven the `StudyInstanceUID` and `SeriesInstanceUID` are unique and the 3D images are inside the SeriesInstanceUID folder I thought this was the patient and the uniqueness of the Study - Seriesimplied no patient participated in more than one Study. Is that right?\n\nThank you!",
      "replies": [
        {
          "id": 1024677,
          "postDate": "2020-09-24T04:18:08.287Z",
          "content": "<p>I actually means StudyInstanceUID, but given StudyInstanceUID and SeriesInstanceUID is 1 to 1 mapping, I guess they are both ok to represent a patient.</p>",
          "rawMarkdown": "I actually means StudyInstanceUID, but given StudyInstanceUID and SeriesInstanceUID is 1 to 1 mapping, I guess they are both ok to represent a patient.",
          "votes": 2
        },
        {
          "id": 1024694,
          "postDate": "2020-09-24T04:30:32.187Z",
          "content": "<p>I thought so too. Thanks!</p>",
          "rawMarkdown": "I thought so too. Thanks!"
        },
        {
          "id": 1024781,
          "postDate": "2020-09-24T06:03:32.307Z",
          "content": "<p>You're welcome :)</p>",
          "rawMarkdown": "You're welcome :)",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1027992,
      "author_name": "Nirjhar Roy",
      "author_url": "",
      "post_date": "2020-09-26T14:07:00.273000",
      "content": "<p><a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a>     Can you please explain me the \"key\" column . I didnt quite get it :(</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1027997,
          "author_name": "khyeh",
          "author_url": "",
          "post_date": "2020-09-26T14:13:40.687000",
          "content": "<p>In the kernel, I did some kind of numerical quantization for the \"image_num_per_patient\" and treat the values as a categorical variable (which column name is \"key\" in my kernel). Then do stratified kfold based on that column. Hope it helps you understand my sharing in both post and kernel better :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1028001,
          "author_name": "Nirjhar Roy",
          "author_url": "",
          "post_date": "2020-09-26T14:18:16.847000",
          "content": "<p>Perfect . Now I am able to understand it .</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1028008,
          "author_name": "khyeh",
          "author_url": "",
          "post_date": "2020-09-26T14:23:58.723000",
          "content": "<p>good to know :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1024074,
      "author_name": "Ronaldo S.A. Batista",
      "author_url": "",
      "post_date": "2020-09-23T16:21:28.323000",
      "content": "<p>Hello Kun Hao Yeh, thank for the validation idea and for sharing the kernel.</p>\n<p>One clarification: When you say patient, you mean the <code>SeriesInstanceUID</code> right?</p>\n<p>Given the <code>StudyInstanceUID</code> and <code>SeriesInstanceUID</code> are unique and the 3D images are inside the SeriesInstanceUID folder I thought this was the patient and the uniqueness of the Study - Seriesimplied no patient participated in more than one Study. Is that right?</p>\n<p>Thank you!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1024677,
          "author_name": "khyeh",
          "author_url": "",
          "post_date": "2020-09-24T04:18:08.287000",
          "content": "<p>I actually means StudyInstanceUID, but given StudyInstanceUID and SeriesInstanceUID is 1 to 1 mapping, I guess they are both ok to represent a patient.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1024694,
          "author_name": "Ronaldo S.A. Batista",
          "author_url": "",
          "post_date": "2020-09-24T04:30:32.187000",
          "content": "<p>I thought so too. Thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1024781,
          "author_name": "khyeh",
          "author_url": "",
          "post_date": "2020-09-24T06:03:32.307000",
          "content": "<p>You're welcome :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1022876": "#### Stratified Validation Strategy Starter\n\nDoing proper validation is important, so **we could trust our cv scores without bombarding the public leaderboard too often and get overfitted to it**. The spirit of doing validation is to simulate how the train-test is split. In this competition, train and test sets are split by different patients, so it might be better to split cv set by patients as well (same patients should be kept in the same validation set)\n\nThere are actually more possible ways to do the validation splits. Considering the number of images per patient, it could be further considered as a way to do the stratified cv splits. I published a kernel showing how to do 20 folds patient-level + stratified validation splits based on the number of image-per-patient. (The motivation is further described in the bottom)\n\n> https://www.kaggle.com/khyeh0719/stratified-validation-strategy\n\nHere is how each fold looks like in the distribution of numbers of images per patient\n\n![](https://i.imgur.com/SfV6Riu.jpg)\n\n#### Further usage of this kernel:\n\n- You could use the kernel output csv file directly to do patient-level subsampling (ex. select fold=1-5 to do 5 fold cross-validation)\n- You could modify FOLD_NUM in the kernel to create a different number of stratified folds yourself\n- You could modify bin_counts+digitize_cols in the kernel to digitize columns with designated bin counts, which will be further incorporated into the new \"key\" to do the stratified validation splits\n\n\n#### Ideas to Deal with Huge Dataset\n\n**1. Patient-Level Subsampling:** In my kernel, I do stratified splits of patients into 20 groups, each group is consists of ~360 patients. To do patient-level subsampling, we could pick patients with fold=1-5 to do stratified 5-fold cross-validation. In each fold, ~360*4=1440 patients are in the train set, ~360 patients are in the validation set, there would be no leakage between train and validation set still (patients will be in the either train or valid, not in both)\n\n**2. Image-Level Subsampling:** We could train our model faster by loading uniformly downsampled images from a single patient. For example, for a patient, we always sample 100 images to feed into the NN. *In this way, the model would always see 100 images from a single patient, without knowing how many images there were in the raw data*. In order to make sure NN could generalize well on validation/test set when we do image-level subsampling, it is better to make sure the distribution we are sampling from is similar between train, validation, and test set. As shown in the kernel, train-test distributions are similar already, we only need to make sure distributions of train-validation are consistent then. This is also the motivation that I would like to make stratified validation splits based on the number of images per patient.\n\nHope you find the kernel and this post helpful as a starter :)",
    "1027992": "@khyeh0719     Can you please explain me the \"key\" column . I didnt quite get it :(",
    "1024074": "Hello Kun Hao Yeh, thank for the validation idea and for sharing the kernel.\n\nOne clarification: When you say patient, you mean the `SeriesInstanceUID` right?\n\nGiven the `StudyInstanceUID` and `SeriesInstanceUID` are unique and the 3D images are inside the SeriesInstanceUID folder I thought this was the patient and the uniqueness of the Study - Seriesimplied no patient participated in more than one Study. Is that right?\n\nThank you!"
  }
}