{
  "id": 183703,
  "title": "PE Pipeline Tensorflow and DICOM",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/183703",
  "author_name": "quadcore/Richard Epstein",
  "post_date": "2020-09-17T18:29:25.001000",
  "votes": 6,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Kagglers,</p>\n<p>I've posted a notebook [<a href=\"https://www.kaggle.com/richardepstein/pe-pipeline-create-tensorflow-with-dicom\" target=\"_blank\">https://www.kaggle.com/richardepstein/pe-pipeline-create-tensorflow-with-dicom</a>]</p>\n<p>Many of the ideas come from previous public notebooks (especially by Chris Deotte). Mainly, I've adapted them to the current competition.</p>\n<p>The Notebook creates TFRecords using the DICOM images and DICOM metadata. It also contains a section to test reading TFRecords that could be adapted into your model data loading.</p>\n<p>Remember that TFRecords, although optimized for TPUs, also work with GPUs and CPUs, So you can convert subsets of the data in the Notebooks and download the smaller TFRecords for local use.</p>\n<p>It is set up to break the data into manageable pieces so you do not have to deal with 1,790,594 files all at once. It should work with both the train and test datasets, although most of my testing was on the train data.</p>\n<p>If found to be useful, it could be used to \"crowdsource\" a public dataset that has the entire train/public test data converted (estimated at 18 nine hour Notebook runs).</p>\n<p>i do not have a full set in this format to upload.</p>\n<p>Areas for exploration include processing the DICOM metadata for a patient and adding patient level metadata to the TFRecords (number of images, normalization across different slice thicknesses). Also, best Window/Level, GrayScale vs Color, using three-channels of image for different variations of the image, image size, cropping, etc.</p>\n<p>I hope you find it useful.</p>\n<p>-Rich</p>",
  "messages": [
    {
      "id": 1014860,
      "postDate": "2020-09-17T18:29:25.003Z",
      "content": "<p>Kagglers,</p>\n<p>I've posted a notebook [<a href=\"https://www.kaggle.com/richardepstein/pe-pipeline-create-tensorflow-with-dicom\" target=\"_blank\">https://www.kaggle.com/richardepstein/pe-pipeline-create-tensorflow-with-dicom</a>]</p>\n<p>Many of the ideas come from previous public notebooks (especially by Chris Deotte). Mainly, I've adapted them to the current competition.</p>\n<p>The Notebook creates TFRecords using the DICOM images and DICOM metadata. It also contains a section to test reading TFRecords that could be adapted into your model data loading.</p>\n<p>Remember that TFRecords, although optimized for TPUs, also work with GPUs and CPUs, So you can convert subsets of the data in the Notebooks and download the smaller TFRecords for local use.</p>\n<p>It is set up to break the data into manageable pieces so you do not have to deal with 1,790,594 files all at once. It should work with both the train and test datasets, although most of my testing was on the train data.</p>\n<p>If found to be useful, it could be used to \"crowdsource\" a public dataset that has the entire train/public test data converted (estimated at 18 nine hour Notebook runs).</p>\n<p>i do not have a full set in this format to upload.</p>\n<p>Areas for exploration include processing the DICOM metadata for a patient and adding patient level metadata to the TFRecords (number of images, normalization across different slice thicknesses). Also, best Window/Level, GrayScale vs Color, using three-channels of image for different variations of the image, image size, cropping, etc.</p>\n<p>I hope you find it useful.</p>\n<p>-Rich</p>",
      "rawMarkdown": "Kagglers,\n\nI've posted a notebook [https://www.kaggle.com/richardepstein/pe-pipeline-create-tensorflow-with-dicom]\n\nMany of the ideas come from previous public notebooks (especially by Chris Deotte). Mainly, I've adapted them to the current competition.\n\nThe Notebook creates TFRecords using the DICOM images and DICOM metadata. It also contains a section to test reading TFRecords that could be adapted into your model data loading.\n\nRemember that TFRecords, although optimized for TPUs, also work with GPUs and CPUs, So you can convert subsets of the data in the Notebooks and download the smaller TFRecords for local use.\n\nIt is set up to break the data into manageable pieces so you do not have to deal with 1,790,594 files all at once. It should work with both the train and test datasets, although most of my testing was on the train data.\n\nIf found to be useful, it could be used to \"crowdsource\" a public dataset that has the entire train/public test data converted (estimated at 18 nine hour Notebook runs).\n\ni do not have a full set in this format to upload.\n\nAreas for exploration include processing the DICOM metadata for a patient and adding patient level metadata to the TFRecords (number of images, normalization across different slice thicknesses). Also, best Window/Level, GrayScale vs Color, using three-channels of image for different variations of the image, image size, cropping, etc.\n\nI hope you find it useful.\n\n-Rich",
      "votes": 6
    },
    {
      "id": 1022097,
      "postDate": "2020-09-22T10:05:34.660Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/richardepstein\" target=\"_blank\">@richardepstein</a>, thanks a lot for your work on this! I was wondering how are you planning to integrate this pipeline in the code competition format, given that the private test set is only available at the submission time? Creating TFREC files for the test images \"on the fly\" before inference seems not feasible given the time constraints…</p>",
      "rawMarkdown": "Hi @richardepstein, thanks a lot for your work on this! I was wondering how are you planning to integrate this pipeline in the code competition format, given that the private test set is only available at the submission time? Creating TFREC files for the test images \"on the fly\" before inference seems not feasible given the time constraints...",
      "replies": [
        {
          "id": 1022290,
          "postDate": "2020-09-22T12:39:26.460Z",
          "content": "<p>You bring up a valid point. During development, it seems easier to keep the same pipeline for training and inference. Eventually I might have to re-write my inference code to skip TFRecs.</p>\n<p>I can also pre-process the \"visible\" test data and only process the \"hidden\" test data in my committed notebook and combine the two. The \"hidden\" test data is about 200,000 [EDIT - exact number unclear] images in addition to the <br>\n\"visible\" test data. But the 9 hour processing time might still be a problem. </p>\n<p>Right now I am testing things in smaller batches and seeing if the score moves in the right direction.</p>",
          "rawMarkdown": "You bring up a valid point. During development, it seems easier to keep the same pipeline for training and inference. Eventually I might have to re-write my inference code to skip TFRecs.\n\nI can also pre-process the \"visible\" test data and only process the \"hidden\" test data in my committed notebook and combine the two. The \"hidden\" test data is about 200,000 [EDIT - exact number unclear] images in addition to the \n\"visible\" test data. But the 9 hour processing time might still be a problem. \n\nRight now I am testing things in smaller batches and seeing if the score moves in the right direction."
        }
      ]
    },
    {
      "id": 1015858,
      "postDate": "2020-09-18T13:42:10.630Z",
      "content": "<p>I found an error where batches that didn't start at zero were failing. Version 5 should fix that. </p>\n<p>Also added GDCM library to avoid jpeg decompression errors.</p>\n<p>I think it runs cleanly now, but takes a long time to test large batches.</p>\n<p>Thanks for the testing.</p>\n<p>The first rule of programming is your program never does what you think it does.</p>",
      "rawMarkdown": "I found an error where batches that didn't start at zero were failing. Version 5 should fix that. \n\nAlso added GDCM library to avoid jpeg decompression errors.\n\nI think it runs cleanly now, but takes a long time to test large batches.\n\nThanks for the testing.\n\nThe first rule of programming is your program never does what you think it does."
    },
    {
      "id": 1015219,
      "postDate": "2020-09-18T03:46:35.870Z",
      "content": "<p>100,000 images takes about 9 hours. So I would expect 20,000 records to take almost 2 hours. If you add your TFRecords to a public notebook I can take a look.</p>\n<p>100,000 images might run over the limit of 9 hours, losing all the work. So maybe 50,000 batches would be most efficient.</p>\n<p>I'll set off a test and see how long a sample takes.</p>",
      "rawMarkdown": "100,000 images takes about 9 hours. So I would expect 20,000 records to take almost 2 hours. If you add your TFRecords to a public notebook I can take a look.\n\n100,000 images might run over the limit of 9 hours, losing all the work. So maybe 50,000 batches would be most efficient.\n\nI'll set off a test and see how long a sample takes.",
      "replies": [
        {
          "id": 1015766,
          "postDate": "2020-09-18T12:04:51.777Z",
          "content": "<p>Sure. You can see it took less than 200 seconds to run.</p>\n<p><a href=\"https://www.kaggle.com/returnofsputnik/dicom-100-700-pt1/\" target=\"_blank\">https://www.kaggle.com/returnofsputnik/dicom-100-700-pt1/</a></p>",
          "rawMarkdown": "Sure. You can see it took less than 200 seconds to run.\n\nhttps://www.kaggle.com/returnofsputnik/dicom-100-700-pt1/"
        },
        {
          "id": 1015791,
          "postDate": "2020-09-18T12:27:05.517Z",
          "content": "<p>i have 10 kernels running now with df_size=50000 each, going all the way up to start=500,000. We'll complete this eventually</p>",
          "rawMarkdown": "i have 10 kernels running now with df_size=50000 each, going all the way up to start=500,000. We'll complete this eventually"
        },
        {
          "id": 1015793,
          "postDate": "2020-09-18T12:29:22.523Z",
          "content": "<p>Actually all of them seem to fail for me.. xD <a href=\"https://www.kaggle.com/returnofsputnik/pe-pipeline-create-tensorflow-with-dicom\" target=\"_blank\">https://www.kaggle.com/returnofsputnik/pe-pipeline-create-tensorflow-with-dicom</a></p>",
          "rawMarkdown": "Actually all of them seem to fail for me.. xD https://www.kaggle.com/returnofsputnik/pe-pipeline-create-tensorflow-with-dicom"
        }
      ]
    },
    {
      "id": 1015184,
      "postDate": "2020-09-18T03:02:25.867Z",
      "content": "<p><a href=\"https://www.kaggle.com/richardepstein\" target=\"_blank\">@richardepstein</a> I am willing to run some notebooks to help create a full TFRecords dataset. However I ran your kernel with df_size=20,000 and it finished very quickly, not the 9 hour you were talking about. Maybe if you can tell me which parts to change and I can kick off some scripts? I probably doing something incorrect.</p>",
      "rawMarkdown": "@richardepstein I am willing to run some notebooks to help create a full TFRecords dataset. However I ran your kernel with df_size=20,000 and it finished very quickly, not the 9 hour you were talking about. Maybe if you can tell me which parts to change and I can kick off some scripts? I probably doing something incorrect."
    },
    {
      "id": 1014908,
      "postDate": "2020-09-17T19:27:57.677Z",
      "content": "<blockquote>\n  <p>estimated at 18 nine hour Notebook runs</p>\n</blockquote>\n<p>Yah, I've noticed the incredible challenge of this competition, the size of the dataset is incredible. I'm almost done creating TFRecords for training 128x128. I estimate about 12-18 nine hour notebook runs myself.</p>",
      "rawMarkdown": "> estimated at 18 nine hour Notebook runs\n\nYah, I've noticed the incredible challenge of this competition, the size of the dataset is incredible. I'm almost done creating TFRecords for training 128x128. I estimate about 12-18 nine hour notebook runs myself."
    }
  ],
  "comments": [
    {
      "id": 1022097,
      "author_name": "Nikita Kozodoi",
      "author_url": "",
      "post_date": "2020-09-22T10:05:34.660000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/richardepstein\" target=\"_blank\">@richardepstein</a>, thanks a lot for your work on this! I was wondering how are you planning to integrate this pipeline in the code competition format, given that the private test set is only available at the submission time? Creating TFREC files for the test images \"on the fly\" before inference seems not feasible given the time constraints…</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1022290,
          "author_name": "quadcore/Richard Epstein",
          "author_url": "",
          "post_date": "2020-09-22T12:39:26.460000",
          "content": "<p>You bring up a valid point. During development, it seems easier to keep the same pipeline for training and inference. Eventually I might have to re-write my inference code to skip TFRecs.</p>\n<p>I can also pre-process the \"visible\" test data and only process the \"hidden\" test data in my committed notebook and combine the two. The \"hidden\" test data is about 200,000 [EDIT - exact number unclear] images in addition to the <br>\n\"visible\" test data. But the 9 hour processing time might still be a problem. </p>\n<p>Right now I am testing things in smaller batches and seeing if the score moves in the right direction.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1015858,
      "author_name": "quadcore/Richard Epstein",
      "author_url": "",
      "post_date": "2020-09-18T13:42:10.630000",
      "content": "<p>I found an error where batches that didn't start at zero were failing. Version 5 should fix that. </p>\n<p>Also added GDCM library to avoid jpeg decompression errors.</p>\n<p>I think it runs cleanly now, but takes a long time to test large batches.</p>\n<p>Thanks for the testing.</p>\n<p>The first rule of programming is your program never does what you think it does.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1015219,
      "author_name": "quadcore/Richard Epstein",
      "author_url": "",
      "post_date": "2020-09-18T03:46:35.870000",
      "content": "<p>100,000 images takes about 9 hours. So I would expect 20,000 records to take almost 2 hours. If you add your TFRecords to a public notebook I can take a look.</p>\n<p>100,000 images might run over the limit of 9 hours, losing all the work. So maybe 50,000 batches would be most efficient.</p>\n<p>I'll set off a test and see how long a sample takes.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1015766,
          "author_name": "CoreyJamesLevinson",
          "author_url": "",
          "post_date": "2020-09-18T12:04:51.777000",
          "content": "<p>Sure. You can see it took less than 200 seconds to run.</p>\n<p><a href=\"https://www.kaggle.com/returnofsputnik/dicom-100-700-pt1/\" target=\"_blank\">https://www.kaggle.com/returnofsputnik/dicom-100-700-pt1/</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1015791,
          "author_name": "CoreyJamesLevinson",
          "author_url": "",
          "post_date": "2020-09-18T12:27:05.517000",
          "content": "<p>i have 10 kernels running now with df_size=50000 each, going all the way up to start=500,000. We'll complete this eventually</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1015793,
          "author_name": "CoreyJamesLevinson",
          "author_url": "",
          "post_date": "2020-09-18T12:29:22.523000",
          "content": "<p>Actually all of them seem to fail for me.. xD <a href=\"https://www.kaggle.com/returnofsputnik/pe-pipeline-create-tensorflow-with-dicom\" target=\"_blank\">https://www.kaggle.com/returnofsputnik/pe-pipeline-create-tensorflow-with-dicom</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1015184,
      "author_name": "CoreyJamesLevinson",
      "author_url": "",
      "post_date": "2020-09-18T03:02:25.867000",
      "content": "<p><a href=\"https://www.kaggle.com/richardepstein\" target=\"_blank\">@richardepstein</a> I am willing to run some notebooks to help create a full TFRecords dataset. However I ran your kernel with df_size=20,000 and it finished very quickly, not the 9 hour you were talking about. Maybe if you can tell me which parts to change and I can kick off some scripts? I probably doing something incorrect.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1014908,
      "author_name": "Tim Yee",
      "author_url": "",
      "post_date": "2020-09-17T19:27:57.677000",
      "content": "<blockquote>\n  <p>estimated at 18 nine hour Notebook runs</p>\n</blockquote>\n<p>Yah, I've noticed the incredible challenge of this competition, the size of the dataset is incredible. I'm almost done creating TFRecords for training 128x128. I estimate about 12-18 nine hour notebook runs myself.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1014860": "Kagglers,\n\nI've posted a notebook [https://www.kaggle.com/richardepstein/pe-pipeline-create-tensorflow-with-dicom]\n\nMany of the ideas come from previous public notebooks (especially by Chris Deotte). Mainly, I've adapted them to the current competition.\n\nThe Notebook creates TFRecords using the DICOM images and DICOM metadata. It also contains a section to test reading TFRecords that could be adapted into your model data loading.\n\nRemember that TFRecords, although optimized for TPUs, also work with GPUs and CPUs, So you can convert subsets of the data in the Notebooks and download the smaller TFRecords for local use.\n\nIt is set up to break the data into manageable pieces so you do not have to deal with 1,790,594 files all at once. It should work with both the train and test datasets, although most of my testing was on the train data.\n\nIf found to be useful, it could be used to \"crowdsource\" a public dataset that has the entire train/public test data converted (estimated at 18 nine hour Notebook runs).\n\ni do not have a full set in this format to upload.\n\nAreas for exploration include processing the DICOM metadata for a patient and adding patient level metadata to the TFRecords (number of images, normalization across different slice thicknesses). Also, best Window/Level, GrayScale vs Color, using three-channels of image for different variations of the image, image size, cropping, etc.\n\nI hope you find it useful.\n\n-Rich",
    "1022097": "Hi @richardepstein, thanks a lot for your work on this! I was wondering how are you planning to integrate this pipeline in the code competition format, given that the private test set is only available at the submission time? Creating TFREC files for the test images \"on the fly\" before inference seems not feasible given the time constraints...",
    "1015858": "I found an error where batches that didn't start at zero were failing. Version 5 should fix that. \n\nAlso added GDCM library to avoid jpeg decompression errors.\n\nI think it runs cleanly now, but takes a long time to test large batches.\n\nThanks for the testing.\n\nThe first rule of programming is your program never does what you think it does.",
    "1015219": "100,000 images takes about 9 hours. So I would expect 20,000 records to take almost 2 hours. If you add your TFRecords to a public notebook I can take a look.\n\n100,000 images might run over the limit of 9 hours, losing all the work. So maybe 50,000 batches would be most efficient.\n\nI'll set off a test and see how long a sample takes.",
    "1015184": "@richardepstein I am willing to run some notebooks to help create a full TFRecords dataset. However I ran your kernel with df_size=20,000 and it finished very quickly, not the 9 hour you were talking about. Maybe if you can tell me which parts to change and I can kick off some scripts? I probably doing something incorrect.",
    "1014908": "> estimated at 18 nine hour Notebook runs\n\nYah, I've noticed the incredible challenge of this competition, the size of the dataset is incredible. I'm almost done creating TFRecords for training 128x128. I estimate about 12-18 nine hour notebook runs myself."
  }
}