{
  "id": 219906,
  "title": "Preprocessing Dicom",
  "url": "/competitions/vinbigdata-chest-xray-abnormalities-detection/discussion/219906",
  "author_name": "Julia L",
  "post_date": "2021-02-16T19:56:03.835000",
  "votes": 1,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi,<br>\nI am trying to take the pixel value of the dicom files by using pixel_array, but I encounter various problems.</p>\n<ol>\n<li>Some files don't have the pixel_array attribute =&gt; Is that the problem of the data set and we just skip those files?</li>\n<li>NotImplementedError: The pixel data with transfer syntax JPEG 2000 Image Compression (Lossless Only), cannot be read because Pillow lacks the JPEG 2000 plugin =&gt; I am not sure how to solve this problem</li>\n</ol>\n<p>Thank you for your help.</p>",
  "messages": [
    {
      "id": 1207208,
      "postDate": "2021-02-17T18:41:35.860Z",
      "content": "<p>I don't think there is really an issue with no image in some .dicom. Here's a <a href=\"https://www.kaggle.com/raddar/convert-dicom-to-np-array-the-correct-way\" target=\"_blank\">great notebook</a> by GM raddar that I used for <a href=\"https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray#7.-Creating-fast-to-read-shelve-file\" target=\"_blank\">my own data pre-processing</a> for reading the image data from the .dicom files.</p>\n<p>You could use PNG for loss-less compression.</p>",
      "rawMarkdown": "I don't think there is really an issue with no image in some .dicom. Here's a [great notebook](https://www.kaggle.com/raddar/convert-dicom-to-np-array-the-correct-way) by GM raddar that I used for [my own data pre-processing](https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray#7.-Creating-fast-to-read-shelve-file) for reading the image data from the .dicom files.\n\nYou could use PNG for loss-less compression.",
      "votes": 1,
      "replies": [
        {
          "id": 1207259,
          "postDate": "2021-02-17T19:11:17.203Z",
          "content": "<p>Thank you for your reply. I saw that notebook. But if there is no pixel_array attribute, you will have an error for calling dicom.pixel_array<br>\nAttributeError: 'FileMetaDataset' object has no attribute 'TransferSyntaxUID'<br>\nHow do you solve that if you didn't exclude them in the training set?</p>",
          "rawMarkdown": "Thank you for your reply. I saw that notebook. But if there is no pixel_array attribute, you will have an error for calling dicom.pixel_array\nAttributeError: 'FileMetaDataset' object has no attribute 'TransferSyntaxUID'\nHow do you solve that if you didn't exclude them in the training set?\n",
          "votes": 1
        },
        {
          "id": 1207336,
          "postDate": "2021-02-17T20:06:00.353Z",
          "content": "<p>Do you have a file name where that's an issue. I don't think I ran out into such an issue, but it's possible I just missed it.</p>",
          "rawMarkdown": "Do you have a file name where that's an issue. I don't think I ran out into such an issue, but it's possible I just missed it."
        },
        {
          "id": 1209005,
          "postDate": "2021-02-18T16:08:15.777Z",
          "content": "<p>There are multiple files. This one (50a418190bc3fb1ef1633bf9678929b3.dicom in the train folder) is an example. I don't know if that is the problem with my download. Could you check it for me? I've been working on this for 1 week.<br>\nThank you for your reply.</p>",
          "rawMarkdown": "There are multiple files. This one (50a418190bc3fb1ef1633bf9678929b3.dicom in the train folder) is an example. I don't know if that is the problem with my download. Could you check it for me? I've been working on this for 1 week.\nThank you for your reply.",
          "votes": 1
        },
        {
          "id": 1209394,
          "postDate": "2021-02-18T21:33:31.223Z",
          "content": "<p>Not sure what's going on for you there, perhaps it's indeed your download. If you create an notebook on the Kaggle site within the competition, you can test this:</p>\n<pre><code>import pydicom\ntrain_dir = '../input/vinbigdata-chest-xray-abnormalities-detection/train'\nfilename = '50a418190bc3fb1ef1633bf9678929b3.dicom'\ndicom = pydicom.read_file(f'{train_dir}/{filename}')\ndicom.pixel_array\n</code></pre>\n<p>and the result is </p>\n<pre><code>array([[16383, 16383, 16383, ...,    74,    66,    70],\n       [16383, 16383, 16383, ...,    84,    80,    85],\n       [16383, 16383, 16383, ...,    85,    78,    84],\n       ...,\n       [   92,    84,    81, ...,    52,    56,    56],\n       [   80,    82,    72, ...,    54,    56,    59],\n       [   89,    91,    87, ...,    62,    64,    61]], dtype=uint16)\n</code></pre>\n<p>for me.</p>",
          "rawMarkdown": "Not sure what's going on for you there, perhaps it's indeed your download. If you create an notebook on the Kaggle site within the competition, you can test this:\n```\nimport pydicom\ntrain_dir = '../input/vinbigdata-chest-xray-abnormalities-detection/train'\nfilename = '50a418190bc3fb1ef1633bf9678929b3.dicom'\ndicom = pydicom.read_file(f'{train_dir}/{filename}')\ndicom.pixel_array\n```\nand the result is \n```\narray([[16383, 16383, 16383, ...,    74,    66,    70],\n       [16383, 16383, 16383, ...,    84,    80,    85],\n       [16383, 16383, 16383, ...,    85,    78,    84],\n       ...,\n       [   92,    84,    81, ...,    52,    56,    56],\n       [   80,    82,    72, ...,    54,    56,    59],\n       [   89,    91,    87, ...,    62,    64,    61]], dtype=uint16)\n```\nfor me."
        },
        {
          "id": 1215578,
          "postDate": "2021-02-23T18:45:41.880Z",
          "content": "<p>Thank you so much. I almost gave up.<br>\nHave a nice day.</p>",
          "rawMarkdown": "Thank you so much. I almost gave up.\nHave a nice day."
        }
      ]
    },
    {
      "id": 1205596,
      "postDate": "2021-02-16T19:56:03.837Z",
      "content": "<p>Hi,<br>\nI am trying to take the pixel value of the dicom files by using pixel_array, but I encounter various problems.</p>\n<ol>\n<li>Some files don't have the pixel_array attribute =&gt; Is that the problem of the data set and we just skip those files?</li>\n<li>NotImplementedError: The pixel data with transfer syntax JPEG 2000 Image Compression (Lossless Only), cannot be read because Pillow lacks the JPEG 2000 plugin =&gt; I am not sure how to solve this problem</li>\n</ol>\n<p>Thank you for your help.</p>",
      "rawMarkdown": "Hi,\nI am trying to take the pixel value of the dicom files by using pixel_array, but I encounter various problems.\n1. Some files don't have the pixel_array attribute => Is that the problem of the data set and we just skip those files?\n2. NotImplementedError: The pixel data with transfer syntax JPEG 2000 Image Compression (Lossless Only), cannot be read because Pillow lacks the JPEG 2000 plugin => I am not sure how to solve this problem\n\nThank you for your help.\n",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1207208,
      "author_name": "Björn",
      "author_url": "",
      "post_date": "2021-02-17T18:41:35.860000",
      "content": "<p>I don't think there is really an issue with no image in some .dicom. Here's a <a href=\"https://www.kaggle.com/raddar/convert-dicom-to-np-array-the-correct-way\" target=\"_blank\">great notebook</a> by GM raddar that I used for <a href=\"https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray#7.-Creating-fast-to-read-shelve-file\" target=\"_blank\">my own data pre-processing</a> for reading the image data from the .dicom files.</p>\n<p>You could use PNG for loss-less compression.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1207259,
          "author_name": "Julia L",
          "author_url": "",
          "post_date": "2021-02-17T19:11:17.203000",
          "content": "<p>Thank you for your reply. I saw that notebook. But if there is no pixel_array attribute, you will have an error for calling dicom.pixel_array<br>\nAttributeError: 'FileMetaDataset' object has no attribute 'TransferSyntaxUID'<br>\nHow do you solve that if you didn't exclude them in the training set?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1207336,
          "author_name": "Björn",
          "author_url": "",
          "post_date": "2021-02-17T20:06:00.353000",
          "content": "<p>Do you have a file name where that's an issue. I don't think I ran out into such an issue, but it's possible I just missed it.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1209005,
          "author_name": "Julia L",
          "author_url": "",
          "post_date": "2021-02-18T16:08:15.777000",
          "content": "<p>There are multiple files. This one (50a418190bc3fb1ef1633bf9678929b3.dicom in the train folder) is an example. I don't know if that is the problem with my download. Could you check it for me? I've been working on this for 1 week.<br>\nThank you for your reply.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1209394,
          "author_name": "Björn",
          "author_url": "",
          "post_date": "2021-02-18T21:33:31.223000",
          "content": "<p>Not sure what's going on for you there, perhaps it's indeed your download. If you create an notebook on the Kaggle site within the competition, you can test this:</p>\n<pre><code>import pydicom\ntrain_dir = '../input/vinbigdata-chest-xray-abnormalities-detection/train'\nfilename = '50a418190bc3fb1ef1633bf9678929b3.dicom'\ndicom = pydicom.read_file(f'{train_dir}/{filename}')\ndicom.pixel_array\n</code></pre>\n<p>and the result is </p>\n<pre><code>array([[16383, 16383, 16383, ...,    74,    66,    70],\n       [16383, 16383, 16383, ...,    84,    80,    85],\n       [16383, 16383, 16383, ...,    85,    78,    84],\n       ...,\n       [   92,    84,    81, ...,    52,    56,    56],\n       [   80,    82,    72, ...,    54,    56,    59],\n       [   89,    91,    87, ...,    62,    64,    61]], dtype=uint16)\n</code></pre>\n<p>for me.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1215578,
          "author_name": "Julia L",
          "author_url": "",
          "post_date": "2021-02-23T18:45:41.880000",
          "content": "<p>Thank you so much. I almost gave up.<br>\nHave a nice day.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1207208": "I don't think there is really an issue with no image in some .dicom. Here's a [great notebook](https://www.kaggle.com/raddar/convert-dicom-to-np-array-the-correct-way) by GM raddar that I used for [my own data pre-processing](https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray#7.-Creating-fast-to-read-shelve-file) for reading the image data from the .dicom files.\n\nYou could use PNG for loss-less compression.",
    "1205596": "Hi,\nI am trying to take the pixel value of the dicom files by using pixel_array, but I encounter various problems.\n1. Some files don't have the pixel_array attribute => Is that the problem of the data set and we just skip those files?\n2. NotImplementedError: The pixel data with transfer syntax JPEG 2000 Image Compression (Lossless Only), cannot be read because Pillow lacks the JPEG 2000 plugin => I am not sure how to solve this problem\n\nThank you for your help.\n"
  }
}