{
  "id": 211097,
  "title": "Extracting header information efficiently from DICOM files",
  "url": "/competitions/vinbigdata-chest-xray-abnormalities-detection/discussion/211097",
  "author_name": "paulreiners",
  "post_date": "2021-01-13T16:15:38.111000",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I'm trying to assemble some of the DICOM header information into a dataset.  However, it takes quite a long time to run.</p>\n<p>Here is the relevant code:</p>\n<pre><code>all_data = pd.read_csv('../input/vinbigdata-chest-xray-abnormalities-detection/train.csv')\nsampled_data = all_data.sample(frac=0.25)\ntrain_dir = '../input/vinbigdata-chest-xray-abnormalities-detection/train'\nsexes = []\nages = []\nweights = []\nsizes = []\nfor index, row in sampled_data.iterrows():\n    image_id = row['image_id']\n    file_path = train_dir + \"/\" + image_id + '.dicom'\n    dcm = pydicom.dcmread(file_path)\n    sex = get_patients_sex(dcm)\n    sexes.append(sex)\n    age = get_patients_age(dcm)\n    ages.append(age)\n    weight = get_patients_weight(dcm)\n    weights.append(weight)\n    size = get_patients_size(dcm)\n    sizes.append(size)\nsampled_data['sex'] = sexes\nsampled_data['age'] = ages\nsampled_data['weight'] = weights\nsampled_data['size'] = sizes\n\nsampled_data.to_csv(\"sampled_data.csv\")\n</code></pre>\n<p>If I sample 1/4 of the data, this will run in an hour, which I consider acceptable.  However, if I run it on all the code, the notebook will go to sleep before finishing.</p>\n<p>Is there a way to make this code more efficient?  The complete notebook is <a href=\"https://www.kaggle.com/paulreiners/prepare-dicom-images-for-ml\" target=\"_blank\">here</a>.</p>",
  "messages": [
    {
      "id": 1152029,
      "postDate": "2021-01-13T17:57:55.320Z",
      "content": "<p><a href=\"https://www.kaggle.com/paulreiners\" target=\"_blank\">@paulreiners</a> if all you want is the dicom metadata, then you can use the <code>stop_before_pixels=True</code> option in the <code>dcmread</code> function so that it doesn't have to load all the image data into memory. So in your code where you read the file, it would become:</p>\n<pre><code>dcm = pydicom.dcmread(file_path, stop_before_pixels=True)\n</code></pre>\n<p>I do the same thing in one of my notebooks where all I wanted was the metadata:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/craigmthomas/localization-of-findings\" target=\"_blank\">https://www.kaggle.com/craigmthomas/localization-of-findings</a></li>\n</ul>\n<p>It speeds up reading the dicom files to the point where I can read all the training dicoms in about 300 seconds.</p>",
      "rawMarkdown": "@paulreiners if all you want is the dicom metadata, then you can use the `stop_before_pixels=True` option in the `dcmread` function so that it doesn't have to load all the image data into memory. So in your code where you read the file, it would become:\n\n    dcm = pydicom.dcmread(file_path, stop_before_pixels=True)\n\nI do the same thing in one of my notebooks where all I wanted was the metadata:\n\n* https://www.kaggle.com/craigmthomas/localization-of-findings\n\nIt speeds up reading the dicom files to the point where I can read all the training dicoms in about 300 seconds.",
      "votes": 5
    },
    {
      "id": 1152137,
      "postDate": "2021-01-13T21:02:38.393Z",
      "content": "<p><a href=\"https://www.kaggle.com/paulreiners\" target=\"_blank\">@paulreiners</a> also, if you want to read the dicoms in a more parallel fashion, you can use the <code>joblib</code> package to help out.  The Kaggle VMs have 2 processor cores and 4 threads, meaning that you can get 4 parallel processes working at once. I have an <code>extract</code> function that takes an <code>image_id</code> and then extracts useful information from the dicom file such as down scaling the pixel data, width, height, etc., and returns it as a tuple. Here's what it looks like:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2051313%2Fbc6ec5cb6e0874dedb25052c599c0973%2Fparallel_normalization.png?generation=1610571528559536&amp;alt=media\" alt=\"\"></p>\n<p>When complete, the variable <code>data</code> will be a list of tuples, with each tuple containing the <code>image_id</code> in position 0, a resized <code>pixel_array</code> in position 1, <code>width</code> in position 2, <code>height</code> in position 3, and other information I extracted from the dicom file. As you can see, it takes about 82 minutes to process all the dicom files. To make things easier to work with, you can convert the resultant list back into a dictionary for easier access by doing:</p>\n<pre><code>data_dict = {\n    i[0]:\n        {\n            \"pixel_array\": i[1],\n            \"width\": i[2],\n            \"height\": i[3]\n            ...\n        } for i in data\n}\n</code></pre>\n<p>Then, for example, to get the pixel data for image id <code>1234567890</code> it would be <code>data_dict[\"1234567890\"][\"pixel_array\"]</code>.  Hope this helps.</p>",
      "rawMarkdown": "@paulreiners also, if you want to read the dicoms in a more parallel fashion, you can use the `joblib` package to help out.  The Kaggle VMs have 2 processor cores and 4 threads, meaning that you can get 4 parallel processes working at once. I have an `extract` function that takes an `image_id` and then extracts useful information from the dicom file such as down scaling the pixel data, width, height, etc., and returns it as a tuple. Here's what it looks like:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2051313%2Fbc6ec5cb6e0874dedb25052c599c0973%2Fparallel_normalization.png?generation=1610571528559536&alt=media)\n\nWhen complete, the variable `data` will be a list of tuples, with each tuple containing the `image_id` in position 0, a resized `pixel_array` in position 1, `width` in position 2, `height` in position 3, and other information I extracted from the dicom file. As you can see, it takes about 82 minutes to process all the dicom files. To make things easier to work with, you can convert the resultant list back into a dictionary for easier access by doing:\n\n    data_dict = {\n        i[0]:\n            {\n                \"pixel_array\": i[1],\n                \"width\": i[2],\n                \"height\": i[3]\n                ...\n            } for i in data\n    }\n\nThen, for example, to get the pixel data for image id `1234567890` it would be `data_dict[\"1234567890\"][\"pixel_array\"]`.  Hope this helps.",
      "votes": 3
    },
    {
      "id": 1151868,
      "postDate": "2021-01-13T16:15:38.113Z",
      "content": "<p>I'm trying to assemble some of the DICOM header information into a dataset.  However, it takes quite a long time to run.</p>\n<p>Here is the relevant code:</p>\n<pre><code>all_data = pd.read_csv('../input/vinbigdata-chest-xray-abnormalities-detection/train.csv')\nsampled_data = all_data.sample(frac=0.25)\ntrain_dir = '../input/vinbigdata-chest-xray-abnormalities-detection/train'\nsexes = []\nages = []\nweights = []\nsizes = []\nfor index, row in sampled_data.iterrows():\n    image_id = row['image_id']\n    file_path = train_dir + \"/\" + image_id + '.dicom'\n    dcm = pydicom.dcmread(file_path)\n    sex = get_patients_sex(dcm)\n    sexes.append(sex)\n    age = get_patients_age(dcm)\n    ages.append(age)\n    weight = get_patients_weight(dcm)\n    weights.append(weight)\n    size = get_patients_size(dcm)\n    sizes.append(size)\nsampled_data['sex'] = sexes\nsampled_data['age'] = ages\nsampled_data['weight'] = weights\nsampled_data['size'] = sizes\n\nsampled_data.to_csv(\"sampled_data.csv\")\n</code></pre>\n<p>If I sample 1/4 of the data, this will run in an hour, which I consider acceptable.  However, if I run it on all the code, the notebook will go to sleep before finishing.</p>\n<p>Is there a way to make this code more efficient?  The complete notebook is <a href=\"https://www.kaggle.com/paulreiners/prepare-dicom-images-for-ml\" target=\"_blank\">here</a>.</p>",
      "rawMarkdown": "I'm trying to assemble some of the DICOM header information into a dataset.  However, it takes quite a long time to run.\n\nHere is the relevant code:\n\n```\nall_data = pd.read_csv('../input/vinbigdata-chest-xray-abnormalities-detection/train.csv')\nsampled_data = all_data.sample(frac=0.25)\ntrain_dir = '../input/vinbigdata-chest-xray-abnormalities-detection/train'\nsexes = []\nages = []\nweights = []\nsizes = []\nfor index, row in sampled_data.iterrows():\n    image_id = row['image_id']\n    file_path = train_dir + \"/\" + image_id + '.dicom'\n    dcm = pydicom.dcmread(file_path)\n    sex = get_patients_sex(dcm)\n    sexes.append(sex)\n    age = get_patients_age(dcm)\n    ages.append(age)\n    weight = get_patients_weight(dcm)\n    weights.append(weight)\n    size = get_patients_size(dcm)\n    sizes.append(size)\nsampled_data['sex'] = sexes\nsampled_data['age'] = ages\nsampled_data['weight'] = weights\nsampled_data['size'] = sizes\n\nsampled_data.to_csv(\"sampled_data.csv\")\n```\n\nIf I sample 1/4 of the data, this will run in an hour, which I consider acceptable.  However, if I run it on all the code, the notebook will go to sleep before finishing.\n\nIs there a way to make this code more efficient?  The complete notebook is [here](https://www.kaggle.com/paulreiners/prepare-dicom-images-for-ml).",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 1152029,
      "author_name": "Craig Thomas",
      "author_url": "",
      "post_date": "2021-01-13T17:57:55.320000",
      "content": "<p><a href=\"https://www.kaggle.com/paulreiners\" target=\"_blank\">@paulreiners</a> if all you want is the dicom metadata, then you can use the <code>stop_before_pixels=True</code> option in the <code>dcmread</code> function so that it doesn't have to load all the image data into memory. So in your code where you read the file, it would become:</p>\n<pre><code>dcm = pydicom.dcmread(file_path, stop_before_pixels=True)\n</code></pre>\n<p>I do the same thing in one of my notebooks where all I wanted was the metadata:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/craigmthomas/localization-of-findings\" target=\"_blank\">https://www.kaggle.com/craigmthomas/localization-of-findings</a></li>\n</ul>\n<p>It speeds up reading the dicom files to the point where I can read all the training dicoms in about 300 seconds.</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 1152137,
      "author_name": "Craig Thomas",
      "author_url": "",
      "post_date": "2021-01-13T21:02:38.393000",
      "content": "<p><a href=\"https://www.kaggle.com/paulreiners\" target=\"_blank\">@paulreiners</a> also, if you want to read the dicoms in a more parallel fashion, you can use the <code>joblib</code> package to help out.  The Kaggle VMs have 2 processor cores and 4 threads, meaning that you can get 4 parallel processes working at once. I have an <code>extract</code> function that takes an <code>image_id</code> and then extracts useful information from the dicom file such as down scaling the pixel data, width, height, etc., and returns it as a tuple. Here's what it looks like:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2051313%2Fbc6ec5cb6e0874dedb25052c599c0973%2Fparallel_normalization.png?generation=1610571528559536&amp;alt=media\" alt=\"\"></p>\n<p>When complete, the variable <code>data</code> will be a list of tuples, with each tuple containing the <code>image_id</code> in position 0, a resized <code>pixel_array</code> in position 1, <code>width</code> in position 2, <code>height</code> in position 3, and other information I extracted from the dicom file. As you can see, it takes about 82 minutes to process all the dicom files. To make things easier to work with, you can convert the resultant list back into a dictionary for easier access by doing:</p>\n<pre><code>data_dict = {\n    i[0]:\n        {\n            \"pixel_array\": i[1],\n            \"width\": i[2],\n            \"height\": i[3]\n            ...\n        } for i in data\n}\n</code></pre>\n<p>Then, for example, to get the pixel data for image id <code>1234567890</code> it would be <code>data_dict[\"1234567890\"][\"pixel_array\"]</code>.  Hope this helps.</p>",
      "votes": 3,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1152029": "@paulreiners if all you want is the dicom metadata, then you can use the `stop_before_pixels=True` option in the `dcmread` function so that it doesn't have to load all the image data into memory. So in your code where you read the file, it would become:\n\n    dcm = pydicom.dcmread(file_path, stop_before_pixels=True)\n\nI do the same thing in one of my notebooks where all I wanted was the metadata:\n\n* https://www.kaggle.com/craigmthomas/localization-of-findings\n\nIt speeds up reading the dicom files to the point where I can read all the training dicoms in about 300 seconds.",
    "1152137": "@paulreiners also, if you want to read the dicoms in a more parallel fashion, you can use the `joblib` package to help out.  The Kaggle VMs have 2 processor cores and 4 threads, meaning that you can get 4 parallel processes working at once. I have an `extract` function that takes an `image_id` and then extracts useful information from the dicom file such as down scaling the pixel data, width, height, etc., and returns it as a tuple. Here's what it looks like:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2051313%2Fbc6ec5cb6e0874dedb25052c599c0973%2Fparallel_normalization.png?generation=1610571528559536&alt=media)\n\nWhen complete, the variable `data` will be a list of tuples, with each tuple containing the `image_id` in position 0, a resized `pixel_array` in position 1, `width` in position 2, `height` in position 3, and other information I extracted from the dicom file. As you can see, it takes about 82 minutes to process all the dicom files. To make things easier to work with, you can convert the resultant list back into a dictionary for easier access by doing:\n\n    data_dict = {\n        i[0]:\n            {\n                \"pixel_array\": i[1],\n                \"width\": i[2],\n                \"height\": i[3]\n                ...\n            } for i in data\n    }\n\nThen, for example, to get the pixel data for image id `1234567890` it would be `data_dict[\"1234567890\"][\"pixel_array\"]`.  Hope this helps.",
    "1151868": "I'm trying to assemble some of the DICOM header information into a dataset.  However, it takes quite a long time to run.\n\nHere is the relevant code:\n\n```\nall_data = pd.read_csv('../input/vinbigdata-chest-xray-abnormalities-detection/train.csv')\nsampled_data = all_data.sample(frac=0.25)\ntrain_dir = '../input/vinbigdata-chest-xray-abnormalities-detection/train'\nsexes = []\nages = []\nweights = []\nsizes = []\nfor index, row in sampled_data.iterrows():\n    image_id = row['image_id']\n    file_path = train_dir + \"/\" + image_id + '.dicom'\n    dcm = pydicom.dcmread(file_path)\n    sex = get_patients_sex(dcm)\n    sexes.append(sex)\n    age = get_patients_age(dcm)\n    ages.append(age)\n    weight = get_patients_weight(dcm)\n    weights.append(weight)\n    size = get_patients_size(dcm)\n    sizes.append(size)\nsampled_data['sex'] = sexes\nsampled_data['age'] = ages\nsampled_data['weight'] = weights\nsampled_data['size'] = sizes\n\nsampled_data.to_csv(\"sampled_data.csv\")\n```\n\nIf I sample 1/4 of the data, this will run in an hour, which I consider acceptable.  However, if I run it on all the code, the notebook will go to sleep before finishing.\n\nIs there a way to make this code more efficient?  The complete notebook is [here](https://www.kaggle.com/paulreiners/prepare-dicom-images-for-ml)."
  }
}