{
  "id": 211010,
  "title": "Efficient storage of images & data loading",
  "url": "/competitions/vinbigdata-chest-xray-abnormalities-detection/discussion/211010",
  "author_name": "Björn",
  "post_date": "2021-01-13T08:38:01.444000",
  "votes": 3,
  "comment_count": 1,
  "views": 0,
  "content": "<p>One obvious issue for the speed of training is loading the images. Loading the images again and again from the <code>.dicom</code> in each epoch is clearly not a good idea. it seems like people more often save 256 by 256 images (e.g. <a href=\"https://www.kaggle.com/xhlulu/vinbigdata-chest-xray-resized-png-256x256\" target=\"_blank\">this dataset</a>) or, of course, some other dimension. I've wondered about doing something else instead, which I've not seen this done much. So, I wondered whether I'm missing much.</p>\n<p>I'm curious whether this approach will work well or whether we loose much in terms of image details:</p>\n<ol>\n<li>Load all images into memory in a loop, turn into 3 channel image</li>\n<li>Transform so that shortest dimension is 600 pixels (so e.g. 1200 by 2400 would go to 600 by 1200 - i.e. original aspect ratio is maintained), while making sure to preserve the bounding boxes using <code>albumentations</code></li>\n<li>Put the resulting uint8 (only one of the channels - as all end up identical) into a dictionary along with <code>image_id</code>, bounding boxes, labels and radiologist IDs (<code>rad_id</code>)</li>\n<li>Save all of this in a <code>shelve</code> file</li>\n</ol>\n<p>I thought <code>shelve</code>would be a good choice here in order to avoid loading all of the dictionary into memory when I just want to get one item (image) as part of a data loader and when I'm only reading, I can also access in parallel, so this should be fast. I've got a simple implementation <a href=\"https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray\" target=\"_blank\">here</a> (Section 7) with the output saved as notebook output, in case anyone wants to use this. </p>\n<p>Is there any downsides to this approach that I'm overlooking or is this a good way to speed up my data-loader?</p>\n<p><strong>Updated as second time</strong>: I've now done (notebook now public <a href=\"https://www.kaggle.com/bjoernholzhauer/vinbigdata-chest-x-ray-comparing-dataloader-speed\" target=\"_blank\">here</a>) a speed comparison using <code>%%timeit</code>(loading 8 batches of 64 images with some basic data augmentation and resizing to 224 by 224) between various approaches.</p>\n<table>\n<thead>\n<tr>\n<th>Approach</th>\n<th>Speed (per batch)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Parallelized (4 workers) loading of batches from shelve</td>\n<td>5.02 s ± 214 ms</td>\n</tr>\n<tr>\n<td>Parallelized (2 workers) loading of batches from shelve</td>\n<td>5.21 s ± 354 ms</td>\n</tr>\n<tr>\n<td>PyTorch DataLoader for images* (num_workers=4)</td>\n<td>14.2 s ± 405 ms</td>\n</tr>\n<tr>\n<td>shelve a batch at a time</td>\n<td>14.7 s ± 1.07 s</td>\n</tr>\n<tr>\n<td>PyTorch DataLoader for images* (num_workers=2)</td>\n<td>20.1 s ± 1.16 s</td>\n</tr>\n<tr>\n<td>PyTorch DataLoader for images* (num_workers=0)</td>\n<td>31.2 s ± 1.99 s</td>\n</tr>\n<tr>\n<td>PyTorch DataLoader for images* (num_workers=1)</td>\n<td>38.3 s ± 1.71 s</td>\n</tr>\n<tr>\n<td>Parallelized (4 workers) processing .dicom on-the-fly</td>\n<td>2min 46s ± 2.22 s</td>\n</tr>\n<tr>\n<td>shelve one item at a time (non-parallelized)</td>\n<td>2min 51s ± 2.61 s</td>\n</tr>\n</tbody>\n</table>\n<p>* Note: the PyTorch DataLoader used 512 by 512 images (so smaller than the ones I saved in the <code>shelve</code> and the dataset someone else prepared did actually not include corrected bounding boxes, so I skipped the augmentation of bounding boxes but that is likely negligible).</p>\n<p>Now that I'm doing more batches, the results are also more useful (with just one batch, parellelized workers did not really make sense, if they worked in a one-worker-per-batch fashion).</p>",
  "messages": [
    {
      "id": 1151318,
      "postDate": "2021-01-13T08:38:01.443Z",
      "content": "<p>One obvious issue for the speed of training is loading the images. Loading the images again and again from the <code>.dicom</code> in each epoch is clearly not a good idea. it seems like people more often save 256 by 256 images (e.g. <a href=\"https://www.kaggle.com/xhlulu/vinbigdata-chest-xray-resized-png-256x256\" target=\"_blank\">this dataset</a>) or, of course, some other dimension. I've wondered about doing something else instead, which I've not seen this done much. So, I wondered whether I'm missing much.</p>\n<p>I'm curious whether this approach will work well or whether we loose much in terms of image details:</p>\n<ol>\n<li>Load all images into memory in a loop, turn into 3 channel image</li>\n<li>Transform so that shortest dimension is 600 pixels (so e.g. 1200 by 2400 would go to 600 by 1200 - i.e. original aspect ratio is maintained), while making sure to preserve the bounding boxes using <code>albumentations</code></li>\n<li>Put the resulting uint8 (only one of the channels - as all end up identical) into a dictionary along with <code>image_id</code>, bounding boxes, labels and radiologist IDs (<code>rad_id</code>)</li>\n<li>Save all of this in a <code>shelve</code> file</li>\n</ol>\n<p>I thought <code>shelve</code>would be a good choice here in order to avoid loading all of the dictionary into memory when I just want to get one item (image) as part of a data loader and when I'm only reading, I can also access in parallel, so this should be fast. I've got a simple implementation <a href=\"https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray\" target=\"_blank\">here</a> (Section 7) with the output saved as notebook output, in case anyone wants to use this. </p>\n<p>Is there any downsides to this approach that I'm overlooking or is this a good way to speed up my data-loader?</p>\n<p><strong>Updated as second time</strong>: I've now done (notebook now public <a href=\"https://www.kaggle.com/bjoernholzhauer/vinbigdata-chest-x-ray-comparing-dataloader-speed\" target=\"_blank\">here</a>) a speed comparison using <code>%%timeit</code>(loading 8 batches of 64 images with some basic data augmentation and resizing to 224 by 224) between various approaches.</p>\n<table>\n<thead>\n<tr>\n<th>Approach</th>\n<th>Speed (per batch)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Parallelized (4 workers) loading of batches from shelve</td>\n<td>5.02 s ± 214 ms</td>\n</tr>\n<tr>\n<td>Parallelized (2 workers) loading of batches from shelve</td>\n<td>5.21 s ± 354 ms</td>\n</tr>\n<tr>\n<td>PyTorch DataLoader for images* (num_workers=4)</td>\n<td>14.2 s ± 405 ms</td>\n</tr>\n<tr>\n<td>shelve a batch at a time</td>\n<td>14.7 s ± 1.07 s</td>\n</tr>\n<tr>\n<td>PyTorch DataLoader for images* (num_workers=2)</td>\n<td>20.1 s ± 1.16 s</td>\n</tr>\n<tr>\n<td>PyTorch DataLoader for images* (num_workers=0)</td>\n<td>31.2 s ± 1.99 s</td>\n</tr>\n<tr>\n<td>PyTorch DataLoader for images* (num_workers=1)</td>\n<td>38.3 s ± 1.71 s</td>\n</tr>\n<tr>\n<td>Parallelized (4 workers) processing .dicom on-the-fly</td>\n<td>2min 46s ± 2.22 s</td>\n</tr>\n<tr>\n<td>shelve one item at a time (non-parallelized)</td>\n<td>2min 51s ± 2.61 s</td>\n</tr>\n</tbody>\n</table>\n<p>* Note: the PyTorch DataLoader used 512 by 512 images (so smaller than the ones I saved in the <code>shelve</code> and the dataset someone else prepared did actually not include corrected bounding boxes, so I skipped the augmentation of bounding boxes but that is likely negligible).</p>\n<p>Now that I'm doing more batches, the results are also more useful (with just one batch, parellelized workers did not really make sense, if they worked in a one-worker-per-batch fashion).</p>",
      "rawMarkdown": "One obvious issue for the speed of training is loading the images. Loading the images again and again from the `.dicom` in each epoch is clearly not a good idea. it seems like people more often save 256 by 256 images (e.g. [this dataset](https://www.kaggle.com/xhlulu/vinbigdata-chest-xray-resized-png-256x256)) or, of course, some other dimension. I've wondered about doing something else instead, which I've not seen this done much. So, I wondered whether I'm missing much.\n\nI'm curious whether this approach will work well or whether we loose much in terms of image details:\n1. Load all images into memory in a loop, turn into 3 channel image\n2. Transform so that shortest dimension is 600 pixels (so e.g. 1200 by 2400 would go to 600 by 1200 - i.e. original aspect ratio is maintained), while making sure to preserve the bounding boxes using `albumentations`\n3. Put the resulting uint8 (only one of the channels - as all end up identical) into a dictionary along with `image_id`, bounding boxes, labels and radiologist IDs (`rad_id`)\n4. Save all of this in a `shelve` file\n\nI thought `shelve`would be a good choice here in order to avoid loading all of the dictionary into memory when I just want to get one item (image) as part of a data loader and when I'm only reading, I can also access in parallel, so this should be fast. I've got a simple implementation [here](https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray) (Section 7) with the output saved as notebook output, in case anyone wants to use this. \n\nIs there any downsides to this approach that I'm overlooking or is this a good way to speed up my data-loader?\n\n**Updated as second time**: I've now done (notebook now public [here](https://www.kaggle.com/bjoernholzhauer/vinbigdata-chest-x-ray-comparing-dataloader-speed)) a speed comparison using `%%timeit`(loading 8 batches of 64 images with some basic data augmentation and resizing to 224 by 224) between various approaches.\n| Approach | Speed (per batch) |\n| --- | --- |\n| Parallelized (4 workers) loading of batches from shelve | 5.02 s ± 214 ms |\n| Parallelized (2 workers) loading of batches from shelve | 5.21 s ± 354 ms |\n| PyTorch DataLoader for images* (num_workers=4) | 14.2 s ± 405 ms |\n| shelve a batch at a time | 14.7 s ± 1.07 s |\n| PyTorch DataLoader for images* (num_workers=2) | 20.1 s ± 1.16 s |\n| PyTorch DataLoader for images* (num_workers=0) | 31.2 s ± 1.99 s|\n| PyTorch DataLoader for images* (num_workers=1) | 38.3 s ± 1.71 s |\n| Parallelized (4 workers) processing .dicom on-the-fly | 2min 46s ± 2.22 s |\n| shelve one item at a time (non-parallelized) | 2min 51s ± 2.61 s |\n\n\n\\* Note: the PyTorch DataLoader used 512 by 512 images (so smaller than the ones I saved in the `shelve` and the dataset someone else prepared did actually not include corrected bounding boxes, so I skipped the augmentation of bounding boxes but that is likely negligible).\n\nNow that I'm doing more batches, the results are also more useful (with just one batch, parellelized workers did not really make sense, if they worked in a one-worker-per-batch fashion).\n",
      "votes": 3
    },
    {
      "id": 1152057,
      "postDate": "2021-01-13T18:45:54.530Z",
      "content": "<p>hi <a href=\"https://www.kaggle.com/bjoernholzhauer\" target=\"_blank\">@bjoernholzhauer</a> ,</p>\n<p>Your exercise is very meticulous. Based on my experience, if multiple processor (CPU /GPU/TPU) is available, it will definitely speed up the process, but at the same time images and model weights will need that much amount of memory to be utilized simultaneously which can lead to either no memory available exception Or internally processor thread will wait for memory to allocate and deallocate which eventually can eat up even more time. I suppose monitoring processor and memory utilization while training can give better idea. See if you can relate to it in further processing. Happy Coding!</p>",
      "rawMarkdown": "hi @bjoernholzhauer ,\n\nYour exercise is very meticulous. Based on my experience, if multiple processor (CPU /GPU/TPU) is available, it will definitely speed up the process, but at the same time images and model weights will need that much amount of memory to be utilized simultaneously which can lead to either no memory available exception Or internally processor thread will wait for memory to allocate and deallocate which eventually can eat up even more time. I suppose monitoring processor and memory utilization while training can give better idea. See if you can relate to it in further processing. Happy Coding!",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1152057,
      "author_name": "supplejade",
      "author_url": "",
      "post_date": "2021-01-13T18:45:54.530000",
      "content": "<p>hi <a href=\"https://www.kaggle.com/bjoernholzhauer\" target=\"_blank\">@bjoernholzhauer</a> ,</p>\n<p>Your exercise is very meticulous. Based on my experience, if multiple processor (CPU /GPU/TPU) is available, it will definitely speed up the process, but at the same time images and model weights will need that much amount of memory to be utilized simultaneously which can lead to either no memory available exception Or internally processor thread will wait for memory to allocate and deallocate which eventually can eat up even more time. I suppose monitoring processor and memory utilization while training can give better idea. See if you can relate to it in further processing. Happy Coding!</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1151318": "One obvious issue for the speed of training is loading the images. Loading the images again and again from the `.dicom` in each epoch is clearly not a good idea. it seems like people more often save 256 by 256 images (e.g. [this dataset](https://www.kaggle.com/xhlulu/vinbigdata-chest-xray-resized-png-256x256)) or, of course, some other dimension. I've wondered about doing something else instead, which I've not seen this done much. So, I wondered whether I'm missing much.\n\nI'm curious whether this approach will work well or whether we loose much in terms of image details:\n1. Load all images into memory in a loop, turn into 3 channel image\n2. Transform so that shortest dimension is 600 pixels (so e.g. 1200 by 2400 would go to 600 by 1200 - i.e. original aspect ratio is maintained), while making sure to preserve the bounding boxes using `albumentations`\n3. Put the resulting uint8 (only one of the channels - as all end up identical) into a dictionary along with `image_id`, bounding boxes, labels and radiologist IDs (`rad_id`)\n4. Save all of this in a `shelve` file\n\nI thought `shelve`would be a good choice here in order to avoid loading all of the dictionary into memory when I just want to get one item (image) as part of a data loader and when I'm only reading, I can also access in parallel, so this should be fast. I've got a simple implementation [here](https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray) (Section 7) with the output saved as notebook output, in case anyone wants to use this. \n\nIs there any downsides to this approach that I'm overlooking or is this a good way to speed up my data-loader?\n\n**Updated as second time**: I've now done (notebook now public [here](https://www.kaggle.com/bjoernholzhauer/vinbigdata-chest-x-ray-comparing-dataloader-speed)) a speed comparison using `%%timeit`(loading 8 batches of 64 images with some basic data augmentation and resizing to 224 by 224) between various approaches.\n| Approach | Speed (per batch) |\n| --- | --- |\n| Parallelized (4 workers) loading of batches from shelve | 5.02 s ± 214 ms |\n| Parallelized (2 workers) loading of batches from shelve | 5.21 s ± 354 ms |\n| PyTorch DataLoader for images* (num_workers=4) | 14.2 s ± 405 ms |\n| shelve a batch at a time | 14.7 s ± 1.07 s |\n| PyTorch DataLoader for images* (num_workers=2) | 20.1 s ± 1.16 s |\n| PyTorch DataLoader for images* (num_workers=0) | 31.2 s ± 1.99 s|\n| PyTorch DataLoader for images* (num_workers=1) | 38.3 s ± 1.71 s |\n| Parallelized (4 workers) processing .dicom on-the-fly | 2min 46s ± 2.22 s |\n| shelve one item at a time (non-parallelized) | 2min 51s ± 2.61 s |\n\n\n\\* Note: the PyTorch DataLoader used 512 by 512 images (so smaller than the ones I saved in the `shelve` and the dataset someone else prepared did actually not include corrected bounding boxes, so I skipped the augmentation of bounding boxes but that is likely negligible).\n\nNow that I'm doing more batches, the results are also more useful (with just one batch, parellelized workers did not really make sense, if they worked in a one-worker-per-batch fashion).\n",
    "1152057": "hi @bjoernholzhauer ,\n\nYour exercise is very meticulous. Based on my experience, if multiple processor (CPU /GPU/TPU) is available, it will definitely speed up the process, but at the same time images and model weights will need that much amount of memory to be utilized simultaneously which can lead to either no memory available exception Or internally processor thread will wait for memory to allocate and deallocate which eventually can eat up even more time. I suppose monitoring processor and memory utilization while training can give better idea. See if you can relate to it in further processing. Happy Coding!"
  }
}