{
  "id": 609292,
  "title": "RSNA Competition: Dataset Loading Taking 2-8 Seconds Per Sample - Optimization Advice Needed",
  "url": "/competitions/rsna-intracranial-aneurysm-detection/discussion/609292",
  "author_name": "Ryuichi-create",
  "post_date": "2025-09-25T16:01:43.985000",
  "votes": 6,
  "comment_count": 4,
  "views": 0,
  "content": "<h3><strong>Problem</strong></h3>\n<p>My <code>RSNADataset</code> is causing severe training bottlenecks due to real-time DICOM processing in Kaggle GPU environment:</p>\n<pre><code> ():\n     ():\n        row = .df.iloc[idx]\n        series_path = \n        volume = process_dicom_series_safe(series_path, target_shape=(, , ))\n        x = torch.from_numpy(volume).()\n         x, y\n</code></pre>\n<h3><strong>Performance Issues</strong></h3>\n<ul>\n<li><strong>Per sample:</strong> 2-8 seconds </li>\n<li><strong>Batch size 8:</strong> 16-64 seconds per batch</li>\n<li><strong>Dataset scale:</strong> 4,348 series × 200-900 DICOM files each</li>\n<li><strong>Training time:</strong> Would take 3-9 hours just for data loading per epoch</li>\n<li><strong>Environment:</strong> Kaggle notebook with Tesla T4 GPU (16GB)<br>\n<strong>Bottleneck breakdown (tested on Kaggle):</strong></li>\n</ul>\n<ol>\n<li><strong>File I/O:</strong> 0.5-2.0s (loading hundreds of DICOM files per series)</li>\n<li><strong>DICOM parsing:</strong> 0.2-0.8s (pydicom processing with <code>pydicom.dcmread()</code>)</li>\n<li><strong>Image processing:</strong> 0.3-1.0s (statistical normalization using <code>np.percentile()</code>, resizing to 384×384)</li>\n<li><strong>3D volume creation:</strong> 0.1-0.3s (stacking slices with <code>scipy.ndimage.zoom</code>)</li>\n</ol>\n<h3><strong>Current DICOM Processing Pipeline</strong></h3>\n<p>My <code>DICOMPreprocessorKaggle</code> class handles:</p>\n<ul>\n<li>Loading 200-900 <code>.dcm</code> files per series with <code>os.walk()</code></li>\n<li>Statistical normalization: <code>p1, p99 = np.percentile(img, [1, 99])</code></li>\n<li>2D resizing: <code>cv2.resize(processed_img, (384, 384))</code></li>\n<li>3D volume creation: <code>ndimage.zoom()</code> for final (32, 384, 384) shape</li>\n<li>Memory cleanup with <code>gc.collect()</code></li>\n</ul>\n<h3><strong>Questions</strong></h3>\n<ol>\n<li><strong>Pre-processing approach:</strong> Should I convert all DICOM series to <code>.npy</code> files beforehand to avoid real-time processing?</li>\n<li><strong>Kaggle GPU optimization:</strong> Any specific strategies for Tesla T4 environment with limited session time?</li>\n<li><strong>Alternative libraries:</strong> Faster alternatives to <code>pydicom</code> + <code>scipy.ndimage</code> combination?</li>\n<li><strong>DataLoader settings:</strong> Best practices for <code>num_workers</code> in Kaggle? (Currently using 0 due to multiprocessing issues)</li>\n</ol>\n<h3><strong>What I've Tried</strong></h3>\n<ul>\n<li><code>num_workers=0</code> (multiprocessing causes issues in Kaggle environment)</li>\n<li>Small batch sizes (8) for GPU memory management</li>\n<li><code>torch.cuda.empty_cache()</code> and memory cleanup</li>\n<li>Mixed precision training to optimize GPU usage</li>\n</ul>\n<h3><strong>Target Goal</strong></h3>\n<p>Looking for practical solutions to get batch loading under 1 second in Kaggle GPU environment. Current GPU utilization is poor due to waiting for data loading.<br>\nAny medical imaging competition veterans with experience optimizing DICOM workflows on Kaggle GPU instances?<br>\n<strong>Tags:</strong> #rsna2024 #dicom #pytorch #dataloader #performance #kaggle-gpu</p>",
  "messages": [
    {
      "id": 3294219,
      "postDate": "2025-09-25T16:01:43.987Z",
      "content": "<h3><strong>Problem</strong></h3>\n<p>My <code>RSNADataset</code> is causing severe training bottlenecks due to real-time DICOM processing in Kaggle GPU environment:</p>\n<pre><code> ():\n     ():\n        row = .df.iloc[idx]\n        series_path = \n        volume = process_dicom_series_safe(series_path, target_shape=(, , ))\n        x = torch.from_numpy(volume).()\n         x, y\n</code></pre>\n<h3><strong>Performance Issues</strong></h3>\n<ul>\n<li><strong>Per sample:</strong> 2-8 seconds </li>\n<li><strong>Batch size 8:</strong> 16-64 seconds per batch</li>\n<li><strong>Dataset scale:</strong> 4,348 series × 200-900 DICOM files each</li>\n<li><strong>Training time:</strong> Would take 3-9 hours just for data loading per epoch</li>\n<li><strong>Environment:</strong> Kaggle notebook with Tesla T4 GPU (16GB)<br>\n<strong>Bottleneck breakdown (tested on Kaggle):</strong></li>\n</ul>\n<ol>\n<li><strong>File I/O:</strong> 0.5-2.0s (loading hundreds of DICOM files per series)</li>\n<li><strong>DICOM parsing:</strong> 0.2-0.8s (pydicom processing with <code>pydicom.dcmread()</code>)</li>\n<li><strong>Image processing:</strong> 0.3-1.0s (statistical normalization using <code>np.percentile()</code>, resizing to 384×384)</li>\n<li><strong>3D volume creation:</strong> 0.1-0.3s (stacking slices with <code>scipy.ndimage.zoom</code>)</li>\n</ol>\n<h3><strong>Current DICOM Processing Pipeline</strong></h3>\n<p>My <code>DICOMPreprocessorKaggle</code> class handles:</p>\n<ul>\n<li>Loading 200-900 <code>.dcm</code> files per series with <code>os.walk()</code></li>\n<li>Statistical normalization: <code>p1, p99 = np.percentile(img, [1, 99])</code></li>\n<li>2D resizing: <code>cv2.resize(processed_img, (384, 384))</code></li>\n<li>3D volume creation: <code>ndimage.zoom()</code> for final (32, 384, 384) shape</li>\n<li>Memory cleanup with <code>gc.collect()</code></li>\n</ul>\n<h3><strong>Questions</strong></h3>\n<ol>\n<li><strong>Pre-processing approach:</strong> Should I convert all DICOM series to <code>.npy</code> files beforehand to avoid real-time processing?</li>\n<li><strong>Kaggle GPU optimization:</strong> Any specific strategies for Tesla T4 environment with limited session time?</li>\n<li><strong>Alternative libraries:</strong> Faster alternatives to <code>pydicom</code> + <code>scipy.ndimage</code> combination?</li>\n<li><strong>DataLoader settings:</strong> Best practices for <code>num_workers</code> in Kaggle? (Currently using 0 due to multiprocessing issues)</li>\n</ol>\n<h3><strong>What I've Tried</strong></h3>\n<ul>\n<li><code>num_workers=0</code> (multiprocessing causes issues in Kaggle environment)</li>\n<li>Small batch sizes (8) for GPU memory management</li>\n<li><code>torch.cuda.empty_cache()</code> and memory cleanup</li>\n<li>Mixed precision training to optimize GPU usage</li>\n</ul>\n<h3><strong>Target Goal</strong></h3>\n<p>Looking for practical solutions to get batch loading under 1 second in Kaggle GPU environment. Current GPU utilization is poor due to waiting for data loading.<br>\nAny medical imaging competition veterans with experience optimizing DICOM workflows on Kaggle GPU instances?<br>\n<strong>Tags:</strong> #rsna2024 #dicom #pytorch #dataloader #performance #kaggle-gpu</p>",
      "rawMarkdown": "### **Problem**\nMy `RSNADataset` is causing severe training bottlenecks due to real-time DICOM processing in Kaggle GPU environment:\n```python\nclass RSNADataset(Dataset):\n    def __getitem__(self, idx):\n        row = self.df.iloc[idx]\n        series_path = f\"{self.series_root}/{row[ID_COL]}\"\n        volume = process_dicom_series_safe(series_path, target_shape=(32, 384, 384))\n        x = torch.from_numpy(volume).float()\n        return x, y\n```\n### **Performance Issues**\n- **Per sample:** 2-8 seconds \n- **Batch size 8:** 16-64 seconds per batch\n- **Dataset scale:** 4,348 series × 200-900 DICOM files each\n- **Training time:** Would take 3-9 hours just for data loading per epoch\n- **Environment:** Kaggle notebook with Tesla T4 GPU (16GB)\n**Bottleneck breakdown (tested on Kaggle):**\n1. **File I/O:** 0.5-2.0s (loading hundreds of DICOM files per series)\n2. **DICOM parsing:** 0.2-0.8s (pydicom processing with `pydicom.dcmread()`)\n3. **Image processing:** 0.3-1.0s (statistical normalization using `np.percentile()`, resizing to 384×384)\n4. **3D volume creation:** 0.1-0.3s (stacking slices with `scipy.ndimage.zoom`)\n### **Current DICOM Processing Pipeline**\nMy `DICOMPreprocessorKaggle` class handles:\n- Loading 200-900 `.dcm` files per series with `os.walk()`\n- Statistical normalization: `p1, p99 = np.percentile(img, [1, 99])`\n- 2D resizing: `cv2.resize(processed_img, (384, 384))`\n- 3D volume creation: `ndimage.zoom()` for final (32, 384, 384) shape\n- Memory cleanup with `gc.collect()`\n### **Questions**\n1. **Pre-processing approach:** Should I convert all DICOM series to `.npy` files beforehand to avoid real-time processing?\n2. **Kaggle GPU optimization:** Any specific strategies for Tesla T4 environment with limited session time?\n3. **Alternative libraries:** Faster alternatives to `pydicom` + `scipy.ndimage` combination?\n4. **DataLoader settings:** Best practices for `num_workers` in Kaggle? (Currently using 0 due to multiprocessing issues)\n### **What I've Tried**\n-  `num_workers=0` (multiprocessing causes issues in Kaggle environment)\n- Small batch sizes (8) for GPU memory management\n-  `torch.cuda.empty_cache()` and memory cleanup\n-  Mixed precision training to optimize GPU usage\n### **Target Goal**\nLooking for practical solutions to get batch loading under 1 second in Kaggle GPU environment. Current GPU utilization is poor due to waiting for data loading.\nAny medical imaging competition veterans with experience optimizing DICOM workflows on Kaggle GPU instances?\n**Tags:** #rsna2024 #dicom #pytorch #dataloader #performance #kaggle-gpu",
      "votes": 6
    },
    {
      "id": 3294471,
      "postDate": "2025-09-26T08:25:30.213Z",
      "content": "<p>Hi, I am struggling with the same issue. I have come across <code>dicomsdl</code>, which at least loads tht pixel-data much faster than <code>pydicom</code>. You can check out this notebook:<br>\n<a href=\"https://www.kaggle.com/code/thalro/test-read-speed\" target=\"_blank\">https://www.kaggle.com/code/thalro/test-read-speed</a><br>\nIt is still painfully slow. If you load all of the slices it takes around 1.5 hours in a kaggle notebook.</p>",
      "rawMarkdown": "Hi, I am struggling with the same issue. I have come across `dicomsdl`, which at least loads tht pixel-data much faster than `pydicom`. You can check out this notebook:\nhttps://www.kaggle.com/code/thalro/test-read-speed\nIt is still painfully slow. If you load all of the slices it takes around 1.5 hours in a kaggle notebook.",
      "votes": 1
    },
    {
      "id": 3295237,
      "postDate": "2025-09-28T05:25:54.357Z",
      "content": "<p>Hi, are you facing this data loading issue while inferencing as well?</p>",
      "rawMarkdown": "Hi, are you facing this data loading issue while inferencing as well?"
    },
    {
      "id": 3295228,
      "postDate": "2025-09-28T04:35:17.377Z",
      "content": "<p>You can convert DCM volumes to NIFTI, then use them to train. </p>",
      "rawMarkdown": "You can convert DCM volumes to NIFTI, then use them to train. "
    },
    {
      "id": 3294935,
      "postDate": "2025-09-27T06:44:48.733Z",
      "content": "<p>Parallel loading of dicom files is possible. Use that to load images, it's much faster</p>",
      "rawMarkdown": "Parallel loading of dicom files is possible. Use that to load images, it's much faster"
    }
  ],
  "comments": [
    {
      "id": 3294471,
      "author_name": "thomas rost",
      "author_url": "",
      "post_date": "2025-09-26T08:25:30.213000",
      "content": "<p>Hi, I am struggling with the same issue. I have come across <code>dicomsdl</code>, which at least loads tht pixel-data much faster than <code>pydicom</code>. You can check out this notebook:<br>\n<a href=\"https://www.kaggle.com/code/thalro/test-read-speed\" target=\"_blank\">https://www.kaggle.com/code/thalro/test-read-speed</a><br>\nIt is still painfully slow. If you load all of the slices it takes around 1.5 hours in a kaggle notebook.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3295237,
      "author_name": "ArjunB",
      "author_url": "",
      "post_date": "2025-09-28T05:25:54.357000",
      "content": "<p>Hi, are you facing this data loading issue while inferencing as well?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3295228,
      "author_name": "Saroj Neupane",
      "author_url": "",
      "post_date": "2025-09-28T04:35:17.377000",
      "content": "<p>You can convert DCM volumes to NIFTI, then use them to train. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3294935,
      "author_name": "jucious",
      "author_url": "",
      "post_date": "2025-09-27T06:44:48.733000",
      "content": "<p>Parallel loading of dicom files is possible. Use that to load images, it's much faster</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3294219": "### **Problem**\nMy `RSNADataset` is causing severe training bottlenecks due to real-time DICOM processing in Kaggle GPU environment:\n```python\nclass RSNADataset(Dataset):\n    def __getitem__(self, idx):\n        row = self.df.iloc[idx]\n        series_path = f\"{self.series_root}/{row[ID_COL]}\"\n        volume = process_dicom_series_safe(series_path, target_shape=(32, 384, 384))\n        x = torch.from_numpy(volume).float()\n        return x, y\n```\n### **Performance Issues**\n- **Per sample:** 2-8 seconds \n- **Batch size 8:** 16-64 seconds per batch\n- **Dataset scale:** 4,348 series × 200-900 DICOM files each\n- **Training time:** Would take 3-9 hours just for data loading per epoch\n- **Environment:** Kaggle notebook with Tesla T4 GPU (16GB)\n**Bottleneck breakdown (tested on Kaggle):**\n1. **File I/O:** 0.5-2.0s (loading hundreds of DICOM files per series)\n2. **DICOM parsing:** 0.2-0.8s (pydicom processing with `pydicom.dcmread()`)\n3. **Image processing:** 0.3-1.0s (statistical normalization using `np.percentile()`, resizing to 384×384)\n4. **3D volume creation:** 0.1-0.3s (stacking slices with `scipy.ndimage.zoom`)\n### **Current DICOM Processing Pipeline**\nMy `DICOMPreprocessorKaggle` class handles:\n- Loading 200-900 `.dcm` files per series with `os.walk()`\n- Statistical normalization: `p1, p99 = np.percentile(img, [1, 99])`\n- 2D resizing: `cv2.resize(processed_img, (384, 384))`\n- 3D volume creation: `ndimage.zoom()` for final (32, 384, 384) shape\n- Memory cleanup with `gc.collect()`\n### **Questions**\n1. **Pre-processing approach:** Should I convert all DICOM series to `.npy` files beforehand to avoid real-time processing?\n2. **Kaggle GPU optimization:** Any specific strategies for Tesla T4 environment with limited session time?\n3. **Alternative libraries:** Faster alternatives to `pydicom` + `scipy.ndimage` combination?\n4. **DataLoader settings:** Best practices for `num_workers` in Kaggle? (Currently using 0 due to multiprocessing issues)\n### **What I've Tried**\n-  `num_workers=0` (multiprocessing causes issues in Kaggle environment)\n- Small batch sizes (8) for GPU memory management\n-  `torch.cuda.empty_cache()` and memory cleanup\n-  Mixed precision training to optimize GPU usage\n### **Target Goal**\nLooking for practical solutions to get batch loading under 1 second in Kaggle GPU environment. Current GPU utilization is poor due to waiting for data loading.\nAny medical imaging competition veterans with experience optimizing DICOM workflows on Kaggle GPU instances?\n**Tags:** #rsna2024 #dicom #pytorch #dataloader #performance #kaggle-gpu",
    "3294471": "Hi, I am struggling with the same issue. I have come across `dicomsdl`, which at least loads tht pixel-data much faster than `pydicom`. You can check out this notebook:\nhttps://www.kaggle.com/code/thalro/test-read-speed\nIt is still painfully slow. If you load all of the slices it takes around 1.5 hours in a kaggle notebook.",
    "3295237": "Hi, are you facing this data loading issue while inferencing as well?",
    "3295228": "You can convert DCM volumes to NIFTI, then use them to train. ",
    "3294935": "Parallel loading of dicom files is possible. Use that to load images, it's much faster"
  }
}