{
  "id": 610741,
  "title": "Help with parallel DICOM loading during inference",
  "url": "/competitions/rsna-intracranial-aneurysm-detection/discussion/610741",
  "author_name": "Ravon",
  "post_date": "2025-10-06T00:45:33.728000",
  "votes": 0,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I've tried to implement parallel DICOM loading to speed up inference, but I don't really see any significant speedup. Not really sure what I'm doing wrong, any help would be greatly appreciated.</p>\n<p>`    def safe_dcmread(self, filepath):<br>\n        try:<br>\n            return pydicom.dcmread(filepath, force=True)<br>\n        except Exception as e:<br>\n            print(f\"Error reading {filepath}: {e}\")<br>\n            return None</p>\n<pre><code> () -&gt; [[pydicom.Dataset], ]:\n    \n    series_path = Path(series_path)\n    series_name = series_path.name\n\n    dicom_files = []\n     root, _, files  os.walk(series_path):\n         file  files:\n             file.endswith():\n                dicom_files.append(os.path.join(root, file))\n\n      dicom_files:\n         ValueError()\n\n    datasets = []\n    \n     ThreadPoolExecutor(max_workers=os.cpu_count())  executor:\n        futures = [executor.submit(.safe_dcmread, fp)  fp  dicom_files]\n         f  as_completed(futures):\n            ds = f.result()\n             ds   :\n                datasets.append(ds)\n\n      datasets:\n         ValueError()\n\n     datasets, series_name\n</code></pre>\n<p>`</p>",
  "messages": [
    {
      "id": 3298607,
      "postDate": "2025-10-06T00:45:33.730Z",
      "content": "<p>I've tried to implement parallel DICOM loading to speed up inference, but I don't really see any significant speedup. Not really sure what I'm doing wrong, any help would be greatly appreciated.</p>\n<p>`    def safe_dcmread(self, filepath):<br>\n        try:<br>\n            return pydicom.dcmread(filepath, force=True)<br>\n        except Exception as e:<br>\n            print(f\"Error reading {filepath}: {e}\")<br>\n            return None</p>\n<pre><code> () -&gt; [[pydicom.Dataset], ]:\n    \n    series_path = Path(series_path)\n    series_name = series_path.name\n\n    dicom_files = []\n     root, _, files  os.walk(series_path):\n         file  files:\n             file.endswith():\n                dicom_files.append(os.path.join(root, file))\n\n      dicom_files:\n         ValueError()\n\n    datasets = []\n    \n     ThreadPoolExecutor(max_workers=os.cpu_count())  executor:\n        futures = [executor.submit(.safe_dcmread, fp)  fp  dicom_files]\n         f  as_completed(futures):\n            ds = f.result()\n             ds   :\n                datasets.append(ds)\n\n      datasets:\n         ValueError()\n\n     datasets, series_name\n</code></pre>\n<p>`</p>",
      "rawMarkdown": "I've tried to implement parallel DICOM loading to speed up inference, but I don't really see any significant speedup. Not really sure what I'm doing wrong, any help would be greatly appreciated.\n\n\n`    def safe_dcmread(self, filepath):\n        try:\n            return pydicom.dcmread(filepath, force=True)\n        except Exception as e:\n            print(f\"Error reading {filepath}: {e}\")\n            return None\n        \n    def load_dicom_series(self, series_path: str) -> Tuple[List[pydicom.Dataset], str]:\n        \"\"\"\n        Load DICOM series (parallelized version)\n        \"\"\"\n        series_path = Path(series_path)\n        series_name = series_path.name\n    \n        dicom_files = []\n        for root, _, files in os.walk(series_path):\n            for file in files:\n                if file.endswith('.dcm'):\n                    dicom_files.append(os.path.join(root, file))\n    \n        if not dicom_files:\n            raise ValueError(f\"No DICOM files found in {series_path}\")\n    \n        datasets = []\n        # adjust max_workers based on CPU count\n        with ThreadPoolExecutor(max_workers=os.cpu_count()) as executor:\n            futures = [executor.submit(self.safe_dcmread, fp) for fp in dicom_files]\n            for f in as_completed(futures):\n                ds = f.result()\n                if ds is not None:\n                    datasets.append(ds)\n    \n        if not datasets:\n            raise ValueError(f\"No valid DICOM files in {series_path}\")\n    \n        return datasets, series_name\n`"
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3298607": "I've tried to implement parallel DICOM loading to speed up inference, but I don't really see any significant speedup. Not really sure what I'm doing wrong, any help would be greatly appreciated.\n\n\n`    def safe_dcmread(self, filepath):\n        try:\n            return pydicom.dcmread(filepath, force=True)\n        except Exception as e:\n            print(f\"Error reading {filepath}: {e}\")\n            return None\n        \n    def load_dicom_series(self, series_path: str) -> Tuple[List[pydicom.Dataset], str]:\n        \"\"\"\n        Load DICOM series (parallelized version)\n        \"\"\"\n        series_path = Path(series_path)\n        series_name = series_path.name\n    \n        dicom_files = []\n        for root, _, files in os.walk(series_path):\n            for file in files:\n                if file.endswith('.dcm'):\n                    dicom_files.append(os.path.join(root, file))\n    \n        if not dicom_files:\n            raise ValueError(f\"No DICOM files found in {series_path}\")\n    \n        datasets = []\n        # adjust max_workers based on CPU count\n        with ThreadPoolExecutor(max_workers=os.cpu_count()) as executor:\n            futures = [executor.submit(self.safe_dcmread, fp) for fp in dicom_files]\n            for f in as_completed(futures):\n                ds = f.result()\n                if ds is not None:\n                    datasets.append(ds)\n    \n        if not datasets:\n            raise ValueError(f\"No valid DICOM files in {series_path}\")\n    \n        return datasets, series_name\n`"
  }
}