{
  "id": 432593,
  "title": "Possibly corrupt .dcm in test set - Can the host/Kaggle staff check this?",
  "url": "/competitions/rsna-2023-abdominal-trauma-detection/discussion/432593",
  "author_name": "Chau YH",
  "post_date": "2023-08-18T04:24:24.966000",
  "votes": 14,
  "comment_count": 25,
  "views": 0,
  "content": "<h1>EDIT:</h1>\n<p><a href=\"https://www.kaggle.com/samfc10\" target=\"_blank\">@samfc10</a> pointed out that we should load the .dcm files with dicomsdl in the comments below. I have tried dicomsdl and it gives the same results as pydicom and has no errors on the hidden test set. However, <a href=\"https://www.kaggle.com/samfc10\" target=\"_blank\">@samfc10</a> pointed out that he is unable to stack the .dcm files together in the hidden test set, and it seems that the shapes of different .dcm files in the same series are inconsistent.</p>\n<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> found out that there is indeed a corrupted image in the test set. The code is able to run if the corrupted image is excluded, confirming that all shapes are consistent:</p>\n<pre><code>count = \n k  ((x)):\n    patient_id = x.loc[x.index[k], ]\n    series_id = x.loc[x.index[k], ]\n\n\n    patient_folder = os.path.join(folder, (patient_id))\n    series_folder = os.path.join(patient_folder, (series_id))\n    instances = os.listdir(series_folder)\n    prev_shape = \n     instance  instances:\n         ((patient_id) == )  ((series_id) == )  ((instance[:-]) == ):\n            \n\n        path = os.path.join(series_folder, instance)\n        dicom_file = dicomsdl.(path)\n        arr = np.array(dicom_file.pixelData(storedvalue=))\n         (arr.shape) == \n        shape = (arr.shape)\n         prev_shape   :\n             shape == prev_shape\n        prev_shape = shape\n\n        count += \n     (count &gt; )   use_test:\n        \n</code></pre>\n<h1>Original post:</h1>\n<p>Dear Competition host/Kaggle Staff,</p>\n<p>Thanks for hosting this awesome competition! I saw that the description of the competition has been updated to include a <a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/overview\" target=\"_blank\">Getting Started notebook</a> to demonstrate how to train a model and do inference, which is great. However, I think there might be some corrupt .dcm in the test set which leads to random errors while submitting notebooks to the competition.</p>\n<h1>Fork of official inference notebook</h1>\n<p>I forked the official inference notebook and removed the model parts from it. I ran two tests, one with STRIDE=1 and one with STRIDE unchanged, and the STRIDE=1 submission threw an error within 1hr while the other submission is successful.</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/chauyh/kerascv-starter-notebook-infer-fork-stride-1\" target=\"_blank\">STRIDE=1</a></li>\n<li><a href=\"https://www.kaggle.com/code/chauyh/kerascv-starter-notebook-infer-fork-stride-10\" target=\"_blank\">STRIDE&gt;1</a></li>\n</ul>\n<h1>Test with other notebooks</h1>\n<p>I created a notebook myself which simply tries to load the .pixel_array using pydicom and gets the shape of the pixel array. The notebook which only loads 100 .dcm files is successful, while the one that loads all the .dcm files crashes within 1hr.</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/chauyh/rsna-pixel-array-100\" target=\"_blank\">100 .dcm</a></li>\n<li><a href=\"https://www.kaggle.com/code/chauyh/rsna-pixel-array-all\" target=\"_blank\">all .dcm</a></li>\n</ul>\n<p>I have more (not so comprehensive) tests in my <a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/431836\" target=\"_blank\">previous discussion</a>.</p>\n<h1>How to fix this</h1>\n<p>I think this is pretty solid evidence that there are some corrupt .dcm files in the test set. Or it could be some unknown bug in my implementation of loading .dicom files. Anyways, could the host/Kaggle Staff modify the official inference notebook to load all the .dcm files with STRIDE=1 instead of a partial list of .dcm files with STRIDE&gt;1 with pydicom?</p>\n<p>Could the host/Kaggle staff also check whether the SOP Class UID of all .dcm files belong to CT Image Storage, not Enhanced CT Image Storage or other SOP Class UIDs? While I'm not a radiologist, I asked GPT and GPT told me that .dcm files belonging to Enhanced CT Image Storage stores 3D data in a single .dcm file, which may also conflict with the usual image loading pipeline.</p>\n<p>This could check whether or not all .dcm files are fine, and moreover a reference for how to load every .dcm file in the test set into numpy arrays properly for new people who wish to join the competition (and for me too).</p>\n<h1>Other stuff</h1>\n<p>Can the host/Kaggle staff change the description of the Data section to say that the visible test folder is not representative of the actual test folder, where the actual test folder contains multiple slices per series instead of a single slice? I know this is addressed officially by <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> in a comment in the discussion forums, but saying that directly in the Data section would make it clearer for new competitors.</p>\n<p>Thanks!<br>\n-YH</p>",
  "messages": [
    {
      "id": 2396106,
      "postDate": "2023-08-18T04:24:24.967Z",
      "content": "<h1>EDIT:</h1>\n<p><a href=\"https://www.kaggle.com/samfc10\" target=\"_blank\">@samfc10</a> pointed out that we should load the .dcm files with dicomsdl in the comments below. I have tried dicomsdl and it gives the same results as pydicom and has no errors on the hidden test set. However, <a href=\"https://www.kaggle.com/samfc10\" target=\"_blank\">@samfc10</a> pointed out that he is unable to stack the .dcm files together in the hidden test set, and it seems that the shapes of different .dcm files in the same series are inconsistent.</p>\n<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> found out that there is indeed a corrupted image in the test set. The code is able to run if the corrupted image is excluded, confirming that all shapes are consistent:</p>\n<pre><code>count = \n k  ((x)):\n    patient_id = x.loc[x.index[k], ]\n    series_id = x.loc[x.index[k], ]\n\n\n    patient_folder = os.path.join(folder, (patient_id))\n    series_folder = os.path.join(patient_folder, (series_id))\n    instances = os.listdir(series_folder)\n    prev_shape = \n     instance  instances:\n         ((patient_id) == )  ((series_id) == )  ((instance[:-]) == ):\n            \n\n        path = os.path.join(series_folder, instance)\n        dicom_file = dicomsdl.(path)\n        arr = np.array(dicom_file.pixelData(storedvalue=))\n         (arr.shape) == \n        shape = (arr.shape)\n         prev_shape   :\n             shape == prev_shape\n        prev_shape = shape\n\n        count += \n     (count &gt; )   use_test:\n        \n</code></pre>\n<h1>Original post:</h1>\n<p>Dear Competition host/Kaggle Staff,</p>\n<p>Thanks for hosting this awesome competition! I saw that the description of the competition has been updated to include a <a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/overview\" target=\"_blank\">Getting Started notebook</a> to demonstrate how to train a model and do inference, which is great. However, I think there might be some corrupt .dcm in the test set which leads to random errors while submitting notebooks to the competition.</p>\n<h1>Fork of official inference notebook</h1>\n<p>I forked the official inference notebook and removed the model parts from it. I ran two tests, one with STRIDE=1 and one with STRIDE unchanged, and the STRIDE=1 submission threw an error within 1hr while the other submission is successful.</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/chauyh/kerascv-starter-notebook-infer-fork-stride-1\" target=\"_blank\">STRIDE=1</a></li>\n<li><a href=\"https://www.kaggle.com/code/chauyh/kerascv-starter-notebook-infer-fork-stride-10\" target=\"_blank\">STRIDE&gt;1</a></li>\n</ul>\n<h1>Test with other notebooks</h1>\n<p>I created a notebook myself which simply tries to load the .pixel_array using pydicom and gets the shape of the pixel array. The notebook which only loads 100 .dcm files is successful, while the one that loads all the .dcm files crashes within 1hr.</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/chauyh/rsna-pixel-array-100\" target=\"_blank\">100 .dcm</a></li>\n<li><a href=\"https://www.kaggle.com/code/chauyh/rsna-pixel-array-all\" target=\"_blank\">all .dcm</a></li>\n</ul>\n<p>I have more (not so comprehensive) tests in my <a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/431836\" target=\"_blank\">previous discussion</a>.</p>\n<h1>How to fix this</h1>\n<p>I think this is pretty solid evidence that there are some corrupt .dcm files in the test set. Or it could be some unknown bug in my implementation of loading .dicom files. Anyways, could the host/Kaggle Staff modify the official inference notebook to load all the .dcm files with STRIDE=1 instead of a partial list of .dcm files with STRIDE&gt;1 with pydicom?</p>\n<p>Could the host/Kaggle staff also check whether the SOP Class UID of all .dcm files belong to CT Image Storage, not Enhanced CT Image Storage or other SOP Class UIDs? While I'm not a radiologist, I asked GPT and GPT told me that .dcm files belonging to Enhanced CT Image Storage stores 3D data in a single .dcm file, which may also conflict with the usual image loading pipeline.</p>\n<p>This could check whether or not all .dcm files are fine, and moreover a reference for how to load every .dcm file in the test set into numpy arrays properly for new people who wish to join the competition (and for me too).</p>\n<h1>Other stuff</h1>\n<p>Can the host/Kaggle staff change the description of the Data section to say that the visible test folder is not representative of the actual test folder, where the actual test folder contains multiple slices per series instead of a single slice? I know this is addressed officially by <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> in a comment in the discussion forums, but saying that directly in the Data section would make it clearer for new competitors.</p>\n<p>Thanks!<br>\n-YH</p>",
      "rawMarkdown": "# EDIT:\n@samfc10 pointed out that we should load the .dcm files with dicomsdl in the comments below. I have tried dicomsdl and it gives the same results as pydicom and has no errors on the hidden test set. However, @samfc10 pointed out that he is unable to stack the .dcm files together in the hidden test set, and it seems that the shapes of different .dcm files in the same series are inconsistent.\n\n@sohier found out that there is indeed a corrupted image in the test set. The code is able to run if the corrupted image is excluded, confirming that all shapes are consistent:\n```python\ncount = 0\nfor k in range(len(x)):\n    patient_id = x.loc[x.index[k], \"patient_id\"]\n    series_id = x.loc[x.index[k], \"series_id\"]\n\n    \n    patient_folder = os.path.join(folder, str(patient_id))\n    series_folder = os.path.join(patient_folder, str(series_id))\n    instances = os.listdir(series_folder)\n    prev_shape = None\n    for instance in instances:\n        if (int(patient_id) == 3124) and (int(series_id) == 5842) and (int(instance[:-4]) == 514):\n            continue\n        \n        path = os.path.join(series_folder, instance)\n        dicom_file = dicomsdl.open(path)\n        arr = np.array(dicom_file.pixelData(storedvalue=True))\n        assert len(arr.shape) == 2\n        shape = str(arr.shape)\n        if prev_shape is not None:\n            assert shape == prev_shape\n        prev_shape = shape\n        \n        count += 1\n    if (count > 100) and not use_test:\n        break\n```\n\n# Original post:\n\nDear Competition host/Kaggle Staff,\n\nThanks for hosting this awesome competition! I saw that the description of the competition has been updated to include a [Getting Started notebook](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/overview) to demonstrate how to train a model and do inference, which is great. However, I think there might be some corrupt .dcm in the test set which leads to random errors while submitting notebooks to the competition.\n\n# Fork of official inference notebook\nI forked the official inference notebook and removed the model parts from it. I ran two tests, one with STRIDE=1 and one with STRIDE unchanged, and the STRIDE=1 submission threw an error within 1hr while the other submission is successful.\n\n- [STRIDE=1](https://www.kaggle.com/code/chauyh/kerascv-starter-notebook-infer-fork-stride-1)\n- [STRIDE>1](https://www.kaggle.com/code/chauyh/kerascv-starter-notebook-infer-fork-stride-10)\n\n# Test with other notebooks\nI created a notebook myself which simply tries to load the .pixel_array using pydicom and gets the shape of the pixel array. The notebook which only loads 100 .dcm files is successful, while the one that loads all the .dcm files crashes within 1hr.\n\n- [100 .dcm](https://www.kaggle.com/code/chauyh/rsna-pixel-array-100)\n- [all .dcm](https://www.kaggle.com/code/chauyh/rsna-pixel-array-all)\n\nI have more (not so comprehensive) tests in my [previous discussion](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/431836).\n\n# How to fix this\nI think this is pretty solid evidence that there are some corrupt .dcm files in the test set. Or it could be some unknown bug in my implementation of loading .dicom files. Anyways, could the host/Kaggle Staff modify the official inference notebook to load all the .dcm files with STRIDE=1 instead of a partial list of .dcm files with STRIDE>1 with pydicom?\n\nCould the host/Kaggle staff also check whether the SOP Class UID of all .dcm files belong to CT Image Storage, not Enhanced CT Image Storage or other SOP Class UIDs? While I'm not a radiologist, I asked GPT and GPT told me that .dcm files belonging to Enhanced CT Image Storage stores 3D data in a single .dcm file, which may also conflict with the usual image loading pipeline.\n\nThis could check whether or not all .dcm files are fine, and moreover a reference for how to load every .dcm file in the test set into numpy arrays properly for new people who wish to join the competition (and for me too).\n\n# Other stuff\nCan the host/Kaggle staff change the description of the Data section to say that the visible test folder is not representative of the actual test folder, where the actual test folder contains multiple slices per series instead of a single slice? I know this is addressed officially by @sohier in a comment in the discussion forums, but saying that directly in the Data section would make it clearer for new competitors.\n\nThanks!\n-YH",
      "votes": 14
    },
    {
      "id": 2412964,
      "postDate": "2023-08-28T15:50:37.713Z",
      "content": "<blockquote>\n  <p>Use dicomsdl instead of pydicom to fix this issue.</p>\n</blockquote>\n<p>Actually dicomsdl doesn't fix the issue. The core problem, i.e corrupted dicoms in test set still remains. Although dicomsdl doesn't raise any error when reading dicoms, there's some error when stacking the tensors to form a 3D tensor, which I'm unable to debug due to the limited error message during submission. Stacking works with pydicom and try-except block during dicom reading.</p>\n<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> Could you please take a look into this issue. It's been 10 days since the discussion has been opened, and there's no response from the hosts, or the kaggle staff. It's sad to see. A lot of competitor's time and submissions have been wasted in trying to debug this issue.</p>",
      "rawMarkdown": "> Use dicomsdl instead of pydicom to fix this issue.\n\nActually dicomsdl doesn't fix the issue. The core problem, i.e corrupted dicoms in test set still remains. Although dicomsdl doesn't raise any error when reading dicoms, there's some error when stacking the tensors to form a 3D tensor, which I'm unable to debug due to the limited error message during submission. Stacking works with pydicom and try-except block during dicom reading.\n\n@sohier Could you please take a look into this issue. It's been 10 days since the discussion has been opened, and there's no response from the hosts, or the kaggle staff. It's sad to see. A lot of competitor's time and submissions have been wasted in trying to debug this issue.",
      "votes": 3,
      "replies": [
        {
          "id": 2413003,
          "postDate": "2023-08-28T16:17:05.003Z",
          "content": "<p>Nice catch. I have updated the post to reflect this. In my test with dicomsdl, I merely accessed the data with<br>\n<code>\ndicomsdl.open(src+dcm_dir).pixelData(storedvalue=True)\n</code><br>\nwhich seems to create a memmap of the .dcm instead of loading the full array. If there are still corrupted dicoms, the exception in your notebook probably occurred in some code that actually loads the memmap into memory. Maybe np.stack is the actual code that loads the array into memory because of numpy internals? I am running a variant of the notebook that tests <br>\n<code>\nnp.array(dicomsdl.open(src+dcm_dir).pixelData(storedvalue=True))\n</code><br>\nto see whether some files are really corrupted or dicomsdl solves the issues.</p>",
          "rawMarkdown": "Nice catch. I have updated the post to reflect this. In my test with dicomsdl, I merely accessed the data with\n``\ndicomsdl.open(src+dcm_dir).pixelData(storedvalue=True)\n``\nwhich seems to create a memmap of the .dcm instead of loading the full array. If there are still corrupted dicoms, the exception in your notebook probably occurred in some code that actually loads the memmap into memory. Maybe np.stack is the actual code that loads the array into memory because of numpy internals? I am running a variant of the notebook that tests \n``\nnp.array(dicomsdl.open(src+dcm_dir).pixelData(storedvalue=True))\n``\nto see whether some files are really corrupted or dicomsdl solves the issues.",
          "replies": [
            {
              "id": 2413093,
              "postDate": "2023-08-28T17:01:59.470Z",
              "content": "<p>just adding another line to my previous code leads to an exception after 1 hour of submission.</p>\n<pre><code>src = \ndcm_dirs = os.listdir(src)\ndcm_dirs.sort(key =  x: (x.split()[]))\narr_list = []\n\n dcm_dir  dcm_dirs:\n    arr = dicomsdl.(src+dcm_dir).pixelData(storedvalue=)\n    arr_list.append(torch.from_numpy(arr.astype(np.float32)))\n\narr_stacked = torch.stack(arr_list, dim=)\n</code></pre>",
              "rawMarkdown": "just adding another line to my previous code leads to an exception after 1 hour of submission.\n\n```python\nsrc = f'/kaggle/input/rsna-2023-abdominal-trauma-detection/test_images/{patient_id}/{series_id}/'\ndcm_dirs = os.listdir(src)\ndcm_dirs.sort(key = lambda x: int(x.split('.')[0]))\narr_list = []\n\nfor dcm_dir in dcm_dirs:\n    arr = dicomsdl.open(src+dcm_dir).pixelData(storedvalue=True)\n    arr_list.append(torch.from_numpy(arr.astype(np.float32)))\n\narr_stacked = torch.stack(arr_list, dim=0)\n```"
            },
            {
              "id": 2413460,
              "postDate": "2023-08-29T00:05:53.113Z",
              "content": "<p>I changed my notebook to the following code:<br>\n``<br>\ncount = 0<br>\nfor k in range(len(x)):<br>\n    patient_id = x.loc[x.index[k], \"patient_id\"]<br>\n    series_id = x.loc[x.index[k], \"series_id\"]</p>\n<pre><code>patient_folder = os.path.str(patient_id))\nseries_folder = os.path.str(series_id))\n= os.listdir(series_folder)\nfor in     path = os.path.    =     arr = np.array(    = str(arr.\n     += \nif ( &gt; ) not use_test:\n    </code></pre>\n<p>``<br>\nand it runs without problems. I am able to obtain the arr from every .dcm file in the hidden test set. Let me check whether the shapes of each .dcm in the same series are consistent, which may lead to errors when stacking.</p>",
              "rawMarkdown": "I changed my notebook to the following code:\n``\ncount = 0\nfor k in range(len(x)):\n    patient_id = x.loc[x.index[k], \"patient_id\"]\n    series_id = x.loc[x.index[k], \"series_id\"]\n\n    \n    patient_folder = os.path.join(folder, str(patient_id))\n    series_folder = os.path.join(patient_folder, str(series_id))\n    instances = os.listdir(series_folder)\n    for instance in instances:\n        path = os.path.join(series_folder, instance)\n        dicom_file = dicomsdl.open(path)\n        arr = np.array(dicom_file.pixelData(storedvalue=True))\n        shape = str(arr.shape)\n        \n        count += 1\n    if (count > 100) and not use_test:\n        break\n``\nand it runs without problems. I am able to obtain the arr from every .dcm file in the hidden test set. Let me check whether the shapes of each .dcm in the same series are consistent, which may lead to errors when stacking."
            },
            {
              "id": 2413489,
              "postDate": "2023-08-29T00:32:20.113Z",
              "content": "<p>As Chau YH mentioned, it might be because of the different image sizes in one sample folder. If that's right, I think the competition host must mention about that.</p>",
              "rawMarkdown": "As Chau YH mentioned, it might be because of the different image sizes in one sample folder. If that's right, I think the competition host must mention about that."
            },
            {
              "id": 2413577,
              "postDate": "2023-08-29T02:53:11.377Z",
              "content": "<p>Indeed, the dcm files in the same series have different shapes in the hidden test set.</p>\n<p>``</p>\n<h1>loop through the folders in the folder</h1>\n<p>count = 0<br>\nfor k in range(len(x)):<br>\n    patient_id = x.loc[x.index[k], \"patient_id\"]<br>\n    series_id = x.loc[x.index[k], \"series_id\"]</p>\n<pre><code>patient_folder = os.path.str(patient_id))\nseries_folder = os.path.str(series_id))\n= os.listdir(series_folder)\nprev_shape = None\nfor in     path = os.path.    =     arr = np.array(    = str(arr.    if prev_shape is not None:\n        assert == prev_shape\n    prev_shape = \n     += \nif ( &gt; ) not use_test:\n    </code></pre>\n<p>``</p>\n<p>Adding an assert statement that checks the shape throws an error during submission. This indeed shows that different .dcm files in the same series may contain different shapes. This also casts doubt on the integrity of the test data, it might be the case that slices belonging to different scans (series) are mixed together in the same series folder.</p>\n<p>This error interferes with pipelines that tries to make use of the whole 3D scan, for example TotalSegmentator as a preprocessing step. I think the best we could do is to loop through every slice in each series and decide which ones to use using a majority vote for the shape within that series, which should remove all the incorrect/corrupted ones.</p>\n<p>Of course, it would be nice if the host/Kaggle staff could run a simple for loop to check for these errors or to provide a sample notebook that loads every .dcm file in every series into a 3D npy array correctly, with a dummy submission.</p>",
              "rawMarkdown": "Indeed, the dcm files in the same series have different shapes in the hidden test set.\n\n``\n# loop through the folders in the folder\ncount = 0\nfor k in range(len(x)):\n    patient_id = x.loc[x.index[k], \"patient_id\"]\n    series_id = x.loc[x.index[k], \"series_id\"]\n\n    \n    patient_folder = os.path.join(folder, str(patient_id))\n    series_folder = os.path.join(patient_folder, str(series_id))\n    instances = os.listdir(series_folder)\n    prev_shape = None\n    for instance in instances:\n        path = os.path.join(series_folder, instance)\n        dicom_file = dicomsdl.open(path)\n        arr = np.array(dicom_file.pixelData(storedvalue=True))\n        shape = str(arr.shape)\n        if prev_shape is not None:\n            assert shape == prev_shape\n        prev_shape = shape\n        \n        count += 1\n    if (count > 100) and not use_test:\n        break\n``\n\nAdding an assert statement that checks the shape throws an error during submission. This indeed shows that different .dcm files in the same series may contain different shapes. This also casts doubt on the integrity of the test data, it might be the case that slices belonging to different scans (series) are mixed together in the same series folder.\n\nThis error interferes with pipelines that tries to make use of the whole 3D scan, for example TotalSegmentator as a preprocessing step. I think the best we could do is to loop through every slice in each series and decide which ones to use using a majority vote for the shape within that series, which should remove all the incorrect/corrupted ones.\n\nOf course, it would be nice if the host/Kaggle staff could run a simple for loop to check for these errors or to provide a sample notebook that loads every .dcm file in every series into a 3D npy array correctly, with a dummy submission.",
              "votes": 1
            },
            {
              "id": 2416360,
              "postDate": "2023-08-31T01:50:13.940Z",
              "content": "<p>Just to add a to the conversation, found this cool function in the code of dicom2nifti that is able to quickly check if the dicom is valid.</p>\n<pre><code> ():\n    \n    file_stream = (filename, )\n    file_stream.seek()\n    data = file_stream.read()\n    file_stream.close()\n     data == :\n         \n     dicom2nifti.settings.pydicom_read_force:\n        :\n            dicom_headers = pydicom.read_file(filename, defer_size=, stop_before_pixels=, force=)\n             dicom_headers   :\n                 \n        :\n            \n     \n</code></pre>\n<p>Source: <a href=\"https://github.com/icometrix/dicom2nifti/blob/main/dicom2nifti/common.py\" target=\"_blank\">https://github.com/icometrix/dicom2nifti/blob/main/dicom2nifti/common.py</a></p>",
              "rawMarkdown": "Just to add a to the conversation, found this cool function in the code of dicom2nifti that is able to quickly check if the dicom is valid.\n\n```python\ndef is_dicom_file(filename):\n    \"\"\"\n    Util function to check if file is a dicom file\n    the first 128 bytes are preamble\n    the next 4 bytes should contain DICM otherwise it is not a dicom\n\n    :param filename: file to check for the DICM header block\n    :type filename: str\n    :returns: True if it is a dicom file\n    \"\"\"\n    file_stream = open(filename, 'rb')\n    file_stream.seek(128)\n    data = file_stream.read(4)\n    file_stream.close()\n    if data == b'DICM':\n        return True\n    if dicom2nifti.settings.pydicom_read_force:\n        try:\n            dicom_headers = pydicom.read_file(filename, defer_size=\"1 KB\", stop_before_pixels=True, force=True)\n            if dicom_headers is not None:\n                return True\n        except:\n            pass\n    return False\n```\n\nSource: https://github.com/icometrix/dicom2nifti/blob/main/dicom2nifti/common.py"
            }
          ]
        },
        {
          "id": 2413479,
          "postDate": "2023-08-29T00:26:39.023Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 2411251,
      "postDate": "2023-08-27T13:59:21.250Z",
      "content": "<p>I did some basic LB probing to validate this. I am able to read all dicom files using dicomsdl, without any try-except block. Also all dicom files are 2D only, same as the train set.</p>",
      "rawMarkdown": "I did some basic LB probing to validate this. I am able to read all dicom files using dicomsdl, without any try-except block. Also all dicom files are 2D only, same as the train set.",
      "votes": 4,
      "replies": [
        {
          "id": 2411295,
          "postDate": "2023-08-27T14:28:53.960Z",
          "content": "<p>Thanks for the information. So it might be a problem with pydicom, or there might be some weird bug in my code. Could you share a notebook on how to successfully load every .dcm in the hidden test set?</p>",
          "rawMarkdown": "Thanks for the information. So it might be a problem with pydicom, or there might be some weird bug in my code. Could you share a notebook on how to successfully load every .dcm in the hidden test set?",
          "replies": [
            {
              "id": 2411464,
              "postDate": "2023-08-27T16:25:28.043Z",
              "content": "<pre><code>src = \ndcm_dirs = os.listdir(src)\ndcm_dirs.sort(key =  x: (x.split()[]))\narr_list = []\n\n dcm_dir  dcm_dirs:\n    arr = dicomsdl.(src+dcm_dir).pixelData(storedvalue=)\n    arr_list.append(torch.from_numpy(arr.astype(np.float32)))\n</code></pre>",
              "rawMarkdown": "```python\nsrc = f'/kaggle/input/rsna-2023-abdominal-trauma-detection/test_images/{patient_id}/{series_id}/'\ndcm_dirs = os.listdir(src)\ndcm_dirs.sort(key = lambda x: int(x.split('.')[0]))\narr_list = []\n    \nfor dcm_dir in dcm_dirs:\n    arr = dicomsdl.open(src+dcm_dir).pixelData(storedvalue=True)\n    arr_list.append(torch.from_numpy(arr.astype(np.float32)))\n```",
              "votes": 3
            },
            {
              "id": 2411828,
              "postDate": "2023-08-27T21:58:37.063Z",
              "content": "<p>Cool! Thank you so much!</p>",
              "rawMarkdown": "Cool! Thank you so much!",
              "votes": 1
            },
            {
              "id": 2412152,
              "postDate": "2023-08-28T05:45:38.717Z",
              "content": "<p>Thanks. I changed to dicomsdl and there are no problems. I have updated my post to reflect that.</p>",
              "rawMarkdown": "Thanks. I changed to dicomsdl and there are no problems. I have updated my post to reflect that."
            }
          ]
        }
      ]
    },
    {
      "id": 2407194,
      "postDate": "2023-08-25T00:00:57.050Z",
      "content": "<p>i suggest some \"kaggle bug report prize\" like swag prize for report like this.<br>\nit takes a lot of time and effort to identify issues like this and more importantly, it helps others a lot!</p>",
      "rawMarkdown": "i suggest some \"kaggle bug report prize\" like swag prize for report like this.\nit takes a lot of time and effort to identify issues like this and more importantly, it helps others a lot!",
      "votes": 1,
      "replies": [
        {
          "id": 2409580,
          "postDate": "2023-08-26T10:24:38.067Z",
          "content": "<p>Cool idea, and I think we really need such incentives. I have submitted too many failed submissions for this competition (more than 10 times). Frustrated…</p>",
          "rawMarkdown": "Cool idea, and I think we really need such incentives. I have submitted too many failed submissions for this competition (more than 10 times). Frustrated...\n\n",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 2414429,
      "postDate": "2023-08-29T16:13:38.687Z",
      "content": "<p>I will take a look but won't be able to follow up until later this week.</p>",
      "rawMarkdown": "I will take a look but won't be able to follow up until later this week.",
      "votes": 2,
      "replies": [
        {
          "id": 2414824,
          "postDate": "2023-08-30T01:33:36.413Z",
          "content": "<p>Thanks so much! I hope this problem can be resolved.</p>",
          "rawMarkdown": "Thanks so much! I hope this problem can be resolved."
        },
        {
          "id": 2416256,
          "postDate": "2023-08-30T22:53:12.150Z",
          "content": "<p>Please see <a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/435815\" target=\"_blank\">this post for my full response</a>; I did find one corrupt image.</p>",
          "rawMarkdown": "Please see [this post for my full response](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/435815); I did find one corrupt image.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2405972,
      "postDate": "2023-08-24T07:10:50.970Z",
      "content": "<p>I also spent more than 10 hrs to debug the problem. After that, I just included a try-except statement to skip loading the problematic dcm files. Because I could not debug it directly with only 'Notebook threw exception' errors.</p>",
      "rawMarkdown": "I also spent more than 10 hrs to debug the problem. After that, I just included a try-except statement to skip loading the problematic dcm files. Because I could not debug it directly with only 'Notebook threw exception' errors.",
      "votes": 2
    },
    {
      "id": 2398066,
      "postDate": "2023-08-19T12:39:43.560Z",
      "content": "<p>I have processed all the images in the dataset and have not found any bugs or corrupt images. I haven't checked the DICOM tags for SOP class, but I have not seen any multi-frame images. I suspect they're all standard CT Image SOP Classes.</p>",
      "rawMarkdown": "I have processed all the images in the dataset and have not found any bugs or corrupt images. I haven't checked the DICOM tags for SOP class, but I have not seen any multi-frame images. I suspect they're all standard CT Image SOP Classes.",
      "votes": 2,
      "replies": [
        {
          "id": 2398410,
          "postDate": "2023-08-19T17:06:27.607Z",
          "content": "<p>Thanks for your reply. I have also successfully translated all .dcm files in the training folder to hdf5, but I could not load every .dcm file with pydicom in the hidden test folder, as indicated in my notebooks above, which crashes during submission. I suspect the error is raised when I use the .pixel_array attribute on a corrupt .hdf5 in the hidden test folder, but I can't confirm that as the error is redacted for submission notebooks.</p>\n<p>If you could load every .dcm in the hidden test folder, that means there's probably some weird bug in my code. If you have the time, could you share publicly a minimal working example of loading every .dcm file with pydicom's .pixel_array attribute in the hidden test set during submission (of course without any models or inference code, just make a dummy submission.csv)? Thanks for your help.</p>\n<p>I am thinking of a 3D model which uses every slice in a series for inference, but I can't load all the slices in the hidden test set.</p>",
          "rawMarkdown": "Thanks for your reply. I have also successfully translated all .dcm files in the training folder to hdf5, but I could not load every .dcm file with pydicom in the hidden test folder, as indicated in my notebooks above, which crashes during submission. I suspect the error is raised when I use the .pixel_array attribute on a corrupt .hdf5 in the hidden test folder, but I can't confirm that as the error is redacted for submission notebooks.\n\nIf you could load every .dcm in the hidden test folder, that means there's probably some weird bug in my code. If you have the time, could you share publicly a minimal working example of loading every .dcm file with pydicom's .pixel_array attribute in the hidden test set during submission (of course without any models or inference code, just make a dummy submission.csv)? Thanks for your help.\n\nI am thinking of a 3D model which uses every slice in a series for inference, but I can't load all the slices in the hidden test set.",
          "votes": 1,
          "replies": [
            {
              "id": 2398848,
              "postDate": "2023-08-20T02:38:46.483Z",
              "content": "<p>Sorry <a href=\"https://www.kaggle.com/chauyh\" target=\"_blank\">@chauyh</a>, I misread your question. I have not made a submission, so I have not processed the test set. I can't confirm there are no corrupt images in the test set.</p>",
              "rawMarkdown": "Sorry @chauyh, I misread your question. I have not made a submission, so I have not processed the test set. I can't confirm there are no corrupt images in the test set.",
              "votes": 1
            },
            {
              "id": 2399751,
              "postDate": "2023-08-20T14:58:58.373Z",
              "content": "<p>No problem. If you happen to be able to load all .dcm files in the test set, I hope that you can share a public notebook with us😁.</p>",
              "rawMarkdown": "No problem. If you happen to be able to load all .dcm files in the test set, I hope that you can share a public notebook with us😁."
            }
          ]
        }
      ]
    },
    {
      "id": 2396456,
      "postDate": "2023-08-18T09:11:39.477Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2396453,
      "postDate": "2023-08-18T09:09:45.680Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2412964,
      "author_name": "Jebastin Nadar",
      "author_url": "",
      "post_date": "2023-08-28T15:50:37.713000",
      "content": "<blockquote>\n  <p>Use dicomsdl instead of pydicom to fix this issue.</p>\n</blockquote>\n<p>Actually dicomsdl doesn't fix the issue. The core problem, i.e corrupted dicoms in test set still remains. Although dicomsdl doesn't raise any error when reading dicoms, there's some error when stacking the tensors to form a 3D tensor, which I'm unable to debug due to the limited error message during submission. Stacking works with pydicom and try-except block during dicom reading.</p>\n<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> Could you please take a look into this issue. It's been 10 days since the discussion has been opened, and there's no response from the hosts, or the kaggle staff. It's sad to see. A lot of competitor's time and submissions have been wasted in trying to debug this issue.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2413003,
          "author_name": "Chau YH",
          "author_url": "",
          "post_date": "2023-08-28T16:17:05.003000",
          "content": "<p>Nice catch. I have updated the post to reflect this. In my test with dicomsdl, I merely accessed the data with<br>\n<code>\ndicomsdl.open(src+dcm_dir).pixelData(storedvalue=True)\n</code><br>\nwhich seems to create a memmap of the .dcm instead of loading the full array. If there are still corrupted dicoms, the exception in your notebook probably occurred in some code that actually loads the memmap into memory. Maybe np.stack is the actual code that loads the array into memory because of numpy internals? I am running a variant of the notebook that tests <br>\n<code>\nnp.array(dicomsdl.open(src+dcm_dir).pixelData(storedvalue=True))\n</code><br>\nto see whether some files are really corrupted or dicomsdl solves the issues.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2413093,
              "author_name": "Jebastin Nadar",
              "author_url": "",
              "post_date": "2023-08-28T17:01:59.470000",
              "content": "<p>just adding another line to my previous code leads to an exception after 1 hour of submission.</p>\n<pre><code>src = \ndcm_dirs = os.listdir(src)\ndcm_dirs.sort(key =  x: (x.split()[]))\narr_list = []\n\n dcm_dir  dcm_dirs:\n    arr = dicomsdl.(src+dcm_dir).pixelData(storedvalue=)\n    arr_list.append(torch.from_numpy(arr.astype(np.float32)))\n\narr_stacked = torch.stack(arr_list, dim=)\n</code></pre>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2413460,
              "author_name": "Chau YH",
              "author_url": "",
              "post_date": "2023-08-29T00:05:53.113000",
              "content": "<p>I changed my notebook to the following code:<br>\n``<br>\ncount = 0<br>\nfor k in range(len(x)):<br>\n    patient_id = x.loc[x.index[k], \"patient_id\"]<br>\n    series_id = x.loc[x.index[k], \"series_id\"]</p>\n<pre><code>patient_folder = os.path.str(patient_id))\nseries_folder = os.path.str(series_id))\n= os.listdir(series_folder)\nfor in     path = os.path.    =     arr = np.array(    = str(arr.\n     += \nif ( &gt; ) not use_test:\n    </code></pre>\n<p>``<br>\nand it runs without problems. I am able to obtain the arr from every .dcm file in the hidden test set. Let me check whether the shapes of each .dcm in the same series are consistent, which may lead to errors when stacking.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2413489,
              "author_name": "junseonglee11",
              "author_url": "",
              "post_date": "2023-08-29T00:32:20.113000",
              "content": "<p>As Chau YH mentioned, it might be because of the different image sizes in one sample folder. If that's right, I think the competition host must mention about that.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2413577,
              "author_name": "Chau YH",
              "author_url": "",
              "post_date": "2023-08-29T02:53:11.377000",
              "content": "<p>Indeed, the dcm files in the same series have different shapes in the hidden test set.</p>\n<p>``</p>\n<h1>loop through the folders in the folder</h1>\n<p>count = 0<br>\nfor k in range(len(x)):<br>\n    patient_id = x.loc[x.index[k], \"patient_id\"]<br>\n    series_id = x.loc[x.index[k], \"series_id\"]</p>\n<pre><code>patient_folder = os.path.str(patient_id))\nseries_folder = os.path.str(series_id))\n= os.listdir(series_folder)\nprev_shape = None\nfor in     path = os.path.    =     arr = np.array(    = str(arr.    if prev_shape is not None:\n        assert == prev_shape\n    prev_shape = \n     += \nif ( &gt; ) not use_test:\n    </code></pre>\n<p>``</p>\n<p>Adding an assert statement that checks the shape throws an error during submission. This indeed shows that different .dcm files in the same series may contain different shapes. This also casts doubt on the integrity of the test data, it might be the case that slices belonging to different scans (series) are mixed together in the same series folder.</p>\n<p>This error interferes with pipelines that tries to make use of the whole 3D scan, for example TotalSegmentator as a preprocessing step. I think the best we could do is to loop through every slice in each series and decide which ones to use using a majority vote for the shape within that series, which should remove all the incorrect/corrupted ones.</p>\n<p>Of course, it would be nice if the host/Kaggle staff could run a simple for loop to check for these errors or to provide a sample notebook that loads every .dcm file in every series into a 3D npy array correctly, with a dummy submission.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2416360,
              "author_name": "Antonio Félix",
              "author_url": "",
              "post_date": "2023-08-31T01:50:13.940000",
              "content": "<p>Just to add a to the conversation, found this cool function in the code of dicom2nifti that is able to quickly check if the dicom is valid.</p>\n<pre><code> ():\n    \n    file_stream = (filename, )\n    file_stream.seek()\n    data = file_stream.read()\n    file_stream.close()\n     data == :\n         \n     dicom2nifti.settings.pydicom_read_force:\n        :\n            dicom_headers = pydicom.read_file(filename, defer_size=, stop_before_pixels=, force=)\n             dicom_headers   :\n                 \n        :\n            \n     \n</code></pre>\n<p>Source: <a href=\"https://github.com/icometrix/dicom2nifti/blob/main/dicom2nifti/common.py\" target=\"_blank\">https://github.com/icometrix/dicom2nifti/blob/main/dicom2nifti/common.py</a></p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2413479,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-08-29T00:26:39.023000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2411251,
      "author_name": "Jebastin Nadar",
      "author_url": "",
      "post_date": "2023-08-27T13:59:21.250000",
      "content": "<p>I did some basic LB probing to validate this. I am able to read all dicom files using dicomsdl, without any try-except block. Also all dicom files are 2D only, same as the train set.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2411295,
          "author_name": "Chau YH",
          "author_url": "",
          "post_date": "2023-08-27T14:28:53.960000",
          "content": "<p>Thanks for the information. So it might be a problem with pydicom, or there might be some weird bug in my code. Could you share a notebook on how to successfully load every .dcm in the hidden test set?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2411464,
              "author_name": "Jebastin Nadar",
              "author_url": "",
              "post_date": "2023-08-27T16:25:28.043000",
              "content": "<pre><code>src = \ndcm_dirs = os.listdir(src)\ndcm_dirs.sort(key =  x: (x.split()[]))\narr_list = []\n\n dcm_dir  dcm_dirs:\n    arr = dicomsdl.(src+dcm_dir).pixelData(storedvalue=)\n    arr_list.append(torch.from_numpy(arr.astype(np.float32)))\n</code></pre>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2411828,
              "author_name": "junseonglee11",
              "author_url": "",
              "post_date": "2023-08-27T21:58:37.063000",
              "content": "<p>Cool! Thank you so much!</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2412152,
              "author_name": "Chau YH",
              "author_url": "",
              "post_date": "2023-08-28T05:45:38.717000",
              "content": "<p>Thanks. I changed to dicomsdl and there are no problems. I have updated my post to reflect that.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2407194,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-08-25T00:00:57.050000",
      "content": "<p>i suggest some \"kaggle bug report prize\" like swag prize for report like this.<br>\nit takes a lot of time and effort to identify issues like this and more importantly, it helps others a lot!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2409580,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-08-26T10:24:38.067000",
          "content": "<p>Cool idea, and I think we really need such incentives. I have submitted too many failed submissions for this competition (more than 10 times). Frustrated…</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2414429,
      "author_name": "Sohier Dane",
      "author_url": "",
      "post_date": "2023-08-29T16:13:38.687000",
      "content": "<p>I will take a look but won't be able to follow up until later this week.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2414824,
          "author_name": "Chau YH",
          "author_url": "",
          "post_date": "2023-08-30T01:33:36.413000",
          "content": "<p>Thanks so much! I hope this problem can be resolved.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2416256,
          "author_name": "Sohier Dane",
          "author_url": "",
          "post_date": "2023-08-30T22:53:12.150000",
          "content": "<p>Please see <a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/435815\" target=\"_blank\">this post for my full response</a>; I did find one corrupt image.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2405972,
      "author_name": "junseonglee11",
      "author_url": "",
      "post_date": "2023-08-24T07:10:50.970000",
      "content": "<p>I also spent more than 10 hrs to debug the problem. After that, I just included a try-except statement to skip loading the problematic dcm files. Because I could not debug it directly with only 'Notebook threw exception' errors.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2398066,
      "author_name": "David Roberts",
      "author_url": "",
      "post_date": "2023-08-19T12:39:43.560000",
      "content": "<p>I have processed all the images in the dataset and have not found any bugs or corrupt images. I haven't checked the DICOM tags for SOP class, but I have not seen any multi-frame images. I suspect they're all standard CT Image SOP Classes.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2398410,
          "author_name": "Chau YH",
          "author_url": "",
          "post_date": "2023-08-19T17:06:27.607000",
          "content": "<p>Thanks for your reply. I have also successfully translated all .dcm files in the training folder to hdf5, but I could not load every .dcm file with pydicom in the hidden test folder, as indicated in my notebooks above, which crashes during submission. I suspect the error is raised when I use the .pixel_array attribute on a corrupt .hdf5 in the hidden test folder, but I can't confirm that as the error is redacted for submission notebooks.</p>\n<p>If you could load every .dcm in the hidden test folder, that means there's probably some weird bug in my code. If you have the time, could you share publicly a minimal working example of loading every .dcm file with pydicom's .pixel_array attribute in the hidden test set during submission (of course without any models or inference code, just make a dummy submission.csv)? Thanks for your help.</p>\n<p>I am thinking of a 3D model which uses every slice in a series for inference, but I can't load all the slices in the hidden test set.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2398848,
              "author_name": "David Roberts",
              "author_url": "",
              "post_date": "2023-08-20T02:38:46.483000",
              "content": "<p>Sorry <a href=\"https://www.kaggle.com/chauyh\" target=\"_blank\">@chauyh</a>, I misread your question. I have not made a submission, so I have not processed the test set. I can't confirm there are no corrupt images in the test set.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2399751,
              "author_name": "Chau YH",
              "author_url": "",
              "post_date": "2023-08-20T14:58:58.373000",
              "content": "<p>No problem. If you happen to be able to load all .dcm files in the test set, I hope that you can share a public notebook with us😁.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2396456,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-08-18T09:11:39.477000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2396453,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-08-18T09:09:45.680000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2396106": "# EDIT:\n@samfc10 pointed out that we should load the .dcm files with dicomsdl in the comments below. I have tried dicomsdl and it gives the same results as pydicom and has no errors on the hidden test set. However, @samfc10 pointed out that he is unable to stack the .dcm files together in the hidden test set, and it seems that the shapes of different .dcm files in the same series are inconsistent.\n\n@sohier found out that there is indeed a corrupted image in the test set. The code is able to run if the corrupted image is excluded, confirming that all shapes are consistent:\n```python\ncount = 0\nfor k in range(len(x)):\n    patient_id = x.loc[x.index[k], \"patient_id\"]\n    series_id = x.loc[x.index[k], \"series_id\"]\n\n    \n    patient_folder = os.path.join(folder, str(patient_id))\n    series_folder = os.path.join(patient_folder, str(series_id))\n    instances = os.listdir(series_folder)\n    prev_shape = None\n    for instance in instances:\n        if (int(patient_id) == 3124) and (int(series_id) == 5842) and (int(instance[:-4]) == 514):\n            continue\n        \n        path = os.path.join(series_folder, instance)\n        dicom_file = dicomsdl.open(path)\n        arr = np.array(dicom_file.pixelData(storedvalue=True))\n        assert len(arr.shape) == 2\n        shape = str(arr.shape)\n        if prev_shape is not None:\n            assert shape == prev_shape\n        prev_shape = shape\n        \n        count += 1\n    if (count > 100) and not use_test:\n        break\n```\n\n# Original post:\n\nDear Competition host/Kaggle Staff,\n\nThanks for hosting this awesome competition! I saw that the description of the competition has been updated to include a [Getting Started notebook](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/overview) to demonstrate how to train a model and do inference, which is great. However, I think there might be some corrupt .dcm in the test set which leads to random errors while submitting notebooks to the competition.\n\n# Fork of official inference notebook\nI forked the official inference notebook and removed the model parts from it. I ran two tests, one with STRIDE=1 and one with STRIDE unchanged, and the STRIDE=1 submission threw an error within 1hr while the other submission is successful.\n\n- [STRIDE=1](https://www.kaggle.com/code/chauyh/kerascv-starter-notebook-infer-fork-stride-1)\n- [STRIDE>1](https://www.kaggle.com/code/chauyh/kerascv-starter-notebook-infer-fork-stride-10)\n\n# Test with other notebooks\nI created a notebook myself which simply tries to load the .pixel_array using pydicom and gets the shape of the pixel array. The notebook which only loads 100 .dcm files is successful, while the one that loads all the .dcm files crashes within 1hr.\n\n- [100 .dcm](https://www.kaggle.com/code/chauyh/rsna-pixel-array-100)\n- [all .dcm](https://www.kaggle.com/code/chauyh/rsna-pixel-array-all)\n\nI have more (not so comprehensive) tests in my [previous discussion](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/431836).\n\n# How to fix this\nI think this is pretty solid evidence that there are some corrupt .dcm files in the test set. Or it could be some unknown bug in my implementation of loading .dicom files. Anyways, could the host/Kaggle Staff modify the official inference notebook to load all the .dcm files with STRIDE=1 instead of a partial list of .dcm files with STRIDE>1 with pydicom?\n\nCould the host/Kaggle staff also check whether the SOP Class UID of all .dcm files belong to CT Image Storage, not Enhanced CT Image Storage or other SOP Class UIDs? While I'm not a radiologist, I asked GPT and GPT told me that .dcm files belonging to Enhanced CT Image Storage stores 3D data in a single .dcm file, which may also conflict with the usual image loading pipeline.\n\nThis could check whether or not all .dcm files are fine, and moreover a reference for how to load every .dcm file in the test set into numpy arrays properly for new people who wish to join the competition (and for me too).\n\n# Other stuff\nCan the host/Kaggle staff change the description of the Data section to say that the visible test folder is not representative of the actual test folder, where the actual test folder contains multiple slices per series instead of a single slice? I know this is addressed officially by @sohier in a comment in the discussion forums, but saying that directly in the Data section would make it clearer for new competitors.\n\nThanks!\n-YH",
    "2412964": "> Use dicomsdl instead of pydicom to fix this issue.\n\nActually dicomsdl doesn't fix the issue. The core problem, i.e corrupted dicoms in test set still remains. Although dicomsdl doesn't raise any error when reading dicoms, there's some error when stacking the tensors to form a 3D tensor, which I'm unable to debug due to the limited error message during submission. Stacking works with pydicom and try-except block during dicom reading.\n\n@sohier Could you please take a look into this issue. It's been 10 days since the discussion has been opened, and there's no response from the hosts, or the kaggle staff. It's sad to see. A lot of competitor's time and submissions have been wasted in trying to debug this issue.",
    "2411251": "I did some basic LB probing to validate this. I am able to read all dicom files using dicomsdl, without any try-except block. Also all dicom files are 2D only, same as the train set.",
    "2407194": "i suggest some \"kaggle bug report prize\" like swag prize for report like this.\nit takes a lot of time and effort to identify issues like this and more importantly, it helps others a lot!",
    "2414429": "I will take a look but won't be able to follow up until later this week.",
    "2405972": "I also spent more than 10 hrs to debug the problem. After that, I just included a try-except statement to skip loading the problematic dcm files. Because I could not debug it directly with only 'Notebook threw exception' errors.",
    "2398066": "I have processed all the images in the dataset and have not found any bugs or corrupt images. I haven't checked the DICOM tags for SOP class, but I have not seen any multi-frame images. I suspect they're all standard CT Image SOP Classes.",
    "2396456": "",
    "2396453": ""
  }
}