{
  "id": 609044,
  "title": "File loading times",
  "url": "/competitions/rsna-intracranial-aneurysm-detection/discussion/609044",
  "author_name": "thomas rost",
  "post_date": "2025-09-23T11:12:08.454000",
  "votes": 2,
  "comment_count": 14,
  "views": 0,
  "content": "<p>How do others handle file loading? On my machine, it takes 1.5 hours to just load all the dicom slices without any processing. This means, a lot of the gpu-time on kaggle would just be used up by this. What is a reasonable approach here? I guess we don't have enough hd-space in kaggle notebooks to cache the whole dataset?<br>\nAny help is appreciated!</p>",
  "messages": [
    {
      "id": 3293259,
      "postDate": "2025-09-23T11:34:35.570Z",
      "content": "<p>For training, preprocess dicoms in a faster acces format. Such raw npy of ROIs or 32x384x384 images, depends on your approach. For submission, optimize reading speed as much as you can with mp tecniques.</p>",
      "rawMarkdown": "For training, preprocess dicoms in a faster acces format. Such raw npy of ROIs or 32x384x384 images, depends on your approach. For submission, optimize reading speed as much as you can with mp tecniques.",
      "votes": 1,
      "replies": [
        {
          "id": 3293260,
          "postDate": "2025-09-23T11:38:19.343Z",
          "content": "<p>ok, thanks. I'll try that. </p>",
          "rawMarkdown": "ok, thanks. I'll try that. "
        }
      ]
    },
    {
      "id": 3293318,
      "postDate": "2025-09-23T14:52:16.347Z",
      "content": "<p>I've preprocessed everything and converted it to npz datasets. Almost no time spent loading , and I can train segmentation in under ~30s per epoch, and classification (entire dataset) ~3 mins per epoch.</p>",
      "rawMarkdown": "I've preprocessed everything and converted it to npz datasets. Almost no time spent loading , and I can train segmentation in under ~30s per epoch, and classification (entire dataset) ~3 mins per epoch.",
      "votes": 2,
      "replies": [
        {
          "id": 3293440,
          "postDate": "2025-09-23T19:56:13.467Z",
          "content": "<p>Ah, interesting. So you work on entire volumes? Thanks for your reply.</p>",
          "rawMarkdown": "Ah, interesting. So you work on entire volumes? Thanks for your reply.",
          "votes": 1
        },
        {
          "id": 3294721,
          "postDate": "2025-09-26T15:29:57.377Z",
          "content": "<p>Do you then save the npz datasets on /kaggle/working? Does it fit?</p>",
          "rawMarkdown": "Do you then save the npz datasets on /kaggle/working? Does it fit?",
          "replies": [
            {
              "id": 3294731,
              "postDate": "2025-09-26T15:45:42.840Z",
              "content": "<p>I save them as uint8 to save some memory. It fits in a single notebook for 128x128x128 volumes. Anything larger,  and I have to split it into multiple runs, I do 500/1000 per notebook depending on size. I am now renting a GPU VM since it saves quite a bit of time doing all of this, and processing larger datasets there, then uploading them to Kaggle as a single dataset.<br>\nMy dataset with float volumes + segmentation masks (128x128x128) is over 90GB.</p>",
              "rawMarkdown": "I save them as uint8 to save some memory. It fits in a single notebook for 128x128x128 volumes. Anything larger,  and I have to split it into multiple runs, I do 500/1000 per notebook depending on size. I am now renting a GPU VM since it saves quite a bit of time doing all of this, and processing larger datasets there, then uploading them to Kaggle as a single dataset.\nMy dataset with float volumes + segmentation masks (128x128x128) is over 90GB.",
              "votes": 1
            },
            {
              "id": 3294733,
              "postDate": "2025-09-26T15:47:11.217Z",
              "content": "<p>Also , I missed your first comment - in reply to that, I save them as entire volumes, but don't exactly use them entirely. In my experience, 3D does not work at all with whatever I have tried!</p>",
              "rawMarkdown": "Also , I missed your first comment - in reply to that, I save them as entire volumes, but don't exactly use them entirely. In my experience, 3D does not work at all with whatever I have tried!",
              "votes": 1
            },
            {
              "id": 3294742,
              "postDate": "2025-09-26T16:06:29.963Z",
              "content": "<p>Ok, thanks for the detailed reply! I would like to avoid renting a VM. But I'll try to create the volume dataset locally and then upload it to kaggle as a dataset. I'm not sure what the size restrictions are for this.</p>",
              "rawMarkdown": "Ok, thanks for the detailed reply! I would like to avoid renting a VM. But I'll try to create the volume dataset locally and then upload it to kaggle as a dataset. I'm not sure what the size restrictions are for this.",
              "votes": 1
            },
            {
              "id": 3294753,
              "postDate": "2025-09-26T16:24:44.190Z",
              "content": "<p>No problem! Don't forget to zip your dataset before uploading to Kaggle, else it usually fails. </p>",
              "rawMarkdown": "No problem! Don't forget to zip your dataset before uploading to Kaggle, else it usually fails. "
            }
          ]
        }
      ]
    },
    {
      "id": 3293251,
      "postDate": "2025-09-23T11:12:08.453Z",
      "content": "<p>How do others handle file loading? On my machine, it takes 1.5 hours to just load all the dicom slices without any processing. This means, a lot of the gpu-time on kaggle would just be used up by this. What is a reasonable approach here? I guess we don't have enough hd-space in kaggle notebooks to cache the whole dataset?<br>\nAny help is appreciated!</p>",
      "rawMarkdown": "How do others handle file loading? On my machine, it takes 1.5 hours to just load all the dicom slices without any processing. This means, a lot of the gpu-time on kaggle would just be used up by this. What is a reasonable approach here? I guess we don't have enough hd-space in kaggle notebooks to cache the whole dataset?\nAny help is appreciated!",
      "votes": 2
    },
    {
      "id": 3293688,
      "postDate": "2025-09-24T12:44:46.303Z",
      "content": "<p>I have come accross the package <code>dicomsdl</code> which is available through <code>pip</code> but not installed in the standard container so it is not availlable at inference time. I have made quick demo notebook here:<br>\n<a href=\"https://www.kaggle.com/code/thalro/test-read-speed\" target=\"_blank\">https://www.kaggle.com/code/thalro/test-read-speed</a><br>\nIt seems to do more or less the same as <code>pydicom</code> but it reads the pixel data 5-10 times faster. <a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a>, I'm not sure who to ask, but would it be possible to install this in the standard environment?</p>",
      "rawMarkdown": "I have come accross the package `dicomsdl` which is available through `pip` but not installed in the standard container so it is not availlable at inference time. I have made quick demo notebook here:\nhttps://www.kaggle.com/code/thalro/test-read-speed\nIt seems to do more or less the same as `pydicom` but it reads the pixel data 5-10 times faster. @ryanholbrook, I'm not sure who to ask, but would it be possible to install this in the standard environment?",
      "replies": [
        {
          "id": 3293698,
          "postDate": "2025-09-24T12:59:53.383Z",
          "content": "<p>You can use kaggle \"scripts\" to install whatever you like during your submission.</p>\n<p>Simply create a notebook with internet access which performs \"!pip install dicomsdl\" and save it as a script.</p>\n<p>Then add this script to the environment of your inference notebook (\"add dataset\") and you'll have the library installed without internet access.</p>",
          "rawMarkdown": "You can use kaggle \"scripts\" to install whatever you like during your submission.\n\nSimply create a notebook with internet access which performs \"!pip install dicomsdl\" and save it as a script.\n\nThen add this script to the environment of your inference notebook (\"add dataset\") and you'll have the library installed without internet access.",
          "votes": 4,
          "replies": [
            {
              "id": 3293742,
              "postDate": "2025-09-24T14:32:39.883Z",
              "content": "<p>Ah, great. Thanks, I didn‘t know that!</p>",
              "rawMarkdown": "Ah, great. Thanks, I didn‘t know that!"
            }
          ]
        },
        {
          "id": 3293713,
          "postDate": "2025-09-24T13:43:15.973Z",
          "content": "<p>In addition to the other solutions mentioned, you can also use the notebook editor's dependency manager. From the editor, do <code>AddOns-&gt;Install Dependencies</code>. You have to turn Internet on, install the dependencies, then turn Internet off again before you submit, so it's a little finicky, but it does keep it all in one place.</p>",
          "rawMarkdown": "In addition to the other solutions mentioned, you can also use the notebook editor's dependency manager. From the editor, do `AddOns->Install Dependencies`. You have to turn Internet on, install the dependencies, then turn Internet off again before you submit, so it's a little finicky, but it does keep it all in one place.",
          "votes": 5,
          "replies": [
            {
              "id": 3293743,
              "postDate": "2025-09-24T14:33:25.140Z",
              "content": "<p>Ok, thanks for the quick reply. I wasn‘t aware of that.</p>",
              "rawMarkdown": "Ok, thanks for the quick reply. I wasn‘t aware of that."
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3293259,
      "author_name": "Ángel Jacinto Sánchez Ruiz",
      "author_url": "",
      "post_date": "2025-09-23T11:34:35.570000",
      "content": "<p>For training, preprocess dicoms in a faster acces format. Such raw npy of ROIs or 32x384x384 images, depends on your approach. For submission, optimize reading speed as much as you can with mp tecniques.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3293260,
          "author_name": "thomas rost",
          "author_url": "",
          "post_date": "2025-09-23T11:38:19.343000",
          "content": "<p>ok, thanks. I'll try that. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3293318,
      "author_name": "Satwik",
      "author_url": "",
      "post_date": "2025-09-23T14:52:16.347000",
      "content": "<p>I've preprocessed everything and converted it to npz datasets. Almost no time spent loading , and I can train segmentation in under ~30s per epoch, and classification (entire dataset) ~3 mins per epoch.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3293440,
          "author_name": "thomas rost",
          "author_url": "",
          "post_date": "2025-09-23T19:56:13.467000",
          "content": "<p>Ah, interesting. So you work on entire volumes? Thanks for your reply.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 3294721,
          "author_name": "thomas rost",
          "author_url": "",
          "post_date": "2025-09-26T15:29:57.377000",
          "content": "<p>Do you then save the npz datasets on /kaggle/working? Does it fit?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3294731,
              "author_name": "Satwik",
              "author_url": "",
              "post_date": "2025-09-26T15:45:42.840000",
              "content": "<p>I save them as uint8 to save some memory. It fits in a single notebook for 128x128x128 volumes. Anything larger,  and I have to split it into multiple runs, I do 500/1000 per notebook depending on size. I am now renting a GPU VM since it saves quite a bit of time doing all of this, and processing larger datasets there, then uploading them to Kaggle as a single dataset.<br>\nMy dataset with float volumes + segmentation masks (128x128x128) is over 90GB.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3294733,
              "author_name": "Satwik",
              "author_url": "",
              "post_date": "2025-09-26T15:47:11.217000",
              "content": "<p>Also , I missed your first comment - in reply to that, I save them as entire volumes, but don't exactly use them entirely. In my experience, 3D does not work at all with whatever I have tried!</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3294742,
              "author_name": "thomas rost",
              "author_url": "",
              "post_date": "2025-09-26T16:06:29.963000",
              "content": "<p>Ok, thanks for the detailed reply! I would like to avoid renting a VM. But I'll try to create the volume dataset locally and then upload it to kaggle as a dataset. I'm not sure what the size restrictions are for this.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3294753,
              "author_name": "Satwik",
              "author_url": "",
              "post_date": "2025-09-26T16:24:44.190000",
              "content": "<p>No problem! Don't forget to zip your dataset before uploading to Kaggle, else it usually fails. </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3293688,
      "author_name": "thomas rost",
      "author_url": "",
      "post_date": "2025-09-24T12:44:46.303000",
      "content": "<p>I have come accross the package <code>dicomsdl</code> which is available through <code>pip</code> but not installed in the standard container so it is not availlable at inference time. I have made quick demo notebook here:<br>\n<a href=\"https://www.kaggle.com/code/thalro/test-read-speed\" target=\"_blank\">https://www.kaggle.com/code/thalro/test-read-speed</a><br>\nIt seems to do more or less the same as <code>pydicom</code> but it reads the pixel data 5-10 times faster. <a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a>, I'm not sure who to ask, but would it be possible to install this in the standard environment?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3293698,
          "author_name": "Optimo",
          "author_url": "",
          "post_date": "2025-09-24T12:59:53.383000",
          "content": "<p>You can use kaggle \"scripts\" to install whatever you like during your submission.</p>\n<p>Simply create a notebook with internet access which performs \"!pip install dicomsdl\" and save it as a script.</p>\n<p>Then add this script to the environment of your inference notebook (\"add dataset\") and you'll have the library installed without internet access.</p>",
          "votes": 4,
          "replies": [
            {
              "id": 3293742,
              "author_name": "thomas rost",
              "author_url": "",
              "post_date": "2025-09-24T14:32:39.883000",
              "content": "<p>Ah, great. Thanks, I didn‘t know that!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 3293713,
          "author_name": "Ryan Holbrook",
          "author_url": "",
          "post_date": "2025-09-24T13:43:15.973000",
          "content": "<p>In addition to the other solutions mentioned, you can also use the notebook editor's dependency manager. From the editor, do <code>AddOns-&gt;Install Dependencies</code>. You have to turn Internet on, install the dependencies, then turn Internet off again before you submit, so it's a little finicky, but it does keep it all in one place.</p>",
          "votes": 5,
          "replies": [
            {
              "id": 3293743,
              "author_name": "thomas rost",
              "author_url": "",
              "post_date": "2025-09-24T14:33:25.140000",
              "content": "<p>Ok, thanks for the quick reply. I wasn‘t aware of that.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3293259": "For training, preprocess dicoms in a faster acces format. Such raw npy of ROIs or 32x384x384 images, depends on your approach. For submission, optimize reading speed as much as you can with mp tecniques.",
    "3293318": "I've preprocessed everything and converted it to npz datasets. Almost no time spent loading , and I can train segmentation in under ~30s per epoch, and classification (entire dataset) ~3 mins per epoch.",
    "3293251": "How do others handle file loading? On my machine, it takes 1.5 hours to just load all the dicom slices without any processing. This means, a lot of the gpu-time on kaggle would just be used up by this. What is a reasonable approach here? I guess we don't have enough hd-space in kaggle notebooks to cache the whole dataset?\nAny help is appreciated!",
    "3293688": "I have come accross the package `dicomsdl` which is available through `pip` but not installed in the standard container so it is not availlable at inference time. I have made quick demo notebook here:\nhttps://www.kaggle.com/code/thalro/test-read-speed\nIt seems to do more or less the same as `pydicom` but it reads the pixel data 5-10 times faster. @ryanholbrook, I'm not sure who to ask, but would it be possible to install this in the standard environment?"
  }
}