{
  "id": 607299,
  "title": "How do you handle very large datasets on cloud GPUs?",
  "url": "/competitions/rsna-intracranial-aneurysm-detection/discussion/607299",
  "author_name": "Naoki Nomurrr",
  "post_date": "2025-09-13T08:54:36.053000",
  "votes": 0,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I'm working with a very large dataset (e.g., &gt;200GB medical images).  <br>\nI've tried zipping and splitting files, but upload limits and transfer time are still a bottleneck.  <br>\nHow do you manage or process such large data in Kaggle Notebooks or other cloud GPU environments?  <br>\nAny recommended workflows or tools (e.g., GCS, AWS S3, DVC, Hugging Face Hub) would be appreciated.</p>",
  "messages": [
    {
      "id": 3289980,
      "postDate": "2025-09-16T22:39:34.747Z",
      "content": "<p>I use webdataset (github.com/webdataset/webdataset). It has a learning curve, but loads of features (streaming, distributed training, shuffling, etc.). I used it for a prior project making a video tokenizer, and was able to carry over a lot of the training codebase.</p>\n<p>I losslessly converted the dataset to about 132gb of webdataset shards and uploaded them to HF. From there, I can stream them to a cloud VM and easily train, or download the dataset prior to training if the networking is slow.</p>",
      "rawMarkdown": "I use webdataset (github.com/webdataset/webdataset). It has a learning curve, but loads of features (streaming, distributed training, shuffling, etc.). I used it for a prior project making a video tokenizer, and was able to carry over a lot of the training codebase.\n\nI losslessly converted the dataset to about 132gb of webdataset shards and uploaded them to HF. From there, I can stream them to a cloud VM and easily train, or download the dataset prior to training if the networking is slow.",
      "votes": 1,
      "replies": [
        {
          "id": 3290584,
          "postDate": "2025-09-18T03:15:44.013Z",
          "content": "<p>Really appreciate you outlining your workflow with webdataset and HF.<br>\nThe tip about losslessly converting the dataset into shards and streaming to a cloud VM is extremely helpful.<br>\nI’ll definitely explore this approach for my own project—thank you!</p>",
          "rawMarkdown": "Really appreciate you outlining your workflow with webdataset and HF.\nThe tip about losslessly converting the dataset into shards and streaming to a cloud VM is extremely helpful.\nI’ll definitely explore this approach for my own project—thank you!\n"
        }
      ]
    },
    {
      "id": 3288197,
      "postDate": "2025-09-13T10:46:40.820Z",
      "content": "<p>Use TPU <a href=\"https://www.kaggle.com/naokinomurrr\" target=\"_blank\">@naokinomurrr</a> </p>",
      "rawMarkdown": "Use TPU @naokinomurrr ",
      "votes": -1
    },
    {
      "id": 3288150,
      "postDate": "2025-09-13T08:54:36.053Z",
      "content": "<p>I'm working with a very large dataset (e.g., &gt;200GB medical images).  <br>\nI've tried zipping and splitting files, but upload limits and transfer time are still a bottleneck.  <br>\nHow do you manage or process such large data in Kaggle Notebooks or other cloud GPU environments?  <br>\nAny recommended workflows or tools (e.g., GCS, AWS S3, DVC, Hugging Face Hub) would be appreciated.</p>",
      "rawMarkdown": "I'm working with a very large dataset (e.g., >200GB medical images).  \nI've tried zipping and splitting files, but upload limits and transfer time are still a bottleneck.  \nHow do you manage or process such large data in Kaggle Notebooks or other cloud GPU environments?  \nAny recommended workflows or tools (e.g., GCS, AWS S3, DVC, Hugging Face Hub) would be appreciated."
    }
  ],
  "comments": [
    {
      "id": 3289980,
      "author_name": "NilanE",
      "author_url": "",
      "post_date": "2025-09-16T22:39:34.747000",
      "content": "<p>I use webdataset (github.com/webdataset/webdataset). It has a learning curve, but loads of features (streaming, distributed training, shuffling, etc.). I used it for a prior project making a video tokenizer, and was able to carry over a lot of the training codebase.</p>\n<p>I losslessly converted the dataset to about 132gb of webdataset shards and uploaded them to HF. From there, I can stream them to a cloud VM and easily train, or download the dataset prior to training if the networking is slow.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3290584,
          "author_name": "Naoki Nomurrr",
          "author_url": "",
          "post_date": "2025-09-18T03:15:44.013000",
          "content": "<p>Really appreciate you outlining your workflow with webdataset and HF.<br>\nThe tip about losslessly converting the dataset into shards and streaming to a cloud VM is extremely helpful.<br>\nI’ll definitely explore this approach for my own project—thank you!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3288197,
      "author_name": "Ravi Ramakrishnan",
      "author_url": "",
      "post_date": "2025-09-13T10:46:40.820000",
      "content": "<p>Use TPU <a href=\"https://www.kaggle.com/naokinomurrr\" target=\"_blank\">@naokinomurrr</a> </p>",
      "votes": -1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3289980": "I use webdataset (github.com/webdataset/webdataset). It has a learning curve, but loads of features (streaming, distributed training, shuffling, etc.). I used it for a prior project making a video tokenizer, and was able to carry over a lot of the training codebase.\n\nI losslessly converted the dataset to about 132gb of webdataset shards and uploaded them to HF. From there, I can stream them to a cloud VM and easily train, or download the dataset prior to training if the networking is slow.",
    "3288197": "Use TPU @naokinomurrr ",
    "3288150": "I'm working with a very large dataset (e.g., >200GB medical images).  \nI've tried zipping and splitting files, but upload limits and transfer time are still a bottleneck.  \nHow do you manage or process such large data in Kaggle Notebooks or other cloud GPU environments?  \nAny recommended workflows or tools (e.g., GCS, AWS S3, DVC, Hugging Face Hub) would be appreciated."
  }
}