{
  "id": 409439,
  "title": "Can someone demonstrate how to handle & play around with datasets of such massive scales through a Notebook? ",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/409439",
  "author_name": "Aryan Garg",
  "post_date": "2023-05-11T03:09:28.616000",
  "votes": 4,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Also, do you ever want to compute statistics for the dataset? (~450 GBs for this one ooof)</p>",
  "messages": [
    {
      "id": 2254478,
      "postDate": "2023-05-11T03:09:28.617Z",
      "content": "<p>Also, do you ever want to compute statistics for the dataset? (~450 GBs for this one ooof)</p>",
      "rawMarkdown": "Also, do you ever want to compute statistics for the dataset? (~450 GBs for this one ooof)",
      "votes": 3
    },
    {
      "id": 2255450,
      "postDate": "2023-05-11T18:07:38.747Z",
      "content": "<p>I wrote a pytorch dataloader that might interest you:<br>\n<a href=\"https://www.kaggle.com/code/thomasrochefort/pytorch-dataloader-example\" target=\"_blank\">https://www.kaggle.com/code/thomasrochefort/pytorch-dataloader-example</a></p>\n<p>It allows for fetching batches from the dataset without loading everything in memory.</p>",
      "rawMarkdown": "I wrote a pytorch dataloader that might interest you:\nhttps://www.kaggle.com/code/thomasrochefort/pytorch-dataloader-example\n\nIt allows for fetching batches from the dataset without loading everything in memory.",
      "votes": 1,
      "replies": [
        {
          "id": 2257102,
          "postDate": "2023-05-13T03:29:12.877Z",
          "content": "<p>Hey Thomas, thanks for the example. I ran your NB but the RAM is overloaded on a simple loop through the dataloader :(</p>\n<p>Is there any workaround?</p>",
          "rawMarkdown": "Hey Thomas, thanks for the example. I ran your NB but the RAM is overloaded on a simple loop through the dataloader :(\n\nIs there any workaround?",
          "replies": [
            {
              "id": 2259061,
              "postDate": "2023-05-14T16:58:20.663Z",
              "content": "<p>You have to load the data batchwise, which means you have to overwrite old batches. The dataloader should work fine in every NB</p>",
              "rawMarkdown": "You have to load the data batchwise, which means you have to overwrite old batches. The dataloader should work fine in every NB"
            },
            {
              "id": 2259149,
              "postDate": "2023-05-14T18:04:22.807Z",
              "content": "<p>I too was having the same issue. Uncommented the loop you wrote to load data, but its too large to handle</p>",
              "rawMarkdown": "I too was having the same issue. Uncommented the loop you wrote to load data, but its too large to handle"
            },
            {
              "id": 2261696,
              "postDate": "2023-05-16T14:02:00.097Z",
              "content": "<p>Hello! The dataloader is meant to be used in a pytorch training loop which will load batches and overwrite old ones. Of course if you keep in memory every batch the RAM will explode. Do you have a code snippet to show how you use it?</p>",
              "rawMarkdown": "Hello! The dataloader is meant to be used in a pytorch training loop which will load batches and overwrite old ones. Of course if you keep in memory every batch the RAM will explode. Do you have a code snippet to show how you use it?"
            },
            {
              "id": 2261887,
              "postDate": "2023-05-16T15:52:07.693Z",
              "content": "<p>I'm using it like this here: <a href=\"https://www.kaggle.com/code/aryangarg01/submission-deeplabv3-on-ash-images-for-contrails\" target=\"_blank\">DeeplabV3+ Submission NB</a></p>\n<p>Let me know if I'm doing it wrong or if there can be optimizations. Thanks in advance :)</p>",
              "rawMarkdown": "I'm using it like this here: [DeeplabV3+ Submission NB](https://www.kaggle.com/code/aryangarg01/submission-deeplabv3-on-ash-images-for-contrails)\n\nLet me know if I'm doing it wrong or if there can be optimizations. Thanks in advance :)"
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2255450,
      "author_name": "Thomas Rochefort-Beaudoin",
      "author_url": "",
      "post_date": "2023-05-11T18:07:38.747000",
      "content": "<p>I wrote a pytorch dataloader that might interest you:<br>\n<a href=\"https://www.kaggle.com/code/thomasrochefort/pytorch-dataloader-example\" target=\"_blank\">https://www.kaggle.com/code/thomasrochefort/pytorch-dataloader-example</a></p>\n<p>It allows for fetching batches from the dataset without loading everything in memory.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2257102,
          "author_name": "Aryan Garg",
          "author_url": "",
          "post_date": "2023-05-13T03:29:12.877000",
          "content": "<p>Hey Thomas, thanks for the example. I ran your NB but the RAM is overloaded on a simple loop through the dataloader :(</p>\n<p>Is there any workaround?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2259061,
              "author_name": "Patchef",
              "author_url": "",
              "post_date": "2023-05-14T16:58:20.663000",
              "content": "<p>You have to load the data batchwise, which means you have to overwrite old batches. The dataloader should work fine in every NB</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2259149,
              "author_name": "Proteek Chaudhuri",
              "author_url": "",
              "post_date": "2023-05-14T18:04:22.807000",
              "content": "<p>I too was having the same issue. Uncommented the loop you wrote to load data, but its too large to handle</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2261696,
              "author_name": "Thomas Rochefort-Beaudoin",
              "author_url": "",
              "post_date": "2023-05-16T14:02:00.097000",
              "content": "<p>Hello! The dataloader is meant to be used in a pytorch training loop which will load batches and overwrite old ones. Of course if you keep in memory every batch the RAM will explode. Do you have a code snippet to show how you use it?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2261887,
              "author_name": "Aryan Garg",
              "author_url": "",
              "post_date": "2023-05-16T15:52:07.693000",
              "content": "<p>I'm using it like this here: <a href=\"https://www.kaggle.com/code/aryangarg01/submission-deeplabv3-on-ash-images-for-contrails\" target=\"_blank\">DeeplabV3+ Submission NB</a></p>\n<p>Let me know if I'm doing it wrong or if there can be optimizations. Thanks in advance :)</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2254478": "Also, do you ever want to compute statistics for the dataset? (~450 GBs for this one ooof)",
    "2255450": "I wrote a pytorch dataloader that might interest you:\nhttps://www.kaggle.com/code/thomasrochefort/pytorch-dataloader-example\n\nIt allows for fetching batches from the dataset without loading everything in memory."
  }
}