{
  "id": 187457,
  "title": "Training on Colab - Lack of Disk Space",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/187457",
  "author_name": "Nitin Datta",
  "post_date": "2020-09-29T03:49:48.681000",
  "votes": 4,
  "comment_count": 22,
  "views": 0,
  "content": "<p>Hi everyone, <br>\nI am trying to import the 52GB dataset made by <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> but I run out of disk space.. <br>\nAny suggestions on what can be done so that we can download the data using Kaggle API and train on Colab.</p>\n<p>Any help will be appreciated. <br>\nThank you!!!</p>",
  "messages": [
    {
      "id": 1030856,
      "postDate": "2020-09-29T03:49:48.680Z",
      "content": "<p>Hi everyone, <br>\nI am trying to import the 52GB dataset made by <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> but I run out of disk space.. <br>\nAny suggestions on what can be done so that we can download the data using Kaggle API and train on Colab.</p>\n<p>Any help will be appreciated. <br>\nThank you!!!</p>",
      "rawMarkdown": "Hi everyone, \nI am trying to import the 52GB dataset made by @vaillant but I run out of disk space.. \nAny suggestions on what can be done so that we can download the data using Kaggle API and train on Colab.\n\nAny help will be appreciated. \nThank you!!!",
      "votes": 4
    },
    {
      "id": 1031468,
      "postDate": "2020-09-29T13:27:57.253Z",
      "content": "<p>Use TPU (colab Pro)</p>\n<p>You will have much more disk space </p>",
      "rawMarkdown": "Use TPU (colab Pro)\n\nYou will have much more disk space ",
      "votes": 3,
      "replies": [
        {
          "id": 1031612,
          "postDate": "2020-09-29T15:02:05.013Z",
          "content": "<p>Thank you for the advice <br>\nBut here in India I dont think colab pro is available</p>",
          "rawMarkdown": "Thank you for the advice \nBut here in India I dont think colab pro is available",
          "votes": 1
        },
        {
          "id": 1053785,
          "postDate": "2020-10-19T11:10:59.977Z",
          "content": "<p><a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> Don't mind me popping the question, how should I load the data in colab pro? I can't seem to do it…</p>",
          "rawMarkdown": "@serigne Don't mind me popping the question, how should I load the data in colab pro? I can't seem to do it...",
          "votes": 1
        },
        {
          "id": 1054329,
          "postDate": "2020-10-19T20:20:12.857Z",
          "content": "<p><a href=\"https://www.kaggle.com/reighns\" target=\"_blank\">@reighns</a> I am using TPU, so I just get the GCS data path from Kaggle, and paste it into a Colab notebook. You can also use the Kaggle API to download the dataset.</p>",
          "rawMarkdown": "@reighns I am using TPU, so I just get the GCS data path from Kaggle, and paste it into a Colab notebook. You can also use the Kaggle API to download the dataset.",
          "votes": 1
        },
        {
          "id": 1054337,
          "postDate": "2020-10-19T20:31:16.997Z",
          "content": "<p><a href=\"https://www.kaggle.com/reighns\" target=\"_blank\">@reighns</a>  You can follow <a href=\"https://medium.com/analytics-vidhya/how-to-fetch-kaggle-datasets-into-google-colab-ea682569851a\" target=\"_blank\">this</a> for instance </p>\n<p><a href=\"https://www.kaggle.com/stanleyjzheng\" target=\"_blank\">@stanleyjzheng</a> I guess you're using TF ?  </p>\n<p>For Pytorch, it's a bit trickier to use GCS paths ^^</p>",
          "rawMarkdown": "@reighns  You can follow [this](https://medium.com/analytics-vidhya/how-to-fetch-kaggle-datasets-into-google-colab-ea682569851a) for instance \n\n @stanleyjzheng I guess you're using TF ?  \n\nFor Pytorch, it's a bit trickier to use GCS paths ^^",
          "votes": 3
        },
        {
          "id": 1054338,
          "postDate": "2020-10-19T20:33:25.490Z",
          "content": "<p>Yes, tf and tfrecords</p>",
          "rawMarkdown": "Yes, tf and tfrecords",
          "votes": 1
        },
        {
          "id": 1054341,
          "postDate": "2020-10-19T20:36:45.220Z",
          "content": "<p>tfrecords make life easier .  But I love too much Pytorch to turn back to TF ^^</p>",
          "rawMarkdown": "tfrecords make life easier .  But I love too much Pytorch to turn back to TF ^^",
          "votes": 2
        },
        {
          "id": 1054570,
          "postDate": "2020-10-20T02:46:09.070Z",
          "content": "<p><a href=\"https://www.kaggle.com/stanleyjzheng\" target=\"_blank\">@stanleyjzheng</a> Thanks a lot bro! </p>",
          "rawMarkdown": "@stanleyjzheng Thanks a lot bro! ",
          "votes": 1
        },
        {
          "id": 1054572,
          "postDate": "2020-10-20T02:46:23.720Z",
          "content": "<p><a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> Thanks, my man!</p>",
          "rawMarkdown": "@serigne Thanks, my man!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1052100,
      "postDate": "2020-10-17T10:04:18.327Z",
      "content": "<p>One can also try to read directly from the zip file , if you only have space for the zip file, using a mounted drive, doesn't increase the disk space.</p>\n<pre><code>!mkdir yyy\n!fuse-zip xxx.zip /yyy/\n</code></pre>",
      "rawMarkdown": "One can also try to read directly from the zip file , if you only have space for the zip file, using a mounted drive, doesn't increase the disk space.\n\n```\n!mkdir yyy\n!fuse-zip xxx.zip /yyy/\n```",
      "votes": 1
    },
    {
      "id": 1039852,
      "postDate": "2020-10-06T20:43:18.413Z",
      "content": "<p>I have a very vague solution that might work:<br>\nif you examine the disk of colab TPU notebook, you get following output for !df -h command.</p>\n<p>Filesystem      Size  Used Avail Use% Mounted on<br>\noverlay         108G   31G   73G  30% /<br>\ntmpfs            64M     0   64M   0% /dev<br>\ntmpfs           6.4G     0  6.4G   0% /sys/fs/cgroup<br>\nshm             5.9G     0  5.9G   0% /dev/shm<br>\ntmpfs           6.4G   12K  6.4G   1% /var/colab<br>\n/dev/sda1       114G   32G   83G  28% /etc/hosts<br>\ntmpfs           6.4G     0  6.4G   0% /proc/acpi<br>\ntmpfs           6.4G     0  6.4G   0% /proc/scsi<br>\ntmpfs           6.4G     0  6.4G   0% /sys/firmware</p>\n<p>here you can use upto 73 GB of disk<br>\nmy solution is create a new dataset split it into 3 parts about 17.6 GB each. Lets say data files are part1.zip, part2.zip and part3.zip. Now download each file, unzip it and delete it one by one. maximum peak storage you will require in this process is 17.6 * 4 = 70.4 GB that is less than available storage. But issue here is it waste lots of time. </p>",
      "rawMarkdown": "I have a very vague solution that might work:\nif you examine the disk of colab TPU notebook, you get following output for !df -h command.\n\nFilesystem      Size  Used Avail Use% Mounted on\noverlay         108G   31G   73G  30% /\ntmpfs            64M     0   64M   0% /dev\ntmpfs           6.4G     0  6.4G   0% /sys/fs/cgroup\nshm             5.9G     0  5.9G   0% /dev/shm\ntmpfs           6.4G   12K  6.4G   1% /var/colab\n/dev/sda1       114G   32G   83G  28% /etc/hosts\ntmpfs           6.4G     0  6.4G   0% /proc/acpi\ntmpfs           6.4G     0  6.4G   0% /proc/scsi\ntmpfs           6.4G     0  6.4G   0% /sys/firmware\n\nhere you can use upto 73 GB of disk\nmy solution is create a new dataset split it into 3 parts about 17.6 GB each. Lets say data files are part1.zip, part2.zip and part3.zip. Now download each file, unzip it and delete it one by one. maximum peak storage you will require in this process is 17.6 * 4 = 70.4 GB that is less than available storage. But issue here is it waste lots of time. \n",
      "votes": 1
    },
    {
      "id": 1034880,
      "postDate": "2020-10-02T09:06:05.200Z",
      "content": "<p>You can attach your google-drive</p>",
      "rawMarkdown": "You can attach your google-drive",
      "votes": 1,
      "replies": [
        {
          "id": 1034884,
          "postDate": "2020-10-02T09:07:20.577Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/leossb\" target=\"_blank\">@leossb</a> , <br>\nBut the problem is that data is 52GB which cannot be stored on drive unless you buy space</p>",
          "rawMarkdown": "Thanks @leossb , \nBut the problem is that data is 52GB which cannot be stored on drive unless you buy space",
          "votes": 1
        },
        {
          "id": 1034892,
          "postDate": "2020-10-02T09:12:23.163Z",
          "content": "<p>google-drive space is very cheap, you can buy 100Gb by $2/month</p>",
          "rawMarkdown": "google-drive space is very cheap, you can buy 100Gb by $2/month",
          "votes": 1
        },
        {
          "id": 1035213,
          "postDate": "2020-10-02T14:39:41.900Z",
          "content": "<p>Laoding huge dataset straight from Drive to your model is a very bad  idea.<br>\nNot only it wil be terribly slow but you will reach quickly the daily quota of data you can load. </p>\n<p>I have 2 To of Drive Premium Storage but I will never use it to load such dataset. </p>\n<p>IMHO, TPU + kaggle API to download the data in colab Disk Space is much better idea. </p>",
          "rawMarkdown": "Laoding huge dataset straight from Drive to your model is a very bad  idea.\nNot only it wil be terribly slow but you will reach quickly the daily quota of data you can load. \n\nI have 2 To of Drive Premium Storage but I will never use it to load such dataset. \n\nIMHO, TPU + kaggle API to download the data in colab Disk Space is much better idea. \n\n",
          "votes": 6
        },
        {
          "id": 1035217,
          "postDate": "2020-10-02T14:42:00.763Z",
          "content": "<p>Yeah I tried using Kaggle API to load the 52 GB data but in free version the dataset is too  large when compared to disk space provided</p>",
          "rawMarkdown": "Yeah I tried using Kaggle API to load the 52 GB data but in free version the dataset is too  large when compared to disk space provided\n"
        }
      ]
    },
    {
      "id": 1031950,
      "postDate": "2020-09-29T19:46:06.423Z",
      "content": "<p>Download it from GCS. I am downloading 72GB with this</p>",
      "rawMarkdown": "Download it from GCS. I am downloading 72GB with this",
      "votes": 2,
      "replies": [
        {
          "id": 1032146,
          "postDate": "2020-09-30T02:02:00.643Z",
          "content": "<p>Yes <a href=\"https://www.kaggle.com/doanquanvietnamca\" target=\"_blank\">@doanquanvietnamca</a> ,<br>\nThat's one way or we can create a 128x128 jpegs dataset… </p>",
          "rawMarkdown": "Yes @doanquanvietnamca ,\nThat's one way or we can create a 128x128 jpegs dataset... "
        }
      ]
    },
    {
      "id": 1031143,
      "postDate": "2020-09-29T09:07:06.347Z",
      "content": "<p>good job ! good</p>",
      "rawMarkdown": "good job ! good",
      "votes": -11,
      "replies": [
        {
          "id": 1031225,
          "postDate": "2020-09-29T10:21:44.903Z",
          "content": "<p>I do not get you</p>",
          "rawMarkdown": "I do not get you",
          "votes": 1
        }
      ]
    },
    {
      "id": 1035848,
      "postDate": "2020-10-03T06:39:33.140Z",
      "content": "<p>Colab has the option to import datsets from google drive. IF you could create a drive and then import that to your colab notebook. That should help. </p>",
      "rawMarkdown": "Colab has the option to import datsets from google drive. IF you could create a drive and then import that to your colab notebook. That should help. ",
      "replies": [
        {
          "id": 1035853,
          "postDate": "2020-10-03T06:43:35.870Z",
          "content": "<p>Already tried it :) <br>\nBut the dataset is 52GB and unless you pay you cannot upload the whole data there </p>",
          "rawMarkdown": "Already tried it :) \nBut the dataset is 52GB and unless you pay you cannot upload the whole data there "
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1031468,
      "author_name": "Serigne ",
      "author_url": "",
      "post_date": "2020-09-29T13:27:57.253000",
      "content": "<p>Use TPU (colab Pro)</p>\n<p>You will have much more disk space </p>",
      "votes": 3,
      "replies": [
        {
          "id": 1031612,
          "author_name": "Nitin Datta",
          "author_url": "",
          "post_date": "2020-09-29T15:02:05.013000",
          "content": "<p>Thank you for the advice <br>\nBut here in India I dont think colab pro is available</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1053785,
          "author_name": "gao-hongnan",
          "author_url": "",
          "post_date": "2020-10-19T11:10:59.977000",
          "content": "<p><a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> Don't mind me popping the question, how should I load the data in colab pro? I can't seem to do it…</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1054329,
          "author_name": "Stanley Zheng",
          "author_url": "",
          "post_date": "2020-10-19T20:20:12.857000",
          "content": "<p><a href=\"https://www.kaggle.com/reighns\" target=\"_blank\">@reighns</a> I am using TPU, so I just get the GCS data path from Kaggle, and paste it into a Colab notebook. You can also use the Kaggle API to download the dataset.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1054337,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2020-10-19T20:31:16.997000",
          "content": "<p><a href=\"https://www.kaggle.com/reighns\" target=\"_blank\">@reighns</a>  You can follow <a href=\"https://medium.com/analytics-vidhya/how-to-fetch-kaggle-datasets-into-google-colab-ea682569851a\" target=\"_blank\">this</a> for instance </p>\n<p><a href=\"https://www.kaggle.com/stanleyjzheng\" target=\"_blank\">@stanleyjzheng</a> I guess you're using TF ?  </p>\n<p>For Pytorch, it's a bit trickier to use GCS paths ^^</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1054338,
          "author_name": "Stanley Zheng",
          "author_url": "",
          "post_date": "2020-10-19T20:33:25.490000",
          "content": "<p>Yes, tf and tfrecords</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1054341,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2020-10-19T20:36:45.220000",
          "content": "<p>tfrecords make life easier .  But I love too much Pytorch to turn back to TF ^^</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1054570,
          "author_name": "gao-hongnan",
          "author_url": "",
          "post_date": "2020-10-20T02:46:09.070000",
          "content": "<p><a href=\"https://www.kaggle.com/stanleyjzheng\" target=\"_blank\">@stanleyjzheng</a> Thanks a lot bro! </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1054572,
          "author_name": "gao-hongnan",
          "author_url": "",
          "post_date": "2020-10-20T02:46:23.720000",
          "content": "<p><a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> Thanks, my man!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1052100,
      "author_name": "Kirderf",
      "author_url": "",
      "post_date": "2020-10-17T10:04:18.327000",
      "content": "<p>One can also try to read directly from the zip file , if you only have space for the zip file, using a mounted drive, doesn't increase the disk space.</p>\n<pre><code>!mkdir yyy\n!fuse-zip xxx.zip /yyy/\n</code></pre>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1039852,
      "author_name": "Raghawendra Singh",
      "author_url": "",
      "post_date": "2020-10-06T20:43:18.413000",
      "content": "<p>I have a very vague solution that might work:<br>\nif you examine the disk of colab TPU notebook, you get following output for !df -h command.</p>\n<p>Filesystem      Size  Used Avail Use% Mounted on<br>\noverlay         108G   31G   73G  30% /<br>\ntmpfs            64M     0   64M   0% /dev<br>\ntmpfs           6.4G     0  6.4G   0% /sys/fs/cgroup<br>\nshm             5.9G     0  5.9G   0% /dev/shm<br>\ntmpfs           6.4G   12K  6.4G   1% /var/colab<br>\n/dev/sda1       114G   32G   83G  28% /etc/hosts<br>\ntmpfs           6.4G     0  6.4G   0% /proc/acpi<br>\ntmpfs           6.4G     0  6.4G   0% /proc/scsi<br>\ntmpfs           6.4G     0  6.4G   0% /sys/firmware</p>\n<p>here you can use upto 73 GB of disk<br>\nmy solution is create a new dataset split it into 3 parts about 17.6 GB each. Lets say data files are part1.zip, part2.zip and part3.zip. Now download each file, unzip it and delete it one by one. maximum peak storage you will require in this process is 17.6 * 4 = 70.4 GB that is less than available storage. But issue here is it waste lots of time. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1034880,
      "author_name": "Leossb",
      "author_url": "",
      "post_date": "2020-10-02T09:06:05.200000",
      "content": "<p>You can attach your google-drive</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1034884,
          "author_name": "Nitin Datta",
          "author_url": "",
          "post_date": "2020-10-02T09:07:20.577000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/leossb\" target=\"_blank\">@leossb</a> , <br>\nBut the problem is that data is 52GB which cannot be stored on drive unless you buy space</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1034892,
          "author_name": "Leossb",
          "author_url": "",
          "post_date": "2020-10-02T09:12:23.163000",
          "content": "<p>google-drive space is very cheap, you can buy 100Gb by $2/month</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1035213,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2020-10-02T14:39:41.900000",
          "content": "<p>Laoding huge dataset straight from Drive to your model is a very bad  idea.<br>\nNot only it wil be terribly slow but you will reach quickly the daily quota of data you can load. </p>\n<p>I have 2 To of Drive Premium Storage but I will never use it to load such dataset. </p>\n<p>IMHO, TPU + kaggle API to download the data in colab Disk Space is much better idea. </p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 1035217,
          "author_name": "Nitin Datta",
          "author_url": "",
          "post_date": "2020-10-02T14:42:00.763000",
          "content": "<p>Yeah I tried using Kaggle API to load the 52 GB data but in free version the dataset is too  large when compared to disk space provided</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1031950,
      "author_name": "Manh Lab",
      "author_url": "",
      "post_date": "2020-09-29T19:46:06.423000",
      "content": "<p>Download it from GCS. I am downloading 72GB with this</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1032146,
          "author_name": "Nitin Datta",
          "author_url": "",
          "post_date": "2020-09-30T02:02:00.643000",
          "content": "<p>Yes <a href=\"https://www.kaggle.com/doanquanvietnamca\" target=\"_blank\">@doanquanvietnamca</a> ,<br>\nThat's one way or we can create a 128x128 jpegs dataset… </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1031143,
      "author_name": "Naim Mhedhbi",
      "author_url": "",
      "post_date": "2020-09-29T09:07:06.347000",
      "content": "<p>good job ! good</p>",
      "votes": -11,
      "replies": [
        {
          "id": 1031225,
          "author_name": "Nitin Datta",
          "author_url": "",
          "post_date": "2020-09-29T10:21:44.903000",
          "content": "<p>I do not get you</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1035848,
      "author_name": "Aayush Kandpal",
      "author_url": "",
      "post_date": "2020-10-03T06:39:33.140000",
      "content": "<p>Colab has the option to import datsets from google drive. IF you could create a drive and then import that to your colab notebook. That should help. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1035853,
          "author_name": "Nitin Datta",
          "author_url": "",
          "post_date": "2020-10-03T06:43:35.870000",
          "content": "<p>Already tried it :) <br>\nBut the dataset is 52GB and unless you pay you cannot upload the whole data there </p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1030856": "Hi everyone, \nI am trying to import the 52GB dataset made by @vaillant but I run out of disk space.. \nAny suggestions on what can be done so that we can download the data using Kaggle API and train on Colab.\n\nAny help will be appreciated. \nThank you!!!",
    "1031468": "Use TPU (colab Pro)\n\nYou will have much more disk space ",
    "1052100": "One can also try to read directly from the zip file , if you only have space for the zip file, using a mounted drive, doesn't increase the disk space.\n\n```\n!mkdir yyy\n!fuse-zip xxx.zip /yyy/\n```",
    "1039852": "I have a very vague solution that might work:\nif you examine the disk of colab TPU notebook, you get following output for !df -h command.\n\nFilesystem      Size  Used Avail Use% Mounted on\noverlay         108G   31G   73G  30% /\ntmpfs            64M     0   64M   0% /dev\ntmpfs           6.4G     0  6.4G   0% /sys/fs/cgroup\nshm             5.9G     0  5.9G   0% /dev/shm\ntmpfs           6.4G   12K  6.4G   1% /var/colab\n/dev/sda1       114G   32G   83G  28% /etc/hosts\ntmpfs           6.4G     0  6.4G   0% /proc/acpi\ntmpfs           6.4G     0  6.4G   0% /proc/scsi\ntmpfs           6.4G     0  6.4G   0% /sys/firmware\n\nhere you can use upto 73 GB of disk\nmy solution is create a new dataset split it into 3 parts about 17.6 GB each. Lets say data files are part1.zip, part2.zip and part3.zip. Now download each file, unzip it and delete it one by one. maximum peak storage you will require in this process is 17.6 * 4 = 70.4 GB that is less than available storage. But issue here is it waste lots of time. \n",
    "1034880": "You can attach your google-drive",
    "1031950": "Download it from GCS. I am downloading 72GB with this",
    "1031143": "good job ! good",
    "1035848": "Colab has the option to import datsets from google drive. IF you could create a drive and then import that to your colab notebook. That should help. "
  }
}