{
  "id": 156470,
  "title": "How are you doing this ? Need Suggestion",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/156470",
  "author_name": "Nirjhar Roy",
  "post_date": "2020-06-06T09:42:07.762000",
  "votes": 4,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I am not a seasoned kaggler or programmer so I might be missing something obvious . So please suggest . </p>\n\n<ol>\n<li><p>The data for the three-four image comp running are huge and there are no way to download necessary data alone via API . How do you get this data in cloud storage ? Downloading locally and uploading to cloud is not feasible due to network problems .</p></li>\n<li><p>The Kaggle output directory space seems to be a problem for me . Right now most of the SOTA models are big (especially for NLP area) , i cant even run a 5 fold at one shot and run my model . It might crash after running for 5 hours due to output size &gt; 5 GB. The same thing happens when I try to generate dataset from large files e.g. jpg from tiff in this comp . Is there a possibility to increase this local space to lets say 20GB or something ? What is the flipside of that ? How are you all doing it ?</p></li>\n<li><p>I bought some Google Drive space of 2 TB during DFDC time , but then when i download data via API i understood that it uses the local diskspace while downloading and hence we cant download &gt;100GB  at one time . So that space is kind of useless.Is there any workaround ?</p></li>\n</ol>",
  "messages": [
    {
      "id": 875934,
      "postDate": "2020-06-06T09:42:07.763Z",
      "content": "<p>I am not a seasoned kaggler or programmer so I might be missing something obvious . So please suggest . </p>\n\n<ol>\n<li><p>The data for the three-four image comp running are huge and there are no way to download necessary data alone via API . How do you get this data in cloud storage ? Downloading locally and uploading to cloud is not feasible due to network problems .</p></li>\n<li><p>The Kaggle output directory space seems to be a problem for me . Right now most of the SOTA models are big (especially for NLP area) , i cant even run a 5 fold at one shot and run my model . It might crash after running for 5 hours due to output size &gt; 5 GB. The same thing happens when I try to generate dataset from large files e.g. jpg from tiff in this comp . Is there a possibility to increase this local space to lets say 20GB or something ? What is the flipside of that ? How are you all doing it ?</p></li>\n<li><p>I bought some Google Drive space of 2 TB during DFDC time , but then when i download data via API i understood that it uses the local diskspace while downloading and hence we cant download &gt;100GB  at one time . So that space is kind of useless.Is there any workaround ?</p></li>\n</ol>",
      "rawMarkdown": "I am not a seasoned kaggler or programmer so I might be missing something obvious . So please suggest . \n\n1. The data for the three-four image comp running are huge and there are no way to download necessary data alone via API . How do you get this data in cloud storage ? Downloading locally and uploading to cloud is not feasible due to network problems .\n\n2. The Kaggle output directory space seems to be a problem for me . Right now most of the SOTA models are big (especially for NLP area) , i cant even run a 5 fold at one shot and run my model . It might crash after running for 5 hours due to output size &gt; 5 GB. The same thing happens when I try to generate dataset from large files e.g. jpg from tiff in this comp . Is there a possibility to increase this local space to lets say 20GB or something ? What is the flipside of that ? How are you all doing it ?\n\n3. I bought some Google Drive space of 2 TB during DFDC time , but then when i download data via API i understood that it uses the local diskspace while downloading and hence we cant download &gt;100GB  at one time . So that space is kind of useless.Is there any workaround ?",
      "votes": 4
    },
    {
      "id": 876661,
      "postDate": "2020-06-06T21:56:41.727Z",
      "content": "<ol>\n<li><p>I don't understand. Just use the kaggle api to download it in a cloud environment. Make sure you have a large enough ssd or cold disk on your instance. Speaking of prostate cancer competition I have done it once on aws and once on gcp.</p></li>\n<li><p>Kaggle kernels are very limited yes. Not much you can do about that.</p></li>\n<li><p>Why are you using Google DRIVE ? Use the real tool for that kind of thing: google cloud.</p></li>\n</ol>\n\n<p>Alternatively use colab it's probably easier to start with it than starting (costly) instances if you have no experience with it.</p>",
      "rawMarkdown": "1. I don't understand. Just use the kaggle api to download it in a cloud environment. Make sure you have a large enough ssd or cold disk on your instance. Speaking of prostate cancer competition I have done it once on aws and once on gcp.\n\n2. Kaggle kernels are very limited yes. Not much you can do about that.\n\n3. Why are you using Google DRIVE ? Use the real tool for that kind of thing: google cloud.\n\nAlternatively use colab it's probably easier to start with it than starting (costly) instances if you have no experience with it.",
      "votes": -1
    }
  ],
  "comments": [
    {
      "id": 876661,
      "author_name": "Arnaud Roussel",
      "author_url": "",
      "post_date": "2020-06-06T21:56:41.727000",
      "content": "<ol>\n<li><p>I don't understand. Just use the kaggle api to download it in a cloud environment. Make sure you have a large enough ssd or cold disk on your instance. Speaking of prostate cancer competition I have done it once on aws and once on gcp.</p></li>\n<li><p>Kaggle kernels are very limited yes. Not much you can do about that.</p></li>\n<li><p>Why are you using Google DRIVE ? Use the real tool for that kind of thing: google cloud.</p></li>\n</ol>\n\n<p>Alternatively use colab it's probably easier to start with it than starting (costly) instances if you have no experience with it.</p>",
      "votes": -1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "875934": "I am not a seasoned kaggler or programmer so I might be missing something obvious . So please suggest . \n\n1. The data for the three-four image comp running are huge and there are no way to download necessary data alone via API . How do you get this data in cloud storage ? Downloading locally and uploading to cloud is not feasible due to network problems .\n\n2. The Kaggle output directory space seems to be a problem for me . Right now most of the SOTA models are big (especially for NLP area) , i cant even run a 5 fold at one shot and run my model . It might crash after running for 5 hours due to output size &gt; 5 GB. The same thing happens when I try to generate dataset from large files e.g. jpg from tiff in this comp . Is there a possibility to increase this local space to lets say 20GB or something ? What is the flipside of that ? How are you all doing it ?\n\n3. I bought some Google Drive space of 2 TB during DFDC time , but then when i download data via API i understood that it uses the local diskspace while downloading and hence we cant download &gt;100GB  at one time . So that space is kind of useless.Is there any workaround ?",
    "876661": "1. I don't understand. Just use the kaggle api to download it in a cloud environment. Make sure you have a large enough ssd or cold disk on your instance. Speaking of prostate cancer competition I have done it once on aws and once on gcp.\n\n2. Kaggle kernels are very limited yes. Not much you can do about that.\n\n3. Why are you using Google DRIVE ? Use the real tool for that kind of thing: google cloud.\n\nAlternatively use colab it's probably easier to start with it than starting (costly) instances if you have no experience with it."
  }
}