{
  "id": 113232,
  "title": "Training Models on GCP by mounting bucket using FUSE",
  "url": "/competitions/rsna-intracranial-hemorrhage-detection/discussion/113232",
  "author_name": "cherring",
  "post_date": "2019-10-18T01:05:12.383000",
  "votes": 4,
  "comment_count": 2,
  "views": 0,
  "content": "<p>HI All!</p>\n\n<p>I am just getting started on this competition, usually I have access to several machines that I can experiment on however they are busy doing other stuff. So I am looking to use GCP.</p>\n\n<p>I am trying to figure out the best way to be able to easily spin up instances and do some training. I have created an image on GCP with everything installed so I can quickly spin up instances that are ready to go.\nI am planning on setting up a GCP bucket and writing some scripts to store logs and artifacts.</p>\n\n<p>However one thing that I am struggling to come up with a good solution for is how to handle the data. Currently I have a 600GB persistent disk that I have downloaded and extracted the data to, but that is obviously not a solution as it is locked to a particular region.</p>\n\n<p>I think the best way is going to be to transfer that data to a GCP bucket. I then have a couple of options:</p>\n\n<ol>\n<li>When an instance is spin up, download all of that data to a local disk before beginning training.</li>\n<li>Add caching to my data generator so I can begin training right away, if it is the first time accessing a dcm then download it from bucket and save it to local disk.</li>\n<li>Use FUSE to mount the bucket as a filesystem.</li>\n</ol>\n\n<p>Can anyone provide any feedback on the viability of each of these - or offer better alternatives :)</p>\n\n<p>Since we are typically only training a couple of epochs, the overhead of option 1 feels like a pretty high cost. Is the first epoch of option 2 and 3 going to have the same latency in loading data?\nIs the latency of loading data from a bucket an actual bottleneck? If so are there ways that I can reduce these bottlenecks (eg formatting the data into TFRecords or something)</p>\n\n<p>I am using Tensorflow (1.14 but maybe I will give 2.0 a shot)</p>",
  "messages": [
    {
      "id": 651808,
      "postDate": "2019-10-18T01:05:12.383Z",
      "content": "<p>HI All!</p>\n\n<p>I am just getting started on this competition, usually I have access to several machines that I can experiment on however they are busy doing other stuff. So I am looking to use GCP.</p>\n\n<p>I am trying to figure out the best way to be able to easily spin up instances and do some training. I have created an image on GCP with everything installed so I can quickly spin up instances that are ready to go.\nI am planning on setting up a GCP bucket and writing some scripts to store logs and artifacts.</p>\n\n<p>However one thing that I am struggling to come up with a good solution for is how to handle the data. Currently I have a 600GB persistent disk that I have downloaded and extracted the data to, but that is obviously not a solution as it is locked to a particular region.</p>\n\n<p>I think the best way is going to be to transfer that data to a GCP bucket. I then have a couple of options:</p>\n\n<ol>\n<li>When an instance is spin up, download all of that data to a local disk before beginning training.</li>\n<li>Add caching to my data generator so I can begin training right away, if it is the first time accessing a dcm then download it from bucket and save it to local disk.</li>\n<li>Use FUSE to mount the bucket as a filesystem.</li>\n</ol>\n\n<p>Can anyone provide any feedback on the viability of each of these - or offer better alternatives :)</p>\n\n<p>Since we are typically only training a couple of epochs, the overhead of option 1 feels like a pretty high cost. Is the first epoch of option 2 and 3 going to have the same latency in loading data?\nIs the latency of loading data from a bucket an actual bottleneck? If so are there ways that I can reduce these bottlenecks (eg formatting the data into TFRecords or something)</p>\n\n<p>I am using Tensorflow (1.14 but maybe I will give 2.0 a shot)</p>",
      "rawMarkdown": "HI All!\n\nI am just getting started on this competition, usually I have access to several machines that I can experiment on however they are busy doing other stuff. So I am looking to use GCP.\n\nI am trying to figure out the best way to be able to easily spin up instances and do some training. I have created an image on GCP with everything installed so I can quickly spin up instances that are ready to go.\nI am planning on setting up a GCP bucket and writing some scripts to store logs and artifacts.\n\nHowever one thing that I am struggling to come up with a good solution for is how to handle the data. Currently I have a 600GB persistent disk that I have downloaded and extracted the data to, but that is obviously not a solution as it is locked to a particular region.\n\nI think the best way is going to be to transfer that data to a GCP bucket. I then have a couple of options:\n\n1. When an instance is spin up, download all of that data to a local disk before beginning training.\n2. Add caching to my data generator so I can begin training right away, if it is the first time accessing a dcm then download it from bucket and save it to local disk.\n3. Use FUSE to mount the bucket as a filesystem.\n\nCan anyone provide any feedback on the viability of each of these - or offer better alternatives :)\n\nSince we are typically only training a couple of epochs, the overhead of option 1 feels like a pretty high cost. Is the first epoch of option 2 and 3 going to have the same latency in loading data?\nIs the latency of loading data from a bucket an actual bottleneck? If so are there ways that I can reduce these bottlenecks (eg formatting the data into TFRecords or something)\n\nI am using Tensorflow (1.14 but maybe I will give 2.0 a shot)",
      "votes": 4
    },
    {
      "id": 663925,
      "postDate": "2019-11-02T21:57:27.013Z",
      "content": "<p>What approach are you using now?</p>",
      "rawMarkdown": "What approach are you using now?\n",
      "votes": -1,
      "replies": [
        {
          "id": 664371,
          "postDate": "2019-11-03T15:19:18.100Z",
          "content": "<p>I ended up just going with option 1. gcloud cp transfers from the bucket at a good 100MiBps, so it doesn't take too long to transfer</p>",
          "rawMarkdown": "I ended up just going with option 1. gcloud cp transfers from the bucket at a good 100MiBps, so it doesn't take too long to transfer",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 663925,
      "author_name": "NitinKshatriya",
      "author_url": "",
      "post_date": "2019-11-02T21:57:27.013000",
      "content": "<p>What approach are you using now?</p>",
      "votes": -1,
      "replies": [
        {
          "id": 664371,
          "author_name": "cherring",
          "author_url": "",
          "post_date": "2019-11-03T15:19:18.100000",
          "content": "<p>I ended up just going with option 1. gcloud cp transfers from the bucket at a good 100MiBps, so it doesn't take too long to transfer</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "651808": "HI All!\n\nI am just getting started on this competition, usually I have access to several machines that I can experiment on however they are busy doing other stuff. So I am looking to use GCP.\n\nI am trying to figure out the best way to be able to easily spin up instances and do some training. I have created an image on GCP with everything installed so I can quickly spin up instances that are ready to go.\nI am planning on setting up a GCP bucket and writing some scripts to store logs and artifacts.\n\nHowever one thing that I am struggling to come up with a good solution for is how to handle the data. Currently I have a 600GB persistent disk that I have downloaded and extracted the data to, but that is obviously not a solution as it is locked to a particular region.\n\nI think the best way is going to be to transfer that data to a GCP bucket. I then have a couple of options:\n\n1. When an instance is spin up, download all of that data to a local disk before beginning training.\n2. Add caching to my data generator so I can begin training right away, if it is the first time accessing a dcm then download it from bucket and save it to local disk.\n3. Use FUSE to mount the bucket as a filesystem.\n\nCan anyone provide any feedback on the viability of each of these - or offer better alternatives :)\n\nSince we are typically only training a couple of epochs, the overhead of option 1 feels like a pretty high cost. Is the first epoch of option 2 and 3 going to have the same latency in loading data?\nIs the latency of loading data from a bucket an actual bottleneck? If so are there ways that I can reduce these bottlenecks (eg formatting the data into TFRecords or something)\n\nI am using Tensorflow (1.14 but maybe I will give 2.0 a shot)",
    "663925": "What approach are you using now?\n"
  }
}