{
  "id": 436877,
  "title": "Training the model with on-demand GPUs",
  "url": "/competitions/rsna-2023-abdominal-trauma-detection/discussion/436877",
  "author_name": "Victor Shlepov",
  "post_date": "2023-09-04T14:31:32.021000",
  "votes": 0,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi! I would appreciate your advise here:</p>\n<p>I'm in a good shape to start the training the final model - I've converted and preprocessed everting to the TFRecord dataset and ready to go. The issue her is the size - about 500Gb - it's too large for Colab Pro (it has the maximum disk space about 160Gb). Training on the local machine is not a good option too.</p>\n<p>Ideally, I'd prefer to cache all the dataset to RAM and process it with a couple of 80GB A100 GPU's. How do you normally handle that, any recommended on-demand services?</p>\n<p>Thanks! And have a good day!</p>",
  "messages": [
    {
      "id": 2423255,
      "postDate": "2023-09-04T14:31:32.020Z",
      "content": "<p>Hi! I would appreciate your advise here:</p>\n<p>I'm in a good shape to start the training the final model - I've converted and preprocessed everting to the TFRecord dataset and ready to go. The issue her is the size - about 500Gb - it's too large for Colab Pro (it has the maximum disk space about 160Gb). Training on the local machine is not a good option too.</p>\n<p>Ideally, I'd prefer to cache all the dataset to RAM and process it with a couple of 80GB A100 GPU's. How do you normally handle that, any recommended on-demand services?</p>\n<p>Thanks! And have a good day!</p>",
      "rawMarkdown": "Hi! I would appreciate your advise here:\n\nI'm in a good shape to start the training the final model - I've converted and preprocessed everting to the TFRecord dataset and ready to go. The issue her is the size - about 500Gb - it's too large for Colab Pro (it has the maximum disk space about 160Gb). Training on the local machine is not a good option too.\n\nIdeally, I'd prefer to cache all the dataset to RAM and process it with a couple of 80GB A100 GPU's. How do you normally handle that, any recommended on-demand services?\n\nThanks! And have a good day!"
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2423255": "Hi! I would appreciate your advise here:\n\nI'm in a good shape to start the training the final model - I've converted and preprocessed everting to the TFRecord dataset and ready to go. The issue her is the size - about 500Gb - it's too large for Colab Pro (it has the maximum disk space about 160Gb). Training on the local machine is not a good option too.\n\nIdeally, I'd prefer to cache all the dataset to RAM and process it with a couple of 80GB A100 GPU's. How do you normally handle that, any recommended on-demand services?\n\nThanks! And have a good day!"
  }
}