{
  "id": 224538,
  "title": "🙋‍♂️Original image data training with Detectron2 [GPU usage] 🙋‍♂️",
  "url": "/competitions/vinbigdata-chest-xray-abnormalities-detection/discussion/224538",
  "author_name": "sourabhsc",
  "post_date": "2021-03-09T00:39:34.596000",
  "votes": 2,
  "comment_count": 5,
  "views": 0,
  "content": "<p>For smaller abnormalities, it makes sense to use larger images. Previous discussions in the forum have also pointed out that original images could work better for training and to identify the bounding boxes with smaller sizes. <br>\nI am using <a href=\"https://www.kaggle.com/corochann/vinbigdata-detectron2-train\" target=\"_blank\">Corochan's notebook</a> for Detectron2 FasterRCNN on original image data. However, I kept getting OOM (Out of memeory) error for the GPU. </p>\n<p>By making some changes to the parameters to reduce GPU memory usage, I managed to run it for 1000 iterations.<br>\nI am choosing ---</p>\n<blockquote>\n  <ul>\n  <li>R50 backbone (<code>config_name = \"faster_rcnn_R_50_FPN_3x.yaml\"</code>)</li>\n  <li><code>ROI_HEADS.BATCH_SIZE_PER_IMAGE = 256</code></li>\n  <li><code>HorizontalFlip\": {\"p\": 0.75}</code></li>\n  <li><code>SOLVER.IMS_PER_BATCH = 1</code></li>\n  <li><code>SOLVER.BASE_LR = 0.005</code></li>\n  </ul>\n  <p>However, this run takes 1 hr 20 minutes on Kaggle. That means for a 9 hour time limit on Kaggle GPUs, I can only run the models for (540/80)*1000 = 6750 iterations. I don't think this would be enough to finish the training.</p>\n</blockquote>\n<p>❓ Can someone please suggest me how I can train the model for higher iterations? <br>\n❓ Or, how I can use a different set of parameters to reduce the memory usage by the GPU?<br>\nI must be missing something straight forward with transfer learning. I am not sure how to set it up for Detectron2. </p>\n<p>Thanks for your help. </p>\n<p>PS- I can't train it on my laptop as it has just a 4GB GPU. :) </p>",
  "messages": [
    {
      "id": 1231961,
      "postDate": "2021-03-09T11:58:51.893Z",
      "content": "<p><a href=\"https://www.kaggle.com/sourabhchauhan\" target=\"_blank\">@sourabhchauhan</a> Colab pro gives you double the RAM size 25 gb (Think so ) . You can upgrade it for 10 usd a month</p>",
      "rawMarkdown": "@sourabhchauhan Colab pro gives you double the RAM size 25 gb (Think so ) . You can upgrade it for 10 usd a month",
      "votes": 1,
      "replies": [
        {
          "id": 1232360,
          "postDate": "2021-03-09T17:34:13.993Z",
          "content": "<p><a href=\"https://www.kaggle.com/usharengaraju\" target=\"_blank\">@usharengaraju</a> But only if you live in the US/Canada, isn't it?</p>",
          "rawMarkdown": "@usharengaraju But only if you live in the US/Canada, isn't it?"
        },
        {
          "id": 1232606,
          "postDate": "2021-03-09T21:20:22.593Z",
          "content": "<p><a href=\"https://www.kaggle.com/hannes82\" target=\"_blank\">@hannes82</a> I'll let you in on a little secret - just put a US postal code and you can get it. </p>",
          "rawMarkdown": "@hannes82 I'll let you in on a little secret - just put a US postal code and you can get it. "
        },
        {
          "id": 1232690,
          "postDate": "2021-03-09T23:46:28.140Z",
          "content": "<p>Ah I see :-) For now I think I will stick to kaggle ressources but thanks for letting me know!</p>",
          "rawMarkdown": "Ah I see :-) For now I think I will stick to kaggle ressources but thanks for letting me know!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1231440,
      "postDate": "2021-03-09T01:44:13.577Z",
      "content": "<p>You can use checkpoint, save the model before kaggle 9 hour limit ends your session, then load the model and continue training. </p>",
      "rawMarkdown": "You can use checkpoint, save the model before kaggle 9 hour limit ends your session, then load the model and continue training. ",
      "votes": 2
    },
    {
      "id": 1231408,
      "postDate": "2021-03-09T00:39:34.597Z",
      "content": "<p>For smaller abnormalities, it makes sense to use larger images. Previous discussions in the forum have also pointed out that original images could work better for training and to identify the bounding boxes with smaller sizes. <br>\nI am using <a href=\"https://www.kaggle.com/corochann/vinbigdata-detectron2-train\" target=\"_blank\">Corochan's notebook</a> for Detectron2 FasterRCNN on original image data. However, I kept getting OOM (Out of memeory) error for the GPU. </p>\n<p>By making some changes to the parameters to reduce GPU memory usage, I managed to run it for 1000 iterations.<br>\nI am choosing ---</p>\n<blockquote>\n  <ul>\n  <li>R50 backbone (<code>config_name = \"faster_rcnn_R_50_FPN_3x.yaml\"</code>)</li>\n  <li><code>ROI_HEADS.BATCH_SIZE_PER_IMAGE = 256</code></li>\n  <li><code>HorizontalFlip\": {\"p\": 0.75}</code></li>\n  <li><code>SOLVER.IMS_PER_BATCH = 1</code></li>\n  <li><code>SOLVER.BASE_LR = 0.005</code></li>\n  </ul>\n  <p>However, this run takes 1 hr 20 minutes on Kaggle. That means for a 9 hour time limit on Kaggle GPUs, I can only run the models for (540/80)*1000 = 6750 iterations. I don't think this would be enough to finish the training.</p>\n</blockquote>\n<p>❓ Can someone please suggest me how I can train the model for higher iterations? <br>\n❓ Or, how I can use a different set of parameters to reduce the memory usage by the GPU?<br>\nI must be missing something straight forward with transfer learning. I am not sure how to set it up for Detectron2. </p>\n<p>Thanks for your help. </p>\n<p>PS- I can't train it on my laptop as it has just a 4GB GPU. :) </p>",
      "rawMarkdown": "For smaller abnormalities, it makes sense to use larger images. Previous discussions in the forum have also pointed out that original images could work better for training and to identify the bounding boxes with smaller sizes. \nI am using [Corochan's notebook] (https://www.kaggle.com/corochann/vinbigdata-detectron2-train) for Detectron2 FasterRCNN on original image data. However, I kept getting OOM (Out of memeory) error for the GPU. \n\nBy making some changes to the parameters to reduce GPU memory usage, I managed to run it for 1000 iterations.\nI am choosing ---\n> \n- R50 backbone (`config_name = \"faster_rcnn_R_50_FPN_3x.yaml\" `)\n- `ROI_HEADS.BATCH_SIZE_PER_IMAGE = 256`\n-  `HorizontalFlip\": {\"p\": 0.75}`\n- `SOLVER.IMS_PER_BATCH = 1 `\n- `SOLVER.BASE_LR = 0.005`\n>\nHowever, this run takes 1 hr 20 minutes on Kaggle. That means for a 9 hour time limit on Kaggle GPUs, I can only run the models for (540/80)*1000 = 6750 iterations. I don't think this would be enough to finish the training.\n\n❓ Can someone please suggest me how I can train the model for higher iterations? \n❓ Or, how I can use a different set of parameters to reduce the memory usage by the GPU?\nI must be missing something straight forward with transfer learning. I am not sure how to set it up for Detectron2. \n\nThanks for your help. \n\nPS- I can't train it on my laptop as it has just a 4GB GPU. :) \n \n\n\n",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 1231961,
      "author_name": "Tensor Girl",
      "author_url": "",
      "post_date": "2021-03-09T11:58:51.893000",
      "content": "<p><a href=\"https://www.kaggle.com/sourabhchauhan\" target=\"_blank\">@sourabhchauhan</a> Colab pro gives you double the RAM size 25 gb (Think so ) . You can upgrade it for 10 usd a month</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1232360,
          "author_name": "Hannes Öhler",
          "author_url": "",
          "post_date": "2021-03-09T17:34:13.993000",
          "content": "<p><a href=\"https://www.kaggle.com/usharengaraju\" target=\"_blank\">@usharengaraju</a> But only if you live in the US/Canada, isn't it?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1232606,
          "author_name": "Stanley Zheng",
          "author_url": "",
          "post_date": "2021-03-09T21:20:22.593000",
          "content": "<p><a href=\"https://www.kaggle.com/hannes82\" target=\"_blank\">@hannes82</a> I'll let you in on a little secret - just put a US postal code and you can get it. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1232690,
          "author_name": "Hannes Öhler",
          "author_url": "",
          "post_date": "2021-03-09T23:46:28.140000",
          "content": "<p>Ah I see :-) For now I think I will stick to kaggle ressources but thanks for letting me know!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1231440,
      "author_name": "Stanley Zheng",
      "author_url": "",
      "post_date": "2021-03-09T01:44:13.577000",
      "content": "<p>You can use checkpoint, save the model before kaggle 9 hour limit ends your session, then load the model and continue training. </p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1231961": "@sourabhchauhan Colab pro gives you double the RAM size 25 gb (Think so ) . You can upgrade it for 10 usd a month",
    "1231440": "You can use checkpoint, save the model before kaggle 9 hour limit ends your session, then load the model and continue training. ",
    "1231408": "For smaller abnormalities, it makes sense to use larger images. Previous discussions in the forum have also pointed out that original images could work better for training and to identify the bounding boxes with smaller sizes. \nI am using [Corochan's notebook] (https://www.kaggle.com/corochann/vinbigdata-detectron2-train) for Detectron2 FasterRCNN on original image data. However, I kept getting OOM (Out of memeory) error for the GPU. \n\nBy making some changes to the parameters to reduce GPU memory usage, I managed to run it for 1000 iterations.\nI am choosing ---\n> \n- R50 backbone (`config_name = \"faster_rcnn_R_50_FPN_3x.yaml\" `)\n- `ROI_HEADS.BATCH_SIZE_PER_IMAGE = 256`\n-  `HorizontalFlip\": {\"p\": 0.75}`\n- `SOLVER.IMS_PER_BATCH = 1 `\n- `SOLVER.BASE_LR = 0.005`\n>\nHowever, this run takes 1 hr 20 minutes on Kaggle. That means for a 9 hour time limit on Kaggle GPUs, I can only run the models for (540/80)*1000 = 6750 iterations. I don't think this would be enough to finish the training.\n\n❓ Can someone please suggest me how I can train the model for higher iterations? \n❓ Or, how I can use a different set of parameters to reduce the memory usage by the GPU?\nI must be missing something straight forward with transfer learning. I am not sure how to set it up for Detectron2. \n\nThanks for your help. \n\nPS- I can't train it on my laptop as it has just a 4GB GPU. :) \n \n\n\n"
  }
}