{
  "id": 371996,
  "title": "OOM while train on gpu",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/371996",
  "author_name": "Lau2664",
  "post_date": "2022-12-13T16:23:38.291000",
  "votes": 2,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Hello guys! I trained my model (efficientnetB4) with tensorflow on GPU P100 with image size of 256X256, and when I use a bigger image size (512/1024) the OOM occurs. I noticed that many of you are using very large image size like 2048, how did you achieve that without OOM ?</p>",
  "messages": [
    {
      "id": 2082098,
      "postDate": "2023-01-01T05:25:49.463Z",
      "content": "<p>You can enable gradient checkpointing like this:</p>\n<pre><code>import timm\nmodel = timm.create_model('efficientnet_b0')\nmodel.set_grad_checkpointing(enable=True)\n</code></pre>\n<p>For more efficient use of gradient checkpointing you can go to <code>timm/models/efficientnet.py</code>, check <code>checkpoint_seq</code> method that has <code>every</code> argument, you can set it to something different from 1 to speed up training (it'll use more memory though, tweak it to use as much VRAM as possible)</p>",
      "rawMarkdown": "You can enable gradient checkpointing like this:\n```\nimport timm\nmodel = timm.create_model('efficientnet_b0')\nmodel.set_grad_checkpointing(enable=True)\n```\n\nFor more efficient use of gradient checkpointing you can go to ```timm/models/efficientnet.py```, check ```checkpoint_seq``` method that has ```every``` argument, you can set it to something different from 1 to speed up training (it'll use more memory though, tweak it to use as much VRAM as possible)",
      "votes": 1
    },
    {
      "id": 2068443,
      "postDate": "2022-12-17T23:06:59.150Z",
      "content": "<p>Gradient checkpointing is also very effective in saving VRAM.</p>",
      "rawMarkdown": "Gradient checkpointing is also very effective in saving VRAM.",
      "votes": 1
    },
    {
      "id": 2064301,
      "postDate": "2022-12-13T16:45:35.250Z",
      "content": "<p>There are many ways to do this.</p>\n<ul>\n<li>Use smaller models.</li>\n<li>Lower batch size.</li>\n<li>Use autocast (fp16) if you are using pytorch. The data size is reduced by half compared to the normally used fp32.</li>\n<li>Use TPU if using tensorflow. Because the memory size of TPU is very large.</li>\n</ul>\n<p>These are just examples, and there may be many more ways to do it, but the simple ways I came up with are above.</p>",
      "rawMarkdown": "There are many ways to do this.\n\n- Use smaller models.\n- Lower batch size.\n- Use autocast (fp16) if you are using pytorch. The data size is reduced by half compared to the normally used fp32.\n- Use TPU if using tensorflow. Because the memory size of TPU is very large.\n\nThese are just examples, and there may be many more ways to do it, but the simple ways I came up with are above.",
      "votes": 1,
      "replies": [
        {
          "id": 2064336,
          "postDate": "2022-12-13T17:20:55.207Z",
          "content": "<p>Thx for replying! <br>When I'm using TPU and running the code <code>from kaggle_datasets import KaggleDatasets</code> , I got another problem: <code>No module named 'kaggle_datasets'</code>. I can't use pip to install this library either.  🙏plz tell me how I can solve this.</p>",
          "rawMarkdown": "Thx for replying! <br/>When I'm using TPU and running the code `from kaggle_datasets import KaggleDatasets` , I got another problem: `No module named 'kaggle_datasets'`. I can't use pip to install this library either.  🙏plz tell me how I can solve this."
        }
      ]
    },
    {
      "id": 2067595,
      "postDate": "2022-12-16T21:28:49.220Z",
      "content": "<p>Hi !</p>\n<p>I recently moved to 1024x1024 for training and this is what helped me to not get OOM :</p>\n<ul>\n<li>Use the T4x2 gpus in kaggle notebook. This allow to get 30Gb of Vram instead of 15Gb with P100.</li>\n<li>Reduce batch size and use gradient accumulation. Reducing the batch size allow to use less Vram, but if the batch size is very small (4 for example) this will have a big impact on convergence. Use gradient accumulation to get better results.</li>\n<li>EfficientNetB4 is a big model. Maybe you can try a smaller version like EfficientNetB0 or even EfficientNetV2-S (for small)</li>\n</ul>\n<p>Hope this will help !</p>",
      "rawMarkdown": "Hi !\n\nI recently moved to 1024x1024 for training and this is what helped me to not get OOM :\n- Use the T4x2 gpus in kaggle notebook. This allow to get 30Gb of Vram instead of 15Gb with P100.\n- Reduce batch size and use gradient accumulation. Reducing the batch size allow to use less Vram, but if the batch size is very small (4 for example) this will have a big impact on convergence. Use gradient accumulation to get better results.\n- EfficientNetB4 is a big model. Maybe you can try a smaller version like EfficientNetB0 or even EfficientNetV2-S (for small)\n\nHope this will help !",
      "votes": 2,
      "replies": [
        {
          "id": 2067856,
          "postDate": "2022-12-17T08:49:12.523Z",
          "content": "<p>I also forget that you can freeze a certain amount of layers if you use a pretrained model. This will also help reduce the amount of Vram used when training -&gt; <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/372550\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/372550</a></p>",
          "rawMarkdown": "I also forget that you can freeze a certain amount of layers if you use a pretrained model. This will also help reduce the amount of Vram used when training -> https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/372550"
        },
        {
          "id": 2082027,
          "postDate": "2023-01-01T02:36:39.783Z",
          "content": "<p>Hi! <a href=\"https://www.kaggle.com/nathanlaoue\" target=\"_blank\">@nathanlaoue</a>. Thanks for replying! I got a problem when using large image size: I used TPU to train EfficientNetB4 with input size of 768x768, and the batch size is 128. The loss and pf1 score changed very slowly during training, sometimes they are even unchanged. Does that indicate that my model has some problems in convergence? How can I fix it?</p>",
          "rawMarkdown": "Hi! @nathanlaoue. Thanks for replying! I got a problem when using large image size: I used TPU to train EfficientNetB4 with input size of 768x768, and the batch size is 128. The loss and pf1 score changed very slowly during training, sometimes they are even unchanged. Does that indicate that my model has some problems in convergence? How can I fix it?",
          "replies": [
            {
              "id": 2082190,
              "postDate": "2023-01-01T08:51:35.287Z",
              "content": "<p>Hi! Do you freeze layer ? Maybe you freeze to much layer and the model can not update the layers that have an impact on learning to detect cancer. Also very basic answer but maybe your learning rate is to low.</p>",
              "rawMarkdown": "Hi! Do you freeze layer ? Maybe you freeze to much layer and the model can not update the layers that have an impact on learning to detect cancer. Also very basic answer but maybe your learning rate is to low."
            },
            {
              "id": 2082252,
              "postDate": "2023-01-01T10:54:12.917Z",
              "content": "<p>I didn't freeze any layer. The whole model is trainable. I also used a cosine learning rate scheduler, the lr is changing during the training.</p>",
              "rawMarkdown": "I didn't freeze any layer. The whole model is trainable. I also used a cosine learning rate scheduler, the lr is changing during the training."
            },
            {
              "id": 2082494,
              "postDate": "2023-01-01T16:18:28.810Z",
              "content": "<p>For how many epochs did you train ? Maybe you didn't wait enough. Also this can be caused by a bug in your code. If you are using PyTorch maybe you have an error in your training loop causing the weights to not update the way they should.</p>",
              "rawMarkdown": "For how many epochs did you train ? Maybe you didn't wait enough. Also this can be caused by a bug in your code. If you are using PyTorch maybe you have an error in your training loop causing the weights to not update the way they should."
            }
          ]
        }
      ]
    },
    {
      "id": 2064261,
      "postDate": "2022-12-13T16:23:38.290Z",
      "content": "<p>Hello guys! I trained my model (efficientnetB4) with tensorflow on GPU P100 with image size of 256X256, and when I use a bigger image size (512/1024) the OOM occurs. I noticed that many of you are using very large image size like 2048, how did you achieve that without OOM ?</p>",
      "rawMarkdown": "Hello guys! I trained my model (efficientnetB4) with tensorflow on GPU P100 with image size of 256X256, and when I use a bigger image size (512/1024) the OOM occurs. I noticed that many of you are using very large image size like 2048, how did you achieve that without OOM ?",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 2082098,
      "author_name": "slime",
      "author_url": "",
      "post_date": "2023-01-01T05:25:49.463000",
      "content": "<p>You can enable gradient checkpointing like this:</p>\n<pre><code>import timm\nmodel = timm.create_model('efficientnet_b0')\nmodel.set_grad_checkpointing(enable=True)\n</code></pre>\n<p>For more efficient use of gradient checkpointing you can go to <code>timm/models/efficientnet.py</code>, check <code>checkpoint_seq</code> method that has <code>every</code> argument, you can set it to something different from 1 to speed up training (it'll use more memory though, tweak it to use as much VRAM as possible)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2068443,
      "author_name": "S. Tomizawa",
      "author_url": "",
      "post_date": "2022-12-17T23:06:59.150000",
      "content": "<p>Gradient checkpointing is also very effective in saving VRAM.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2064301,
      "author_name": "YYama",
      "author_url": "",
      "post_date": "2022-12-13T16:45:35.250000",
      "content": "<p>There are many ways to do this.</p>\n<ul>\n<li>Use smaller models.</li>\n<li>Lower batch size.</li>\n<li>Use autocast (fp16) if you are using pytorch. The data size is reduced by half compared to the normally used fp32.</li>\n<li>Use TPU if using tensorflow. Because the memory size of TPU is very large.</li>\n</ul>\n<p>These are just examples, and there may be many more ways to do it, but the simple ways I came up with are above.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2064336,
          "author_name": "Lau2664",
          "author_url": "",
          "post_date": "2022-12-13T17:20:55.207000",
          "content": "<p>Thx for replying! <br>When I'm using TPU and running the code <code>from kaggle_datasets import KaggleDatasets</code> , I got another problem: <code>No module named 'kaggle_datasets'</code>. I can't use pip to install this library either.  🙏plz tell me how I can solve this.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2067595,
      "author_name": "Natyu",
      "author_url": "",
      "post_date": "2022-12-16T21:28:49.220000",
      "content": "<p>Hi !</p>\n<p>I recently moved to 1024x1024 for training and this is what helped me to not get OOM :</p>\n<ul>\n<li>Use the T4x2 gpus in kaggle notebook. This allow to get 30Gb of Vram instead of 15Gb with P100.</li>\n<li>Reduce batch size and use gradient accumulation. Reducing the batch size allow to use less Vram, but if the batch size is very small (4 for example) this will have a big impact on convergence. Use gradient accumulation to get better results.</li>\n<li>EfficientNetB4 is a big model. Maybe you can try a smaller version like EfficientNetB0 or even EfficientNetV2-S (for small)</li>\n</ul>\n<p>Hope this will help !</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2067856,
          "author_name": "Natyu",
          "author_url": "",
          "post_date": "2022-12-17T08:49:12.523000",
          "content": "<p>I also forget that you can freeze a certain amount of layers if you use a pretrained model. This will also help reduce the amount of Vram used when training -&gt; <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/372550\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/372550</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2082027,
          "author_name": "Lau2664",
          "author_url": "",
          "post_date": "2023-01-01T02:36:39.783000",
          "content": "<p>Hi! <a href=\"https://www.kaggle.com/nathanlaoue\" target=\"_blank\">@nathanlaoue</a>. Thanks for replying! I got a problem when using large image size: I used TPU to train EfficientNetB4 with input size of 768x768, and the batch size is 128. The loss and pf1 score changed very slowly during training, sometimes they are even unchanged. Does that indicate that my model has some problems in convergence? How can I fix it?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2082190,
              "author_name": "Natyu",
              "author_url": "",
              "post_date": "2023-01-01T08:51:35.287000",
              "content": "<p>Hi! Do you freeze layer ? Maybe you freeze to much layer and the model can not update the layers that have an impact on learning to detect cancer. Also very basic answer but maybe your learning rate is to low.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2082252,
              "author_name": "Lau2664",
              "author_url": "",
              "post_date": "2023-01-01T10:54:12.917000",
              "content": "<p>I didn't freeze any layer. The whole model is trainable. I also used a cosine learning rate scheduler, the lr is changing during the training.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2082494,
              "author_name": "Natyu",
              "author_url": "",
              "post_date": "2023-01-01T16:18:28.810000",
              "content": "<p>For how many epochs did you train ? Maybe you didn't wait enough. Also this can be caused by a bug in your code. If you are using PyTorch maybe you have an error in your training loop causing the weights to not update the way they should.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2082098": "You can enable gradient checkpointing like this:\n```\nimport timm\nmodel = timm.create_model('efficientnet_b0')\nmodel.set_grad_checkpointing(enable=True)\n```\n\nFor more efficient use of gradient checkpointing you can go to ```timm/models/efficientnet.py```, check ```checkpoint_seq``` method that has ```every``` argument, you can set it to something different from 1 to speed up training (it'll use more memory though, tweak it to use as much VRAM as possible)",
    "2068443": "Gradient checkpointing is also very effective in saving VRAM.",
    "2064301": "There are many ways to do this.\n\n- Use smaller models.\n- Lower batch size.\n- Use autocast (fp16) if you are using pytorch. The data size is reduced by half compared to the normally used fp32.\n- Use TPU if using tensorflow. Because the memory size of TPU is very large.\n\nThese are just examples, and there may be many more ways to do it, but the simple ways I came up with are above.",
    "2067595": "Hi !\n\nI recently moved to 1024x1024 for training and this is what helped me to not get OOM :\n- Use the T4x2 gpus in kaggle notebook. This allow to get 30Gb of Vram instead of 15Gb with P100.\n- Reduce batch size and use gradient accumulation. Reducing the batch size allow to use less Vram, but if the batch size is very small (4 for example) this will have a big impact on convergence. Use gradient accumulation to get better results.\n- EfficientNetB4 is a big model. Maybe you can try a smaller version like EfficientNetB0 or even EfficientNetV2-S (for small)\n\nHope this will help !",
    "2064261": "Hello guys! I trained my model (efficientnetB4) with tensorflow on GPU P100 with image size of 256X256, and when I use a bigger image size (512/1024) the OOM occurs. I noticed that many of you are using very large image size like 2048, how did you achieve that without OOM ?"
  }
}