{
  "id": 372192,
  "title": "How can we train models with 1024px images using Kaggle?",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/372192",
  "author_name": "moth",
  "post_date": "2022-12-14T17:08:47.033000",
  "votes": 5,
  "comment_count": 13,
  "views": 0,
  "content": "<p>I was wondering if anyone has managed to train a model on 1024 images using only Kaggle kernels. I keep running into OOM 🤷🏼.</p>\n<p>I've tried using <strong>AMP (Automatic Mixed Precision)</strong> to reduce the memory limitations and <strong>gradient accumulation</strong> (decrease batch size but still backpropagate on \"larger\" batches). Haven't tried <strong>gradient checkpointing</strong> yet though.</p>\n<p>In addition, as some of these tricks make training much longer it seems unfeasible to train a model on 1024 with the weekly hour limit. I guess that I may be able to run 3-5 experiments weekly :(</p>\n<p>Are there any more tricks to reduce VRAM limitations? </p>\n<p>My code:</p>\n<pre><code>...\n\nwith tqdm(train_loader, unit=\"train_batch\") as tqdm_train_loader:\n    for step, (idx, images, labels) in enumerate(tqdm_train_loader):\n        images = torch.tensor(images, device=device, dtype=torch.float32)\n        labels = torch.tensor(labels, device=device, dtype=torch.float32)\n        with torch.cuda.amp.autocast(enabled=True):\n            out = model(images)\n            loss = criterion(out, labels.unsqueeze(1))\n        if config.GRADIENT_ACCUMULATION_STEPS &gt; 1:\n            loss = loss / config.GRADIENT_ACCUMULATION_STEPS\n        scaler.scale(loss).backward() # backward propagation pass\n        if (step + 1) % config.GRADIENT_ACCUMULATION_STEPS == 0:\n            scaler.step(optimizer) # update optimizer parameters\n            scaler.update()\n            optimizer.zero_grad()\n            ...\n</code></pre>\n<p><strong>Edit 1:</strong> TPU accelerates training but not resolved OOM. Does not require much code refactoring but requires a lot of configuration regarding conflicting package versions and memory handling via <code>gc.collect()</code>.<br>\n<strong>Edit 2:</strong> Freezing layers does help both with OOM and computation time:<br>\nExperiment:</p>\n<ul>\n<li><strong>Architecture:</strong> EfficientNetB2.</li>\n<li><strong>Image resolution:</strong> 1024px.</li>\n<li><strong>Train Batch Size:</strong> 8 images per batch.</li>\n<li><strong>Number of Frozen Layers:</strong> 100 out of 300 (~33% of the model).</li>\n<li><strong>Consumed GPU's VRAM:</strong> 4.7 out of 15.9 GB.</li>\n<li><strong>Training time for 1 EPOCH:</strong> 43k images on ~1:03 hrs.</li>\n<li><strong>Gradient Accumulation:</strong> True. Every 4 steps.</li>\n</ul>",
  "messages": [
    {
      "id": 2065452,
      "postDate": "2022-12-14T17:08:47.033Z",
      "content": "<p>I was wondering if anyone has managed to train a model on 1024 images using only Kaggle kernels. I keep running into OOM 🤷🏼.</p>\n<p>I've tried using <strong>AMP (Automatic Mixed Precision)</strong> to reduce the memory limitations and <strong>gradient accumulation</strong> (decrease batch size but still backpropagate on \"larger\" batches). Haven't tried <strong>gradient checkpointing</strong> yet though.</p>\n<p>In addition, as some of these tricks make training much longer it seems unfeasible to train a model on 1024 with the weekly hour limit. I guess that I may be able to run 3-5 experiments weekly :(</p>\n<p>Are there any more tricks to reduce VRAM limitations? </p>\n<p>My code:</p>\n<pre><code>...\n\nwith tqdm(train_loader, unit=\"train_batch\") as tqdm_train_loader:\n    for step, (idx, images, labels) in enumerate(tqdm_train_loader):\n        images = torch.tensor(images, device=device, dtype=torch.float32)\n        labels = torch.tensor(labels, device=device, dtype=torch.float32)\n        with torch.cuda.amp.autocast(enabled=True):\n            out = model(images)\n            loss = criterion(out, labels.unsqueeze(1))\n        if config.GRADIENT_ACCUMULATION_STEPS &gt; 1:\n            loss = loss / config.GRADIENT_ACCUMULATION_STEPS\n        scaler.scale(loss).backward() # backward propagation pass\n        if (step + 1) % config.GRADIENT_ACCUMULATION_STEPS == 0:\n            scaler.step(optimizer) # update optimizer parameters\n            scaler.update()\n            optimizer.zero_grad()\n            ...\n</code></pre>\n<p><strong>Edit 1:</strong> TPU accelerates training but not resolved OOM. Does not require much code refactoring but requires a lot of configuration regarding conflicting package versions and memory handling via <code>gc.collect()</code>.<br>\n<strong>Edit 2:</strong> Freezing layers does help both with OOM and computation time:<br>\nExperiment:</p>\n<ul>\n<li><strong>Architecture:</strong> EfficientNetB2.</li>\n<li><strong>Image resolution:</strong> 1024px.</li>\n<li><strong>Train Batch Size:</strong> 8 images per batch.</li>\n<li><strong>Number of Frozen Layers:</strong> 100 out of 300 (~33% of the model).</li>\n<li><strong>Consumed GPU's VRAM:</strong> 4.7 out of 15.9 GB.</li>\n<li><strong>Training time for 1 EPOCH:</strong> 43k images on ~1:03 hrs.</li>\n<li><strong>Gradient Accumulation:</strong> True. Every 4 steps.</li>\n</ul>",
      "rawMarkdown": "I was wondering if anyone has managed to train a model on 1024 images using only Kaggle kernels. I keep running into OOM 🤷🏼.\n\n I've tried using **AMP (Automatic Mixed Precision)** to reduce the memory limitations and **gradient accumulation** (decrease batch size but still backpropagate on \"larger\" batches). Haven't tried **gradient checkpointing** yet though.\n\nIn addition, as some of these tricks make training much longer it seems unfeasible to train a model on 1024 with the weekly hour limit. I guess that I may be able to run 3-5 experiments weekly :(\n\nAre there any more tricks to reduce VRAM limitations? \n\nMy code:\n```\n...\n\nwith tqdm(train_loader, unit=\"train_batch\") as tqdm_train_loader:\n    for step, (idx, images, labels) in enumerate(tqdm_train_loader):\n        images = torch.tensor(images, device=device, dtype=torch.float32)\n        labels = torch.tensor(labels, device=device, dtype=torch.float32)\n        with torch.cuda.amp.autocast(enabled=True):\n            out = model(images)\n            loss = criterion(out, labels.unsqueeze(1))\n        if config.GRADIENT_ACCUMULATION_STEPS > 1:\n            loss = loss / config.GRADIENT_ACCUMULATION_STEPS\n        scaler.scale(loss).backward() # backward propagation pass\n        if (step + 1) % config.GRADIENT_ACCUMULATION_STEPS == 0:\n            scaler.step(optimizer) # update optimizer parameters\n            scaler.update()\n            optimizer.zero_grad()\n            ...\n```\n\n**Edit 1:** TPU accelerates training but not resolved OOM. Does not require much code refactoring but requires a lot of configuration regarding conflicting package versions and memory handling via `gc.collect()`.\n**Edit 2:** Freezing layers does help both with OOM and computation time:\nExperiment:\n- **Architecture:** EfficientNetB2.\n- **Image resolution:** 1024px.\n- **Train Batch Size:** 8 images per batch.\n- **Number of Frozen Layers:** 100 out of 300 (~33% of the model).\n- **Consumed GPU's VRAM:** 4.7 out of 15.9 GB.\n- **Training time for 1 EPOCH:** 43k images on ~1:03 hrs.\n- **Gradient Accumulation:** True. Every 4 steps.",
      "votes": 5
    },
    {
      "id": 2065638,
      "postDate": "2022-12-14T23:07:27.330Z",
      "content": "<p>When training on GPU in Kaggle notebooks, you can freeze layers to avoid memory error.</p>",
      "rawMarkdown": "When training on GPU in Kaggle notebooks, you can freeze layers to avoid memory error.",
      "votes": 4,
      "replies": [
        {
          "id": 2065661,
          "postDate": "2022-12-15T00:24:02.467Z",
          "content": "<p>my understanding is that when you freeze layers backpropagation does not affect those layers, wouldn't that mean you don't end up training those layers? This would decrease model performance. Of course, maybe there's a trade-off</p>",
          "rawMarkdown": "my understanding is that when you freeze layers backpropagation does not affect those layers, wouldn't that mean you don't end up training those layers? This would decrease model performance. Of course, maybe there's a trade-off",
          "votes": 2
        },
        {
          "id": 2065667,
          "postDate": "2022-12-15T00:30:57.203Z",
          "content": "<p>Sometimes freezing layers produces a more accurate model for Kaggle competition, sometimes not. We need to compute CV and LB to see.</p>\n<p>The first layers learn low level image features like lines, dots, circles etc. The last layer features learn high level image features like cat's faces, car tires etc (when trained on imagenet dataset). Many times the learnings of the low layers already generalized to Kaggle competition and does not need to be retrained. (So we can freeze beginning layers and train ending layers).</p>\n<p>At any rate, by freezing (some or all) layers it requires less GPU RAM during training because we do not need to save those activations during back propagation nor need to update those weights. This is a trick to train very large models using ordinary GPU.</p>\n<p>Note that Giba won 1st place in Kaggle's Pet competition freezing the entire model and not doing any training. He just took the last layer activations and trained a RAPIDS SVR on those embeddings. Write up <a href=\"https://www.kaggle.com/competitions/petfinder-pawpularity-score/discussion/301686\" target=\"_blank\">here</a></p>",
          "rawMarkdown": "Sometimes freezing layers produces a more accurate model for Kaggle competition, sometimes not. We need to compute CV and LB to see.\n\nThe first layers learn low level image features like lines, dots, circles etc. The last layer features learn high level image features like cat's faces, car tires etc (when trained on imagenet dataset). Many times the learnings of the low layers already generalized to Kaggle competition and does not need to be retrained. (So we can freeze beginning layers and train ending layers).\n\nAt any rate, by freezing (some or all) layers it requires less GPU RAM during training because we do not need to save those activations during back propagation nor need to update those weights. This is a trick to train very large models using ordinary GPU.\n\nNote that Giba won 1st place in Kaggle's Pet competition freezing the entire model and not doing any training. He just took the last layer activations and trained a RAPIDS SVR on those embeddings. Write up [here][1]\n\n[1]: https://www.kaggle.com/competitions/petfinder-pawpularity-score/discussion/301686",
          "votes": 13,
          "replies": [
            {
              "id": 2066170,
              "postDate": "2022-12-15T13:06:25.140Z",
              "content": "<p>Thanks Chris, I guess from what you say that first layers capture low-level representations so it makes sense to just fine-tune the upper layers. Must try</p>",
              "rawMarkdown": "Thanks Chris, I guess from what you say that first layers capture low-level representations so it makes sense to just fine-tune the upper layers. Must try",
              "votes": 2
            },
            {
              "id": 2067032,
              "postDate": "2022-12-16T09:56:45.007Z",
              "content": "<p>i wish there are some code implementations that can freeze say (random) 50% of the weights at each layer at each training iteration</p>",
              "rawMarkdown": "i wish there are some code implementations that can freeze say (random) 50% of the weights at each layer at each training iteration"
            },
            {
              "id": 2067298,
              "postDate": "2022-12-16T15:17:32.650Z",
              "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> The best I can come up with: at each epoch load your model, freeze randomly the number of layers you want. Train the model for 1 epoch and save the model's weights. On the next epoch load the model with the last epoch's weights and freeze again some layers in a new random way. Repeat. </p>\n<p>Pseudo-code (I assume you use PyTorch):</p>\n<pre><code>for epoch in epochs:\n\n    model = load_model_function(PATH_MODEL_WEIGHTS)\n\n    NUM_TOTAL_LAYERS = 300 # determine beforehand the model's total number of layers\n    NUM_FREEZE_LAYERS = 10 # the number of layers you want to freeze\n    MODEL_PARAMS = list(model.named_parameters())\n    LAYERS_TO_FREEZE = [random.randint(0, NUM_TOTAL_LAYERS) for iter in range(NUM_FREEZE_LAYERS)]\n    RANDOM_LAYERS = [MODEL_PARAMS[i] for i in LAYERS_TO_FREEZE]\n\n    for name, param in RANDOM_LAYERS:     \n        param.requires_grad = False\n\n    for step, (images, labels) in train_data_loader:\n        ...\n\n    torch.save(model.state_dict())\n</code></pre>",
              "rawMarkdown": "@hengck23 The best I can come up with: at each epoch load your model, freeze randomly the number of layers you want. Train the model for 1 epoch and save the model's weights. On the next epoch load the model with the last epoch's weights and freeze again some layers in a new random way. Repeat. \n\nPseudo-code (I assume you use PyTorch):\n```\nfor epoch in epochs:\n    \n    model = load_model_function(PATH_MODEL_WEIGHTS)\n    \n    NUM_TOTAL_LAYERS = 300 # determine beforehand the model's total number of layers\n    NUM_FREEZE_LAYERS = 10 # the number of layers you want to freeze\n    MODEL_PARAMS = list(model.named_parameters())\n    LAYERS_TO_FREEZE = [random.randint(0, NUM_TOTAL_LAYERS) for iter in range(NUM_FREEZE_LAYERS)]\n    RANDOM_LAYERS = [MODEL_PARAMS[i] for i in LAYERS_TO_FREEZE]\n\n    for name, param in RANDOM_LAYERS:     \n        param.requires_grad = False\n        \n    for step, (images, labels) in train_data_loader:\n        ...\n        \n    torch.save(model.state_dict())\n```"
            },
            {
              "id": 2067314,
              "postDate": "2022-12-16T15:33:19.873Z",
              "content": "<p>check figure 1 and 2 of the papers:<br>\n<a href=\"https://www.ecva.net/papers/eccv_2020/papers_ECCV/papers/123720188.pdf\" target=\"_blank\">https://www.ecva.net/papers/eccv_2020/papers_ECCV/papers/123720188.pdf</a></p>",
              "rawMarkdown": "check figure 1 and 2 of the papers:\nhttps://www.ecva.net/papers/eccv_2020/papers_ECCV/papers/123720188.pdf"
            },
            {
              "id": 2076121,
              "postDate": "2022-12-26T06:29:30.517Z",
              "content": "<p>You can easily do it with <code>requires_grad = False</code> on every batch. <br>\nThis is a known trick to boost performance, called <a href=\"https://arxiv.org/abs/1603.09382\" target=\"_blank\">stochastic depth</a>.</p>\n<p>I think that in the context of reducing memory this won't be just enough.. If you want to reduce the memory even more you might have to look at int8 training and some \"not fun at all\" tricks.. </p>",
              "rawMarkdown": "You can easily do it with `requires_grad = False` on every batch. \nThis is a known trick to boost performance, called [stochastic depth](https://arxiv.org/abs/1603.09382).\n\nI think that in the context of reducing memory this won't be just enough.. If you want to reduce the memory even more you might have to look at int8 training and some \"not fun at all\" tricks.. "
            }
          ]
        }
      ]
    },
    {
      "id": 2076250,
      "postDate": "2022-12-26T09:30:38.923Z",
      "content": "<p><a href=\"https://www.kaggle.com/alejopaullier\" target=\"_blank\">@alejopaullier</a> You could try ROI patches, this will allow you to keep original resolution and train on much smaller images.  eg:  train model which looks for ROI, extract them as patches (frequently quite small, masses don't seem get larger than 512x512) , and then train a classifier on those patches</p>\n<p>For experiments, you'll want to subset the train data into something small but hopefully fairly representative of the problem set and train/test on that with a local CV. Once you're satisfied you've made a significant improvement, train on the complete set. </p>",
      "rawMarkdown": "@alejopaullier You could try ROI patches, this will allow you to keep original resolution and train on much smaller images.  eg:  train model which looks for ROI, extract them as patches (frequently quite small, masses don't seem get larger than 512x512) , and then train a classifier on those patches\n\nFor experiments, you'll want to subset the train data into something small but hopefully fairly representative of the problem set and train/test on that with a local CV. Once you're satisfied you've made a significant improvement, train on the complete set. \n\n"
    },
    {
      "id": 2065453,
      "postDate": "2022-12-14T17:10:55.053Z",
      "content": "<p><a href=\"https://www.kaggle.com/alejopaullier\" target=\"_blank\">@alejopaullier</a> Have you tried this on the new TPU 1VM Accelerator?<br>\nIt should support torch xla and has considerable memory/cores.</p>",
      "rawMarkdown": "@alejopaullier Have you tried this on the new TPU 1VM Accelerator?\nIt should support torch xla and has considerable memory/cores.",
      "replies": [
        {
          "id": 2065458,
          "postDate": "2022-12-14T17:15:07.393Z",
          "content": "<p>Hi Dustin, thanks for the reply. I've never tried a TPU before! Guess it's something new to learn. </p>\n<p>Noob question, does it require a lot of code refactoring or is it as easy as changing the <code>torch.device()</code> from <code>cuda</code> to something like <code>tpu</code>?</p>",
          "rawMarkdown": "Hi Dustin, thanks for the reply. I've never tried a TPU before! Guess it's something new to learn. \n\nNoob question, does it require a lot of code refactoring or is it as easy as changing the `torch.device()` from `cuda` to something like `tpu`?"
        },
        {
          "id": 2065461,
          "postDate": "2022-12-14T17:18:32.993Z",
          "content": "<p><a href=\"https://www.kaggle.com/alejopaullier\" target=\"_blank\">@alejopaullier</a> I have a couple examples on the announcement post: <a href=\"https://www.kaggle.com/discussions/product-feedback/369338\" target=\"_blank\">https://www.kaggle.com/discussions/product-feedback/369338</a></p>\n<p>For the most part yes, though to take advantage of all 8 TPU cores you do have to use xmp.spawn as well. <br>\nI'm just a SWE though so my data science expertise is limited sorry!</p>",
          "rawMarkdown": "@alejopaullier I have a couple examples on the announcement post: https://www.kaggle.com/discussions/product-feedback/369338\n\nFor the most part yes, though to take advantage of all 8 TPU cores you do have to use xmp.spawn as well. \nI'm just a SWE though so my data science expertise is limited sorry!",
          "votes": 3
        },
        {
          "id": 2065462,
          "postDate": "2022-12-14T17:20:38.670Z",
          "content": "<p>Thanks for the learning source!</p>",
          "rawMarkdown": "Thanks for the learning source!"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2065638,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2022-12-14T23:07:27.330000",
      "content": "<p>When training on GPU in Kaggle notebooks, you can freeze layers to avoid memory error.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2065661,
          "author_name": "moth",
          "author_url": "",
          "post_date": "2022-12-15T00:24:02.467000",
          "content": "<p>my understanding is that when you freeze layers backpropagation does not affect those layers, wouldn't that mean you don't end up training those layers? This would decrease model performance. Of course, maybe there's a trade-off</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2065667,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-12-15T00:30:57.203000",
          "content": "<p>Sometimes freezing layers produces a more accurate model for Kaggle competition, sometimes not. We need to compute CV and LB to see.</p>\n<p>The first layers learn low level image features like lines, dots, circles etc. The last layer features learn high level image features like cat's faces, car tires etc (when trained on imagenet dataset). Many times the learnings of the low layers already generalized to Kaggle competition and does not need to be retrained. (So we can freeze beginning layers and train ending layers).</p>\n<p>At any rate, by freezing (some or all) layers it requires less GPU RAM during training because we do not need to save those activations during back propagation nor need to update those weights. This is a trick to train very large models using ordinary GPU.</p>\n<p>Note that Giba won 1st place in Kaggle's Pet competition freezing the entire model and not doing any training. He just took the last layer activations and trained a RAPIDS SVR on those embeddings. Write up <a href=\"https://www.kaggle.com/competitions/petfinder-pawpularity-score/discussion/301686\" target=\"_blank\">here</a></p>",
          "votes": 13,
          "replies": [
            {
              "id": 2066170,
              "author_name": "moth",
              "author_url": "",
              "post_date": "2022-12-15T13:06:25.140000",
              "content": "<p>Thanks Chris, I guess from what you say that first layers capture low-level representations so it makes sense to just fine-tune the upper layers. Must try</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2067032,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2022-12-16T09:56:45.007000",
              "content": "<p>i wish there are some code implementations that can freeze say (random) 50% of the weights at each layer at each training iteration</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2067298,
              "author_name": "moth",
              "author_url": "",
              "post_date": "2022-12-16T15:17:32.650000",
              "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> The best I can come up with: at each epoch load your model, freeze randomly the number of layers you want. Train the model for 1 epoch and save the model's weights. On the next epoch load the model with the last epoch's weights and freeze again some layers in a new random way. Repeat. </p>\n<p>Pseudo-code (I assume you use PyTorch):</p>\n<pre><code>for epoch in epochs:\n\n    model = load_model_function(PATH_MODEL_WEIGHTS)\n\n    NUM_TOTAL_LAYERS = 300 # determine beforehand the model's total number of layers\n    NUM_FREEZE_LAYERS = 10 # the number of layers you want to freeze\n    MODEL_PARAMS = list(model.named_parameters())\n    LAYERS_TO_FREEZE = [random.randint(0, NUM_TOTAL_LAYERS) for iter in range(NUM_FREEZE_LAYERS)]\n    RANDOM_LAYERS = [MODEL_PARAMS[i] for i in LAYERS_TO_FREEZE]\n\n    for name, param in RANDOM_LAYERS:     \n        param.requires_grad = False\n\n    for step, (images, labels) in train_data_loader:\n        ...\n\n    torch.save(model.state_dict())\n</code></pre>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2067314,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2022-12-16T15:33:19.873000",
              "content": "<p>check figure 1 and 2 of the papers:<br>\n<a href=\"https://www.ecva.net/papers/eccv_2020/papers_ECCV/papers/123720188.pdf\" target=\"_blank\">https://www.ecva.net/papers/eccv_2020/papers_ECCV/papers/123720188.pdf</a></p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2076121,
              "author_name": "The Devastator",
              "author_url": "",
              "post_date": "2022-12-26T06:29:30.517000",
              "content": "<p>You can easily do it with <code>requires_grad = False</code> on every batch. <br>\nThis is a known trick to boost performance, called <a href=\"https://arxiv.org/abs/1603.09382\" target=\"_blank\">stochastic depth</a>.</p>\n<p>I think that in the context of reducing memory this won't be just enough.. If you want to reduce the memory even more you might have to look at int8 training and some \"not fun at all\" tricks.. </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2076250,
      "author_name": "@kaggleqrdl",
      "author_url": "",
      "post_date": "2022-12-26T09:30:38.923000",
      "content": "<p><a href=\"https://www.kaggle.com/alejopaullier\" target=\"_blank\">@alejopaullier</a> You could try ROI patches, this will allow you to keep original resolution and train on much smaller images.  eg:  train model which looks for ROI, extract them as patches (frequently quite small, masses don't seem get larger than 512x512) , and then train a classifier on those patches</p>\n<p>For experiments, you'll want to subset the train data into something small but hopefully fairly representative of the problem set and train/test on that with a local CV. Once you're satisfied you've made a significant improvement, train on the complete set. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2065453,
      "author_name": "Dustin",
      "author_url": "",
      "post_date": "2022-12-14T17:10:55.053000",
      "content": "<p><a href=\"https://www.kaggle.com/alejopaullier\" target=\"_blank\">@alejopaullier</a> Have you tried this on the new TPU 1VM Accelerator?<br>\nIt should support torch xla and has considerable memory/cores.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2065458,
          "author_name": "moth",
          "author_url": "",
          "post_date": "2022-12-14T17:15:07.393000",
          "content": "<p>Hi Dustin, thanks for the reply. I've never tried a TPU before! Guess it's something new to learn. </p>\n<p>Noob question, does it require a lot of code refactoring or is it as easy as changing the <code>torch.device()</code> from <code>cuda</code> to something like <code>tpu</code>?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2065461,
          "author_name": "Dustin",
          "author_url": "",
          "post_date": "2022-12-14T17:18:32.993000",
          "content": "<p><a href=\"https://www.kaggle.com/alejopaullier\" target=\"_blank\">@alejopaullier</a> I have a couple examples on the announcement post: <a href=\"https://www.kaggle.com/discussions/product-feedback/369338\" target=\"_blank\">https://www.kaggle.com/discussions/product-feedback/369338</a></p>\n<p>For the most part yes, though to take advantage of all 8 TPU cores you do have to use xmp.spawn as well. <br>\nI'm just a SWE though so my data science expertise is limited sorry!</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 2065462,
          "author_name": "moth",
          "author_url": "",
          "post_date": "2022-12-14T17:20:38.670000",
          "content": "<p>Thanks for the learning source!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2065452": "I was wondering if anyone has managed to train a model on 1024 images using only Kaggle kernels. I keep running into OOM 🤷🏼.\n\n I've tried using **AMP (Automatic Mixed Precision)** to reduce the memory limitations and **gradient accumulation** (decrease batch size but still backpropagate on \"larger\" batches). Haven't tried **gradient checkpointing** yet though.\n\nIn addition, as some of these tricks make training much longer it seems unfeasible to train a model on 1024 with the weekly hour limit. I guess that I may be able to run 3-5 experiments weekly :(\n\nAre there any more tricks to reduce VRAM limitations? \n\nMy code:\n```\n...\n\nwith tqdm(train_loader, unit=\"train_batch\") as tqdm_train_loader:\n    for step, (idx, images, labels) in enumerate(tqdm_train_loader):\n        images = torch.tensor(images, device=device, dtype=torch.float32)\n        labels = torch.tensor(labels, device=device, dtype=torch.float32)\n        with torch.cuda.amp.autocast(enabled=True):\n            out = model(images)\n            loss = criterion(out, labels.unsqueeze(1))\n        if config.GRADIENT_ACCUMULATION_STEPS > 1:\n            loss = loss / config.GRADIENT_ACCUMULATION_STEPS\n        scaler.scale(loss).backward() # backward propagation pass\n        if (step + 1) % config.GRADIENT_ACCUMULATION_STEPS == 0:\n            scaler.step(optimizer) # update optimizer parameters\n            scaler.update()\n            optimizer.zero_grad()\n            ...\n```\n\n**Edit 1:** TPU accelerates training but not resolved OOM. Does not require much code refactoring but requires a lot of configuration regarding conflicting package versions and memory handling via `gc.collect()`.\n**Edit 2:** Freezing layers does help both with OOM and computation time:\nExperiment:\n- **Architecture:** EfficientNetB2.\n- **Image resolution:** 1024px.\n- **Train Batch Size:** 8 images per batch.\n- **Number of Frozen Layers:** 100 out of 300 (~33% of the model).\n- **Consumed GPU's VRAM:** 4.7 out of 15.9 GB.\n- **Training time for 1 EPOCH:** 43k images on ~1:03 hrs.\n- **Gradient Accumulation:** True. Every 4 steps.",
    "2065638": "When training on GPU in Kaggle notebooks, you can freeze layers to avoid memory error.",
    "2076250": "@alejopaullier You could try ROI patches, this will allow you to keep original resolution and train on much smaller images.  eg:  train model which looks for ROI, extract them as patches (frequently quite small, masses don't seem get larger than 512x512) , and then train a classifier on those patches\n\nFor experiments, you'll want to subset the train data into something small but hopefully fairly representative of the problem set and train/test on that with a local CV. Once you're satisfied you've made a significant improvement, train on the complete set. \n\n",
    "2065453": "@alejopaullier Have you tried this on the new TPU 1VM Accelerator?\nIt should support torch xla and has considerable memory/cores."
  }
}