{
  "id": 414700,
  "title": "For Those Who Consider Big Ensembles (Not)",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/414700",
  "author_name": "SSS",
  "post_date": "2023-06-02T19:37:39.639000",
  "votes": 18,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Dear all, </p>\n<p>Having tried to average predictions for two of my models, I faced a <strong>Notebook Out Of Memory error</strong>.<br>\nThen it came to me that we got only 13 GB RAM, silly me. </p>\n<p>I came across some troubleshooting tips <a href=\"https://www.kaggle.com/getting-started/188347\" target=\"_blank\">here</a>, but unfortunately, they weren't applicable to my specific situation.</p>\n<p>However, the fix was quite simple I send everything to CUDA…</p>\n<h2><strong>Here is a pseudocode.</strong></h2>\n<pre><code>test_preds6 = torch.cat(get_predictions(ckpt_exp6_path, test_loader6)).cuda()\ntest_preds8 = torch.cat(get_predictions(ckpt_exp8_path, test_loader8)).cuda()\ntest_preds8 = test_preds6* + test_preds8*\n</code></pre>\n<p>It's both frustrating and exciting that memory optimization techniques are required to overcome such limitations. While they may limit us from employing our lovely \"kaggle-tricks\" of stacking numerous models, though, it makes me think about efficiency.</p>\n<h2><strong>Hope you avoid the failed submissions after reading this.</strong></h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F8aef3afea124231f700e74261fc10b1a%2Fmemory.png?generation=1685734642198202&amp;alt=media\" alt=\"\"></p>\n<p>Have fun!</p>",
  "messages": [
    {
      "id": 2285539,
      "postDate": "2023-06-02T19:37:39.640Z",
      "content": "<p>Dear all, </p>\n<p>Having tried to average predictions for two of my models, I faced a <strong>Notebook Out Of Memory error</strong>.<br>\nThen it came to me that we got only 13 GB RAM, silly me. </p>\n<p>I came across some troubleshooting tips <a href=\"https://www.kaggle.com/getting-started/188347\" target=\"_blank\">here</a>, but unfortunately, they weren't applicable to my specific situation.</p>\n<p>However, the fix was quite simple I send everything to CUDA…</p>\n<h2><strong>Here is a pseudocode.</strong></h2>\n<pre><code>test_preds6 = torch.cat(get_predictions(ckpt_exp6_path, test_loader6)).cuda()\ntest_preds8 = torch.cat(get_predictions(ckpt_exp8_path, test_loader8)).cuda()\ntest_preds8 = test_preds6* + test_preds8*\n</code></pre>\n<p>It's both frustrating and exciting that memory optimization techniques are required to overcome such limitations. While they may limit us from employing our lovely \"kaggle-tricks\" of stacking numerous models, though, it makes me think about efficiency.</p>\n<h2><strong>Hope you avoid the failed submissions after reading this.</strong></h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F8aef3afea124231f700e74261fc10b1a%2Fmemory.png?generation=1685734642198202&amp;alt=media\" alt=\"\"></p>\n<p>Have fun!</p>",
      "rawMarkdown": "Dear all, \n\nHaving tried to average predictions for two of my models, I faced a **Notebook Out Of Memory error**.\nThen it came to me that we got only 13 GB RAM, silly me. \n\nI came across some troubleshooting tips [here](https://www.kaggle.com/getting-started/188347), but unfortunately, they weren't applicable to my specific situation.\n\nHowever, the fix was quite simple I send everything to CUDA...\n\n##  **Here is a pseudocode.**\n```python\ntest_preds6 = torch.cat(get_predictions(ckpt_exp6_path, test_loader6)).cuda()\ntest_preds8 = torch.cat(get_predictions(ckpt_exp8_path, test_loader8)).cuda()\ntest_preds8 = test_preds6*0.7 + test_preds8*0.3\n```\n\nIt's both frustrating and exciting that memory optimization techniques are required to overcome such limitations. While they may limit us from employing our lovely \"kaggle-tricks\" of stacking numerous models, though, it makes me think about efficiency.\n\n##  **Hope you avoid the failed submissions after reading this.**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F8aef3afea124231f700e74261fc10b1a%2Fmemory.png?generation=1685734642198202&alt=media)\n\nHave fun!",
      "votes": 17
    },
    {
      "id": 2286448,
      "postDate": "2023-06-03T13:18:49.263Z",
      "content": "<p>save each predictions as file (for example out = (torch.sigmoid(logits)*255).astype(np.uint8)) to /kaggle/working/your-predictions-dir, and average everything in the end</p>",
      "rawMarkdown": "save each predictions as file (for example out = (torch.sigmoid(logits)*255).astype(np.uint8)) to /kaggle/working/your-predictions-dir, and average everything in the end",
      "votes": 5,
      "replies": [
        {
          "id": 2286488,
          "postDate": "2023-06-03T13:53:32.877Z",
          "content": "<p>..and dont forget to clean-up the directory once averaging and submission file generation is done ..</p>",
          "rawMarkdown": "..and dont forget to clean-up the directory once averaging and submission file generation is done ..",
          "votes": 2,
          "replies": [
            {
              "id": 2287320,
              "postDate": "2023-06-04T10:33:46.947Z",
              "content": "<p>Hey <a href=\"https://www.kaggle.com/phoenix9032\" target=\"_blank\">@phoenix9032</a> <br>\nI don't quite understand what would happen if that wasn't done, would you mind explaining it quickly please ?</p>",
              "rawMarkdown": "Hey @phoenix9032 \nI don't quite understand what would happen if that wasn't done, would you mind explaining it quickly please ?"
            },
            {
              "id": 2287340,
              "postDate": "2023-06-04T11:04:58.350Z",
              "content": "<p>Your submission might give \"Submission file not found\" error , if the working directory is not clean</p>",
              "rawMarkdown": "Your submission might give \"Submission file not found\" error , if the working directory is not clean",
              "votes": 1
            },
            {
              "id": 2287977,
              "postDate": "2023-06-05T05:02:39.063Z",
              "content": "<p>Are you sure ? I put all the files in there and then call them with a dataloader, and I don't get any issues..</p>",
              "rawMarkdown": "Are you sure ? I put all the files in there and then call them with a dataloader, and I don't get any issues..",
              "votes": 1
            },
            {
              "id": 2291291,
              "postDate": "2023-06-07T12:59:55.420Z",
              "content": "<p>this error comes from time to time. <br>\nso it is better to clean up :) </p>",
              "rawMarkdown": "this error comes from time to time. \nso it is better to clean up :) ",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2285751,
      "postDate": "2023-06-03T01:07:50.613Z",
      "content": "<p>Didn't even consider RAM ! I was thinking about 150models ensemble seeing the amount of time on GPU we had, silly me too..</p>",
      "rawMarkdown": "Didn't even consider RAM ! I was thinking about 150models ensemble seeing the amount of time on GPU we had, silly me too..",
      "votes": 1,
      "replies": [
        {
          "id": 2285775,
          "postDate": "2023-06-03T01:56:06.443Z",
          "content": "<p>yes, it seems like the best one (maybe few) model competition + pre/postprocessing magic limited by our abilities only.</p>",
          "rawMarkdown": "yes, it seems like the best one (maybe few) model competition + pre/postprocessing magic limited by our abilities only.",
          "votes": 1,
          "replies": [
            {
              "id": 2285784,
              "postDate": "2023-06-03T02:09:12.983Z",
              "content": "<p>update: ha, it's just came to me that we can use generators inside the eval loop, blend predictions of multiple models there and add them to the submission file by applying rle right away :D. </p>",
              "rawMarkdown": "update: ha, it's just came to me that we can use generators inside the eval loop, blend predictions of multiple models there and add them to the submission file by applying rle right away :D. "
            },
            {
              "id": 2285816,
              "postDate": "2023-06-03T02:49:57.533Z",
              "content": "<p>yeah I think this can be tackled in the loop rather than saving everything and ensembling later, but I think 150 is a bit too much to hope for aha</p>",
              "rawMarkdown": "yeah I think this can be tackled in the loop rather than saving everything and ensembling later, but I think 150 is a bit too much to hope for aha"
            }
          ]
        }
      ]
    },
    {
      "id": 2287806,
      "postDate": "2023-06-04T23:30:46.223Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2286448,
      "author_name": "Kostiantyn Maksymov",
      "author_url": "",
      "post_date": "2023-06-03T13:18:49.263000",
      "content": "<p>save each predictions as file (for example out = (torch.sigmoid(logits)*255).astype(np.uint8)) to /kaggle/working/your-predictions-dir, and average everything in the end</p>",
      "votes": 5,
      "replies": [
        {
          "id": 2286488,
          "author_name": "Nirjhar Roy",
          "author_url": "",
          "post_date": "2023-06-03T13:53:32.877000",
          "content": "<p>..and dont forget to clean-up the directory once averaging and submission file generation is done ..</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2287320,
              "author_name": "JEANMPIA",
              "author_url": "",
              "post_date": "2023-06-04T10:33:46.947000",
              "content": "<p>Hey <a href=\"https://www.kaggle.com/phoenix9032\" target=\"_blank\">@phoenix9032</a> <br>\nI don't quite understand what would happen if that wasn't done, would you mind explaining it quickly please ?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2287340,
              "author_name": "Nirjhar Roy",
              "author_url": "",
              "post_date": "2023-06-04T11:04:58.350000",
              "content": "<p>Your submission might give \"Submission file not found\" error , if the working directory is not clean</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2287977,
              "author_name": "JEANMPIA",
              "author_url": "",
              "post_date": "2023-06-05T05:02:39.063000",
              "content": "<p>Are you sure ? I put all the files in there and then call them with a dataloader, and I don't get any issues..</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2291291,
              "author_name": "Igor Krashenyi",
              "author_url": "",
              "post_date": "2023-06-07T12:59:55.420000",
              "content": "<p>this error comes from time to time. <br>\nso it is better to clean up :) </p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2285751,
      "author_name": "JEANMPIA",
      "author_url": "",
      "post_date": "2023-06-03T01:07:50.613000",
      "content": "<p>Didn't even consider RAM ! I was thinking about 150models ensemble seeing the amount of time on GPU we had, silly me too..</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2285775,
          "author_name": "SSS",
          "author_url": "",
          "post_date": "2023-06-03T01:56:06.443000",
          "content": "<p>yes, it seems like the best one (maybe few) model competition + pre/postprocessing magic limited by our abilities only.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2285784,
              "author_name": "SSS",
              "author_url": "",
              "post_date": "2023-06-03T02:09:12.983000",
              "content": "<p>update: ha, it's just came to me that we can use generators inside the eval loop, blend predictions of multiple models there and add them to the submission file by applying rle right away :D. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2285816,
              "author_name": "JEANMPIA",
              "author_url": "",
              "post_date": "2023-06-03T02:49:57.533000",
              "content": "<p>yeah I think this can be tackled in the loop rather than saving everything and ensembling later, but I think 150 is a bit too much to hope for aha</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2287806,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-06-04T23:30:46.223000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2285539": "Dear all, \n\nHaving tried to average predictions for two of my models, I faced a **Notebook Out Of Memory error**.\nThen it came to me that we got only 13 GB RAM, silly me. \n\nI came across some troubleshooting tips [here](https://www.kaggle.com/getting-started/188347), but unfortunately, they weren't applicable to my specific situation.\n\nHowever, the fix was quite simple I send everything to CUDA...\n\n##  **Here is a pseudocode.**\n```python\ntest_preds6 = torch.cat(get_predictions(ckpt_exp6_path, test_loader6)).cuda()\ntest_preds8 = torch.cat(get_predictions(ckpt_exp8_path, test_loader8)).cuda()\ntest_preds8 = test_preds6*0.7 + test_preds8*0.3\n```\n\nIt's both frustrating and exciting that memory optimization techniques are required to overcome such limitations. While they may limit us from employing our lovely \"kaggle-tricks\" of stacking numerous models, though, it makes me think about efficiency.\n\n##  **Hope you avoid the failed submissions after reading this.**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F8aef3afea124231f700e74261fc10b1a%2Fmemory.png?generation=1685734642198202&alt=media)\n\nHave fun!",
    "2286448": "save each predictions as file (for example out = (torch.sigmoid(logits)*255).astype(np.uint8)) to /kaggle/working/your-predictions-dir, and average everything in the end",
    "2285751": "Didn't even consider RAM ! I was thinking about 150models ensemble seeing the amount of time on GPU we had, silly me too..",
    "2287806": ""
  }
}