{
  "id": 377524,
  "title": "PyTorch TPU notebook huge slowdown",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/377524",
  "author_name": "Boris Polishchuk",
  "post_date": "2023-01-11T13:07:49.775000",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Dear all,</p>\n<p>I found that learning loop of my PyTorch notebook is running infinitely slowly when using TPU in comparison with GPU. Besides, it allows only few steps (not epochs!) depending on the batch size and eventually crashes due to excessive  memory consumption (typical error message is like: \"RuntimeError: Resource exhausted: From /job:tpu_worker/replica:0/task:0:<br>\nRan out of memory in memory space hbm. Used 24.61G of 15.98G hbm. Exceeded hbm capacity by 8.63G).</p>\n<p>I''m using some public TFRecords dataset of cropped mammograms (namely, /kaggle/input/rsna-preprocessing-tfrecords-640x512-dataset-pub). The fork of some TensorFlow notebook used this dataset is running just fine and swiftly.</p>\n<p>Do anybody in this competition have an experience of training PyTorch model on TPU and TFRecords and could give me any clue or example of public notebook?</p>\n<p>Thanks in advance!</p>",
  "messages": [
    {
      "id": 2095583,
      "postDate": "2023-01-11T13:34:01.590Z",
      "content": "<p>For me its taking around 30-35 mins for a single epoch with image size 1024x1024 with a batch size of 8*8 total for 8 cores , memory consumption is around  15.06376 gb  average for a single core and 1.170736 is free , the question is what batch size are u using here and are u using 8 cores ?</p>",
      "rawMarkdown": "For me its taking around 30-35 mins for a single epoch with image size 1024x1024 with a batch size of 8*8 total for 8 cores , memory consumption is around  15.06376 gb  average for a single core and 1.170736 is free , the question is what batch size are u using here and are u using 8 cores ?",
      "votes": 1,
      "replies": [
        {
          "id": 2096754,
          "postDate": "2023-01-12T08:37:00.063Z",
          "content": "<p>I use batch_size = 64 (TPU v3-8).<br>\nAs for a number of cores, I set nothing in my code and rely on default settings.</p>\n<p>For 1024x1024 it crashes on the first step due to memory consumption, for smaller sizes it crashes a bit later. :(</p>\n<p>Could you please reveal the parts of your notebook concerning TPU settings?</p>",
          "rawMarkdown": "I use batch_size = 64 (TPU v3-8).\nAs for a number of cores, I set nothing in my code and rely on default settings.\n\nFor 1024x1024 it crashes on the first step due to memory consumption, for smaller sizes it crashes a bit later. :(\n\nCould you please reveal the parts of your notebook concerning TPU settings?",
          "votes": 1,
          "replies": [
            {
              "id": 2096943,
              "postDate": "2023-01-12T11:25:25.737Z",
              "content": "<p>are you explicetly setting nprocs = 8 here in the map function xmp.spawn(_mp_fn, args=(FLAGS,), nprocs=8, start_method=\"fork\") ?? and also one more thing to keep in mind is you need to get xla device inside map function ( device = xm.xla_device() )  and also if you are setting batch size to 64 in your parallel loaders then it will set 64 batch size for each 8 cores so there you will saw OOM for 1024x1024x3 image size best is to set probably around batch size of 8 which is set to for each core so total will be 64 and also it will depend upon your model size right now i am using effv2_m it works fine with these settings . </p>",
              "rawMarkdown": "are you explicetly setting nprocs = 8 here in the map function xmp.spawn(_mp_fn, args=(FLAGS,), nprocs=8, start_method=\"fork\") ?? and also one more thing to keep in mind is you need to get xla device inside map function ( device = xm.xla_device() )  and also if you are setting batch size to 64 in your parallel loaders then it will set 64 batch size for each 8 cores so there you will saw OOM for 1024x1024x3 image size best is to set probably around batch size of 8 which is set to for each core so total will be 64 and also it will depend upon your model size right now i am using effv2_m it works fine with these settings . "
            }
          ]
        }
      ]
    },
    {
      "id": 2095559,
      "postDate": "2023-01-11T13:07:49.777Z",
      "content": "<p>Dear all,</p>\n<p>I found that learning loop of my PyTorch notebook is running infinitely slowly when using TPU in comparison with GPU. Besides, it allows only few steps (not epochs!) depending on the batch size and eventually crashes due to excessive  memory consumption (typical error message is like: \"RuntimeError: Resource exhausted: From /job:tpu_worker/replica:0/task:0:<br>\nRan out of memory in memory space hbm. Used 24.61G of 15.98G hbm. Exceeded hbm capacity by 8.63G).</p>\n<p>I''m using some public TFRecords dataset of cropped mammograms (namely, /kaggle/input/rsna-preprocessing-tfrecords-640x512-dataset-pub). The fork of some TensorFlow notebook used this dataset is running just fine and swiftly.</p>\n<p>Do anybody in this competition have an experience of training PyTorch model on TPU and TFRecords and could give me any clue or example of public notebook?</p>\n<p>Thanks in advance!</p>",
      "rawMarkdown": "Dear all,\n\nI found that learning loop of my PyTorch notebook is running infinitely slowly when using TPU in comparison with GPU. Besides, it allows only few steps (not epochs!) depending on the batch size and eventually crashes due to excessive  memory consumption (typical error message is like: \"RuntimeError: Resource exhausted: From /job:tpu_worker/replica:0/task:0:\nRan out of memory in memory space hbm. Used 24.61G of 15.98G hbm. Exceeded hbm capacity by 8.63G).\n\nI''m using some public TFRecords dataset of cropped mammograms (namely, /kaggle/input/rsna-preprocessing-tfrecords-640x512-dataset-pub). The fork of some TensorFlow notebook used this dataset is running just fine and swiftly.\n\nDo anybody in this competition have an experience of training PyTorch model on TPU and TFRecords and could give me any clue or example of public notebook?\n\nThanks in advance!",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 2095583,
      "author_name": "Shubham Thapa",
      "author_url": "",
      "post_date": "2023-01-11T13:34:01.590000",
      "content": "<p>For me its taking around 30-35 mins for a single epoch with image size 1024x1024 with a batch size of 8*8 total for 8 cores , memory consumption is around  15.06376 gb  average for a single core and 1.170736 is free , the question is what batch size are u using here and are u using 8 cores ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2096754,
          "author_name": "Boris Polishchuk",
          "author_url": "",
          "post_date": "2023-01-12T08:37:00.063000",
          "content": "<p>I use batch_size = 64 (TPU v3-8).<br>\nAs for a number of cores, I set nothing in my code and rely on default settings.</p>\n<p>For 1024x1024 it crashes on the first step due to memory consumption, for smaller sizes it crashes a bit later. :(</p>\n<p>Could you please reveal the parts of your notebook concerning TPU settings?</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2096943,
              "author_name": "Shubham Thapa",
              "author_url": "",
              "post_date": "2023-01-12T11:25:25.737000",
              "content": "<p>are you explicetly setting nprocs = 8 here in the map function xmp.spawn(_mp_fn, args=(FLAGS,), nprocs=8, start_method=\"fork\") ?? and also one more thing to keep in mind is you need to get xla device inside map function ( device = xm.xla_device() )  and also if you are setting batch size to 64 in your parallel loaders then it will set 64 batch size for each 8 cores so there you will saw OOM for 1024x1024x3 image size best is to set probably around batch size of 8 which is set to for each core so total will be 64 and also it will depend upon your model size right now i am using effv2_m it works fine with these settings . </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2095583": "For me its taking around 30-35 mins for a single epoch with image size 1024x1024 with a batch size of 8*8 total for 8 cores , memory consumption is around  15.06376 gb  average for a single core and 1.170736 is free , the question is what batch size are u using here and are u using 8 cores ?",
    "2095559": "Dear all,\n\nI found that learning loop of my PyTorch notebook is running infinitely slowly when using TPU in comparison with GPU. Besides, it allows only few steps (not epochs!) depending on the batch size and eventually crashes due to excessive  memory consumption (typical error message is like: \"RuntimeError: Resource exhausted: From /job:tpu_worker/replica:0/task:0:\nRan out of memory in memory space hbm. Used 24.61G of 15.98G hbm. Exceeded hbm capacity by 8.63G).\n\nI''m using some public TFRecords dataset of cropped mammograms (namely, /kaggle/input/rsna-preprocessing-tfrecords-640x512-dataset-pub). The fork of some TensorFlow notebook used this dataset is running just fine and swiftly.\n\nDo anybody in this competition have an experience of training PyTorch model on TPU and TFRecords and could give me any clue or example of public notebook?\n\nThanks in advance!"
  }
}