{
  "id": 185969,
  "title": "pytorch lightning tpu",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/185969",
  "author_name": "Artyom Lyan",
  "post_date": "2020-09-22T17:52:09.245000",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>I'm trying to make a pipeline with tpu and pytorch lightning, but run out of ram when I'm using 8 tpu cores. How to solve that problem? What is best practice working with pytorch and tpu (tf is possible the next step)</p>\n<p>thanks</p>",
  "messages": [
    {
      "id": 1024109,
      "postDate": "2020-09-23T16:50:45.847Z",
      "content": "<p>I'm thinking to use the same approach.</p>\n<p>We don't have batch size problems using TPU with Tensorflow but with Pytorch Lightning we do.<br>\nGiven the memory is the same available to use as the GPU, as a first approach, I would try to use a batch size at least 8 times lower on the TPU from the one I can run successfully with the GPU then I would to steadily increase. I did that in the Melanoma Competition and could run in the TPU with a little less than half of the GPU batch size.</p>\n<p>Some options to try in the <code>Trainer</code> if you are not already using:</p>\n<ul>\n<li><code>num_workers=1</code> when using the TPU, I don't recall the discussion explaining this but that is the practice</li>\n<li><code>precision=16</code> ( The new version of Pytorch Lightning support half-precision for TPU also )</li>\n<li><code>deterministic=False</code> </li>\n<li><code>benchmark=True</code> ( This makes the GPU training faster if images are all the same size, as are those in this competition,  according to the docs, I don't think it applies to TPU but it doesn't hurt trying.</li>\n</ul>\n<p>Let me know what works.</p>",
      "rawMarkdown": "I'm thinking to use the same approach.\n\nWe don't have batch size problems using TPU with Tensorflow but with Pytorch Lightning we do.\nGiven the memory is the same available to use as the GPU, as a first approach, I would try to use a batch size at least 8 times lower on the TPU from the one I can run successfully with the GPU then I would to steadily increase. I did that in the Melanoma Competition and could run in the TPU with a little less than half of the GPU batch size.\n\nSome options to try in the `Trainer` if you are not already using:\n- `num_workers=1` when using the TPU, I don't recall the discussion explaining this but that is the practice\n- `precision=16` ( The new version of Pytorch Lightning support half-precision for TPU also )\n- `deterministic=False` \n- `benchmark=True` ( This makes the GPU training faster if images are all the same size, as are those in this competition,  according to the docs, I don't think it applies to TPU but it doesn't hurt trying.\n\nLet me know what works.",
      "votes": 1,
      "replies": [
        {
          "id": 1024394,
          "postDate": "2020-09-23T20:25:44.417Z",
          "content": "<p>it's not failing on a single tpu core, so I believe this happening because pytorch lightning spawns 8 dataloaders which cause oom</p>",
          "rawMarkdown": "it's not failing on a single tpu core, so I believe this happening because pytorch lightning spawns 8 dataloaders which cause oom"
        },
        {
          "id": 1024437,
          "postDate": "2020-09-23T21:29:05.190Z",
          "content": "<p>I guess that's why the <code>num_workers=1</code> is necessary.</p>",
          "rawMarkdown": "I guess that's why the `num_workers=1` is necessary."
        }
      ]
    },
    {
      "id": 1022744,
      "postDate": "2020-09-22T17:52:09.247Z",
      "content": "<p>Hi everyone,</p>\n<p>I'm trying to make a pipeline with tpu and pytorch lightning, but run out of ram when I'm using 8 tpu cores. How to solve that problem? What is best practice working with pytorch and tpu (tf is possible the next step)</p>\n<p>thanks</p>",
      "rawMarkdown": "Hi everyone,\n\nI'm trying to make a pipeline with tpu and pytorch lightning, but run out of ram when I'm using 8 tpu cores. How to solve that problem? What is best practice working with pytorch and tpu (tf is possible the next step)\n\nthanks",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 1024109,
      "author_name": "Ronaldo S.A. Batista",
      "author_url": "",
      "post_date": "2020-09-23T16:50:45.847000",
      "content": "<p>I'm thinking to use the same approach.</p>\n<p>We don't have batch size problems using TPU with Tensorflow but with Pytorch Lightning we do.<br>\nGiven the memory is the same available to use as the GPU, as a first approach, I would try to use a batch size at least 8 times lower on the TPU from the one I can run successfully with the GPU then I would to steadily increase. I did that in the Melanoma Competition and could run in the TPU with a little less than half of the GPU batch size.</p>\n<p>Some options to try in the <code>Trainer</code> if you are not already using:</p>\n<ul>\n<li><code>num_workers=1</code> when using the TPU, I don't recall the discussion explaining this but that is the practice</li>\n<li><code>precision=16</code> ( The new version of Pytorch Lightning support half-precision for TPU also )</li>\n<li><code>deterministic=False</code> </li>\n<li><code>benchmark=True</code> ( This makes the GPU training faster if images are all the same size, as are those in this competition,  according to the docs, I don't think it applies to TPU but it doesn't hurt trying.</li>\n</ul>\n<p>Let me know what works.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1024394,
          "author_name": "Artyom Lyan",
          "author_url": "",
          "post_date": "2020-09-23T20:25:44.417000",
          "content": "<p>it's not failing on a single tpu core, so I believe this happening because pytorch lightning spawns 8 dataloaders which cause oom</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1024437,
          "author_name": "Ronaldo S.A. Batista",
          "author_url": "",
          "post_date": "2020-09-23T21:29:05.190000",
          "content": "<p>I guess that's why the <code>num_workers=1</code> is necessary.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1024109": "I'm thinking to use the same approach.\n\nWe don't have batch size problems using TPU with Tensorflow but with Pytorch Lightning we do.\nGiven the memory is the same available to use as the GPU, as a first approach, I would try to use a batch size at least 8 times lower on the TPU from the one I can run successfully with the GPU then I would to steadily increase. I did that in the Melanoma Competition and could run in the TPU with a little less than half of the GPU batch size.\n\nSome options to try in the `Trainer` if you are not already using:\n- `num_workers=1` when using the TPU, I don't recall the discussion explaining this but that is the practice\n- `precision=16` ( The new version of Pytorch Lightning support half-precision for TPU also )\n- `deterministic=False` \n- `benchmark=True` ( This makes the GPU training faster if images are all the same size, as are those in this competition,  according to the docs, I don't think it applies to TPU but it doesn't hurt trying.\n\nLet me know what works.",
    "1022744": "Hi everyone,\n\nI'm trying to make a pipeline with tpu and pytorch lightning, but run out of ram when I'm using 8 tpu cores. How to solve that problem? What is best practice working with pytorch and tpu (tf is possible the next step)\n\nthanks"
  }
}