{
  "id": 72119,
  "title": "how to estimate the GPU memory needed for model?",
  "url": "/competitions/quickdraw-doodle-recognition/discussion/72119",
  "author_name": "Salaryman",
  "post_date": "2018-11-20T14:44:33.474000",
  "votes": 3,
  "comment_count": 7,
  "views": 0,
  "content": "<p>When train with large dataset I always encounter the GPU OOM problem, I don't know how to calculate the max batch size/ the amount of GPU memory i need for the model.\nFor example if according to model summary there are:</p>\n\n<p>Total params: 20,000,00\nTrainable params: 15,000,000</p>\n\n<p>And lets say the img tensor each is (512,512,3)</p>\n\n<p>What is the max batch size I can feed into 1080ti with 11G ram?</p>\n\n<p>Sorry for this rookie question, I tried google but come with so many confusing answers.\nHope someone can help!</p>\n\n<p>thanks so much</p>",
  "messages": [
    {
      "id": 424697,
      "postDate": "2018-11-20T14:44:33.473Z",
      "content": "<p>When train with large dataset I always encounter the GPU OOM problem, I don't know how to calculate the max batch size/ the amount of GPU memory i need for the model.\nFor example if according to model summary there are:</p>\n\n<p>Total params: 20,000,00\nTrainable params: 15,000,000</p>\n\n<p>And lets say the img tensor each is (512,512,3)</p>\n\n<p>What is the max batch size I can feed into 1080ti with 11G ram?</p>\n\n<p>Sorry for this rookie question, I tried google but come with so many confusing answers.\nHope someone can help!</p>\n\n<p>thanks so much</p>",
      "rawMarkdown": "When train with large dataset I always encounter the GPU OOM problem, I don't know how to calculate the max batch size/ the amount of GPU memory i need for the model.\nFor example if according to model summary there are:\n\nTotal params: 20,000,00\nTrainable params: 15,000,000\n\nAnd lets say the img tensor each is (512,512,3)\n\nWhat is the max batch size I can feed into 1080ti with 11G ram?\n\nSorry for this rookie question, I tried google but come with so many confusing answers.\nHope someone can help!\n\nthanks so much\n\n\n\n",
      "votes": 3
    },
    {
      "id": 425028,
      "postDate": "2018-11-21T02:52:56.213Z",
      "content": "<p>For a specific model, binary search the batch size is enough, till the batch size saturate the GPU memory.</p>",
      "rawMarkdown": "For a specific model, binary search the batch size is enough, till the batch size saturate the GPU memory.",
      "votes": 1,
      "replies": [
        {
          "id": 425396,
          "postDate": "2018-11-21T14:53:14.387Z",
          "content": "<p>thanks, just curious if there is more convenient way to better utilize the resources</p>",
          "rawMarkdown": "thanks, just curious if there is more convenient way to better utilize the resources"
        },
        {
          "id": 427302,
          "postDate": "2018-11-25T06:35:20.643Z",
          "content": "<p>I also use binary search to find the maximal batch size that below the border of OOM</p>",
          "rawMarkdown": "I also use binary search to find the maximal batch size that below the border of OOM"
        }
      ]
    },
    {
      "id": 424745,
      "postDate": "2018-11-20T15:54:44.980Z",
      "content": "<p>I hope someone can provide you with an answer, but thinking that batch size depends on more than just the model size.  For example, the size of the image is important in vision.</p>\n\n<p>I use the artillery method - shoot and adjust.  I start with batch size I am sure will work - run long enough to confirm that the first epoch works.  Stop - increase size of batch.  Run again - keeping increasing until OOM.\nOn my local PC this is a pretty painful process in Jupyter notebooks since death of the kernel seems to occur before an OOM error message. \nI use Microsoft Visual Code to run a py version of the file.  VS Code provides more information on state of GPU memory than Jupyter does.  When getting close to OOM VS Code will report an issue running out of memory \"trying to allocation some amount of GB.  The caller indicates that is is not a failure....\"  At this point I either run the script or back off on batch size to avoid that message.  I keep track of results and for a given problem can form a simple chart after a reasonable number of experiments that involve changes to batch size and image size.  The chart is than used in further work in Jupyter notebooks - but only for that specific challenge.</p>\n\n<p>Hopefully someone has a better tool than shoot and adjust.  Since almost everything I do uses tensorflow most of the tools that report GPU memory have not worked for me.</p>\n\n<p>Check out this paper for good starting point for your batch size - page 6.\n<a href=\"https://arxiv.org/pdf/1810.00736.pdf\">https://arxiv.org/pdf/1810.00736.pdf</a></p>",
      "rawMarkdown": "I hope someone can provide you with an answer, but thinking that batch size depends on more than just the model size.  For example, the size of the image is important in vision.\n\nI use the artillery method - shoot and adjust.  I start with batch size I am sure will work - run long enough to confirm that the first epoch works.  Stop - increase size of batch.  Run again - keeping increasing until OOM.\nOn my local PC this is a pretty painful process in Jupyter notebooks since death of the kernel seems to occur before an OOM error message. \nI use Microsoft Visual Code to run a py version of the file.  VS Code provides more information on state of GPU memory than Jupyter does.  When getting close to OOM VS Code will report an issue running out of memory \"trying to allocation some amount of GB.  The caller indicates that is is not a failure....\"  At this point I either run the script or back off on batch size to avoid that message.  I keep track of results and for a given problem can form a simple chart after a reasonable number of experiments that involve changes to batch size and image size.  The chart is than used in further work in Jupyter notebooks - but only for that specific challenge.\n\nHopefully someone has a better tool than shoot and adjust.  Since almost everything I do uses tensorflow most of the tools that report GPU memory have not worked for me.\n\nCheck out this paper for good starting point for your batch size - page 6.\nhttps://arxiv.org/pdf/1810.00736.pdf\n",
      "votes": 2,
      "replies": [
        {
          "id": 425395,
          "postDate": "2018-11-21T14:52:36.367Z",
          "content": "<p>thanks for your info jimmy. Will try your method in Microsoft Visual Code!</p>",
          "rawMarkdown": "thanks for your info jimmy. Will try your method in Microsoft Visual Code!"
        }
      ]
    },
    {
      "id": 425249,
      "postDate": "2018-11-21T10:36:42.087Z",
      "content": "<p>512 is too big for this competition.\n224/256 is enough to be suitable for the model and this task.</p>\n\n<p>I remember resnet50 has 20,000,00 \nfor 224 image size, I suppose a batch size lower than 64 is able to get started to run.</p>",
      "rawMarkdown": "512 is too big for this competition.\n224/256 is enough to be suitable for the model and this task.\n\nI remember resnet50 has 20,000,00 \nfor 224 image size, I suppose a batch size lower than 64 is able to get started to run.",
      "replies": [
        {
          "id": 425399,
          "postDate": "2018-11-21T14:54:46.343Z",
          "content": "<p>thanks, my question is not specific to this competition.  Just want to know if there is more convenient way to better utilize the GPU resources in future.</p>",
          "rawMarkdown": "thanks, my question is not specific to this competition.  Just want to know if there is more convenient way to better utilize the GPU resources in future."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 425028,
      "author_name": "wh1te",
      "author_url": "",
      "post_date": "2018-11-21T02:52:56.213000",
      "content": "<p>For a specific model, binary search the batch size is enough, till the batch size saturate the GPU memory.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 425396,
          "author_name": "Salaryman",
          "author_url": "",
          "post_date": "2018-11-21T14:53:14.387000",
          "content": "<p>thanks, just curious if there is more convenient way to better utilize the resources</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 427302,
          "author_name": "[he.ai]soulmachine",
          "author_url": "",
          "post_date": "2018-11-25T06:35:20.643000",
          "content": "<p>I also use binary search to find the maximal batch size that below the border of OOM</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 424745,
      "author_name": "PC Jimmmy",
      "author_url": "",
      "post_date": "2018-11-20T15:54:44.980000",
      "content": "<p>I hope someone can provide you with an answer, but thinking that batch size depends on more than just the model size.  For example, the size of the image is important in vision.</p>\n\n<p>I use the artillery method - shoot and adjust.  I start with batch size I am sure will work - run long enough to confirm that the first epoch works.  Stop - increase size of batch.  Run again - keeping increasing until OOM.\nOn my local PC this is a pretty painful process in Jupyter notebooks since death of the kernel seems to occur before an OOM error message. \nI use Microsoft Visual Code to run a py version of the file.  VS Code provides more information on state of GPU memory than Jupyter does.  When getting close to OOM VS Code will report an issue running out of memory \"trying to allocation some amount of GB.  The caller indicates that is is not a failure....\"  At this point I either run the script or back off on batch size to avoid that message.  I keep track of results and for a given problem can form a simple chart after a reasonable number of experiments that involve changes to batch size and image size.  The chart is than used in further work in Jupyter notebooks - but only for that specific challenge.</p>\n\n<p>Hopefully someone has a better tool than shoot and adjust.  Since almost everything I do uses tensorflow most of the tools that report GPU memory have not worked for me.</p>\n\n<p>Check out this paper for good starting point for your batch size - page 6.\n<a href=\"https://arxiv.org/pdf/1810.00736.pdf\">https://arxiv.org/pdf/1810.00736.pdf</a></p>",
      "votes": 2,
      "replies": [
        {
          "id": 425395,
          "author_name": "Salaryman",
          "author_url": "",
          "post_date": "2018-11-21T14:52:36.367000",
          "content": "<p>thanks for your info jimmy. Will try your method in Microsoft Visual Code!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 425249,
      "author_name": "Zineng Tang",
      "author_url": "",
      "post_date": "2018-11-21T10:36:42.087000",
      "content": "<p>512 is too big for this competition.\n224/256 is enough to be suitable for the model and this task.</p>\n\n<p>I remember resnet50 has 20,000,00 \nfor 224 image size, I suppose a batch size lower than 64 is able to get started to run.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 425399,
          "author_name": "Salaryman",
          "author_url": "",
          "post_date": "2018-11-21T14:54:46.343000",
          "content": "<p>thanks, my question is not specific to this competition.  Just want to know if there is more convenient way to better utilize the GPU resources in future.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "424697": "When train with large dataset I always encounter the GPU OOM problem, I don't know how to calculate the max batch size/ the amount of GPU memory i need for the model.\nFor example if according to model summary there are:\n\nTotal params: 20,000,00\nTrainable params: 15,000,000\n\nAnd lets say the img tensor each is (512,512,3)\n\nWhat is the max batch size I can feed into 1080ti with 11G ram?\n\nSorry for this rookie question, I tried google but come with so many confusing answers.\nHope someone can help!\n\nthanks so much\n\n\n\n",
    "425028": "For a specific model, binary search the batch size is enough, till the batch size saturate the GPU memory.",
    "424745": "I hope someone can provide you with an answer, but thinking that batch size depends on more than just the model size.  For example, the size of the image is important in vision.\n\nI use the artillery method - shoot and adjust.  I start with batch size I am sure will work - run long enough to confirm that the first epoch works.  Stop - increase size of batch.  Run again - keeping increasing until OOM.\nOn my local PC this is a pretty painful process in Jupyter notebooks since death of the kernel seems to occur before an OOM error message. \nI use Microsoft Visual Code to run a py version of the file.  VS Code provides more information on state of GPU memory than Jupyter does.  When getting close to OOM VS Code will report an issue running out of memory \"trying to allocation some amount of GB.  The caller indicates that is is not a failure....\"  At this point I either run the script or back off on batch size to avoid that message.  I keep track of results and for a given problem can form a simple chart after a reasonable number of experiments that involve changes to batch size and image size.  The chart is than used in further work in Jupyter notebooks - but only for that specific challenge.\n\nHopefully someone has a better tool than shoot and adjust.  Since almost everything I do uses tensorflow most of the tools that report GPU memory have not worked for me.\n\nCheck out this paper for good starting point for your batch size - page 6.\nhttps://arxiv.org/pdf/1810.00736.pdf\n",
    "425249": "512 is too big for this competition.\n224/256 is enough to be suitable for the model and this task.\n\nI remember resnet50 has 20,000,00 \nfor 224 image size, I suppose a batch size lower than 64 is able to get started to run."
  }
}