{
  "id": 72503,
  "title": "Large Batch Sizes",
  "url": "/competitions/quickdraw-doodle-recognition/discussion/72503",
  "author_name": "phun",
  "post_date": "2018-11-23T23:45:55.610000",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I've read how large batch sizes can combat noise in the data.</p>\n\n<p>I've used gradient accumulation (Keras version here: <a href=\"https://github.com/keras-team/keras/issues/3556#issuecomment-440638517\">https://github.com/keras-team/keras/issues/3556#issuecomment-440638517</a>) across 8 batches of 60 images of 256x256 each batch. (Xception and ResNet50).</p>\n\n<p>How are others achieving large batch sizes? The GPU memory seems to max out at 16 GB.</p>",
  "messages": [
    {
      "id": 426819,
      "postDate": "2018-11-23T23:45:55.610Z",
      "content": "<p>I've read how large batch sizes can combat noise in the data.</p>\n\n<p>I've used gradient accumulation (Keras version here: <a href=\"https://github.com/keras-team/keras/issues/3556#issuecomment-440638517\">https://github.com/keras-team/keras/issues/3556#issuecomment-440638517</a>) across 8 batches of 60 images of 256x256 each batch. (Xception and ResNet50).</p>\n\n<p>How are others achieving large batch sizes? The GPU memory seems to max out at 16 GB.</p>",
      "rawMarkdown": "I've read how large batch sizes can combat noise in the data.\n\nI've used gradient accumulation (Keras version here: https://github.com/keras-team/keras/issues/3556#issuecomment-440638517) across 8 batches of 60 images of 256x256 each batch. (Xception and ResNet50).\n\nHow are others achieving large batch sizes? The GPU memory seems to max out at 16 GB.",
      "votes": 1
    },
    {
      "id": 427300,
      "postDate": "2018-11-25T06:24:07.763Z",
      "content": "<p>GPU memory consumption is linear with batch_size*width*width, so you have to sacrifice resolution for larger batch size. On Tesla V100, I tried many times and got the maximal batch size as the  following:\n<code>\nModel           Image Width   Batch Size\nMobileNet          128            512\nMobileNet           256          144\nMobileNetV2        128             256\nMobileNetV2        256             119\nResNet50            128           256\n</code></p>",
      "rawMarkdown": "GPU memory consumption is linear with batch_size*width*width, so you have to sacrifice resolution for larger batch size. On Tesla V100, I tried many times and got the maximal batch size as the  following:\n```\nModel           Image Width   Batch Size\nMobileNet          128            512\nMobileNet           256          144\nMobileNetV2        128             256\nMobileNetV2        256             119\nResNet50            128           256\n```",
      "votes": 2
    },
    {
      "id": 427287,
      "postDate": "2018-11-25T05:04:33.747Z",
      "content": "<p>Right now I'm sacrificing resolution for batch size. Currently 600 images at 128x128 on two GPUs. Could probably up that with fp16. My problems seem to be with convergence in general. I'm taking a conservative lrplateau strategy that's still working, bur really slow...</p>",
      "rawMarkdown": "Right now I'm sacrificing resolution for batch size. Currently 600 images at 128x128 on two GPUs. Could probably up that with fp16. My problems seem to be with convergence in general. I'm taking a conservative lrplateau strategy that's still working, bur really slow...",
      "replies": [
        {
          "id": 427505,
          "postDate": "2018-11-25T16:41:26.783Z",
          "content": "<p>try increasing the learning rate if you are using large batches. You can also reduce the image size further to 64*64</p>",
          "rawMarkdown": "try increasing the learning rate if you are using large batches. You can also reduce the image size further to 64*64",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 427300,
      "author_name": "[he.ai]soulmachine",
      "author_url": "",
      "post_date": "2018-11-25T06:24:07.763000",
      "content": "<p>GPU memory consumption is linear with batch_size*width*width, so you have to sacrifice resolution for larger batch size. On Tesla V100, I tried many times and got the maximal batch size as the  following:\n<code>\nModel           Image Width   Batch Size\nMobileNet          128            512\nMobileNet           256          144\nMobileNetV2        128             256\nMobileNetV2        256             119\nResNet50            128           256\n</code></p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 427287,
      "author_name": "Erik Gaasedelen",
      "author_url": "",
      "post_date": "2018-11-25T05:04:33.747000",
      "content": "<p>Right now I'm sacrificing resolution for batch size. Currently 600 images at 128x128 on two GPUs. Could probably up that with fp16. My problems seem to be with convergence in general. I'm taking a conservative lrplateau strategy that's still working, bur really slow...</p>",
      "votes": 0,
      "replies": [
        {
          "id": 427505,
          "author_name": "Vishnu Subramanian",
          "author_url": "",
          "post_date": "2018-11-25T16:41:26.783000",
          "content": "<p>try increasing the learning rate if you are using large batches. You can also reduce the image size further to 64*64</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "426819": "I've read how large batch sizes can combat noise in the data.\n\nI've used gradient accumulation (Keras version here: https://github.com/keras-team/keras/issues/3556#issuecomment-440638517) across 8 batches of 60 images of 256x256 each batch. (Xception and ResNet50).\n\nHow are others achieving large batch sizes? The GPU memory seems to max out at 16 GB.",
    "427300": "GPU memory consumption is linear with batch_size*width*width, so you have to sacrifice resolution for larger batch size. On Tesla V100, I tried many times and got the maximal batch size as the  following:\n```\nModel           Image Width   Batch Size\nMobileNet          128            512\nMobileNet           256          144\nMobileNetV2        128             256\nMobileNetV2        256             119\nResNet50            128           256\n```",
    "427287": "Right now I'm sacrificing resolution for batch size. Currently 600 images at 128x128 on two GPUs. Could probably up that with fp16. My problems seem to be with convergence in general. I'm taking a conservative lrplateau strategy that's still working, bur really slow..."
  }
}