{
  "id": 362769,
  "title": "what is your batch size and  LB?",
  "url": "/competitions/g2net-detecting-continuous-gravitational-waves/discussion/362769",
  "author_name": "dragon zhang",
  "post_date": "2022-10-29T02:12:10.592000",
  "votes": 6,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I use batch size 8 , 4 for bigger model on kaggle.</p>\n<p>if bigger ram and batch size can improve LB score much higher?</p>",
  "messages": [
    {
      "id": 2015615,
      "postDate": "2022-11-03T12:23:10.290Z",
      "content": "<p>Not really, high batch size does not mean greater results. <br>\nIn my experience, when the total number of instances is high, then a large batch size is good. But in small datasets like this, smaller batch sizes like 8 works very well. <br>\nIn general high batch size can lead to over-fitting and local optimum stacks especially in small datasets. <br>\nIn contrast, very small batch size e.g. 1 to 4 lead to unstable training.</p>\n<p>In conclusion, for me the 8 bs gave me the best results by average (for 800 instances after balancing the 600 init ones).</p>",
      "rawMarkdown": "Not really, high batch size does not mean greater results. \nIn my experience, when the total number of instances is high, then a large batch size is good. But in small datasets like this, smaller batch sizes like 8 works very well. \nIn general high batch size can lead to over-fitting and local optimum stacks especially in small datasets. \nIn contrast, very small batch size e.g. 1 to 4 lead to unstable training.\n\nIn conclusion, for me the 8 bs gave me the best results by average (for 800 instances after balancing the 600 init ones).",
      "votes": 5,
      "replies": [
        {
          "id": 2015726,
          "postDate": "2022-11-03T13:58:02.677Z",
          "content": "<p>But that is also very depended on the used learning rate. If you change your batch size you will also have to change your learning rate. Higher batch size means less weight changes to compensate that you will need to increase your learning rate.</p>\n<p>In my experience 32 batch size and 0.001 learning rate mostly delivered the best results.</p>",
          "rawMarkdown": "But that is also very depended on the used learning rate. If you change your batch size you will also have to change your learning rate. Higher batch size means less weight changes to compensate that you will need to increase your learning rate.\n\nIn my experience 32 batch size and 0.001 learning rate mostly delivered the best results.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2008380,
      "postDate": "2022-10-29T02:12:10.593Z",
      "content": "<p>I use batch size 8 , 4 for bigger model on kaggle.</p>\n<p>if bigger ram and batch size can improve LB score much higher?</p>",
      "rawMarkdown": "I use batch size 8 , 4 for bigger model on kaggle.\n\nif bigger ram and batch size can improve LB score much higher?\n\n",
      "votes": 6
    },
    {
      "id": 2011865,
      "postDate": "2022-10-31T21:40:34.087Z",
      "content": "<p>I'm using 16-32 also. It's all I can fit. I would use bigger if I could. I haven't tried <a href=\"https://www.deepspeed.ai/tutorials/zero/\" target=\"_blank\">this</a> though.</p>",
      "rawMarkdown": "I'm using 16-32 also. It's all I can fit. I would use bigger if I could. I haven't tried [this](https://www.deepspeed.ai/tutorials/zero/) though."
    },
    {
      "id": 2010571,
      "postDate": "2022-10-30T21:31:42.587Z",
      "content": "<p>I use 16-32 depending on the model size:</p>\n<blockquote>\n  <p>Friends dont let friends use minibatches larger than 32.</p>\n</blockquote>\n<p><a href=\"https://twitter.com/ylecun/status/989610208497360896\" target=\"_blank\">Quote - Yann LeCun</a></p>",
      "rawMarkdown": "I use 16-32 depending on the model size:\n\n> Friends dont let friends use minibatches larger than 32.\n\n[Quote - Yann LeCun](https://twitter.com/ylecun/status/989610208497360896)",
      "replies": [
        {
          "id": 2010693,
          "postDate": "2022-10-31T03:17:52.250Z",
          "content": "<p>thx for your reply.  </p>",
          "rawMarkdown": "thx for your reply.  "
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2015615,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-11-03T12:23:10.290000",
      "content": "<p>Not really, high batch size does not mean greater results. <br>\nIn my experience, when the total number of instances is high, then a large batch size is good. But in small datasets like this, smaller batch sizes like 8 works very well. <br>\nIn general high batch size can lead to over-fitting and local optimum stacks especially in small datasets. <br>\nIn contrast, very small batch size e.g. 1 to 4 lead to unstable training.</p>\n<p>In conclusion, for me the 8 bs gave me the best results by average (for 800 instances after balancing the 600 init ones).</p>",
      "votes": 5,
      "replies": [
        {
          "id": 2015726,
          "author_name": "Ali Abdin",
          "author_url": "",
          "post_date": "2022-11-03T13:58:02.677000",
          "content": "<p>But that is also very depended on the used learning rate. If you change your batch size you will also have to change your learning rate. Higher batch size means less weight changes to compensate that you will need to increase your learning rate.</p>\n<p>In my experience 32 batch size and 0.001 learning rate mostly delivered the best results.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2011865,
      "author_name": "Will Rice",
      "author_url": "",
      "post_date": "2022-10-31T21:40:34.087000",
      "content": "<p>I'm using 16-32 also. It's all I can fit. I would use bigger if I could. I haven't tried <a href=\"https://www.deepspeed.ai/tutorials/zero/\" target=\"_blank\">this</a> though.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2010571,
      "author_name": "Ali Abdin",
      "author_url": "",
      "post_date": "2022-10-30T21:31:42.587000",
      "content": "<p>I use 16-32 depending on the model size:</p>\n<blockquote>\n  <p>Friends dont let friends use minibatches larger than 32.</p>\n</blockquote>\n<p><a href=\"https://twitter.com/ylecun/status/989610208497360896\" target=\"_blank\">Quote - Yann LeCun</a></p>",
      "votes": 0,
      "replies": [
        {
          "id": 2010693,
          "author_name": "dragon zhang",
          "author_url": "",
          "post_date": "2022-10-31T03:17:52.250000",
          "content": "<p>thx for your reply.  </p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2015615": "Not really, high batch size does not mean greater results. \nIn my experience, when the total number of instances is high, then a large batch size is good. But in small datasets like this, smaller batch sizes like 8 works very well. \nIn general high batch size can lead to over-fitting and local optimum stacks especially in small datasets. \nIn contrast, very small batch size e.g. 1 to 4 lead to unstable training.\n\nIn conclusion, for me the 8 bs gave me the best results by average (for 800 instances after balancing the 600 init ones).",
    "2008380": "I use batch size 8 , 4 for bigger model on kaggle.\n\nif bigger ram and batch size can improve LB score much higher?\n\n",
    "2011865": "I'm using 16-32 also. It's all I can fit. I would use bigger if I could. I haven't tried [this](https://www.deepspeed.ai/tutorials/zero/) though.",
    "2010571": "I use 16-32 depending on the model size:\n\n> Friends dont let friends use minibatches larger than 32.\n\n[Quote - Yann LeCun](https://twitter.com/ylecun/status/989610208497360896)"
  }
}