{
  "id": 154572,
  "title": "What learning rate are you using?",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/154572",
  "author_name": "Claudio Fanconi",
  "post_date": "2020-05-28T22:13:36.865000",
  "votes": 5,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Hi there again!</p>\n\n<p>I wanted to ask you guys what learning rates you guys are using to train transfer learning on your models? I originally used 1e-3, but after switching to a regression problem with MSELoss, I just had the feeling that my validation score would continue to fluctuate. \nThus, I believed that it's a too large learning rate, but I might be wrong.</p>\n\n<p>I ran the fastai learning rate finder, and it suggests me a loss of around 1e-8, which in my opinion seems incredibly low. (See image below). Nonetheless, I am still running the experiments.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2411547%2Fc390fea597d401133107ece73e680a85%2Fdownload.png?generation=1590703728798221&amp;alt=media\" alt=\"\"></p>\n\n<p>What learning rates are you guys using, if I may ask? And should optimizers like Rectified Adam, or Over9000 converge to an optimum, inspite different starting learning rates?</p>\n\n<p>Cheers and take care!</p>",
  "messages": [
    {
      "id": 865792,
      "postDate": "2020-05-28T22:13:36.867Z",
      "content": "<p>Hi there again!</p>\n\n<p>I wanted to ask you guys what learning rates you guys are using to train transfer learning on your models? I originally used 1e-3, but after switching to a regression problem with MSELoss, I just had the feeling that my validation score would continue to fluctuate. \nThus, I believed that it's a too large learning rate, but I might be wrong.</p>\n\n<p>I ran the fastai learning rate finder, and it suggests me a loss of around 1e-8, which in my opinion seems incredibly low. (See image below). Nonetheless, I am still running the experiments.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2411547%2Fc390fea597d401133107ece73e680a85%2Fdownload.png?generation=1590703728798221&amp;alt=media\" alt=\"\"></p>\n\n<p>What learning rates are you guys using, if I may ask? And should optimizers like Rectified Adam, or Over9000 converge to an optimum, inspite different starting learning rates?</p>\n\n<p>Cheers and take care!</p>",
      "rawMarkdown": "Hi there again!\n\nI wanted to ask you guys what learning rates you guys are using to train transfer learning on your models? I originally used 1e-3, but after switching to a regression problem with MSELoss, I just had the feeling that my validation score would continue to fluctuate. \nThus, I believed that it's a too large learning rate, but I might be wrong.\n\nI ran the fastai learning rate finder, and it suggests me a loss of around 1e-8, which in my opinion seems incredibly low. (See image below). Nonetheless, I am still running the experiments.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2411547%2Fc390fea597d401133107ece73e680a85%2Fdownload.png?generation=1590703728798221&amp;alt=media)\n\nWhat learning rates are you guys using, if I may ask? And should optimizers like Rectified Adam, or Over9000 converge to an optimum, inspite different starting learning rates?\n\nCheers and take care!\n",
      "votes": 5
    },
    {
      "id": 866697,
      "postDate": "2020-05-29T15:57:57.470Z",
      "content": "<p>I currently use 1e-4 for backbone and 1e-3 for head. 1e-3 overall gives me bad result and 1e-4 is just slower. Can also try 1e-5 1e-3 I think.</p>\n\n<p>In general the recommendation from lrfind be it fastai's or pytorch lightning's are awful because they are based on an automated analysis of the graph that is very naive. You should look at the graph and pick manually.</p>",
      "rawMarkdown": "I currently use 1e-4 for backbone and 1e-3 for head. 1e-3 overall gives me bad result and 1e-4 is just slower. Can also try 1e-5 1e-3 I think.\n\nIn general the recommendation from lrfind be it fastai's or pytorch lightning's are awful because they are based on an automated analysis of the graph that is very naive. You should look at the graph and pick manually.",
      "votes": 4,
      "replies": [
        {
          "id": 866726,
          "postDate": "2020-05-29T16:24:33.323Z",
          "content": "<p>Thank you very much for your recommendation. That sounds reasonable - i had no considered splitting the LR for backbone and head, but it makes completely sense!</p>",
          "rawMarkdown": "Thank you very much for your recommendation. That sounds reasonable - i had no considered splitting the LR for backbone and head, but it makes completely sense!"
        }
      ]
    },
    {
      "id": 869391,
      "postDate": "2020-06-01T01:47:00.853Z",
      "content": "<p>I believe 3e-4 is the best LR for Adam optim.</p>",
      "rawMarkdown": "I believe 3e-4 is the best LR for Adam optim.",
      "votes": 1
    },
    {
      "id": 865824,
      "postDate": "2020-05-28T23:12:29.880Z",
      "content": "<p>Its highly dependent on network ... and other stuff... It is hard to give advice. Bet I use somewhere between <code>1e-3</code> to <code>1e-4</code></p>",
      "rawMarkdown": "Its highly dependent on network ... and other stuff... It is hard to give advice. Bet I use somewhere between `1e-3` to `1e-4`",
      "votes": 2,
      "replies": [
        {
          "id": 865839,
          "postDate": "2020-05-28T23:37:02.240Z",
          "content": "<p>That is true! Currently I am mostly testing with an EffNetB0. I am still experimenting a bit around.\n1e-8 converged way too slow. Whereas 1e-2, had low train loss, but the valid loss almost exploded...</p>",
          "rawMarkdown": "That is true! Currently I am mostly testing with an EffNetB0. I am still experimenting a bit around.\n1e-8 converged way too slow. Whereas 1e-2, had low train loss, but the valid loss almost exploded...",
          "votes": 1
        },
        {
          "id": 865845,
          "postDate": "2020-05-28T23:45:19.600Z",
          "content": "<p><code>B0</code> is too big for my poor <code>gpu</code>... I am still doing everything on <code>resnet34</code> =) But good luck! =) </p>",
          "rawMarkdown": "`B0` is too big for my poor `gpu`... I am still doing everything on `resnet34` =) But good luck! =) ",
          "votes": 3
        },
        {
          "id": 866129,
          "postDate": "2020-05-29T06:15:56.867Z",
          "content": "<p>If I am not mistaken, EffnetB0 has significantly less parameters than Resnet34! The first one has around 5 mio. whereas the later has around 20 mio. parameters.\nWhich is kinda the whole point of EfficientNets. You should give it a shot! :) </p>",
          "rawMarkdown": "If I am not mistaken, EffnetB0 has significantly less parameters than Resnet34! The first one has around 5 mio. whereas the later has around 20 mio. parameters.\nWhich is kinda the whole point of EfficientNets. You should give it a shot! :) ",
          "votes": 1
        },
        {
          "id": 866213,
          "postDate": "2020-05-29T07:44:12.680Z",
          "content": "<p>EffnetB0 has less parameters and achieved better performances than ResNet34 (at least on Imagenet).\n<img src=\"https://raw.githubusercontent.com/tensorflow/tpu/master/models/official/efficientnet/g3doc/params.png\" width=\"400\">()</p>",
          "rawMarkdown": "EffnetB0 has less parameters and achieved better performances than ResNet34 (at least on Imagenet).\n<img src=\"https://raw.githubusercontent.com/tensorflow/tpu/master/models/official/efficientnet/g3doc/params.png\" width=\"400\">()",
          "votes": 1
        },
        {
          "id": 866536,
          "postDate": "2020-05-29T13:27:17.340Z",
          "content": "<p>Less parameters  doesn't mean correlates with efficient of gpu usage. Efficient net uses more expensive activation function and many other stuff ...  There is nice comparison and more in depth analysis here <a href=\"https://github.com/rwightman/pytorch-image-models/blob/master/notebooks/EffResNetComparison.ipynb\">https://github.com/rwightman/pytorch-image-models/blob/master/notebooks/EffResNetComparison.ipynb</a> . But yes in general they are more accurate than resnets =) </p>",
          "rawMarkdown": "Less parameters  doesn't mean correlates with efficient of gpu usage. Efficient net uses more expensive activation function and many other stuff ...  There is nice comparison and more in depth analysis here https://github.com/rwightman/pytorch-image-models/blob/master/notebooks/EffResNetComparison.ipynb . But yes in general they are more accurate than resnets =) ",
          "votes": 6
        },
        {
          "id": 866543,
          "postDate": "2020-05-29T13:33:16.110Z",
          "content": "<p>That is true, the swish activation is more complex, as well as the mobile conv blocks! thanks for pointing this out! :)</p>",
          "rawMarkdown": "That is true, the swish activation is more complex, as well as the mobile conv blocks! thanks for pointing this out! :)",
          "votes": 1
        },
        {
          "id": 866839,
          "postDate": "2020-05-29T18:22:12.877Z",
          "content": "<p>I was not aware of that, thank you! :)</p>",
          "rawMarkdown": "I was not aware of that, thank you! :)",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 866697,
      "author_name": "Arnaud Roussel",
      "author_url": "",
      "post_date": "2020-05-29T15:57:57.470000",
      "content": "<p>I currently use 1e-4 for backbone and 1e-3 for head. 1e-3 overall gives me bad result and 1e-4 is just slower. Can also try 1e-5 1e-3 I think.</p>\n\n<p>In general the recommendation from lrfind be it fastai's or pytorch lightning's are awful because they are based on an automated analysis of the graph that is very naive. You should look at the graph and pick manually.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 866726,
          "author_name": "Claudio Fanconi",
          "author_url": "",
          "post_date": "2020-05-29T16:24:33.323000",
          "content": "<p>Thank you very much for your recommendation. That sounds reasonable - i had no considered splitting the LR for backbone and head, but it makes completely sense!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 869391,
      "author_name": "seefun",
      "author_url": "",
      "post_date": "2020-06-01T01:47:00.853000",
      "content": "<p>I believe 3e-4 is the best LR for Adam optim.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 865824,
      "author_name": "DrHB",
      "author_url": "",
      "post_date": "2020-05-28T23:12:29.880000",
      "content": "<p>Its highly dependent on network ... and other stuff... It is hard to give advice. Bet I use somewhere between <code>1e-3</code> to <code>1e-4</code></p>",
      "votes": 2,
      "replies": [
        {
          "id": 865839,
          "author_name": "Claudio Fanconi",
          "author_url": "",
          "post_date": "2020-05-28T23:37:02.240000",
          "content": "<p>That is true! Currently I am mostly testing with an EffNetB0. I am still experimenting a bit around.\n1e-8 converged way too slow. Whereas 1e-2, had low train loss, but the valid loss almost exploded...</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 865845,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-05-28T23:45:19.600000",
          "content": "<p><code>B0</code> is too big for my poor <code>gpu</code>... I am still doing everything on <code>resnet34</code> =) But good luck! =) </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 866129,
          "author_name": "Claudio Fanconi",
          "author_url": "",
          "post_date": "2020-05-29T06:15:56.867000",
          "content": "<p>If I am not mistaken, EffnetB0 has significantly less parameters than Resnet34! The first one has around 5 mio. whereas the later has around 20 mio. parameters.\nWhich is kinda the whole point of EfficientNets. You should give it a shot! :) </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 866213,
          "author_name": "Pasquale",
          "author_url": "",
          "post_date": "2020-05-29T07:44:12.680000",
          "content": "<p>EffnetB0 has less parameters and achieved better performances than ResNet34 (at least on Imagenet).\n<img src=\"https://raw.githubusercontent.com/tensorflow/tpu/master/models/official/efficientnet/g3doc/params.png\" width=\"400\">()</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 866536,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-05-29T13:27:17.340000",
          "content": "<p>Less parameters  doesn't mean correlates with efficient of gpu usage. Efficient net uses more expensive activation function and many other stuff ...  There is nice comparison and more in depth analysis here <a href=\"https://github.com/rwightman/pytorch-image-models/blob/master/notebooks/EffResNetComparison.ipynb\">https://github.com/rwightman/pytorch-image-models/blob/master/notebooks/EffResNetComparison.ipynb</a> . But yes in general they are more accurate than resnets =) </p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 866543,
          "author_name": "Claudio Fanconi",
          "author_url": "",
          "post_date": "2020-05-29T13:33:16.110000",
          "content": "<p>That is true, the swish activation is more complex, as well as the mobile conv blocks! thanks for pointing this out! :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 866839,
          "author_name": "Pasquale",
          "author_url": "",
          "post_date": "2020-05-29T18:22:12.877000",
          "content": "<p>I was not aware of that, thank you! :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "865792": "Hi there again!\n\nI wanted to ask you guys what learning rates you guys are using to train transfer learning on your models? I originally used 1e-3, but after switching to a regression problem with MSELoss, I just had the feeling that my validation score would continue to fluctuate. \nThus, I believed that it's a too large learning rate, but I might be wrong.\n\nI ran the fastai learning rate finder, and it suggests me a loss of around 1e-8, which in my opinion seems incredibly low. (See image below). Nonetheless, I am still running the experiments.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2411547%2Fc390fea597d401133107ece73e680a85%2Fdownload.png?generation=1590703728798221&amp;alt=media)\n\nWhat learning rates are you guys using, if I may ask? And should optimizers like Rectified Adam, or Over9000 converge to an optimum, inspite different starting learning rates?\n\nCheers and take care!\n",
    "866697": "I currently use 1e-4 for backbone and 1e-3 for head. 1e-3 overall gives me bad result and 1e-4 is just slower. Can also try 1e-5 1e-3 I think.\n\nIn general the recommendation from lrfind be it fastai's or pytorch lightning's are awful because they are based on an automated analysis of the graph that is very naive. You should look at the graph and pick manually.",
    "869391": "I believe 3e-4 is the best LR for Adam optim.",
    "865824": "Its highly dependent on network ... and other stuff... It is hard to give advice. Bet I use somewhere between `1e-3` to `1e-4`"
  }
}