{
  "id": 155886,
  "title": "Start very high QWK but Does not improve much....",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/155886",
  "author_name": "Harshit Sheoran",
  "post_date": "2020-06-03T12:17:50.012000",
  "votes": 0,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Disclaimer: I have already tried various combinations of @iafoss dataset, and I tried the models on my dataset too, even when I lower the learning rate a lot, the model does not go further so, I used annealing learning rate, start from 5e-5 then decrease 5e-6 every epoch, reset after 10 epochs, but did not get any improvement even after 30 epochs of training, the result was 0.76 but when I increase the learning rate to 1e-4 I get 0.76 on my first epoch and it increases to 0.80 in first 10 epochs.</p>\n\n<p>I have no idea why lowering the learning rate and making my model train for longer is making even worse results.</p>\n\n<p>Thank you for answering in advance.</p>",
  "messages": [
    {
      "id": 872811,
      "postDate": "2020-06-03T14:21:10.627Z",
      "content": "<p>If your learning rate is too small it is not surprising that the results are bad. Consider annealing, and all the things that go in an adaptive optimizer like Adam. Consider also that you probably have a randomly initialized head on top of a pretrained backbone. Small learning rates can quickly end up in a local minima.</p>\n\n<p>High learning rates allow for quicker exploration of the \"landscape\". After all that is the main argument for Leslie's OneCycle fit where you ramp it up to high values in order to end up in a \"flatter\" area of the landscape before doing the small adjustments with small learning rates.</p>",
      "rawMarkdown": "If your learning rate is too small it is not surprising that the results are bad. Consider annealing, and all the things that go in an adaptive optimizer like Adam. Consider also that you probably have a randomly initialized head on top of a pretrained backbone. Small learning rates can quickly end up in a local minima.\n\nHigh learning rates allow for quicker exploration of the \"landscape\". After all that is the main argument for Leslie's OneCycle fit where you ramp it up to high values in order to end up in a \"flatter\" area of the landscape before doing the small adjustments with small learning rates.",
      "votes": 1
    },
    {
      "id": 872676,
      "postDate": "2020-06-03T12:17:50.013Z",
      "content": "<p>Disclaimer: I have already tried various combinations of @iafoss dataset, and I tried the models on my dataset too, even when I lower the learning rate a lot, the model does not go further so, I used annealing learning rate, start from 5e-5 then decrease 5e-6 every epoch, reset after 10 epochs, but did not get any improvement even after 30 epochs of training, the result was 0.76 but when I increase the learning rate to 1e-4 I get 0.76 on my first epoch and it increases to 0.80 in first 10 epochs.</p>\n\n<p>I have no idea why lowering the learning rate and making my model train for longer is making even worse results.</p>\n\n<p>Thank you for answering in advance.</p>",
      "rawMarkdown": "Disclaimer: I have already tried various combinations of @iafoss dataset, and I tried the models on my dataset too, even when I lower the learning rate a lot, the model does not go further so, I used annealing learning rate, start from 5e-5 then decrease 5e-6 every epoch, reset after 10 epochs, but did not get any improvement even after 30 epochs of training, the result was 0.76 but when I increase the learning rate to 1e-4 I get 0.76 on my first epoch and it increases to 0.80 in first 10 epochs.\n\nI have no idea why lowering the learning rate and making my model train for longer is making even worse results.\n\nThank you for answering in advance."
    }
  ],
  "comments": [
    {
      "id": 872811,
      "author_name": "Arnaud Roussel",
      "author_url": "",
      "post_date": "2020-06-03T14:21:10.627000",
      "content": "<p>If your learning rate is too small it is not surprising that the results are bad. Consider annealing, and all the things that go in an adaptive optimizer like Adam. Consider also that you probably have a randomly initialized head on top of a pretrained backbone. Small learning rates can quickly end up in a local minima.</p>\n\n<p>High learning rates allow for quicker exploration of the \"landscape\". After all that is the main argument for Leslie's OneCycle fit where you ramp it up to high values in order to end up in a \"flatter\" area of the landscape before doing the small adjustments with small learning rates.</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "872811": "If your learning rate is too small it is not surprising that the results are bad. Consider annealing, and all the things that go in an adaptive optimizer like Adam. Consider also that you probably have a randomly initialized head on top of a pretrained backbone. Small learning rates can quickly end up in a local minima.\n\nHigh learning rates allow for quicker exploration of the \"landscape\". After all that is the main argument for Leslie's OneCycle fit where you ramp it up to high values in order to end up in a \"flatter\" area of the landscape before doing the small adjustments with small learning rates.",
    "872676": "Disclaimer: I have already tried various combinations of @iafoss dataset, and I tried the models on my dataset too, even when I lower the learning rate a lot, the model does not go further so, I used annealing learning rate, start from 5e-5 then decrease 5e-6 every epoch, reset after 10 epochs, but did not get any improvement even after 30 epochs of training, the result was 0.76 but when I increase the learning rate to 1e-4 I get 0.76 on my first epoch and it increases to 0.80 in first 10 epochs.\n\nI have no idea why lowering the learning rate and making my model train for longer is making even worse results.\n\nThank you for answering in advance."
  }
}