{
  "id": 160588,
  "title": "Optimizers",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/160588",
  "author_name": "Arnaud Roussel",
  "post_date": "2020-06-21T20:04:58.638000",
  "votes": 7,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Wanted to share the optimizers I've tried (without much success :( ) and wondered if someone had a boost from just a different optimizing cycle:\nOver 30 epochs:</p>\n\n<p>=&gt; Adam 3e-4 + One Fit Cycle (start 4e-5 and downward cycle starts after 1 epoch). Similar to the top kernel but I update every batch instead of every epoch.\nCV = 0.9014 (LB 0.89). Some peaks/valley pattern.</p>\n\n<p>=&gt; Ranger 3e-4 Flat followed by CosineAnnealing schedule (annealing starts at 22 epoch)\nCV = 0.8943. More stable than regular Adam.</p>\n\n<p>=&gt; Ranger 3e-4 Flat followed by CosineAnnealing schedule (annealing starts at 12 epoch)\nCV = 0.8974. More stable than regular Adam.</p>\n\n<p>=&gt; Ranger 3e-4 Flat followed by CosineAnnealing schedule (annealing starts at 6 epoch)\nCV = 0.8915  More stable than regular Adam.</p>\n\n<p>=&gt; Over9000 3e-4 Flat followed by CosineAnnealing schedule (annealing starts at 22 epoch)\nCV = 0.8850. Sometimes the CV was also dropping considerably (like 0.05) so I abandoned it.</p>\n\n<p>Architecture Efficientnet-b0</p>",
  "messages": [
    {
      "id": 896080,
      "postDate": "2020-06-21T20:04:58.640Z",
      "content": "<p>Wanted to share the optimizers I've tried (without much success :( ) and wondered if someone had a boost from just a different optimizing cycle:\nOver 30 epochs:</p>\n\n<p>=&gt; Adam 3e-4 + One Fit Cycle (start 4e-5 and downward cycle starts after 1 epoch). Similar to the top kernel but I update every batch instead of every epoch.\nCV = 0.9014 (LB 0.89). Some peaks/valley pattern.</p>\n\n<p>=&gt; Ranger 3e-4 Flat followed by CosineAnnealing schedule (annealing starts at 22 epoch)\nCV = 0.8943. More stable than regular Adam.</p>\n\n<p>=&gt; Ranger 3e-4 Flat followed by CosineAnnealing schedule (annealing starts at 12 epoch)\nCV = 0.8974. More stable than regular Adam.</p>\n\n<p>=&gt; Ranger 3e-4 Flat followed by CosineAnnealing schedule (annealing starts at 6 epoch)\nCV = 0.8915  More stable than regular Adam.</p>\n\n<p>=&gt; Over9000 3e-4 Flat followed by CosineAnnealing schedule (annealing starts at 22 epoch)\nCV = 0.8850. Sometimes the CV was also dropping considerably (like 0.05) so I abandoned it.</p>\n\n<p>Architecture Efficientnet-b0</p>",
      "rawMarkdown": "Wanted to share the optimizers I've tried (without much success :( ) and wondered if someone had a boost from just a different optimizing cycle:\nOver 30 epochs:\n\n=&gt; Adam 3e-4 + One Fit Cycle (start 4e-5 and downward cycle starts after 1 epoch). Similar to the top kernel but I update every batch instead of every epoch.\nCV = 0.9014 (LB 0.89). Some peaks/valley pattern.\n\n=&gt; Ranger 3e-4 Flat followed by CosineAnnealing schedule (annealing starts at 22 epoch)\nCV = 0.8943. More stable than regular Adam.\n\n=&gt; Ranger 3e-4 Flat followed by CosineAnnealing schedule (annealing starts at 12 epoch)\nCV = 0.8974. More stable than regular Adam.\n\n=&gt; Ranger 3e-4 Flat followed by CosineAnnealing schedule (annealing starts at 6 epoch)\nCV = 0.8915  More stable than regular Adam.\n\n=&gt; Over9000 3e-4 Flat followed by CosineAnnealing schedule (annealing starts at 22 epoch)\nCV = 0.8850. Sometimes the CV was also dropping considerably (like 0.05) so I abandoned it.\n\nArchitecture Efficientnet-b0",
      "votes": 7
    },
    {
      "id": 901050,
      "postDate": "2020-06-25T08:17:29.443Z",
      "content": "<p>Can you provide LBs for all tries? You have provided only for the first one as of now. Thanks</p>",
      "rawMarkdown": "Can you provide LBs for all tries? You have provided only for the first one as of now. Thanks"
    },
    {
      "id": 897901,
      "postDate": "2020-06-23T07:08:53.443Z",
      "content": "<p>thank you for sharing.. For clarify, your meaning \"Similar to the top kernel but I update every batch instead of every epoch.\"is that you call scheduler().step() in every batch?.. Can I ask why?.. .. Thank you :)</p>",
      "rawMarkdown": "thank you for sharing.. For clarify, your meaning \"Similar to the top kernel but I update every batch instead of every epoch.\"is that you call scheduler().step() in every batch?.. Can I ask why?.. .. Thank you :)",
      "replies": [
        {
          "id": 898519,
          "postDate": "2020-06-23T15:09:11.653Z",
          "content": "<p>I use torch.optim.lr_scheduler.OneCycleLR.\nAs per pytorch doc this scheduler is designed to be called at every batch over a number of steps = epochs * len(dataloader).</p>\n\n<p>One could also simply use it for total_steps = epochs and call it at every epoch so that the lr stays exactly the same during a single epoch.</p>",
          "rawMarkdown": "I use torch.optim.lr_scheduler.OneCycleLR.\nAs per pytorch doc this scheduler is designed to be called at every batch over a number of steps = epochs * len(dataloader).\n\nOne could also simply use it for total_steps = epochs and call it at every epoch so that the lr stays exactly the same during a single epoch."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 901050,
      "author_name": "Viraj Bagal",
      "author_url": "",
      "post_date": "2020-06-25T08:17:29.443000",
      "content": "<p>Can you provide LBs for all tries? You have provided only for the first one as of now. Thanks</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 897901,
      "author_name": "JJShadow",
      "author_url": "",
      "post_date": "2020-06-23T07:08:53.443000",
      "content": "<p>thank you for sharing.. For clarify, your meaning \"Similar to the top kernel but I update every batch instead of every epoch.\"is that you call scheduler().step() in every batch?.. Can I ask why?.. .. Thank you :)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 898519,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2020-06-23T15:09:11.653000",
          "content": "<p>I use torch.optim.lr_scheduler.OneCycleLR.\nAs per pytorch doc this scheduler is designed to be called at every batch over a number of steps = epochs * len(dataloader).</p>\n\n<p>One could also simply use it for total_steps = epochs and call it at every epoch so that the lr stays exactly the same during a single epoch.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "896080": "Wanted to share the optimizers I've tried (without much success :( ) and wondered if someone had a boost from just a different optimizing cycle:\nOver 30 epochs:\n\n=&gt; Adam 3e-4 + One Fit Cycle (start 4e-5 and downward cycle starts after 1 epoch). Similar to the top kernel but I update every batch instead of every epoch.\nCV = 0.9014 (LB 0.89). Some peaks/valley pattern.\n\n=&gt; Ranger 3e-4 Flat followed by CosineAnnealing schedule (annealing starts at 22 epoch)\nCV = 0.8943. More stable than regular Adam.\n\n=&gt; Ranger 3e-4 Flat followed by CosineAnnealing schedule (annealing starts at 12 epoch)\nCV = 0.8974. More stable than regular Adam.\n\n=&gt; Ranger 3e-4 Flat followed by CosineAnnealing schedule (annealing starts at 6 epoch)\nCV = 0.8915  More stable than regular Adam.\n\n=&gt; Over9000 3e-4 Flat followed by CosineAnnealing schedule (annealing starts at 22 epoch)\nCV = 0.8850. Sometimes the CV was also dropping considerably (like 0.05) so I abandoned it.\n\nArchitecture Efficientnet-b0",
    "901050": "Can you provide LBs for all tries? You have provided only for the first one as of now. Thanks",
    "897901": "thank you for sharing.. For clarify, your meaning \"Similar to the top kernel but I update every batch instead of every epoch.\"is that you call scheduler().step() in every batch?.. Can I ask why?.. .. Thank you :)"
  }
}