{
  "id": 348949,
  "title": "Optimizers AdamP and Ranger21",
  "url": "/competitions/mayo-clinic-strip-ai/discussion/348949",
  "author_name": "Kirderf",
  "post_date": "2022-08-30T17:34:47.143000",
  "votes": 5,
  "comment_count": 4,
  "views": 0,
  "content": "<p>If someone is using Adam, Ranger or AdamW in training, maybe AdamP and Ranger21 can help even better, anyway can be worth a try :)</p>\n<p><strong>AdamP</strong></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3924325%2F0123b313bdc492661d572fc4c3dab6ed%2Fsdgp.JPG?generation=1596093943582809&amp;alt=media\" alt=\"\"></p>\n<p><a href=\"https://github.com/clovaai/AdamP\" target=\"_blank\">https://github.com/clovaai/AdamP</a><br>\n<a href=\"https://arxiv.org/abs/2006.08217\" target=\"_blank\">https://arxiv.org/abs/2006.08217</a></p>\n<p><strong>Ranger21</strong></p>\n<pre><code>Ranger21 - integrating the latest deep learning components into a single optimizer\nA rewrite of the Ranger deep learning optimizer to integrate newer optimization ideas and, in particular:\n\nuses the AdamW optimizer as its core (or, optionally, MadGrad)\nAdaptive gradient clipping\nGradient centralization\nPositive-Negative momentum\nNorm loss\nStable weight decay\nLinear learning rate warm-up\nExplore-exploit learning rate schedule\nLookahead\nSoftplus transformation\nGradient Normalization\n</code></pre>\n<p><a href=\"https://github.com/lessw2020/Ranger21\" target=\"_blank\">https://github.com/lessw2020/Ranger21</a><br>\n<a href=\"https://arxiv.org/abs/2106.13731\" target=\"_blank\">https://arxiv.org/abs/2106.13731</a></p>",
  "messages": [
    {
      "id": 1919826,
      "postDate": "2022-08-30T17:34:47.143Z",
      "content": "<p>If someone is using Adam, Ranger or AdamW in training, maybe AdamP and Ranger21 can help even better, anyway can be worth a try :)</p>\n<p><strong>AdamP</strong></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3924325%2F0123b313bdc492661d572fc4c3dab6ed%2Fsdgp.JPG?generation=1596093943582809&amp;alt=media\" alt=\"\"></p>\n<p><a href=\"https://github.com/clovaai/AdamP\" target=\"_blank\">https://github.com/clovaai/AdamP</a><br>\n<a href=\"https://arxiv.org/abs/2006.08217\" target=\"_blank\">https://arxiv.org/abs/2006.08217</a></p>\n<p><strong>Ranger21</strong></p>\n<pre><code>Ranger21 - integrating the latest deep learning components into a single optimizer\nA rewrite of the Ranger deep learning optimizer to integrate newer optimization ideas and, in particular:\n\nuses the AdamW optimizer as its core (or, optionally, MadGrad)\nAdaptive gradient clipping\nGradient centralization\nPositive-Negative momentum\nNorm loss\nStable weight decay\nLinear learning rate warm-up\nExplore-exploit learning rate schedule\nLookahead\nSoftplus transformation\nGradient Normalization\n</code></pre>\n<p><a href=\"https://github.com/lessw2020/Ranger21\" target=\"_blank\">https://github.com/lessw2020/Ranger21</a><br>\n<a href=\"https://arxiv.org/abs/2106.13731\" target=\"_blank\">https://arxiv.org/abs/2106.13731</a></p>",
      "rawMarkdown": "If someone is using Adam, Ranger or AdamW in training, maybe AdamP and Ranger21 can help even better, anyway can be worth a try :)\n\n**AdamP**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3924325%2F0123b313bdc492661d572fc4c3dab6ed%2Fsdgp.JPG?generation=1596093943582809&amp;alt=media)\n\nhttps://github.com/clovaai/AdamP\nhttps://arxiv.org/abs/2006.08217\n\n**Ranger21**\n\n```\nRanger21 - integrating the latest deep learning components into a single optimizer\nA rewrite of the Ranger deep learning optimizer to integrate newer optimization ideas and, in particular:\n\nuses the AdamW optimizer as its core (or, optionally, MadGrad)\nAdaptive gradient clipping\nGradient centralization\nPositive-Negative momentum\nNorm loss\nStable weight decay\nLinear learning rate warm-up\nExplore-exploit learning rate schedule\nLookahead\nSoftplus transformation\nGradient Normalization\n```\n\nhttps://github.com/lessw2020/Ranger21\nhttps://arxiv.org/abs/2106.13731",
      "votes": 5
    },
    {
      "id": 1954838,
      "postDate": "2022-09-25T13:55:49.357Z",
      "content": "<p>Many participants seem to have given up because they could not find the signal on the dataset.<br>\nDid your model find the any signal?</p>",
      "rawMarkdown": "Many participants seem to have given up because they could not find the signal on the dataset.\nDid your model find the any signal?"
    },
    {
      "id": 1922039,
      "postDate": "2022-09-01T08:30:58.207Z",
      "content": "<p>Did you try them? Any score improvement? </p>",
      "rawMarkdown": "Did you try them? Any score improvement? ",
      "replies": [
        {
          "id": 1922078,
          "postDate": "2022-09-01T08:54:52.603Z",
          "content": "<p>I always try AdamP when Adam is in play, and in general get better result.<br>\nRanger21 is quite new with lots of SOTA features, so with the some knowledge how it works I think it can be very useful. I have used it before with good results, one can almost see it as a automl optimizer+scheduler with all the features.</p>\n<p>Why I post it before testing it here, that's because I always post finding in the end of a competition due to lack of time but have also started posting what I usually use and best practice in the start. Maybe it can help someone better, having them early one get more time testing things instead later in the competition….And maybe it works very well after testing and tuning here, then I can't post it close to the end of the competition, not so fair. But it's only some optimizers, maybe not a huge LB mover ;) Anyway good optimizer :)</p>",
          "rawMarkdown": "I always try AdamP when Adam is in play, and in general get better result.\nRanger21 is quite new with lots of SOTA features, so with the some knowledge how it works I think it can be very useful. I have used it before with good results, one can almost see it as a automl optimizer+scheduler with all the features.\n\nWhy I post it before testing it here, that's because I always post finding in the end of a competition due to lack of time but have also started posting what I usually use and best practice in the start. Maybe it can help someone better, having them early one get more time testing things instead later in the competition....And maybe it works very well after testing and tuning here, then I can't post it close to the end of the competition, not so fair. But it's only some optimizers, maybe not a huge LB mover ;) Anyway good optimizer :)",
          "votes": 3
        }
      ]
    },
    {
      "id": 1923259,
      "postDate": "2022-09-02T05:12:03.200Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1954838,
      "author_name": "ParkSom",
      "author_url": "",
      "post_date": "2022-09-25T13:55:49.357000",
      "content": "<p>Many participants seem to have given up because they could not find the signal on the dataset.<br>\nDid your model find the any signal?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1922039,
      "author_name": "The Devastator",
      "author_url": "",
      "post_date": "2022-09-01T08:30:58.207000",
      "content": "<p>Did you try them? Any score improvement? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1922078,
          "author_name": "Kirderf",
          "author_url": "",
          "post_date": "2022-09-01T08:54:52.603000",
          "content": "<p>I always try AdamP when Adam is in play, and in general get better result.<br>\nRanger21 is quite new with lots of SOTA features, so with the some knowledge how it works I think it can be very useful. I have used it before with good results, one can almost see it as a automl optimizer+scheduler with all the features.</p>\n<p>Why I post it before testing it here, that's because I always post finding in the end of a competition due to lack of time but have also started posting what I usually use and best practice in the start. Maybe it can help someone better, having them early one get more time testing things instead later in the competition….And maybe it works very well after testing and tuning here, then I can't post it close to the end of the competition, not so fair. But it's only some optimizers, maybe not a huge LB mover ;) Anyway good optimizer :)</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1923259,
      "author_name": "BettyD",
      "author_url": "",
      "post_date": "2022-09-02T05:12:03.200000",
      "content": "<p>Thanks for sharing!</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1919826": "If someone is using Adam, Ranger or AdamW in training, maybe AdamP and Ranger21 can help even better, anyway can be worth a try :)\n\n**AdamP**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3924325%2F0123b313bdc492661d572fc4c3dab6ed%2Fsdgp.JPG?generation=1596093943582809&amp;alt=media)\n\nhttps://github.com/clovaai/AdamP\nhttps://arxiv.org/abs/2006.08217\n\n**Ranger21**\n\n```\nRanger21 - integrating the latest deep learning components into a single optimizer\nA rewrite of the Ranger deep learning optimizer to integrate newer optimization ideas and, in particular:\n\nuses the AdamW optimizer as its core (or, optionally, MadGrad)\nAdaptive gradient clipping\nGradient centralization\nPositive-Negative momentum\nNorm loss\nStable weight decay\nLinear learning rate warm-up\nExplore-exploit learning rate schedule\nLookahead\nSoftplus transformation\nGradient Normalization\n```\n\nhttps://github.com/lessw2020/Ranger21\nhttps://arxiv.org/abs/2106.13731",
    "1954838": "Many participants seem to have given up because they could not find the signal on the dataset.\nDid your model find the any signal?",
    "1922039": "Did you try them? Any score improvement? ",
    "1923259": "Thanks for sharing!"
  }
}