{
  "id": 72400,
  "title": "How to escape a local minimum?",
  "url": "/competitions/quickdraw-doodle-recognition/discussion/72400",
  "author_name": "HuyenNguyen",
  "post_date": "2018-11-22T23:24:05.946000",
  "votes": 5,
  "comment_count": 26,
  "views": 0,
  "content": "<p>In a lot of my models, the accuracy gets really stuck when it gets to about the 0.8-0.82 range. I tried cyclic learning rate, reduce on plateau, exponential decay or even occasionally do a learning rate finder to find the optimal learning rate, but it can't escape this minimum. I use Adam. What's everyone's experience with this? </p>",
  "messages": [
    {
      "id": 426227,
      "postDate": "2018-11-22T23:24:05.947Z",
      "content": "<p>In a lot of my models, the accuracy gets really stuck when it gets to about the 0.8-0.82 range. I tried cyclic learning rate, reduce on plateau, exponential decay or even occasionally do a learning rate finder to find the optimal learning rate, but it can't escape this minimum. I use Adam. What's everyone's experience with this? </p>",
      "rawMarkdown": "In a lot of my models, the accuracy gets really stuck when it gets to about the 0.8-0.82 range. I tried cyclic learning rate, reduce on plateau, exponential decay or even occasionally do a learning rate finder to find the optimal learning rate, but it can't escape this minimum. I use Adam. What's everyone's experience with this? ",
      "votes": 5
    },
    {
      "id": 429494,
      "postDate": "2018-11-29T00:07:29.433Z",
      "content": "<p>In reply to <a href=\"/fiyeroleung\">@fiyeroleung</a> in private email: I gave cyclic LR another try. It knocked off the accuracy but eventually recovered and got higher. </p>",
      "rawMarkdown": "In reply to @fiyeroleung in private email: I gave cyclic LR another try. It knocked off the accuracy but eventually recovered and got higher. ",
      "votes": 1,
      "replies": [
        {
          "id": 429579,
          "postDate": "2018-11-29T03:22:32.853Z",
          "content": "<p>@HuyenNguyen, how long did it take to recover? My model is only in epoch 2 right now but the loss went above 1.0 and accuracy below 0.9, is that the kind of things you saw with cyclic LR?</p>",
          "rawMarkdown": "@HuyenNguyen, how long did it take to recover? My model is only in epoch 2 right now but the loss went above 1.0 and accuracy below 0.9, is that the kind of things you saw with cyclic LR?"
        },
        {
          "id": 429836,
          "postDate": "2018-11-29T12:38:28.297Z",
          "content": "<p>Took quite a long time to recover, for me it was about 80 epochs * 2000 steps * 500 batchsize. </p>",
          "rawMarkdown": "Took quite a long time to recover, for me it was about 80 epochs * 2000 steps * 500 batchsize. "
        },
        {
          "id": 429838,
          "postDate": "2018-11-29T12:38:51.107Z",
          "content": "<p>so 1.6 epoch of full dataset. </p>",
          "rawMarkdown": "so 1.6 epoch of full dataset. "
        },
        {
          "id": 429951,
          "postDate": "2018-11-29T15:28:44.593Z",
          "content": "<p>Wow! Is that with image size 128x128 ? </p>",
          "rawMarkdown": "Wow! Is that with image size 128x128 ? "
        },
        {
          "id": 430136,
          "postDate": "2018-11-29T22:06:09.850Z",
          "content": "<p>Yeah, but the total training time was long. Took ages to get to 0.925 and then 1.6 epochs to get to 0.931. The magic has run out now, I am stuck at 0.931 no matter what I do. </p>",
          "rawMarkdown": "Yeah, but the total training time was long. Took ages to get to 0.925 and then 1.6 epochs to get to 0.931. The magic has run out now, I am stuck at 0.931 no matter what I do. ",
          "votes": 1
        },
        {
          "id": 430175,
          "postDate": "2018-11-30T00:47:29.043Z",
          "content": "<p>Perhaps you can try fine-tuning with larger images size, like 144 or 168, I also think there's no more thing that can improve the result</p>",
          "rawMarkdown": "Perhaps you can try fine-tuning with larger images size, like 144 or 168, I also think there's no more thing that can improve the result"
        },
        {
          "id": 430781,
          "postDate": "2018-11-30T23:45:47.813Z",
          "content": "<p>Mine took a long time too but reached a peak of 0.919 only. I used 150k per class for training. </p>\n\n<p>How do you train on 1.6 epochs? I believe you meant 1.6 epoch training on the whole data. So if you are using something similar to your kernel, how did you use 1.6 epochs?</p>",
          "rawMarkdown": "Mine took a long time too but reached a peak of 0.919 only. I used 150k per class for training. \n\nHow do you train on 1.6 epochs? I believe you meant 1.6 epoch training on the whole data. So if you are using something similar to your kernel, how did you use 1.6 epochs?"
        },
        {
          "id": 430787,
          "postDate": "2018-12-01T00:11:37.967Z",
          "content": "<p>80 epochs * 2000 steps * 500 batchsize, I didn't run in on the kernel though. </p>",
          "rawMarkdown": "80 epochs * 2000 steps * 500 batchsize, I didn't run in on the kernel though. "
        },
        {
          "id": 430811,
          "postDate": "2018-12-01T01:42:31.053Z",
          "content": "<blockquote>\n  <p>80 epochs * 2000 steps * 500</p>\n</blockquote>\n\n<p>This is 80,000,000 samples. This is 1.6% of the whole data samples? Or am i misunderstanding what 1.6 epochs mean here?</p>",
          "rawMarkdown": "&gt; 80 epochs * 2000 steps * 500\n\nThis is 80,000,000 samples. This is 1.6% of the whole data samples? Or am i misunderstanding what 1.6 epochs mean here?"
        },
        {
          "id": 430966,
          "postDate": "2018-12-01T10:51:35.217Z",
          "content": "<p>1.6 epochs means going through the whole dataset 1.6 times. </p>",
          "rawMarkdown": "1.6 epochs means going through the whole dataset 1.6 times. "
        },
        {
          "id": 430972,
          "postDate": "2018-12-01T11:02:47.780Z",
          "content": "<p>Thanks. I have not used the whole data set yet i.e used only up to 500,000 per class. How many samples are there in the whole data?</p>",
          "rawMarkdown": "Thanks. I have not used the whole data set yet i.e used only up to 500,000 per class. How many samples are there in the whole data?",
          "votes": -1
        }
      ]
    },
    {
      "id": 426287,
      "postDate": "2018-11-23T02:25:02.133Z",
      "content": "<p>Generally, I get stuck with train-accuracy in the range of 0.92-0.93, and I train it for over 70-80 epochs. After that, nothing will be helpful to increasing the accuracy.\nI also use Adam, and first cyclic lr and then reduce on plateau lr. Sometimes I trained over 100 epochs.</p>",
      "rawMarkdown": "Generally, I get stuck with train-accuracy in the range of 0.92-0.93, and I train it for over 70-80 epochs. After that, nothing will be helpful to increasing the accuracy.\nI also use Adam, and first cyclic lr and then reduce on plateau lr. Sometimes I trained over 100 epochs.",
      "votes": 1,
      "replies": [
        {
          "id": 426314,
          "postDate": "2018-11-23T03:45:59.667Z",
          "content": "<p>yes, I have exactly the same problem, stuck at 0.925, train more at various learning rates or image siz, but can't escape it. </p>",
          "rawMarkdown": "yes, I have exactly the same problem, stuck at 0.925, train more at various learning rates or image siz, but can't escape it. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 426296,
      "postDate": "2018-11-23T02:56:10.137Z",
      "content": "<p>Use large learning rate(i.e. 0.002) at beginning, then use ReduceLRonPlateau while the patience has been exausted.</p>\n\n<p>However, when the learning rate has reduced to about 5e-5, the local validation still improving, but public LB stoped. I guess it is because the model start fitting the noise label rather normal label.</p>\n\n<p>I don't think cyclic LR is a good idea, because it will fit noise data in each cycle rather than normal data.</p>",
      "rawMarkdown": "Use large learning rate(i.e. 0.002) at beginning, then use ReduceLRonPlateau while the patience has been exausted.\n\nHowever, when the learning rate has reduced to about 5e-5, the local validation still improving, but public LB stoped. I guess it is because the model start fitting the noise label rather normal label.\n\nI don't think cyclic LR is a good idea, because it will fit noise data in each cycle rather than normal data.",
      "votes": 2,
      "replies": [
        {
          "id": 426297,
          "postDate": "2018-11-23T03:01:40.297Z",
          "content": "<p>you are right. model start fitting the noise label rather normal label with small lr.</p>",
          "rawMarkdown": "you are right. model start fitting the noise label rather normal label with small lr."
        },
        {
          "id": 426315,
          "postDate": "2018-11-23T03:47:28.613Z",
          "content": "<p>So how do we avoid this? can we exclude the crappy images in the late stage of training? I tried excluding them from the beginning, but it made no difference compared to including all images. </p>",
          "rawMarkdown": "So how do we avoid this? can we exclude the crappy images in the late stage of training? I tried excluding them from the beginning, but it made no difference compared to including all images. "
        },
        {
          "id": 426451,
          "postDate": "2018-11-23T09:30:08.260Z",
          "content": "<p>Use large batchsize as heng's idea, it's really effective :)</p>",
          "rawMarkdown": "Use large batchsize as heng's idea, it's really effective :)"
        },
        {
          "id": 426456,
          "postDate": "2018-11-23T09:37:44.083Z",
          "content": "<p>But then we must decrease the input size of pictures...Otherwise it will generate 'Resource Exhausted Error.'</p>",
          "rawMarkdown": "But then we must decrease the input size of pictures...Otherwise it will generate 'Resource Exhausted Error.'"
        },
        {
          "id": 426493,
          "postDate": "2018-11-23T10:51:44.953Z",
          "content": "<p>img size 128 is enough</p>",
          "rawMarkdown": "img size 128 is enough",
          "votes": 1
        },
        {
          "id": 426516,
          "postDate": "2018-11-23T11:35:20.500Z",
          "content": "<p>I use the gradient accumulator here <a href=\"https://github.com/keras-team/keras/issues/3556\">https://github.com/keras-team/keras/issues/3556</a>, batchsize of 100, accumulate 20 iterations, still stuck in this minimum/saddle.</p>",
          "rawMarkdown": "I use the gradient accumulator here https://github.com/keras-team/keras/issues/3556, batchsize of 100, accumulate 20 iterations, still stuck in this minimum/saddle."
        },
        {
          "id": 426730,
          "postDate": "2018-11-23T18:08:25.730Z",
          "content": "<p>Bigger batch size might help.\nIn my case, batch size 1024 with 80x80 image size gets LB 0.934 in less than 2 epochs from scratch. The small batch size could not achieve that score.\nJust as Heng mentioned in <a href=\"https://www.kaggle.com/c/quickdraw-doodle-recognition/discussion/70558#418793\">other discussions</a>, the bigger batch size seems to help limiting the negative impact of noise in data.</p>",
          "rawMarkdown": "Bigger batch size might help.\nIn my case, batch size 1024 with 80x80 image size gets LB 0.934 in less than 2 epochs from scratch. The small batch size could not achieve that score.\nJust as Heng mentioned in [other discussions][1], the bigger batch size seems to help limiting the negative impact of noise in data.\n\n\n  [1]: https://www.kaggle.com/c/quickdraw-doodle-recognition/discussion/70558#418793",
          "votes": 1
        },
        {
          "id": 426796,
          "postDate": "2018-11-23T21:47:47.410Z",
          "content": "<p>and what's your learning rate schedule?</p>",
          "rawMarkdown": "and what's your learning rate schedule?"
        },
        {
          "id": 426915,
          "postDate": "2018-11-24T05:39:01.640Z",
          "content": "<p>Reducelronplateau(from 3e-3) with restarts.\nIf a large batch size (100*20?) does not help, it might be different reasons such as a way of encoding stroke information or network I guess. </p>",
          "rawMarkdown": "Reducelronplateau(from 3e-3) with restarts.\nIf a large batch size (100*20?) does not help, it might be different reasons such as a way of encoding stroke information or network I guess. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 426258,
      "postDate": "2018-11-23T01:21:11.147Z",
      "content": "<p>you can use big lr to  excape the local minumum.</p>",
      "rawMarkdown": "you can use big lr to  excape the local minumum.",
      "replies": [
        {
          "id": 426265,
          "postDate": "2018-11-23T01:35:50.900Z",
          "content": "<p>I did that with cyclic lr, but the results aren't convincing :-/</p>",
          "rawMarkdown": "I did that with cyclic lr, but the results aren't convincing :-/"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 429494,
      "author_name": "HuyenNguyen",
      "author_url": "",
      "post_date": "2018-11-29T00:07:29.433000",
      "content": "<p>In reply to <a href=\"/fiyeroleung\">@fiyeroleung</a> in private email: I gave cyclic LR another try. It knocked off the accuracy but eventually recovered and got higher. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 429579,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-11-29T03:22:32.853000",
          "content": "<p>@HuyenNguyen, how long did it take to recover? My model is only in epoch 2 right now but the loss went above 1.0 and accuracy below 0.9, is that the kind of things you saw with cyclic LR?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 429836,
          "author_name": "HuyenNguyen",
          "author_url": "",
          "post_date": "2018-11-29T12:38:28.297000",
          "content": "<p>Took quite a long time to recover, for me it was about 80 epochs * 2000 steps * 500 batchsize. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 429838,
          "author_name": "HuyenNguyen",
          "author_url": "",
          "post_date": "2018-11-29T12:38:51.107000",
          "content": "<p>so 1.6 epoch of full dataset. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 429951,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-11-29T15:28:44.593000",
          "content": "<p>Wow! Is that with image size 128x128 ? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 430136,
          "author_name": "HuyenNguyen",
          "author_url": "",
          "post_date": "2018-11-29T22:06:09.850000",
          "content": "<p>Yeah, but the total training time was long. Took ages to get to 0.925 and then 1.6 epochs to get to 0.931. The magic has run out now, I am stuck at 0.931 no matter what I do. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 430175,
          "author_name": "Dilapsky Lee",
          "author_url": "",
          "post_date": "2018-11-30T00:47:29.043000",
          "content": "<p>Perhaps you can try fine-tuning with larger images size, like 144 or 168, I also think there's no more thing that can improve the result</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 430781,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-11-30T23:45:47.813000",
          "content": "<p>Mine took a long time too but reached a peak of 0.919 only. I used 150k per class for training. </p>\n\n<p>How do you train on 1.6 epochs? I believe you meant 1.6 epoch training on the whole data. So if you are using something similar to your kernel, how did you use 1.6 epochs?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 430787,
          "author_name": "HuyenNguyen",
          "author_url": "",
          "post_date": "2018-12-01T00:11:37.967000",
          "content": "<p>80 epochs * 2000 steps * 500 batchsize, I didn't run in on the kernel though. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 430811,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-12-01T01:42:31.053000",
          "content": "<blockquote>\n  <p>80 epochs * 2000 steps * 500</p>\n</blockquote>\n\n<p>This is 80,000,000 samples. This is 1.6% of the whole data samples? Or am i misunderstanding what 1.6 epochs mean here?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 430966,
          "author_name": "HuyenNguyen",
          "author_url": "",
          "post_date": "2018-12-01T10:51:35.217000",
          "content": "<p>1.6 epochs means going through the whole dataset 1.6 times. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 430972,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-12-01T11:02:47.780000",
          "content": "<p>Thanks. I have not used the whole data set yet i.e used only up to 500,000 per class. How many samples are there in the whole data?</p>",
          "votes": -1,
          "replies": []
        }
      ]
    },
    {
      "id": 426287,
      "author_name": "Dilapsky Lee",
      "author_url": "",
      "post_date": "2018-11-23T02:25:02.133000",
      "content": "<p>Generally, I get stuck with train-accuracy in the range of 0.92-0.93, and I train it for over 70-80 epochs. After that, nothing will be helpful to increasing the accuracy.\nI also use Adam, and first cyclic lr and then reduce on plateau lr. Sometimes I trained over 100 epochs.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 426314,
          "author_name": "HuyenNguyen",
          "author_url": "",
          "post_date": "2018-11-23T03:45:59.667000",
          "content": "<p>yes, I have exactly the same problem, stuck at 0.925, train more at various learning rates or image siz, but can't escape it. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 426296,
      "author_name": "wh1te",
      "author_url": "",
      "post_date": "2018-11-23T02:56:10.137000",
      "content": "<p>Use large learning rate(i.e. 0.002) at beginning, then use ReduceLRonPlateau while the patience has been exausted.</p>\n\n<p>However, when the learning rate has reduced to about 5e-5, the local validation still improving, but public LB stoped. I guess it is because the model start fitting the noise label rather normal label.</p>\n\n<p>I don't think cyclic LR is a good idea, because it will fit noise data in each cycle rather than normal data.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 426297,
          "author_name": "Gary",
          "author_url": "",
          "post_date": "2018-11-23T03:01:40.297000",
          "content": "<p>you are right. model start fitting the noise label rather normal label with small lr.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 426315,
          "author_name": "HuyenNguyen",
          "author_url": "",
          "post_date": "2018-11-23T03:47:28.613000",
          "content": "<p>So how do we avoid this? can we exclude the crappy images in the late stage of training? I tried excluding them from the beginning, but it made no difference compared to including all images. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 426451,
          "author_name": "wh1te",
          "author_url": "",
          "post_date": "2018-11-23T09:30:08.260000",
          "content": "<p>Use large batchsize as heng's idea, it's really effective :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 426456,
          "author_name": "Dilapsky Lee",
          "author_url": "",
          "post_date": "2018-11-23T09:37:44.083000",
          "content": "<p>But then we must decrease the input size of pictures...Otherwise it will generate 'Resource Exhausted Error.'</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 426493,
          "author_name": "wh1te",
          "author_url": "",
          "post_date": "2018-11-23T10:51:44.953000",
          "content": "<p>img size 128 is enough</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 426516,
          "author_name": "HuyenNguyen",
          "author_url": "",
          "post_date": "2018-11-23T11:35:20.500000",
          "content": "<p>I use the gradient accumulator here <a href=\"https://github.com/keras-team/keras/issues/3556\">https://github.com/keras-team/keras/issues/3556</a>, batchsize of 100, accumulate 20 iterations, still stuck in this minimum/saddle.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 426730,
          "author_name": "Appian",
          "author_url": "",
          "post_date": "2018-11-23T18:08:25.730000",
          "content": "<p>Bigger batch size might help.\nIn my case, batch size 1024 with 80x80 image size gets LB 0.934 in less than 2 epochs from scratch. The small batch size could not achieve that score.\nJust as Heng mentioned in <a href=\"https://www.kaggle.com/c/quickdraw-doodle-recognition/discussion/70558#418793\">other discussions</a>, the bigger batch size seems to help limiting the negative impact of noise in data.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 426796,
          "author_name": "HuyenNguyen",
          "author_url": "",
          "post_date": "2018-11-23T21:47:47.410000",
          "content": "<p>and what's your learning rate schedule?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 426915,
          "author_name": "Appian",
          "author_url": "",
          "post_date": "2018-11-24T05:39:01.640000",
          "content": "<p>Reducelronplateau(from 3e-3) with restarts.\nIf a large batch size (100*20?) does not help, it might be different reasons such as a way of encoding stroke information or network I guess. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 426258,
      "author_name": "Gary",
      "author_url": "",
      "post_date": "2018-11-23T01:21:11.147000",
      "content": "<p>you can use big lr to  excape the local minumum.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 426265,
          "author_name": "HuyenNguyen",
          "author_url": "",
          "post_date": "2018-11-23T01:35:50.900000",
          "content": "<p>I did that with cyclic lr, but the results aren't convincing :-/</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "426227": "In a lot of my models, the accuracy gets really stuck when it gets to about the 0.8-0.82 range. I tried cyclic learning rate, reduce on plateau, exponential decay or even occasionally do a learning rate finder to find the optimal learning rate, but it can't escape this minimum. I use Adam. What's everyone's experience with this? ",
    "429494": "In reply to @fiyeroleung in private email: I gave cyclic LR another try. It knocked off the accuracy but eventually recovered and got higher. ",
    "426287": "Generally, I get stuck with train-accuracy in the range of 0.92-0.93, and I train it for over 70-80 epochs. After that, nothing will be helpful to increasing the accuracy.\nI also use Adam, and first cyclic lr and then reduce on plateau lr. Sometimes I trained over 100 epochs.",
    "426296": "Use large learning rate(i.e. 0.002) at beginning, then use ReduceLRonPlateau while the patience has been exausted.\n\nHowever, when the learning rate has reduced to about 5e-5, the local validation still improving, but public LB stoped. I guess it is because the model start fitting the noise label rather normal label.\n\nI don't think cyclic LR is a good idea, because it will fit noise data in each cycle rather than normal data.",
    "426258": "you can use big lr to  excape the local minumum."
  }
}