{
  "id": 71366,
  "title": "Pytorch Bug - Sudden Loss Spike",
  "url": "/competitions/quickdraw-doodle-recognition/discussion/71366",
  "author_name": "Moiz Saifee",
  "post_date": "2018-11-13T04:18:16.976000",
  "votes": 10,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I was having this weird issue where my model training would go as expected for a few hours and then suddenly validation loss would spike 10-20 times and would remain there. </p>\n\n<p>\"step\": 1001, \"valid_loss\": 1.7431628095829272, \"valid_metric\": 78.19093322753906\n\"step\": 2001, \"valid_loss\": 1.4513265029396243, \"valid_metric\": 82.96835327148438\n\"step\": 3001, \"valid_loss\": 1.3451246505655596, \"valid_metric\": 84.70668029785156\n\"step\": 4001, \"valid_loss\": 1.3000356848435024, \"valid_metric\": 85.25215911865234\n\"step\": 5001, \"valid_loss\": 1.270157439324557, \"valid_metric\": 85.66276550292969\n\"step\": 6001, \"valid_loss\": 1.2508466885522809, \"valid_metric\": 85.9275131225586\n\"step\": 7001, \"valid_loss\": 1.1894207958827543, \"valid_metric\": 86.96851348876953\n\"step\": 8001, \"valid_loss\": 1.1928023867442479, \"valid_metric\": 86.73373413085938\n\"step\": 9001, \"valid_loss\": 1.2461847584418324, \"valid_metric\": 86.09734344482422\n\"step\": 10001, \"valid_loss\": 1.1466158477546613, \"valid_metric\": 87.60589599609375\n\"step\": 11001, \"valid_loss\": 9.312969134591729, \"valid_metric\": 1.01902174949646\n\"step\": 12001, \"valid_loss\": 11.223968137560599, \"valid_metric\": 0.8671675324440002\n\"step\": 13001, \"valid_loss\": 14.000899806351917, \"valid_metric\": 0.9960438013076782\n\"step\": 14001, \"valid_loss\": 15.858244545929267, \"valid_metric\": 0.9960438013076782</p>\n\n<p>After going crazy with this issue for a few days, it looks to be a bug in PyTorch BN layer. The fix suggested by djangogo at <a href=\"https://github.com/bearpaw/pytorch-pose/issues/33\">https://github.com/bearpaw/pytorch-pose/issues/33</a> seems to work for me.  Hope it helps someone who encountered a similar issue. </p>",
  "messages": [
    {
      "id": 420104,
      "postDate": "2018-11-13T04:18:16.977Z",
      "content": "<p>I was having this weird issue where my model training would go as expected for a few hours and then suddenly validation loss would spike 10-20 times and would remain there. </p>\n\n<p>\"step\": 1001, \"valid_loss\": 1.7431628095829272, \"valid_metric\": 78.19093322753906\n\"step\": 2001, \"valid_loss\": 1.4513265029396243, \"valid_metric\": 82.96835327148438\n\"step\": 3001, \"valid_loss\": 1.3451246505655596, \"valid_metric\": 84.70668029785156\n\"step\": 4001, \"valid_loss\": 1.3000356848435024, \"valid_metric\": 85.25215911865234\n\"step\": 5001, \"valid_loss\": 1.270157439324557, \"valid_metric\": 85.66276550292969\n\"step\": 6001, \"valid_loss\": 1.2508466885522809, \"valid_metric\": 85.9275131225586\n\"step\": 7001, \"valid_loss\": 1.1894207958827543, \"valid_metric\": 86.96851348876953\n\"step\": 8001, \"valid_loss\": 1.1928023867442479, \"valid_metric\": 86.73373413085938\n\"step\": 9001, \"valid_loss\": 1.2461847584418324, \"valid_metric\": 86.09734344482422\n\"step\": 10001, \"valid_loss\": 1.1466158477546613, \"valid_metric\": 87.60589599609375\n\"step\": 11001, \"valid_loss\": 9.312969134591729, \"valid_metric\": 1.01902174949646\n\"step\": 12001, \"valid_loss\": 11.223968137560599, \"valid_metric\": 0.8671675324440002\n\"step\": 13001, \"valid_loss\": 14.000899806351917, \"valid_metric\": 0.9960438013076782\n\"step\": 14001, \"valid_loss\": 15.858244545929267, \"valid_metric\": 0.9960438013076782</p>\n\n<p>After going crazy with this issue for a few days, it looks to be a bug in PyTorch BN layer. The fix suggested by djangogo at <a href=\"https://github.com/bearpaw/pytorch-pose/issues/33\">https://github.com/bearpaw/pytorch-pose/issues/33</a> seems to work for me.  Hope it helps someone who encountered a similar issue. </p>",
      "rawMarkdown": "I was having this weird issue where my model training would go as expected for a few hours and then suddenly validation loss would spike 10-20 times and would remain there. \n\n\"step\": 1001, \"valid_loss\": 1.7431628095829272, \"valid_metric\": 78.19093322753906\n\"step\": 2001, \"valid_loss\": 1.4513265029396243, \"valid_metric\": 82.96835327148438\n\"step\": 3001, \"valid_loss\": 1.3451246505655596, \"valid_metric\": 84.70668029785156\n\"step\": 4001, \"valid_loss\": 1.3000356848435024, \"valid_metric\": 85.25215911865234\n\"step\": 5001, \"valid_loss\": 1.270157439324557, \"valid_metric\": 85.66276550292969\n\"step\": 6001, \"valid_loss\": 1.2508466885522809, \"valid_metric\": 85.9275131225586\n\"step\": 7001, \"valid_loss\": 1.1894207958827543, \"valid_metric\": 86.96851348876953\n\"step\": 8001, \"valid_loss\": 1.1928023867442479, \"valid_metric\": 86.73373413085938\n\"step\": 9001, \"valid_loss\": 1.2461847584418324, \"valid_metric\": 86.09734344482422\n\"step\": 10001, \"valid_loss\": 1.1466158477546613, \"valid_metric\": 87.60589599609375\n\"step\": 11001, \"valid_loss\": 9.312969134591729, \"valid_metric\": 1.01902174949646\n\"step\": 12001, \"valid_loss\": 11.223968137560599, \"valid_metric\": 0.8671675324440002\n\"step\": 13001, \"valid_loss\": 14.000899806351917, \"valid_metric\": 0.9960438013076782\n\"step\": 14001, \"valid_loss\": 15.858244545929267, \"valid_metric\": 0.9960438013076782\n\nAfter going crazy with this issue for a few days, it looks to be a bug in PyTorch BN layer. The fix suggested by djangogo at https://github.com/bearpaw/pytorch-pose/issues/33 seems to work for me.  Hope it helps someone who encountered a similar issue. \n",
      "votes": 10
    },
    {
      "id": 420395,
      "postDate": "2018-11-13T15:05:12.720Z",
      "content": "<p>Wow, I also have this problem when train on se resnet 50</p>",
      "rawMarkdown": "Wow, I also have this problem when train on se resnet 50",
      "replies": [
        {
          "id": 420675,
          "postDate": "2018-11-14T01:26:26.083Z",
          "content": "<p>Me too, my resnet50 donot convergence.</p>",
          "rawMarkdown": "Me too, my resnet50 donot convergence."
        }
      ]
    },
    {
      "id": 420162,
      "postDate": "2018-11-13T06:59:46.530Z",
      "content": "<p>Yes, set that to False. Didn't affect speed by a lot</p>",
      "rawMarkdown": "Yes, set that to False. Didn't affect speed by a lot",
      "replies": [
        {
          "id": 420419,
          "postDate": "2018-11-13T15:44:06.247Z",
          "content": "<p>Still have the same problem... Wondering if upgrading to 1.0 will solve the problem or not?\nLet me restart everything to try it again.</p>",
          "rawMarkdown": "Still have the same problem... Wondering if upgrading to 1.0 will solve the problem or not?\nLet me restart everything to try it again."
        }
      ]
    },
    {
      "id": 420139,
      "postDate": "2018-11-13T05:56:36.100Z",
      "content": "<p>Man, you saved my days.... I am having extract same issue as yours and checking my codes for several days, looking for bugs. Thank you so much to point this out! Did you fix it by changing torch.backends.cudnn.enabled to False in function batchnorm? Does it affect speed of training? Thanks again.</p>",
      "rawMarkdown": "Man, you saved my days.... I am having extract same issue as yours and checking my codes for several days, looking for bugs. Thank you so much to point this out! Did you fix it by changing torch.backends.cudnn.enabled to False in function batchnorm? Does it affect speed of training? Thanks again."
    }
  ],
  "comments": [
    {
      "id": 420395,
      "author_name": "Strideradu",
      "author_url": "",
      "post_date": "2018-11-13T15:05:12.720000",
      "content": "<p>Wow, I also have this problem when train on se resnet 50</p>",
      "votes": 0,
      "replies": [
        {
          "id": 420675,
          "author_name": "Finlay",
          "author_url": "",
          "post_date": "2018-11-14T01:26:26.083000",
          "content": "<p>Me too, my resnet50 donot convergence.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 420162,
      "author_name": "Moiz Saifee",
      "author_url": "",
      "post_date": "2018-11-13T06:59:46.530000",
      "content": "<p>Yes, set that to False. Didn't affect speed by a lot</p>",
      "votes": 0,
      "replies": [
        {
          "id": 420419,
          "author_name": "YoungLamb",
          "author_url": "",
          "post_date": "2018-11-13T15:44:06.247000",
          "content": "<p>Still have the same problem... Wondering if upgrading to 1.0 will solve the problem or not?\nLet me restart everything to try it again.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 420139,
      "author_name": "YoungLamb",
      "author_url": "",
      "post_date": "2018-11-13T05:56:36.100000",
      "content": "<p>Man, you saved my days.... I am having extract same issue as yours and checking my codes for several days, looking for bugs. Thank you so much to point this out! Did you fix it by changing torch.backends.cudnn.enabled to False in function batchnorm? Does it affect speed of training? Thanks again.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "420104": "I was having this weird issue where my model training would go as expected for a few hours and then suddenly validation loss would spike 10-20 times and would remain there. \n\n\"step\": 1001, \"valid_loss\": 1.7431628095829272, \"valid_metric\": 78.19093322753906\n\"step\": 2001, \"valid_loss\": 1.4513265029396243, \"valid_metric\": 82.96835327148438\n\"step\": 3001, \"valid_loss\": 1.3451246505655596, \"valid_metric\": 84.70668029785156\n\"step\": 4001, \"valid_loss\": 1.3000356848435024, \"valid_metric\": 85.25215911865234\n\"step\": 5001, \"valid_loss\": 1.270157439324557, \"valid_metric\": 85.66276550292969\n\"step\": 6001, \"valid_loss\": 1.2508466885522809, \"valid_metric\": 85.9275131225586\n\"step\": 7001, \"valid_loss\": 1.1894207958827543, \"valid_metric\": 86.96851348876953\n\"step\": 8001, \"valid_loss\": 1.1928023867442479, \"valid_metric\": 86.73373413085938\n\"step\": 9001, \"valid_loss\": 1.2461847584418324, \"valid_metric\": 86.09734344482422\n\"step\": 10001, \"valid_loss\": 1.1466158477546613, \"valid_metric\": 87.60589599609375\n\"step\": 11001, \"valid_loss\": 9.312969134591729, \"valid_metric\": 1.01902174949646\n\"step\": 12001, \"valid_loss\": 11.223968137560599, \"valid_metric\": 0.8671675324440002\n\"step\": 13001, \"valid_loss\": 14.000899806351917, \"valid_metric\": 0.9960438013076782\n\"step\": 14001, \"valid_loss\": 15.858244545929267, \"valid_metric\": 0.9960438013076782\n\nAfter going crazy with this issue for a few days, it looks to be a bug in PyTorch BN layer. The fix suggested by djangogo at https://github.com/bearpaw/pytorch-pose/issues/33 seems to work for me.  Hope it helps someone who encountered a similar issue. \n",
    "420395": "Wow, I also have this problem when train on se resnet 50",
    "420162": "Yes, set that to False. Didn't affect speed by a lot",
    "420139": "Man, you saved my days.... I am having extract same issue as yours and checking my codes for several days, looking for bugs. Thank you so much to point this out! Did you fix it by changing torch.backends.cudnn.enabled to False in function batchnorm? Does it affect speed of training? Thanks again."
  }
}