{
  "id": 71731,
  "title": "What make you think the model is ready? instead of keep training and get overfit",
  "url": "/competitions/quickdraw-doodle-recognition/discussion/71731",
  "author_name": "Salaryman",
  "post_date": "2018-11-16T03:30:58.043000",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi I am new to Machine Learning.\nI am curious about in real life how should we \"decide\" whether the algo is well trained?\nIn books/articles for rookie I learn I should compare the validation score and training score.</p>\n\n<p>For example if the validation score is already not bad (eg 0.9) and seems cant get improvement by using small learning rate, I may stop training to avoid overfitting. </p>\n\n<p>In kaggle, we have public leader board score so I know I should continue training if the score is far from the top performers. (so I will realise validation score 0.9 is not enough, either keep tuning the parameters/ switch model).</p>\n\n<p>But in real life there is no such kaggle score for us to evaluate. What make you think the model is ready? instead of keep training and get overfit?</p>\n\n<p>Thanks so much guys</p>",
  "messages": [
    {
      "id": 422609,
      "postDate": "2018-11-16T13:56:23.540Z",
      "content": "<p>If you are using a deep learning network, usually drawing  loss and accuracy figures of both training and validation will give you good clues about the effectiveness of your network. I recommend section 7 of these <a href=\"http://lamda.nju.edu.cn/weixs/project/CNNTricks/CNNTricks.html\">cnntricks</a> webpage to interpret your figures.</p>",
      "rawMarkdown": "If you are using a deep learning network, usually drawing  loss and accuracy figures of both training and validation will give you good clues about the effectiveness of your network. I recommend section 7 of these [cnntricks](http://lamda.nju.edu.cn/weixs/project/CNNTricks/CNNTricks.html) webpage to interpret your figures.",
      "votes": 1
    },
    {
      "id": 422304,
      "postDate": "2018-11-16T03:30:58.043Z",
      "content": "<p>Hi I am new to Machine Learning.\nI am curious about in real life how should we \"decide\" whether the algo is well trained?\nIn books/articles for rookie I learn I should compare the validation score and training score.</p>\n\n<p>For example if the validation score is already not bad (eg 0.9) and seems cant get improvement by using small learning rate, I may stop training to avoid overfitting. </p>\n\n<p>In kaggle, we have public leader board score so I know I should continue training if the score is far from the top performers. (so I will realise validation score 0.9 is not enough, either keep tuning the parameters/ switch model).</p>\n\n<p>But in real life there is no such kaggle score for us to evaluate. What make you think the model is ready? instead of keep training and get overfit?</p>\n\n<p>Thanks so much guys</p>",
      "rawMarkdown": "Hi I am new to Machine Learning.\nI am curious about in real life how should we \"decide\" whether the algo is well trained?\nIn books/articles for rookie I learn I should compare the validation score and training score.\n\nFor example if the validation score is already not bad (eg 0.9) and seems cant get improvement by using small learning rate, I may stop training to avoid overfitting. \n\nIn kaggle, we have public leader board score so I know I should continue training if the score is far from the top performers. (so I will realise validation score 0.9 is not enough, either keep tuning the parameters/ switch model).\n\nBut in real life there is no such kaggle score for us to evaluate. What make you think the model is ready? instead of keep training and get overfit?\n\nThanks so much guys\n",
      "votes": 2
    },
    {
      "id": 422929,
      "postDate": "2018-11-17T05:02:54.573Z",
      "content": "<p>Agree with Eric that you should be doing the section 7 stuff all the time.  Work with data and model when the plots tell you a different story on the training vs validation plots.</p>\n\n<p>For all of the Kaggle challenges I have worked on I probably spent 100 times the hours trying to improve than I would do in real life.  On the Mercedes challenge my first model hit R2 of 0.52 - three hundred hours of coding later I was at 0.54.  In real life I would have stopped at the first result and sought to improve the data set.  I think the top score was around 0.57.  A huge difference in Kaggle, but absolutely NO difference in real life.  </p>\n\n<p>When your training and validation plots are decent matches than you need to ask yourself what have I learned?  Lots of real world data sets are crap - the thing I most often learned was that my data was not sufficient to answer the question.  Deep learning is powerful but it will not pull a diamond from a pile of crap.</p>",
      "rawMarkdown": "Agree with Eric that you should be doing the section 7 stuff all the time.  Work with data and model when the plots tell you a different story on the training vs validation plots.\n\nFor all of the Kaggle challenges I have worked on I probably spent 100 times the hours trying to improve than I would do in real life.  On the Mercedes challenge my first model hit R2 of 0.52 - three hundred hours of coding later I was at 0.54.  In real life I would have stopped at the first result and sought to improve the data set.  I think the top score was around 0.57.  A huge difference in Kaggle, but absolutely NO difference in real life.  \n\nWhen your training and validation plots are decent matches than you need to ask yourself what have I learned?  Lots of real world data sets are crap - the thing I most often learned was that my data was not sufficient to answer the question.  Deep learning is powerful but it will not pull a diamond from a pile of crap."
    },
    {
      "id": 423139,
      "postDate": "2018-11-17T15:45:16.397Z",
      "content": "<p>thanks so much</p>",
      "rawMarkdown": "thanks so much"
    }
  ],
  "comments": [
    {
      "id": 422609,
      "author_name": "Eric Bouteillon",
      "author_url": "",
      "post_date": "2018-11-16T13:56:23.540000",
      "content": "<p>If you are using a deep learning network, usually drawing  loss and accuracy figures of both training and validation will give you good clues about the effectiveness of your network. I recommend section 7 of these <a href=\"http://lamda.nju.edu.cn/weixs/project/CNNTricks/CNNTricks.html\">cnntricks</a> webpage to interpret your figures.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 422929,
      "author_name": "PC Jimmmy",
      "author_url": "",
      "post_date": "2018-11-17T05:02:54.573000",
      "content": "<p>Agree with Eric that you should be doing the section 7 stuff all the time.  Work with data and model when the plots tell you a different story on the training vs validation plots.</p>\n\n<p>For all of the Kaggle challenges I have worked on I probably spent 100 times the hours trying to improve than I would do in real life.  On the Mercedes challenge my first model hit R2 of 0.52 - three hundred hours of coding later I was at 0.54.  In real life I would have stopped at the first result and sought to improve the data set.  I think the top score was around 0.57.  A huge difference in Kaggle, but absolutely NO difference in real life.  </p>\n\n<p>When your training and validation plots are decent matches than you need to ask yourself what have I learned?  Lots of real world data sets are crap - the thing I most often learned was that my data was not sufficient to answer the question.  Deep learning is powerful but it will not pull a diamond from a pile of crap.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 423139,
      "author_name": "Salaryman",
      "author_url": "",
      "post_date": "2018-11-17T15:45:16.397000",
      "content": "<p>thanks so much</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "422609": "If you are using a deep learning network, usually drawing  loss and accuracy figures of both training and validation will give you good clues about the effectiveness of your network. I recommend section 7 of these [cnntricks](http://lamda.nju.edu.cn/weixs/project/CNNTricks/CNNTricks.html) webpage to interpret your figures.",
    "422304": "Hi I am new to Machine Learning.\nI am curious about in real life how should we \"decide\" whether the algo is well trained?\nIn books/articles for rookie I learn I should compare the validation score and training score.\n\nFor example if the validation score is already not bad (eg 0.9) and seems cant get improvement by using small learning rate, I may stop training to avoid overfitting. \n\nIn kaggle, we have public leader board score so I know I should continue training if the score is far from the top performers. (so I will realise validation score 0.9 is not enough, either keep tuning the parameters/ switch model).\n\nBut in real life there is no such kaggle score for us to evaluate. What make you think the model is ready? instead of keep training and get overfit?\n\nThanks so much guys\n",
    "422929": "Agree with Eric that you should be doing the section 7 stuff all the time.  Work with data and model when the plots tell you a different story on the training vs validation plots.\n\nFor all of the Kaggle challenges I have worked on I probably spent 100 times the hours trying to improve than I would do in real life.  On the Mercedes challenge my first model hit R2 of 0.52 - three hundred hours of coding later I was at 0.54.  In real life I would have stopped at the first result and sought to improve the data set.  I think the top score was around 0.57.  A huge difference in Kaggle, but absolutely NO difference in real life.  \n\nWhen your training and validation plots are decent matches than you need to ask yourself what have I learned?  Lots of real world data sets are crap - the thing I most often learned was that my data was not sufficient to answer the question.  Deep learning is powerful but it will not pull a diamond from a pile of crap.",
    "423139": "thanks so much"
  }
}