{
  "id": 72341,
  "title": "myths about pretrained model?",
  "url": "/competitions/quickdraw-doodle-recognition/discussion/72341",
  "author_name": "hengck23",
  "post_date": "2018-11-22T11:26:40.798000",
  "votes": 23,
  "comment_count": 10,
  "views": 0,
  "content": "<p><a href=\"https://arxiv.org/abs/1811.08883\">https://arxiv.org/abs/1811.08883</a></p>\n\n<p>\"Rethinking ImageNet Pre-training\" - Kaiming He, Ross Girshick, Piotr Dollár, arxiv 2018</p>\n\n<p>\"Experiments show that ImageNet pre-training speeds up convergence early in training, but does not necessarily provide regularization or improve final target task accuracy\"</p>\n\n<p>\"Training from scratch can be no worse than its ImageNet pre-training counterparts under many circumstances, down to as few as 10k COCO images\"</p>",
  "messages": [
    {
      "id": 425978,
      "postDate": "2018-11-22T11:26:40.800Z",
      "content": "<p><a href=\"https://arxiv.org/abs/1811.08883\">https://arxiv.org/abs/1811.08883</a></p>\n\n<p>\"Rethinking ImageNet Pre-training\" - Kaiming He, Ross Girshick, Piotr Dollár, arxiv 2018</p>\n\n<p>\"Experiments show that ImageNet pre-training speeds up convergence early in training, but does not necessarily provide regularization or improve final target task accuracy\"</p>\n\n<p>\"Training from scratch can be no worse than its ImageNet pre-training counterparts under many circumstances, down to as few as 10k COCO images\"</p>",
      "rawMarkdown": "https://arxiv.org/abs/1811.08883\n\n\"Rethinking ImageNet Pre-training\" - Kaiming He, Ross Girshick, Piotr Dollár, arxiv 2018\n\n\n\"Experiments show that ImageNet pre-training speeds up convergence early in training, but does not necessarily provide regularization or improve final target task accuracy\"\n\n\"Training from scratch can be no worse than its ImageNet pre-training counterparts under many circumstances, down to as few as 10k COCO images\"\n",
      "votes": 23
    },
    {
      "id": 426278,
      "postDate": "2018-11-23T02:10:53.507Z",
      "content": "<p>the conclusion from the paper should be:</p>\n\n<ol>\n<li><p>train from scratch or using pretrained imagenet leads to almost  same accuracy results</p></li>\n<li><p>However, using pretrained imagenet would speed up convergence. (i.e. obtain the results with much fewer iterations)</p></li>\n</ol>\n\n<p>Where is the myth? There has been discussion previously that pretrained model leads to \"better\" results. But this paper shows that this is not true. So long  as your target data is sufficiently large enough, random initialisation works as well pretrain models. But pretrained  model still have the advantage of faster convergence.</p>\n\n<p>As an example, typical COCO dataset has 118k train images. the paper shows:</p>\n\n<p>at 118K, random initialisation = pretrain  results.</p>\n\n<p>even at 10 k, random initialisation = pretrain  results (this is surprising)</p>\n\n<p>but at 1 k, pretrain becomes much better. There is much overfitting due to small data for large model.</p>",
      "rawMarkdown": "the conclusion from the paper should be:\n\n1. train from scratch or using pretrained imagenet leads to almost  same accuracy results\n\n2. However, using pretrained imagenet would speed up convergence. (i.e. obtain the results with much fewer iterations)\n\nWhere is the myth? There has been discussion previously that pretrained model leads to \"better\" results. But this paper shows that this is not true. So long  as your target data is sufficiently large enough, random initialisation works as well pretrain models. But pretrained  model still have the advantage of faster convergence.\n\nAs an example, typical COCO dataset has 118k train images. the paper shows:\n\nat 118K, random initialisation = pretrain  results.\n\neven at 10 k, random initialisation = pretrain  results (this is surprising)\n\nbut at 1 k, pretrain becomes much better. There is much overfitting due to small data for large model.",
      "votes": 5
    },
    {
      "id": 426272,
      "postDate": "2018-11-23T01:57:21.287Z",
      "content": "<p>a few days ago , tried imagenet weight, and that's not good for me. I trained from scratch.</p>",
      "rawMarkdown": "a few days ago , tried imagenet weight, and that's not good for me. I trained from scratch.",
      "votes": 1
    },
    {
      "id": 427846,
      "postDate": "2018-11-26T08:58:14.983Z",
      "content": "<p>how do you even use a pretrained model anyway? The input has different number of channels.</p>",
      "rawMarkdown": "how do you even use a pretrained model anyway? The input has different number of channels.",
      "replies": [
        {
          "id": 427998,
          "postDate": "2018-11-26T15:28:01.953Z",
          "content": "<p>You can just create an image with 3 channels</p>",
          "rawMarkdown": "You can just create an image with 3 channels"
        }
      ]
    },
    {
      "id": 426870,
      "postDate": "2018-11-24T03:34:38.053Z",
      "content": "<p>From my experience on this match, maybe model from scratch is comparable to the fine tune on pretrained model, however time is the biggest concern as we only have limited time to finish this competition</p>",
      "rawMarkdown": "From my experience on this match, maybe model from scratch is comparable to the fine tune on pretrained model, however time is the biggest concern as we only have limited time to finish this competition"
    },
    {
      "id": 426255,
      "postDate": "2018-11-23T01:10:30.623Z",
      "content": "<p>Same issue in this competition.</p>",
      "rawMarkdown": "Same issue in this competition."
    },
    {
      "id": 426093,
      "postDate": "2018-11-22T15:37:05.557Z",
      "content": "<p>Same here, model with pretrained weights is performing worse than training a model from scratch.\nAlthough, I feel that I am doing something wrong.</p>",
      "rawMarkdown": "Same here, model with pretrained weights is performing worse than training a model from scratch.\nAlthough, I feel that I am doing something wrong."
    },
    {
      "id": 426035,
      "postDate": "2018-11-22T13:27:20.620Z",
      "content": "<p>Very interesting. It reminds me of the credibility theory within actuarial science, where we blend universal experience with more specific experience. The more credible (more samples) your data, the more you want to rely on the specific experience instead of the general.</p>",
      "rawMarkdown": "Very interesting. It reminds me of the credibility theory within actuarial science, where we blend universal experience with more specific experience. The more credible (more samples) your data, the more you want to rely on the specific experience instead of the general."
    },
    {
      "id": 426015,
      "postDate": "2018-11-22T12:43:23.357Z",
      "content": "<p>I had the worse score with pretrained imagenet weights than training from scratch. I haven't trained long enough tho.</p>",
      "rawMarkdown": "I had the worse score with pretrained imagenet weights than training from scratch. I haven't trained long enough tho."
    },
    {
      "id": 425987,
      "postDate": "2018-11-22T11:48:45.263Z",
      "content": "<p>As for the small dataset, pretrained model can avoid overfitting. But this competition, maybe we can train a model from scratch because of the big dataset.</p>",
      "rawMarkdown": "As for the small dataset, pretrained model can avoid overfitting. But this competition, maybe we can train a model from scratch because of the big dataset."
    }
  ],
  "comments": [
    {
      "id": 426278,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-11-23T02:10:53.507000",
      "content": "<p>the conclusion from the paper should be:</p>\n\n<ol>\n<li><p>train from scratch or using pretrained imagenet leads to almost  same accuracy results</p></li>\n<li><p>However, using pretrained imagenet would speed up convergence. (i.e. obtain the results with much fewer iterations)</p></li>\n</ol>\n\n<p>Where is the myth? There has been discussion previously that pretrained model leads to \"better\" results. But this paper shows that this is not true. So long  as your target data is sufficiently large enough, random initialisation works as well pretrain models. But pretrained  model still have the advantage of faster convergence.</p>\n\n<p>As an example, typical COCO dataset has 118k train images. the paper shows:</p>\n\n<p>at 118K, random initialisation = pretrain  results.</p>\n\n<p>even at 10 k, random initialisation = pretrain  results (this is surprising)</p>\n\n<p>but at 1 k, pretrain becomes much better. There is much overfitting due to small data for large model.</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 426272,
      "author_name": "yyqing",
      "author_url": "",
      "post_date": "2018-11-23T01:57:21.287000",
      "content": "<p>a few days ago , tried imagenet weight, and that's not good for me. I trained from scratch.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 427846,
      "author_name": "Master",
      "author_url": "",
      "post_date": "2018-11-26T08:58:14.983000",
      "content": "<p>how do you even use a pretrained model anyway? The input has different number of channels.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 427998,
          "author_name": "Strideradu",
          "author_url": "",
          "post_date": "2018-11-26T15:28:01.953000",
          "content": "<p>You can just create an image with 3 channels</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 426870,
      "author_name": "Strideradu",
      "author_url": "",
      "post_date": "2018-11-24T03:34:38.053000",
      "content": "<p>From my experience on this match, maybe model from scratch is comparable to the fine tune on pretrained model, however time is the biggest concern as we only have limited time to finish this competition</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 426255,
      "author_name": "good good study",
      "author_url": "",
      "post_date": "2018-11-23T01:10:30.623000",
      "content": "<p>Same issue in this competition.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 426093,
      "author_name": "Rajath",
      "author_url": "",
      "post_date": "2018-11-22T15:37:05.557000",
      "content": "<p>Same here, model with pretrained weights is performing worse than training a model from scratch.\nAlthough, I feel that I am doing something wrong.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 426035,
      "author_name": "Allen",
      "author_url": "",
      "post_date": "2018-11-22T13:27:20.620000",
      "content": "<p>Very interesting. It reminds me of the credibility theory within actuarial science, where we blend universal experience with more specific experience. The more credible (more samples) your data, the more you want to rely on the specific experience instead of the general.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 426015,
      "author_name": "Ankit Sati",
      "author_url": "",
      "post_date": "2018-11-22T12:43:23.357000",
      "content": "<p>I had the worse score with pretrained imagenet weights than training from scratch. I haven't trained long enough tho.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 425987,
      "author_name": "Gary",
      "author_url": "",
      "post_date": "2018-11-22T11:48:45.263000",
      "content": "<p>As for the small dataset, pretrained model can avoid overfitting. But this competition, maybe we can train a model from scratch because of the big dataset.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "425978": "https://arxiv.org/abs/1811.08883\n\n\"Rethinking ImageNet Pre-training\" - Kaiming He, Ross Girshick, Piotr Dollár, arxiv 2018\n\n\n\"Experiments show that ImageNet pre-training speeds up convergence early in training, but does not necessarily provide regularization or improve final target task accuracy\"\n\n\"Training from scratch can be no worse than its ImageNet pre-training counterparts under many circumstances, down to as few as 10k COCO images\"\n",
    "426278": "the conclusion from the paper should be:\n\n1. train from scratch or using pretrained imagenet leads to almost  same accuracy results\n\n2. However, using pretrained imagenet would speed up convergence. (i.e. obtain the results with much fewer iterations)\n\nWhere is the myth? There has been discussion previously that pretrained model leads to \"better\" results. But this paper shows that this is not true. So long  as your target data is sufficiently large enough, random initialisation works as well pretrain models. But pretrained  model still have the advantage of faster convergence.\n\nAs an example, typical COCO dataset has 118k train images. the paper shows:\n\nat 118K, random initialisation = pretrain  results.\n\neven at 10 k, random initialisation = pretrain  results (this is surprising)\n\nbut at 1 k, pretrain becomes much better. There is much overfitting due to small data for large model.",
    "426272": "a few days ago , tried imagenet weight, and that's not good for me. I trained from scratch.",
    "427846": "how do you even use a pretrained model anyway? The input has different number of channels.",
    "426870": "From my experience on this match, maybe model from scratch is comparable to the fine tune on pretrained model, however time is the biggest concern as we only have limited time to finish this competition",
    "426255": "Same issue in this competition.",
    "426093": "Same here, model with pretrained weights is performing worse than training a model from scratch.\nAlthough, I feel that I am doing something wrong.",
    "426035": "Very interesting. It reminds me of the credibility theory within actuarial science, where we blend universal experience with more specific experience. The more credible (more samples) your data, the more you want to rely on the specific experience instead of the general.",
    "426015": "I had the worse score with pretrained imagenet weights than training from scratch. I haven't trained long enough tho.",
    "425987": "As for the small dataset, pretrained model can avoid overfitting. But this competition, maybe we can train a model from scratch because of the big dataset."
  }
}