{
  "id": 73761,
  "title": "14th place solution ",
  "url": "/competitions/quickdraw-doodle-recognition/discussion/73761",
  "author_name": "pudae",
  "post_date": "2018-12-05T12:36:48.925000",
  "votes": 23,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Because I joined this competition with just two weeks left,  I decided to train only basic classifier.\nSo,  It was really lucky to have achieved a relatively high score.</p>\n\n<p>After reading discussions (especially <a href=\"https://www.kaggle.com/hengck23\">Heng</a>'s posts), I decided to use following settings.</p>\n\n<h2>Dataset</h2>\n\n<ul>\n<li><em>Train</em>: All images including unrecognized</li>\n<li><em>Validation</em>: 500 images per class</li>\n<li>3 channel, each are 1/3, 2/3, 3/3 of the total strokes.</li>\n<li>224 x 224</li>\n</ul>\n\n<h2>Training</h2>\n\n<ul>\n<li>cross entropy loss</li>\n<li>batch size: 256 for se_resnext50, 128 for se_resnext101 and xception.</li>\n<li>adam optimizer</li>\n<li>learning rate 0.00025</li>\n<li>reduce learning rate when MAP@3 has stopped improving by half</li>\n<li>no augmentation</li>\n</ul>\n\n<h2>Inference</h2>\n\n<ul>\n<li>average last ten weights (I saved checkpoints every 5000 step)</li>\n<li>horizontal flip tta with weight 0.5</li>\n</ul>\n\n<p>In total I trained 3 models:</p>\n\n<ul>\n<li>se_resnext50, se_resnext101, xception</li>\n</ul>\n\n<p>All of them have similar Public LB scores, 0.947x.\nAfter ensembling all of them Public LB 0.950x.</p>\n\n<p>Congratulations to the winners and thanks for all the participants!</p>",
  "messages": [
    {
      "id": 433741,
      "postDate": "2018-12-05T12:36:48.927Z",
      "content": "<p>Because I joined this competition with just two weeks left,  I decided to train only basic classifier.\nSo,  It was really lucky to have achieved a relatively high score.</p>\n\n<p>After reading discussions (especially <a href=\"https://www.kaggle.com/hengck23\">Heng</a>'s posts), I decided to use following settings.</p>\n\n<h2>Dataset</h2>\n\n<ul>\n<li><em>Train</em>: All images including unrecognized</li>\n<li><em>Validation</em>: 500 images per class</li>\n<li>3 channel, each are 1/3, 2/3, 3/3 of the total strokes.</li>\n<li>224 x 224</li>\n</ul>\n\n<h2>Training</h2>\n\n<ul>\n<li>cross entropy loss</li>\n<li>batch size: 256 for se_resnext50, 128 for se_resnext101 and xception.</li>\n<li>adam optimizer</li>\n<li>learning rate 0.00025</li>\n<li>reduce learning rate when MAP@3 has stopped improving by half</li>\n<li>no augmentation</li>\n</ul>\n\n<h2>Inference</h2>\n\n<ul>\n<li>average last ten weights (I saved checkpoints every 5000 step)</li>\n<li>horizontal flip tta with weight 0.5</li>\n</ul>\n\n<p>In total I trained 3 models:</p>\n\n<ul>\n<li>se_resnext50, se_resnext101, xception</li>\n</ul>\n\n<p>All of them have similar Public LB scores, 0.947x.\nAfter ensembling all of them Public LB 0.950x.</p>\n\n<p>Congratulations to the winners and thanks for all the participants!</p>",
      "rawMarkdown": "Because I joined this competition with just two weeks left,  I decided to train only basic classifier.\nSo,  It was really lucky to have achieved a relatively high score.\n\nAfter reading discussions (especially [Heng](https://www.kaggle.com/hengck23)'s posts), I decided to use following settings.\n\n## Dataset\n- *Train*: All images including unrecognized\n- *Validation*: 500 images per class\n- 3 channel, each are 1/3, 2/3, 3/3 of the total strokes.\n- 224 x 224\n\n## Training\n- cross entropy loss\n- batch size: 256 for se_resnext50, 128 for se_resnext101 and xception.\n- adam optimizer\n- learning rate 0.00025\n- reduce learning rate when MAP@3 has stopped improving by half\n- no augmentation\n\n## Inference\n- average last ten weights (I saved checkpoints every 5000 step)\n- horizontal flip tta with weight 0.5\n\nIn total I trained 3 models:\n\n- se_resnext50, se_resnext101, xception\n\nAll of them have similar Public LB scores, 0.947x.\nAfter ensembling all of them Public LB 0.950x.\n\nCongratulations to the winners and thanks for all the participants!",
      "votes": 23
    },
    {
      "id": 434106,
      "postDate": "2018-12-05T22:38:28.290Z",
      "content": "<p>good work!</p>\n\n<p>\"3 channel, each are 1/3, 2/3, 3/3 of the total strokes.\"  - this gives good improvement over single baseline of total strokes for all channels.  </p>\n\n<p>I wonder if more channels gives better results, e.g. 1/4,2/4,3/4,4/4 for 4 channels. (you can simply just change the first convolution to accept 4 input channel and use the pretrain model weights)</p>",
      "rawMarkdown": "good work!\n\n\"3 channel, each are 1/3, 2/3, 3/3 of the total strokes.\"  - this gives good improvement over single baseline of total strokes for all channels.  \n\nI wonder if more channels gives better results, e.g. 1/4,2/4,3/4,4/4 for 4 channels. (you can simply just change the first convolution to accept 4 input channel and use the pretrain model weights)",
      "votes": 3
    },
    {
      "id": 435077,
      "postDate": "2018-12-07T13:13:59.247Z",
      "content": "<p>wow. only two weeks. </p>\n\n<p>do you use early-stopping? when the lr is reduced to a tiny number, the weight don't change much. </p>",
      "rawMarkdown": "wow. only two weeks. \n\ndo you use early-stopping? when the lr is reduced to a tiny number, the weight don't change much. "
    },
    {
      "id": 434340,
      "postDate": "2018-12-06T08:15:32.097Z",
      "content": "<p>Thanks for clear explanation and congratulations! A little pitty that you missed the gold medal.</p>",
      "rawMarkdown": "Thanks for clear explanation and congratulations! A little pitty that you missed the gold medal."
    },
    {
      "id": 433859,
      "postDate": "2018-12-05T15:16:26.540Z",
      "content": "<p>Congratulation and thanks for sharing! It is amazing that with only 2 weeks, you manage to reach that high level performance of three models. I have tried similar schemes but due to my inexperience, I couldn’t reach that level of LB. Could you please clarify further on these points : </p>\n\n<ul>\n<li>Learning rate of  0.00025 : did you use it from the start? without tuning at all ? </li>\n<li>Batch size and GPU memory : did you try to vary batch size, and how did you manage the batch size to fit in the GPU memory (especially because you use 224x224 which is quite big) ?</li>\n<li>Could you please give me a pointer of the SE-ResNext implementation that you use ? (I tried both Xception and ResNext, but only Xception that I was able to train to get a good LB)</li>\n</ul>",
      "rawMarkdown": "Congratulation and thanks for sharing! It is amazing that with only 2 weeks, you manage to reach that high level performance of three models. I have tried similar schemes but due to my inexperience, I couldn’t reach that level of LB. Could you please clarify further on these points : \n\n- Learning rate of  0.00025 : did you use it from the start? without tuning at all ? \n- Batch size and GPU memory : did you try to vary batch size, and how did you manage the batch size to fit in the GPU memory (especially because you use 224x224 which is quite big) ?\n- Could you please give me a pointer of the SE-ResNext implementation that you use ? (I tried both Xception and ResNext, but only Xception that I was able to train to get a good LB)",
      "replies": [
        {
          "id": 433900,
          "postDate": "2018-12-05T16:21:34.573Z",
          "content": "<ul>\n<li>I used resnet34 with 128x128 input to explore proper learning rate.\nI tried 0.001, 0.0005, 0.00025, 0.0001. After training for 1 hour each, the best one were selected.</li>\n<li>I chose the largest batch size available. I have 4 Titan X, but available gpus are varied because my work. I didn't try gradient accumulation.</li>\n<li><a href=\"https://github.com/Cadene/pretrained-models.pytorch\">Cadene/pretrained-models.pytorch</a> is used.</li>\n</ul>",
          "rawMarkdown": "- I used resnet34 with 128x128 input to explore proper learning rate.\n  I tried 0.001, 0.0005, 0.00025, 0.0001. After training for 1 hour each, the best one were selected.\n- I chose the largest batch size available. I have 4 Titan X, but available gpus are varied because my work. I didn't try gradient accumulation.\n- [Cadene/pretrained-models.pytorch](https://github.com/Cadene/pretrained-models.pytorch) is used.",
          "votes": 3
        },
        {
          "id": 434212,
          "postDate": "2018-12-06T03:53:15.177Z",
          "content": "<p>Thank you!  See you again in the next round :)</p>",
          "rawMarkdown": "Thank you!  See you again in the next round :)"
        },
        {
          "id": 435075,
          "postDate": "2018-12-07T13:10:00.050Z",
          "content": "<p>when you select the lr, do you look at the validation score to choose the best lr, rather the train loss right? </p>",
          "rawMarkdown": "when you select the lr, do you look at the validation score to choose the best lr, rather the train loss right? "
        },
        {
          "id": 435629,
          "postDate": "2018-12-08T12:14:47.640Z",
          "content": "<p>Yes, I used validation map@3 to choose learning rate.</p>",
          "rawMarkdown": "Yes, I used validation map@3 to choose learning rate."
        },
        {
          "id": 436738,
          "postDate": "2018-12-10T21:05:46.983Z",
          "content": "<p>I tried to your approach for finding the best LR, but what I found is, the larger the LR, the best the map@3 is... </p>",
          "rawMarkdown": "I tried to your approach for finding the best LR, but what I found is, the larger the LR, the best the map@3 is... "
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 434106,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-12-05T22:38:28.290000",
      "content": "<p>good work!</p>\n\n<p>\"3 channel, each are 1/3, 2/3, 3/3 of the total strokes.\"  - this gives good improvement over single baseline of total strokes for all channels.  </p>\n\n<p>I wonder if more channels gives better results, e.g. 1/4,2/4,3/4,4/4 for 4 channels. (you can simply just change the first convolution to accept 4 input channel and use the pretrain model weights)</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 435077,
      "author_name": "yimacs",
      "author_url": "",
      "post_date": "2018-12-07T13:13:59.247000",
      "content": "<p>wow. only two weeks. </p>\n\n<p>do you use early-stopping? when the lr is reduced to a tiny number, the weight don't change much. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 434340,
      "author_name": "Tommy Jiang",
      "author_url": "",
      "post_date": "2018-12-06T08:15:32.097000",
      "content": "<p>Thanks for clear explanation and congratulations! A little pitty that you missed the gold medal.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 433859,
      "author_name": "Neuron Engineer",
      "author_url": "",
      "post_date": "2018-12-05T15:16:26.540000",
      "content": "<p>Congratulation and thanks for sharing! It is amazing that with only 2 weeks, you manage to reach that high level performance of three models. I have tried similar schemes but due to my inexperience, I couldn’t reach that level of LB. Could you please clarify further on these points : </p>\n\n<ul>\n<li>Learning rate of  0.00025 : did you use it from the start? without tuning at all ? </li>\n<li>Batch size and GPU memory : did you try to vary batch size, and how did you manage the batch size to fit in the GPU memory (especially because you use 224x224 which is quite big) ?</li>\n<li>Could you please give me a pointer of the SE-ResNext implementation that you use ? (I tried both Xception and ResNext, but only Xception that I was able to train to get a good LB)</li>\n</ul>",
      "votes": 0,
      "replies": [
        {
          "id": 433900,
          "author_name": "pudae",
          "author_url": "",
          "post_date": "2018-12-05T16:21:34.573000",
          "content": "<ul>\n<li>I used resnet34 with 128x128 input to explore proper learning rate.\nI tried 0.001, 0.0005, 0.00025, 0.0001. After training for 1 hour each, the best one were selected.</li>\n<li>I chose the largest batch size available. I have 4 Titan X, but available gpus are varied because my work. I didn't try gradient accumulation.</li>\n<li><a href=\"https://github.com/Cadene/pretrained-models.pytorch\">Cadene/pretrained-models.pytorch</a> is used.</li>\n</ul>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 434212,
          "author_name": "Neuron Engineer",
          "author_url": "",
          "post_date": "2018-12-06T03:53:15.177000",
          "content": "<p>Thank you!  See you again in the next round :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 435075,
          "author_name": "yimacs",
          "author_url": "",
          "post_date": "2018-12-07T13:10:00.050000",
          "content": "<p>when you select the lr, do you look at the validation score to choose the best lr, rather the train loss right? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 435629,
          "author_name": "pudae",
          "author_url": "",
          "post_date": "2018-12-08T12:14:47.640000",
          "content": "<p>Yes, I used validation map@3 to choose learning rate.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 436738,
          "author_name": "yimacs",
          "author_url": "",
          "post_date": "2018-12-10T21:05:46.983000",
          "content": "<p>I tried to your approach for finding the best LR, but what I found is, the larger the LR, the best the map@3 is... </p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "433741": "Because I joined this competition with just two weeks left,  I decided to train only basic classifier.\nSo,  It was really lucky to have achieved a relatively high score.\n\nAfter reading discussions (especially [Heng](https://www.kaggle.com/hengck23)'s posts), I decided to use following settings.\n\n## Dataset\n- *Train*: All images including unrecognized\n- *Validation*: 500 images per class\n- 3 channel, each are 1/3, 2/3, 3/3 of the total strokes.\n- 224 x 224\n\n## Training\n- cross entropy loss\n- batch size: 256 for se_resnext50, 128 for se_resnext101 and xception.\n- adam optimizer\n- learning rate 0.00025\n- reduce learning rate when MAP@3 has stopped improving by half\n- no augmentation\n\n## Inference\n- average last ten weights (I saved checkpoints every 5000 step)\n- horizontal flip tta with weight 0.5\n\nIn total I trained 3 models:\n\n- se_resnext50, se_resnext101, xception\n\nAll of them have similar Public LB scores, 0.947x.\nAfter ensembling all of them Public LB 0.950x.\n\nCongratulations to the winners and thanks for all the participants!",
    "434106": "good work!\n\n\"3 channel, each are 1/3, 2/3, 3/3 of the total strokes.\"  - this gives good improvement over single baseline of total strokes for all channels.  \n\nI wonder if more channels gives better results, e.g. 1/4,2/4,3/4,4/4 for 4 channels. (you can simply just change the first convolution to accept 4 input channel and use the pretrain model weights)",
    "435077": "wow. only two weeks. \n\ndo you use early-stopping? when the lr is reduced to a tiny number, the weight don't change much. ",
    "434340": "Thanks for clear explanation and congratulations! A little pitty that you missed the gold medal.",
    "433859": "Congratulation and thanks for sharing! It is amazing that with only 2 weeks, you manage to reach that high level performance of three models. I have tried similar schemes but due to my inexperience, I couldn’t reach that level of LB. Could you please clarify further on these points : \n\n- Learning rate of  0.00025 : did you use it from the start? without tuning at all ? \n- Batch size and GPU memory : did you try to vary batch size, and how did you manage the batch size to fit in the GPU memory (especially because you use 224x224 which is quite big) ?\n- Could you please give me a pointer of the SE-ResNext implementation that you use ? (I tried both Xception and ResNext, but only Xception that I was able to train to get a good LB)"
  }
}