{
  "id": 73808,
  "title": "11th place solution with limited hardware resources up to 2xP40",
  "url": "/competitions/quickdraw-doodle-recognition/discussion/73808",
  "author_name": "Hai Nam Nguyen",
  "post_date": "2018-12-05T17:36:28.209000",
  "votes": 27,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Our solution is pretty simple.\n <strong>1. CNN</strong>\n- We tried some relatively small models with 100k/class and 64x64, 128x128 and 224x224 input size first.\n- Then we retrained with full data + weighted loss on se-resnext101 (.944),  se-resnext50 (.943), se-resnet50 (.942),  resnet50 (.942), densenet169 (.939), xception(.938), densenet121(.934) with batch size &gt;= 400, 128x128 input size.\n- It took 4 weeks.</p>\n\n<p><strong>2. RNN</strong>\n - We tried some public kernels and modified them (deeper, bigger and stronger, replacing LSTM by attention/GRU, using timestamp on raw data) with 100k/class first.\n - Then, we trained some best of them with full data and got best results around 0.93x.\n - It took nearly 2 weeks.</p>\n\n<p><strong>3. Inference</strong>\n - TTA (hflip) + Ensemble (0.8 * CNN + 0.2 * RNN) + optimization(Secret Sauce + magic wand).\n - Private LB without optimization: 0.94701</p>\n\n<p><strong>4. Our weakness</strong>\n -  Limited on something as in the article, we were not able to try some bigger input size with big enough batch size.\n - We did not find a suitable way of filtering the training dataset on unrecognized images.\n - Only depend on optimization made our model roughly overfit. However, the best model on the public LB was our best model on private LB, too. Big thank to God for that!</p>",
  "messages": [
    {
      "id": 433958,
      "postDate": "2018-12-05T17:36:28.210Z",
      "content": "<p>Our solution is pretty simple.\n <strong>1. CNN</strong>\n- We tried some relatively small models with 100k/class and 64x64, 128x128 and 224x224 input size first.\n- Then we retrained with full data + weighted loss on se-resnext101 (.944),  se-resnext50 (.943), se-resnet50 (.942),  resnet50 (.942), densenet169 (.939), xception(.938), densenet121(.934) with batch size &gt;= 400, 128x128 input size.\n- It took 4 weeks.</p>\n\n<p><strong>2. RNN</strong>\n - We tried some public kernels and modified them (deeper, bigger and stronger, replacing LSTM by attention/GRU, using timestamp on raw data) with 100k/class first.\n - Then, we trained some best of them with full data and got best results around 0.93x.\n - It took nearly 2 weeks.</p>\n\n<p><strong>3. Inference</strong>\n - TTA (hflip) + Ensemble (0.8 * CNN + 0.2 * RNN) + optimization(Secret Sauce + magic wand).\n - Private LB without optimization: 0.94701</p>\n\n<p><strong>4. Our weakness</strong>\n -  Limited on something as in the article, we were not able to try some bigger input size with big enough batch size.\n - We did not find a suitable way of filtering the training dataset on unrecognized images.\n - Only depend on optimization made our model roughly overfit. However, the best model on the public LB was our best model on private LB, too. Big thank to God for that!</p>",
      "rawMarkdown": "Our solution is pretty simple.\n **1. CNN**\n- We tried some relatively small models with 100k/class and 64x64, 128x128 and 224x224 input size first.\n- Then we retrained with full data + weighted loss on se-resnext101 (.944),  se-resnext50 (.943), se-resnet50 (.942),  resnet50 (.942), densenet169 (.939), xception(.938), densenet121(.934) with batch size &gt;= 400, 128x128 input size.\n- It took 4 weeks.\n\n**2. RNN**\n - We tried some public kernels and modified them (deeper, bigger and stronger, replacing LSTM by attention/GRU, using timestamp on raw data) with 100k/class first.\n - Then, we trained some best of them with full data and got best results around 0.93x.\n - It took nearly 2 weeks.\n\n**3. Inference**\n - TTA (hflip) + Ensemble (0.8 * CNN + 0.2 * RNN) + optimization(Secret Sauce + magic wand).\n - Private LB without optimization: 0.94701\n\n**4. Our weakness**\n -  Limited on something as in the article, we were not able to try some bigger input size with big enough batch size.\n - We did not find a suitable way of filtering the training dataset on unrecognized images.\n - Only depend on optimization made our model roughly overfit. However, the best model on the public LB was our best model on private LB, too. Big thank to God for that!",
      "votes": 27
    },
    {
      "id": 434571,
      "postDate": "2018-12-06T15:55:21.143Z",
      "content": "<p>Congratulations and Thanks for sharing. </p>\n\n<p>Beginner question: How do you decide which architectures to pick? Is it based on experimentation and what seems more promising or is there another methodology? </p>",
      "rawMarkdown": "Congratulations and Thanks for sharing. \n\nBeginner question: How do you decide which architectures to pick? Is it based on experimentation and what seems more promising or is there another methodology? ",
      "votes": 1,
      "replies": [
        {
          "id": 434621,
          "postDate": "2018-12-06T17:24:40.480Z",
          "content": "<p>Good question, I think that it depends on not only engineering skills but also hardware resources.\nFrom my point of view, I picked our best models based on Input size, Batch size, model architecture, and training tricks.\nAlso, there is a good paper helping to choose model architecture: <a href=\"https://arxiv.org/pdf/1810.00736.pdf\">https://arxiv.org/pdf/1810.00736.pdf</a></p>",
          "rawMarkdown": "Good question, I think that it depends on not only engineering skills but also hardware resources.\nFrom my point of view, I picked our best models based on Input size, Batch size, model architecture, and training tricks.\nAlso, there is a good paper helping to choose model architecture: https://arxiv.org/pdf/1810.00736.pdf\n",
          "votes": 2
        }
      ]
    },
    {
      "id": 434337,
      "postDate": "2018-12-06T08:08:35.643Z",
      "content": "<p>Congratulation and thanks for sharing</p>",
      "rawMarkdown": "Congratulation and thanks for sharing",
      "votes": 1,
      "replies": [
        {
          "id": 434377,
          "postDate": "2018-12-06T09:51:28.203Z",
          "content": "<p>I wish one day we could take part in a challenge as a team!</p>",
          "rawMarkdown": "I wish one day we could take part in a challenge as a team!",
          "votes": 1
        },
        {
          "id": 434380,
          "postDate": "2018-12-06T09:55:05.907Z",
          "content": "<p>I'm in, whenever you want ;)</p>",
          "rawMarkdown": "I'm in, whenever you want ;)"
        }
      ]
    },
    {
      "id": 434329,
      "postDate": "2018-12-06T07:55:19.610Z",
      "content": "<p>Thanks for sharing and I really love your avatar (Louis Cha).</p>",
      "rawMarkdown": "Thanks for sharing and I really love your avatar (Louis Cha).",
      "votes": 1,
      "replies": [
        {
          "id": 434480,
          "postDate": "2018-12-06T13:24:35.870Z",
          "content": "<p>Yeah, even though I'm Vietnamese, he has been my idol for almost 20 years.</p>",
          "rawMarkdown": "Yeah, even though I'm Vietnamese, he has been my idol for almost 20 years."
        }
      ]
    },
    {
      "id": 434210,
      "postDate": "2018-12-06T03:51:20.970Z",
      "content": "<p>Hello <a href=\"/hainamnguyen\">@hainamnguyen</a>, thank you and congratulation! Could you please clarify some details on your \"weighted loss\" ?</p>",
      "rawMarkdown": "Hello @hainamnguyen, thank you and congratulation! Could you please clarify some details on your \"weighted loss\" ?",
      "votes": 2,
      "replies": [
        {
          "id": 434284,
          "postDate": "2018-12-06T06:45:52.077Z",
          "content": "<p>It's like class-weight in fit-generator function on Keras with class_weight was the inverted number distribution of training dataset.</p>",
          "rawMarkdown": "It's like class-weight in fit-generator function on Keras with class_weight was the inverted number distribution of training dataset.",
          "votes": 4
        }
      ]
    },
    {
      "id": 434396,
      "postDate": "2018-12-06T10:13:08.557Z",
      "content": "<p>Impressive results even though you were limited !</p>",
      "rawMarkdown": "Impressive results even though you were limited !"
    }
  ],
  "comments": [
    {
      "id": 434571,
      "author_name": "Sanyam Bhutani",
      "author_url": "",
      "post_date": "2018-12-06T15:55:21.143000",
      "content": "<p>Congratulations and Thanks for sharing. </p>\n\n<p>Beginner question: How do you decide which architectures to pick? Is it based on experimentation and what seems more promising or is there another methodology? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 434621,
          "author_name": "Hai Nam Nguyen",
          "author_url": "",
          "post_date": "2018-12-06T17:24:40.480000",
          "content": "<p>Good question, I think that it depends on not only engineering skills but also hardware resources.\nFrom my point of view, I picked our best models based on Input size, Batch size, model architecture, and training tricks.\nAlso, there is a good paper helping to choose model architecture: <a href=\"https://arxiv.org/pdf/1810.00736.pdf\">https://arxiv.org/pdf/1810.00736.pdf</a></p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 434337,
      "author_name": "Hoàng Tùng Lâm",
      "author_url": "",
      "post_date": "2018-12-06T08:08:35.643000",
      "content": "<p>Congratulation and thanks for sharing</p>",
      "votes": 1,
      "replies": [
        {
          "id": 434377,
          "author_name": "Hai Nam Nguyen",
          "author_url": "",
          "post_date": "2018-12-06T09:51:28.203000",
          "content": "<p>I wish one day we could take part in a challenge as a team!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 434380,
          "author_name": "Hoàng Tùng Lâm",
          "author_url": "",
          "post_date": "2018-12-06T09:55:05.907000",
          "content": "<p>I'm in, whenever you want ;)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 434329,
      "author_name": "Tommy Jiang",
      "author_url": "",
      "post_date": "2018-12-06T07:55:19.610000",
      "content": "<p>Thanks for sharing and I really love your avatar (Louis Cha).</p>",
      "votes": 1,
      "replies": [
        {
          "id": 434480,
          "author_name": "Hai Nam Nguyen",
          "author_url": "",
          "post_date": "2018-12-06T13:24:35.870000",
          "content": "<p>Yeah, even though I'm Vietnamese, he has been my idol for almost 20 years.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 434210,
      "author_name": "Neuron Engineer",
      "author_url": "",
      "post_date": "2018-12-06T03:51:20.970000",
      "content": "<p>Hello <a href=\"/hainamnguyen\">@hainamnguyen</a>, thank you and congratulation! Could you please clarify some details on your \"weighted loss\" ?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 434284,
          "author_name": "Hai Nam Nguyen",
          "author_url": "",
          "post_date": "2018-12-06T06:45:52.077000",
          "content": "<p>It's like class-weight in fit-generator function on Keras with class_weight was the inverted number distribution of training dataset.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 434396,
      "author_name": "Theo Viel",
      "author_url": "",
      "post_date": "2018-12-06T10:13:08.557000",
      "content": "<p>Impressive results even though you were limited !</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "433958": "Our solution is pretty simple.\n **1. CNN**\n- We tried some relatively small models with 100k/class and 64x64, 128x128 and 224x224 input size first.\n- Then we retrained with full data + weighted loss on se-resnext101 (.944),  se-resnext50 (.943), se-resnet50 (.942),  resnet50 (.942), densenet169 (.939), xception(.938), densenet121(.934) with batch size &gt;= 400, 128x128 input size.\n- It took 4 weeks.\n\n**2. RNN**\n - We tried some public kernels and modified them (deeper, bigger and stronger, replacing LSTM by attention/GRU, using timestamp on raw data) with 100k/class first.\n - Then, we trained some best of them with full data and got best results around 0.93x.\n - It took nearly 2 weeks.\n\n**3. Inference**\n - TTA (hflip) + Ensemble (0.8 * CNN + 0.2 * RNN) + optimization(Secret Sauce + magic wand).\n - Private LB without optimization: 0.94701\n\n**4. Our weakness**\n -  Limited on something as in the article, we were not able to try some bigger input size with big enough batch size.\n - We did not find a suitable way of filtering the training dataset on unrecognized images.\n - Only depend on optimization made our model roughly overfit. However, the best model on the public LB was our best model on private LB, too. Big thank to God for that!",
    "434571": "Congratulations and Thanks for sharing. \n\nBeginner question: How do you decide which architectures to pick? Is it based on experimentation and what seems more promising or is there another methodology? ",
    "434337": "Congratulation and thanks for sharing",
    "434329": "Thanks for sharing and I really love your avatar (Louis Cha).",
    "434210": "Hello @hainamnguyen, thank you and congratulation! Could you please clarify some details on your \"weighted loss\" ?",
    "434396": "Impressive results even though you were limited !"
  }
}