{
  "id": 73754,
  "title": "21st place solution [LB 0.948] on simplified data only",
  "url": "/competitions/quickdraw-doodle-recognition/discussion/73754",
  "author_name": "Pavel Tsai",
  "post_date": "2018-12-05T11:47:37.010000",
  "votes": 18,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Thanks everybody for the competition! The task was really interesting and had huge amounts of data. Thanks very much to my teammate who writes really perfect code on Python. You should understand me, he writes <strong>comments</strong> to functions and uses <strong>types</strong> in Python! Also thanks to <a href=\"https://www.kaggle.com/scitator\"></a><a href=\"/scitator\">@scitator</a> for his perfect training ML framework <a href=\"https://github.com/Scitator/catalyst\">Catalyst</a>, all models were trained with the help of it.</p>\n\n<p>Now I’d like to tell you about my and <a href=\"https://www.kaggle.com/artyomp\">Artyom Palvelev</a> solution using simplified data only.</p>\n\n<h1>data preprocessing</h1>\n\n<p>Pictures of size 128x128 gave the best combination of score and training time. I don’t have time information, so I encoded the following data in three channels:</p>\n\n<ol>\n<li>The index of line (linearly from 10 for the first line to 255 for the last one)</li>\n<li>The number of strokes in the line</li>\n<li>Just constant 255\n<h1>Our models</h1></li>\n</ol>\n\n<p>Firstly, I tried models from forum like MobileNetV2, but they showed poor performance, so I trained something deeper. My first good model was SE_ResNext50 0.942 LB. I used CosineAnnealingLR and averaged top4 checkpoints.\nThen we merged with Artyom and tried different model architectures. Blend with his models gave 0.944 LB. In the final submission we had:</p>\n\n<ol>\n<li>SE_ResNext50 (~0.942LB)</li>\n<li>SE_ResNext101 (~0.944LB)</li>\n<li>NASNet-A-Large (~0.944LB)</li>\n<li>SENet154 (~0.945LB)</li>\n<li>CBAM_ResNet50 (~0.941LB)</li>\n</ol>\n\n<p>We also used <em>gradient accumulation</em> to increase batch size to 1024 because data was noisy and bigger batch size gave better performance.</p>\n\n<h1>LGBM</h1>\n\n<p>We decided to use LGBM to ensemble our models because it is fast and usually gives better results than other methods. We used same idea as <a href=\"https://www.kaggle.com/pavelost\">Pavel Ostyakov</a> described in his <a href=\"https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/45733\">5th place Cdiscount solution</a>. We predicted top10 classes by our best network, concatenated probabilities of other networks, added class_id feature and gave binary label: whether the class_id is correct or not. This method resulted in validation score 0.001 higher than a simple average. So, if you have enough time, always try LGBM to ensemble.\nFinal submit with LGBM -&gt; 0.948 LB</p>\n\n<h1>What we tried and didn’t work</h1>\n\n<ul>\n<li>CatBoost, RF, XGBoost, ensemble of 3rd level</li>\n<li>Tuning models on clean data. We predicted the train dataset with our best models and dropped out pictures with small probability for the correct class (about 1M samples). It gave small boost on validation, but we didn’t have enough time, so we only trained a couple of models for about 2-3 epochs. So, I suppose, it is also a good idea for noisy data.</li>\n</ul>\n\n<h1>What we didn’t try but it worked</h1>\n\n<ol>\n<li>We didn’t noticed that the test is balanced (we could use same technique as in <a href=\"https://www.kaggle.com/c/sp-society-camera-model-identification\">Camera Model Identification</a> and get gold) Others say, it give plus ~0.7% score to any submission.</li>\n<li>TTA with deleting 20% strokes (people on forum said it also improved score). I tried just light augmentations like flips and shift_scale_rotate but they made score on validation even worse.</li>\n</ol>",
  "messages": [
    {
      "id": 433716,
      "postDate": "2018-12-05T11:47:37.010Z",
      "content": "<p>Thanks everybody for the competition! The task was really interesting and had huge amounts of data. Thanks very much to my teammate who writes really perfect code on Python. You should understand me, he writes <strong>comments</strong> to functions and uses <strong>types</strong> in Python! Also thanks to <a href=\"https://www.kaggle.com/scitator\"></a><a href=\"/scitator\">@scitator</a> for his perfect training ML framework <a href=\"https://github.com/Scitator/catalyst\">Catalyst</a>, all models were trained with the help of it.</p>\n\n<p>Now I’d like to tell you about my and <a href=\"https://www.kaggle.com/artyomp\">Artyom Palvelev</a> solution using simplified data only.</p>\n\n<h1>data preprocessing</h1>\n\n<p>Pictures of size 128x128 gave the best combination of score and training time. I don’t have time information, so I encoded the following data in three channels:</p>\n\n<ol>\n<li>The index of line (linearly from 10 for the first line to 255 for the last one)</li>\n<li>The number of strokes in the line</li>\n<li>Just constant 255\n<h1>Our models</h1></li>\n</ol>\n\n<p>Firstly, I tried models from forum like MobileNetV2, but they showed poor performance, so I trained something deeper. My first good model was SE_ResNext50 0.942 LB. I used CosineAnnealingLR and averaged top4 checkpoints.\nThen we merged with Artyom and tried different model architectures. Blend with his models gave 0.944 LB. In the final submission we had:</p>\n\n<ol>\n<li>SE_ResNext50 (~0.942LB)</li>\n<li>SE_ResNext101 (~0.944LB)</li>\n<li>NASNet-A-Large (~0.944LB)</li>\n<li>SENet154 (~0.945LB)</li>\n<li>CBAM_ResNet50 (~0.941LB)</li>\n</ol>\n\n<p>We also used <em>gradient accumulation</em> to increase batch size to 1024 because data was noisy and bigger batch size gave better performance.</p>\n\n<h1>LGBM</h1>\n\n<p>We decided to use LGBM to ensemble our models because it is fast and usually gives better results than other methods. We used same idea as <a href=\"https://www.kaggle.com/pavelost\">Pavel Ostyakov</a> described in his <a href=\"https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/45733\">5th place Cdiscount solution</a>. We predicted top10 classes by our best network, concatenated probabilities of other networks, added class_id feature and gave binary label: whether the class_id is correct or not. This method resulted in validation score 0.001 higher than a simple average. So, if you have enough time, always try LGBM to ensemble.\nFinal submit with LGBM -&gt; 0.948 LB</p>\n\n<h1>What we tried and didn’t work</h1>\n\n<ul>\n<li>CatBoost, RF, XGBoost, ensemble of 3rd level</li>\n<li>Tuning models on clean data. We predicted the train dataset with our best models and dropped out pictures with small probability for the correct class (about 1M samples). It gave small boost on validation, but we didn’t have enough time, so we only trained a couple of models for about 2-3 epochs. So, I suppose, it is also a good idea for noisy data.</li>\n</ul>\n\n<h1>What we didn’t try but it worked</h1>\n\n<ol>\n<li>We didn’t noticed that the test is balanced (we could use same technique as in <a href=\"https://www.kaggle.com/c/sp-society-camera-model-identification\">Camera Model Identification</a> and get gold) Others say, it give plus ~0.7% score to any submission.</li>\n<li>TTA with deleting 20% strokes (people on forum said it also improved score). I tried just light augmentations like flips and shift_scale_rotate but they made score on validation even worse.</li>\n</ol>",
      "rawMarkdown": "Thanks everybody for the competition! The task was really interesting and had huge amounts of data. Thanks very much to my teammate who writes really perfect code on Python. You should understand me, he writes **comments** to functions and uses **types** in Python! Also thanks to [@scitator][1] for his perfect training ML framework [Catalyst][2], all models were trained with the help of it.\n \nNow I’d like to tell you about my and [Artyom Palvelev][3] solution using simplified data only.\n#data preprocessing#\nPictures of size 128x128 gave the best combination of score and training time. I don’t have time information, so I encoded the following data in three channels:\n\n 1. The index of line (linearly from 10 for the first line to 255 for the last one)\n 2. The number of strokes in the line\n 3. Just constant 255\n#Our models#\nFirstly, I tried models from forum like MobileNetV2, but they showed poor performance, so I trained something deeper. My first good model was SE_ResNext50 0.942 LB. I used CosineAnnealingLR and averaged top4 checkpoints.\nThen we merged with Artyom and tried different model architectures. Blend with his models gave 0.944 LB. In the final submission we had:\n\n1. SE_ResNext50 (~0.942LB)\n2. SE_ResNext101 (~0.944LB)\n3. NASNet-A-Large (~0.944LB)\n4. SENet154 (~0.945LB)\n5. CBAM_ResNet50 (~0.941LB)\n\nWe also used *gradient accumulation* to increase batch size to 1024 because data was noisy and bigger batch size gave better performance.\n#LGBM#\nWe decided to use LGBM to ensemble our models because it is fast and usually gives better results than other methods. We used same idea as [Pavel Ostyakov][4] described in his [5th place Cdiscount solution][5]. We predicted top10 classes by our best network, concatenated probabilities of other networks, added class_id feature and gave binary label: whether the class_id is correct or not. This method resulted in validation score 0.001 higher than a simple average. So, if you have enough time, always try LGBM to ensemble.\nFinal submit with LGBM -&gt; 0.948 LB\n#What we tried and didn’t work#\n\n- CatBoost, RF, XGBoost, ensemble of 3rd level\n- Tuning models on clean data. We predicted the train dataset with our best models and dropped out pictures with small probability for the correct class (about 1M samples). It gave small boost on validation, but we didn’t have enough time, so we only trained a couple of models for about 2-3 epochs. So, I suppose, it is also a good idea for noisy data.\n\n#What we didn’t try but it worked#\n\n1. We didn’t noticed that the test is balanced (we could use same technique as in [Camera Model Identification][6] and get gold) Others say, it give plus ~0.7% score to any submission.\n2. TTA with deleting 20% strokes (people on forum said it also improved score). I tried just light augmentations like flips and shift_scale_rotate but they made score on validation even worse.\n\n\n  [1]: https://www.kaggle.com/scitator\n  [2]: https://github.com/Scitator/catalyst\n  [3]: https://www.kaggle.com/artyomp\n  [4]: https://www.kaggle.com/pavelost\n  [5]: https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/45733\n  [6]: https://www.kaggle.com/c/sp-society-camera-model-identification",
      "votes": 18
    },
    {
      "id": 435080,
      "postDate": "2018-12-07T13:19:20.077Z",
      "content": "<p>thanks for sharing. </p>\n\n<p>for CosineAnnealingLR, I wonder what parameters you use? I tried to use but I can't decide what to use for number of cycles, and iterations. </p>",
      "rawMarkdown": "thanks for sharing. \n\nfor CosineAnnealingLR, I wonder what parameters you use? I tried to use but I can't decide what to use for number of cycles, and iterations. ",
      "votes": 1,
      "replies": [
        {
          "id": 437233,
          "postDate": "2018-12-11T15:29:54.807Z",
          "content": "<p>Hi! I have used SGD optimizer, starting from 0.1 LR. Scheduler params were: T_max=7, eta_min=0.001. I divided the whole dataset on 70 epochs for annealing stage.</p>",
          "rawMarkdown": "Hi! I have used SGD optimizer, starting from 0.1 LR. Scheduler params were: T_max=7, eta_min=0.001. I divided the whole dataset on 70 epochs for annealing stage."
        },
        {
          "id": 437425,
          "postDate": "2018-12-11T22:24:30.270Z",
          "content": "<p>Thanks for sharing. how do you use etamin in your implementation? do you use min(etamin, eta)? </p>\n\n<p>also, min LR 0.001 seems really big.. </p>",
          "rawMarkdown": "Thanks for sharing. how do you use etamin in your implementation? do you use min(etamin, eta)? \n\nalso, min LR 0.001 seems really big.. \n\n"
        }
      ]
    },
    {
      "id": 434283,
      "postDate": "2018-12-06T06:44:34.363Z",
      "content": "<p>Hi, thanks for your sharing. Would you please explain your LGBM method in detail? How did you get training data for \n LGBM? How did you do in the inference stage? Thanks. </p>",
      "rawMarkdown": "Hi, thanks for your sharing. Would you please explain your LGBM method in detail? How did you get training data for \n LGBM? How did you do in the inference stage? Thanks. ",
      "votes": 1,
      "replies": [
        {
          "id": 437238,
          "postDate": "2018-12-11T15:37:28.493Z",
          "content": "<p>Hi! Sorry for such a long answer. At first, I took my best network and predicted validation with it. Then, I took only top10 classes by probability of prediction. I concatenated predictions of other networks and added <strong>class_id</strong> as a feature. Target feature was <strong>is it true, that target_id is class_id</strong>. So I got binary classification problem and for every sample from validation and test I got 10 samples. In the inference stage I just took top3 probabilities of predictions of my LGBM.</p>",
          "rawMarkdown": "Hi! Sorry for such a long answer. At first, I took my best network and predicted validation with it. Then, I took only top10 classes by probability of prediction. I concatenated predictions of other networks and added **class_id** as a feature. Target feature was **is it true, that target_id is class_id**. So I got binary classification problem and for every sample from validation and test I got 10 samples. In the inference stage I just took top3 probabilities of predictions of my LGBM.",
          "votes": 1
        }
      ]
    },
    {
      "id": 434344,
      "postDate": "2018-12-06T08:30:01.230Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 435080,
      "author_name": "yimacs",
      "author_url": "",
      "post_date": "2018-12-07T13:19:20.077000",
      "content": "<p>thanks for sharing. </p>\n\n<p>for CosineAnnealingLR, I wonder what parameters you use? I tried to use but I can't decide what to use for number of cycles, and iterations. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 437233,
          "author_name": "Pavel Tsai",
          "author_url": "",
          "post_date": "2018-12-11T15:29:54.807000",
          "content": "<p>Hi! I have used SGD optimizer, starting from 0.1 LR. Scheduler params were: T_max=7, eta_min=0.001. I divided the whole dataset on 70 epochs for annealing stage.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 437425,
          "author_name": "yimacs",
          "author_url": "",
          "post_date": "2018-12-11T22:24:30.270000",
          "content": "<p>Thanks for sharing. how do you use etamin in your implementation? do you use min(etamin, eta)? </p>\n\n<p>also, min LR 0.001 seems really big.. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 434283,
      "author_name": "good good study",
      "author_url": "",
      "post_date": "2018-12-06T06:44:34.363000",
      "content": "<p>Hi, thanks for your sharing. Would you please explain your LGBM method in detail? How did you get training data for \n LGBM? How did you do in the inference stage? Thanks. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 437238,
          "author_name": "Pavel Tsai",
          "author_url": "",
          "post_date": "2018-12-11T15:37:28.493000",
          "content": "<p>Hi! Sorry for such a long answer. At first, I took my best network and predicted validation with it. Then, I took only top10 classes by probability of prediction. I concatenated predictions of other networks and added <strong>class_id</strong> as a feature. Target feature was <strong>is it true, that target_id is class_id</strong>. So I got binary classification problem and for every sample from validation and test I got 10 samples. In the inference stage I just took top3 probabilities of predictions of my LGBM.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 434344,
      "author_name": "Tommy Jiang",
      "author_url": "",
      "post_date": "2018-12-06T08:30:01.230000",
      "content": "<p>Thanks for sharing!</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "433716": "Thanks everybody for the competition! The task was really interesting and had huge amounts of data. Thanks very much to my teammate who writes really perfect code on Python. You should understand me, he writes **comments** to functions and uses **types** in Python! Also thanks to [@scitator][1] for his perfect training ML framework [Catalyst][2], all models were trained with the help of it.\n \nNow I’d like to tell you about my and [Artyom Palvelev][3] solution using simplified data only.\n#data preprocessing#\nPictures of size 128x128 gave the best combination of score and training time. I don’t have time information, so I encoded the following data in three channels:\n\n 1. The index of line (linearly from 10 for the first line to 255 for the last one)\n 2. The number of strokes in the line\n 3. Just constant 255\n#Our models#\nFirstly, I tried models from forum like MobileNetV2, but they showed poor performance, so I trained something deeper. My first good model was SE_ResNext50 0.942 LB. I used CosineAnnealingLR and averaged top4 checkpoints.\nThen we merged with Artyom and tried different model architectures. Blend with his models gave 0.944 LB. In the final submission we had:\n\n1. SE_ResNext50 (~0.942LB)\n2. SE_ResNext101 (~0.944LB)\n3. NASNet-A-Large (~0.944LB)\n4. SENet154 (~0.945LB)\n5. CBAM_ResNet50 (~0.941LB)\n\nWe also used *gradient accumulation* to increase batch size to 1024 because data was noisy and bigger batch size gave better performance.\n#LGBM#\nWe decided to use LGBM to ensemble our models because it is fast and usually gives better results than other methods. We used same idea as [Pavel Ostyakov][4] described in his [5th place Cdiscount solution][5]. We predicted top10 classes by our best network, concatenated probabilities of other networks, added class_id feature and gave binary label: whether the class_id is correct or not. This method resulted in validation score 0.001 higher than a simple average. So, if you have enough time, always try LGBM to ensemble.\nFinal submit with LGBM -&gt; 0.948 LB\n#What we tried and didn’t work#\n\n- CatBoost, RF, XGBoost, ensemble of 3rd level\n- Tuning models on clean data. We predicted the train dataset with our best models and dropped out pictures with small probability for the correct class (about 1M samples). It gave small boost on validation, but we didn’t have enough time, so we only trained a couple of models for about 2-3 epochs. So, I suppose, it is also a good idea for noisy data.\n\n#What we didn’t try but it worked#\n\n1. We didn’t noticed that the test is balanced (we could use same technique as in [Camera Model Identification][6] and get gold) Others say, it give plus ~0.7% score to any submission.\n2. TTA with deleting 20% strokes (people on forum said it also improved score). I tried just light augmentations like flips and shift_scale_rotate but they made score on validation even worse.\n\n\n  [1]: https://www.kaggle.com/scitator\n  [2]: https://github.com/Scitator/catalyst\n  [3]: https://www.kaggle.com/artyomp\n  [4]: https://www.kaggle.com/pavelost\n  [5]: https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/45733\n  [6]: https://www.kaggle.com/c/sp-society-camera-model-identification",
    "435080": "thanks for sharing. \n\nfor CosineAnnealingLR, I wonder what parameters you use? I tried to use but I can't decide what to use for number of cycles, and iterations. ",
    "434283": "Hi, thanks for your sharing. Would you please explain your LGBM method in detail? How did you get training data for \n LGBM? How did you do in the inference stage? Thanks. ",
    "434344": "Thanks for sharing!"
  }
}