{
  "id": 73708,
  "title": "5th place solution",
  "url": "/competitions/quickdraw-doodle-recognition/discussion/73708",
  "author_name": "Artur Ispiriants",
  "post_date": "2018-12-05T01:53:59.095000",
  "votes": 52,
  "comment_count": 33,
  "views": 0,
  "content": "<p>First of all, congrats everyone with the end of the competition. It was an exciting experience, and we want to share our approach.</p>\n\n<p><strong>Handling the data</strong></p>\n\n<p>Both simplified and raw data were used. To be able to read any random image, all strokes from CSV files were separated to one image per binary file. It took about 400GB on SSD, but it allowed to start different experiments very fast.</p>\n\n<p><strong>Our models</strong></p>\n\n<p>In total we trained 3 main models:</p>\n\n<ol>\n<li>Se-Resnext50</li>\n<li>DPN-92</li>\n<li>Se-Resnext101</li>\n</ol>\n\n<p>All of them were pretrained on imagenet. We used different image sizes, 128 -&gt; 192 -&gt; 224 -&gt; 256. It was clear almost from the beginning - the bigger image size -  the bigger score in both local validation and public lb. It was hard to train 256px due to limited GPU resources. Using fit predict and 128px image it was pretty straightforward to get 0.944 public LB.</p>\n\n<p>In the middle of the competition after merging with @firenero, we had ~ 0.948 public lb score. To move forward, it was essential to use time information which only exists in the full dataset. We encoded each stroke using 3 channels:</p>\n\n<ul>\n<li>Delay value scaled to 0-255.</li>\n<li>Draw time per stroke scaled to 0-255</li>\n<li>Number of strokes scaled to 0-255</li>\n</ul>\n\n<p>It gave a significant boost in local validation and gave us ~0.951 public LB.\nAnother important thing is batch size; we tuned all our models with huge batch size increasing it with each snapshot up to 10K.</p>\n\n<p>@firenero is going to tell more about the final phase and how we achieved 0.953.</p>",
  "messages": [
    {
      "id": 433353,
      "postDate": "2018-12-05T01:53:59.097Z",
      "content": "<p>First of all, congrats everyone with the end of the competition. It was an exciting experience, and we want to share our approach.</p>\n\n<p><strong>Handling the data</strong></p>\n\n<p>Both simplified and raw data were used. To be able to read any random image, all strokes from CSV files were separated to one image per binary file. It took about 400GB on SSD, but it allowed to start different experiments very fast.</p>\n\n<p><strong>Our models</strong></p>\n\n<p>In total we trained 3 main models:</p>\n\n<ol>\n<li>Se-Resnext50</li>\n<li>DPN-92</li>\n<li>Se-Resnext101</li>\n</ol>\n\n<p>All of them were pretrained on imagenet. We used different image sizes, 128 -&gt; 192 -&gt; 224 -&gt; 256. It was clear almost from the beginning - the bigger image size -  the bigger score in both local validation and public lb. It was hard to train 256px due to limited GPU resources. Using fit predict and 128px image it was pretty straightforward to get 0.944 public LB.</p>\n\n<p>In the middle of the competition after merging with @firenero, we had ~ 0.948 public lb score. To move forward, it was essential to use time information which only exists in the full dataset. We encoded each stroke using 3 channels:</p>\n\n<ul>\n<li>Delay value scaled to 0-255.</li>\n<li>Draw time per stroke scaled to 0-255</li>\n<li>Number of strokes scaled to 0-255</li>\n</ul>\n\n<p>It gave a significant boost in local validation and gave us ~0.951 public LB.\nAnother important thing is batch size; we tuned all our models with huge batch size increasing it with each snapshot up to 10K.</p>\n\n<p>@firenero is going to tell more about the final phase and how we achieved 0.953.</p>",
      "rawMarkdown": "First of all, congrats everyone with the end of the competition. It was an exciting experience, and we want to share our approach.\n\n**Handling the data**\n\nBoth simplified and raw data were used. To be able to read any random image, all strokes from CSV files were separated to one image per binary file. It took about 400GB on SSD, but it allowed to start different experiments very fast.\n\n**Our models**\n\nIn total we trained 3 main models:\n\n 1. Se-Resnext50\n 2. DPN-92\n 3. Se-Resnext101\n\nAll of them were pretrained on imagenet. We used different image sizes, 128 -&gt; 192 -&gt; 224 -&gt; 256. It was clear almost from the beginning - the bigger image size -  the bigger score in both local validation and public lb. It was hard to train 256px due to limited GPU resources. Using fit predict and 128px image it was pretty straightforward to get 0.944 public LB.\n\nIn the middle of the competition after merging with @firenero, we had ~ 0.948 public lb score. To move forward, it was essential to use time information which only exists in the full dataset. We encoded each stroke using 3 channels:\n\n - Delay value scaled to 0-255.\n - Draw time per stroke scaled to 0-255\n - Number of strokes scaled to 0-255\n\nIt gave a significant boost in local validation and gave us ~0.951 public LB.\nAnother important thing is batch size; we tuned all our models with huge batch size increasing it with each snapshot up to 10K.\n\n@firenero is going to tell more about the final phase and how we achieved 0.953.",
      "votes": 52
    },
    {
      "id": 433354,
      "postDate": "2018-12-05T01:57:42.400Z",
      "content": "<p>Congrats everyone with good results and ending of this exciting competition!</p>\n\n<p>Early in competition becomes clear that deeper models and bigger images lead to better results, so I've trained SE-ResNext101 that gave 0.947 public LB before merging with <a href=\"/aispiriants\">@aispiriants</a>. \nFirst of all, I've never trained full epoch at once and divided full train set into smaller groups to be able to save models more often. Those checkpoints' size varied depending on train image size, batch size, available time. In general, I tried to save one checkpoint every 4-6 hours of training.</p>\n\n<p>Training flow:</p>\n\n<ol>\n<li>Train on 128x128 images until net converges. It took around 20 checkpoints. It was enough to score around 0.945 on public LB.</li>\n<li>After training on 128x128 stops giving any improvements, I used aggregated batches technique to use batch size around 3-4K images which gave a boost on local validation.</li>\n<li>Further improvements gained from the fine-tuning model on bigger images (192x192 and then 256x256)</li>\n</ol>\n\n<p><strong>Some more tricks we used in our solution:</strong></p>\n\n<p><strong>Validation</strong>\nAfter merging we stratified all train images into buckets by a country/class that gave us many buckets. For validation, we took 25 images per each bucket that give us a pretty solid validation score and allowed us to blend our models easily.</p>\n\n<p><strong>Ensembling</strong>\nSaving checkpoints often gives a free boost by using snapshot ensembling. Final solution contains several snapshots of each model. All models were blended by the simple mean.</p>\n\n<p><strong>Pseudo-labeling</strong>\nThat's the main score booster in the later steps of competition. Using blend that scored 0.951 on public LB we generated predicts for the test dataset and used them to fine-tune all of our models around 10 epochs using only pseudo-labeled data. It gave us around 0.952.\nThen using 0.952 blend, we generated pseudo-labels again and fine-tune again :) Which gives us 0.953 public LB. For some reason, we stopped doing this anymore, but probably we could get even higher score just re-pseudo-labeling :)</p>",
      "rawMarkdown": "Congrats everyone with good results and ending of this exciting competition!\n\nEarly in competition becomes clear that deeper models and bigger images lead to better results, so I've trained SE-ResNext101 that gave 0.947 public LB before merging with @aispiriants. \nFirst of all, I've never trained full epoch at once and divided full train set into smaller groups to be able to save models more often. Those checkpoints' size varied depending on train image size, batch size, available time. In general, I tried to save one checkpoint every 4-6 hours of training.\n\nTraining flow:\n\n 1. Train on 128x128 images until net converges. It took around 20 checkpoints. It was enough to score around 0.945 on public LB.\n 2. After training on 128x128 stops giving any improvements, I used aggregated batches technique to use batch size around 3-4K images which gave a boost on local validation.\n 3. Further improvements gained from the fine-tuning model on bigger images (192x192 and then 256x256)\n \n**Some more tricks we used in our solution:**\n\n**Validation**\nAfter merging we stratified all train images into buckets by a country/class that gave us many buckets. For validation, we took 25 images per each bucket that give us a pretty solid validation score and allowed us to blend our models easily.\n\n**Ensembling**\nSaving checkpoints often gives a free boost by using snapshot ensembling. Final solution contains several snapshots of each model. All models were blended by the simple mean.\n\n**Pseudo-labeling**\nThat's the main score booster in the later steps of competition. Using blend that scored 0.951 on public LB we generated predicts for the test dataset and used them to fine-tune all of our models around 10 epochs using only pseudo-labeled data. It gave us around 0.952.\nThen using 0.952 blend, we generated pseudo-labels again and fine-tune again :) Which gives us 0.953 public LB. For some reason, we stopped doing this anymore, but probably we could get even higher score just re-pseudo-labeling :)",
      "votes": 11,
      "replies": [
        {
          "id": 433360,
          "postDate": "2018-12-05T02:07:22.307Z",
          "content": "<p>do you try pseudo label the unrecognised train images as well?</p>",
          "rawMarkdown": "do you try pseudo label the unrecognised train images as well?"
        },
        {
          "id": 433365,
          "postDate": "2018-12-05T02:11:50.470Z",
          "content": "<p>No, just test. We trained on all unrecognized images as is. \nBut that sounds like a good idea to correct unrecognized ones with pseudo labels, and I would like to try it if I thought about it earlier.</p>",
          "rawMarkdown": "No, just test. We trained on all unrecognized images as is. \nBut that sounds like a good idea to correct unrecognized ones with pseudo labels, and I would like to try it if I thought about it earlier.",
          "votes": 2
        },
        {
          "id": 433373,
          "postDate": "2018-12-05T02:26:51.730Z",
          "content": "<p>Could you tell me how to fine-tune your model after pseudo-labeling test data ? Is there any risk of overfiting?</p>",
          "rawMarkdown": "Could you tell me how to fine-tune your model after pseudo-labeling test data ? Is there any risk of overfiting?",
          "votes": 1
        },
        {
          "id": 433401,
          "postDate": "2018-12-05T03:11:39Z",
          "content": "<p>In this competition, we just load model and then trained it only on pseudo-labeled data 10 epochs (chose randomly) and selected the best 3 of them for a blend. And yes, it's a quite dangerous technique in terms of overfitting so you need to be extra careful with it. \nHowever, there are many more approaches to apply pseudo-labeling like mixing them into train dataset or train completely new model on pseudo-labeled images and then fine-tune on regular train dataset. </p>",
          "rawMarkdown": "In this competition, we just load model and then trained it only on pseudo-labeled data 10 epochs (chose randomly) and selected the best 3 of them for a blend. And yes, it's a quite dangerous technique in terms of overfitting so you need to be extra careful with it. \nHowever, there are many more approaches to apply pseudo-labeling like mixing them into train dataset or train completely new model on pseudo-labeled images and then fine-tune on regular train dataset. ",
          "votes": 1
        },
        {
          "id": 433420,
          "postDate": "2018-12-05T03:55:50.563Z",
          "content": "<p>Congrats and thx for sharing! For pseudo label, which kind of label do you use, top1 predict or probabilities of all 340 classes? </p>",
          "rawMarkdown": "Congrats and thx for sharing! For pseudo label, which kind of label do you use, top1 predict or probabilities of all 340 classes? "
        },
        {
          "id": 433727,
          "postDate": "2018-12-05T12:07:24.733Z",
          "content": "<p>Just top1 prediction</p>",
          "rawMarkdown": "Just top1 prediction",
          "votes": 1
        },
        {
          "id": 433875,
          "postDate": "2018-12-05T15:31:45.130Z",
          "content": "<p>Is it possible to aggregate batches in Keras? The couldn't get my batch size up that high for SeResNext50, so it really limited my success. </p>",
          "rawMarkdown": "Is it possible to aggregate batches in Keras? The couldn't get my batch size up that high for SeResNext50, so it really limited my success. "
        },
        {
          "id": 434149,
          "postDate": "2018-12-06T00:59:10.540Z",
          "content": "<p>We used Pytorch but someone said in comments that you can find Keras code in this thread: <a href=\"https://www.kaggle.com/c/quickdraw-doodle-recognition/discussion/73701\">https://www.kaggle.com/c/quickdraw-doodle-recognition/discussion/73701</a></p>",
          "rawMarkdown": "We used Pytorch but someone said in comments that you can find Keras code in this thread: https://www.kaggle.com/c/quickdraw-doodle-recognition/discussion/73701",
          "votes": 1
        }
      ]
    },
    {
      "id": 433517,
      "postDate": "2018-12-05T06:36:57.567Z",
      "content": "<p>Congrats and thanks for the sharing, it is great inspiration for future competition for starting out in kaggle </p>",
      "rawMarkdown": "Congrats and thanks for the sharing, it is great inspiration for future competition for starting out in kaggle ",
      "votes": 1
    },
    {
      "id": 433395,
      "postDate": "2018-12-05T02:58:38.617Z",
      "content": "<p>Congratulation and thanks for sharing <a href=\"/aispiriants\">@aispiriants</a> and <a href=\"/firenero\">@firenero</a>!!!</p>\n\n<p>Could you please share further on the following points:\n- Is there any tricks for selecting the optimizer or its schedule learning rate ? (or just use plain Adam / SGD)\n- Is there any useful pointers to help newbies (like myself) to implement batch accumulation ? (my team tried Adam Batch Accumulation but usually failed because the Out-of-memory problem)\n- So you use “all” snapshot models to make “snapshot ensemble” regardless of their accuracies ? (In my case, sometimes kicking-out the weak snapshots helps improve the overall ensemble a little bit)\n- which GPUs you did you use?</p>",
      "rawMarkdown": "Congratulation and thanks for sharing @aispiriants and @firenero!!!\n\nCould you please share further on the following points:\n- Is there any tricks for selecting the optimizer or its schedule learning rate ? (or just use plain Adam / SGD)\n- Is there any useful pointers to help newbies (like myself) to implement batch accumulation ? (my team tried Adam Batch Accumulation but usually failed because the Out-of-memory problem)\n- So you use “all” snapshot models to make “snapshot ensemble” regardless of their accuracies ? (In my case, sometimes kicking-out the weak snapshots helps improve the overall ensemble a little bit)\n- which GPUs you did you use?",
      "votes": 1,
      "replies": [
        {
          "id": 433403,
          "postDate": "2018-12-05T03:20:42.147Z",
          "content": "<ul>\n<li>We just used Adam as it usually the best one and easier to get good results with. Choosing an LR is not an easy task for me too and I would like to hear any suggestions about it too :) For this competition, I did some experiments and came to 0.0001 LR which was reduced to 0.00001 in later steps of training.</li>\n<li>We've just found some explanation on pytorch forum and implemented it. I don't have a link anymore but maybe @aispiriats has.</li>\n<li>We chose \"best\" snapshots too as weak snapshots were making score worse. Best snapshots were found by manual replacing models and comparing val score.</li>\n<li>Each of us has 2 1080ti</li>\n</ul>",
          "rawMarkdown": " - We just used Adam as it usually the best one and easier to get good results with. Choosing an LR is not an easy task for me too and I would like to hear any suggestions about it too :) For this competition, I did some experiments and came to 0.0001 LR which was reduced to 0.00001 in later steps of training.\n - We've just found some explanation on pytorch forum and implemented it. I don't have a link anymore but maybe @aispiriats has.\n - We chose \"best\" snapshots too as weak snapshots were making score worse. Best snapshots were found by manual replacing models and comparing val score.\n - Each of us has 2 1080ti",
          "votes": 2
        },
        {
          "id": 433444,
          "postDate": "2018-12-05T04:57:01.223Z",
          "content": "<p>HI, @The Neuron Engineer，if you use pytorch, you can  refer to <a href=\"https://discuss.pytorch.org/t/how-to-implement-accumulated-gradient/3822\">here</a> for batch accumulation. </p>",
          "rawMarkdown": "HI, @The Neuron Engineer，if you use pytorch, you can  refer to [here](https://discuss.pytorch.org/t/how-to-implement-accumulated-gradient/3822) for batch accumulation. ",
          "votes": 2
        },
        {
          "id": 433463,
          "postDate": "2018-12-05T05:23:59.387Z",
          "content": "<p>Thanks <a href=\"/firenero\">@firenero</a> for your kind answer! and thanks @upup for your suggestion too! (we use Keras, but we will try pytorch in the future)\nIn our case, to finetune LR, it seems that the lr_finder trick of fast.ai is quite practical to us. From your explanation, the data preparation process is super important; we will make sure to do better next time.</p>\n\n<p>— EDIT — we found the Keras Batch Accumulation code kindly shared in the “24th solution” discussion.</p>",
          "rawMarkdown": "Thanks @firenero for your kind answer! and thanks @upup for your suggestion too! (we use Keras, but we will try pytorch in the future)\nIn our case, to finetune LR, it seems that the lr_finder trick of fast.ai is quite practical to us. From your explanation, the data preparation process is super important; we will make sure to do better next time.\n\n— EDIT — we found the Keras Batch Accumulation code kindly shared in the “24th solution” discussion.",
          "votes": 1
        }
      ]
    },
    {
      "id": 433584,
      "postDate": "2018-12-05T07:49:21.773Z",
      "content": "<p>Congrats and thanks for sharing !\nI have tried similar encoding with you but fail to get a decent performance gain :( \nI want to ask that what do you mean by using <strong>number of strokes scaled to 0-255 ?</strong> Is that means the sequence number of the stroke ?</p>",
      "rawMarkdown": "Congrats and thanks for sharing !\nI have tried similar encoding with you but fail to get a decent performance gain :( \nI want to ask that what do you mean by using **number of strokes scaled to 0-255 ?** Is that means the sequence number of the stroke ?",
      "votes": 2,
      "replies": [
        {
          "id": 433708,
          "postDate": "2018-12-05T11:35:36.373Z",
          "content": "<p>Yes, sequence number</p>",
          "rawMarkdown": "Yes, sequence number",
          "votes": 2
        },
        {
          "id": 433758,
          "postDate": "2018-12-05T12:52:16.550Z",
          "content": "<p>Thanks for your reply ! I have another question \nDid you use weight decay ? I removed the weight-decay to get faster convergence since the dataset size is huge, however my se-resnxt-50 only get ~ 0.940 LB. Wondering if it is the problem. </p>",
          "rawMarkdown": "Thanks for your reply ! I have another question \nDid you use weight decay ? I removed the weight-decay to get faster convergence since the dataset size is huge, however my se-resnxt-50 only get ~ 0.940 LB. Wondering if it is the problem. \n"
        }
      ]
    },
    {
      "id": 433682,
      "postDate": "2018-12-05T10:44:37.820Z",
      "content": "<p>Thanks for sharing, Artur! Could you elaborate on how you encode time information into 3 channels?</p>",
      "rawMarkdown": "Thanks for sharing, Artur! Could you elaborate on how you encode time information into 3 channels?"
    },
    {
      "id": 433625,
      "postDate": "2018-12-05T08:49:11.020Z",
      "content": "<p>Is there a good working Se-Resnext implementation for keras?</p>",
      "rawMarkdown": "Is there a good working Se-Resnext implementation for keras?",
      "replies": [
        {
          "id": 433631,
          "postDate": "2018-12-05T09:07:01.163Z",
          "content": "<p>Yes, see this link:\n<a href=\"https://github.com/titu1994/keras-squeeze-excite-network\">https://github.com/titu1994/keras-squeeze-excite-network</a></p>",
          "rawMarkdown": "Yes, see this link:\nhttps://github.com/titu1994/keras-squeeze-excite-network",
          "votes": 1
        },
        {
          "id": 433679,
          "postDate": "2018-12-05T10:40:54.763Z",
          "content": "<p>I tried this SE-ResNeXt implementation, however, it only scored 0.727, I used 128x128 image size, trained on full dataset.</p>",
          "rawMarkdown": "I tried this SE-ResNeXt implementation, however, it only scored 0.727, I used 128x128 image size, trained on full dataset.",
          "votes": 2
        },
        {
          "id": 433706,
          "postDate": "2018-12-05T11:34:42.267Z",
          "content": "<p><a href=\"/soulmachine\">@soulmachine</a> same here ! &gt;__&lt;</p>",
          "rawMarkdown": "@soulmachine same here ! &gt;__&lt;"
        },
        {
          "id": 434310,
          "postDate": "2018-12-06T07:23:18.743Z",
          "content": "<p>Maybe your batch size is not large enough?</p>",
          "rawMarkdown": "Maybe your batch size is not large enough?"
        }
      ]
    },
    {
      "id": 433552,
      "postDate": "2018-12-05T07:04:50.433Z",
      "content": "<p>Hi <a href=\"/aispiriants\">@aispiriants</a>, could you please clarify this sentence : </p>\n\n<p>“To be able to read any random image, all strokes from CSV files were separated to one image per binary file.”</p>\n\n<p>One stroke per file?</p>\n\n<p>Many thanks again!</p>",
      "rawMarkdown": "Hi @aispiriants, could you please clarify this sentence : \n\n“To be able to read any random image, all strokes from CSV files were separated to one image per binary file.”\n\nOne stroke per file?\n\nMany thanks again!",
      "replies": [
        {
          "id": 433712,
          "postDate": "2018-12-05T11:38:41.133Z",
          "content": "<p>No, one image contains multiple strokes. So one file contains an array of strokes for one image. It simplifies the reading process, because you can shuffle ids and then by id it is easy to load any random image. It also solved memory problem, because it was pretty hard to fit all strokes in memory.</p>",
          "rawMarkdown": "No, one image contains multiple strokes. So one file contains an array of strokes for one image. It simplifies the reading process, because you can shuffle ids and then by id it is easy to load any random image. It also solved memory problem, because it was pretty hard to fit all strokes in memory.",
          "votes": 2
        },
        {
          "id": 434148,
          "postDate": "2018-12-06T00:57:31.887Z",
          "content": "<p>thanks so much!</p>",
          "rawMarkdown": "thanks so much!"
        },
        {
          "id": 439321,
          "postDate": "2018-12-15T06:29:14.347Z",
          "content": "<p>So you simply use the <code>key_id</code> as the file name, and each file stores the strokes in binary file, which is a 3D list of integers. Do I understand it correct?</p>",
          "rawMarkdown": "So you simply use the `key_id` as the file name, and each file stores the strokes in binary file, which is a 3D list of integers. Do I understand it correct?"
        }
      ]
    },
    {
      "id": 433369,
      "postDate": "2018-12-05T02:18:03.867Z",
      "content": "<p>I would have used the SE-ResNet, but I found that the structure of se-resnet is strange. For example: from 64*64, then change it to 1*1, then expand to 16*16, then 1*1 , etc. \nWhat's the advantages of this kind of structure? \nThx!</p>",
      "rawMarkdown": "I would have used the SE-ResNet, but I found that the structure of se-resnet is strange. For example: from 64*64, then change it to 1*1, then expand to 16*16, then 1*1 , etc. \nWhat's the advantages of this kind of structure? \nThx!"
    },
    {
      "id": 439308,
      "postDate": "2018-12-15T06:14:36.423Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 433734,
      "postDate": "2018-12-05T12:31:06.883Z",
      "rawMarkdown": "",
      "votes": 8,
      "isDeleted": true
    },
    {
      "id": 434395,
      "postDate": "2018-12-06T10:11:45.040Z",
      "content": "<p>Great job, thanks for sharing !</p>",
      "rawMarkdown": "Great job, thanks for sharing !"
    },
    {
      "id": 434326,
      "postDate": "2018-12-06T07:49:38.273Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!"
    },
    {
      "id": 433562,
      "postDate": "2018-12-05T07:23:49.163Z",
      "content": "<p>Congrats and thanks for sharing.</p>",
      "rawMarkdown": "Congrats and thanks for sharing."
    }
  ],
  "comments": [
    {
      "id": 433354,
      "author_name": "Mykhailo Matviiv",
      "author_url": "",
      "post_date": "2018-12-05T01:57:42.400000",
      "content": "<p>Congrats everyone with good results and ending of this exciting competition!</p>\n\n<p>Early in competition becomes clear that deeper models and bigger images lead to better results, so I've trained SE-ResNext101 that gave 0.947 public LB before merging with <a href=\"/aispiriants\">@aispiriants</a>. \nFirst of all, I've never trained full epoch at once and divided full train set into smaller groups to be able to save models more often. Those checkpoints' size varied depending on train image size, batch size, available time. In general, I tried to save one checkpoint every 4-6 hours of training.</p>\n\n<p>Training flow:</p>\n\n<ol>\n<li>Train on 128x128 images until net converges. It took around 20 checkpoints. It was enough to score around 0.945 on public LB.</li>\n<li>After training on 128x128 stops giving any improvements, I used aggregated batches technique to use batch size around 3-4K images which gave a boost on local validation.</li>\n<li>Further improvements gained from the fine-tuning model on bigger images (192x192 and then 256x256)</li>\n</ol>\n\n<p><strong>Some more tricks we used in our solution:</strong></p>\n\n<p><strong>Validation</strong>\nAfter merging we stratified all train images into buckets by a country/class that gave us many buckets. For validation, we took 25 images per each bucket that give us a pretty solid validation score and allowed us to blend our models easily.</p>\n\n<p><strong>Ensembling</strong>\nSaving checkpoints often gives a free boost by using snapshot ensembling. Final solution contains several snapshots of each model. All models were blended by the simple mean.</p>\n\n<p><strong>Pseudo-labeling</strong>\nThat's the main score booster in the later steps of competition. Using blend that scored 0.951 on public LB we generated predicts for the test dataset and used them to fine-tune all of our models around 10 epochs using only pseudo-labeled data. It gave us around 0.952.\nThen using 0.952 blend, we generated pseudo-labels again and fine-tune again :) Which gives us 0.953 public LB. For some reason, we stopped doing this anymore, but probably we could get even higher score just re-pseudo-labeling :)</p>",
      "votes": 11,
      "replies": [
        {
          "id": 433360,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2018-12-05T02:07:22.307000",
          "content": "<p>do you try pseudo label the unrecognised train images as well?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 433365,
          "author_name": "Mykhailo Matviiv",
          "author_url": "",
          "post_date": "2018-12-05T02:11:50.470000",
          "content": "<p>No, just test. We trained on all unrecognized images as is. \nBut that sounds like a good idea to correct unrecognized ones with pseudo labels, and I would like to try it if I thought about it earlier.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 433373,
          "author_name": "upup",
          "author_url": "",
          "post_date": "2018-12-05T02:26:51.730000",
          "content": "<p>Could you tell me how to fine-tune your model after pseudo-labeling test data ? Is there any risk of overfiting?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 433401,
          "author_name": "Mykhailo Matviiv",
          "author_url": "",
          "post_date": "2018-12-05T03:11:39",
          "content": "<p>In this competition, we just load model and then trained it only on pseudo-labeled data 10 epochs (chose randomly) and selected the best 3 of them for a blend. And yes, it's a quite dangerous technique in terms of overfitting so you need to be extra careful with it. \nHowever, there are many more approaches to apply pseudo-labeling like mixing them into train dataset or train completely new model on pseudo-labeled images and then fine-tune on regular train dataset. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 433420,
          "author_name": "dingwoai",
          "author_url": "",
          "post_date": "2018-12-05T03:55:50.563000",
          "content": "<p>Congrats and thx for sharing! For pseudo label, which kind of label do you use, top1 predict or probabilities of all 340 classes? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 433727,
          "author_name": "Mykhailo Matviiv",
          "author_url": "",
          "post_date": "2018-12-05T12:07:24.733000",
          "content": "<p>Just top1 prediction</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 433875,
          "author_name": "DanBerman",
          "author_url": "",
          "post_date": "2018-12-05T15:31:45.130000",
          "content": "<p>Is it possible to aggregate batches in Keras? The couldn't get my batch size up that high for SeResNext50, so it really limited my success. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 434149,
          "author_name": "Mykhailo Matviiv",
          "author_url": "",
          "post_date": "2018-12-06T00:59:10.540000",
          "content": "<p>We used Pytorch but someone said in comments that you can find Keras code in this thread: <a href=\"https://www.kaggle.com/c/quickdraw-doodle-recognition/discussion/73701\">https://www.kaggle.com/c/quickdraw-doodle-recognition/discussion/73701</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 433517,
      "author_name": "Gal Peled",
      "author_url": "",
      "post_date": "2018-12-05T06:36:57.567000",
      "content": "<p>Congrats and thanks for the sharing, it is great inspiration for future competition for starting out in kaggle </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 433395,
      "author_name": "Neuron Engineer",
      "author_url": "",
      "post_date": "2018-12-05T02:58:38.617000",
      "content": "<p>Congratulation and thanks for sharing <a href=\"/aispiriants\">@aispiriants</a> and <a href=\"/firenero\">@firenero</a>!!!</p>\n\n<p>Could you please share further on the following points:\n- Is there any tricks for selecting the optimizer or its schedule learning rate ? (or just use plain Adam / SGD)\n- Is there any useful pointers to help newbies (like myself) to implement batch accumulation ? (my team tried Adam Batch Accumulation but usually failed because the Out-of-memory problem)\n- So you use “all” snapshot models to make “snapshot ensemble” regardless of their accuracies ? (In my case, sometimes kicking-out the weak snapshots helps improve the overall ensemble a little bit)\n- which GPUs you did you use?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 433403,
          "author_name": "Mykhailo Matviiv",
          "author_url": "",
          "post_date": "2018-12-05T03:20:42.147000",
          "content": "<ul>\n<li>We just used Adam as it usually the best one and easier to get good results with. Choosing an LR is not an easy task for me too and I would like to hear any suggestions about it too :) For this competition, I did some experiments and came to 0.0001 LR which was reduced to 0.00001 in later steps of training.</li>\n<li>We've just found some explanation on pytorch forum and implemented it. I don't have a link anymore but maybe @aispiriats has.</li>\n<li>We chose \"best\" snapshots too as weak snapshots were making score worse. Best snapshots were found by manual replacing models and comparing val score.</li>\n<li>Each of us has 2 1080ti</li>\n</ul>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 433444,
          "author_name": "upup",
          "author_url": "",
          "post_date": "2018-12-05T04:57:01.223000",
          "content": "<p>HI, @The Neuron Engineer，if you use pytorch, you can  refer to <a href=\"https://discuss.pytorch.org/t/how-to-implement-accumulated-gradient/3822\">here</a> for batch accumulation. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 433463,
          "author_name": "Neuron Engineer",
          "author_url": "",
          "post_date": "2018-12-05T05:23:59.387000",
          "content": "<p>Thanks <a href=\"/firenero\">@firenero</a> for your kind answer! and thanks @upup for your suggestion too! (we use Keras, but we will try pytorch in the future)\nIn our case, to finetune LR, it seems that the lr_finder trick of fast.ai is quite practical to us. From your explanation, the data preparation process is super important; we will make sure to do better next time.</p>\n\n<p>— EDIT — we found the Keras Batch Accumulation code kindly shared in the “24th solution” discussion.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 433584,
      "author_name": "tkuanlun350",
      "author_url": "",
      "post_date": "2018-12-05T07:49:21.773000",
      "content": "<p>Congrats and thanks for sharing !\nI have tried similar encoding with you but fail to get a decent performance gain :( \nI want to ask that what do you mean by using <strong>number of strokes scaled to 0-255 ?</strong> Is that means the sequence number of the stroke ?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 433708,
          "author_name": "Artur Ispiriants",
          "author_url": "",
          "post_date": "2018-12-05T11:35:36.373000",
          "content": "<p>Yes, sequence number</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 433758,
          "author_name": "tkuanlun350",
          "author_url": "",
          "post_date": "2018-12-05T12:52:16.550000",
          "content": "<p>Thanks for your reply ! I have another question \nDid you use weight decay ? I removed the weight-decay to get faster convergence since the dataset size is huge, however my se-resnxt-50 only get ~ 0.940 LB. Wondering if it is the problem. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 433682,
      "author_name": "[he.ai]soulmachine",
      "author_url": "",
      "post_date": "2018-12-05T10:44:37.820000",
      "content": "<p>Thanks for sharing, Artur! Could you elaborate on how you encode time information into 3 channels?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 433625,
      "author_name": "Nickolay Safronov",
      "author_url": "",
      "post_date": "2018-12-05T08:49:11.020000",
      "content": "<p>Is there a good working Se-Resnext implementation for keras?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 433631,
          "author_name": "Dilapsky Lee",
          "author_url": "",
          "post_date": "2018-12-05T09:07:01.163000",
          "content": "<p>Yes, see this link:\n<a href=\"https://github.com/titu1994/keras-squeeze-excite-network\">https://github.com/titu1994/keras-squeeze-excite-network</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 433679,
          "author_name": "[he.ai]soulmachine",
          "author_url": "",
          "post_date": "2018-12-05T10:40:54.763000",
          "content": "<p>I tried this SE-ResNeXt implementation, however, it only scored 0.727, I used 128x128 image size, trained on full dataset.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 433706,
          "author_name": "Neuron Engineer",
          "author_url": "",
          "post_date": "2018-12-05T11:34:42.267000",
          "content": "<p><a href=\"/soulmachine\">@soulmachine</a> same here ! &gt;__&lt;</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 434310,
          "author_name": "Tommy Jiang",
          "author_url": "",
          "post_date": "2018-12-06T07:23:18.743000",
          "content": "<p>Maybe your batch size is not large enough?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 433552,
      "author_name": "Neuron Engineer",
      "author_url": "",
      "post_date": "2018-12-05T07:04:50.433000",
      "content": "<p>Hi <a href=\"/aispiriants\">@aispiriants</a>, could you please clarify this sentence : </p>\n\n<p>“To be able to read any random image, all strokes from CSV files were separated to one image per binary file.”</p>\n\n<p>One stroke per file?</p>\n\n<p>Many thanks again!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 433712,
          "author_name": "Artur Ispiriants",
          "author_url": "",
          "post_date": "2018-12-05T11:38:41.133000",
          "content": "<p>No, one image contains multiple strokes. So one file contains an array of strokes for one image. It simplifies the reading process, because you can shuffle ids and then by id it is easy to load any random image. It also solved memory problem, because it was pretty hard to fit all strokes in memory.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 434148,
          "author_name": "Neuron Engineer",
          "author_url": "",
          "post_date": "2018-12-06T00:57:31.887000",
          "content": "<p>thanks so much!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 439321,
          "author_name": "[he.ai]soulmachine",
          "author_url": "",
          "post_date": "2018-12-15T06:29:14.347000",
          "content": "<p>So you simply use the <code>key_id</code> as the file name, and each file stores the strokes in binary file, which is a 3D list of integers. Do I understand it correct?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 433369,
      "author_name": "Dilapsky Lee",
      "author_url": "",
      "post_date": "2018-12-05T02:18:03.867000",
      "content": "<p>I would have used the SE-ResNet, but I found that the structure of se-resnet is strange. For example: from 64*64, then change it to 1*1, then expand to 16*16, then 1*1 , etc. \nWhat's the advantages of this kind of structure? \nThx!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 439308,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-15T06:14:36.423000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 433734,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-05T12:31:06.883000",
      "content": "",
      "votes": 8,
      "replies": []
    },
    {
      "id": 434395,
      "author_name": "Theo Viel",
      "author_url": "",
      "post_date": "2018-12-06T10:11:45.040000",
      "content": "<p>Great job, thanks for sharing !</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 434326,
      "author_name": "Tommy Jiang",
      "author_url": "",
      "post_date": "2018-12-06T07:49:38.273000",
      "content": "<p>Thanks for sharing!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 433562,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2018-12-05T07:23:49.163000",
      "content": "<p>Congrats and thanks for sharing.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "433353": "First of all, congrats everyone with the end of the competition. It was an exciting experience, and we want to share our approach.\n\n**Handling the data**\n\nBoth simplified and raw data were used. To be able to read any random image, all strokes from CSV files were separated to one image per binary file. It took about 400GB on SSD, but it allowed to start different experiments very fast.\n\n**Our models**\n\nIn total we trained 3 main models:\n\n 1. Se-Resnext50\n 2. DPN-92\n 3. Se-Resnext101\n\nAll of them were pretrained on imagenet. We used different image sizes, 128 -&gt; 192 -&gt; 224 -&gt; 256. It was clear almost from the beginning - the bigger image size -  the bigger score in both local validation and public lb. It was hard to train 256px due to limited GPU resources. Using fit predict and 128px image it was pretty straightforward to get 0.944 public LB.\n\nIn the middle of the competition after merging with @firenero, we had ~ 0.948 public lb score. To move forward, it was essential to use time information which only exists in the full dataset. We encoded each stroke using 3 channels:\n\n - Delay value scaled to 0-255.\n - Draw time per stroke scaled to 0-255\n - Number of strokes scaled to 0-255\n\nIt gave a significant boost in local validation and gave us ~0.951 public LB.\nAnother important thing is batch size; we tuned all our models with huge batch size increasing it with each snapshot up to 10K.\n\n@firenero is going to tell more about the final phase and how we achieved 0.953.",
    "433354": "Congrats everyone with good results and ending of this exciting competition!\n\nEarly in competition becomes clear that deeper models and bigger images lead to better results, so I've trained SE-ResNext101 that gave 0.947 public LB before merging with @aispiriants. \nFirst of all, I've never trained full epoch at once and divided full train set into smaller groups to be able to save models more often. Those checkpoints' size varied depending on train image size, batch size, available time. In general, I tried to save one checkpoint every 4-6 hours of training.\n\nTraining flow:\n\n 1. Train on 128x128 images until net converges. It took around 20 checkpoints. It was enough to score around 0.945 on public LB.\n 2. After training on 128x128 stops giving any improvements, I used aggregated batches technique to use batch size around 3-4K images which gave a boost on local validation.\n 3. Further improvements gained from the fine-tuning model on bigger images (192x192 and then 256x256)\n \n**Some more tricks we used in our solution:**\n\n**Validation**\nAfter merging we stratified all train images into buckets by a country/class that gave us many buckets. For validation, we took 25 images per each bucket that give us a pretty solid validation score and allowed us to blend our models easily.\n\n**Ensembling**\nSaving checkpoints often gives a free boost by using snapshot ensembling. Final solution contains several snapshots of each model. All models were blended by the simple mean.\n\n**Pseudo-labeling**\nThat's the main score booster in the later steps of competition. Using blend that scored 0.951 on public LB we generated predicts for the test dataset and used them to fine-tune all of our models around 10 epochs using only pseudo-labeled data. It gave us around 0.952.\nThen using 0.952 blend, we generated pseudo-labels again and fine-tune again :) Which gives us 0.953 public LB. For some reason, we stopped doing this anymore, but probably we could get even higher score just re-pseudo-labeling :)",
    "433517": "Congrats and thanks for the sharing, it is great inspiration for future competition for starting out in kaggle ",
    "433395": "Congratulation and thanks for sharing @aispiriants and @firenero!!!\n\nCould you please share further on the following points:\n- Is there any tricks for selecting the optimizer or its schedule learning rate ? (or just use plain Adam / SGD)\n- Is there any useful pointers to help newbies (like myself) to implement batch accumulation ? (my team tried Adam Batch Accumulation but usually failed because the Out-of-memory problem)\n- So you use “all” snapshot models to make “snapshot ensemble” regardless of their accuracies ? (In my case, sometimes kicking-out the weak snapshots helps improve the overall ensemble a little bit)\n- which GPUs you did you use?",
    "433584": "Congrats and thanks for sharing !\nI have tried similar encoding with you but fail to get a decent performance gain :( \nI want to ask that what do you mean by using **number of strokes scaled to 0-255 ?** Is that means the sequence number of the stroke ?",
    "433682": "Thanks for sharing, Artur! Could you elaborate on how you encode time information into 3 channels?",
    "433625": "Is there a good working Se-Resnext implementation for keras?",
    "433552": "Hi @aispiriants, could you please clarify this sentence : \n\n“To be able to read any random image, all strokes from CSV files were separated to one image per binary file.”\n\nOne stroke per file?\n\nMany thanks again!",
    "433369": "I would have used the SE-ResNet, but I found that the structure of se-resnet is strange. For example: from 64*64, then change it to 1*1, then expand to 16*16, then 1*1 , etc. \nWhat's the advantages of this kind of structure? \nThx!",
    "439308": "",
    "433734": "",
    "434395": "Great job, thanks for sharing !",
    "434326": "Thanks for sharing!",
    "433562": "Congrats and thanks for sharing."
  }
}