{
  "id": 73738,
  "title": "1st place solution",
  "url": "/competitions/quickdraw-doodle-recognition/discussion/73738",
  "author_name": "",
  "post_date": "2018-12-05T08:02:46.948000",
  "votes": 201,
  "comment_count": 108,
  "views": 0,
  "content": "<p>Big thanks to Google for hosting this flawless competition and collecting such a great dataset. I'm also very excited to become top-5 in overall user ranking and even more excited for my teammate <a href=\"https://www.kaggle.com/pavelost\">Pavel Ostyakov</a> who got his second 1st place in a row! </p>\n\n<h2>CNN</h2>\n\n<p>First of all, Pavel did what he does best - trained a bunch of pytorch classification models. Here is the list of architectures: resnet18, resnet34, resnet50, resnet101, resnet152, resnext50, resnext101, densenet121, densenet201, vgg11, pnasnet, incresnet, polynet, nasnetmobile, senet154, seresnet50, seresnext50, seresnext101. </p>\n\n<p>One and three channels preprocessing were used as well as different image sizes starting from 112 and up to 256. The best model got 0.946 score, in total there were around 40 models. However, the gold could be achieved with a single model.</p>\n\n<h2>RNN</h2>\n\n<p>I trained a couple of LSTM models based on the <a href=\"https://www.kaggle.com/huyenvyvy/bidirectional-lstm-using-data-generator-lb-0-825\">best public kernel</a>. Tweaked the architecture a bit, got rid of dropouts and achieved 0.893 score. Would love to hear in the comments how you got better results.</p>\n\n<h2>LightGBM</h2>\n\n<p>How do you ensemble models with too many classes? This issue has been already resolved during <a href=\"https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/45733\">Cdiscount’s Image Classification Challenge</a>. The idea is the following: for each sample and for each model you collect top 10 probabilities with the labels, then convert them into 10 samples with the binary outcome - whether this is a correct label or not (9 negative examples + 1 positive). It's easy to feed such a dataset to any booster because the number of features will be small (equal to the number of models). On top of that, I also added some time-specific features. The most significant was maximum timestamp from the raw representations of the strokes.</p>\n\n<h2>Secret sauce (aka \"щепотка табака\")</h2>\n\n<p>As it was mentioned by <a href=\"https://www.kaggle.com/hengck23\">Heng CherKeng</a> a month ago <a href=\"https://www.kaggle.com/c/quickdraw-doodle-recognition/discussion/70540#416772\">classes in a test set were equally distributed</a>. It was a very important clue which seemed to be lost in the depths of the forum. I also did not see this comment but arrived at the same conclusion by noting that (112199+1)/340=330 (number of samples in the test set plus one is divisible by the number of classes). Knowing the structure of the test set gave us an average boost of 0.7% for every model. </p>\n\n<p>The algorithm behind postprocessing is the following: for the most popular class decrease all the probabilities iteratively by the same small value until it is no longer the most popular, repeat this procedure until all classes become equal. This technique was also used in <a href=\"https://www.kaggle.com/c/sp-society-camera-model-identification/discussion/49334\">one of the previous competitions</a> (see github link for the code).</p>\n\n<h2>Blending</h2>\n\n<p>After struggling for a week and producing 17 different balanced submits Pavel left me with the 5 last attempts to improve our public score of 0.956. I used <a href=\"https://www.kaggle.com/paulorzp/ensemble-weighted-voting\">this public ensembling kernel</a> and scored 0.957 after the first attempt. Changing weights from 5-i to 1/(i+1) gave us a slight additional boost (it mimics map3 weights) and the 1st place. </p>\n\n<h2>Data</h2>\n\n<p>We used 34000 random samples as the overall holdout set and 1 mln samples for building second layer models. All first layer models were trained on 49 mln simplified samples. Raw data features were only added to LightGBM model.</p>\n\n<h2>Key takeaways</h2>\n\n<ul>\n<li>Read forum carefully, especially when <a href=\"https://www.kaggle.com/hengck23\">Heng CherKeng</a> is present</li>\n<li>Study past solutions from similar competitions</li>\n</ul>",
  "messages": [
    {
      "id": 433596,
      "postDate": "2018-12-05T08:02:46.950Z",
      "content": "<p>Big thanks to Google for hosting this flawless competition and collecting such a great dataset. I'm also very excited to become top-5 in overall user ranking and even more excited for my teammate <a href=\"https://www.kaggle.com/pavelost\">Pavel Ostyakov</a> who got his second 1st place in a row! </p>\n\n<h2>CNN</h2>\n\n<p>First of all, Pavel did what he does best - trained a bunch of pytorch classification models. Here is the list of architectures: resnet18, resnet34, resnet50, resnet101, resnet152, resnext50, resnext101, densenet121, densenet201, vgg11, pnasnet, incresnet, polynet, nasnetmobile, senet154, seresnet50, seresnext50, seresnext101. </p>\n\n<p>One and three channels preprocessing were used as well as different image sizes starting from 112 and up to 256. The best model got 0.946 score, in total there were around 40 models. However, the gold could be achieved with a single model.</p>\n\n<h2>RNN</h2>\n\n<p>I trained a couple of LSTM models based on the <a href=\"https://www.kaggle.com/huyenvyvy/bidirectional-lstm-using-data-generator-lb-0-825\">best public kernel</a>. Tweaked the architecture a bit, got rid of dropouts and achieved 0.893 score. Would love to hear in the comments how you got better results.</p>\n\n<h2>LightGBM</h2>\n\n<p>How do you ensemble models with too many classes? This issue has been already resolved during <a href=\"https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/45733\">Cdiscount’s Image Classification Challenge</a>. The idea is the following: for each sample and for each model you collect top 10 probabilities with the labels, then convert them into 10 samples with the binary outcome - whether this is a correct label or not (9 negative examples + 1 positive). It's easy to feed such a dataset to any booster because the number of features will be small (equal to the number of models). On top of that, I also added some time-specific features. The most significant was maximum timestamp from the raw representations of the strokes.</p>\n\n<h2>Secret sauce (aka \"щепотка табака\")</h2>\n\n<p>As it was mentioned by <a href=\"https://www.kaggle.com/hengck23\">Heng CherKeng</a> a month ago <a href=\"https://www.kaggle.com/c/quickdraw-doodle-recognition/discussion/70540#416772\">classes in a test set were equally distributed</a>. It was a very important clue which seemed to be lost in the depths of the forum. I also did not see this comment but arrived at the same conclusion by noting that (112199+1)/340=330 (number of samples in the test set plus one is divisible by the number of classes). Knowing the structure of the test set gave us an average boost of 0.7% for every model. </p>\n\n<p>The algorithm behind postprocessing is the following: for the most popular class decrease all the probabilities iteratively by the same small value until it is no longer the most popular, repeat this procedure until all classes become equal. This technique was also used in <a href=\"https://www.kaggle.com/c/sp-society-camera-model-identification/discussion/49334\">one of the previous competitions</a> (see github link for the code).</p>\n\n<h2>Blending</h2>\n\n<p>After struggling for a week and producing 17 different balanced submits Pavel left me with the 5 last attempts to improve our public score of 0.956. I used <a href=\"https://www.kaggle.com/paulorzp/ensemble-weighted-voting\">this public ensembling kernel</a> and scored 0.957 after the first attempt. Changing weights from 5-i to 1/(i+1) gave us a slight additional boost (it mimics map3 weights) and the 1st place. </p>\n\n<h2>Data</h2>\n\n<p>We used 34000 random samples as the overall holdout set and 1 mln samples for building second layer models. All first layer models were trained on 49 mln simplified samples. Raw data features were only added to LightGBM model.</p>\n\n<h2>Key takeaways</h2>\n\n<ul>\n<li>Read forum carefully, especially when <a href=\"https://www.kaggle.com/hengck23\">Heng CherKeng</a> is present</li>\n<li>Study past solutions from similar competitions</li>\n</ul>",
      "rawMarkdown": "Big thanks to Google for hosting this flawless competition and collecting such a great dataset. I'm also very excited to become top-5 in overall user ranking and even more excited for my teammate [Pavel Ostyakov][1] who got his second 1st place in a row! \n\nCNN\n---\n\nFirst of all, Pavel did what he does best - trained a bunch of pytorch classification models. Here is the list of architectures: resnet18, resnet34, resnet50, resnet101, resnet152, resnext50, resnext101, densenet121, densenet201, vgg11, pnasnet, incresnet, polynet, nasnetmobile, senet154, seresnet50, seresnext50, seresnext101. \n\nOne and three channels preprocessing were used as well as different image sizes starting from 112 and up to 256. The best model got 0.946 score, in total there were around 40 models. However, the gold could be achieved with a single model.\n\nRNN\n---\nI trained a couple of LSTM models based on the [best public kernel][3]. Tweaked the architecture a bit, got rid of dropouts and achieved 0.893 score. Would love to hear in the comments how you got better results.\n\nLightGBM\n--------\n\nHow do you ensemble models with too many classes? This issue has been already resolved during [Cdiscount’s Image Classification Challenge][4]. The idea is the following: for each sample and for each model you collect top 10 probabilities with the labels, then convert them into 10 samples with the binary outcome - whether this is a correct label or not (9 negative examples + 1 positive). It's easy to feed such a dataset to any booster because the number of features will be small (equal to the number of models). On top of that, I also added some time-specific features. The most significant was maximum timestamp from the raw representations of the strokes.\n\nSecret sauce (aka \"щепотка табака\")\n-----\nAs it was mentioned by [Heng CherKeng][5] a month ago [classes in a test set were equally distributed][6]. It was a very important clue which seemed to be lost in the depths of the forum. I also did not see this comment but arrived at the same conclusion by noting that (112199+1)/340=330 (number of samples in the test set plus one is divisible by the number of classes). Knowing the structure of the test set gave us an average boost of 0.7% for every model. \n\nThe algorithm behind postprocessing is the following: for the most popular class decrease all the probabilities iteratively by the same small value until it is no longer the most popular, repeat this procedure until all classes become equal. This technique was also used in [one of the previous competitions][7] (see github link for the code).\n\nBlending\n-----\n\nAfter struggling for a week and producing 17 different balanced submits Pavel left me with the 5 last attempts to improve our public score of 0.956. I used [this public ensembling kernel][9] and scored 0.957 after the first attempt. Changing weights from 5-i to 1/(i+1) gave us a slight additional boost (it mimics map3 weights) and the 1st place. \n\nData\n----\n\nWe used 34000 random samples as the overall holdout set and 1 mln samples for building second layer models. All first layer models were trained on 49 mln simplified samples. Raw data features were only added to LightGBM model.\n\nKey takeaways\n-------------\n\n - Read forum carefully, especially when [Heng CherKeng][10] is present\n - Study past solutions from similar competitions\n\n\n  [1]: https://www.kaggle.com/pavelost \"Pavel\"\n  [2]: https://www.kaggle.com/pavelost \"Pavel\"\n  [3]: https://www.kaggle.com/huyenvyvy/bidirectional-lstm-using-data-generator-lb-0-825\n  [4]: https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/45733\n  [5]: https://www.kaggle.com/hengck23\n  [6]: https://www.kaggle.com/c/quickdraw-doodle-recognition/discussion/70540#416772\n  [7]: https://www.kaggle.com/c/sp-society-camera-model-identification/discussion/49334\n  [8]: https://www.kaggle.com/pavelost \"Pavel\"\n  [9]: https://www.kaggle.com/paulorzp/ensemble-weighted-voting\n  [10]: https://www.kaggle.com/hengck23",
      "votes": 201
    },
    {
      "id": 433717,
      "postDate": "2018-12-05T11:49:26.937Z",
      "content": "<p>@Pavel Pleskov</p>\n\n<p>Thanks for the solution and congrats for winning the first place! </p>\n\n<p>\"I also did not see this comment but arrived at the same conclusion by noting that (112199+1)/340=330 (number of samples in the test set plus one is divisible by the number of classes). Knowing the structure of the test set gave us an average boost of 0.7% for every model.\"</p>\n\n<p>What is the estimated score of your final solution if this trick of test class balancing is \"not applied\"? Would it be around  of 0.950?</p>",
      "rawMarkdown": "@Pavel Pleskov\n\nThanks for the solution and congrats for winning the first place! \n\n\"I also did not see this comment but arrived at the same conclusion by noting that (112199+1)/340=330 (number of samples in the test set plus one is divisible by the number of classes). Knowing the structure of the test set gave us an average boost of 0.7% for every model.\"\n\nWhat is the estimated score of your final solution if this trick of test class balancing is \"not applied\"? Would it be around  of 0.950?",
      "votes": 10,
      "replies": [
        {
          "id": 433729,
          "postDate": "2018-12-05T12:10:26.477Z",
          "content": "<p>0.94996 on private, you are close :)</p>",
          "rawMarkdown": "0.94996 on private, you are close :)",
          "votes": 2
        },
        {
          "id": 433844,
          "postDate": "2018-12-05T14:57:47.200Z",
          "content": "<p>@Heng CherKeng</p>\n\n<p>You are allways present, anytime, anywhere ;)</p>",
          "rawMarkdown": "@Heng CherKeng\n\nYou are allways present, anytime, anywhere ;)",
          "votes": 3
        },
        {
          "id": 433877,
          "postDate": "2018-12-05T15:36:14.820Z",
          "content": "<p><a href=\"https://www.kaggle.com/ppleskov\">@Pavel Pleskov</a> Could you please check the alternative solution to your bruteforce postprocessing search? If you have test probabilities <code>Y[112199x340]</code> how much you will get after this: <code>Y = Y / Y.mean(axis=0) ** (1 / temperature)</code> with <code>temperature</code> equal to 0.5 for example. Would it be better than your algorithm?</p>",
          "rawMarkdown": "[@Pavel Pleskov](https://www.kaggle.com/ppleskov) Could you please check the alternative solution to your bruteforce postprocessing search? If you have test probabilities <code>Y[112199x340]</code> how much you will get after this: <code>Y = Y / Y.mean(axis=0) ** (1 / temperature)</code> with <code>temperature</code> equal to 0.5 for example. Would it be better than your algorithm?",
          "votes": 1
        },
        {
          "id": 433904,
          "postDate": "2018-12-05T16:24:33.720Z",
          "content": "<p>I checked it on a random checkpoint. It seems, our variant is better :)</p>\n\n<p><img src=\"https://i.imgur.com/Q2hA0YH.png\"></p>",
          "rawMarkdown": "I checked it on a random checkpoint. It seems, our variant is better :)\n\n<img src=\"https://i.imgur.com/Q2hA0YH.png\">",
          "votes": 3
        },
        {
          "id": 433948,
          "postDate": "2018-12-05T17:24:08.567Z",
          "content": "<p>Thank you! One more, what is the initial score for this checkpoint?</p>",
          "rawMarkdown": "Thank you! One more, what is the initial score for this checkpoint?"
        },
        {
          "id": 433968,
          "postDate": "2018-12-05T17:52:10.597Z",
          "content": "<p>0.94514 Public and 0.94350 Private</p>",
          "rawMarkdown": "0.94514 Public and 0.94350 Private"
        }
      ]
    },
    {
      "id": 433618,
      "postDate": "2018-12-05T08:32:12.127Z",
      "content": "<p>Congrats to your team and Thank you for these valuable insights! :-)</p>",
      "rawMarkdown": "Congrats to your team and Thank you for these valuable insights! :-)\n",
      "votes": 8,
      "replies": [
        {
          "id": 433620,
          "postDate": "2018-12-05T08:38:25.857Z",
          "content": "<p>I'm curious, how much computation powers your team have in order to train many different models?</p>",
          "rawMarkdown": "I'm curious, how much computation powers your team have in order to train many different models?",
          "votes": 15
        },
        {
          "id": 433646,
          "postDate": "2018-12-05T09:40:46.343Z",
          "content": "<p>in order not to discourage other participants let's say we had much more gpus than needed to win this competition</p>",
          "rawMarkdown": "in order not to discourage other participants let's say we had much more gpus than needed to win this competition",
          "votes": 8
        }
      ]
    },
    {
      "id": 435273,
      "postDate": "2018-12-07T19:27:11.320Z",
      "content": "<p>I'm glad my script ( www.kaggle.com/paulorzp/ensemble-weighted-voting )  was useful. Congratulations!</p>",
      "rawMarkdown": "I'm glad my script ( www.kaggle.com/paulorzp/ensemble-weighted-voting )  was useful. Congratulations!",
      "votes": 6,
      "replies": [
        {
          "id": 435317,
          "postDate": "2018-12-07T20:52:42.880Z",
          "content": "<p>as usual :)</p>",
          "rawMarkdown": "as usual :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 434135,
      "postDate": "2018-12-06T00:16:43.287Z",
      "content": "<p>Great job! It is true that \"the devil is in the details.\"</p>",
      "rawMarkdown": "Great job! It is true that \"the devil is in the details.\"",
      "votes": 3
    },
    {
      "id": 433915,
      "postDate": "2018-12-05T16:43:48.710Z",
      "content": "<p>just trying to pass novice with a comment xD </p>",
      "rawMarkdown": "just trying to pass novice with a comment xD \n",
      "votes": 4,
      "replies": [
        {
          "id": 550028,
          "postDate": "2019-06-11T08:29:08.573Z",
          "content": "<p>me too)</p>",
          "rawMarkdown": "me too)",
          "replies": [
            {
              "id": 2435191,
              "postDate": "2023-09-12T18:58:08.507Z",
              "content": "<p>me too                                 </p>",
              "rawMarkdown": "me too                                 "
            }
          ]
        }
      ]
    },
    {
      "id": 433700,
      "postDate": "2018-12-05T11:19:28.233Z",
      "content": "<p>Congrats Pavel and Pavel  for both your winning and current kaggle ranking .  Thanks for sharing such great insights. </p>\n\n<blockquote>\n  <p>I trained a couple of LSTM models based on the best public kernel.\n  Tweaked the architecture a bit, got rid of dropouts and achieved 0.893\n  score. Would love to hear in the comments how you got better results.</p>\n</blockquote>\n\n<p>The LSTM model we used was built by my teamate @Luudactam.  It scored 0.936 on LB</p>",
      "rawMarkdown": "Congrats Pavel and Pavel  for both your winning and current kaggle ranking .  Thanks for sharing such great insights. \n\n&gt; I trained a couple of LSTM models based on the best public kernel.\n&gt; Tweaked the architecture a bit, got rid of dropouts and achieved 0.893\n&gt; score. Would love to hear in the comments how you got better results.\n\n\nThe LSTM model we used was built by my teamate @Luudactam.  It scored 0.936 on LB",
      "votes": 4,
      "replies": [
        {
          "id": 433704,
          "postDate": "2018-12-05T11:23:20.547Z",
          "content": "<p>That's interesting! we would like to hear your team approaches on how you achieved such 0.936 score on LSTM model? is it based on keras or pytorch?</p>",
          "rawMarkdown": "That's interesting! we would like to hear your team approaches on how you achieved such 0.936 score on LSTM model? is it based on keras or pytorch?\n",
          "votes": 7
        },
        {
          "id": 433740,
          "postDate": "2018-12-05T12:35:54.393Z",
          "content": "<p>As for me,  A simple Bi-LSTM + Attention gave me 0.905 on LB in the very early days of the competition and using merely 15k for each class.\nI stopped working on it and then focused on CNN when I teamed up with @Luudactam as he already had a very good LSTM model.  I would let him a talk a bit about his approach. </p>",
          "rawMarkdown": "As for me,  A simple Bi-LSTM + Attention gave me 0.905 on LB in the very early days of the competition and using merely 15k for each class.\nI stopped working on it and then focused on CNN when I teamed up with @Luudactam as he already had a very good LSTM model.  I would let him a talk a bit about his approach. ",
          "votes": 2
        },
        {
          "id": 433807,
          "postDate": "2018-12-05T13:54:50.910Z",
          "content": "<p>wow, looking forward to your teammate's RNN model.</p>",
          "rawMarkdown": "wow, looking forward to your teammate's RNN model."
        }
      ]
    },
    {
      "id": 437162,
      "postDate": "2018-12-11T13:28:08.677Z",
      "content": "<p>great</p>",
      "rawMarkdown": "great",
      "votes": 1
    },
    {
      "id": 434180,
      "postDate": "2018-12-06T02:49:33.807Z",
      "content": "<p>Fantastic write up - clear, succinct, and interesting.  Will be taking some of this stuff to the whales competition.</p>",
      "rawMarkdown": "Fantastic write up - clear, succinct, and interesting.  Will be taking some of this stuff to the whales competition.",
      "votes": 1
    },
    {
      "id": 434123,
      "postDate": "2018-12-05T23:19:42.523Z",
      "content": "<p>Congrats Pavel and Pavel. Thanks for sharing.</p>\n\n<p>I am still confused about the ensemble. If I had 5 models, for each sample I got 50 pobabilities and labels where each model had 10. Then how did I use them to construct 9 negative examples and 1 positive example?</p>",
      "rawMarkdown": "Congrats Pavel and Pavel. Thanks for sharing.\n\nI am still confused about the ensemble. If I had 5 models, for each sample I got 50 pobabilities and labels where each model had 10. Then how did I use them to construct 9 negative examples and 1 positive example?",
      "votes": 1,
      "replies": [
        {
          "id": 434300,
          "postDate": "2018-12-06T07:09:35.770Z",
          "content": "<p>each model predicts 10 labels only one of which is correct, thus, we have a binary problem with 10 examples</p>",
          "rawMarkdown": "each model predicts 10 labels only one of which is correct, thus, we have a binary problem with 10 examples"
        },
        {
          "id": 434546,
          "postDate": "2018-12-06T15:08:48.417Z",
          "content": "<p>So all the 10 examples have the same features except one feature which is possible class, and if the possible class equals the ground truth it will be positive example otherwise negative? But why did you say \"the number of features will be small (equal to the number of models)\"? I think each model uses multiple probabilities and relevant classes as features?</p>",
          "rawMarkdown": "So all the 10 examples have the same features except one feature which is possible class, and if the possible class equals the ground truth it will be positive example otherwise negative? But why did you say \"the number of features will be small (equal to the number of models)\"? I think each model uses multiple probabilities and relevant classes as features?"
        },
        {
          "id": 434570,
          "postDate": "2018-12-06T15:53:47Z",
          "content": "<p>how do you feed the categorical variables (class ids) to the model? you would have top-10 ids + 1 id of potential true class - how do you encode them for the model to be able to use this information?</p>",
          "rawMarkdown": "how do you feed the categorical variables (class ids) to the model? you would have top-10 ids + 1 id of potential true class - how do you encode them for the model to be able to use this information?",
          "votes": 1
        }
      ]
    },
    {
      "id": 433992,
      "postDate": "2018-12-05T18:35:24.077Z",
      "content": "<p>Congratulations, and thanks for sharing these important ideas.</p>",
      "rawMarkdown": "Congratulations, and thanks for sharing these important ideas.",
      "votes": 1
    },
    {
      "id": 433675,
      "postDate": "2018-12-05T10:30:33.283Z",
      "content": "<p>Hi Pavel, Congrats! I have a few questions:</p>\n\n<ol>\n<li>How do you encode temporal information into images? Is it similar to Beluga's <code>draw_cv2()</code> in this notebook, <a href=\"https://www.kaggle.com/gaborfodor/greyscale-mobilenet-lb-0-892\">https://www.kaggle.com/gaborfodor/greyscale-mobilenet-lb-0-892</a> ?</li>\n<li>Since you have tried so many models, could you list all the models  you used in the ensemble phase?</li>\n<li>With such a big dataset, do you generate training data on the fly, like the Beluga's notebook above(<code>image_generator_xd ()</code>)? Or do you preprocess data and save them into TFRecords for later usage?</li>\n</ol>",
      "rawMarkdown": "Hi Pavel, Congrats! I have a few questions:\n\n1. How do you encode temporal information into images? Is it similar to Beluga's `draw_cv2()` in this notebook, https://www.kaggle.com/gaborfodor/greyscale-mobilenet-lb-0-892 ?\n2. Since you have tried so many models, could you list all the models  you used in the ensemble phase?\n3. With such a big dataset, do you generate training data on the fly, like the Beluga's notebook above(`image_generator_xd ()`)? Or do you preprocess data and save them into TFRecords for later usage?",
      "votes": 1,
      "replies": [
        {
          "id": 433683,
          "postDate": "2018-12-05T10:46:22.893Z",
          "content": "<ol>\n<li>we used different variations of Beluga's pre-processing, no silver bullets, just wanted to increase diversity</li>\n<li>all models went to lgbm ensembling</li>\n<li>Pavel put all data into RAM for pytorch, I used chunks for keras as it was done in public kernels</li>\n</ol>",
          "rawMarkdown": "1. we used different variations of Beluga's pre-processing, no silver bullets, just wanted to increase diversity\n2. all models went to lgbm ensembling\n3. Pavel put all data into RAM for pytorch, I used chunks for keras as it was done in public kernels",
          "votes": 2
        },
        {
          "id": 434002,
          "postDate": "2018-12-05T18:50:54.157Z",
          "content": "<p>Thanks for these answers! For the third answer, I guess Pavel load all samples from <code>train_simplified.zip</code>, and then convert strokes to images on the fly during training, right? Because it's not possible to put all images in memory.</p>\n\n<p>For example, there are 49707579 training samples in total, even you use 128x128 grayscale images, it would consume <code>49707579*128*128/(1024**3)=758.47GB</code> memory.</p>",
          "rawMarkdown": "Thanks for these answers! For the third answer, I guess Pavel load all samples from `train_simplified.zip`, and then convert strokes to images on the fly during training, right? Because it's not possible to put all images in memory.\n\nFor example, there are 49707579 training samples in total, even you use 128x128 grayscale images, it would consume `49707579*128*128/(1024**3)=758.47GB` memory."
        },
        {
          "id": 434023,
          "postDate": "2018-12-05T19:38:21.453Z",
          "content": "<p>You are right!</p>",
          "rawMarkdown": "You are right!",
          "votes": 2
        },
        {
          "id": 434094,
          "postDate": "2018-12-05T21:44:47.120Z",
          "content": "<p>Thanks!</p>",
          "rawMarkdown": "Thanks!"
        }
      ]
    },
    {
      "id": 433661,
      "postDate": "2018-12-05T10:02:45.453Z",
      "content": "<p>Congrats Pavel and Pavel : )\nOne question, you mentioned both LGBM method and the blend kernel, how do these methods been used together?</p>",
      "rawMarkdown": "Congrats Pavel and Pavel : )\nOne question, you mentioned both LGBM method and the blend kernel, how do these methods been used together?",
      "votes": 1,
      "replies": [
        {
          "id": 433680,
          "postDate": "2018-12-05T10:43:07.670Z",
          "content": "<p>first we used several boosters (lgbm, xgb, cat) to ensemble CNN and RNN models, then post-processed each solution to balance classes, then blended all submits together</p>",
          "rawMarkdown": "first we used several boosters (lgbm, xgb, cat) to ensemble CNN and RNN models, then post-processed each solution to balance classes, then blended all submits together",
          "votes": 4
        },
        {
          "id": 433689,
          "postDate": "2018-12-05T10:56:41.450Z",
          "content": "<p>Get it, thanks!</p>",
          "rawMarkdown": "Get it, thanks!"
        },
        {
          "id": 434197,
          "postDate": "2018-12-06T03:25:02.223Z",
          "content": "<p>so, blending is on all submitted csv? not the probabilities?</p>",
          "rawMarkdown": "so, blending is on all submitted csv? not the probabilities?",
          "votes": 1
        },
        {
          "id": 434244,
          "postDate": "2018-12-06T05:06:47.007Z",
          "content": "<p>Yes, right </p>",
          "rawMarkdown": "Yes, right ",
          "votes": 1
        }
      ]
    },
    {
      "id": 433639,
      "postDate": "2018-12-05T09:22:57.840Z",
      "content": "<p>\" for each sample and for each model you collect top 10 probabilities with the labels <strong>than</strong> convert them into 10 sample with the binary outcome \" \nIs it than or then?</p>",
      "rawMarkdown": "\" for each sample and for each model you collect top 10 probabilities with the labels **than** convert them into 10 sample with the binary outcome \" \nIs it than or then?",
      "votes": 1,
      "replies": [
        {
          "id": 433642,
          "postDate": "2018-12-05T09:37:56.157Z",
          "content": "<p>great catch, thank you</p>",
          "rawMarkdown": "great catch, thank you"
        }
      ]
    },
    {
      "id": 433638,
      "postDate": "2018-12-05T09:20:08.370Z",
      "content": "<p>It is a great opportunity to learn from your overall idea. Many ideas seem beyond my current knowledge, so I have to slowly digest them. Thanks for sharing and of course congratulation!!!</p>",
      "rawMarkdown": "It is a great opportunity to learn from your overall idea. Many ideas seem beyond my current knowledge, so I have to slowly digest them. Thanks for sharing and of course congratulation!!!",
      "votes": 1,
      "replies": [
        {
          "id": 433685,
          "postDate": "2018-12-05T10:47:18.493Z",
          "content": "<p>we are glad that it was usefull</p>",
          "rawMarkdown": "we are glad that it was usefull",
          "votes": 1
        }
      ]
    },
    {
      "id": 433617,
      "postDate": "2018-12-05T08:31:23.037Z",
      "content": "<p>Thank you for sharing famous 'щепотка табака' trick</p>",
      "rawMarkdown": "Thank you for sharing famous 'щепотка табака' trick",
      "votes": 1,
      "replies": [
        {
          "id": 433645,
          "postDate": "2018-12-05T09:39:35.260Z",
          "content": "<p>coined by famous GM Vladimir Iglovikov</p>",
          "rawMarkdown": "coined by famous GM Vladimir Iglovikov"
        }
      ]
    },
    {
      "id": 434578,
      "postDate": "2018-12-06T16:22:56.153Z",
      "content": "<p>A huge congrats Pavel &amp; Pavel!</p>",
      "rawMarkdown": "A huge congrats Pavel &amp; Pavel!",
      "votes": 2
    },
    {
      "id": 433800,
      "postDate": "2018-12-05T13:45:09.677Z",
      "content": "<p>Did you have multithreading issues with PyTorch? It seems that on my machine all the available RAM is consumed during a single training epoch when <code>num_workers&gt;0</code> in <code>DataLoader</code> classes. We're discussing this issue on fastai forums and can't figure out if the problem comes from nightly version of PyTorch, fastai library or somewhere else. </p>\n\n<p>Though probably you had a lot of RAM if all the data was fit into the memory. Having a lot of hardware seems to be very helpful with this competition =)</p>",
      "rawMarkdown": "Did you have multithreading issues with PyTorch? It seems that on my machine all the available RAM is consumed during a single training epoch when `num_workers&gt;0` in `DataLoader` classes. We're discussing this issue on fastai forums and can't figure out if the problem comes from nightly version of PyTorch, fastai library or somewhere else. \n\nThough probably you had a lot of RAM if all the data was fit into the memory. Having a lot of hardware seems to be very helpful with this competition =)",
      "votes": 2,
      "replies": [
        {
          "id": 433819,
          "postDate": "2018-12-05T14:17:21.213Z",
          "content": "<p>I believe Pavel had 256 gb and used pytorch version 0.4.1 without any issues of this kind</p>",
          "rawMarkdown": "I believe Pavel had 256 gb and used pytorch version 0.4.1 without any issues of this kind",
          "votes": 2
        },
        {
          "id": 433906,
          "postDate": "2018-12-05T16:29:23.697Z",
          "content": "<p>Hey, there is a problem with Python multiprocessing in general (not only in PyTorch). For example, if you use a Python dict like a shared object, there will be some memory leakage even during reading values from this dict from another process (it's expected behaviour, you may google it for more information). To avoid this problem and to train multiple networks while loading the data only ones I used Pyro4 lib: <a href=\"https://github.com/irmen/Pyro4\">https://github.com/irmen/Pyro4</a></p>",
          "rawMarkdown": "Hey, there is a problem with Python multiprocessing in general (not only in PyTorch). For example, if you use a Python dict like a shared object, there will be some memory leakage even during reading values from this dict from another process (it's expected behaviour, you may google it for more information). To avoid this problem and to train multiple networks while loading the data only ones I used Pyro4 lib: https://github.com/irmen/Pyro4",
          "votes": 5
        },
        {
          "id": 433926,
          "postDate": "2018-12-05T16:57:45.527Z",
          "content": "<p>@PavelPleskov \nOk, got it! Yeah, 256 gb should be enough to read everything at once without troubles.</p>\n\n<p>@PavelOstyakov\nHm, that's an interesting observation. I was thinking that it should be possible to train a single ResNet model on a single machine. I have 32 gb on my local machine but it is not enough to train something with more then 20-25K per class. What a bummer! It sounds like a joke to have <code>DataLoader</code> class that has claims parallel execution capabilitie​s but without a chance to really use right out of the box.</p>\n\n<p>I think PyTorch team should put a <em>huge</em> disclaimer about this stuff onto their landing page =)</p>\n\n<p>For future reference, here is how the memory consumption plot looks like when I was trying to use plain <code>DataLoader</code> classes with 50K samples per class:</p>\n\n<p><img src=\"https://i.ibb.co/hdhsDX8/plot.png\"></p>\n\n<p>It really changes my understanding of how parallel things work in Python and in training models on huge datasets :D</p>",
          "rawMarkdown": "@PavelPleskov \nOk, got it! Yeah, 256 gb should be enough to read everything at once without troubles.\n\n@PavelOstyakov\nHm, that's an interesting observation. I was thinking that it should be possible to train a single ResNet model on a single machine. I have 32 gb on my local machine but it is not enough to train something with more then 20-25K per class. What a bummer! It sounds like a joke to have `DataLoader` class that has claims parallel execution capabilitie​s but without a chance to really use right out of the box.\n\nI think PyTorch team should put a _huge_ disclaimer about this stuff onto their landing page =)\n\nFor future reference, here is how the memory consumption plot looks like when I was trying to use plain `DataLoader` classes with 50K samples per class:\n\n<img src=\"https://i.ibb.co/hdhsDX8/plot.png\">\n\nIt really changes my understanding of how parallel things work in Python and in training models on huge datasets :D"
        },
        {
          "id": 433933,
          "postDate": "2018-12-05T17:07:53.430Z",
          "content": "<p>@devforfu,</p>\n\n<p>~20 GB RAM is enough for reading all the simplified train data using a trick with Pyro4</p>",
          "rawMarkdown": "@devforfu,\n\n~20 GB RAM is enough for reading all the simplified train data using a trick with Pyro4",
          "votes": 2
        },
        {
          "id": 433939,
          "postDate": "2018-12-05T17:14:05.623Z",
          "content": "<p>@PavelOstyakov That's cool! Thank you for sharing the insight about this library. I guess that Data Science competitions still not that straightforward as it could be thought =)</p>",
          "rawMarkdown": "@PavelOstyakov That's cool! Thank you for sharing the insight about this library. I guess that Data Science competitions still not that straightforward as it could be thought =)"
        },
        {
          "id": 433960,
          "postDate": "2018-12-05T17:37:49.850Z",
          "content": "<p>@devforfu I ran into the same issue using fast.ai and setting the num_workers to 0 did work as bandaid but wasn't really an option given the time left in the competition. My epochs went from 15 hours to over 40. I have ~200g of ram and noticed at some point the system would start to use the pagefile then eventually die. I did find this issue on the Pytorch forums which may be related. <a href=\"https://discuss.pytorch.org/t/dataloader-always-failed-with-interrupted-system-call/30242\">https://discuss.pytorch.org/t/dataloader-always-failed-with-interrupted-system-call/30242</a></p>",
          "rawMarkdown": "@devforfu I ran into the same issue using fast.ai and setting the num_workers to 0 did work as bandaid but wasn't really an option given the time left in the competition. My epochs went from 15 hours to over 40. I have ~200g of ram and noticed at some point the system would start to use the pagefile then eventually die. I did find this issue on the Pytorch forums which may be related. https://discuss.pytorch.org/t/dataloader-always-failed-with-interrupted-system-call/30242"
        },
        {
          "id": 434115,
          "postDate": "2018-12-05T22:52:22.137Z",
          "content": "<p>Hi Pavel, I'm new to Pyro4, what does it do? For example, I need an immutable, efficient dict sharable between multiple processes, for example, I load all training data into dict,  <code>key_id</code> as key,  <code>strokes</code> as values. Can Pyro4 do this for me?</p>",
          "rawMarkdown": "Hi Pavel, I'm new to Pyro4, what does it do? For example, I need an immutable, efficient dict sharable between multiple processes, for example, I load all training data into dict,  `key_id` as key,  `strokes` as values. Can Pyro4 do this for me?"
        },
        {
          "id": 434188,
          "postDate": "2018-12-06T03:02:30.177Z",
          "content": "<p>@SamAriabod Yeah, you're right, num_workers zero is a solution but it really drops the performance if you have a lot of cores. Same for me, the training time dropped drastically with my i7-6800K. Interesting that the error you've linked seems to have even different nature from what I have though still related to multiprocessing stuff.</p>",
          "rawMarkdown": "@SamAriabod Yeah, you're right, num_workers zero is a solution but it really drops the performance if you have a lot of cores. Same for me, the training time dropped drastically with my i7-6800K. Interesting that the error you've linked seems to have even different nature from what I have though still related to multiprocessing stuff."
        },
        {
          "id": 436463,
          "postDate": "2018-12-10T11:23:24.370Z",
          "content": "<p>Hi Pavel, what's the trick to load all simplified training data to Pyro4? I tried to load all training data into a dict inside Pyro4, but the memory consumption is too huge, even I have 52GB RAM it still not enough, my dict structure is <code>key_id -&gt; (strokes, y)</code>. </p>\n\n<p>The following is my Python code of <code>pyro4-server.py</code>:</p>\n\n<p>```python\nfrom <strong>future</strong> import print_function\nimport gc\nimport os\nimport Pyro4\nimport pandas as pd\nimport pyarrow.parquet\nfrom tqdm import tqdm\nimport time</p>\n\n<p>NUM_FOLDS = 100</p>\n\n<p>@Pyro4.expose\n@Pyro4.behavior(instance_mode=\"single\")\nclass QuickDrawData(object):\n    def <strong>init</strong>(self, file_path='/data/doodle/quickdraw.hdf5'):\n        self.key_id_to_strokes = {}</p>\n\n<pre><code>    if file_path.endswith('.hdf5'):\n        hdf_store = pd.HDFStore(file_path, mode='r')\n        assert len(hdf_store.keys()) == NUM_FOLDS\n    elif file_path.endswith('.parquet'):\n        parquet_file = pyarrow.parquet.ParquetFile(file_path)\n        assert parquet_file.num_row_groups == NUM_FOLDS\n\n    begin_time = time.time()\n    for k in tqdm(range(NUM_FOLDS)):\n        if file_path.endswith('.hdf5'):\n            df = hdf_store.get(key='fold_{0:02d}'.format(k))\n        elif file_path.endswith('.parquet'):\n            df = parquet_file.read_row_group(k, columns=['key_id', 'drawing', 'y']).to_pandas()\n\n        for index, row in df.iterrows():\n            key_id = row['key_id']\n            strokes = row['drawing']\n            y = row['y']\n            self.key_id_to_strokes[key_id] = (strokes, y)\n        del df\n        gc.collect()\n\n    if file_path.endswith('.hdf5'):\n        hdf_store.close()\n    print('Took {}s to read {}'.format(int(time.time()-begin_time), file_path))\n\ndef get(self, key_id):\n    return self.key_id_to_strokes[key_id]\n</code></pre>\n\n<p>def main():\n    Pyro4.Daemon.serveSimple(\n            {\n                QuickDrawData: \"quickdraw.data\"\n            },\n            ns = True)</p>\n\n<p>if <strong>name</strong>==\"<strong>main</strong>\":\n    main()\n```</p>",
          "rawMarkdown": "Hi Pavel, what's the trick to load all simplified training data to Pyro4? I tried to load all training data into a dict inside Pyro4, but the memory consumption is too huge, even I have 52GB RAM it still not enough, my dict structure is `key_id -&gt; (strokes, y)`. \n\nThe following is my Python code of `pyro4-server.py`:\n\n```python\nfrom __future__ import print_function\nimport gc\nimport os\nimport Pyro4\nimport pandas as pd\nimport pyarrow.parquet\nfrom tqdm import tqdm\nimport time\n\nNUM_FOLDS = 100\n\n\n@Pyro4.expose\n@Pyro4.behavior(instance_mode=\"single\")\nclass QuickDrawData(object):\n    def __init__(self, file_path='/data/doodle/quickdraw.hdf5'):\n        self.key_id_to_strokes = {}\n        \n        if file_path.endswith('.hdf5'):\n            hdf_store = pd.HDFStore(file_path, mode='r')\n            assert len(hdf_store.keys()) == NUM_FOLDS\n        elif file_path.endswith('.parquet'):\n            parquet_file = pyarrow.parquet.ParquetFile(file_path)\n            assert parquet_file.num_row_groups == NUM_FOLDS\n\n        begin_time = time.time()\n        for k in tqdm(range(NUM_FOLDS)):\n            if file_path.endswith('.hdf5'):\n                df = hdf_store.get(key='fold_{0:02d}'.format(k))\n            elif file_path.endswith('.parquet'):\n                df = parquet_file.read_row_group(k, columns=['key_id', 'drawing', 'y']).to_pandas()\n\n            for index, row in df.iterrows():\n                key_id = row['key_id']\n                strokes = row['drawing']\n                y = row['y']\n                self.key_id_to_strokes[key_id] = (strokes, y)\n            del df\n            gc.collect()\n\n        if file_path.endswith('.hdf5'):\n            hdf_store.close()\n        print('Took {}s to read {}'.format(int(time.time()-begin_time), file_path))\n\n    def get(self, key_id):\n        return self.key_id_to_strokes[key_id]\n\ndef main():\n    Pyro4.Daemon.serveSimple(\n            {\n                QuickDrawData: \"quickdraw.data\"\n            },\n            ns = True)\n\nif __name__==\"__main__\":\n    main()\n```",
          "votes": 1
        }
      ]
    },
    {
      "id": 433641,
      "postDate": "2018-12-05T09:33:31.593Z",
      "content": "<p>Excellent work and congratulations ! It is an Awesome work and very good systematic approach.</p>\n\n<p>I reached 0.935 with a single recurrent network that has the following structure:\n1d time conv. (Kernel size 3)\n1d time conv (kernel size 5)\nBi directional gru with stabilization layers</p>\n\n<p>These blocks are concatenated 3 times with increasing feature size. Last layer takes the last sequence element features and fed it to a dense layer.</p>",
      "rawMarkdown": "Excellent work and congratulations ! It is an Awesome work and very good systematic approach.\n\nI reached 0.935 with a single recurrent network that has the following structure:\n1d time conv. (Kernel size 3)\n1d time conv (kernel size 5)\nBi directional gru with stabilization layers\n\nThese blocks are concatenated 3 times with increasing feature size. Last layer takes the last sequence element features and fed it to a dense layer.",
      "votes": 2,
      "replies": [
        {
          "id": 433643,
          "postDate": "2018-12-05T09:38:26.910Z",
          "content": "<p>thanks for sharing!</p>",
          "rawMarkdown": "thanks for sharing!"
        }
      ]
    },
    {
      "id": 438254,
      "postDate": "2018-12-13T11:07:56.720Z",
      "content": "<p>Impressive...</p>",
      "rawMarkdown": "Impressive..."
    },
    {
      "id": 438253,
      "postDate": "2018-12-13T10:59:43.427Z",
      "content": "<p>Congratutlations @Pravel Plestov .\nJust a query how much Training Top 3 accuracy, loss did you get on your best model?</p>\n\n<p>Trying something different so wanted a benchmark to compare to.</p>",
      "rawMarkdown": "Congratutlations @Pravel Plestov .\nJust a query how much Training Top 3 accuracy, loss did you get on your best model?\n\nTrying something different so wanted a benchmark to compare to."
    },
    {
      "id": 437371,
      "postDate": "2018-12-11T19:29:33.257Z",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!"
    },
    {
      "id": 437228,
      "postDate": "2018-12-11T15:24:43.870Z",
      "content": "<p>Who runs this bich?</p>",
      "rawMarkdown": "Who runs this bich?"
    },
    {
      "id": 437214,
      "postDate": "2018-12-11T14:58:35.630Z",
      "content": "<p>Congrats ! You ran all of this on Kaggle Kernel or your own servers ?</p>",
      "rawMarkdown": "Congrats ! You ran all of this on Kaggle Kernel or your own servers ?",
      "replies": [
        {
          "id": 437351,
          "postDate": "2018-12-11T18:57:08.737Z",
          "content": "<p>on our own of course</p>",
          "rawMarkdown": "on our own of course",
          "votes": -1
        }
      ]
    },
    {
      "id": 437136,
      "postDate": "2018-12-11T12:25:34.633Z",
      "content": "<p>How someone can get 100% in accurecy?</p>",
      "rawMarkdown": "How someone can get 100% in accurecy?",
      "replies": [
        {
          "id": 437352,
          "postDate": "2018-12-11T18:58:00.037Z",
          "content": "<p>due to the noise in the data, it is highly unlikely</p>",
          "rawMarkdown": "due to the noise in the data, it is highly unlikely",
          "votes": -1
        }
      ]
    },
    {
      "id": 436699,
      "postDate": "2018-12-10T19:13:17.300Z",
      "content": "<p>This is awesome, congratulations and thank you for the write up </p>",
      "rawMarkdown": "This is awesome, congratulations and thank you for the write up "
    },
    {
      "id": 436541,
      "postDate": "2018-12-10T13:58:57.927Z",
      "content": "<p>Congrat your team, Many thanks for your sharing!</p>",
      "rawMarkdown": "Congrat your team, Many thanks for your sharing!"
    },
    {
      "id": 436419,
      "postDate": "2018-12-10T09:08:18.967Z",
      "content": "<p>make firend</p>",
      "rawMarkdown": "make firend"
    },
    {
      "id": 435989,
      "postDate": "2018-12-09T08:19:54.960Z",
      "content": "<p>Thank you so much for sharing this knowledge with us. </p>",
      "rawMarkdown": "Thank you so much for sharing this knowledge with us. "
    },
    {
      "id": 435670,
      "postDate": "2018-12-08T14:05:56.300Z",
      "content": "<p>Learned so much! Amazing approaches, congratulations and thanks for sharing! </p>",
      "rawMarkdown": "Learned so much! Amazing approaches, congratulations and thanks for sharing! "
    },
    {
      "id": 435162,
      "postDate": "2018-12-07T15:51:09.797Z",
      "content": "<p>Congrats</p>",
      "rawMarkdown": "Congrats"
    },
    {
      "id": 434884,
      "postDate": "2018-12-07T05:18:26.450Z",
      "content": "<p>Congrats Pavel and Pavel. Thank you for sharing.\nCan you share the models you have trained? I want to use this model to extract some of the features of the data.\nthanks!</p>",
      "rawMarkdown": "Congrats Pavel and Pavel. Thank you for sharing.\nCan you share the models you have trained? I want to use this model to extract some of the features of the data.\nthanks!"
    },
    {
      "id": 434795,
      "postDate": "2018-12-07T00:45:39.987Z",
      "content": "<p>Great write-up, congrats!</p>",
      "rawMarkdown": "Great write-up, congrats!"
    },
    {
      "id": 434766,
      "postDate": "2018-12-06T22:56:04.490Z",
      "content": "<p>Congratulations! </p>",
      "rawMarkdown": "Congratulations! "
    },
    {
      "id": 434748,
      "postDate": "2018-12-06T21:59:05.627Z",
      "content": "<p>Congrats Pavel and Pavel</p>",
      "rawMarkdown": "Congrats Pavel and Pavel"
    },
    {
      "id": 434393,
      "postDate": "2018-12-06T10:10:20.020Z",
      "content": "<p>Congratz, and thanks for sharing ! I definitely enjoyed reading this !</p>",
      "rawMarkdown": "Congratz, and thanks for sharing ! I definitely enjoyed reading this !"
    },
    {
      "id": 434331,
      "postDate": "2018-12-06T08:03:08.810Z",
      "content": "<p>Congrats to your team and Thank you for these valuable insights &lt;3. BTW, what is the best model that got 0.946 score ? How big your batch size and is and how much memory did you use to train it ?</p>",
      "rawMarkdown": "Congrats to your team and Thank you for these valuable insights &lt;3. BTW, what is the best model that got 0.946 score ? How big your batch size and is and how much memory did you use to train it ?",
      "replies": [
        {
          "id": 434456,
          "postDate": "2018-12-06T12:32:37.310Z",
          "content": "<p>PNASNet5Large, 128 size on 8 GPUs</p>",
          "rawMarkdown": "PNASNet5Large, 128 size on 8 GPUs"
        },
        {
          "id": 438918,
          "postDate": "2018-12-14T12:12:08.743Z",
          "content": "<p>8 Tesla V100 ?</p>",
          "rawMarkdown": "8 Tesla V100 ?"
        }
      ]
    },
    {
      "id": 434297,
      "postDate": "2018-12-06T07:03:53.710Z",
      "content": "<p>Congratulations! How did you choose your hyper-parameters for your LightGBM model? </p>",
      "rawMarkdown": "Congratulations! How did you choose your hyper-parameters for your LightGBM model? ",
      "replies": [
        {
          "id": 434301,
          "postDate": "2018-12-06T07:11:35.497Z",
          "content": "<p>it turned out that model was not sensitive to the parameters, only feature generation added to the quality slightly</p>",
          "rawMarkdown": "it turned out that model was not sensitive to the parameters, only feature generation added to the quality slightly"
        },
        {
          "id": 434322,
          "postDate": "2018-12-06T07:41:10.897Z",
          "content": "<p>I'm actually finding that out the hard way with the Elo competition. Tried a million different param tunings, but the only thing that seems to make a difference is feature engineering. </p>",
          "rawMarkdown": "I'm actually finding that out the hard way with the Elo competition. Tried a million different param tunings, but the only thing that seems to make a difference is feature engineering. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 434288,
      "postDate": "2018-12-06T06:49:07.787Z",
      "content": "<p>Thanks for sharing. Would you please explain your LGBM method in detail? How did you get training data for \nLGBM? How did you do in the inference stage? Thanks.</p>",
      "rawMarkdown": "Thanks for sharing. Would you please explain your LGBM method in detail? How did you get training data for \nLGBM? How did you do in the inference stage? Thanks."
    },
    {
      "id": 433841,
      "postDate": "2018-12-05T14:55:31.210Z",
      "content": "<p>@Pavel Pleskov</p>\n\n<p>Excellent 1st place and excellent information about your <em>щепотка табака</em>, among other stuff. </p>\n\n<p>Congratulations !</p>",
      "rawMarkdown": "@Pavel Pleskov\n\nExcellent 1st place and excellent information about your *щепотка табака*, among other stuff. \n\nCongratulations !"
    },
    {
      "id": 433788,
      "postDate": "2018-12-05T13:33:03.777Z",
      "content": "<p>Congrats for your win and great learning for all of us from your notes !!</p>",
      "rawMarkdown": "Congrats for your win and great learning for all of us from your notes !!"
    },
    {
      "id": 435448,
      "postDate": "2018-12-08T04:01:58.453Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true
    },
    {
      "id": 434528,
      "postDate": "2018-12-06T14:40:25.380Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 434317,
      "postDate": "2018-12-06T07:39:30.793Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 434251,
      "postDate": "2018-12-06T05:25:32.490Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 433733,
      "postDate": "2018-12-05T12:28:11.500Z",
      "rawMarkdown": "",
      "votes": 8,
      "isDeleted": true
    },
    {
      "id": 455803,
      "postDate": "2019-01-14T16:01:11.073Z",
      "content": "<p>Congrats and thanks for the solution</p>",
      "rawMarkdown": "Congrats and thanks for the solution",
      "votes": 1
    },
    {
      "id": 433681,
      "postDate": "2018-12-05T10:43:50.423Z",
      "content": "<p>Congratulations and thanks for sharing.</p>",
      "rawMarkdown": "Congratulations and thanks for sharing.",
      "votes": 1
    },
    {
      "id": 433651,
      "postDate": "2018-12-05T09:47:11.157Z",
      "content": "<p>Congratulations and thanks for sharing.</p>",
      "rawMarkdown": "Congratulations and thanks for sharing.",
      "votes": 1
    },
    {
      "id": 437920,
      "postDate": "2018-12-12T19:00:04.887Z",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": 2
    },
    {
      "id": 435302,
      "postDate": "2018-12-07T20:20:58.050Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": 2
    },
    {
      "id": 438259,
      "postDate": "2018-12-13T11:25:14.827Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing"
    },
    {
      "id": 437844,
      "postDate": "2018-12-12T15:54:27.560Z",
      "content": "<p>Good job, thanks for sharing!</p>",
      "rawMarkdown": "Good job, thanks for sharing!"
    },
    {
      "id": 437728,
      "postDate": "2018-12-12T11:34:17.863Z",
      "content": "<p>Congratulations and thanks for sharing.</p>",
      "rawMarkdown": "Congratulations and thanks for sharing."
    },
    {
      "id": 437315,
      "postDate": "2018-12-11T17:48:53.727Z",
      "content": "<p>Congratulations! Thanks for sharing!</p>",
      "rawMarkdown": "Congratulations! Thanks for sharing!"
    },
    {
      "id": 436255,
      "postDate": "2018-12-10T01:56:27.560Z",
      "content": "<p>Thank you for sharing!</p>",
      "rawMarkdown": "Thank you for sharing!"
    },
    {
      "id": 436097,
      "postDate": "2018-12-09T15:12:34.323Z",
      "content": "<p>Congrats\nand many thanks for sharing :D</p>",
      "rawMarkdown": "Congrats\nand many thanks for sharing :D"
    },
    {
      "id": 435931,
      "postDate": "2018-12-09T05:09:39.770Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!"
    },
    {
      "id": 435640,
      "postDate": "2018-12-08T12:44:14.350Z",
      "content": "<p>Congratulations and thanks for sharing.</p>",
      "rawMarkdown": "Congratulations and thanks for sharing."
    },
    {
      "id": 435298,
      "postDate": "2018-12-07T20:11:25.347Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!"
    },
    {
      "id": 435207,
      "postDate": "2018-12-07T17:14:04.617Z",
      "content": "<p>Great write up, thanks for sharing!</p>",
      "rawMarkdown": "Great write up, thanks for sharing!"
    },
    {
      "id": 434540,
      "postDate": "2018-12-06T14:57:29.400Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing"
    },
    {
      "id": 434252,
      "postDate": "2018-12-06T05:26:24.243Z",
      "content": "<p>Congrats!! Thanks for sharing!</p>",
      "rawMarkdown": "Congrats!! Thanks for sharing!"
    },
    {
      "id": 434215,
      "postDate": "2018-12-06T03:57:25.283Z",
      "content": "<p>Congratulations and thanks for sharing!</p>",
      "rawMarkdown": "Congratulations and thanks for sharing!"
    },
    {
      "id": 434176,
      "postDate": "2018-12-06T02:31:31.107Z",
      "content": "<p>Congratulations and thanks for sharing.</p>",
      "rawMarkdown": "Congratulations and thanks for sharing."
    }
  ],
  "comments": [
    {
      "id": 433717,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-12-05T11:49:26.937000",
      "content": "<p>@Pavel Pleskov</p>\n\n<p>Thanks for the solution and congrats for winning the first place! </p>\n\n<p>\"I also did not see this comment but arrived at the same conclusion by noting that (112199+1)/340=330 (number of samples in the test set plus one is divisible by the number of classes). Knowing the structure of the test set gave us an average boost of 0.7% for every model.\"</p>\n\n<p>What is the estimated score of your final solution if this trick of test class balancing is \"not applied\"? Would it be around  of 0.950?</p>",
      "votes": 10,
      "replies": [
        {
          "id": 433729,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-05T12:10:26.477000",
          "content": "<p>0.94996 on private, you are close :)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 433844,
          "author_name": "mezoganet",
          "author_url": "",
          "post_date": "2018-12-05T14:57:47.200000",
          "content": "<p>@Heng CherKeng</p>\n\n<p>You are allways present, anytime, anywhere ;)</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 433877,
          "author_name": "ivan",
          "author_url": "",
          "post_date": "2018-12-05T15:36:14.820000",
          "content": "<p><a href=\"https://www.kaggle.com/ppleskov\">@Pavel Pleskov</a> Could you please check the alternative solution to your bruteforce postprocessing search? If you have test probabilities <code>Y[112199x340]</code> how much you will get after this: <code>Y = Y / Y.mean(axis=0) ** (1 / temperature)</code> with <code>temperature</code> equal to 0.5 for example. Would it be better than your algorithm?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 433904,
          "author_name": "Pavel Ostyakov",
          "author_url": "",
          "post_date": "2018-12-05T16:24:33.720000",
          "content": "<p>I checked it on a random checkpoint. It seems, our variant is better :)</p>\n\n<p><img src=\"https://i.imgur.com/Q2hA0YH.png\"></p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 433948,
          "author_name": "ivan",
          "author_url": "",
          "post_date": "2018-12-05T17:24:08.567000",
          "content": "<p>Thank you! One more, what is the initial score for this checkpoint?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 433968,
          "author_name": "Pavel Ostyakov",
          "author_url": "",
          "post_date": "2018-12-05T17:52:10.597000",
          "content": "<p>0.94514 Public and 0.94350 Private</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 433618,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-05T08:32:12.127000",
      "content": "<p>Congrats to your team and Thank you for these valuable insights! :-)</p>",
      "votes": 8,
      "replies": [
        {
          "id": 433620,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-05T08:38:25.857000",
          "content": "<p>I'm curious, how much computation powers your team have in order to train many different models?</p>",
          "votes": 15,
          "replies": []
        },
        {
          "id": 433646,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-05T09:40:46.343000",
          "content": "<p>in order not to discourage other participants let's say we had much more gpus than needed to win this competition</p>",
          "votes": 8,
          "replies": []
        }
      ]
    },
    {
      "id": 435273,
      "author_name": "Paulo Pinto",
      "author_url": "",
      "post_date": "2018-12-07T19:27:11.320000",
      "content": "<p>I'm glad my script ( www.kaggle.com/paulorzp/ensemble-weighted-voting )  was useful. Congratulations!</p>",
      "votes": 6,
      "replies": [
        {
          "id": 435317,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-07T20:52:42.880000",
          "content": "<p>as usual :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 434135,
      "author_name": "Peiyuan Liao",
      "author_url": "",
      "post_date": "2018-12-06T00:16:43.287000",
      "content": "<p>Great job! It is true that \"the devil is in the details.\"</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 433915,
      "author_name": "Houssam Boudiar",
      "author_url": "",
      "post_date": "2018-12-05T16:43:48.710000",
      "content": "<p>just trying to pass novice with a comment xD </p>",
      "votes": 4,
      "replies": [
        {
          "id": 550028,
          "author_name": "Darkhan Zhursin",
          "author_url": "",
          "post_date": "2019-06-11T08:29:08.573000",
          "content": "<p>me too)</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2435191,
              "author_name": "Pumpkin panda",
              "author_url": "",
              "post_date": "2023-09-12T18:58:08.507000",
              "content": "<p>me too                                 </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 433700,
      "author_name": "Serigne ",
      "author_url": "",
      "post_date": "2018-12-05T11:19:28.233000",
      "content": "<p>Congrats Pavel and Pavel  for both your winning and current kaggle ranking .  Thanks for sharing such great insights. </p>\n\n<blockquote>\n  <p>I trained a couple of LSTM models based on the best public kernel.\n  Tweaked the architecture a bit, got rid of dropouts and achieved 0.893\n  score. Would love to hear in the comments how you got better results.</p>\n</blockquote>\n\n<p>The LSTM model we used was built by my teamate @Luudactam.  It scored 0.936 on LB</p>",
      "votes": 4,
      "replies": [
        {
          "id": 433704,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-05T11:23:20.547000",
          "content": "<p>That's interesting! we would like to hear your team approaches on how you achieved such 0.936 score on LSTM model? is it based on keras or pytorch?</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 433740,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2018-12-05T12:35:54.393000",
          "content": "<p>As for me,  A simple Bi-LSTM + Attention gave me 0.905 on LB in the very early days of the competition and using merely 15k for each class.\nI stopped working on it and then focused on CNN when I teamed up with @Luudactam as he already had a very good LSTM model.  I would let him a talk a bit about his approach. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 433807,
          "author_name": "upup",
          "author_url": "",
          "post_date": "2018-12-05T13:54:50.910000",
          "content": "<p>wow, looking forward to your teammate's RNN model.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 437162,
      "author_name": "Alexa",
      "author_url": "",
      "post_date": "2018-12-11T13:28:08.677000",
      "content": "<p>great</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 434180,
      "author_name": "Matthew Anderson",
      "author_url": "",
      "post_date": "2018-12-06T02:49:33.807000",
      "content": "<p>Fantastic write up - clear, succinct, and interesting.  Will be taking some of this stuff to the whales competition.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 434123,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-05T23:19:42.523000",
      "content": "<p>Congrats Pavel and Pavel. Thanks for sharing.</p>\n\n<p>I am still confused about the ensemble. If I had 5 models, for each sample I got 50 pobabilities and labels where each model had 10. Then how did I use them to construct 9 negative examples and 1 positive example?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 434300,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-06T07:09:35.770000",
          "content": "<p>each model predicts 10 labels only one of which is correct, thus, we have a binary problem with 10 examples</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 434546,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-06T15:08:48.417000",
          "content": "<p>So all the 10 examples have the same features except one feature which is possible class, and if the possible class equals the ground truth it will be positive example otherwise negative? But why did you say \"the number of features will be small (equal to the number of models)\"? I think each model uses multiple probabilities and relevant classes as features?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 434570,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2018-12-06T15:53:47",
          "content": "<p>how do you feed the categorical variables (class ids) to the model? you would have top-10 ids + 1 id of potential true class - how do you encode them for the model to be able to use this information?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 433992,
      "author_name": "JM100",
      "author_url": "",
      "post_date": "2018-12-05T18:35:24.077000",
      "content": "<p>Congratulations, and thanks for sharing these important ideas.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 433675,
      "author_name": "[he.ai]soulmachine",
      "author_url": "",
      "post_date": "2018-12-05T10:30:33.283000",
      "content": "<p>Hi Pavel, Congrats! I have a few questions:</p>\n\n<ol>\n<li>How do you encode temporal information into images? Is it similar to Beluga's <code>draw_cv2()</code> in this notebook, <a href=\"https://www.kaggle.com/gaborfodor/greyscale-mobilenet-lb-0-892\">https://www.kaggle.com/gaborfodor/greyscale-mobilenet-lb-0-892</a> ?</li>\n<li>Since you have tried so many models, could you list all the models  you used in the ensemble phase?</li>\n<li>With such a big dataset, do you generate training data on the fly, like the Beluga's notebook above(<code>image_generator_xd ()</code>)? Or do you preprocess data and save them into TFRecords for later usage?</li>\n</ol>",
      "votes": 1,
      "replies": [
        {
          "id": 433683,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-05T10:46:22.893000",
          "content": "<ol>\n<li>we used different variations of Beluga's pre-processing, no silver bullets, just wanted to increase diversity</li>\n<li>all models went to lgbm ensembling</li>\n<li>Pavel put all data into RAM for pytorch, I used chunks for keras as it was done in public kernels</li>\n</ol>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 434002,
          "author_name": "[he.ai]soulmachine",
          "author_url": "",
          "post_date": "2018-12-05T18:50:54.157000",
          "content": "<p>Thanks for these answers! For the third answer, I guess Pavel load all samples from <code>train_simplified.zip</code>, and then convert strokes to images on the fly during training, right? Because it's not possible to put all images in memory.</p>\n\n<p>For example, there are 49707579 training samples in total, even you use 128x128 grayscale images, it would consume <code>49707579*128*128/(1024**3)=758.47GB</code> memory.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 434023,
          "author_name": "Pavel Ostyakov",
          "author_url": "",
          "post_date": "2018-12-05T19:38:21.453000",
          "content": "<p>You are right!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 434094,
          "author_name": "[he.ai]soulmachine",
          "author_url": "",
          "post_date": "2018-12-05T21:44:47.120000",
          "content": "<p>Thanks!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 433661,
      "author_name": "Yiheng Wang",
      "author_url": "",
      "post_date": "2018-12-05T10:02:45.453000",
      "content": "<p>Congrats Pavel and Pavel : )\nOne question, you mentioned both LGBM method and the blend kernel, how do these methods been used together?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 433680,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-05T10:43:07.670000",
          "content": "<p>first we used several boosters (lgbm, xgb, cat) to ensemble CNN and RNN models, then post-processed each solution to balance classes, then blended all submits together</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 433689,
          "author_name": "Yiheng Wang",
          "author_url": "",
          "post_date": "2018-12-05T10:56:41.450000",
          "content": "<p>Get it, thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 434197,
          "author_name": "good good study",
          "author_url": "",
          "post_date": "2018-12-06T03:25:02.223000",
          "content": "<p>so, blending is on all submitted csv? not the probabilities?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 434244,
          "author_name": "Pavel Ostyakov",
          "author_url": "",
          "post_date": "2018-12-06T05:06:47.007000",
          "content": "<p>Yes, right </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 433639,
      "author_name": "Hai Nam Nguyen",
      "author_url": "",
      "post_date": "2018-12-05T09:22:57.840000",
      "content": "<p>\" for each sample and for each model you collect top 10 probabilities with the labels <strong>than</strong> convert them into 10 sample with the binary outcome \" \nIs it than or then?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 433642,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-05T09:37:56.157000",
          "content": "<p>great catch, thank you</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 433638,
      "author_name": "Neuron Engineer",
      "author_url": "",
      "post_date": "2018-12-05T09:20:08.370000",
      "content": "<p>It is a great opportunity to learn from your overall idea. Many ideas seem beyond my current knowledge, so I have to slowly digest them. Thanks for sharing and of course congratulation!!!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 433685,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-05T10:47:18.493000",
          "content": "<p>we are glad that it was usefull</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 433617,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-05T08:31:23.037000",
      "content": "<p>Thank you for sharing famous 'щепотка табака' trick</p>",
      "votes": 1,
      "replies": [
        {
          "id": 433645,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-05T09:39:35.260000",
          "content": "<p>coined by famous GM Vladimir Iglovikov</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 434578,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2018-12-06T16:22:56.153000",
      "content": "<p>A huge congrats Pavel &amp; Pavel!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 433800,
      "author_name": "Ilia Zaitsev",
      "author_url": "",
      "post_date": "2018-12-05T13:45:09.677000",
      "content": "<p>Did you have multithreading issues with PyTorch? It seems that on my machine all the available RAM is consumed during a single training epoch when <code>num_workers&gt;0</code> in <code>DataLoader</code> classes. We're discussing this issue on fastai forums and can't figure out if the problem comes from nightly version of PyTorch, fastai library or somewhere else. </p>\n\n<p>Though probably you had a lot of RAM if all the data was fit into the memory. Having a lot of hardware seems to be very helpful with this competition =)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 433819,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-05T14:17:21.213000",
          "content": "<p>I believe Pavel had 256 gb and used pytorch version 0.4.1 without any issues of this kind</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 433906,
          "author_name": "Pavel Ostyakov",
          "author_url": "",
          "post_date": "2018-12-05T16:29:23.697000",
          "content": "<p>Hey, there is a problem with Python multiprocessing in general (not only in PyTorch). For example, if you use a Python dict like a shared object, there will be some memory leakage even during reading values from this dict from another process (it's expected behaviour, you may google it for more information). To avoid this problem and to train multiple networks while loading the data only ones I used Pyro4 lib: <a href=\"https://github.com/irmen/Pyro4\">https://github.com/irmen/Pyro4</a></p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 433926,
          "author_name": "Ilia Zaitsev",
          "author_url": "",
          "post_date": "2018-12-05T16:57:45.527000",
          "content": "<p>@PavelPleskov \nOk, got it! Yeah, 256 gb should be enough to read everything at once without troubles.</p>\n\n<p>@PavelOstyakov\nHm, that's an interesting observation. I was thinking that it should be possible to train a single ResNet model on a single machine. I have 32 gb on my local machine but it is not enough to train something with more then 20-25K per class. What a bummer! It sounds like a joke to have <code>DataLoader</code> class that has claims parallel execution capabilitie​s but without a chance to really use right out of the box.</p>\n\n<p>I think PyTorch team should put a <em>huge</em> disclaimer about this stuff onto their landing page =)</p>\n\n<p>For future reference, here is how the memory consumption plot looks like when I was trying to use plain <code>DataLoader</code> classes with 50K samples per class:</p>\n\n<p><img src=\"https://i.ibb.co/hdhsDX8/plot.png\"></p>\n\n<p>It really changes my understanding of how parallel things work in Python and in training models on huge datasets :D</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 433933,
          "author_name": "Pavel Ostyakov",
          "author_url": "",
          "post_date": "2018-12-05T17:07:53.430000",
          "content": "<p>@devforfu,</p>\n\n<p>~20 GB RAM is enough for reading all the simplified train data using a trick with Pyro4</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 433939,
          "author_name": "Ilia Zaitsev",
          "author_url": "",
          "post_date": "2018-12-05T17:14:05.623000",
          "content": "<p>@PavelOstyakov That's cool! Thank you for sharing the insight about this library. I guess that Data Science competitions still not that straightforward as it could be thought =)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 433960,
          "author_name": "Sam Ariabod",
          "author_url": "",
          "post_date": "2018-12-05T17:37:49.850000",
          "content": "<p>@devforfu I ran into the same issue using fast.ai and setting the num_workers to 0 did work as bandaid but wasn't really an option given the time left in the competition. My epochs went from 15 hours to over 40. I have ~200g of ram and noticed at some point the system would start to use the pagefile then eventually die. I did find this issue on the Pytorch forums which may be related. <a href=\"https://discuss.pytorch.org/t/dataloader-always-failed-with-interrupted-system-call/30242\">https://discuss.pytorch.org/t/dataloader-always-failed-with-interrupted-system-call/30242</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 434115,
          "author_name": "[he.ai]soulmachine",
          "author_url": "",
          "post_date": "2018-12-05T22:52:22.137000",
          "content": "<p>Hi Pavel, I'm new to Pyro4, what does it do? For example, I need an immutable, efficient dict sharable between multiple processes, for example, I load all training data into dict,  <code>key_id</code> as key,  <code>strokes</code> as values. Can Pyro4 do this for me?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 434188,
          "author_name": "Ilia Zaitsev",
          "author_url": "",
          "post_date": "2018-12-06T03:02:30.177000",
          "content": "<p>@SamAriabod Yeah, you're right, num_workers zero is a solution but it really drops the performance if you have a lot of cores. Same for me, the training time dropped drastically with my i7-6800K. Interesting that the error you've linked seems to have even different nature from what I have though still related to multiprocessing stuff.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 436463,
          "author_name": "[he.ai]soulmachine",
          "author_url": "",
          "post_date": "2018-12-10T11:23:24.370000",
          "content": "<p>Hi Pavel, what's the trick to load all simplified training data to Pyro4? I tried to load all training data into a dict inside Pyro4, but the memory consumption is too huge, even I have 52GB RAM it still not enough, my dict structure is <code>key_id -&gt; (strokes, y)</code>. </p>\n\n<p>The following is my Python code of <code>pyro4-server.py</code>:</p>\n\n<p>```python\nfrom <strong>future</strong> import print_function\nimport gc\nimport os\nimport Pyro4\nimport pandas as pd\nimport pyarrow.parquet\nfrom tqdm import tqdm\nimport time</p>\n\n<p>NUM_FOLDS = 100</p>\n\n<p>@Pyro4.expose\n@Pyro4.behavior(instance_mode=\"single\")\nclass QuickDrawData(object):\n    def <strong>init</strong>(self, file_path='/data/doodle/quickdraw.hdf5'):\n        self.key_id_to_strokes = {}</p>\n\n<pre><code>    if file_path.endswith('.hdf5'):\n        hdf_store = pd.HDFStore(file_path, mode='r')\n        assert len(hdf_store.keys()) == NUM_FOLDS\n    elif file_path.endswith('.parquet'):\n        parquet_file = pyarrow.parquet.ParquetFile(file_path)\n        assert parquet_file.num_row_groups == NUM_FOLDS\n\n    begin_time = time.time()\n    for k in tqdm(range(NUM_FOLDS)):\n        if file_path.endswith('.hdf5'):\n            df = hdf_store.get(key='fold_{0:02d}'.format(k))\n        elif file_path.endswith('.parquet'):\n            df = parquet_file.read_row_group(k, columns=['key_id', 'drawing', 'y']).to_pandas()\n\n        for index, row in df.iterrows():\n            key_id = row['key_id']\n            strokes = row['drawing']\n            y = row['y']\n            self.key_id_to_strokes[key_id] = (strokes, y)\n        del df\n        gc.collect()\n\n    if file_path.endswith('.hdf5'):\n        hdf_store.close()\n    print('Took {}s to read {}'.format(int(time.time()-begin_time), file_path))\n\ndef get(self, key_id):\n    return self.key_id_to_strokes[key_id]\n</code></pre>\n\n<p>def main():\n    Pyro4.Daemon.serveSimple(\n            {\n                QuickDrawData: \"quickdraw.data\"\n            },\n            ns = True)</p>\n\n<p>if <strong>name</strong>==\"<strong>main</strong>\":\n    main()\n```</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 433641,
      "author_name": "jeandebleau",
      "author_url": "",
      "post_date": "2018-12-05T09:33:31.593000",
      "content": "<p>Excellent work and congratulations ! It is an Awesome work and very good systematic approach.</p>\n\n<p>I reached 0.935 with a single recurrent network that has the following structure:\n1d time conv. (Kernel size 3)\n1d time conv (kernel size 5)\nBi directional gru with stabilization layers</p>\n\n<p>These blocks are concatenated 3 times with increasing feature size. Last layer takes the last sequence element features and fed it to a dense layer.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 433643,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-05T09:38:26.910000",
          "content": "<p>thanks for sharing!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 438254,
      "author_name": "Douglas Baldwin",
      "author_url": "",
      "post_date": "2018-12-13T11:07:56.720000",
      "content": "<p>Impressive...</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 438253,
      "author_name": "AbdurRafae",
      "author_url": "",
      "post_date": "2018-12-13T10:59:43.427000",
      "content": "<p>Congratutlations @Pravel Plestov .\nJust a query how much Training Top 3 accuracy, loss did you get on your best model?</p>\n\n<p>Trying something different so wanted a benchmark to compare to.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 437371,
      "author_name": "David Peletz",
      "author_url": "",
      "post_date": "2018-12-11T19:29:33.257000",
      "content": "<p>Congratulations!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 437228,
      "author_name": "Grupper",
      "author_url": "",
      "post_date": "2018-12-11T15:24:43.870000",
      "content": "<p>Who runs this bich?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 437214,
      "author_name": "Sebastien Merchez",
      "author_url": "",
      "post_date": "2018-12-11T14:58:35.630000",
      "content": "<p>Congrats ! You ran all of this on Kaggle Kernel or your own servers ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 437351,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-11T18:57:08.737000",
          "content": "<p>on our own of course</p>",
          "votes": -1,
          "replies": []
        }
      ]
    },
    {
      "id": 437136,
      "author_name": "Fouzi Takelait",
      "author_url": "",
      "post_date": "2018-12-11T12:25:34.633000",
      "content": "<p>How someone can get 100% in accurecy?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 437352,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-11T18:58:00.037000",
          "content": "<p>due to the noise in the data, it is highly unlikely</p>",
          "votes": -1,
          "replies": []
        }
      ]
    },
    {
      "id": 436699,
      "author_name": "Romtein Rostami",
      "author_url": "",
      "post_date": "2018-12-10T19:13:17.300000",
      "content": "<p>This is awesome, congratulations and thank you for the write up </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 436541,
      "author_name": "Sani Kamal",
      "author_url": "",
      "post_date": "2018-12-10T13:58:57.927000",
      "content": "<p>Congrat your team, Many thanks for your sharing!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 436419,
      "author_name": "小宁china",
      "author_url": "",
      "post_date": "2018-12-10T09:08:18.967000",
      "content": "<p>make firend</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 435989,
      "author_name": "jhajhria",
      "author_url": "",
      "post_date": "2018-12-09T08:19:54.960000",
      "content": "<p>Thank you so much for sharing this knowledge with us. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 435670,
      "author_name": "Atanas Atanasov",
      "author_url": "",
      "post_date": "2018-12-08T14:05:56.300000",
      "content": "<p>Learned so much! Amazing approaches, congratulations and thanks for sharing! </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 435162,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-07T15:51:09.797000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 434884,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-07T05:18:26.450000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 434795,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-07T00:45:39.987000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 434766,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-06T22:56:04.490000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 434748,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-06T21:59:05.627000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 434393,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-06T10:10:20.020000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 434331,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-06T08:03:08.810000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 434456,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-06T12:32:37.310000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 438918,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-14T12:12:08.743000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 434297,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-06T07:03:53.710000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 434301,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-06T07:11:35.497000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 434322,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-06T07:41:10.897000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 434288,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-06T06:49:07.787000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 433841,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-05T14:55:31.210000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 433788,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-05T13:33:03.777000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 435448,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-08T04:01:58.453000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 434528,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-06T14:40:25.380000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 434317,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-06T07:39:30.793000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 434251,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-06T05:25:32.490000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 433733,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-05T12:28:11.500000",
      "content": "",
      "votes": 8,
      "replies": []
    },
    {
      "id": 455803,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-14T16:01:11.073000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 433681,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-05T10:43:50.423000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 433651,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-05T09:47:11.157000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 437920,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-12T19:00:04.887000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 435302,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-07T20:20:58.050000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 438259,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-13T11:25:14.827000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 437844,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-12T15:54:27.560000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 437728,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-12T11:34:17.863000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 437315,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-11T17:48:53.727000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 436255,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-10T01:56:27.560000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 436097,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-09T15:12:34.323000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 435931,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-09T05:09:39.770000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 435640,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-08T12:44:14.350000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 435298,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-07T20:11:25.347000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 435207,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-07T17:14:04.617000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 434540,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-06T14:57:29.400000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 434252,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-06T05:26:24.243000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 434215,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-06T03:57:25.283000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 434176,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-06T02:31:31.107000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "433596": "Big thanks to Google for hosting this flawless competition and collecting such a great dataset. I'm also very excited to become top-5 in overall user ranking and even more excited for my teammate [Pavel Ostyakov][1] who got his second 1st place in a row! \n\nCNN\n---\n\nFirst of all, Pavel did what he does best - trained a bunch of pytorch classification models. Here is the list of architectures: resnet18, resnet34, resnet50, resnet101, resnet152, resnext50, resnext101, densenet121, densenet201, vgg11, pnasnet, incresnet, polynet, nasnetmobile, senet154, seresnet50, seresnext50, seresnext101. \n\nOne and three channels preprocessing were used as well as different image sizes starting from 112 and up to 256. The best model got 0.946 score, in total there were around 40 models. However, the gold could be achieved with a single model.\n\nRNN\n---\nI trained a couple of LSTM models based on the [best public kernel][3]. Tweaked the architecture a bit, got rid of dropouts and achieved 0.893 score. Would love to hear in the comments how you got better results.\n\nLightGBM\n--------\n\nHow do you ensemble models with too many classes? This issue has been already resolved during [Cdiscount’s Image Classification Challenge][4]. The idea is the following: for each sample and for each model you collect top 10 probabilities with the labels, then convert them into 10 samples with the binary outcome - whether this is a correct label or not (9 negative examples + 1 positive). It's easy to feed such a dataset to any booster because the number of features will be small (equal to the number of models). On top of that, I also added some time-specific features. The most significant was maximum timestamp from the raw representations of the strokes.\n\nSecret sauce (aka \"щепотка табака\")\n-----\nAs it was mentioned by [Heng CherKeng][5] a month ago [classes in a test set were equally distributed][6]. It was a very important clue which seemed to be lost in the depths of the forum. I also did not see this comment but arrived at the same conclusion by noting that (112199+1)/340=330 (number of samples in the test set plus one is divisible by the number of classes). Knowing the structure of the test set gave us an average boost of 0.7% for every model. \n\nThe algorithm behind postprocessing is the following: for the most popular class decrease all the probabilities iteratively by the same small value until it is no longer the most popular, repeat this procedure until all classes become equal. This technique was also used in [one of the previous competitions][7] (see github link for the code).\n\nBlending\n-----\n\nAfter struggling for a week and producing 17 different balanced submits Pavel left me with the 5 last attempts to improve our public score of 0.956. I used [this public ensembling kernel][9] and scored 0.957 after the first attempt. Changing weights from 5-i to 1/(i+1) gave us a slight additional boost (it mimics map3 weights) and the 1st place. \n\nData\n----\n\nWe used 34000 random samples as the overall holdout set and 1 mln samples for building second layer models. All first layer models were trained on 49 mln simplified samples. Raw data features were only added to LightGBM model.\n\nKey takeaways\n-------------\n\n - Read forum carefully, especially when [Heng CherKeng][10] is present\n - Study past solutions from similar competitions\n\n\n  [1]: https://www.kaggle.com/pavelost \"Pavel\"\n  [2]: https://www.kaggle.com/pavelost \"Pavel\"\n  [3]: https://www.kaggle.com/huyenvyvy/bidirectional-lstm-using-data-generator-lb-0-825\n  [4]: https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/45733\n  [5]: https://www.kaggle.com/hengck23\n  [6]: https://www.kaggle.com/c/quickdraw-doodle-recognition/discussion/70540#416772\n  [7]: https://www.kaggle.com/c/sp-society-camera-model-identification/discussion/49334\n  [8]: https://www.kaggle.com/pavelost \"Pavel\"\n  [9]: https://www.kaggle.com/paulorzp/ensemble-weighted-voting\n  [10]: https://www.kaggle.com/hengck23",
    "433717": "@Pavel Pleskov\n\nThanks for the solution and congrats for winning the first place! \n\n\"I also did not see this comment but arrived at the same conclusion by noting that (112199+1)/340=330 (number of samples in the test set plus one is divisible by the number of classes). Knowing the structure of the test set gave us an average boost of 0.7% for every model.\"\n\nWhat is the estimated score of your final solution if this trick of test class balancing is \"not applied\"? Would it be around  of 0.950?",
    "433618": "Congrats to your team and Thank you for these valuable insights! :-)\n",
    "435273": "I'm glad my script ( www.kaggle.com/paulorzp/ensemble-weighted-voting )  was useful. Congratulations!",
    "434135": "Great job! It is true that \"the devil is in the details.\"",
    "433915": "just trying to pass novice with a comment xD \n",
    "433700": "Congrats Pavel and Pavel  for both your winning and current kaggle ranking .  Thanks for sharing such great insights. \n\n&gt; I trained a couple of LSTM models based on the best public kernel.\n&gt; Tweaked the architecture a bit, got rid of dropouts and achieved 0.893\n&gt; score. Would love to hear in the comments how you got better results.\n\n\nThe LSTM model we used was built by my teamate @Luudactam.  It scored 0.936 on LB",
    "437162": "great",
    "434180": "Fantastic write up - clear, succinct, and interesting.  Will be taking some of this stuff to the whales competition.",
    "434123": "Congrats Pavel and Pavel. Thanks for sharing.\n\nI am still confused about the ensemble. If I had 5 models, for each sample I got 50 pobabilities and labels where each model had 10. Then how did I use them to construct 9 negative examples and 1 positive example?",
    "433992": "Congratulations, and thanks for sharing these important ideas.",
    "433675": "Hi Pavel, Congrats! I have a few questions:\n\n1. How do you encode temporal information into images? Is it similar to Beluga's `draw_cv2()` in this notebook, https://www.kaggle.com/gaborfodor/greyscale-mobilenet-lb-0-892 ?\n2. Since you have tried so many models, could you list all the models  you used in the ensemble phase?\n3. With such a big dataset, do you generate training data on the fly, like the Beluga's notebook above(`image_generator_xd ()`)? Or do you preprocess data and save them into TFRecords for later usage?",
    "433661": "Congrats Pavel and Pavel : )\nOne question, you mentioned both LGBM method and the blend kernel, how do these methods been used together?",
    "433639": "\" for each sample and for each model you collect top 10 probabilities with the labels **than** convert them into 10 sample with the binary outcome \" \nIs it than or then?",
    "433638": "It is a great opportunity to learn from your overall idea. Many ideas seem beyond my current knowledge, so I have to slowly digest them. Thanks for sharing and of course congratulation!!!",
    "433617": "Thank you for sharing famous 'щепотка табака' trick",
    "434578": "A huge congrats Pavel &amp; Pavel!",
    "433800": "Did you have multithreading issues with PyTorch? It seems that on my machine all the available RAM is consumed during a single training epoch when `num_workers&gt;0` in `DataLoader` classes. We're discussing this issue on fastai forums and can't figure out if the problem comes from nightly version of PyTorch, fastai library or somewhere else. \n\nThough probably you had a lot of RAM if all the data was fit into the memory. Having a lot of hardware seems to be very helpful with this competition =)",
    "433641": "Excellent work and congratulations ! It is an Awesome work and very good systematic approach.\n\nI reached 0.935 with a single recurrent network that has the following structure:\n1d time conv. (Kernel size 3)\n1d time conv (kernel size 5)\nBi directional gru with stabilization layers\n\nThese blocks are concatenated 3 times with increasing feature size. Last layer takes the last sequence element features and fed it to a dense layer.",
    "438254": "Impressive...",
    "438253": "Congratutlations @Pravel Plestov .\nJust a query how much Training Top 3 accuracy, loss did you get on your best model?\n\nTrying something different so wanted a benchmark to compare to.",
    "437371": "Congratulations!",
    "437228": "Who runs this bich?",
    "437214": "Congrats ! You ran all of this on Kaggle Kernel or your own servers ?",
    "437136": "How someone can get 100% in accurecy?",
    "436699": "This is awesome, congratulations and thank you for the write up ",
    "436541": "Congrat your team, Many thanks for your sharing!",
    "436419": "make firend",
    "435989": "Thank you so much for sharing this knowledge with us. ",
    "435670": "Learned so much! Amazing approaches, congratulations and thanks for sharing! ",
    "435162": "Congrats",
    "434884": "Congrats Pavel and Pavel. Thank you for sharing.\nCan you share the models you have trained? I want to use this model to extract some of the features of the data.\nthanks!",
    "434795": "Great write-up, congrats!",
    "434766": "Congratulations! ",
    "434748": "Congrats Pavel and Pavel",
    "434393": "Congratz, and thanks for sharing ! I definitely enjoyed reading this !",
    "434331": "Congrats to your team and Thank you for these valuable insights &lt;3. BTW, what is the best model that got 0.946 score ? How big your batch size and is and how much memory did you use to train it ?",
    "434297": "Congratulations! How did you choose your hyper-parameters for your LightGBM model? ",
    "434288": "Thanks for sharing. Would you please explain your LGBM method in detail? How did you get training data for \nLGBM? How did you do in the inference stage? Thanks.",
    "433841": "@Pavel Pleskov\n\nExcellent 1st place and excellent information about your *щепотка табака*, among other stuff. \n\nCongratulations !",
    "433788": "Congrats for your win and great learning for all of us from your notes !!",
    "435448": "",
    "434528": "",
    "434317": "",
    "434251": "",
    "433733": "",
    "455803": "Congrats and thanks for the solution",
    "433681": "Congratulations and thanks for sharing.",
    "433651": "Congratulations and thanks for sharing.",
    "437920": "Thanks!",
    "435302": "Thanks for sharing!",
    "438259": "Thanks for sharing",
    "437844": "Good job, thanks for sharing!",
    "437728": "Congratulations and thanks for sharing.",
    "437315": "Congratulations! Thanks for sharing!",
    "436255": "Thank you for sharing!",
    "436097": "Congrats\nand many thanks for sharing :D",
    "435931": "Thanks for sharing!",
    "435640": "Congratulations and thanks for sharing.",
    "435298": "Thanks for sharing!",
    "435207": "Great write up, thanks for sharing!",
    "434540": "Thanks for sharing",
    "434252": "Congrats!! Thanks for sharing!",
    "434215": "Congratulations and thanks for sharing!",
    "434176": "Congratulations and thanks for sharing."
  }
}