{
  "id": 73712,
  "title": "First Kaggle Competition Experience (Team: rm-rf / | Private LB: 385)",
  "url": "/competitions/quickdraw-doodle-recognition/discussion/73712",
  "author_name": "Sanyam Bhutani",
  "post_date": "2018-12-05T04:44:47.270000",
  "votes": 18,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hi Everyone, \nThis was my First ever kaggle competition and these are a few notes for a few beginners like me. (Dear Experts, please excuse this thread if it doesn't have many interesting insights.) </p>\n\n<p>Thanks to GoogleAI team for hosting the competition.\nCongratulations to Team Pablos, <a href=\"/wowfattie\">@wowfattie</a> and Team mgchbot for the Top 3 finishes. </p>\n\n<p>Special Thanks to Master <a href=\"/radek1\">@radek1</a>, Grandmaster <a href=\"/hengck23\">@hengck23</a> and Grandmaster <a href=\"/gaborfodor\">@gaborfodor</a> for the amazing starter pack, the amazing discussions and the great starter kernel!</p>\n\n<p><strong>My Team Rank:</strong> 385 Private,  391 Public. </p>\n\n<p><strong>Team Name:</strong> \"rm-rf /\", with <a href=\"/init927\">@init927</a> my buisness partner. </p>\n\n<p>My First kaggle competition felt like a 100 Mile sprint where you are competing against people on Supercars (GrandMasters with a LOT of experience) while I was running barefoot.</p>\n\n<p>Our approach relied on the starter kernel shared by Master Radek and the Kernel shared by GrandMaster Beluga. The best submission was a \"blend\" of a MobileNet from the Kernel and a ResNet 152 Model trained using fastai library</p>\n\n<p>We had jumped into the competition past mid-way since its launch and personally it was fascinating and completely overwhelming to keep up with the overflow of the ideas-I was surprised that even the Top performers are generous with their ideas and share them publically. </p>\n\n<p>As a First competition, We'd work our way to a submission each night, wake up to have lost the submission by 20 ranks, rework towards a better submission and repeat!</p>\n\n<p>Ideas that worked:</p>\n\n<ul>\n<li><p>Training on 1% of the data with 256 image size, fine-tuning to 5% of the data with 128 image size, fine-tuning to 20% of the data with 64 image size. </p>\n\n<ul><li>ResNet 18 &lt; ResNet 34 &lt; ResNet 50 &lt; ResNet 152 showed a consistent increase in performance (Pre-Trained models, fine-tuned using fastai)</li></ul></li>\n</ul>\n\n<p>Compute: Starting out, I wasn't sure if I had enough compute for the competition, turns we had more than sufficient-20 (10+10) kaggle kernels for small experiements, a 1070 based \"laptop\" for bigger experiments and p3 instances on AWS for extra experiementation, while documenting our approach with Google Sheets.</p>\n\n<p>Personally, It was an amazing learning experience, I learnt much more than I had ever learnt via a MOOC. I'd definitely love to participate in more competition and slowly work my way upwards. </p>\n\n<p>In the end, Our best submission was limited by our experience to make a better submission with our hardware setup. If anyone is still not sure about taking part in a kaggle comp, I'd say just jump in. \nMake your first submission, see it fall down on the LB and try to keep up!</p>\n\n<p>PS: If anyone has any suggestions towards our approach or any comments on how to have better approached the competition, I'd be very thankful.</p>",
  "messages": [
    {
      "id": 433439,
      "postDate": "2018-12-05T04:44:47.270Z",
      "content": "<p>Hi Everyone, \nThis was my First ever kaggle competition and these are a few notes for a few beginners like me. (Dear Experts, please excuse this thread if it doesn't have many interesting insights.) </p>\n\n<p>Thanks to GoogleAI team for hosting the competition.\nCongratulations to Team Pablos, <a href=\"/wowfattie\">@wowfattie</a> and Team mgchbot for the Top 3 finishes. </p>\n\n<p>Special Thanks to Master <a href=\"/radek1\">@radek1</a>, Grandmaster <a href=\"/hengck23\">@hengck23</a> and Grandmaster <a href=\"/gaborfodor\">@gaborfodor</a> for the amazing starter pack, the amazing discussions and the great starter kernel!</p>\n\n<p><strong>My Team Rank:</strong> 385 Private,  391 Public. </p>\n\n<p><strong>Team Name:</strong> \"rm-rf /\", with <a href=\"/init927\">@init927</a> my buisness partner. </p>\n\n<p>My First kaggle competition felt like a 100 Mile sprint where you are competing against people on Supercars (GrandMasters with a LOT of experience) while I was running barefoot.</p>\n\n<p>Our approach relied on the starter kernel shared by Master Radek and the Kernel shared by GrandMaster Beluga. The best submission was a \"blend\" of a MobileNet from the Kernel and a ResNet 152 Model trained using fastai library</p>\n\n<p>We had jumped into the competition past mid-way since its launch and personally it was fascinating and completely overwhelming to keep up with the overflow of the ideas-I was surprised that even the Top performers are generous with their ideas and share them publically. </p>\n\n<p>As a First competition, We'd work our way to a submission each night, wake up to have lost the submission by 20 ranks, rework towards a better submission and repeat!</p>\n\n<p>Ideas that worked:</p>\n\n<ul>\n<li><p>Training on 1% of the data with 256 image size, fine-tuning to 5% of the data with 128 image size, fine-tuning to 20% of the data with 64 image size. </p>\n\n<ul><li>ResNet 18 &lt; ResNet 34 &lt; ResNet 50 &lt; ResNet 152 showed a consistent increase in performance (Pre-Trained models, fine-tuned using fastai)</li></ul></li>\n</ul>\n\n<p>Compute: Starting out, I wasn't sure if I had enough compute for the competition, turns we had more than sufficient-20 (10+10) kaggle kernels for small experiements, a 1070 based \"laptop\" for bigger experiments and p3 instances on AWS for extra experiementation, while documenting our approach with Google Sheets.</p>\n\n<p>Personally, It was an amazing learning experience, I learnt much more than I had ever learnt via a MOOC. I'd definitely love to participate in more competition and slowly work my way upwards. </p>\n\n<p>In the end, Our best submission was limited by our experience to make a better submission with our hardware setup. If anyone is still not sure about taking part in a kaggle comp, I'd say just jump in. \nMake your first submission, see it fall down on the LB and try to keep up!</p>\n\n<p>PS: If anyone has any suggestions towards our approach or any comments on how to have better approached the competition, I'd be very thankful.</p>",
      "rawMarkdown": "Hi Everyone, \nThis was my First ever kaggle competition and these are a few notes for a few beginners like me. (Dear Experts, please excuse this thread if it doesn't have many interesting insights.) \n\nThanks to GoogleAI team for hosting the competition.\nCongratulations to Team Pablos, @wowfattie and Team mgchbot for the Top 3 finishes. \n\nSpecial Thanks to Master @radek1, Grandmaster @hengck23 and Grandmaster @gaborfodor for the amazing starter pack, the amazing discussions and the great starter kernel!\n\n**My Team Rank:** 385 Private,  391 Public. \n\n**Team Name:** \"rm-rf /\", with @init927 my buisness partner. \n\nMy First kaggle competition felt like a 100 Mile sprint where you are competing against people on Supercars (GrandMasters with a LOT of experience) while I was running barefoot.\n\nOur approach relied on the starter kernel shared by Master Radek and the Kernel shared by GrandMaster Beluga. The best submission was a \"blend\" of a MobileNet from the Kernel and a ResNet 152 Model trained using fastai library\n\nWe had jumped into the competition past mid-way since its launch and personally it was fascinating and completely overwhelming to keep up with the overflow of the ideas-I was surprised that even the Top performers are generous with their ideas and share them publically. \n\nAs a First competition, We'd work our way to a submission each night, wake up to have lost the submission by 20 ranks, rework towards a better submission and repeat!\n\nIdeas that worked:\n\n - Training on 1% of the data with 256 image size, fine-tuning to 5% of the data with 128 image size, fine-tuning to 20% of the data with 64 image size. \n\n- ResNet 18 &lt; ResNet 34 &lt; ResNet 50 &lt; ResNet 152 showed a consistent increase in performance (Pre-Trained models, fine-tuned using fastai)\n\nCompute: Starting out, I wasn't sure if I had enough compute for the competition, turns we had more than sufficient-20 (10+10) kaggle kernels for small experiements, a 1070 based \"laptop\" for bigger experiments and p3 instances on AWS for extra experiementation, while documenting our approach with Google Sheets.\n\nPersonally, It was an amazing learning experience, I learnt much more than I had ever learnt via a MOOC. I'd definitely love to participate in more competition and slowly work my way upwards. \n\nIn the end, Our best submission was limited by our experience to make a better submission with our hardware setup. If anyone is still not sure about taking part in a kaggle comp, I'd say just jump in. \nMake your first submission, see it fall down on the LB and try to keep up!\n\n\nPS: If anyone has any suggestions towards our approach or any comments on how to have better approached the competition, I'd be very thankful.",
      "votes": 18
    },
    {
      "id": 433485,
      "postDate": "2018-12-05T05:55:54.510Z",
      "content": "<p>Congrat <a href=\"/init27\">@init27</a> for your (and my) great experience!  Nothing is more important than getting start. We can learn and do better each time. According to your writing, next time, we may perhaps be able to ride a horse against those supercars! :D</p>",
      "rawMarkdown": "Congrat @init27 for your (and my) great experience!  Nothing is more important than getting start. We can learn and do better each time. According to your writing, next time, we may perhaps be able to ride a horse against those supercars! :D",
      "votes": 3,
      "replies": [
        {
          "id": 433558,
          "postDate": "2018-12-05T07:19:24.350Z",
          "content": "<p>Thanks! And yes, I'd definitely love to team up in the future.</p>",
          "rawMarkdown": "Thanks! And yes, I'd definitely love to team up in the future.",
          "votes": 1
        }
      ]
    },
    {
      "id": 434206,
      "postDate": "2018-12-06T03:40:50.780Z",
      "content": "<p>Congratulations Dear !!!\nLooking forward to team up  and let's drive An Autopilot Tesla's Car against Those Falcon's !</p>",
      "rawMarkdown": "Congratulations Dear !!!\nLooking forward to team up  and let's drive An Autopilot Tesla's Car against Those Falcon's !",
      "votes": 1,
      "replies": [
        {
          "id": 434533,
          "postDate": "2018-12-06T14:47:48.540Z",
          "content": "<p>Thanks and that's a very motivating analogy :)</p>",
          "rawMarkdown": "Thanks and that's a very motivating analogy :)"
        }
      ]
    },
    {
      "id": 433605,
      "postDate": "2018-12-05T08:11:22.137Z",
      "content": "<p>You can train all data, neither 20% nor 1% or them, and you will make extremely good improvement. \nFor me, I get 0.915 on Public LB when using 20% data, but get 0.931 on Public LB when using all data</p>",
      "rawMarkdown": "You can train all data, neither 20% nor 1% or them, and you will make extremely good improvement. \nFor me, I get 0.915 on Public LB when using 20% data, but get 0.931 on Public LB when using all data",
      "votes": 1,
      "replies": [
        {
          "id": 433684,
          "postDate": "2018-12-05T10:47:09.623Z",
          "content": "<p>Thanks. I was aware of that but my approach relied on the \"draw\" function to first draw the images, which in itself would take a lot of hours to generate and followed by training which would be even for longer duration on my GTX-1070. \nCould you share your approach? How did you train on the 100% of the data?</p>",
          "rawMarkdown": "Thanks. I was aware of that but my approach relied on the \"draw\" function to first draw the images, which in itself would take a lot of hours to generate and followed by training which would be even for longer duration on my GTX-1070. \nCould you share your approach? How did you train on the 100% of the data?",
          "votes": 1
        },
        {
          "id": 434214,
          "postDate": "2018-12-06T03:56:42.707Z",
          "content": "<p>Well, I write a callback here:\n<a href=\"https://www.kaggle.com/c/quickdraw-doodle-recognition/discussion/73704\">https://www.kaggle.com/c/quickdraw-doodle-recognition/discussion/73704</a>\nBeluga's Kernel is a good example! I just learn from him:\n<a href=\"https://www.kaggle.com/gaborfodor/greyscale-mobilenet-lb-0-892\">https://www.kaggle.com/gaborfodor/greyscale-mobilenet-lb-0-892</a> </p>",
          "rawMarkdown": "Well, I write a callback here:\nhttps://www.kaggle.com/c/quickdraw-doodle-recognition/discussion/73704\nBeluga's Kernel is a good example! I just learn from him:\nhttps://www.kaggle.com/gaborfodor/greyscale-mobilenet-lb-0-892 ",
          "votes": 1
        },
        {
          "id": 434508,
          "postDate": "2018-12-06T14:15:42.543Z",
          "rawMarkdown": ""
        }
      ]
    },
    {
      "id": 437042,
      "postDate": "2018-12-11T09:40:57.040Z",
      "content": "<p>Congrats, it is indeed a great achievement on your first competition especially with such a large dataset. Wish you all the best for your next competitions</p>",
      "rawMarkdown": "Congrats, it is indeed a great achievement on your first competition especially with such a large dataset. Wish you all the best for your next competitions"
    }
  ],
  "comments": [
    {
      "id": 433485,
      "author_name": "Neuron Engineer",
      "author_url": "",
      "post_date": "2018-12-05T05:55:54.510000",
      "content": "<p>Congrat <a href=\"/init27\">@init27</a> for your (and my) great experience!  Nothing is more important than getting start. We can learn and do better each time. According to your writing, next time, we may perhaps be able to ride a horse against those supercars! :D</p>",
      "votes": 3,
      "replies": [
        {
          "id": 433558,
          "author_name": "Sanyam Bhutani",
          "author_url": "",
          "post_date": "2018-12-05T07:19:24.350000",
          "content": "<p>Thanks! And yes, I'd definitely love to team up in the future.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 434206,
      "author_name": "Aditya Soni",
      "author_url": "",
      "post_date": "2018-12-06T03:40:50.780000",
      "content": "<p>Congratulations Dear !!!\nLooking forward to team up  and let's drive An Autopilot Tesla's Car against Those Falcon's !</p>",
      "votes": 1,
      "replies": [
        {
          "id": 434533,
          "author_name": "Sanyam Bhutani",
          "author_url": "",
          "post_date": "2018-12-06T14:47:48.540000",
          "content": "<p>Thanks and that's a very motivating analogy :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 433605,
      "author_name": "Dilapsky Lee",
      "author_url": "",
      "post_date": "2018-12-05T08:11:22.137000",
      "content": "<p>You can train all data, neither 20% nor 1% or them, and you will make extremely good improvement. \nFor me, I get 0.915 on Public LB when using 20% data, but get 0.931 on Public LB when using all data</p>",
      "votes": 1,
      "replies": [
        {
          "id": 433684,
          "author_name": "Sanyam Bhutani",
          "author_url": "",
          "post_date": "2018-12-05T10:47:09.623000",
          "content": "<p>Thanks. I was aware of that but my approach relied on the \"draw\" function to first draw the images, which in itself would take a lot of hours to generate and followed by training which would be even for longer duration on my GTX-1070. \nCould you share your approach? How did you train on the 100% of the data?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 434214,
          "author_name": "Dilapsky Lee",
          "author_url": "",
          "post_date": "2018-12-06T03:56:42.707000",
          "content": "<p>Well, I write a callback here:\n<a href=\"https://www.kaggle.com/c/quickdraw-doodle-recognition/discussion/73704\">https://www.kaggle.com/c/quickdraw-doodle-recognition/discussion/73704</a>\nBeluga's Kernel is a good example! I just learn from him:\n<a href=\"https://www.kaggle.com/gaborfodor/greyscale-mobilenet-lb-0-892\">https://www.kaggle.com/gaborfodor/greyscale-mobilenet-lb-0-892</a> </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 434508,
          "author_name": "Kushal Sharma",
          "author_url": "",
          "post_date": "2018-12-06T14:15:42.543000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 437042,
      "author_name": "Karthick",
      "author_url": "",
      "post_date": "2018-12-11T09:40:57.040000",
      "content": "<p>Congrats, it is indeed a great achievement on your first competition especially with such a large dataset. Wish you all the best for your next competitions</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "433439": "Hi Everyone, \nThis was my First ever kaggle competition and these are a few notes for a few beginners like me. (Dear Experts, please excuse this thread if it doesn't have many interesting insights.) \n\nThanks to GoogleAI team for hosting the competition.\nCongratulations to Team Pablos, @wowfattie and Team mgchbot for the Top 3 finishes. \n\nSpecial Thanks to Master @radek1, Grandmaster @hengck23 and Grandmaster @gaborfodor for the amazing starter pack, the amazing discussions and the great starter kernel!\n\n**My Team Rank:** 385 Private,  391 Public. \n\n**Team Name:** \"rm-rf /\", with @init927 my buisness partner. \n\nMy First kaggle competition felt like a 100 Mile sprint where you are competing against people on Supercars (GrandMasters with a LOT of experience) while I was running barefoot.\n\nOur approach relied on the starter kernel shared by Master Radek and the Kernel shared by GrandMaster Beluga. The best submission was a \"blend\" of a MobileNet from the Kernel and a ResNet 152 Model trained using fastai library\n\nWe had jumped into the competition past mid-way since its launch and personally it was fascinating and completely overwhelming to keep up with the overflow of the ideas-I was surprised that even the Top performers are generous with their ideas and share them publically. \n\nAs a First competition, We'd work our way to a submission each night, wake up to have lost the submission by 20 ranks, rework towards a better submission and repeat!\n\nIdeas that worked:\n\n - Training on 1% of the data with 256 image size, fine-tuning to 5% of the data with 128 image size, fine-tuning to 20% of the data with 64 image size. \n\n- ResNet 18 &lt; ResNet 34 &lt; ResNet 50 &lt; ResNet 152 showed a consistent increase in performance (Pre-Trained models, fine-tuned using fastai)\n\nCompute: Starting out, I wasn't sure if I had enough compute for the competition, turns we had more than sufficient-20 (10+10) kaggle kernels for small experiements, a 1070 based \"laptop\" for bigger experiments and p3 instances on AWS for extra experiementation, while documenting our approach with Google Sheets.\n\nPersonally, It was an amazing learning experience, I learnt much more than I had ever learnt via a MOOC. I'd definitely love to participate in more competition and slowly work my way upwards. \n\nIn the end, Our best submission was limited by our experience to make a better submission with our hardware setup. If anyone is still not sure about taking part in a kaggle comp, I'd say just jump in. \nMake your first submission, see it fall down on the LB and try to keep up!\n\n\nPS: If anyone has any suggestions towards our approach or any comments on how to have better approached the competition, I'd be very thankful.",
    "433485": "Congrat @init27 for your (and my) great experience!  Nothing is more important than getting start. We can learn and do better each time. According to your writing, next time, we may perhaps be able to ride a horse against those supercars! :D",
    "434206": "Congratulations Dear !!!\nLooking forward to team up  and let's drive An Autopilot Tesla's Car against Those Falcon's !",
    "433605": "You can train all data, neither 20% nor 1% or them, and you will make extremely good improvement. \nFor me, I get 0.915 on Public LB when using 20% data, but get 0.931 on Public LB when using all data",
    "437042": "Congrats, it is indeed a great achievement on your first competition especially with such a large dataset. Wish you all the best for your next competitions"
  }
}