{
  "id": 73704,
  "title": "My First Medal Callback: From Novice to 106th place",
  "url": "/competitions/quickdraw-doodle-recognition/discussion/73704",
  "author_name": "Dilapsky Lee",
  "post_date": "2018-12-05T01:26:40.095000",
  "votes": 23,
  "comment_count": 12,
  "views": 0,
  "content": "<p>I am an amateur whose major is Mechanical Engineering, and this is the first time I enrolled in a Featured Kaggle Competition.\nHere I have some experience to sharing with you guys(good for kaggle beginners):</p>\n\n<ol>\n<li><p>Save your model frequently.\nGoogle Doodle has a large datasets(About 50000000), and you will spend several days in training this datasets. \nFor me, 128*128size with resnet50, 2 epochs takes 3-4 days.\nOne day, I planed to train the data for 12 hours. However, the server shutdown accidently, making me disappointed and exhausted.\nSo save your model frequently, model.save('model.h5') and model.save_weights('weights.h5')</p></li>\n<li><p>Nohup jupyter notebook\nThere will be some network connection error and shutdown your jupyter notebook accidently if you don't use nohup commmand.\nIf you use nohup command, your program still run even if you close the network connection.</p></li>\n<li><p>Get GPU from Kaggle Kernel\nIn discussion, most of you say that you have 1-4 gpus. If you want to try different neutral networks, your limited gpu resource will be an obstacles.\nIn fact you can use Kaggle Kernel gpu, with model.save('model.h5') and model.save_weights('weights.h5') method stated on 1, you can train more networks.</p></li>\n<li><p>Vote and check forked notebook from others\nAt first, I voted and forked beluga's notebook. Then I tried to change the neutral network from mobilenet to resnet50, densenet121, etc.\nWhen the public LB stucked at about 0.915 for several days, I decided to do data augmentation.\nWhen I checked beluga's shuffle-csv notebook, I found pd.read_csv(..., nrows = 34000)!\nI only used 1/5 of the datasets!\nIn previous, I only focused on neutral network improvement.\nI changed it to the total datasets, and the accuracy improved a lot.\nSo after you forked other's notebook, check everything at the beginning!</p></li>\n</ol>\n\n<p>For my last result, I use 128*128 size image, about 1-2 epochs for full dataset, train 2 models: densenet121(batchsize: 170) and resnet50(batchsize: 250).\nsize(validation dataset) = 0.3 * size(full dataset)\nThe best resnet50 model: 0.931 Public LB\nThe best densenet model: 0.934 Public LB and 0.932 Public LB\nEnsemble: 0.934_densenet(weight = 2), 0.932_densenet(weight = 1.8), 0.931_resnet(weight = 1.65) result = 0.937 Public LB\nIt seems that various neutral network models can ensemble a higher accuracy.\nTTA: symmetric deformation of test_simplified.csv, but no help in accuracy.</p>\n\n<p>At last,I am so happy to both win a medal and get help from you guys.:-)</p>",
  "messages": [
    {
      "id": 433329,
      "postDate": "2018-12-05T01:26:40.097Z",
      "content": "<p>I am an amateur whose major is Mechanical Engineering, and this is the first time I enrolled in a Featured Kaggle Competition.\nHere I have some experience to sharing with you guys(good for kaggle beginners):</p>\n\n<ol>\n<li><p>Save your model frequently.\nGoogle Doodle has a large datasets(About 50000000), and you will spend several days in training this datasets. \nFor me, 128*128size with resnet50, 2 epochs takes 3-4 days.\nOne day, I planed to train the data for 12 hours. However, the server shutdown accidently, making me disappointed and exhausted.\nSo save your model frequently, model.save('model.h5') and model.save_weights('weights.h5')</p></li>\n<li><p>Nohup jupyter notebook\nThere will be some network connection error and shutdown your jupyter notebook accidently if you don't use nohup commmand.\nIf you use nohup command, your program still run even if you close the network connection.</p></li>\n<li><p>Get GPU from Kaggle Kernel\nIn discussion, most of you say that you have 1-4 gpus. If you want to try different neutral networks, your limited gpu resource will be an obstacles.\nIn fact you can use Kaggle Kernel gpu, with model.save('model.h5') and model.save_weights('weights.h5') method stated on 1, you can train more networks.</p></li>\n<li><p>Vote and check forked notebook from others\nAt first, I voted and forked beluga's notebook. Then I tried to change the neutral network from mobilenet to resnet50, densenet121, etc.\nWhen the public LB stucked at about 0.915 for several days, I decided to do data augmentation.\nWhen I checked beluga's shuffle-csv notebook, I found pd.read_csv(..., nrows = 34000)!\nI only used 1/5 of the datasets!\nIn previous, I only focused on neutral network improvement.\nI changed it to the total datasets, and the accuracy improved a lot.\nSo after you forked other's notebook, check everything at the beginning!</p></li>\n</ol>\n\n<p>For my last result, I use 128*128 size image, about 1-2 epochs for full dataset, train 2 models: densenet121(batchsize: 170) and resnet50(batchsize: 250).\nsize(validation dataset) = 0.3 * size(full dataset)\nThe best resnet50 model: 0.931 Public LB\nThe best densenet model: 0.934 Public LB and 0.932 Public LB\nEnsemble: 0.934_densenet(weight = 2), 0.932_densenet(weight = 1.8), 0.931_resnet(weight = 1.65) result = 0.937 Public LB\nIt seems that various neutral network models can ensemble a higher accuracy.\nTTA: symmetric deformation of test_simplified.csv, but no help in accuracy.</p>\n\n<p>At last,I am so happy to both win a medal and get help from you guys.:-)</p>",
      "rawMarkdown": "I am an amateur whose major is Mechanical Engineering, and this is the first time I enrolled in a Featured Kaggle Competition.\nHere I have some experience to sharing with you guys(good for kaggle beginners):\n\n1. Save your model frequently.\n  Google Doodle has a large datasets(About 50000000), and you will spend several days in training this datasets. \n  For me, 128*128size with resnet50, 2 epochs takes 3-4 days.\n  One day, I planed to train the data for 12 hours. However, the server shutdown accidently, making me disappointed and exhausted.\n  So save your model frequently, model.save('model.h5') and model.save_weights('weights.h5')\n  \n2. Nohup jupyter notebook\n  There will be some network connection error and shutdown your jupyter notebook accidently if you don't use nohup commmand.\n  If you use nohup command, your program still run even if you close the network connection.\n  \n3. Get GPU from Kaggle Kernel\n  In discussion, most of you say that you have 1-4 gpus. If you want to try different neutral networks, your limited gpu resource will be an obstacles.\n  In fact you can use Kaggle Kernel gpu, with model.save('model.h5') and model.save_weights('weights.h5') method stated on 1, you can train more networks.\n\n4. Vote and check forked notebook from others\n  At first, I voted and forked beluga's notebook. Then I tried to change the neutral network from mobilenet to resnet50, densenet121, etc.\n  When the public LB stucked at about 0.915 for several days, I decided to do data augmentation.\n  When I checked beluga's shuffle-csv notebook, I found pd.read_csv(..., nrows = 34000)!\n  I only used 1/5 of the datasets!\n  In previous, I only focused on neutral network improvement.\n  I changed it to the total datasets, and the accuracy improved a lot.\n  So after you forked other's notebook, check everything at the beginning!\n  \nFor my last result, I use 128*128 size image, about 1-2 epochs for full dataset, train 2 models: densenet121(batchsize: 170) and resnet50(batchsize: 250).\nsize(validation dataset) = 0.3 * size(full dataset)\nThe best resnet50 model: 0.931 Public LB\nThe best densenet model: 0.934 Public LB and 0.932 Public LB\nEnsemble: 0.934_densenet(weight = 2), 0.932_densenet(weight = 1.8), 0.931_resnet(weight = 1.65) result = 0.937 Public LB\nIt seems that various neutral network models can ensemble a higher accuracy.\nTTA: symmetric deformation of test_simplified.csv, but no help in accuracy.\n\nAt last,I am so happy to both win a medal and get help from you guys.:-)",
      "votes": 23
    },
    {
      "id": 434537,
      "postDate": "2018-12-06T14:53:26.933Z",
      "content": "<p>Congratulations and Thanks for sharing. \nThe mistake that I made was tweaking the MobileNet for quite a few days and then trying DenseNet. However, instead of expermimenting with smaller subsets, I decided to use the complete dataset due to time constraints-the results were not good. \nTwo things I've learnt: Do not enter a competition after its half way through, experiment well before running a full-blown experiment.</p>",
      "rawMarkdown": "Congratulations and Thanks for sharing. \nThe mistake that I made was tweaking the MobileNet for quite a few days and then trying DenseNet. However, instead of expermimenting with smaller subsets, I decided to use the complete dataset due to time constraints-the results were not good. \nTwo things I've learnt: Do not enter a competition after its half way through, experiment well before running a full-blown experiment.",
      "votes": 1,
      "replies": [
        {
          "id": 434794,
          "postDate": "2018-12-07T00:41:29.853Z",
          "content": "<p>Thanks for your experience!\nBefore we run a full-dataset experiment, we usually use small dataset with small batchsize to \"pretrain\" it, and then we use fine-tuning and graduately increase the image size and batch size to make convergence.</p>",
          "rawMarkdown": "Thanks for your experience!\nBefore we run a full-dataset experiment, we usually use small dataset with small batchsize to \"pretrain\" it, and then we use fine-tuning and graduately increase the image size and batch size to make convergence.",
          "votes": 1
        }
      ]
    },
    {
      "id": 433799,
      "postDate": "2018-12-05T13:42:22.047Z",
      "content": "<p>Congratulation on finishing in 106 as a novice, It is indeed inspiring !!</p>",
      "rawMarkdown": "Congratulation on finishing in 106 as a novice, It is indeed inspiring !!",
      "votes": 1,
      "replies": [
        {
          "id": 434152,
          "postDate": "2018-12-06T01:01:27.010Z",
          "content": "<p>Thanks! And wish you can get good grades very soon!</p>",
          "rawMarkdown": "Thanks! And wish you can get good grades very soon!"
        }
      ]
    },
    {
      "id": 433711,
      "postDate": "2018-12-05T11:38:06.713Z",
      "content": "<p>@Dilapsky Lee</p>\n\n<p><em>\"For me, 128x128size with resnet50, 2 epochs takes 3-4 days.\"</em></p>\n\n<p>I understand my modest score : 1 GPU with 4Gb RAM, and never more than ONE night max for training : I need my computer during daytime.</p>\n\n<p>I was also disappointed by the poor results of my GRU's trials.</p>\n\n<p>Anyway, once again very good job for a starter and very good advices above !</p>",
      "rawMarkdown": "@Dilapsky Lee\n\n*\"For me, 128x128size with resnet50, 2 epochs takes 3-4 days.\"*\n\nI understand my modest score : 1 GPU with 4Gb RAM, and never more than ONE night max for training : I need my computer during daytime.\n\nI was also disappointed by the poor results of my GRU's trials.\n\nAnyway, once again very good job for a starter and very good advices above !",
      "votes": 1,
      "replies": [
        {
          "id": 434150,
          "postDate": "2018-12-06T00:59:56.503Z",
          "content": "<p>Thank you for your congratulation!\nImage-based competition do requires GPU hardware, so I recommend you to use kaggle kernel gpu and model.save skill to continuously train your model.  You can also ask if there's anyone wants to form a team.</p>\n\n<p>In addition, there are lots of machine-learning competition in kaggle where gpu is not strictly needed.</p>",
          "rawMarkdown": "Thank you for your congratulation!\nImage-based competition do requires GPU hardware, so I recommend you to use kaggle kernel gpu and model.save skill to continuously train your model.  You can also ask if there's anyone wants to form a team.\n\nIn addition, there are lots of machine-learning competition in kaggle where gpu is not strictly needed.\n"
        }
      ]
    },
    {
      "id": 433498,
      "postDate": "2018-12-05T06:07:39.317Z",
      "content": "<p>Congratulations and thanks for these practical tips. Hope to see you at the top of a leaderboard very soon.</p>",
      "rawMarkdown": "Congratulations and thanks for these practical tips. Hope to see you at the top of a leaderboard very soon.",
      "votes": 1,
      "replies": [
        {
          "id": 433597,
          "postDate": "2018-12-05T08:03:25.510Z",
          "content": "<p>Thanks! And the same with you, too. </p>",
          "rawMarkdown": "Thanks! And the same with you, too. "
        }
      ]
    },
    {
      "id": 433491,
      "postDate": "2018-12-05T06:01:37.583Z",
      "content": "<p>Congrat and thank you for sharing! We have similar experience, in our first week, we just realize that we use only 1/5 of the data as well! Our score is closed, our best single LB is also closed (We got 0.933LB using Xception)</p>",
      "rawMarkdown": "Congrat and thank you for sharing! We have similar experience, in our first week, we just realize that we use only 1/5 of the data as well! Our score is closed, our best single LB is also closed (We got 0.933LB using Xception)",
      "votes": 1,
      "replies": [
        {
          "id": 433595,
          "postDate": "2018-12-05T08:00:33.423Z",
          "content": "<p>Congrat for your great job, too :-D . We can share experience for each other in future kaggle competition!</p>",
          "rawMarkdown": "Congrat for your great job, too :-D . We can share experience for each other in future kaggle competition!"
        }
      ]
    },
    {
      "id": 433516,
      "postDate": "2018-12-05T06:36:07.803Z",
      "content": "<p>Congratulations ! </p>\n\n<p>You now can apply your knowledge to mechanical drawing :)</p>",
      "rawMarkdown": "Congratulations ! \n\nYou now can apply your knowledge to mechanical drawing :)",
      "votes": 2,
      "replies": [
        {
          "id": 433599,
          "postDate": "2018-12-05T08:05:08.857Z",
          "content": "<p>Thanks! That sounds a good idea! But mechanical drawing should follow strict rules, not doodle. :-&gt;</p>",
          "rawMarkdown": "Thanks! That sounds a good idea! But mechanical drawing should follow strict rules, not doodle. :-&gt;"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 434537,
      "author_name": "Sanyam Bhutani",
      "author_url": "",
      "post_date": "2018-12-06T14:53:26.933000",
      "content": "<p>Congratulations and Thanks for sharing. \nThe mistake that I made was tweaking the MobileNet for quite a few days and then trying DenseNet. However, instead of expermimenting with smaller subsets, I decided to use the complete dataset due to time constraints-the results were not good. \nTwo things I've learnt: Do not enter a competition after its half way through, experiment well before running a full-blown experiment.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 434794,
          "author_name": "Dilapsky Lee",
          "author_url": "",
          "post_date": "2018-12-07T00:41:29.853000",
          "content": "<p>Thanks for your experience!\nBefore we run a full-dataset experiment, we usually use small dataset with small batchsize to \"pretrain\" it, and then we use fine-tuning and graduately increase the image size and batch size to make convergence.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 433799,
      "author_name": "Vishy",
      "author_url": "",
      "post_date": "2018-12-05T13:42:22.047000",
      "content": "<p>Congratulation on finishing in 106 as a novice, It is indeed inspiring !!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 434152,
          "author_name": "Dilapsky Lee",
          "author_url": "",
          "post_date": "2018-12-06T01:01:27.010000",
          "content": "<p>Thanks! And wish you can get good grades very soon!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 433711,
      "author_name": "mezoganet",
      "author_url": "",
      "post_date": "2018-12-05T11:38:06.713000",
      "content": "<p>@Dilapsky Lee</p>\n\n<p><em>\"For me, 128x128size with resnet50, 2 epochs takes 3-4 days.\"</em></p>\n\n<p>I understand my modest score : 1 GPU with 4Gb RAM, and never more than ONE night max for training : I need my computer during daytime.</p>\n\n<p>I was also disappointed by the poor results of my GRU's trials.</p>\n\n<p>Anyway, once again very good job for a starter and very good advices above !</p>",
      "votes": 1,
      "replies": [
        {
          "id": 434150,
          "author_name": "Dilapsky Lee",
          "author_url": "",
          "post_date": "2018-12-06T00:59:56.503000",
          "content": "<p>Thank you for your congratulation!\nImage-based competition do requires GPU hardware, so I recommend you to use kaggle kernel gpu and model.save skill to continuously train your model.  You can also ask if there's anyone wants to form a team.</p>\n\n<p>In addition, there are lots of machine-learning competition in kaggle where gpu is not strictly needed.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 433498,
      "author_name": "vbookshelf",
      "author_url": "",
      "post_date": "2018-12-05T06:07:39.317000",
      "content": "<p>Congratulations and thanks for these practical tips. Hope to see you at the top of a leaderboard very soon.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 433597,
          "author_name": "Dilapsky Lee",
          "author_url": "",
          "post_date": "2018-12-05T08:03:25.510000",
          "content": "<p>Thanks! And the same with you, too. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 433491,
      "author_name": "Neuron Engineer",
      "author_url": "",
      "post_date": "2018-12-05T06:01:37.583000",
      "content": "<p>Congrat and thank you for sharing! We have similar experience, in our first week, we just realize that we use only 1/5 of the data as well! Our score is closed, our best single LB is also closed (We got 0.933LB using Xception)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 433595,
          "author_name": "Dilapsky Lee",
          "author_url": "",
          "post_date": "2018-12-05T08:00:33.423000",
          "content": "<p>Congrat for your great job, too :-D . We can share experience for each other in future kaggle competition!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 433516,
      "author_name": "mezoganet",
      "author_url": "",
      "post_date": "2018-12-05T06:36:07.803000",
      "content": "<p>Congratulations ! </p>\n\n<p>You now can apply your knowledge to mechanical drawing :)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 433599,
          "author_name": "Dilapsky Lee",
          "author_url": "",
          "post_date": "2018-12-05T08:05:08.857000",
          "content": "<p>Thanks! That sounds a good idea! But mechanical drawing should follow strict rules, not doodle. :-&gt;</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "433329": "I am an amateur whose major is Mechanical Engineering, and this is the first time I enrolled in a Featured Kaggle Competition.\nHere I have some experience to sharing with you guys(good for kaggle beginners):\n\n1. Save your model frequently.\n  Google Doodle has a large datasets(About 50000000), and you will spend several days in training this datasets. \n  For me, 128*128size with resnet50, 2 epochs takes 3-4 days.\n  One day, I planed to train the data for 12 hours. However, the server shutdown accidently, making me disappointed and exhausted.\n  So save your model frequently, model.save('model.h5') and model.save_weights('weights.h5')\n  \n2. Nohup jupyter notebook\n  There will be some network connection error and shutdown your jupyter notebook accidently if you don't use nohup commmand.\n  If you use nohup command, your program still run even if you close the network connection.\n  \n3. Get GPU from Kaggle Kernel\n  In discussion, most of you say that you have 1-4 gpus. If you want to try different neutral networks, your limited gpu resource will be an obstacles.\n  In fact you can use Kaggle Kernel gpu, with model.save('model.h5') and model.save_weights('weights.h5') method stated on 1, you can train more networks.\n\n4. Vote and check forked notebook from others\n  At first, I voted and forked beluga's notebook. Then I tried to change the neutral network from mobilenet to resnet50, densenet121, etc.\n  When the public LB stucked at about 0.915 for several days, I decided to do data augmentation.\n  When I checked beluga's shuffle-csv notebook, I found pd.read_csv(..., nrows = 34000)!\n  I only used 1/5 of the datasets!\n  In previous, I only focused on neutral network improvement.\n  I changed it to the total datasets, and the accuracy improved a lot.\n  So after you forked other's notebook, check everything at the beginning!\n  \nFor my last result, I use 128*128 size image, about 1-2 epochs for full dataset, train 2 models: densenet121(batchsize: 170) and resnet50(batchsize: 250).\nsize(validation dataset) = 0.3 * size(full dataset)\nThe best resnet50 model: 0.931 Public LB\nThe best densenet model: 0.934 Public LB and 0.932 Public LB\nEnsemble: 0.934_densenet(weight = 2), 0.932_densenet(weight = 1.8), 0.931_resnet(weight = 1.65) result = 0.937 Public LB\nIt seems that various neutral network models can ensemble a higher accuracy.\nTTA: symmetric deformation of test_simplified.csv, but no help in accuracy.\n\nAt last,I am so happy to both win a medal and get help from you guys.:-)",
    "434537": "Congratulations and Thanks for sharing. \nThe mistake that I made was tweaking the MobileNet for quite a few days and then trying DenseNet. However, instead of expermimenting with smaller subsets, I decided to use the complete dataset due to time constraints-the results were not good. \nTwo things I've learnt: Do not enter a competition after its half way through, experiment well before running a full-blown experiment.",
    "433799": "Congratulation on finishing in 106 as a novice, It is indeed inspiring !!",
    "433711": "@Dilapsky Lee\n\n*\"For me, 128x128size with resnet50, 2 epochs takes 3-4 days.\"*\n\nI understand my modest score : 1 GPU with 4Gb RAM, and never more than ONE night max for training : I need my computer during daytime.\n\nI was also disappointed by the poor results of my GRU's trials.\n\nAnyway, once again very good job for a starter and very good advices above !",
    "433498": "Congratulations and thanks for these practical tips. Hope to see you at the top of a leaderboard very soon.",
    "433491": "Congrat and thank you for sharing! We have similar experience, in our first week, we just realize that we use only 1/5 of the data as well! Our score is closed, our best single LB is also closed (We got 0.933LB using Xception)",
    "433516": "Congratulations ! \n\nYou now can apply your knowledge to mechanical drawing :)"
  }
}