{
  "id": 70558,
  "title": "some tricks for getting LB 0.945",
  "url": "/competitions/quickdraw-doodle-recognition/discussion/70558",
  "author_name": "hengck23",
  "post_date": "2018-11-05T07:39:47.632000",
  "votes": 88,
  "comment_count": 143,
  "views": 0,
  "content": "<p>if you are using CNN:</p>\n\n<ol>\n<li><p>use all train images</p></li>\n<li><p>use large image</p></li>\n<li><p>use correct network</p></li>\n<li><p>use larger batch size</p></li>\n<li><p>train for many iterations</p></li>\n</ol>\n\n<p>you you don't see improvement, change the way you render the strokes to images. <strong>(you have 3 channel rgb, make good use of it! try to encode more information)</strong></p>\n\n<p>you should be able to get to 0.941 without trick, or ensemble ...</p>",
  "messages": [
    {
      "id": 415474,
      "postDate": "2018-11-05T07:39:47.633Z",
      "content": "<p>if you are using CNN:</p>\n\n<ol>\n<li><p>use all train images</p></li>\n<li><p>use large image</p></li>\n<li><p>use correct network</p></li>\n<li><p>use larger batch size</p></li>\n<li><p>train for many iterations</p></li>\n</ol>\n\n<p>you you don't see improvement, change the way you render the strokes to images. <strong>(you have 3 channel rgb, make good use of it! try to encode more information)</strong></p>\n\n<p>you should be able to get to 0.941 without trick, or ensemble ...</p>",
      "rawMarkdown": "if you are using CNN:\n\n1. use all train images\n\n2. use large image\n\n3. use correct network\n\n4. use larger batch size\n\n5. train for many iterations\n\nyou you don't see improvement, change the way you render the strokes to images. **(you have 3 channel rgb, make good use of it! try to encode more information)**\n\nyou should be able to get to 0.941 without trick, or ensemble ...",
      "votes": 88
    },
    {
      "id": 423426,
      "postDate": "2018-11-18T08:05:51.523Z",
      "content": "<p>baseline performance:</p>\n\n<p>LB: 0.944</p>\n\n<p>se-resnext-50 (256x256 input, 200 samples per batch), single model, no TTA , no ensmble</p>\n\n<p>all train samples (simplified only, recognized and non-recognized): random 80 images per class are selected as validation, the rest are training samples.</p>\n\n<p>validation loss : ce_loss, top1, top3,  (map@3)</p>\n\n<p>0.607  0.842  0.946  (0.890)*  </p>\n\n<p>train loss : ce_loss, top1, top3,  (map@3)</p>\n\n<p>0.590  0.844  0.948  (0.892) </p>\n\n<hr>\n\n<p>LB 0.928</p>\n\n<p>customized LSTM\n(Max seq length = 600, 200 samples per batch), single model, no TTA , no ensmble</p>\n\n<p>same train samples as above</p>\n\n<p>validation loss : ce_loss, top1, top3,  (map@3)</p>\n\n<p>0.637  0.834  0.943  (0.885)*  </p>\n\n<p>train loss : ce_loss, top1, top3,  (map@3)</p>\n\n<p>0.595  0.841  0.949  (0.892)</p>\n\n<hr>\n\n<p>LB 0.945</p>\n\n<p>stacking via a fuse classifier on features from cnn and lstm. , no TTA , no ensmble</p>\n\n<p>same train samples as above. batch size increase to 512</p>\n\n<p>validation loss : ce_loss, top1, top3,  (map@3)</p>\n\n<p>0.588  0.850  0.947  (0.895)*  </p>\n\n<p>train loss : ce_loss, top1, top3,  (map@3)</p>\n\n<p>0.549  0.850  0.949  (0.896)</p>",
      "rawMarkdown": "baseline performance:\n\nLB: 0.944\n\nse-resnext-50 (256x256 input, 200 samples per batch), single model, no TTA , no ensmble\n\nall train samples (simplified only, recognized and non-recognized): random 80 images per class are selected as validation, the rest are training samples.\n\nvalidation loss : ce_loss, top1, top3,  (map@3)\n\n0.607  0.842  0.946  (0.890)*  \n\ntrain loss : ce_loss, top1, top3,  (map@3)\n\n0.590  0.844  0.948  (0.892) \n\n----\nLB 0.928\n\ncustomized LSTM\n(Max seq length = 600, 200 samples per batch), single model, no TTA , no ensmble\n\nsame train samples as above\n\n\nvalidation loss : ce_loss, top1, top3,  (map@3)\n\n0.637  0.834  0.943  (0.885)*  \n\ntrain loss : ce_loss, top1, top3,  (map@3)\n\n 0.595  0.841  0.949  (0.892)\n\n----\nLB 0.945\n\nstacking via a fuse classifier on features from cnn and lstm. , no TTA , no ensmble\n\n\nsame train samples as above. batch size increase to 512\n\n\nvalidation loss : ce_loss, top1, top3,  (map@3)\n\n0.588  0.850  0.947  (0.895)*  \n\ntrain loss : ce_loss, top1, top3,  (map@3)\n\n 0.549  0.850  0.949  (0.896)\n",
      "votes": 22,
      "replies": [
        {
          "id": 423428,
          "postDate": "2018-11-18T08:09:42.780Z",
          "content": "<p>Hi~\nHeng\nWhat is your batch_size</p>",
          "rawMarkdown": "Hi~\nHeng\nWhat is your batch_size",
          "votes": -4
        },
        {
          "id": 423443,
          "postDate": "2018-11-18T08:47:13.543Z",
          "content": "<p>200 samples per batch...</p>",
          "rawMarkdown": "200 samples per batch...",
          "votes": 2
        },
        {
          "id": 423444,
          "postDate": "2018-11-18T08:48:30.520Z",
          "content": "<p>Hi, Heng, how much time did your model spend for a epoch ?</p>",
          "rawMarkdown": "Hi, Heng, how much time did your model spend for a epoch ?"
        },
        {
          "id": 423450,
          "postDate": "2018-11-18T09:21:04.663Z",
          "content": "<pre><code>state_dict[key] = pretrain_state_dict[key.replace('resnet.layer2.','layer2.')]\n</code></pre>\n\n<p>KeyError: 'layer2.0.bn1.num_batches_tracked'</p>\n\n<p>Hi~ Heng \nWhen I use your code I got this error at load_pretrained file in the run_check_net function</p>",
          "rawMarkdown": "    state_dict[key] = pretrain_state_dict[key.replace('resnet.layer2.','layer2.')]\n\nKeyError: 'layer2.0.bn1.num_batches_tracked'\n\nHi~ Heng \nWhen I use your code I got this error at load_pretrained file in the run_check_net function"
        },
        {
          "id": 423477,
          "postDate": "2018-11-18T10:48:54.167Z",
          "content": "<p>this is due to different version of pytorch.</p>\n\n<p><a href=\"https://discuss.pytorch.org/t/unexpected-key-in-state-dict-bn1-num-batches-tracked/29454\">https://discuss.pytorch.org/t/unexpected-key-in-state-dict-bn1-num-batches-tracked/29454</a></p>\n\n<p>just ignore the key 'numbatchestracked'</p>",
          "rawMarkdown": "this is due to different version of pytorch.\n\nhttps://discuss.pytorch.org/t/unexpected-key-in-state-dict-bn1-num-batches-tracked/29454\n\n\njust ignore the key 'numbatchestracked'"
        },
        {
          "id": 423491,
          "postDate": "2018-11-18T11:35:59.933Z",
          "content": "<p>Hey Heng, what kind of dataloader are you using? drawing them on the fly or saving images and then loading them while training. \nThanks. </p>",
          "rawMarkdown": "Hey Heng, what kind of dataloader are you using? drawing them on the fly or saving images and then loading them while training. \nThanks. ",
          "votes": 1
        },
        {
          "id": 423575,
          "postDate": "2018-11-18T15:28:57.923Z",
          "content": "<p>For the LSTM, do you using the raw data instead of simplified data?</p>",
          "rawMarkdown": "For the LSTM, do you using the raw data instead of simplified data?",
          "votes": 1
        },
        {
          "id": 423732,
          "postDate": "2018-11-19T00:25:41.037Z",
          "content": "<p>Thanks Heng, what learning rate schedule do you use? </p>",
          "rawMarkdown": "Thanks Heng, what learning rate schedule do you use? "
        },
        {
          "id": 423845,
          "postDate": "2018-11-19T06:17:25.347Z",
          "content": "<p>@Heng CherKeng， Thanks for all the tips that are very useful. I saw you only used batch_size 200, which seems quite small considering the dataset is quite noisy.  Were your trainset manually corrected or you just use the original simplified dataset? Thank you so much.</p>",
          "rawMarkdown": "@Heng CherKeng， Thanks for all the tips that are very useful. I saw you only used batch_size 200, which seems quite small considering the dataset is quite noisy.  Were your trainset manually corrected or you just use the original simplified dataset? Thank you so much.",
          "votes": 1
        },
        {
          "id": 425317,
          "postDate": "2018-11-21T12:44:53.793Z",
          "content": "<p>I tried running this with Keras on a p3.2xlarge AWS instance and it took 2 hours to finish 1 epoch of 250k images, so 400 hours to go through the entire dataset! :( how can you train this in a reasonable amount of time??</p>",
          "rawMarkdown": "I tried running this with Keras on a p3.2xlarge AWS instance and it took 2 hours to finish 1 epoch of 250k images, so 400 hours to go through the entire dataset! :( how can you train this in a reasonable amount of time??"
        },
        {
          "id": 429024,
          "postDate": "2018-11-28T08:35:56.853Z",
          "content": "<p>Hi, i trained 256*256 se-resnext50 with batch size 400. i have got better validation results on local LB (0.894 map@3, which is better than 0.890), but I can not get 0.944(as you have stated) on public LB, public LB is only 0.942. What can be the possible reason? It would be great help if you can give me some advice. Thanks a lot.</p>\n\n<p>ps: I use your pytorch start code as my code base. Thanks again.</p>",
          "rawMarkdown": "Hi, i trained 256*256 se-resnext50 with batch size 400. i have got better validation results on local LB (0.894 map@3, which is better than 0.890), but I can not get 0.944(as you have stated) on public LB, public LB is only 0.942. What can be the possible reason? It would be great help if you can give me some advice. Thanks a lot.\n\nps: I use your pytorch start code as my code base. Thanks again."
        },
        {
          "id": 429271,
          "postDate": "2018-11-28T15:55:13.050Z",
          "content": "<p>@ 凉宫ハルヒ</p>\n\n<p>what is the stroke encoding and lr schedule that you use ?</p>",
          "rawMarkdown": "@ 凉宫ハルヒ\n\nwhat is the stroke encoding and lr schedule that you use ?"
        },
        {
          "id": 429537,
          "postDate": "2018-11-29T01:55:09.747Z",
          "content": "<p>I use the stroke encoding provided by Heng's start code. \nlr schedule: SGD decay 0.5 at 800k steps</p>",
          "rawMarkdown": "I use the stroke encoding provided by Heng's start code. \nlr schedule: SGD decay 0.5 at 800k steps",
          "votes": 1
        },
        {
          "id": 429931,
          "postDate": "2018-11-29T15:05:43.437Z",
          "content": "<p>@凉宫ハルヒ</p>\n\n<p>Thanks for your valuable insight. </p>",
          "rawMarkdown": "@凉宫ハルヒ\n\nThanks for your valuable insight. "
        }
      ]
    },
    {
      "id": 423502,
      "postDate": "2018-11-18T12:12:59.193Z",
      "content": "<p>Regarding the efficient usage of input channels (3 or even more if you use a custom model), the following paper might give some inspiration: <a href=\"https://arxiv.org/pdf/1501.07873.pdf\">https://arxiv.org/pdf/1501.07873.pdf</a></p>",
      "rawMarkdown": "Regarding the efficient usage of input channels (3 or even more if you use a custom model), the following paper might give some inspiration: https://arxiv.org/pdf/1501.07873.pdf",
      "votes": 11,
      "replies": [
        {
          "id": 423894,
          "postDate": "2018-11-19T08:02:28.170Z",
          "content": "<p>HI, Omallo. Did you use 6 channels as input from paper? And if you don't mind. Coud you share what the best model did you use?</p>",
          "rawMarkdown": "HI, Omallo. Did you use 6 channels as input from paper? And if you don't mind. Coud you share what the best model did you use?",
          "votes": 1
        },
        {
          "id": 424280,
          "postDate": "2018-11-19T20:39:01.163Z",
          "content": "<p>I'm currently using 2 models, one self designed model which is weaker but much faster for training and one deeper model (SeResNext50).</p>\n\n<p>With my custom model, I use the 6 channels as described in the paper. Compared with using just 1 channel, this gave me a LB boost of +0.015.</p>\n\n<p>With the SeResNext50 model, I use only 3 channels where I draw 1/3, 2/3, and 3/3 of the strokes, respectively. Compared with using just 1 channel, this gave me a LB boost of +0.013.</p>",
          "rawMarkdown": "I'm currently using 2 models, one self designed model which is weaker but much faster for training and one deeper model (SeResNext50).\n\nWith my custom model, I use the 6 channels as described in the paper. Compared with using just 1 channel, this gave me a LB boost of +0.015.\n\nWith the SeResNext50 model, I use only 3 channels where I draw 1/3, 2/3, and 3/3 of the strokes, respectively. Compared with using just 1 channel, this gave me a LB boost of +0.013.",
          "votes": 3
        },
        {
          "id": 424299,
          "postDate": "2018-11-19T21:36:34.407Z",
          "content": "<p>I am wondering when you predict the test set, are u still using these method or you just have 3 copy of same image?</p>",
          "rawMarkdown": "I am wondering when you predict the test set, are u still using these method or you just have 3 copy of same image?"
        },
        {
          "id": 424313,
          "postDate": "2018-11-19T21:58:19.133Z",
          "content": "<p>I draw the images on the fly from the stroke data while training and while evaluating on the validation and test sets.</p>",
          "rawMarkdown": "I draw the images on the fly from the stroke data while training and while evaluating on the validation and test sets.",
          "votes": 2
        },
        {
          "id": 424356,
          "postDate": "2018-11-20T01:16:22.017Z",
          "content": "<p>nice, thank you, Did you use lr schedule？what schedule?</p>",
          "rawMarkdown": "nice, thank you, Did you use lr schedule？what schedule?"
        },
        {
          "id": 424454,
          "postDate": "2018-11-20T06:30:42.487Z",
          "content": "<p>Omallo, what's your experience with the speed of training of SeResnext50? I tried running it on the Kaggle kernel and it was extremely slow. </p>",
          "rawMarkdown": "Omallo, what's your experience with the speed of training of SeResnext50? I tried running it on the Kaggle kernel and it was extremely slow. ",
          "votes": 1
        },
        {
          "id": 424529,
          "postDate": "2018-11-20T09:14:06.267Z",
          "content": "<p>HuyenNguyen\nI still run SeResnext 50 on 256*256 img, I estimated it will take 2 weeks to go through a epoch on gtx 1070.\nyou can try half precision first, then multiple GPU if possible. Hopely I can finished in 5 days.</p>",
          "rawMarkdown": "HuyenNguyen\nI still run SeResnext 50 on 256*256 img, I estimated it will take 2 weeks to go through a epoch on gtx 1070.\nyou can try half precision first, then multiple GPU if possible. Hopely I can finished in 5 days.",
          "votes": 2
        },
        {
          "id": 424567,
          "postDate": "2018-11-20T10:30:41.550Z",
          "content": "<p>omg, 1 week for 1 epoch @.@ how do you do half precision?</p>",
          "rawMarkdown": "omg, 1 week for 1 epoch @.@ how do you do half precision?"
        },
        {
          "id": 424674,
          "postDate": "2018-11-20T14:18:05.513Z",
          "content": "<p>Im using pytorch, pytorch make it easy to do just call, <strong>half()</strong>  or <strong>float()</strong> to convert to 16fp and 32fp\n<a href=\"https://discuss.pytorch.org/t/training-with-half-precision/11815\">https://discuss.pytorch.org/t/training-with-half-precision/11815</a>\njust beware the adam optimizer, \n<a href=\"https://discuss.pytorch.org/t/adam-half-precision-nans/1765\">https://discuss.pytorch.org/t/adam-half-precision-nans/1765</a></p>",
          "rawMarkdown": "Im using pytorch, pytorch make it easy to do just call, **half()**  or **float()** to convert to 16fp and 32fp\nhttps://discuss.pytorch.org/t/training-with-half-precision/11815\njust beware the adam optimizer, \nhttps://discuss.pytorch.org/t/adam-half-precision-nans/1765\n",
          "votes": 2
        },
        {
          "id": 424759,
          "postDate": "2018-11-20T16:12:11.763Z",
          "content": "<p>@Gary-DeepLearning I'm using a cyclical lr scheduling using cosine annealing as described in the SGDR paper: <a href=\"https://arxiv.org/pdf/1608.03983.pdf\">https://arxiv.org/pdf/1608.03983.pdf</a></p>",
          "rawMarkdown": "@Gary-DeepLearning I'm using a cyclical lr scheduling using cosine annealing as described in the SGDR paper: https://arxiv.org/pdf/1608.03983.pdf",
          "votes": 1
        },
        {
          "id": 424762,
          "postDate": "2018-11-20T16:19:52.660Z",
          "content": "<p>@HuyenNguyen I'm using 128x128 images and 90% of the samples for training. Under these conditions, it takes about 30h for training 1 epoch on a NVIDIA Quadro P6000.</p>\n\n<p>Using 256x256 images, it takes around 4x longer so I sticked with 128x128 images.</p>",
          "rawMarkdown": "@HuyenNguyen I'm using 128x128 images and 90% of the samples for training. Under these conditions, it takes about 30h for training 1 epoch on a NVIDIA Quadro P6000.\n\nUsing 256x256 images, it takes around 4x longer so I sticked with 128x128 images.",
          "votes": 1
        },
        {
          "id": 424948,
          "postDate": "2018-11-20T23:25:38.283Z",
          "content": "<p>Does cyclical lr help compared to just ReduceonPlateau?</p>",
          "rawMarkdown": "Does cyclical lr help compared to just ReduceonPlateau?",
          "votes": 1
        },
        {
          "id": 424970,
          "postDate": "2018-11-21T00:19:14.497Z",
          "content": "<p>Hi HuyenNguyen,\nMay I update my suggestion:\nfp16 is slow compare to fp32, i speculate older gpu don't have many fp16 unit, maybe latest rtx gpu will do the job better,\nand fp16 also cause many problem because you have to take care numerical stability yourself, \nDistributed training might be the way to go</p>",
          "rawMarkdown": "Hi HuyenNguyen,\nMay I update my suggestion:\nfp16 is slow compare to fp32, i speculate older gpu don't have many fp16 unit, maybe latest rtx gpu will do the job better,\nand fp16 also cause many problem because you have to take care numerical stability yourself, \nDistributed training might be the way to go",
          "votes": 2
        },
        {
          "id": 425006,
          "postDate": "2018-11-21T01:55:44.477Z",
          "content": "<p>Oh, I see. Thank you.</p>",
          "rawMarkdown": "Oh, I see. Thank you."
        },
        {
          "id": 425159,
          "postDate": "2018-11-21T08:01:48.757Z",
          "content": "<p>@HuyenNguyen In general, with cyclical lr scheduling models tend to converge faster even though you might end up with slightly worse performance of the model. Often I use cyclical lr scheduling and, at the end, one run with a constant lr or use ReduceOnPlateau to do fine tuning. One thing I like about cyclical lr scheduling is that you can easily combine it with snapshot ensembling.</p>",
          "rawMarkdown": "@HuyenNguyen In general, with cyclical lr scheduling models tend to converge faster even though you might end up with slightly worse performance of the model. Often I use cyclical lr scheduling and, at the end, one run with a constant lr or use ReduceOnPlateau to do fine tuning. One thing I like about cyclical lr scheduling is that you can easily combine it with snapshot ensembling.",
          "votes": 2
        },
        {
          "id": 425322,
          "postDate": "2018-11-21T12:48:23.030Z",
          "content": "<p>Great, thank you for the tips. </p>",
          "rawMarkdown": "Great, thank you for the tips. "
        },
        {
          "id": 425323,
          "postDate": "2018-11-21T12:49:16Z",
          "content": "<p>Hi Omallo, when you say \"I draw 1/3, 2/3, and 3/3 of the strokes\" do you still give stroke a different colour? </p>",
          "rawMarkdown": "Hi Omallo, when you say \"I draw 1/3, 2/3, and 3/3 of the strokes\" do you still give stroke a different colour? "
        },
        {
          "id": 426052,
          "postDate": "2018-11-22T14:13:51.647Z",
          "content": "<p>@HuyenNguyen Yes, I sill use different colors for the strokes. Here is an example of using 6 channels as described in the paper: <a href=\"https://ibb.co/dGizNA\">https://ibb.co/dGizNA</a></p>",
          "rawMarkdown": "@HuyenNguyen Yes, I sill use different colors for the strokes. Here is an example of using 6 channels as described in the paper: https://ibb.co/dGizNA",
          "votes": 1
        },
        {
          "id": 427017,
          "postDate": "2018-11-24T10:53:12.533Z",
          "content": "<p>Hi, omallo. After treated the image data like this,1/3, 2/3, 3/3. Does it cost you more steps or more epoch to get the result?</p>",
          "rawMarkdown": "Hi, omallo. After treated the image data like this,1/3, 2/3, 3/3. Does it cost you more steps or more epoch to get the result?"
        },
        {
          "id": 427028,
          "postDate": "2018-11-24T11:24:34.897Z",
          "content": "<p>If I remember correctly, there was no noticeable difference, especially if you use a pretrained model</p>",
          "rawMarkdown": "If I remember correctly, there was no noticeable difference, especially if you use a pretrained model",
          "votes": 1
        },
        {
          "id": 427065,
          "postDate": "2018-11-24T13:31:21.950Z",
          "content": "<p>Thanks, omallo. After prepocessing the img data as you do, when I train my model, I find the loss is large than the 1 channels model in the early stage. This condition makes me feel very strange. Maybe some of my preprocessing work has some mistakes in it.</p>",
          "rawMarkdown": "Thanks, omallo. After prepocessing the img data as you do, when I train my model, I find the loss is large than the 1 channels model in the early stage. This condition makes me feel very strange. Maybe some of my preprocessing work has some mistakes in it."
        },
        {
          "id": 427103,
          "postDate": "2018-11-24T15:29:29.550Z",
          "content": "<p>Hi guys! I have tried taht and give me worst results (no so much) using 6 channels in White/Black... I divided the points / number of channels (doing this too for getting part1+part2 and so on)... How you did it in colour? The number of strokes is not uniform and for example If I want 3 fragments there are images with 1 \"movement\" for all the picture.</p>",
          "rawMarkdown": "Hi guys! I have tried taht and give me worst results (no so much) using 6 channels in White/Black... I divided the points / number of channels (doing this too for getting part1+part2 and so on)... How you did it in colour? The number of strokes is not uniform and for example If I want 3 fragments there are images with 1 \"movement\" for all the picture."
        },
        {
          "id": 427116,
          "postDate": "2018-11-24T16:22:50.203Z",
          "content": "<p>So... for this <a href=\"https://ibb.co/dGizNA\">https://ibb.co/dGizNA</a> you have 6 images * 3 channels =&gt; 18 channels as input</p>",
          "rawMarkdown": "So... for this https://ibb.co/dGizNA you have 6 images * 3 channels =&gt; 18 channels as input"
        },
        {
          "id": 427130,
          "postDate": "2018-11-24T17:11:01.037Z",
          "content": "<p>Mario, I just realized that my image might be somewhat confusing. Actually, for the example with 6 images, the colors are just different grayscale values between 0 and 240 so it's not a RGB color. The images look colored since I used the default color map of matplotlib to render them.</p>\n\n<p>This means that you would only need 6 channels in that case. Sorry for the confusion.</p>",
          "rawMarkdown": "Mario, I just realized that my image might be somewhat confusing. Actually, for the example with 6 images, the colors are just different grayscale values between 0 and 240 so it's not a RGB color. The images look colored since I used the default color map of matplotlib to render them.\n\nThis means that you would only need 6 channels in that case. Sorry for the confusion."
        },
        {
          "id": 427132,
          "postDate": "2018-11-24T17:21:08.383Z",
          "content": "<p>Thanks omallo! Unfourtunately I tried and obtained slightly worst results. But thanks anyway, here I am to learn :D</p>",
          "rawMarkdown": "Thanks omallo! Unfourtunately I tried and obtained slightly worst results. But thanks anyway, here I am to learn :D"
        }
      ]
    },
    {
      "id": 429119,
      "postDate": "2018-11-28T11:29:47.570Z",
      "content": "<p>another reference results</p>\n\n<p>LB: did not submit (estimated to be 0.942)</p>\n\n<p>senet154 (128x128 input, 256 samples per batch), single model, no TTA , no ensmble</p>\n\n<p>all train samples (simplified only, recognized and non-recognized): random 80 images per class are selected as validation, the rest are training samples.</p>\n\n<p>validation loss : ce_loss, top1, top3, (map@3)</p>\n\n<p>0.622  0.839  0.944  (0.888)* </p>\n\n<p>train loss : ce_loss, top1, top3, (map@3)</p>\n\n<p>0.582  0.848  0.946  (0.893)  </p>",
      "rawMarkdown": "another reference results\n\nLB: did not submit (estimated to be 0.942)\n\nsenet154 (128x128 input, 256 samples per batch), single model, no TTA , no ensmble\n\nall train samples (simplified only, recognized and non-recognized): random 80 images per class are selected as validation, the rest are training samples.\n\nvalidation loss : ce_loss, top1, top3, (map@3)\n\n0.622  0.839  0.944  (0.888)* \n\ntrain loss : ce_loss, top1, top3, (map@3)\n\n0.582  0.848  0.946  (0.893)  ",
      "votes": 5,
      "replies": [
        {
          "id": 429265,
          "postDate": "2018-11-28T15:50:23.573Z",
          "content": "<p>@Heng</p>\n\n<p>What is your lr schedule ? and stroke encoding ?</p>",
          "rawMarkdown": "@Heng\n\nWhat is your lr schedule ? and stroke encoding ?",
          "votes": 1
        }
      ]
    },
    {
      "id": 415742,
      "postDate": "2018-11-05T15:42:24.273Z",
      "content": "<p>Thanks Heng </p>\n\n<p>I still use 64*64 ....may be it's time to use larger image  :)</p>\n\n<p>I am also  working on encoding more information using all the 3 channels. </p>\n\n<p>BTW , do you think there are some augmentations that help ? </p>",
      "rawMarkdown": "Thanks Heng \n\n I still use 64*64 ....may be it's time to use larger image  :)\n\nI am also  working on encoding more information using all the 3 channels. \n\nBTW , do you think there are some augmentations that help ? ",
      "votes": 5,
      "replies": [
        {
          "id": 418494,
          "postDate": "2018-11-10T01:33:30.223Z",
          "content": "<p>Hi, how did you encode more information with using all the 3 channels?</p>",
          "rawMarkdown": "Hi, how did you encode more information with using all the 3 channels?",
          "votes": 1
        }
      ]
    },
    {
      "id": 415551,
      "postDate": "2018-11-05T09:57:38.987Z",
      "content": "<p>Now I use \n1. 50k/class\n2. img_size 224*224\n3. res50\n4. 64*8\ncould got 0.936 and 0.938 with ensemble</p>\n\n<p>Due to some bugs, I did not got any new models with new setting for two weeks. Sign~</p>\n\n<p>Now I will try:\n1. all imgs\n2. img_size: 448*448\n3. Senet?\n4. 256*8</p>",
      "rawMarkdown": "Now I use \n1. 50k/class\n2. img_size 224*224\n3. res50\n4. 64*8\ncould got 0.936 and 0.938 with ensemble\n\nDue to some bugs, I did not got any new models with new setting for two weeks. Sign~\n\nNow I will try:\n1. all imgs\n2. img_size: 448*448\n3. Senet?\n4. 256*8",
      "votes": 5,
      "replies": [
        {
          "id": 415559,
          "postDate": "2018-11-05T10:12:19.887Z",
          "content": "<p>Don't you think 448 is too large? 256 should be more than enough according to me.</p>",
          "rawMarkdown": "Don't you think 448 is too large? 256 should be more than enough according to me."
        },
        {
          "id": 415574,
          "postDate": "2018-11-05T10:45:44.940Z",
          "content": "<p>This is so much compute, also if have such specs, consider using NasNet then. It should give better results as per benchamarks.</p>",
          "rawMarkdown": "This is so much compute, also if have such specs, consider using NasNet then. It should give better results as per benchamarks.",
          "votes": 1
        },
        {
          "id": 415580,
          "postDate": "2018-11-05T10:55:08.187Z",
          "content": "<p>recently, i wrote a distributed framework, so 448*448 do not take a lot of time</p>",
          "rawMarkdown": "recently, i wrote a distributed framework, so 448*448 do not take a lot of time"
        },
        {
          "id": 415624,
          "postDate": "2018-11-05T12:04:27.533Z",
          "content": "<p>Are are you using multiple GPUs?</p>",
          "rawMarkdown": "Are are you using multiple GPUs?",
          "votes": 1
        },
        {
          "id": 415633,
          "postDate": "2018-11-05T12:25:29.467Z",
          "content": "<p>@Li Bin</p>\n\n<p>e.g.</p>\n\n<p>1 epoch of \"50,000k/class\"  may be better than 100 epoch of \"50k/class\"</p>",
          "rawMarkdown": "@Li Bin\n\ne.g.\n\n1 epoch of \"50,000k/class\"  may be better than 100 epoch of \"50k/class\"",
          "votes": 6
        },
        {
          "id": 415638,
          "postDate": "2018-11-05T12:31:53.637Z",
          "content": "<p>you should be able to get lb ~0.940 with 224x224. 256x256 is better</p>",
          "rawMarkdown": "you should be able to get lb ~0.940 with 224x224. 256x256 is better",
          "votes": 3
        },
        {
          "id": 415668,
          "postDate": "2018-11-05T13:20:24.993Z",
          "content": "<p>Thank your for your suggestion. It is very useful for me.</p>",
          "rawMarkdown": "Thank your for your suggestion. It is very useful for me."
        },
        {
          "id": 415693,
          "postDate": "2018-11-05T14:11:17.997Z",
          "content": "<blockquote>\n  <p>1 epoch of \"50,000k/class\" may be better than 100 epoch of \"50k/class\"</p>\n</blockquote>\n\n<p>@Heng do you mean use high number of steps_per_epoch? </p>",
          "rawMarkdown": "&gt; 1 epoch of \"50,000k/class\" may be better than 100 epoch of \"50k/class\"\n\n@Heng do you mean use high number of steps_per_epoch? ",
          "votes": 1
        },
        {
          "id": 415964,
          "postDate": "2018-11-06T01:23:34.107Z",
          "content": "<p>Training Resnet50 on that amount data of that image size will take forever. what machine are you using? </p>",
          "rawMarkdown": "Training Resnet50 on that amount data of that image size will take forever. what machine are you using? ",
          "votes": 2
        },
        {
          "id": 416191,
          "postDate": "2018-11-06T11:19:37.557Z",
          "content": "<p>try GCP preemptible instances with free trial</p>",
          "rawMarkdown": "try GCP preemptible instances with free trial"
        },
        {
          "id": 416388,
          "postDate": "2018-11-06T15:56:44.130Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 416522,
          "postDate": "2018-11-06T19:58:15.117Z",
          "content": "<p>I'm showing a single epoch at that image size (224, 224, 3) and all layers trainable at appx 2 days for ResNet50 - this is on an AWS gpu instance with one GPU. This is will all available data from the Shuffle CSV data sets. </p>",
          "rawMarkdown": "I'm showing a single epoch at that image size (224, 224, 3) and all layers trainable at appx 2 days for ResNet50 - this is on an AWS gpu instance with one GPU. This is will all available data from the Shuffle CSV data sets. ",
          "votes": 1
        },
        {
          "id": 416596,
          "postDate": "2018-11-06T23:59:56.687Z",
          "content": "<p>A batchsize of 2048 images of size 224x224x3, how can you fit that in memory?</p>",
          "rawMarkdown": "A batchsize of 2048 images of size 224x224x3, how can you fit that in memory?",
          "votes": 1
        },
        {
          "id": 416598,
          "postDate": "2018-11-07T00:17:52.373Z",
          "content": "<p>my batch size is only 64. I could go a little higher but not much. That is another reason for the slower training that I reported.</p>",
          "rawMarkdown": "my batch size is only 64. I could go a little higher but not much. That is another reason for the slower training that I reported.",
          "votes": 1
        },
        {
          "id": 416612,
          "postDate": "2018-11-07T01:21:40.907Z",
          "content": "<p>@Li Bin - just curious how many epochs did it take to reach that point? </p>",
          "rawMarkdown": "@Li Bin - just curious how many epochs did it take to reach that point? "
        },
        {
          "id": 416631,
          "postDate": "2018-11-07T02:46:44.580Z",
          "content": "<p>12 epoch</p>",
          "rawMarkdown": "12 epoch"
        },
        {
          "id": 416654,
          "postDate": "2018-11-07T03:30:33.283Z",
          "content": "<p>what does 256*8 mean? I am a new-bee. Thx. XD</p>",
          "rawMarkdown": "what does 256*8 mean? I am a new-bee. Thx. XD"
        },
        {
          "id": 416667,
          "postDate": "2018-11-07T03:53:47.523Z",
          "content": "<p>256 batch per gpu, and I have 8 gpus</p>",
          "rawMarkdown": "256 batch per gpu, and I have 8 gpus",
          "votes": 2
        },
        {
          "id": 416672,
          "postDate": "2018-11-07T04:05:58.750Z",
          "content": "<p>64 batch size is too small</p>",
          "rawMarkdown": "64 batch size is too small"
        },
        {
          "id": 416674,
          "postDate": "2018-11-07T04:14:41.413Z",
          "content": "<p>@CherKeng 256 batch/gpu, the batchsize totally is 256*8</p>",
          "rawMarkdown": "@CherKeng 256 batch/gpu, the batchsize totally is 256*8",
          "votes": 1
        },
        {
          "id": 416683,
          "postDate": "2018-11-07T04:54:04.040Z",
          "content": "<p>@CherKeng - yes I figured that this might be to small to train in a reasonable amount of time. Do you have any suggestions to be able to increase my batch size on a single GPU?</p>",
          "rawMarkdown": "@CherKeng - yes I figured that this might be to small to train in a reasonable amount of time. Do you have any suggestions to be able to increase my batch size on a single GPU?",
          "votes": 1
        },
        {
          "id": 416854,
          "postDate": "2018-11-07T11:21:03.500Z",
          "content": "<p>I  use two TitanXP and set the batch size = 256 then keras report me OOM. :( <br>\n(image size is 256*256 and keras resnet50)\nSo am I set something wrong?  could you tell me what kind of GPU do you used? Thx  </p>",
          "rawMarkdown": "I  use two TitanXP and set the batch size = 256 then keras report me OOM. :(  \n(image size is 256*256 and keras resnet50)\nSo am I set something wrong?  could you tell me what kind of GPU do you used? Thx  ",
          "votes": 1
        },
        {
          "id": 416868,
          "postDate": "2018-11-07T11:33:21.347Z",
          "content": "<p>@RDizzl3</p>\n\n<p>May be you should set very large Steps per Epoch ? </p>\n\n<p>I think 224*224 is way too large for this data, but given people say it gives better results, why not ! :)</p>\n\n<p>In the model I am currently runnung , I set 128*128 for image size  and batchsize 448.  That's the best compromise I could find.</p>",
          "rawMarkdown": "@RDizzl3\n\n\nMay be you should set very large Steps per Epoch ? \n\nI think 224*224 is way too large for this data, but given people say it gives better results, why not ! :)\n\nIn the model I am currently runnung , I set 128*128 for image size  and batchsize 448.  That's the best compromise I could find.",
          "votes": 3
        },
        {
          "id": 425327,
          "postDate": "2018-11-21T12:56:47.087Z",
          "content": "<p>Li Bin, are you saying 12 epochs, and go through the full dataset in each epoch?</p>",
          "rawMarkdown": "Li Bin, are you saying 12 epochs, and go through the full dataset in each epoch?"
        }
      ]
    },
    {
      "id": 415477,
      "postDate": "2018-11-05T07:48:53.463Z",
      "content": "<p>typical training log:\n(1 iter = 1000 iterations, see my pytorch starter kit)</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/415477/10604/train_log.png\" alt=\"enter image description here\"></p>",
      "rawMarkdown": "typical training log:\n(1 iter = 1000 iterations, see my pytorch starter kit)\n\n\n  ![enter image description here][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/415477/10604/train_log.png",
      "votes": 6,
      "replies": [
        {
          "id": 415528,
          "postDate": "2018-11-05T09:14:29.693Z",
          "content": "<p>thank you, sir!</p>",
          "rawMarkdown": "thank you, sir!"
        },
        {
          "id": 416202,
          "postDate": "2018-11-06T11:49:53.510Z",
          "content": "<p>Hi CherKeng, why your log have such high top 3 acc (0.945)but low lb(0.889), and loss corresponds to cross entropy softmax loss?</p>",
          "rawMarkdown": "Hi CherKeng, why your log have such high top 3 acc (0.945)but low lb(0.889), and loss corresponds to cross entropy softmax loss?",
          "votes": 2
        },
        {
          "id": 417291,
          "postDate": "2018-11-08T03:22:20.787Z",
          "content": "<p>Thank you !</p>",
          "rawMarkdown": "Thank you !"
        },
        {
          "id": 418821,
          "postDate": "2018-11-10T17:05:07.363Z",
          "content": "<p>Hi Heng,\nis this log starting from just a normal imagenet pretrained model or continueing from some checkpoint?</p>",
          "rawMarkdown": "Hi Heng,\nis this log starting from just a normal imagenet pretrained model or continueing from some checkpoint?",
          "votes": 2
        },
        {
          "id": 425072,
          "postDate": "2018-11-21T05:10:07.137Z",
          "content": "<p>It is continuing from checkpoint because 0.61 val loss will give leader board score of around 0.938. </p>",
          "rawMarkdown": "It is continuing from checkpoint because 0.61 val loss will give leader board score of around 0.938. "
        },
        {
          "id": 426306,
          "postDate": "2018-11-23T03:22:05.780Z",
          "content": "<p>Do you think there is a gap between local score and public LB score? If so, why? In all of my models, I'm getting local validation score of ~0.05 lower than public LB.</p>",
          "rawMarkdown": "Do you think there is a gap between local score and public LB score? If so, why? In all of my models, I'm getting local validation score of ~0.05 lower than public LB."
        }
      ]
    },
    {
      "id": 421429,
      "postDate": "2018-11-15T01:54:35.503Z",
      "content": "<p>Hi Heng, thanks for your great tips! I've trained ResNet101 for 3 epochs with 64x64 scale and 1024 batch size and it scored 0.925. Does the image size really matter?</p>",
      "rawMarkdown": "Hi Heng, thanks for your great tips! I've trained ResNet101 for 3 epochs with 64x64 scale and 1024 batch size and it scored 0.925. Does the image size really matter?",
      "votes": 3
    },
    {
      "id": 419753,
      "postDate": "2018-11-12T14:14:07.340Z",
      "content": "<p>Hi Heng,\nI am currently running on gtx 1080 8GB, using resnet 50, with 64x64 image size, 128 batch size, \nI am using 90% data as training dataset. it take 1 day to go through 1 epoch.\nMy questions are:\n1. Does the training time look right? or can you suggest a training time baseline for 1 epoch?\n2. Is it better to train using 10K per class with 10 epoch, rather than 100K per class with 1 epoch? \nBecause my gpu is quite slow.\nThanks</p>",
      "rawMarkdown": "Hi Heng,\nI am currently running on gtx 1080 8GB, using resnet 50, with 64x64 image size, 128 batch size, \nI am using 90% data as training dataset. it take 1 day to go through 1 epoch.\nMy questions are:\n1. Does the training time look right? or can you suggest a training time baseline for 1 epoch?\n2. Is it better to train using 10K per class with 10 epoch, rather than 100K per class with 1 epoch? \nBecause my gpu is quite slow.\nThanks",
      "votes": 3,
      "replies": [
        {
          "id": 419805,
          "postDate": "2018-11-12T15:26:11.157Z",
          "content": "<p>my validation set is 80 per class. All the rest are train images. Whatever you do, you should get</p>\n\n<p>0.886 for local LB  to have a 0.940 LB.\nThis is about local Top1-acc = 0.838, Top3-acc = 0.943,</p>\n\n<hr>\n\n<p>(however, i don't think 64x64 can reach that accuracy. best of 64x64 i can get is about LB 0.934, but i do not use resnet50)</p>\n\n<p>image-net resnet uses batch of 256 for 1000 classes. maybe 128 batch is ok (i haven't tried batch size 128, as i use larger than 128) </p>\n\n<hr>\n\n<p>about 2 epoch my be sufficient. With such large dataset, i do not use epoch to monitor my training. instead i use training iterations. i.e the learning rate are decreased based on number of iterations and not epoch</p>\n\n<p>You can save your model at every e.g. 1 hours and record the validation loss. Then you decide what to make submission. you should be able to submit after training for 24 to 48 hrs?</p>",
          "rawMarkdown": "my validation set is 80 per class. All the rest are train images. Whatever you do, you should get\n\n0.886 for local LB  to have a 0.940 LB.\nThis is about local Top1-acc = 0.838, Top3-acc = 0.943,\n\n---\n(however, i don't think 64x64 can reach that accuracy. best of 64x64 i can get is about LB 0.934, but i do not use resnet50)\n\nimage-net resnet uses batch of 256 for 1000 classes. maybe 128 batch is ok (i haven't tried batch size 128, as i use larger than 128) \n\n---\n\nabout 2 epoch my be sufficient. With such large dataset, i do not use epoch to monitor my training. instead i use training iterations. i.e the learning rate are decreased based on number of iterations and not epoch\n\nYou can save your model at every e.g. 1 hours and record the validation loss. Then you decide what to make submission. you should be able to submit after training for 24 to 48 hrs?",
          "votes": 8
        },
        {
          "id": 419808,
          "postDate": "2018-11-12T15:31:47.003Z",
          "content": "<p>Hi, heng, your best score was based on CNN?</p>",
          "rawMarkdown": "Hi, heng, your best score was based on CNN?"
        },
        {
          "id": 419810,
          "postDate": "2018-11-12T15:33:58.570Z",
          "content": "<p>LB 0.943~0.944 without ensemble</p>",
          "rawMarkdown": "LB 0.943~0.944 without ensemble"
        },
        {
          "id": 419812,
          "postDate": "2018-11-12T15:35:43.090Z",
          "content": "<p>ok. thank you.</p>",
          "rawMarkdown": "ok. thank you."
        },
        {
          "id": 419830,
          "postDate": "2018-11-12T15:56:32.083Z",
          "content": "<p>I have some idea where i might do wrong,\nThanks you very much CherKeng</p>",
          "rawMarkdown": "I have some idea where i might do wrong,\nThanks you very much CherKeng"
        }
      ]
    },
    {
      "id": 419730,
      "postDate": "2018-11-12T13:36:29.190Z",
      "content": "<p>encode by time(left)</p>\n\n<p>encode by stroke/part (right)</p>\n\n<p>you can tnink of a way to encode \"pause time\"</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/419730/10659/0.73_9000297423888044_baseball_bat_bat_baseball.png\" alt=\"enter image description here\"></p>",
      "rawMarkdown": "encode by time(left)\n\nencode by stroke/part (right)\n\n\nyou can tnink of a way to encode \"pause time\"\n\n   ![enter image description here][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/419730/10659/0.73_9000297423888044_baseball_bat_bat_baseball.png",
      "votes": 4
    },
    {
      "id": 418071,
      "postDate": "2018-11-09T08:46:43.257Z",
      "content": "<p>Hi Heng,</p>\n\n<p>Amazing tips as usual. I am currently running on gtx 1060 6GB, and so I cannot quite increase both image size and batch size as it would get OOM. Which of the two should I priortize? Imagesize or batch?</p>",
      "rawMarkdown": "Hi Heng,\n\nAmazing tips as usual. I am currently running on gtx 1060 6GB, and so I cannot quite increase both image size and batch size as it would get OOM. Which of the two should I priortize? Imagesize or batch?",
      "votes": 2,
      "replies": [
        {
          "id": 418762,
          "postDate": "2018-11-10T15:12:57.867Z",
          "content": "<p>There is trick that you can accumulate the gradient for several batches and then update the gradient, so the actual batch size is much bigger</p>",
          "rawMarkdown": "There is trick that you can accumulate the gradient for several batches and then update the gradient, so the actual batch size is much bigger",
          "votes": 6
        },
        {
          "id": 418978,
          "postDate": "2018-11-11T02:36:05.400Z",
          "content": "<p>Interesting, found this thread about it so far: <a href=\"https://github.com/keras-team/keras/issues/3556\">https://github.com/keras-team/keras/issues/3556</a></p>",
          "rawMarkdown": "Interesting, found this thread about it so far: https://github.com/keras-team/keras/issues/3556",
          "votes": 2
        }
      ]
    },
    {
      "id": 425650,
      "postDate": "2018-11-21T22:48:58.020Z",
      "content": "<p>Hi Heng,\nMay I ask what optimizer you used?\nDo you think Adam + Cyclic LR is good idea, \nMy Top 3 accu can easy get 0.96 using this combination but with poor public LB score......</p>",
      "rawMarkdown": "Hi Heng,\nMay I ask what optimizer you used?\nDo you think Adam + Cyclic LR is good idea, \nMy Top 3 accu can easy get 0.96 using this combination but with poor public LB score......\n",
      "votes": 1,
      "replies": [
        {
          "id": 425749,
          "postDate": "2018-11-22T03:37:27.693Z",
          "content": "<p>Did you use all dataset?</p>",
          "rawMarkdown": "Did you use all dataset?",
          "votes": 1
        },
        {
          "id": 426016,
          "postDate": "2018-11-22T12:43:56.303Z",
          "content": "<p>Yes, 80 samples/class for validation, rest are training data\nAnd it can reach 0.96 top 3 accu, only training on 0.1 epoch,\nI also check the train val dataset, there is no overlap between.\nI don't see any leakage, but cannot 100% sure,\nanyone have the same experience?</p>",
          "rawMarkdown": "Yes, 80 samples/class for validation, rest are training data\nAnd it can reach 0.96 top 3 accu, only training on 0.1 epoch,\nI also check the train val dataset, there is no overlap between.\nI don't see any leakage, but cannot 100% sure,\nanyone have the same experience?"
        },
        {
          "id": 426282,
          "postDate": "2018-11-23T02:16:45.587Z",
          "content": "<p>I have.\nIf you use multiple GPUs and do not shuffle data you feed to your gpus, you should.\nIn this competition we have 340 classes. I splitted them into 4 pieces and fed each of them to each gpu. I mean,  I fed samples of The Effel Tower, ..., Drill to gpu A,  samples of drums, ..., laptop to gpu B, the same hereafter.\nThe deep learning library would have averaged losses calculated by all gpus,  however, resulting model unfairly predicted classes that fed to gpu A.\nShuffling data fed to gpus solved this problem, at least for me.</p>",
          "rawMarkdown": "I have.\nIf you use multiple GPUs and do not shuffle data you feed to your gpus, you should.\nIn this competition we have 340 classes. I splitted them into 4 pieces and fed each of them to each gpu. I mean,  I fed samples of The Effel Tower, ..., Drill to gpu A,  samples of drums, ..., laptop to gpu B, the same hereafter.\nThe deep learning library would have averaged losses calculated by all gpus,  however, resulting model unfairly predicted classes that fed to gpu A.\nShuffling data fed to gpus solved this problem, at least for me.",
          "votes": 1
        }
      ]
    },
    {
      "id": 421562,
      "postDate": "2018-11-15T06:03:57.720Z",
      "content": "<p>Hi~ Heng\nHow to get the 3 channel rgb?\nShould we  draw the every line with different color?</p>",
      "rawMarkdown": "Hi~ Heng\nHow to get the 3 channel rgb?\nShould we  draw the every line with different color?",
      "votes": 1
    },
    {
      "id": 418503,
      "postDate": "2018-11-10T01:54:20.640Z",
      "content": "<p>Hi Heng,</p>\n\n<p>Could you tell more about \"you have 3 channel rgb, make good use of it! try to encode more information\"? you mean use the bigger filters?</p>",
      "rawMarkdown": "Hi Heng,\n\nCould you tell more about \"you have 3 channel rgb, make good use of it! try to encode more information\"? you mean use the bigger filters?",
      "votes": 1,
      "replies": [
        {
          "id": 418891,
          "postDate": "2018-11-10T20:14:08.247Z",
          "content": "<p>Try using stroke velocity in 2nd channel and perhaps stroke acceleration in the 3rd channel.</p>",
          "rawMarkdown": "Try using stroke velocity in 2nd channel and perhaps stroke acceleration in the 3rd channel.",
          "votes": 1
        },
        {
          "id": 418933,
          "postDate": "2018-11-10T22:45:50.580Z",
          "content": "<p>There's no temporal info about the different points on a stroke though. </p>",
          "rawMarkdown": "There's no temporal info about the different points on a stroke though. "
        },
        {
          "id": 420187,
          "postDate": "2018-11-13T08:29:32.383Z",
          "content": "<p>There are time stamps for each point in raw data.</p>",
          "rawMarkdown": "There are time stamps for each point in raw data.",
          "votes": 1
        },
        {
          "id": 425174,
          "postDate": "2018-11-21T08:42:44.770Z",
          "content": "<p>I see that time stamp is for each row, not each point in each row. \nPlease clarify this. </p>\n\n<p>how can we use timestamp as part of training data?</p>",
          "rawMarkdown": "I see that time stamp is for each row, not each point in each row. \nPlease clarify this. \n\nhow can we use timestamp as part of training data?"
        }
      ]
    },
    {
      "id": 416720,
      "postDate": "2018-11-07T06:58:56.077Z",
      "content": "<p>What is the batch size did you use?  I could not get higher than 64 for 224*224 images for 100000/class. </p>",
      "rawMarkdown": "What is the batch size did you use?  I could not get higher than 64 for 224*224 images for 100000/class. ",
      "votes": 1,
      "replies": [
        {
          "id": 429576,
          "postDate": "2018-11-29T03:11:45.860Z",
          "content": "<p>Use smaller image, priortize batch size first </p>",
          "rawMarkdown": "Use smaller image, priortize batch size first "
        }
      ]
    },
    {
      "id": 415835,
      "postDate": "2018-11-05T19:22:44.960Z",
      "content": "<p>Hi all, thanks for the great discussion.</p>\n\n<p>I am curious how many batch size a Titan X can handle, for example, when using 224x224x3 and ResNet50. I guessed around ~16. Does this make sense or it could be more?</p>\n\n<p>@Li bin mentioned using 256*8 of 224x224x3 input (and also mentioned using a distributed framework.) What would be a typical capacity for a Titan X (as an example)?</p>",
      "rawMarkdown": "Hi all, thanks for the great discussion.\n\nI am curious how many batch size a Titan X can handle, for example, when using 224x224x3 and ResNet50. I guessed around ~16. Does this make sense or it could be more?\n\n@Li bin mentioned using 256*8 of 224x224x3 input (and also mentioned using a distributed framework.) What would be a typical capacity for a Titan X (as an example)?",
      "votes": 1,
      "replies": [
        {
          "id": 415850,
          "postDate": "2018-11-05T19:59:08.810Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 416811,
          "postDate": "2018-11-07T10:01:01.983Z",
          "content": "<p>Sorry， I made a big mistake. For res34 we could have 256batch/gpu, for res50, we could only have 64batch/gpu.</p>",
          "rawMarkdown": "Sorry， I made a big mistake. For res34 we could have 256batch/gpu, for res50, we could only have 64batch/gpu.",
          "votes": 1
        }
      ]
    },
    {
      "id": 415681,
      "postDate": "2018-11-05T13:35:10.773Z",
      "content": "<p>Just curious why large batch size helps, some paper mentioned that smaller batch size should generalize better than larger batch size...</p>",
      "rawMarkdown": "Just curious why large batch size helps, some paper mentioned that smaller batch size should generalize better than larger batch size...\n",
      "votes": 1,
      "replies": [
        {
          "id": 415744,
          "postDate": "2018-11-05T15:46:08.313Z",
          "content": "<p>noise in data. some people draw totally wrong object</p>",
          "rawMarkdown": "noise in data. some people draw totally wrong object",
          "votes": 14
        },
        {
          "id": 415756,
          "postDate": "2018-11-05T16:19:46.090Z",
          "content": "<p>That's the core reason why even after so much data, models are having difficulties in generalization</p>",
          "rawMarkdown": "That's the core reason why even after so much data, models are having difficulties in generalization"
        },
        {
          "id": 415762,
          "postDate": "2018-11-05T16:45:11.623Z",
          "content": "<p>It is true the smaller batch size is better, but here we have a lot of data so we can use more data with the larger batch size. I think the more used data helps not the bigger batch size itself.</p>",
          "rawMarkdown": "It is true the smaller batch size is better, but here we have a lot of data so we can use more data with the larger batch size. I think the more used data helps not the bigger batch size itself."
        },
        {
          "id": 427054,
          "postDate": "2018-11-24T12:49:39.647Z",
          "content": "<p>Refers to this code:\n<a href=\"https://discuss.pytorch.org/t/simulation-of-large-batch-size/3842/4\">https://discuss.pytorch.org/t/simulation-of-large-batch-size/3842/4</a></p>",
          "rawMarkdown": "Refers to this code:\nhttps://discuss.pytorch.org/t/simulation-of-large-batch-size/3842/4"
        }
      ]
    },
    {
      "id": 421529,
      "postDate": "2018-11-15T05:21:04.840Z",
      "content": "<p>Hi~ Heng\nWhat Data Augmentation do you use?</p>",
      "rawMarkdown": "Hi~ Heng\nWhat Data Augmentation do you use?",
      "votes": 2
    },
    {
      "id": 419505,
      "postDate": "2018-11-12T04:39:42.143Z",
      "content": "<p>What size image is large enough? 256x256? Is there any gain trying larger, like 512?</p>",
      "rawMarkdown": "What size image is large enough? 256x256? Is there any gain trying larger, like 512?",
      "votes": 2,
      "replies": [
        {
          "id": 419818,
          "postDate": "2018-11-12T15:42:31.923Z",
          "content": "<p>you need large batch size to combat noise, so 512x512 is difficult to train.</p>\n\n<p>There will be no gain or even worse results is the batch size is too small.</p>\n\n<p>imagenet uses batch size of 256.</p>\n\n<p>note that NASNET uses size up to about 331x331? some of the inception uses 229x229</p>\n\n<hr>\n\n<p>to clean the non-recognised drawings, you can train a smaller classifier first (e.g. 128x128) to label them.</p>\n\n<p>manual annotation is also a possible solution  but it would take some time.</p>\n\n<p>the best way is to use a classifier to rank the non-recognised drawings, and them use visual inspection to clean up the data. (human in the loop active learning)</p>",
          "rawMarkdown": "you need large batch size to combat noise, so 512x512 is difficult to train.\n\nThere will be no gain or even worse results is the batch size is too small.\n\nimagenet uses batch size of 256.\n\n\nnote that NASNET uses size up to about 331x331? some of the inception uses 229x229\n\n---\n\nto clean the non-recognised drawings, you can train a smaller classifier first (e.g. 128x128) to label them.\n\nmanual annotation is also a possible solution  but it would take some time.\n\nthe best way is to use a classifier to rank the non-recognised drawings, and them use visual inspection to clean up the data. (human in the loop active learning)",
          "votes": 4
        },
        {
          "id": 419844,
          "postDate": "2018-11-12T16:17:08.657Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 422921,
          "postDate": "2018-11-17T03:44:31.033Z",
          "content": "<p>Hi Heng is this cleaning meant for raw data ? Thanks</p>",
          "rawMarkdown": "Hi Heng is this cleaning meant for raw data ? Thanks"
        }
      ]
    },
    {
      "id": 417261,
      "postDate": "2018-11-08T02:15:53.040Z",
      "content": "<p>@CherKeng - just wondering you mentioned that 50,000k per class for 1 epoch could possibly be better than 50k per class for 100 epochs. In practice, if someone were to train only a single epoch is it better to start with a much small learning rate than if you were to train 100 epochs?</p>",
      "rawMarkdown": "@CherKeng - just wondering you mentioned that 50,000k per class for 1 epoch could possibly be better than 50k per class for 100 epochs. In practice, if someone were to train only a single epoch is it better to start with a much small learning rate than if you were to train 100 epochs?",
      "votes": 2,
      "replies": [
        {
          "id": 417412,
          "postDate": "2018-11-08T08:11:19.463Z",
          "content": "<p>Keng, I don't get the logic of why few long epochs are better than many short ones? If the total number of steps is still the same. </p>",
          "rawMarkdown": "Keng, I don't get the logic of why few long epochs are better than many short ones? If the total number of steps is still the same. ",
          "votes": 2
        },
        {
          "id": 418529,
          "postDate": "2018-11-10T03:35:03.967Z",
          "content": "<p>In this case, as Heng told, there's much noise in the data (a significant portion images are not made correctly), running over small epochs may hinder model's progress because that sample of data (from minibatch) might introduce more gradient flow than others. Including more samples/cycle helps distribute it to some extent (Not all samples are created equal).</p>",
          "rawMarkdown": "In this case, as Heng told, there's much noise in the data (a significant portion images are not made correctly), running over small epochs may hinder model's progress because that sample of data (from minibatch) might introduce more gradient flow than others. Including more samples/cycle helps distribute it to some extent (Not all samples are created equal).",
          "votes": 3
        },
        {
          "id": 418793,
          "postDate": "2018-11-10T16:12:00.603Z",
          "content": "<p>first, seeing more train samples is better than seeing less samples + their augmentation.</p>\n\n<p>for learning rate and batch size relationship and experimental results, refer to:</p>\n\n<p>\"Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour\"\n<a href=\"https://arxiv.org/pdf/1706.02677.pdf\">https://arxiv.org/pdf/1706.02677.pdf</a></p>\n\n<p>\"Linear Scaling Rule: When the minibatch size is\nmultiplied by k, multiply the learning rate by k.\"</p>\n\n<p>\"DON’T DECAY THE LEARNING RATE,   INCREASE THE BATCH SIZE\"</p>\n\n<p><a href=\"https://openreview.net/pdf?id=B1Yy1BxCZ\">https://openreview.net/pdf?id=B1Yy1BxCZ</a></p>\n\n<p>for batch size and noise relationship, there are some papers that on that (but i currently forget the title. you may google for it)</p>\n\n<p><a href=\"https://arxiv.org/pdf/1705.10694.pdf\">https://arxiv.org/pdf/1705.10694.pdf</a>\n\"Deep Learning is Robust to Massive Label Noise\"</p>\n\n<p>\"Finally, we have observed that noisy labels reduce\nthe effective batch size, an effect that can be mitigated by\nlarger batch sizes and downscaling the learning rate.\"</p>",
          "rawMarkdown": "first, seeing more train samples is better than seeing less samples + their augmentation.\n\nfor learning rate and batch size relationship and experimental results, refer to:\n\n\"Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour\"\nhttps://arxiv.org/pdf/1706.02677.pdf\n\n\"Linear Scaling Rule: When the minibatch size is\nmultiplied by k, multiply the learning rate by k.\"\n\n\"DON’T DECAY THE LEARNING RATE,   INCREASE THE BATCH SIZE\"\n\nhttps://openreview.net/pdf?id=B1Yy1BxCZ\n\n\n\nfor batch size and noise relationship, there are some papers that on that (but i currently forget the title. you may google for it)\n\nhttps://arxiv.org/pdf/1705.10694.pdf\n\"Deep Learning is Robust to Massive Label Noise\"\n\n\"Finally, we have observed that noisy labels reduce\nthe effective batch size, an effect that can be mitigated by\nlarger batch sizes and downscaling the learning rate.\"\n",
          "votes": 13
        },
        {
          "id": 419388,
          "postDate": "2018-11-11T21:25:03.633Z",
          "content": "<p>For example, if we cannot process 512 images as batch (out of memory), in Pytorch we could accumulate the cost in a variable and then when we have processed 512 do the \"cost.backward()\"??</p>\n\n<p>You suggest then use for example... 10 epochs with 256 batch size -&gt; 10 epochs with 512 batch size... and not 20 epochs with 256 batch size and modify every each 10 epochs the learning rate (annealing)... True?</p>",
          "rawMarkdown": "For example, if we cannot process 512 images as batch (out of memory), in Pytorch we could accumulate the cost in a variable and then when we have processed 512 do the \"cost.backward()\"??\n\nYou suggest then use for example... 10 epochs with 256 batch size -&gt; 10 epochs with 512 batch size... and not 20 epochs with 256 batch size and modify every each 10 epochs the learning rate (annealing)... True?",
          "votes": 1
        }
      ]
    },
    {
      "id": 416785,
      "postDate": "2018-11-07T09:24:57.140Z",
      "content": "<p>Thank you very much for your tips! One question. I wanted to try relatively large batch size (&gt;=256) and came across the problem - I get the \"ResourceExhaustedError\" when I run the code at my local machine with GPU. Seems like I run out of memory on GPU. What could you suggest? Any ideas how could I approach this problem? Of course I could run the code on Kaggle kernel, but in my opinion running the code at local machine is more comfortable.</p>\n\n<p>My settings:\nGeForce GTX1050 Ti Gigabyte PCIE 4096Mb, driver version 398.11\nTensorflow: version 1.8.0\nKeras: version 2.1.6</p>",
      "rawMarkdown": "Thank you very much for your tips! One question. I wanted to try relatively large batch size (&gt;=256) and came across the problem - I get the \"ResourceExhaustedError\" when I run the code at my local machine with GPU. Seems like I run out of memory on GPU. What could you suggest? Any ideas how could I approach this problem? Of course I could run the code on Kaggle kernel, but in my opinion running the code at local machine is more comfortable.\n\nMy settings:\nGeForce GTX1050 Ti Gigabyte PCIE 4096Mb, driver version 398.11\nTensorflow: version 1.8.0\nKeras: version 2.1.6",
      "votes": 2,
      "replies": [
        {
          "id": 417308,
          "postDate": "2018-11-08T04:26:08.197Z",
          "content": "<p>I think the limit is your GPU, sir.</p>",
          "rawMarkdown": "I think the limit is your GPU, sir."
        }
      ]
    },
    {
      "id": 415579,
      "postDate": "2018-11-05T10:54:45.190Z",
      "content": "<p>How much time does it take for you to train on entire dataset ?</p>",
      "rawMarkdown": "How much time does it take for you to train on entire dataset ?",
      "votes": 2
    },
    {
      "id": 427832,
      "postDate": "2018-11-26T08:32:03.260Z",
      "content": "<p>24g gpu memory is not enough for se resnext50 model</p>",
      "rawMarkdown": "24g gpu memory is not enough for se resnext50 model",
      "replies": [
        {
          "id": 427997,
          "postDate": "2018-11-26T15:26:22.070Z",
          "content": "<p>24G is already a very large GPU and from my experience, with img size 128, you should be OK to run se resnext50 model with batch size at least 256</p>",
          "rawMarkdown": "24G is already a very large GPU and from my experience, with img size 128, you should be OK to run se resnext50 model with batch size at least 256"
        },
        {
          "id": 429140,
          "postDate": "2018-11-28T11:57:37.740Z",
          "content": "<p>how big ur image and batch size is ?</p>",
          "rawMarkdown": "how big ur image and batch size is ?",
          "votes": 1
        },
        {
          "id": 429550,
          "postDate": "2018-11-29T02:15:52.510Z",
          "content": "<p>im using xception model! batch size=200 image size=128</p>",
          "rawMarkdown": "im using xception model! batch size=200 image size=128",
          "votes": -1
        },
        {
          "id": 429572,
          "postDate": "2018-11-29T03:08:28.707Z",
          "content": "<p><a href=\"/jonney\">@jonney</a> You can try 64x64, I get up to 0.925 with just 64x64 images</p>",
          "rawMarkdown": "@jonney You can try 64x64, I get up to 0.925 with just 64x64 images"
        }
      ]
    },
    {
      "id": 427114,
      "postDate": "2018-11-24T16:11:36.230Z",
      "content": "<p>hi, i'm new to data science. I do not have a powerful gpu and I teach at google colaboratory, but when I have a picture size of 256 * 256, I can’t install a large batch size, it gives an out of memory error. tell me where I can train large images online with a great batchsize? thanks</p>",
      "rawMarkdown": "hi, i'm new to data science. I do not have a powerful gpu and I teach at google colaboratory, but when I have a picture size of 256 * 256, I can’t install a large batch size, it gives an out of memory error. tell me where I can train large images online with a great batchsize? thanks",
      "replies": [
        {
          "id": 429734,
          "postDate": "2018-11-29T09:16:26.593Z",
          "content": "<p>Hi Matskevich,\nYou can try accumulating gradients,</p>\n\n<p>Some idea <a href=\"https://medium.com/huggingface/training-larger-batches-practical-tips-on-1-gpu-multi-gpu-distributed-setups-ec88c3e51255\">https://medium.com/huggingface/training-larger-batches-practical-tips-on-1-gpu-multi-gpu-distributed-setups-ec88c3e51255</a></p>\n\n<p>pytorch <a href=\"https://discuss.pytorch.org/t/how-to-implement-accumulated-gradient/3822\">https://discuss.pytorch.org/t/how-to-implement-accumulated-gradient/3822</a>\nkeras <a href=\"https://github.com/keras-team/keras/issues/3556\">https://github.com/keras-team/keras/issues/3556</a></p>",
          "rawMarkdown": "Hi Matskevich,\nYou can try accumulating gradients,\n\nSome idea https://medium.com/huggingface/training-larger-batches-practical-tips-on-1-gpu-multi-gpu-distributed-setups-ec88c3e51255\n\npytorch https://discuss.pytorch.org/t/how-to-implement-accumulated-gradient/3822\nkeras https://github.com/keras-team/keras/issues/3556\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 427052,
      "postDate": "2018-11-24T12:47:53.997Z",
      "content": "<p>Heng CherKeng thank you for sharing </p>\n\n<p>Questions:\n1. It means to use all epoch and test samples?\n2. should we use raw images?\n3. MobileNet is the right network to use? or refers to speed connection?\n4. Which batch size should we use? 4096K?</p>",
      "rawMarkdown": "Heng CherKeng thank you for sharing \n\nQuestions:\n1. It means to use all epoch and test samples?\n2. should we use raw images?\n3. MobileNet is the right network to use? or refers to speed connection?\n4. Which batch size should we use? 4096K?\n\n\n\n"
    },
    {
      "id": 422556,
      "postDate": "2018-11-16T11:59:08.097Z",
      "content": "<p>Hi, heng! Did you use multi GPUs for larger batch size? My batch size of ResNet50 can only reach to 64 (image224x224x3) on an TITANX GPU(12G).</p>",
      "rawMarkdown": "Hi, heng! Did you use multi GPUs for larger batch size? My batch size of ResNet50 can only reach to 64 (image224x224x3) on an TITANX GPU(12G).",
      "replies": [
        {
          "id": 422914,
          "postDate": "2018-11-17T03:20:35.747Z",
          "content": "<p>If you are using pytorch, you can accumulate the loss for 2 or 3 batch, then calculate the gradient</p>",
          "rawMarkdown": "If you are using pytorch, you can accumulate the loss for 2 or 3 batch, then calculate the gradient",
          "votes": 6
        },
        {
          "id": 423050,
          "postDate": "2018-11-17T10:56:13.707Z",
          "content": "<p>That's a good idea, thanks!</p>",
          "rawMarkdown": "That's a good idea, thanks!"
        },
        {
          "id": 423272,
          "postDate": "2018-11-17T20:35:28.767Z",
          "content": "<p>How you did it Strideradu? I write my suggestion here <a href=\"https://discuss.pytorch.org/t/simulation-of-large-batch-size/3842/4\">https://discuss.pytorch.org/t/simulation-of-large-batch-size/3842/4</a> but not getting same results 64-&gt;64 batch as 128 batch updates on weights</p>",
          "rawMarkdown": "How you did it Strideradu? I write my suggestion here https://discuss.pytorch.org/t/simulation-of-large-batch-size/3842/4 but not getting same results 64-&gt;64 batch as 128 batch updates on weights"
        },
        {
          "id": 423324,
          "postDate": "2018-11-17T23:19:46.110Z",
          "content": "<p><a href=\"https://discuss.pytorch.org/t/pytorch-gradients/884/2\">https://discuss.pytorch.org/t/pytorch-gradients/884/2</a> I didn't try this for this competetion since now I have accecss to a large memory GPU. But I remembered many people using this for the Carvana competetion</p>",
          "rawMarkdown": "https://discuss.pytorch.org/t/pytorch-gradients/884/2 I didn't try this for this competetion since now I have accecss to a large memory GPU. But I remembered many people using this for the Carvana competetion"
        }
      ]
    },
    {
      "id": 421567,
      "postDate": "2018-11-15T06:17:54.780Z",
      "content": "<p>Large image means reize (64,64) -&gt;(128,128) or even bigger\nor pad_and_resize?</p>",
      "rawMarkdown": "Large image means reize (64,64) -&gt;(128,128) or even bigger\nor pad_and_resize?"
    },
    {
      "id": 415516,
      "postDate": "2018-11-05T08:50:44.980Z",
      "content": "<p>Thanks for sharing this , but 2 and 3 seems kinda ambiguous, could be provide more info regarding size (you mean 256)?</p>",
      "rawMarkdown": "Thanks for sharing this , but 2 and 3 seems kinda ambiguous, could be provide more info regarding size (you mean 256)?",
      "replies": [
        {
          "id": 415621,
          "postDate": "2018-11-05T11:59:42.933Z",
          "content": "<p>@Prajjwal, I was able to get 0.928 on LB with image size of 128. Trying out with 256, 224 now. </p>",
          "rawMarkdown": "@Prajjwal, I was able to get 0.928 on LB with image size of 128. Trying out with 256, 224 now. ",
          "votes": 2
        },
        {
          "id": 415635,
          "postDate": "2018-11-05T12:29:52.070Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 415965,
          "postDate": "2018-11-06T01:28:26.670Z",
          "content": "<p>How much data are you using? I used Grayscale Mobilenet on 128*128 images, but my score didn't improved :-/ 30k images/class compared to 64*64</p>",
          "rawMarkdown": "How much data are you using? I used Grayscale Mobilenet on 128*128 images, but my score didn't improved :-/ 30k images/class compared to 64*64"
        }
      ]
    },
    {
      "id": 430211,
      "postDate": "2018-11-30T03:09:24.987Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true
    },
    {
      "id": 423805,
      "postDate": "2018-11-19T05:00:13.213Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true
    },
    {
      "id": 416161,
      "postDate": "2018-11-06T09:49:22.493Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2451789,
      "postDate": "2023-09-22T19:23:53.913Z",
      "content": "<p>thanks for this code</p>",
      "rawMarkdown": "thanks for this code"
    }
  ],
  "comments": [
    {
      "id": 423426,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-11-18T08:05:51.523000",
      "content": "<p>baseline performance:</p>\n\n<p>LB: 0.944</p>\n\n<p>se-resnext-50 (256x256 input, 200 samples per batch), single model, no TTA , no ensmble</p>\n\n<p>all train samples (simplified only, recognized and non-recognized): random 80 images per class are selected as validation, the rest are training samples.</p>\n\n<p>validation loss : ce_loss, top1, top3,  (map@3)</p>\n\n<p>0.607  0.842  0.946  (0.890)*  </p>\n\n<p>train loss : ce_loss, top1, top3,  (map@3)</p>\n\n<p>0.590  0.844  0.948  (0.892) </p>\n\n<hr>\n\n<p>LB 0.928</p>\n\n<p>customized LSTM\n(Max seq length = 600, 200 samples per batch), single model, no TTA , no ensmble</p>\n\n<p>same train samples as above</p>\n\n<p>validation loss : ce_loss, top1, top3,  (map@3)</p>\n\n<p>0.637  0.834  0.943  (0.885)*  </p>\n\n<p>train loss : ce_loss, top1, top3,  (map@3)</p>\n\n<p>0.595  0.841  0.949  (0.892)</p>\n\n<hr>\n\n<p>LB 0.945</p>\n\n<p>stacking via a fuse classifier on features from cnn and lstm. , no TTA , no ensmble</p>\n\n<p>same train samples as above. batch size increase to 512</p>\n\n<p>validation loss : ce_loss, top1, top3,  (map@3)</p>\n\n<p>0.588  0.850  0.947  (0.895)*  </p>\n\n<p>train loss : ce_loss, top1, top3,  (map@3)</p>\n\n<p>0.549  0.850  0.949  (0.896)</p>",
      "votes": 22,
      "replies": [
        {
          "id": 423428,
          "author_name": "TPloveYXT520",
          "author_url": "",
          "post_date": "2018-11-18T08:09:42.780000",
          "content": "<p>Hi~\nHeng\nWhat is your batch_size</p>",
          "votes": -4,
          "replies": []
        },
        {
          "id": 423443,
          "author_name": "upup",
          "author_url": "",
          "post_date": "2018-11-18T08:47:13.543000",
          "content": "<p>200 samples per batch...</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 423444,
          "author_name": "upup",
          "author_url": "",
          "post_date": "2018-11-18T08:48:30.520000",
          "content": "<p>Hi, Heng, how much time did your model spend for a epoch ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 423450,
          "author_name": "TPloveYXT520",
          "author_url": "",
          "post_date": "2018-11-18T09:21:04.663000",
          "content": "<pre><code>state_dict[key] = pretrain_state_dict[key.replace('resnet.layer2.','layer2.')]\n</code></pre>\n\n<p>KeyError: 'layer2.0.bn1.num_batches_tracked'</p>\n\n<p>Hi~ Heng \nWhen I use your code I got this error at load_pretrained file in the run_check_net function</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 423477,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2018-11-18T10:48:54.167000",
          "content": "<p>this is due to different version of pytorch.</p>\n\n<p><a href=\"https://discuss.pytorch.org/t/unexpected-key-in-state-dict-bn1-num-batches-tracked/29454\">https://discuss.pytorch.org/t/unexpected-key-in-state-dict-bn1-num-batches-tracked/29454</a></p>\n\n<p>just ignore the key 'numbatchestracked'</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 423491,
          "author_name": "Nazim Girach",
          "author_url": "",
          "post_date": "2018-11-18T11:35:59.933000",
          "content": "<p>Hey Heng, what kind of dataloader are you using? drawing them on the fly or saving images and then loading them while training. \nThanks. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 423575,
          "author_name": "Strideradu",
          "author_url": "",
          "post_date": "2018-11-18T15:28:57.923000",
          "content": "<p>For the LSTM, do you using the raw data instead of simplified data?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 423732,
          "author_name": "HuyenNguyen",
          "author_url": "",
          "post_date": "2018-11-19T00:25:41.037000",
          "content": "<p>Thanks Heng, what learning rate schedule do you use? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 423845,
          "author_name": "Yibing Wu",
          "author_url": "",
          "post_date": "2018-11-19T06:17:25.347000",
          "content": "<p>@Heng CherKeng， Thanks for all the tips that are very useful. I saw you only used batch_size 200, which seems quite small considering the dataset is quite noisy.  Were your trainset manually corrected or you just use the original simplified dataset? Thank you so much.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 425317,
          "author_name": "HuyenNguyen",
          "author_url": "",
          "post_date": "2018-11-21T12:44:53.793000",
          "content": "<p>I tried running this with Keras on a p3.2xlarge AWS instance and it took 2 hours to finish 1 epoch of 250k images, so 400 hours to go through the entire dataset! :( how can you train this in a reasonable amount of time??</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 429024,
          "author_name": "good good study",
          "author_url": "",
          "post_date": "2018-11-28T08:35:56.853000",
          "content": "<p>Hi, i trained 256*256 se-resnext50 with batch size 400. i have got better validation results on local LB (0.894 map@3, which is better than 0.890), but I can not get 0.944(as you have stated) on public LB, public LB is only 0.942. What can be the possible reason? It would be great help if you can give me some advice. Thanks a lot.</p>\n\n<p>ps: I use your pytorch start code as my code base. Thanks again.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 429271,
          "author_name": "remidi",
          "author_url": "",
          "post_date": "2018-11-28T15:55:13.050000",
          "content": "<p>@ 凉宫ハルヒ</p>\n\n<p>what is the stroke encoding and lr schedule that you use ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 429537,
          "author_name": "good good study",
          "author_url": "",
          "post_date": "2018-11-29T01:55:09.747000",
          "content": "<p>I use the stroke encoding provided by Heng's start code. \nlr schedule: SGD decay 0.5 at 800k steps</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 429931,
          "author_name": "remidi",
          "author_url": "",
          "post_date": "2018-11-29T15:05:43.437000",
          "content": "<p>@凉宫ハルヒ</p>\n\n<p>Thanks for your valuable insight. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 423502,
      "author_name": "omallo",
      "author_url": "",
      "post_date": "2018-11-18T12:12:59.193000",
      "content": "<p>Regarding the efficient usage of input channels (3 or even more if you use a custom model), the following paper might give some inspiration: <a href=\"https://arxiv.org/pdf/1501.07873.pdf\">https://arxiv.org/pdf/1501.07873.pdf</a></p>",
      "votes": 11,
      "replies": [
        {
          "id": 423894,
          "author_name": "Gary",
          "author_url": "",
          "post_date": "2018-11-19T08:02:28.170000",
          "content": "<p>HI, Omallo. Did you use 6 channels as input from paper? And if you don't mind. Coud you share what the best model did you use?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 424280,
          "author_name": "omallo",
          "author_url": "",
          "post_date": "2018-11-19T20:39:01.163000",
          "content": "<p>I'm currently using 2 models, one self designed model which is weaker but much faster for training and one deeper model (SeResNext50).</p>\n\n<p>With my custom model, I use the 6 channels as described in the paper. Compared with using just 1 channel, this gave me a LB boost of +0.015.</p>\n\n<p>With the SeResNext50 model, I use only 3 channels where I draw 1/3, 2/3, and 3/3 of the strokes, respectively. Compared with using just 1 channel, this gave me a LB boost of +0.013.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 424299,
          "author_name": "Strideradu",
          "author_url": "",
          "post_date": "2018-11-19T21:36:34.407000",
          "content": "<p>I am wondering when you predict the test set, are u still using these method or you just have 3 copy of same image?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 424313,
          "author_name": "omallo",
          "author_url": "",
          "post_date": "2018-11-19T21:58:19.133000",
          "content": "<p>I draw the images on the fly from the stroke data while training and while evaluating on the validation and test sets.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 424356,
          "author_name": "Gary",
          "author_url": "",
          "post_date": "2018-11-20T01:16:22.017000",
          "content": "<p>nice, thank you, Did you use lr schedule？what schedule?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 424454,
          "author_name": "HuyenNguyen",
          "author_url": "",
          "post_date": "2018-11-20T06:30:42.487000",
          "content": "<p>Omallo, what's your experience with the speed of training of SeResnext50? I tried running it on the Kaggle kernel and it was extremely slow. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 424529,
          "author_name": "gody7334",
          "author_url": "",
          "post_date": "2018-11-20T09:14:06.267000",
          "content": "<p>HuyenNguyen\nI still run SeResnext 50 on 256*256 img, I estimated it will take 2 weeks to go through a epoch on gtx 1070.\nyou can try half precision first, then multiple GPU if possible. Hopely I can finished in 5 days.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 424567,
          "author_name": "HuyenNguyen",
          "author_url": "",
          "post_date": "2018-11-20T10:30:41.550000",
          "content": "<p>omg, 1 week for 1 epoch @.@ how do you do half precision?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 424674,
          "author_name": "gody7334",
          "author_url": "",
          "post_date": "2018-11-20T14:18:05.513000",
          "content": "<p>Im using pytorch, pytorch make it easy to do just call, <strong>half()</strong>  or <strong>float()</strong> to convert to 16fp and 32fp\n<a href=\"https://discuss.pytorch.org/t/training-with-half-precision/11815\">https://discuss.pytorch.org/t/training-with-half-precision/11815</a>\njust beware the adam optimizer, \n<a href=\"https://discuss.pytorch.org/t/adam-half-precision-nans/1765\">https://discuss.pytorch.org/t/adam-half-precision-nans/1765</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 424759,
          "author_name": "omallo",
          "author_url": "",
          "post_date": "2018-11-20T16:12:11.763000",
          "content": "<p>@Gary-DeepLearning I'm using a cyclical lr scheduling using cosine annealing as described in the SGDR paper: <a href=\"https://arxiv.org/pdf/1608.03983.pdf\">https://arxiv.org/pdf/1608.03983.pdf</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 424762,
          "author_name": "omallo",
          "author_url": "",
          "post_date": "2018-11-20T16:19:52.660000",
          "content": "<p>@HuyenNguyen I'm using 128x128 images and 90% of the samples for training. Under these conditions, it takes about 30h for training 1 epoch on a NVIDIA Quadro P6000.</p>\n\n<p>Using 256x256 images, it takes around 4x longer so I sticked with 128x128 images.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 424948,
          "author_name": "HuyenNguyen",
          "author_url": "",
          "post_date": "2018-11-20T23:25:38.283000",
          "content": "<p>Does cyclical lr help compared to just ReduceonPlateau?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 424970,
          "author_name": "gody7334",
          "author_url": "",
          "post_date": "2018-11-21T00:19:14.497000",
          "content": "<p>Hi HuyenNguyen,\nMay I update my suggestion:\nfp16 is slow compare to fp32, i speculate older gpu don't have many fp16 unit, maybe latest rtx gpu will do the job better,\nand fp16 also cause many problem because you have to take care numerical stability yourself, \nDistributed training might be the way to go</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 425006,
          "author_name": "Gary",
          "author_url": "",
          "post_date": "2018-11-21T01:55:44.477000",
          "content": "<p>Oh, I see. Thank you.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 425159,
          "author_name": "omallo",
          "author_url": "",
          "post_date": "2018-11-21T08:01:48.757000",
          "content": "<p>@HuyenNguyen In general, with cyclical lr scheduling models tend to converge faster even though you might end up with slightly worse performance of the model. Often I use cyclical lr scheduling and, at the end, one run with a constant lr or use ReduceOnPlateau to do fine tuning. One thing I like about cyclical lr scheduling is that you can easily combine it with snapshot ensembling.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 425322,
          "author_name": "HuyenNguyen",
          "author_url": "",
          "post_date": "2018-11-21T12:48:23.030000",
          "content": "<p>Great, thank you for the tips. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 425323,
          "author_name": "HuyenNguyen",
          "author_url": "",
          "post_date": "2018-11-21T12:49:16",
          "content": "<p>Hi Omallo, when you say \"I draw 1/3, 2/3, and 3/3 of the strokes\" do you still give stroke a different colour? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 426052,
          "author_name": "omallo",
          "author_url": "",
          "post_date": "2018-11-22T14:13:51.647000",
          "content": "<p>@HuyenNguyen Yes, I sill use different colors for the strokes. Here is an example of using 6 channels as described in the paper: <a href=\"https://ibb.co/dGizNA\">https://ibb.co/dGizNA</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 427017,
          "author_name": "Siyuan Dang",
          "author_url": "",
          "post_date": "2018-11-24T10:53:12.533000",
          "content": "<p>Hi, omallo. After treated the image data like this,1/3, 2/3, 3/3. Does it cost you more steps or more epoch to get the result?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 427028,
          "author_name": "omallo",
          "author_url": "",
          "post_date": "2018-11-24T11:24:34.897000",
          "content": "<p>If I remember correctly, there was no noticeable difference, especially if you use a pretrained model</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 427065,
          "author_name": "Siyuan Dang",
          "author_url": "",
          "post_date": "2018-11-24T13:31:21.950000",
          "content": "<p>Thanks, omallo. After prepocessing the img data as you do, when I train my model, I find the loss is large than the 1 channels model in the early stage. This condition makes me feel very strange. Maybe some of my preprocessing work has some mistakes in it.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 427103,
          "author_name": "Mario Parreño Lara",
          "author_url": "",
          "post_date": "2018-11-24T15:29:29.550000",
          "content": "<p>Hi guys! I have tried taht and give me worst results (no so much) using 6 channels in White/Black... I divided the points / number of channels (doing this too for getting part1+part2 and so on)... How you did it in colour? The number of strokes is not uniform and for example If I want 3 fragments there are images with 1 \"movement\" for all the picture.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 427116,
          "author_name": "Mario Parreño Lara",
          "author_url": "",
          "post_date": "2018-11-24T16:22:50.203000",
          "content": "<p>So... for this <a href=\"https://ibb.co/dGizNA\">https://ibb.co/dGizNA</a> you have 6 images * 3 channels =&gt; 18 channels as input</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 427130,
          "author_name": "omallo",
          "author_url": "",
          "post_date": "2018-11-24T17:11:01.037000",
          "content": "<p>Mario, I just realized that my image might be somewhat confusing. Actually, for the example with 6 images, the colors are just different grayscale values between 0 and 240 so it's not a RGB color. The images look colored since I used the default color map of matplotlib to render them.</p>\n\n<p>This means that you would only need 6 channels in that case. Sorry for the confusion.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 427132,
          "author_name": "Mario Parreño Lara",
          "author_url": "",
          "post_date": "2018-11-24T17:21:08.383000",
          "content": "<p>Thanks omallo! Unfourtunately I tried and obtained slightly worst results. But thanks anyway, here I am to learn :D</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 429119,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-11-28T11:29:47.570000",
      "content": "<p>another reference results</p>\n\n<p>LB: did not submit (estimated to be 0.942)</p>\n\n<p>senet154 (128x128 input, 256 samples per batch), single model, no TTA , no ensmble</p>\n\n<p>all train samples (simplified only, recognized and non-recognized): random 80 images per class are selected as validation, the rest are training samples.</p>\n\n<p>validation loss : ce_loss, top1, top3, (map@3)</p>\n\n<p>0.622  0.839  0.944  (0.888)* </p>\n\n<p>train loss : ce_loss, top1, top3, (map@3)</p>\n\n<p>0.582  0.848  0.946  (0.893)  </p>",
      "votes": 5,
      "replies": [
        {
          "id": 429265,
          "author_name": "remidi",
          "author_url": "",
          "post_date": "2018-11-28T15:50:23.573000",
          "content": "<p>@Heng</p>\n\n<p>What is your lr schedule ? and stroke encoding ?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 415742,
      "author_name": "Serigne ",
      "author_url": "",
      "post_date": "2018-11-05T15:42:24.273000",
      "content": "<p>Thanks Heng </p>\n\n<p>I still use 64*64 ....may be it's time to use larger image  :)</p>\n\n<p>I am also  working on encoding more information using all the 3 channels. </p>\n\n<p>BTW , do you think there are some augmentations that help ? </p>",
      "votes": 5,
      "replies": [
        {
          "id": 418494,
          "author_name": "Gary",
          "author_url": "",
          "post_date": "2018-11-10T01:33:30.223000",
          "content": "<p>Hi, how did you encode more information with using all the 3 channels?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 415551,
      "author_name": "The16thRoute",
      "author_url": "",
      "post_date": "2018-11-05T09:57:38.987000",
      "content": "<p>Now I use \n1. 50k/class\n2. img_size 224*224\n3. res50\n4. 64*8\ncould got 0.936 and 0.938 with ensemble</p>\n\n<p>Due to some bugs, I did not got any new models with new setting for two weeks. Sign~</p>\n\n<p>Now I will try:\n1. all imgs\n2. img_size: 448*448\n3. Senet?\n4. 256*8</p>",
      "votes": 5,
      "replies": [
        {
          "id": 415559,
          "author_name": "Nazim Girach",
          "author_url": "",
          "post_date": "2018-11-05T10:12:19.887000",
          "content": "<p>Don't you think 448 is too large? 256 should be more than enough according to me.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 415574,
          "author_name": "Prajjwal",
          "author_url": "",
          "post_date": "2018-11-05T10:45:44.940000",
          "content": "<p>This is so much compute, also if have such specs, consider using NasNet then. It should give better results as per benchamarks.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 415580,
          "author_name": "The16thRoute",
          "author_url": "",
          "post_date": "2018-11-05T10:55:08.187000",
          "content": "<p>recently, i wrote a distributed framework, so 448*448 do not take a lot of time</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 415624,
          "author_name": "Nazim Girach",
          "author_url": "",
          "post_date": "2018-11-05T12:04:27.533000",
          "content": "<p>Are are you using multiple GPUs?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 415633,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2018-11-05T12:25:29.467000",
          "content": "<p>@Li Bin</p>\n\n<p>e.g.</p>\n\n<p>1 epoch of \"50,000k/class\"  may be better than 100 epoch of \"50k/class\"</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 415638,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2018-11-05T12:31:53.637000",
          "content": "<p>you should be able to get lb ~0.940 with 224x224. 256x256 is better</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 415668,
          "author_name": "The16thRoute",
          "author_url": "",
          "post_date": "2018-11-05T13:20:24.993000",
          "content": "<p>Thank your for your suggestion. It is very useful for me.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 415693,
          "author_name": "Nazim Girach",
          "author_url": "",
          "post_date": "2018-11-05T14:11:17.997000",
          "content": "<blockquote>\n  <p>1 epoch of \"50,000k/class\" may be better than 100 epoch of \"50k/class\"</p>\n</blockquote>\n\n<p>@Heng do you mean use high number of steps_per_epoch? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 415964,
          "author_name": "HuyenNguyen",
          "author_url": "",
          "post_date": "2018-11-06T01:23:34.107000",
          "content": "<p>Training Resnet50 on that amount data of that image size will take forever. what machine are you using? </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 416191,
          "author_name": "sblroid",
          "author_url": "",
          "post_date": "2018-11-06T11:19:37.557000",
          "content": "<p>try GCP preemptible instances with free trial</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 416388,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-11-06T15:56:44.130000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 416522,
          "author_name": "RDizzl3",
          "author_url": "",
          "post_date": "2018-11-06T19:58:15.117000",
          "content": "<p>I'm showing a single epoch at that image size (224, 224, 3) and all layers trainable at appx 2 days for ResNet50 - this is on an AWS gpu instance with one GPU. This is will all available data from the Shuffle CSV data sets. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 416596,
          "author_name": "HuyenNguyen",
          "author_url": "",
          "post_date": "2018-11-06T23:59:56.687000",
          "content": "<p>A batchsize of 2048 images of size 224x224x3, how can you fit that in memory?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 416598,
          "author_name": "RDizzl3",
          "author_url": "",
          "post_date": "2018-11-07T00:17:52.373000",
          "content": "<p>my batch size is only 64. I could go a little higher but not much. That is another reason for the slower training that I reported.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 416612,
          "author_name": "RDizzl3",
          "author_url": "",
          "post_date": "2018-11-07T01:21:40.907000",
          "content": "<p>@Li Bin - just curious how many epochs did it take to reach that point? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 416631,
          "author_name": "The16thRoute",
          "author_url": "",
          "post_date": "2018-11-07T02:46:44.580000",
          "content": "<p>12 epoch</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 416654,
          "author_name": "chen",
          "author_url": "",
          "post_date": "2018-11-07T03:30:33.283000",
          "content": "<p>what does 256*8 mean? I am a new-bee. Thx. XD</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 416667,
          "author_name": "The16thRoute",
          "author_url": "",
          "post_date": "2018-11-07T03:53:47.523000",
          "content": "<p>256 batch per gpu, and I have 8 gpus</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 416672,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2018-11-07T04:05:58.750000",
          "content": "<p>64 batch size is too small</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 416674,
          "author_name": "The16thRoute",
          "author_url": "",
          "post_date": "2018-11-07T04:14:41.413000",
          "content": "<p>@CherKeng 256 batch/gpu, the batchsize totally is 256*8</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 416683,
          "author_name": "RDizzl3",
          "author_url": "",
          "post_date": "2018-11-07T04:54:04.040000",
          "content": "<p>@CherKeng - yes I figured that this might be to small to train in a reasonable amount of time. Do you have any suggestions to be able to increase my batch size on a single GPU?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 416854,
          "author_name": "Leon Chang",
          "author_url": "",
          "post_date": "2018-11-07T11:21:03.500000",
          "content": "<p>I  use two TitanXP and set the batch size = 256 then keras report me OOM. :( <br>\n(image size is 256*256 and keras resnet50)\nSo am I set something wrong?  could you tell me what kind of GPU do you used? Thx  </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 416868,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2018-11-07T11:33:21.347000",
          "content": "<p>@RDizzl3</p>\n\n<p>May be you should set very large Steps per Epoch ? </p>\n\n<p>I think 224*224 is way too large for this data, but given people say it gives better results, why not ! :)</p>\n\n<p>In the model I am currently runnung , I set 128*128 for image size  and batchsize 448.  That's the best compromise I could find.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 425327,
          "author_name": "HuyenNguyen",
          "author_url": "",
          "post_date": "2018-11-21T12:56:47.087000",
          "content": "<p>Li Bin, are you saying 12 epochs, and go through the full dataset in each epoch?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 415477,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-11-05T07:48:53.463000",
      "content": "<p>typical training log:\n(1 iter = 1000 iterations, see my pytorch starter kit)</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/415477/10604/train_log.png\" alt=\"enter image description here\"></p>",
      "votes": 6,
      "replies": [
        {
          "id": 415528,
          "author_name": "Finlay",
          "author_url": "",
          "post_date": "2018-11-05T09:14:29.693000",
          "content": "<p>thank you, sir!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 416202,
          "author_name": "Joe Ho",
          "author_url": "",
          "post_date": "2018-11-06T11:49:53.510000",
          "content": "<p>Hi CherKeng, why your log have such high top 3 acc (0.945)but low lb(0.889), and loss corresponds to cross entropy softmax loss?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 417291,
          "author_name": "Vipul Rai",
          "author_url": "",
          "post_date": "2018-11-08T03:22:20.787000",
          "content": "<p>Thank you !</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 418821,
          "author_name": "Tim Joseph",
          "author_url": "",
          "post_date": "2018-11-10T17:05:07.363000",
          "content": "<p>Hi Heng,\nis this log starting from just a normal imagenet pretrained model or continueing from some checkpoint?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 425072,
          "author_name": "Sanjay Kumar",
          "author_url": "",
          "post_date": "2018-11-21T05:10:07.137000",
          "content": "<p>It is continuing from checkpoint because 0.61 val loss will give leader board score of around 0.938. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 426306,
          "author_name": "kdaqlkgjnmaklvmnkankl",
          "author_url": "",
          "post_date": "2018-11-23T03:22:05.780000",
          "content": "<p>Do you think there is a gap between local score and public LB score? If so, why? In all of my models, I'm getting local validation score of ~0.05 lower than public LB.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 421429,
      "author_name": "Jungwoo Park",
      "author_url": "",
      "post_date": "2018-11-15T01:54:35.503000",
      "content": "<p>Hi Heng, thanks for your great tips! I've trained ResNet101 for 3 epochs with 64x64 scale and 1024 batch size and it scored 0.925. Does the image size really matter?</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 419753,
      "author_name": "gody7334",
      "author_url": "",
      "post_date": "2018-11-12T14:14:07.340000",
      "content": "<p>Hi Heng,\nI am currently running on gtx 1080 8GB, using resnet 50, with 64x64 image size, 128 batch size, \nI am using 90% data as training dataset. it take 1 day to go through 1 epoch.\nMy questions are:\n1. Does the training time look right? or can you suggest a training time baseline for 1 epoch?\n2. Is it better to train using 10K per class with 10 epoch, rather than 100K per class with 1 epoch? \nBecause my gpu is quite slow.\nThanks</p>",
      "votes": 3,
      "replies": [
        {
          "id": 419805,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2018-11-12T15:26:11.157000",
          "content": "<p>my validation set is 80 per class. All the rest are train images. Whatever you do, you should get</p>\n\n<p>0.886 for local LB  to have a 0.940 LB.\nThis is about local Top1-acc = 0.838, Top3-acc = 0.943,</p>\n\n<hr>\n\n<p>(however, i don't think 64x64 can reach that accuracy. best of 64x64 i can get is about LB 0.934, but i do not use resnet50)</p>\n\n<p>image-net resnet uses batch of 256 for 1000 classes. maybe 128 batch is ok (i haven't tried batch size 128, as i use larger than 128) </p>\n\n<hr>\n\n<p>about 2 epoch my be sufficient. With such large dataset, i do not use epoch to monitor my training. instead i use training iterations. i.e the learning rate are decreased based on number of iterations and not epoch</p>\n\n<p>You can save your model at every e.g. 1 hours and record the validation loss. Then you decide what to make submission. you should be able to submit after training for 24 to 48 hrs?</p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 419808,
          "author_name": "Gary",
          "author_url": "",
          "post_date": "2018-11-12T15:31:47.003000",
          "content": "<p>Hi, heng, your best score was based on CNN?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 419810,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2018-11-12T15:33:58.570000",
          "content": "<p>LB 0.943~0.944 without ensemble</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 419812,
          "author_name": "Gary",
          "author_url": "",
          "post_date": "2018-11-12T15:35:43.090000",
          "content": "<p>ok. thank you.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 419830,
          "author_name": "gody7334",
          "author_url": "",
          "post_date": "2018-11-12T15:56:32.083000",
          "content": "<p>I have some idea where i might do wrong,\nThanks you very much CherKeng</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 419730,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-11-12T13:36:29.190000",
      "content": "<p>encode by time(left)</p>\n\n<p>encode by stroke/part (right)</p>\n\n<p>you can tnink of a way to encode \"pause time\"</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/419730/10659/0.73_9000297423888044_baseball_bat_bat_baseball.png\" alt=\"enter image description here\"></p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 418071,
      "author_name": "Brian Lee",
      "author_url": "",
      "post_date": "2018-11-09T08:46:43.257000",
      "content": "<p>Hi Heng,</p>\n\n<p>Amazing tips as usual. I am currently running on gtx 1060 6GB, and so I cannot quite increase both image size and batch size as it would get OOM. Which of the two should I priortize? Imagesize or batch?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 418762,
          "author_name": "Strideradu",
          "author_url": "",
          "post_date": "2018-11-10T15:12:57.867000",
          "content": "<p>There is trick that you can accumulate the gradient for several batches and then update the gradient, so the actual batch size is much bigger</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 418978,
          "author_name": "Brian Lee",
          "author_url": "",
          "post_date": "2018-11-11T02:36:05.400000",
          "content": "<p>Interesting, found this thread about it so far: <a href=\"https://github.com/keras-team/keras/issues/3556\">https://github.com/keras-team/keras/issues/3556</a></p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 425650,
      "author_name": "gody7334",
      "author_url": "",
      "post_date": "2018-11-21T22:48:58.020000",
      "content": "<p>Hi Heng,\nMay I ask what optimizer you used?\nDo you think Adam + Cyclic LR is good idea, \nMy Top 3 accu can easy get 0.96 using this combination but with poor public LB score......</p>",
      "votes": 1,
      "replies": [
        {
          "id": 425749,
          "author_name": "upup",
          "author_url": "",
          "post_date": "2018-11-22T03:37:27.693000",
          "content": "<p>Did you use all dataset?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 426016,
          "author_name": "gody7334",
          "author_url": "",
          "post_date": "2018-11-22T12:43:56.303000",
          "content": "<p>Yes, 80 samples/class for validation, rest are training data\nAnd it can reach 0.96 top 3 accu, only training on 0.1 epoch,\nI also check the train val dataset, there is no overlap between.\nI don't see any leakage, but cannot 100% sure,\nanyone have the same experience?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 426282,
          "author_name": "MamadaNaoya",
          "author_url": "",
          "post_date": "2018-11-23T02:16:45.587000",
          "content": "<p>I have.\nIf you use multiple GPUs and do not shuffle data you feed to your gpus, you should.\nIn this competition we have 340 classes. I splitted them into 4 pieces and fed each of them to each gpu. I mean,  I fed samples of The Effel Tower, ..., Drill to gpu A,  samples of drums, ..., laptop to gpu B, the same hereafter.\nThe deep learning library would have averaged losses calculated by all gpus,  however, resulting model unfairly predicted classes that fed to gpu A.\nShuffling data fed to gpus solved this problem, at least for me.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 421562,
      "author_name": "TPloveYXT520",
      "author_url": "",
      "post_date": "2018-11-15T06:03:57.720000",
      "content": "<p>Hi~ Heng\nHow to get the 3 channel rgb?\nShould we  draw the every line with different color?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 418503,
      "author_name": "Huang, Shuang",
      "author_url": "",
      "post_date": "2018-11-10T01:54:20.640000",
      "content": "<p>Hi Heng,</p>\n\n<p>Could you tell more about \"you have 3 channel rgb, make good use of it! try to encode more information\"? you mean use the bigger filters?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 418891,
          "author_name": "Paul Jurczak",
          "author_url": "",
          "post_date": "2018-11-10T20:14:08.247000",
          "content": "<p>Try using stroke velocity in 2nd channel and perhaps stroke acceleration in the 3rd channel.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 418933,
          "author_name": "HuyenNguyen",
          "author_url": "",
          "post_date": "2018-11-10T22:45:50.580000",
          "content": "<p>There's no temporal info about the different points on a stroke though. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 420187,
          "author_name": "Paul Jurczak",
          "author_url": "",
          "post_date": "2018-11-13T08:29:32.383000",
          "content": "<p>There are time stamps for each point in raw data.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 425174,
          "author_name": "chunduri",
          "author_url": "",
          "post_date": "2018-11-21T08:42:44.770000",
          "content": "<p>I see that time stamp is for each row, not each point in each row. \nPlease clarify this. </p>\n\n<p>how can we use timestamp as part of training data?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 416720,
      "author_name": "Nazim Girach",
      "author_url": "",
      "post_date": "2018-11-07T06:58:56.077000",
      "content": "<p>What is the batch size did you use?  I could not get higher than 64 for 224*224 images for 100000/class. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 429576,
          "author_name": "Hoàng Tùng Lâm",
          "author_url": "",
          "post_date": "2018-11-29T03:11:45.860000",
          "content": "<p>Use smaller image, priortize batch size first </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 415835,
      "author_name": "soulkwon",
      "author_url": "",
      "post_date": "2018-11-05T19:22:44.960000",
      "content": "<p>Hi all, thanks for the great discussion.</p>\n\n<p>I am curious how many batch size a Titan X can handle, for example, when using 224x224x3 and ResNet50. I guessed around ~16. Does this make sense or it could be more?</p>\n\n<p>@Li bin mentioned using 256*8 of 224x224x3 input (and also mentioned using a distributed framework.) What would be a typical capacity for a Titan X (as an example)?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 415850,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-11-05T19:59:08.810000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 416811,
          "author_name": "The16thRoute",
          "author_url": "",
          "post_date": "2018-11-07T10:01:01.983000",
          "content": "<p>Sorry， I made a big mistake. For res34 we could have 256batch/gpu, for res50, we could only have 64batch/gpu.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 415681,
      "author_name": "Joe Ho",
      "author_url": "",
      "post_date": "2018-11-05T13:35:10.773000",
      "content": "<p>Just curious why large batch size helps, some paper mentioned that smaller batch size should generalize better than larger batch size...</p>",
      "votes": 1,
      "replies": [
        {
          "id": 415744,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2018-11-05T15:46:08.313000",
          "content": "<p>noise in data. some people draw totally wrong object</p>",
          "votes": 14,
          "replies": []
        },
        {
          "id": 415756,
          "author_name": "Prajjwal",
          "author_url": "",
          "post_date": "2018-11-05T16:19:46.090000",
          "content": "<p>That's the core reason why even after so much data, models are having difficulties in generalization</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 415762,
          "author_name": "Yibing Wu",
          "author_url": "",
          "post_date": "2018-11-05T16:45:11.623000",
          "content": "<p>It is true the smaller batch size is better, but here we have a lot of data so we can use more data with the larger batch size. I think the more used data helps not the bigger batch size itself.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 427054,
          "author_name": "Diego Perez",
          "author_url": "",
          "post_date": "2018-11-24T12:49:39.647000",
          "content": "<p>Refers to this code:\n<a href=\"https://discuss.pytorch.org/t/simulation-of-large-batch-size/3842/4\">https://discuss.pytorch.org/t/simulation-of-large-batch-size/3842/4</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 421529,
      "author_name": "我滴个神啊",
      "author_url": "",
      "post_date": "2018-11-15T05:21:04.840000",
      "content": "<p>Hi~ Heng\nWhat Data Augmentation do you use?</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 419505,
      "author_name": "William Horton",
      "author_url": "",
      "post_date": "2018-11-12T04:39:42.143000",
      "content": "<p>What size image is large enough? 256x256? Is there any gain trying larger, like 512?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 419818,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2018-11-12T15:42:31.923000",
          "content": "<p>you need large batch size to combat noise, so 512x512 is difficult to train.</p>\n\n<p>There will be no gain or even worse results is the batch size is too small.</p>\n\n<p>imagenet uses batch size of 256.</p>\n\n<p>note that NASNET uses size up to about 331x331? some of the inception uses 229x229</p>\n\n<hr>\n\n<p>to clean the non-recognised drawings, you can train a smaller classifier first (e.g. 128x128) to label them.</p>\n\n<p>manual annotation is also a possible solution  but it would take some time.</p>\n\n<p>the best way is to use a classifier to rank the non-recognised drawings, and them use visual inspection to clean up the data. (human in the loop active learning)</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 419844,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-11-12T16:17:08.657000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 422921,
          "author_name": "Kerem Turgutlu",
          "author_url": "",
          "post_date": "2018-11-17T03:44:31.033000",
          "content": "<p>Hi Heng is this cleaning meant for raw data ? Thanks</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 417261,
      "author_name": "RDizzl3",
      "author_url": "",
      "post_date": "2018-11-08T02:15:53.040000",
      "content": "<p>@CherKeng - just wondering you mentioned that 50,000k per class for 1 epoch could possibly be better than 50k per class for 100 epochs. In practice, if someone were to train only a single epoch is it better to start with a much small learning rate than if you were to train 100 epochs?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 417412,
          "author_name": "HuyenNguyen",
          "author_url": "",
          "post_date": "2018-11-08T08:11:19.463000",
          "content": "<p>Keng, I don't get the logic of why few long epochs are better than many short ones? If the total number of steps is still the same. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 418529,
          "author_name": "Prajjwal",
          "author_url": "",
          "post_date": "2018-11-10T03:35:03.967000",
          "content": "<p>In this case, as Heng told, there's much noise in the data (a significant portion images are not made correctly), running over small epochs may hinder model's progress because that sample of data (from minibatch) might introduce more gradient flow than others. Including more samples/cycle helps distribute it to some extent (Not all samples are created equal).</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 418793,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2018-11-10T16:12:00.603000",
          "content": "<p>first, seeing more train samples is better than seeing less samples + their augmentation.</p>\n\n<p>for learning rate and batch size relationship and experimental results, refer to:</p>\n\n<p>\"Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour\"\n<a href=\"https://arxiv.org/pdf/1706.02677.pdf\">https://arxiv.org/pdf/1706.02677.pdf</a></p>\n\n<p>\"Linear Scaling Rule: When the minibatch size is\nmultiplied by k, multiply the learning rate by k.\"</p>\n\n<p>\"DON’T DECAY THE LEARNING RATE,   INCREASE THE BATCH SIZE\"</p>\n\n<p><a href=\"https://openreview.net/pdf?id=B1Yy1BxCZ\">https://openreview.net/pdf?id=B1Yy1BxCZ</a></p>\n\n<p>for batch size and noise relationship, there are some papers that on that (but i currently forget the title. you may google for it)</p>\n\n<p><a href=\"https://arxiv.org/pdf/1705.10694.pdf\">https://arxiv.org/pdf/1705.10694.pdf</a>\n\"Deep Learning is Robust to Massive Label Noise\"</p>\n\n<p>\"Finally, we have observed that noisy labels reduce\nthe effective batch size, an effect that can be mitigated by\nlarger batch sizes and downscaling the learning rate.\"</p>",
          "votes": 13,
          "replies": []
        },
        {
          "id": 419388,
          "author_name": "Mario Parreño Lara",
          "author_url": "",
          "post_date": "2018-11-11T21:25:03.633000",
          "content": "<p>For example, if we cannot process 512 images as batch (out of memory), in Pytorch we could accumulate the cost in a variable and then when we have processed 512 do the \"cost.backward()\"??</p>\n\n<p>You suggest then use for example... 10 epochs with 256 batch size -&gt; 10 epochs with 512 batch size... and not 20 epochs with 256 batch size and modify every each 10 epochs the learning rate (annealing)... True?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 416785,
      "author_name": "Evgeny Kovalev",
      "author_url": "",
      "post_date": "2018-11-07T09:24:57.140000",
      "content": "<p>Thank you very much for your tips! One question. I wanted to try relatively large batch size (&gt;=256) and came across the problem - I get the \"ResourceExhaustedError\" when I run the code at my local machine with GPU. Seems like I run out of memory on GPU. What could you suggest? Any ideas how could I approach this problem? Of course I could run the code on Kaggle kernel, but in my opinion running the code at local machine is more comfortable.</p>\n\n<p>My settings:\nGeForce GTX1050 Ti Gigabyte PCIE 4096Mb, driver version 398.11\nTensorflow: version 1.8.0\nKeras: version 2.1.6</p>",
      "votes": 2,
      "replies": [
        {
          "id": 417308,
          "author_name": "Finlay",
          "author_url": "",
          "post_date": "2018-11-08T04:26:08.197000",
          "content": "<p>I think the limit is your GPU, sir.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 415579,
      "author_name": "Prajjwal",
      "author_url": "",
      "post_date": "2018-11-05T10:54:45.190000",
      "content": "<p>How much time does it take for you to train on entire dataset ?</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 427832,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-26T08:32:03.260000",
      "content": "<p>24g gpu memory is not enough for se resnext50 model</p>",
      "votes": 0,
      "replies": [
        {
          "id": 427997,
          "author_name": "Strideradu",
          "author_url": "",
          "post_date": "2018-11-26T15:26:22.070000",
          "content": "<p>24G is already a very large GPU and from my experience, with img size 128, you should be OK to run se resnext50 model with batch size at least 256</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 429140,
          "author_name": "Hoàng Tùng Lâm",
          "author_url": "",
          "post_date": "2018-11-28T11:57:37.740000",
          "content": "<p>how big ur image and batch size is ?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 429550,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-11-29T02:15:52.510000",
          "content": "<p>im using xception model! batch size=200 image size=128</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 429572,
          "author_name": "Hoàng Tùng Lâm",
          "author_url": "",
          "post_date": "2018-11-29T03:08:28.707000",
          "content": "<p><a href=\"/jonney\">@jonney</a> You can try 64x64, I get up to 0.925 with just 64x64 images</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 427114,
      "author_name": "Matskevich Ivan",
      "author_url": "",
      "post_date": "2018-11-24T16:11:36.230000",
      "content": "<p>hi, i'm new to data science. I do not have a powerful gpu and I teach at google colaboratory, but when I have a picture size of 256 * 256, I can’t install a large batch size, it gives an out of memory error. tell me where I can train large images online with a great batchsize? thanks</p>",
      "votes": 0,
      "replies": [
        {
          "id": 429734,
          "author_name": "gody7334",
          "author_url": "",
          "post_date": "2018-11-29T09:16:26.593000",
          "content": "<p>Hi Matskevich,\nYou can try accumulating gradients,</p>\n\n<p>Some idea <a href=\"https://medium.com/huggingface/training-larger-batches-practical-tips-on-1-gpu-multi-gpu-distributed-setups-ec88c3e51255\">https://medium.com/huggingface/training-larger-batches-practical-tips-on-1-gpu-multi-gpu-distributed-setups-ec88c3e51255</a></p>\n\n<p>pytorch <a href=\"https://discuss.pytorch.org/t/how-to-implement-accumulated-gradient/3822\">https://discuss.pytorch.org/t/how-to-implement-accumulated-gradient/3822</a>\nkeras <a href=\"https://github.com/keras-team/keras/issues/3556\">https://github.com/keras-team/keras/issues/3556</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 427052,
      "author_name": "Diego Perez",
      "author_url": "",
      "post_date": "2018-11-24T12:47:53.997000",
      "content": "<p>Heng CherKeng thank you for sharing </p>\n\n<p>Questions:\n1. It means to use all epoch and test samples?\n2. should we use raw images?\n3. MobileNet is the right network to use? or refers to speed connection?\n4. Which batch size should we use? 4096K?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 422556,
      "author_name": "upup",
      "author_url": "",
      "post_date": "2018-11-16T11:59:08.097000",
      "content": "<p>Hi, heng! Did you use multi GPUs for larger batch size? My batch size of ResNet50 can only reach to 64 (image224x224x3) on an TITANX GPU(12G).</p>",
      "votes": 0,
      "replies": [
        {
          "id": 422914,
          "author_name": "Strideradu",
          "author_url": "",
          "post_date": "2018-11-17T03:20:35.747000",
          "content": "<p>If you are using pytorch, you can accumulate the loss for 2 or 3 batch, then calculate the gradient</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 423050,
          "author_name": "upup",
          "author_url": "",
          "post_date": "2018-11-17T10:56:13.707000",
          "content": "<p>That's a good idea, thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 423272,
          "author_name": "Mario Parreño Lara",
          "author_url": "",
          "post_date": "2018-11-17T20:35:28.767000",
          "content": "<p>How you did it Strideradu? I write my suggestion here <a href=\"https://discuss.pytorch.org/t/simulation-of-large-batch-size/3842/4\">https://discuss.pytorch.org/t/simulation-of-large-batch-size/3842/4</a> but not getting same results 64-&gt;64 batch as 128 batch updates on weights</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 423324,
          "author_name": "Strideradu",
          "author_url": "",
          "post_date": "2018-11-17T23:19:46.110000",
          "content": "<p><a href=\"https://discuss.pytorch.org/t/pytorch-gradients/884/2\">https://discuss.pytorch.org/t/pytorch-gradients/884/2</a> I didn't try this for this competetion since now I have accecss to a large memory GPU. But I remembered many people using this for the Carvana competetion</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 421567,
      "author_name": "TPloveYXT520",
      "author_url": "",
      "post_date": "2018-11-15T06:17:54.780000",
      "content": "<p>Large image means reize (64,64) -&gt;(128,128) or even bigger\nor pad_and_resize?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 415516,
      "author_name": "Prajjwal",
      "author_url": "",
      "post_date": "2018-11-05T08:50:44.980000",
      "content": "<p>Thanks for sharing this , but 2 and 3 seems kinda ambiguous, could be provide more info regarding size (you mean 256)?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 415621,
          "author_name": "Rajesh Shreedhar",
          "author_url": "",
          "post_date": "2018-11-05T11:59:42.933000",
          "content": "<p>@Prajjwal, I was able to get 0.928 on LB with image size of 128. Trying out with 256, 224 now. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 415635,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-11-05T12:29:52.070000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 415965,
          "author_name": "HuyenNguyen",
          "author_url": "",
          "post_date": "2018-11-06T01:28:26.670000",
          "content": "<p>How much data are you using? I used Grayscale Mobilenet on 128*128 images, but my score didn't improved :-/ 30k images/class compared to 64*64</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 430211,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-30T03:09:24.987000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 423805,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-19T05:00:13.213000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 416161,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-06T09:49:22.493000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2451789,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-09-22T19:23:53.913000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "415474": "if you are using CNN:\n\n1. use all train images\n\n2. use large image\n\n3. use correct network\n\n4. use larger batch size\n\n5. train for many iterations\n\nyou you don't see improvement, change the way you render the strokes to images. **(you have 3 channel rgb, make good use of it! try to encode more information)**\n\nyou should be able to get to 0.941 without trick, or ensemble ...",
    "423426": "baseline performance:\n\nLB: 0.944\n\nse-resnext-50 (256x256 input, 200 samples per batch), single model, no TTA , no ensmble\n\nall train samples (simplified only, recognized and non-recognized): random 80 images per class are selected as validation, the rest are training samples.\n\nvalidation loss : ce_loss, top1, top3,  (map@3)\n\n0.607  0.842  0.946  (0.890)*  \n\ntrain loss : ce_loss, top1, top3,  (map@3)\n\n0.590  0.844  0.948  (0.892) \n\n----\nLB 0.928\n\ncustomized LSTM\n(Max seq length = 600, 200 samples per batch), single model, no TTA , no ensmble\n\nsame train samples as above\n\n\nvalidation loss : ce_loss, top1, top3,  (map@3)\n\n0.637  0.834  0.943  (0.885)*  \n\ntrain loss : ce_loss, top1, top3,  (map@3)\n\n 0.595  0.841  0.949  (0.892)\n\n----\nLB 0.945\n\nstacking via a fuse classifier on features from cnn and lstm. , no TTA , no ensmble\n\n\nsame train samples as above. batch size increase to 512\n\n\nvalidation loss : ce_loss, top1, top3,  (map@3)\n\n0.588  0.850  0.947  (0.895)*  \n\ntrain loss : ce_loss, top1, top3,  (map@3)\n\n 0.549  0.850  0.949  (0.896)\n",
    "423502": "Regarding the efficient usage of input channels (3 or even more if you use a custom model), the following paper might give some inspiration: https://arxiv.org/pdf/1501.07873.pdf",
    "429119": "another reference results\n\nLB: did not submit (estimated to be 0.942)\n\nsenet154 (128x128 input, 256 samples per batch), single model, no TTA , no ensmble\n\nall train samples (simplified only, recognized and non-recognized): random 80 images per class are selected as validation, the rest are training samples.\n\nvalidation loss : ce_loss, top1, top3, (map@3)\n\n0.622  0.839  0.944  (0.888)* \n\ntrain loss : ce_loss, top1, top3, (map@3)\n\n0.582  0.848  0.946  (0.893)  ",
    "415742": "Thanks Heng \n\n I still use 64*64 ....may be it's time to use larger image  :)\n\nI am also  working on encoding more information using all the 3 channels. \n\nBTW , do you think there are some augmentations that help ? ",
    "415551": "Now I use \n1. 50k/class\n2. img_size 224*224\n3. res50\n4. 64*8\ncould got 0.936 and 0.938 with ensemble\n\nDue to some bugs, I did not got any new models with new setting for two weeks. Sign~\n\nNow I will try:\n1. all imgs\n2. img_size: 448*448\n3. Senet?\n4. 256*8",
    "415477": "typical training log:\n(1 iter = 1000 iterations, see my pytorch starter kit)\n\n\n  ![enter image description here][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/415477/10604/train_log.png",
    "421429": "Hi Heng, thanks for your great tips! I've trained ResNet101 for 3 epochs with 64x64 scale and 1024 batch size and it scored 0.925. Does the image size really matter?",
    "419753": "Hi Heng,\nI am currently running on gtx 1080 8GB, using resnet 50, with 64x64 image size, 128 batch size, \nI am using 90% data as training dataset. it take 1 day to go through 1 epoch.\nMy questions are:\n1. Does the training time look right? or can you suggest a training time baseline for 1 epoch?\n2. Is it better to train using 10K per class with 10 epoch, rather than 100K per class with 1 epoch? \nBecause my gpu is quite slow.\nThanks",
    "419730": "encode by time(left)\n\nencode by stroke/part (right)\n\n\nyou can tnink of a way to encode \"pause time\"\n\n   ![enter image description here][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/419730/10659/0.73_9000297423888044_baseball_bat_bat_baseball.png",
    "418071": "Hi Heng,\n\nAmazing tips as usual. I am currently running on gtx 1060 6GB, and so I cannot quite increase both image size and batch size as it would get OOM. Which of the two should I priortize? Imagesize or batch?",
    "425650": "Hi Heng,\nMay I ask what optimizer you used?\nDo you think Adam + Cyclic LR is good idea, \nMy Top 3 accu can easy get 0.96 using this combination but with poor public LB score......\n",
    "421562": "Hi~ Heng\nHow to get the 3 channel rgb?\nShould we  draw the every line with different color?",
    "418503": "Hi Heng,\n\nCould you tell more about \"you have 3 channel rgb, make good use of it! try to encode more information\"? you mean use the bigger filters?",
    "416720": "What is the batch size did you use?  I could not get higher than 64 for 224*224 images for 100000/class. ",
    "415835": "Hi all, thanks for the great discussion.\n\nI am curious how many batch size a Titan X can handle, for example, when using 224x224x3 and ResNet50. I guessed around ~16. Does this make sense or it could be more?\n\n@Li bin mentioned using 256*8 of 224x224x3 input (and also mentioned using a distributed framework.) What would be a typical capacity for a Titan X (as an example)?",
    "415681": "Just curious why large batch size helps, some paper mentioned that smaller batch size should generalize better than larger batch size...\n",
    "421529": "Hi~ Heng\nWhat Data Augmentation do you use?",
    "419505": "What size image is large enough? 256x256? Is there any gain trying larger, like 512?",
    "417261": "@CherKeng - just wondering you mentioned that 50,000k per class for 1 epoch could possibly be better than 50k per class for 100 epochs. In practice, if someone were to train only a single epoch is it better to start with a much small learning rate than if you were to train 100 epochs?",
    "416785": "Thank you very much for your tips! One question. I wanted to try relatively large batch size (&gt;=256) and came across the problem - I get the \"ResourceExhaustedError\" when I run the code at my local machine with GPU. Seems like I run out of memory on GPU. What could you suggest? Any ideas how could I approach this problem? Of course I could run the code on Kaggle kernel, but in my opinion running the code at local machine is more comfortable.\n\nMy settings:\nGeForce GTX1050 Ti Gigabyte PCIE 4096Mb, driver version 398.11\nTensorflow: version 1.8.0\nKeras: version 2.1.6",
    "415579": "How much time does it take for you to train on entire dataset ?",
    "427832": "24g gpu memory is not enough for se resnext50 model",
    "427114": "hi, i'm new to data science. I do not have a powerful gpu and I teach at google colaboratory, but when I have a picture size of 256 * 256, I can’t install a large batch size, it gives an out of memory error. tell me where I can train large images online with a great batchsize? thanks",
    "427052": "Heng CherKeng thank you for sharing \n\nQuestions:\n1. It means to use all epoch and test samples?\n2. should we use raw images?\n3. MobileNet is the right network to use? or refers to speed connection?\n4. Which batch size should we use? 4096K?\n\n\n\n",
    "422556": "Hi, heng! Did you use multi GPUs for larger batch size? My batch size of ResNet50 can only reach to 64 (image224x224x3) on an TITANX GPU(12G).",
    "421567": "Large image means reize (64,64) -&gt;(128,128) or even bigger\nor pad_and_resize?",
    "415516": "Thanks for sharing this , but 2 and 3 seems kinda ambiguous, could be provide more info regarding size (you mean 256)?",
    "430211": "",
    "423805": "",
    "416161": "",
    "2451789": "thanks for this code"
  }
}