{
  "id": 71559,
  "title": "Sensible approach",
  "url": "/competitions/quickdraw-doodle-recognition/discussion/71559",
  "author_name": "Brian Lee",
  "post_date": "2018-11-14T19:39:05.996000",
  "votes": 6,
  "comment_count": 29,
  "views": 0,
  "content": "<p>Hello fellow kagglers,</p>\n\n<p>Am I the only one really missing out on how to approach this problem? Ever since I reached 0.906 with Xception and size 128*128, I have been really stuck on how to improve. Increasing size doesn't help me as much since I got a 1060, but then increasing both size and batch by using cumulative gradient descent doesn't seem to help me converge. I tried going back to MobileNet but it seems to not work either. </p>\n\n<p>Is it simply matter of getting better hardware for increased batch size and image size? </p>",
  "messages": [
    {
      "id": 421263,
      "postDate": "2018-11-14T20:14:04.367Z",
      "content": "<p>I am now at 0.931 LB using MobileNet 128x128 images. I have a 1070 and the max batch size I could reach is 256. With Kaggle Kernels I was able to fit a higher batch size but it was a bit more tricky due to training time limitations per session. So for now I don't think you should consider larger images yet.</p>\n\n<p>How many images are you using? I would definitely suggest training on all images, and the best way to do that would be setting up a good generator and load images that way so its really just a matter of how many steps &amp; epochs you train to get through all of the images. Check Beluga's great Kernel for reference on how to do that. Using this method you should be able to reach at least 0.93 on LB even with a lightweight CNN like MobileNet at image size of 128x128. </p>\n\n<p>Bigger image size i.e. 224x224, deeper models and ensembling/averaging the model preds should get you beyond 0.93.</p>",
      "rawMarkdown": "I am now at 0.931 LB using MobileNet 128x128 images. I have a 1070 and the max batch size I could reach is 256. With Kaggle Kernels I was able to fit a higher batch size but it was a bit more tricky due to training time limitations per session. So for now I don't think you should consider larger images yet.\n\nHow many images are you using? I would definitely suggest training on all images, and the best way to do that would be setting up a good generator and load images that way so its really just a matter of how many steps &amp; epochs you train to get through all of the images. Check Beluga's great Kernel for reference on how to do that. Using this method you should be able to reach at least 0.93 on LB even with a lightweight CNN like MobileNet at image size of 128x128. \n\nBigger image size i.e. 224x224, deeper models and ensembling/averaging the model preds should get you beyond 0.93.",
      "votes": 7,
      "replies": [
        {
          "id": 421310,
          "postDate": "2018-11-14T22:16:18.350Z",
          "content": "<p>I am using all images w beluga's shuffle csv script. I am doing grayscale encoding and concatenating it for imagenet pretrained weights keras. </p>\n\n<p>I am concatenating grayscale encodings to make 3 channel input for keras preteained weights. Could that be a problem? Idk how to encode for rgb .</p>",
          "rawMarkdown": "I am using all images w beluga's shuffle csv script. I am doing grayscale encoding and concatenating it for imagenet pretrained weights keras. \n\nI am concatenating grayscale encodings to make 3 channel input for keras preteained weights. Could that be a problem? Idk how to encode for rgb ."
        },
        {
          "id": 421692,
          "postDate": "2018-11-15T09:18:50.690Z",
          "content": "<p>How many steps/epochs you are training your mobilenet?</p>",
          "rawMarkdown": "How many steps/epochs you are training your mobilenet?",
          "votes": 1
        },
        {
          "id": 422182,
          "postDate": "2018-11-15T22:03:02.523Z",
          "content": "<p>Hi James, how many epochs, steps and what batchsize do you use? I've been running MobileNet for days and only got up to 0.925 :-/</p>",
          "rawMarkdown": "Hi James, how many epochs, steps and what batchsize do you use? I've been running MobileNet for days and only got up to 0.925 :-/",
          "votes": 1
        },
        {
          "id": 422188,
          "postDate": "2018-11-15T22:08:02.740Z",
          "content": "<p>Hi James, some info on batchsize and epochs would be very helpful for me too</p>",
          "rawMarkdown": "Hi James, some info on batchsize and epochs would be very helpful for me too"
        },
        {
          "id": 422259,
          "postDate": "2018-11-16T01:35:51.180Z",
          "content": "<p>Around 50 epochs - 8000 steps per epoch w/ batchsize of 256</p>",
          "rawMarkdown": "Around 50 epochs - 8000 steps per epoch w/ batchsize of 256",
          "votes": 1
        },
        {
          "id": 422840,
          "postDate": "2018-11-16T21:27:23.550Z",
          "content": "<p>8000*256/340=6024, so you are roughly training with 6k images per class right?</p>",
          "rawMarkdown": "8000*256/340=6024, so you are roughly training with 6k images per class right?",
          "votes": 1
        },
        {
          "id": 422869,
          "postDate": "2018-11-16T22:46:41.903Z",
          "content": "<p>Per epoch yes, but I am generating images from the full dataset so across multiple epochs the model will end up seeing more than 6k per class.</p>",
          "rawMarkdown": "Per epoch yes, but I am generating images from the full dataset so across multiple epochs the model will end up seeing more than 6k per class.",
          "votes": 2
        },
        {
          "id": 423753,
          "postDate": "2018-11-19T02:01:24.287Z",
          "content": "<p>In the shuffle CSV script, I believe the nrows was set to 30 thousand. To use all images, did you just change that to nrows=None?</p>\n\n<blockquote>\n  <p><strong>JoonHo Lee wrote</strong></p>\n  \n  <blockquote>\n    <p>I am using all images w beluga's shuffle csv script. I am doing grayscale encoding and concatenating it for imagenet pretrained weights keras. </p>\n  </blockquote>\n  \n  <p>I am concatenating grayscale encodings to make 3 channel input for keras preteained weights. Could that be a problem? Idk how to encode for rgb .</p>\n</blockquote>",
          "rawMarkdown": " In the shuffle CSV script, I believe the nrows was set to 30 thousand. To use all images, did you just change that to nrows=None?\n\n&gt; **JoonHo Lee wrote**\n&gt; \n&gt; &gt; I am using all images w beluga's shuffle csv script. I am doing grayscale encoding and concatenating it for imagenet pretrained weights keras. \n&gt; \n&gt; I am concatenating grayscale encodings to make 3 channel input for keras preteained weights. Could that be a problem? Idk how to encode for rgb .",
          "votes": 1
        },
        {
          "id": 424300,
          "postDate": "2018-11-19T21:37:04.100Z",
          "content": "<p>Hi, I believe just not including the \"rows\" argument automatically reads everything.</p>",
          "rawMarkdown": "Hi, I believe just not including the \"rows\" argument automatically reads everything."
        },
        {
          "id": 425919,
          "postDate": "2018-11-22T10:00:49.867Z",
          "content": "<p>correct</p>",
          "rawMarkdown": "correct"
        }
      ]
    },
    {
      "id": 421251,
      "postDate": "2018-11-14T19:39:05.997Z",
      "content": "<p>Hello fellow kagglers,</p>\n\n<p>Am I the only one really missing out on how to approach this problem? Ever since I reached 0.906 with Xception and size 128*128, I have been really stuck on how to improve. Increasing size doesn't help me as much since I got a 1060, but then increasing both size and batch by using cumulative gradient descent doesn't seem to help me converge. I tried going back to MobileNet but it seems to not work either. </p>\n\n<p>Is it simply matter of getting better hardware for increased batch size and image size? </p>",
      "rawMarkdown": "Hello fellow kagglers,\n\nAm I the only one really missing out on how to approach this problem? Ever since I reached 0.906 with Xception and size 128*128, I have been really stuck on how to improve. Increasing size doesn't help me as much since I got a 1060, but then increasing both size and batch by using cumulative gradient descent doesn't seem to help me converge. I tried going back to MobileNet but it seems to not work either. \n\nIs it simply matter of getting better hardware for increased batch size and image size? ",
      "votes": 6
    },
    {
      "id": 421449,
      "postDate": "2018-11-15T02:38:35.807Z",
      "content": "<p>Hello! In my case, I've got about 9.19 using grayscale and the score increased to 9.25 when I added time information to images. I think this should be helpful: <a href=\"https://www.kaggle.com/gaborfodor/greyscale-mobilenet-lb-0-892\">https://www.kaggle.com/gaborfodor/greyscale-mobilenet-lb-0-892</a></p>",
      "rawMarkdown": "Hello! In my case, I've got about 9.19 using grayscale and the score increased to 9.25 when I added time information to images. I think this should be helpful: https://www.kaggle.com/gaborfodor/greyscale-mobilenet-lb-0-892",
      "votes": 2,
      "replies": [
        {
          "id": 421453,
          "postDate": "2018-11-15T02:58:16.207Z",
          "content": "<p>Interesting! I am following that kernel, guess I'll go back to single channel input. Is there any resources for how to add time information?</p>",
          "rawMarkdown": "Interesting! I am following that kernel, guess I'll go back to single channel input. Is there any resources for how to add time information?",
          "votes": 1
        },
        {
          "id": 421464,
          "postDate": "2018-11-15T03:33:10.613Z",
          "content": "<p>Oh, that's actually \"stroke\" information, not \"time\". Add stroke group information. That's enough.</p>",
          "rawMarkdown": "Oh, that's actually \"stroke\" information, not \"time\". Add stroke group information. That's enough.",
          "votes": 1
        },
        {
          "id": 422742,
          "postDate": "2018-11-16T17:51:30.207Z",
          "content": "<p>So does that mean you have a 2 channel input (grayscale image + stroke information)? I'm not sure how to add stroke information. (Still a beginner so I'm struggling with encoding information a lot)</p>",
          "rawMarkdown": "So does that mean you have a 2 channel input (grayscale image + stroke information)? I'm not sure how to add stroke information. (Still a beginner so I'm struggling with encoding information a lot)",
          "votes": 1
        },
        {
          "id": 424866,
          "postDate": "2018-11-20T19:43:36.543Z",
          "content": "<p>I think it's just one channel, but instead of being binary 0 is not draw, 1 is draw for each pixel.</p>\n\n<p>you use a gray scale (so values between 0 an 1.0) to indicate the stroke group...</p>\n\n<p>for example, all points of the first stroke are values 1.0, second strokes are 0.9 third stroke are 0.8, and eveything that is beyond the 10nth stroke is just 0.1</p>\n\n<p>This way the tensor has more information about how the traces tight together.</p>",
          "rawMarkdown": "I think it's just one channel, but instead of being binary 0 is not draw, 1 is draw for each pixel.\n\nyou use a gray scale (so values between 0 an 1.0) to indicate the stroke group...\n\n\nfor example, all points of the first stroke are values 1.0, second strokes are 0.9 third stroke are 0.8, and eveything that is beyond the 10nth stroke is just 0.1\n\nThis way the tensor has more information about how the traces tight together.",
          "votes": 1
        }
      ]
    },
    {
      "id": 424468,
      "postDate": "2018-11-20T06:55:59.353Z",
      "content": "<p>Update: Xceptionnet with all images and accumulative gradient got to 0.919. I may have stopped it a bit early but my losses were going back up. </p>",
      "rawMarkdown": "Update: Xceptionnet with all images and accumulative gradient got to 0.919. I may have stopped it a bit early but my losses were going back up. ",
      "replies": [
        {
          "id": 425400,
          "postDate": "2018-11-21T14:54:51.143Z",
          "content": "<p>@ JoonHo Lee </p>\n\n<p>I was able to get 0.919 with mobilenet, <code>input_shape=(128, 128, 1), batchsize=300,  all_images</code></p>\n\n<p><strong>Hardware:</strong> <code>GTX 1080Ti</code> </p>\n\n<p><strong>Time Taken:</strong> <code>24hr +</code></p>",
          "rawMarkdown": "@ JoonHo Lee \n\nI was able to get 0.919 with mobilenet, `input_shape=(128, 128, 1), batchsize=300,  all_images`\n\n**Hardware:** `GTX 1080Ti` \n\n**Time Taken:** `24hr + `",
          "votes": 1
        },
        {
          "id": 425571,
          "postDate": "2018-11-21T19:42:57.533Z",
          "content": "<p>May I know what your step size was? Were you going through all images per epoch?</p>",
          "rawMarkdown": "May I know what your step size was? Were you going through all images per epoch?"
        },
        {
          "id": 425812,
          "postDate": "2018-11-22T06:43:03.307Z",
          "content": "<p>@ JoonHo Lee I have separated the steps and total number of images so my epochs are calculated as below:\n<code>EPOCHS = math.ceil(total_samples / (batchsize * STEPS) * num_cycles)</code></p>\n\n<p>so my <strong>STEPS</strong> are variable but for above I have used  <code>steps = 1000</code></p>",
          "rawMarkdown": "@ JoonHo Lee I have separated the steps and total number of images so my epochs are calculated as below:\n`EPOCHS = math.ceil(total_samples / (batchsize * STEPS) * num_cycles)`\n\nso my **STEPS** are variable but for above I have used  `steps = 1000`\n"
        }
      ]
    },
    {
      "id": 421762,
      "postDate": "2018-11-15T11:42:32.700Z",
      "content": "<p>Hi JoonHo,\nMay I ask how long it take to train 3 epoch on 1060?\nThanks,</p>",
      "rawMarkdown": "Hi JoonHo,\nMay I ask how long it take to train 3 epoch on 1060?\nThanks,",
      "replies": [
        {
          "id": 422171,
          "postDate": "2018-11-15T21:44:10.040Z",
          "content": "<p>Hi,\n2500 steps per epoch w all images (image size 96 and batchsize 64 * 8) tkaes about 1000s per epoch. I probably need a larger step size to go through all dataet though.</p>",
          "rawMarkdown": "Hi,\n2500 steps per epoch w all images (image size 96 and batchsize 64 * 8) tkaes about 1000s per epoch. I probably need a larger step size to go through all dataet though."
        },
        {
          "id": 422303,
          "postDate": "2018-11-16T03:28:59.343Z",
          "content": "<p>1000 s per epoch?</p>",
          "rawMarkdown": "1000 s per epoch?"
        },
        {
          "id": 422501,
          "postDate": "2018-11-16T10:15:16.347Z",
          "content": "<p>Hi JoonHo, thanks for sharing\nI will reasonably guess its 1000s per step?\nSimply calculate training time, it take (2500 * 1000s) = 2500000s = 416000mins = 690hours = 28days\nIs that correct?</p>",
          "rawMarkdown": "Hi JoonHo, thanks for sharing\nI will reasonably guess its 1000s per step?\nSimply calculate training time, it take (2500 * 1000s) = 2500000s = 416000mins = 690hours = 28days\nIs that correct?",
          "votes": 1
        },
        {
          "id": 422730,
          "postDate": "2018-11-16T17:33:16.630Z",
          "content": "<p>Hi gody,</p>\n\n<p>I meant it's 1000s for the entire epoch(So about 1s per step). I'm not running full amount of epoches e.g. stop at 80/100 so it takes me 80*1000/3600 = 22 hours.</p>",
          "rawMarkdown": "Hi gody,\n\nI meant it's 1000s for the entire epoch(So about 1s per step). I'm not running full amount of epoches e.g. stop at 80/100 so it takes me 80*1000/3600 = 22 hours.",
          "votes": 1
        }
      ]
    },
    {
      "id": 421473,
      "postDate": "2018-11-15T03:44:07.257Z",
      "content": "<p>Wondering when you tring to use cumulative gradient descent . what is your equivalent batch_size</p>",
      "rawMarkdown": "Wondering when you tring to use cumulative gradient descent . what is your equivalent batch_size",
      "replies": [
        {
          "id": 421511,
          "postDate": "2018-11-15T04:52:47.620Z",
          "content": "<p>I've been using batchsize 64 with 4 iterations, so I guess 256? Using code from <a href=\"https://github.com/keras-team/keras/issues/3556\">https://github.com/keras-team/keras/issues/3556</a></p>",
          "rawMarkdown": "I've been using batchsize 64 with 4 iterations, so I guess 256? Using code from https://github.com/keras-team/keras/issues/3556",
          "votes": 1
        },
        {
          "id": 421975,
          "postDate": "2018-11-15T16:14:40.997Z",
          "content": "<p>Ok, I see. I am not familiar with how to do it in Keras</p>",
          "rawMarkdown": "Ok, I see. I am not familiar with how to do it in Keras"
        },
        {
          "id": 422743,
          "postDate": "2018-11-16T17:52:02.427Z",
          "content": "<p>I see. What are you using for this competition, pytorch?</p>",
          "rawMarkdown": "I see. What are you using for this competition, pytorch?"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 421263,
      "author_name": "James Requa",
      "author_url": "",
      "post_date": "2018-11-14T20:14:04.367000",
      "content": "<p>I am now at 0.931 LB using MobileNet 128x128 images. I have a 1070 and the max batch size I could reach is 256. With Kaggle Kernels I was able to fit a higher batch size but it was a bit more tricky due to training time limitations per session. So for now I don't think you should consider larger images yet.</p>\n\n<p>How many images are you using? I would definitely suggest training on all images, and the best way to do that would be setting up a good generator and load images that way so its really just a matter of how many steps &amp; epochs you train to get through all of the images. Check Beluga's great Kernel for reference on how to do that. Using this method you should be able to reach at least 0.93 on LB even with a lightweight CNN like MobileNet at image size of 128x128. </p>\n\n<p>Bigger image size i.e. 224x224, deeper models and ensembling/averaging the model preds should get you beyond 0.93.</p>",
      "votes": 7,
      "replies": [
        {
          "id": 421310,
          "author_name": "Brian Lee",
          "author_url": "",
          "post_date": "2018-11-14T22:16:18.350000",
          "content": "<p>I am using all images w beluga's shuffle csv script. I am doing grayscale encoding and concatenating it for imagenet pretrained weights keras. </p>\n\n<p>I am concatenating grayscale encodings to make 3 channel input for keras preteained weights. Could that be a problem? Idk how to encode for rgb .</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 421692,
          "author_name": "Ankit Sati",
          "author_url": "",
          "post_date": "2018-11-15T09:18:50.690000",
          "content": "<p>How many steps/epochs you are training your mobilenet?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 422182,
          "author_name": "HuyenNguyen",
          "author_url": "",
          "post_date": "2018-11-15T22:03:02.523000",
          "content": "<p>Hi James, how many epochs, steps and what batchsize do you use? I've been running MobileNet for days and only got up to 0.925 :-/</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 422188,
          "author_name": "Brian Lee",
          "author_url": "",
          "post_date": "2018-11-15T22:08:02.740000",
          "content": "<p>Hi James, some info on batchsize and epochs would be very helpful for me too</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 422259,
          "author_name": "James Requa",
          "author_url": "",
          "post_date": "2018-11-16T01:35:51.180000",
          "content": "<p>Around 50 epochs - 8000 steps per epoch w/ batchsize of 256</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 422840,
          "author_name": "Yuming Huang",
          "author_url": "",
          "post_date": "2018-11-16T21:27:23.550000",
          "content": "<p>8000*256/340=6024, so you are roughly training with 6k images per class right?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 422869,
          "author_name": "James Requa",
          "author_url": "",
          "post_date": "2018-11-16T22:46:41.903000",
          "content": "<p>Per epoch yes, but I am generating images from the full dataset so across multiple epochs the model will end up seeing more than 6k per class.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 423753,
          "author_name": "kdaqlkgjnmaklvmnkankl",
          "author_url": "",
          "post_date": "2018-11-19T02:01:24.287000",
          "content": "<p>In the shuffle CSV script, I believe the nrows was set to 30 thousand. To use all images, did you just change that to nrows=None?</p>\n\n<blockquote>\n  <p><strong>JoonHo Lee wrote</strong></p>\n  \n  <blockquote>\n    <p>I am using all images w beluga's shuffle csv script. I am doing grayscale encoding and concatenating it for imagenet pretrained weights keras. </p>\n  </blockquote>\n  \n  <p>I am concatenating grayscale encodings to make 3 channel input for keras preteained weights. Could that be a problem? Idk how to encode for rgb .</p>\n</blockquote>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 424300,
          "author_name": "Brian Lee",
          "author_url": "",
          "post_date": "2018-11-19T21:37:04.100000",
          "content": "<p>Hi, I believe just not including the \"rows\" argument automatically reads everything.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 425919,
          "author_name": "[he.ai]soulmachine",
          "author_url": "",
          "post_date": "2018-11-22T10:00:49.867000",
          "content": "<p>correct</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 421449,
      "author_name": "Jungwoo Park",
      "author_url": "",
      "post_date": "2018-11-15T02:38:35.807000",
      "content": "<p>Hello! In my case, I've got about 9.19 using grayscale and the score increased to 9.25 when I added time information to images. I think this should be helpful: <a href=\"https://www.kaggle.com/gaborfodor/greyscale-mobilenet-lb-0-892\">https://www.kaggle.com/gaborfodor/greyscale-mobilenet-lb-0-892</a></p>",
      "votes": 2,
      "replies": [
        {
          "id": 421453,
          "author_name": "Brian Lee",
          "author_url": "",
          "post_date": "2018-11-15T02:58:16.207000",
          "content": "<p>Interesting! I am following that kernel, guess I'll go back to single channel input. Is there any resources for how to add time information?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 421464,
          "author_name": "Jungwoo Park",
          "author_url": "",
          "post_date": "2018-11-15T03:33:10.613000",
          "content": "<p>Oh, that's actually \"stroke\" information, not \"time\". Add stroke group information. That's enough.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 422742,
          "author_name": "Brian Lee",
          "author_url": "",
          "post_date": "2018-11-16T17:51:30.207000",
          "content": "<p>So does that mean you have a 2 channel input (grayscale image + stroke information)? I'm not sure how to add stroke information. (Still a beginner so I'm struggling with encoding information a lot)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 424866,
          "author_name": "JordiCoscolla",
          "author_url": "",
          "post_date": "2018-11-20T19:43:36.543000",
          "content": "<p>I think it's just one channel, but instead of being binary 0 is not draw, 1 is draw for each pixel.</p>\n\n<p>you use a gray scale (so values between 0 an 1.0) to indicate the stroke group...</p>\n\n<p>for example, all points of the first stroke are values 1.0, second strokes are 0.9 third stroke are 0.8, and eveything that is beyond the 10nth stroke is just 0.1</p>\n\n<p>This way the tensor has more information about how the traces tight together.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 424468,
      "author_name": "Brian Lee",
      "author_url": "",
      "post_date": "2018-11-20T06:55:59.353000",
      "content": "<p>Update: Xceptionnet with all images and accumulative gradient got to 0.919. I may have stopped it a bit early but my losses were going back up. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 425400,
          "author_name": "remidi",
          "author_url": "",
          "post_date": "2018-11-21T14:54:51.143000",
          "content": "<p>@ JoonHo Lee </p>\n\n<p>I was able to get 0.919 with mobilenet, <code>input_shape=(128, 128, 1), batchsize=300,  all_images</code></p>\n\n<p><strong>Hardware:</strong> <code>GTX 1080Ti</code> </p>\n\n<p><strong>Time Taken:</strong> <code>24hr +</code></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 425571,
          "author_name": "Brian Lee",
          "author_url": "",
          "post_date": "2018-11-21T19:42:57.533000",
          "content": "<p>May I know what your step size was? Were you going through all images per epoch?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 425812,
          "author_name": "remidi",
          "author_url": "",
          "post_date": "2018-11-22T06:43:03.307000",
          "content": "<p>@ JoonHo Lee I have separated the steps and total number of images so my epochs are calculated as below:\n<code>EPOCHS = math.ceil(total_samples / (batchsize * STEPS) * num_cycles)</code></p>\n\n<p>so my <strong>STEPS</strong> are variable but for above I have used  <code>steps = 1000</code></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 421762,
      "author_name": "gody7334",
      "author_url": "",
      "post_date": "2018-11-15T11:42:32.700000",
      "content": "<p>Hi JoonHo,\nMay I ask how long it take to train 3 epoch on 1060?\nThanks,</p>",
      "votes": 0,
      "replies": [
        {
          "id": 422171,
          "author_name": "Brian Lee",
          "author_url": "",
          "post_date": "2018-11-15T21:44:10.040000",
          "content": "<p>Hi,\n2500 steps per epoch w all images (image size 96 and batchsize 64 * 8) tkaes about 1000s per epoch. I probably need a larger step size to go through all dataet though.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 422303,
          "author_name": "Strideradu",
          "author_url": "",
          "post_date": "2018-11-16T03:28:59.343000",
          "content": "<p>1000 s per epoch?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 422501,
          "author_name": "gody7334",
          "author_url": "",
          "post_date": "2018-11-16T10:15:16.347000",
          "content": "<p>Hi JoonHo, thanks for sharing\nI will reasonably guess its 1000s per step?\nSimply calculate training time, it take (2500 * 1000s) = 2500000s = 416000mins = 690hours = 28days\nIs that correct?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 422730,
          "author_name": "Brian Lee",
          "author_url": "",
          "post_date": "2018-11-16T17:33:16.630000",
          "content": "<p>Hi gody,</p>\n\n<p>I meant it's 1000s for the entire epoch(So about 1s per step). I'm not running full amount of epoches e.g. stop at 80/100 so it takes me 80*1000/3600 = 22 hours.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 421473,
      "author_name": "Strideradu",
      "author_url": "",
      "post_date": "2018-11-15T03:44:07.257000",
      "content": "<p>Wondering when you tring to use cumulative gradient descent . what is your equivalent batch_size</p>",
      "votes": 0,
      "replies": [
        {
          "id": 421511,
          "author_name": "Brian Lee",
          "author_url": "",
          "post_date": "2018-11-15T04:52:47.620000",
          "content": "<p>I've been using batchsize 64 with 4 iterations, so I guess 256? Using code from <a href=\"https://github.com/keras-team/keras/issues/3556\">https://github.com/keras-team/keras/issues/3556</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 421975,
          "author_name": "Strideradu",
          "author_url": "",
          "post_date": "2018-11-15T16:14:40.997000",
          "content": "<p>Ok, I see. I am not familiar with how to do it in Keras</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 422743,
          "author_name": "Brian Lee",
          "author_url": "",
          "post_date": "2018-11-16T17:52:02.427000",
          "content": "<p>I see. What are you using for this competition, pytorch?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "421263": "I am now at 0.931 LB using MobileNet 128x128 images. I have a 1070 and the max batch size I could reach is 256. With Kaggle Kernels I was able to fit a higher batch size but it was a bit more tricky due to training time limitations per session. So for now I don't think you should consider larger images yet.\n\nHow many images are you using? I would definitely suggest training on all images, and the best way to do that would be setting up a good generator and load images that way so its really just a matter of how many steps &amp; epochs you train to get through all of the images. Check Beluga's great Kernel for reference on how to do that. Using this method you should be able to reach at least 0.93 on LB even with a lightweight CNN like MobileNet at image size of 128x128. \n\nBigger image size i.e. 224x224, deeper models and ensembling/averaging the model preds should get you beyond 0.93.",
    "421251": "Hello fellow kagglers,\n\nAm I the only one really missing out on how to approach this problem? Ever since I reached 0.906 with Xception and size 128*128, I have been really stuck on how to improve. Increasing size doesn't help me as much since I got a 1060, but then increasing both size and batch by using cumulative gradient descent doesn't seem to help me converge. I tried going back to MobileNet but it seems to not work either. \n\nIs it simply matter of getting better hardware for increased batch size and image size? ",
    "421449": "Hello! In my case, I've got about 9.19 using grayscale and the score increased to 9.25 when I added time information to images. I think this should be helpful: https://www.kaggle.com/gaborfodor/greyscale-mobilenet-lb-0-892",
    "424468": "Update: Xceptionnet with all images and accumulative gradient got to 0.919. I may have stopped it a bit early but my losses were going back up. ",
    "421762": "Hi JoonHo,\nMay I ask how long it take to train 3 epoch on 1060?\nThanks,",
    "421473": "Wondering when you tring to use cumulative gradient descent . what is your equivalent batch_size"
  }
}