{
  "id": 72685,
  "title": "Steps Per Epoch with Keras generator",
  "url": "/competitions/quickdraw-doodle-recognition/discussion/72685",
  "author_name": "Ankit Sati",
  "post_date": "2018-11-26T08:32:02.413000",
  "votes": 2,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Hi Guys,<br>\nI am wondering how would you set steps per epoch when you are using keras data generator to generate batches. I am using shuffle csv's and can't figure out a way to set the right amount of steps per epoch. Can anyone help with this?<br>\nThanks,</p>",
  "messages": [
    {
      "id": 427841,
      "postDate": "2018-11-26T08:44:13.890Z",
      "content": "<p>Steps in Keras is meant for epoch performance check. I am using <code>STEPS=1000</code> so for every 1000 steps of <code>batchsize=96(example)</code> I would check validation score/error/performance.\nI am just ensuring that all the samples are completed in <strong>epochs * steps * batchsize</strong></p>\n\n<pre><code> STEPS = 1000 \n size = 128\n batchsize = 96\n total_samples = 49673579\n EPOCHS = math.ceil(total_samples / (batchsize * STEPS) * 1)\n valid_steps= math.ceil(34000/batchsize)\n</code></pre>",
      "rawMarkdown": "Steps in Keras is meant for epoch performance check. I am using `STEPS=1000` so for every 1000 steps of `batchsize=96(example)` I would check validation score/error/performance.\nI am just ensuring that all the samples are completed in **epochs * steps * batchsize**\n\n \n\n     STEPS = 1000 \n     size = 128\n     batchsize = 96\n     total_samples = 49673579\n     EPOCHS = math.ceil(total_samples / (batchsize * STEPS) * 1)\n     valid_steps= math.ceil(34000/batchsize)",
      "votes": 2,
      "replies": [
        {
          "id": 427842,
          "postDate": "2018-11-26T08:50:30.767Z",
          "content": "<p>Do your total samples cover all of the images in train_simplified?</p>",
          "rawMarkdown": "Do your total samples cover all of the images in train_simplified?"
        },
        {
          "id": 427862,
          "postDate": "2018-11-26T09:56:55.673Z",
          "content": "<p>yes; Please note that I use <code>100 samples per class for validation</code></p>\n\n<p><code>total_samples = 49673579</code> is actually for training <em>(ncsvs)</em> and <code>validation_samples = 34000</code>;  a separate file in <em>valid.csv</em></p>\n\n<p>Please refer: <a href=\"https://www.kaggle.com/remidi/robust-train-and-valid-split\">robust-train-and-valid-split</a></p>",
          "rawMarkdown": "yes; Please note that I use `100 samples per class for validation`\n\n`total_samples = 49673579` is actually for training *(ncsvs)* and `validation_samples = 34000`;  a separate file in *valid.csv*\n\nPlease refer: [robust-train-and-valid-split](https://www.kaggle.com/remidi/robust-train-and-valid-split)",
          "votes": 1
        },
        {
          "id": 427871,
          "postDate": "2018-11-26T10:11:17.573Z",
          "content": "<p>Thanks <a href=\"/remidi\">@remidi</a> for the help.</p>",
          "rawMarkdown": "Thanks @remidi for the help."
        },
        {
          "id": 428349,
          "postDate": "2018-11-27T06:00:15.903Z",
          "content": "<p><a href=\"/remidi\">@remidi</a>, so your model is trained on a total data of (49673579 - 34000)? Where 34000 is the validation data?</p>\n\n<p>How long did your model take to train for image size of 128?</p>",
          "rawMarkdown": "@remidi, so your model is trained on a total data of (49673579 - 34000)? Where 34000 is the validation data?\n\nHow long did your model take to train for image size of 128?"
        },
        {
          "id": 428387,
          "postDate": "2018-11-27T07:28:39.523Z",
          "content": "<p>@YaGana Sheriff-Hussaini </p>\n\n<p>It takes more than ~ 48 hrs for se-resnext50 but I am doing something fundamentally wrong so my score is not improving beyond <code>0.922 public LB</code>. According to other competitors it should reach ~ 0.94 maybe poor <strong>lr schedule</strong>. And also I run for only 1 epoch maybe I need to further <strong>fine-tune</strong>.</p>",
          "rawMarkdown": "@YaGana Sheriff-Hussaini \n\nIt takes more than ~ 48 hrs for se-resnext50 but I am doing something fundamentally wrong so my score is not improving beyond `0.922 public LB`. According to other competitors it should reach ~ 0.94 maybe poor **lr schedule**. And also I run for only 1 epoch maybe I need to further **fine-tune**.",
          "votes": 1
        },
        {
          "id": 428401,
          "postDate": "2018-11-27T08:02:55.353Z",
          "content": "<p>Thanks <a href=\"/remidi\">@remidi</a> for answering my question. I got a better result of 0.913 with image size of 80. After increasing it to 128x128 and burning over 50hrs of training in the cloud, it only got to 0.905 on the LB. I only did that because many people mentioned getting 0.94+ with image size 128. It depends on the type of model I guess.</p>",
          "rawMarkdown": "Thanks @remidi for answering my question. I got a better result of 0.913 with image size of 80. After increasing it to 128x128 and burning over 50hrs of training in the cloud, it only got to 0.905 on the LB. I only did that because many people mentioned getting 0.94+ with image size 128. It depends on the type of model I guess."
        },
        {
          "id": 428413,
          "postDate": "2018-11-27T08:32:42.553Z",
          "content": "<p>@YaGana Sheriff-Hussaini</p>\n\n<p>True, I guess I have to experiment more with models, Xception is something I have not tried still and there are some papers suggested by @Heng which claim some promising results. I am trying to implement them too.</p>\n\n<p>reference: </p>\n\n<ul>\n<li><a href=\"https://arxiv.org/abs/1501.07873\">https://arxiv.org/abs/1501.07873</a></li>\n<li><a href=\"https://arxiv.org/abs/1811.08170\">https://arxiv.org/abs/1811.08170</a></li>\n</ul>",
          "rawMarkdown": "@YaGana Sheriff-Hussaini\n\nTrue, I guess I have to experiment more with models, Xception is something I have not tried still and there are some papers suggested by @Heng which claim some promising results. I am trying to implement them too.\n\nreference: \n\n+ https://arxiv.org/abs/1501.07873\n+ https://arxiv.org/abs/1811.08170"
        },
        {
          "id": 428465,
          "postDate": "2018-11-27T10:27:35.857Z",
          "content": "<p>Thanks <a href=\"/remidi\">@remidi</a>, I have not seen the 2nd paper before. It looks promising.</p>",
          "rawMarkdown": "Thanks @remidi, I have not seen the 2nd paper before. It looks promising."
        }
      ]
    },
    {
      "id": 427831,
      "postDate": "2018-11-26T08:32:02.413Z",
      "content": "<p>Hi Guys,<br>\nI am wondering how would you set steps per epoch when you are using keras data generator to generate batches. I am using shuffle csv's and can't figure out a way to set the right amount of steps per epoch. Can anyone help with this?<br>\nThanks,</p>",
      "rawMarkdown": "Hi Guys,<br>\nI am wondering how would you set steps per epoch when you are using keras data generator to generate batches. I am using shuffle csv's and can't figure out a way to set the right amount of steps per epoch. Can anyone help with this?<br>\nThanks,",
      "votes": 2
    },
    {
      "id": 427854,
      "postDate": "2018-11-26T09:24:57.590Z",
      "content": "<p>Sorry I also have a rookie question. \nIf a train set have 10000 samples.</p>\n\n<p>Is there any difference on:\nVersion 1: Step=100, batch_size=100, epoch=1\nand\nVersion 2: Step=10, batch_size=100, epoch=10 ?</p>\n\n<p>Is version 2 just repeatedly training on the SAME 1000 data for 10 times? \nthanks</p>",
      "rawMarkdown": "Sorry I also have a rookie question. \nIf a train set have 10000 samples.\n\nIs there any difference on:\nVersion 1: Step=100, batch_size=100, epoch=1\nand\nVersion 2: Step=10, batch_size=100, epoch=10 ?\n\nIs version 2 just repeatedly training on the SAME 1000 data for 10 times? \nthanks\n\n",
      "replies": [
        {
          "id": 427865,
          "postDate": "2018-11-26T10:05:08.150Z",
          "content": "<p>model trains on the <strong>same number of samples</strong> in both the cases but in case of <em>version: 2</em> you are monitoring more number of times and calculating valid loss and accuaracy for 10 times rather than 1 time in <em>version: 1</em></p>",
          "rawMarkdown": "model trains on the **same number of samples** in both the cases but in case of *version: 2* you are monitoring more number of times and calculating valid loss and accuaracy for 10 times rather than 1 time in *version: 1*",
          "votes": 1
        },
        {
          "id": 427873,
          "postDate": "2018-11-26T10:17:07.163Z",
          "content": "<p>thanks so much remidi</p>",
          "rawMarkdown": "thanks so much remidi"
        },
        {
          "id": 429610,
          "postDate": "2018-11-29T04:10:21.303Z",
          "content": "<p>remidi - is the net result of more monitoring/calculating any improvement in the results - or are you just able to see a bad model earlier?</p>",
          "rawMarkdown": "remidi - is the net result of more monitoring/calculating any improvement in the results - or are you just able to see a bad model earlier?"
        },
        {
          "id": 429942,
          "postDate": "2018-11-29T15:14:53.407Z",
          "content": "<p>@PCJimmmy \nI see no improvement in the model by doing more validation; just a better monitoring technique but keep in mind that more epochs will take more time because of validation for each epoch.\n<code>extra_time = validation_time * epochs</code></p>",
          "rawMarkdown": "@PCJimmmy \nI see no improvement in the model by doing more validation; just a better monitoring technique but keep in mind that more epochs will take more time because of validation for each epoch.\n` extra_time = validation_time * epochs`"
        },
        {
          "id": 450664,
          "postDate": "2019-01-05T13:52:48.257Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 427841,
      "author_name": "remidi",
      "author_url": "",
      "post_date": "2018-11-26T08:44:13.890000",
      "content": "<p>Steps in Keras is meant for epoch performance check. I am using <code>STEPS=1000</code> so for every 1000 steps of <code>batchsize=96(example)</code> I would check validation score/error/performance.\nI am just ensuring that all the samples are completed in <strong>epochs * steps * batchsize</strong></p>\n\n<pre><code> STEPS = 1000 \n size = 128\n batchsize = 96\n total_samples = 49673579\n EPOCHS = math.ceil(total_samples / (batchsize * STEPS) * 1)\n valid_steps= math.ceil(34000/batchsize)\n</code></pre>",
      "votes": 2,
      "replies": [
        {
          "id": 427842,
          "author_name": "Ankit Sati",
          "author_url": "",
          "post_date": "2018-11-26T08:50:30.767000",
          "content": "<p>Do your total samples cover all of the images in train_simplified?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 427862,
          "author_name": "remidi",
          "author_url": "",
          "post_date": "2018-11-26T09:56:55.673000",
          "content": "<p>yes; Please note that I use <code>100 samples per class for validation</code></p>\n\n<p><code>total_samples = 49673579</code> is actually for training <em>(ncsvs)</em> and <code>validation_samples = 34000</code>;  a separate file in <em>valid.csv</em></p>\n\n<p>Please refer: <a href=\"https://www.kaggle.com/remidi/robust-train-and-valid-split\">robust-train-and-valid-split</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 427871,
          "author_name": "Ankit Sati",
          "author_url": "",
          "post_date": "2018-11-26T10:11:17.573000",
          "content": "<p>Thanks <a href=\"/remidi\">@remidi</a> for the help.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 428349,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-11-27T06:00:15.903000",
          "content": "<p><a href=\"/remidi\">@remidi</a>, so your model is trained on a total data of (49673579 - 34000)? Where 34000 is the validation data?</p>\n\n<p>How long did your model take to train for image size of 128?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 428387,
          "author_name": "remidi",
          "author_url": "",
          "post_date": "2018-11-27T07:28:39.523000",
          "content": "<p>@YaGana Sheriff-Hussaini </p>\n\n<p>It takes more than ~ 48 hrs for se-resnext50 but I am doing something fundamentally wrong so my score is not improving beyond <code>0.922 public LB</code>. According to other competitors it should reach ~ 0.94 maybe poor <strong>lr schedule</strong>. And also I run for only 1 epoch maybe I need to further <strong>fine-tune</strong>.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 428401,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-11-27T08:02:55.353000",
          "content": "<p>Thanks <a href=\"/remidi\">@remidi</a> for answering my question. I got a better result of 0.913 with image size of 80. After increasing it to 128x128 and burning over 50hrs of training in the cloud, it only got to 0.905 on the LB. I only did that because many people mentioned getting 0.94+ with image size 128. It depends on the type of model I guess.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 428413,
          "author_name": "remidi",
          "author_url": "",
          "post_date": "2018-11-27T08:32:42.553000",
          "content": "<p>@YaGana Sheriff-Hussaini</p>\n\n<p>True, I guess I have to experiment more with models, Xception is something I have not tried still and there are some papers suggested by @Heng which claim some promising results. I am trying to implement them too.</p>\n\n<p>reference: </p>\n\n<ul>\n<li><a href=\"https://arxiv.org/abs/1501.07873\">https://arxiv.org/abs/1501.07873</a></li>\n<li><a href=\"https://arxiv.org/abs/1811.08170\">https://arxiv.org/abs/1811.08170</a></li>\n</ul>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 428465,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-11-27T10:27:35.857000",
          "content": "<p>Thanks <a href=\"/remidi\">@remidi</a>, I have not seen the 2nd paper before. It looks promising.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 427854,
      "author_name": "Salaryman",
      "author_url": "",
      "post_date": "2018-11-26T09:24:57.590000",
      "content": "<p>Sorry I also have a rookie question. \nIf a train set have 10000 samples.</p>\n\n<p>Is there any difference on:\nVersion 1: Step=100, batch_size=100, epoch=1\nand\nVersion 2: Step=10, batch_size=100, epoch=10 ?</p>\n\n<p>Is version 2 just repeatedly training on the SAME 1000 data for 10 times? \nthanks</p>",
      "votes": 0,
      "replies": [
        {
          "id": 427865,
          "author_name": "remidi",
          "author_url": "",
          "post_date": "2018-11-26T10:05:08.150000",
          "content": "<p>model trains on the <strong>same number of samples</strong> in both the cases but in case of <em>version: 2</em> you are monitoring more number of times and calculating valid loss and accuaracy for 10 times rather than 1 time in <em>version: 1</em></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 427873,
          "author_name": "Salaryman",
          "author_url": "",
          "post_date": "2018-11-26T10:17:07.163000",
          "content": "<p>thanks so much remidi</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 429610,
          "author_name": "PC Jimmmy",
          "author_url": "",
          "post_date": "2018-11-29T04:10:21.303000",
          "content": "<p>remidi - is the net result of more monitoring/calculating any improvement in the results - or are you just able to see a bad model earlier?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 429942,
          "author_name": "remidi",
          "author_url": "",
          "post_date": "2018-11-29T15:14:53.407000",
          "content": "<p>@PCJimmmy \nI see no improvement in the model by doing more validation; just a better monitoring technique but keep in mind that more epochs will take more time because of validation for each epoch.\n<code>extra_time = validation_time * epochs</code></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 450664,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-01-05T13:52:48.257000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "427841": "Steps in Keras is meant for epoch performance check. I am using `STEPS=1000` so for every 1000 steps of `batchsize=96(example)` I would check validation score/error/performance.\nI am just ensuring that all the samples are completed in **epochs * steps * batchsize**\n\n \n\n     STEPS = 1000 \n     size = 128\n     batchsize = 96\n     total_samples = 49673579\n     EPOCHS = math.ceil(total_samples / (batchsize * STEPS) * 1)\n     valid_steps= math.ceil(34000/batchsize)",
    "427831": "Hi Guys,<br>\nI am wondering how would you set steps per epoch when you are using keras data generator to generate batches. I am using shuffle csv's and can't figure out a way to set the right amount of steps per epoch. Can anyone help with this?<br>\nThanks,",
    "427854": "Sorry I also have a rookie question. \nIf a train set have 10000 samples.\n\nIs there any difference on:\nVersion 1: Step=100, batch_size=100, epoch=1\nand\nVersion 2: Step=10, batch_size=100, epoch=10 ?\n\nIs version 2 just repeatedly training on the SAME 1000 data for 10 times? \nthanks\n\n"
  }
}