{
  "id": 70972,
  "title": "Speedup for LSTM training",
  "url": "/competitions/quickdraw-doodle-recognition/discussion/70972",
  "author_name": "Kees van Rooijen",
  "post_date": "2018-11-08T21:55:32.658000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Just wanted to share a pretty straight-forward optimization. For training a RNN/LSTM we need to pad each sequence in a batch with zeroes so that every sequence has the same length. However, the length of sequences varies greatly. Now if we change:</p>\n\n<pre><code>batch = load_batch(batch_size)\nbatch = pad(batch)\n</code></pre>\n\n<p>to:</p>\n\n<pre><code>N = 50 # for example\nsuperbatch = load_batch(batch_size*N)\nsuperbatch = superbatch.sort_by(sequence_length)\nfor i in range(N):\n  batch = superbatch(i*batch_size:(i+1)*batch_size)\n  batch = pad(batch)\n</code></pre>\n\n<p>We create batches where the amount of space used for padding is relatively low. For me this gave a 3x speedup when used with N=50 or so. </p>",
  "messages": [
    {
      "id": 417859,
      "postDate": "2018-11-08T21:55:32.660Z",
      "content": "<p>Just wanted to share a pretty straight-forward optimization. For training a RNN/LSTM we need to pad each sequence in a batch with zeroes so that every sequence has the same length. However, the length of sequences varies greatly. Now if we change:</p>\n\n<pre><code>batch = load_batch(batch_size)\nbatch = pad(batch)\n</code></pre>\n\n<p>to:</p>\n\n<pre><code>N = 50 # for example\nsuperbatch = load_batch(batch_size*N)\nsuperbatch = superbatch.sort_by(sequence_length)\nfor i in range(N):\n  batch = superbatch(i*batch_size:(i+1)*batch_size)\n  batch = pad(batch)\n</code></pre>\n\n<p>We create batches where the amount of space used for padding is relatively low. For me this gave a 3x speedup when used with N=50 or so. </p>",
      "rawMarkdown": "Just wanted to share a pretty straight-forward optimization. For training a RNN/LSTM we need to pad each sequence in a batch with zeroes so that every sequence has the same length. However, the length of sequences varies greatly. Now if we change:\n\n    batch = load_batch(batch_size)\n    batch = pad(batch)\n\nto:\n\n    N = 50 # for example\n    superbatch = load_batch(batch_size*N)\n    superbatch = superbatch.sort_by(sequence_length)\n    for i in range(N):\n      batch = superbatch(i*batch_size:(i+1)*batch_size)\n      batch = pad(batch)\n\nWe create batches where the amount of space used for padding is relatively low. For me this gave a 3x speedup when used with N=50 or so. ",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "417859": "Just wanted to share a pretty straight-forward optimization. For training a RNN/LSTM we need to pad each sequence in a batch with zeroes so that every sequence has the same length. However, the length of sequences varies greatly. Now if we change:\n\n    batch = load_batch(batch_size)\n    batch = pad(batch)\n\nto:\n\n    N = 50 # for example\n    superbatch = load_batch(batch_size*N)\n    superbatch = superbatch.sort_by(sequence_length)\n    for i in range(N):\n      batch = superbatch(i*batch_size:(i+1)*batch_size)\n      batch = pad(batch)\n\nWe create batches where the amount of space used for padding is relatively low. For me this gave a 3x speedup when used with N=50 or so. "
  }
}