{
  "id": 73816,
  "title": "RNNs that worked alright (LB 0.941 by themselves)",
  "url": "/competitions/quickdraw-doodle-recognition/discussion/73816",
  "author_name": "RossWightman",
  "post_date": "2018-12-05T20:19:14.395000",
  "votes": 18,
  "comment_count": 10,
  "views": 0,
  "content": "<p>With a busy Oct and Nov I jumped into this challenge very late. I really wanted to try something a little different so I started with an RNN based model. Inspired a bit by a Tensorflow example I saw, <a href=\"https://github.com/tensorflow/models/blob/master/tutorials/rnn/quickdraw/train_model.py\">https://github.com/tensorflow/models/blob/master/tutorials/rnn/quickdraw/train_model.py</a>, I created a 'basic' 1d CNN + LSTM model in PyTorch. I created a more complex one by incorporating SE blocks from the SEResNext network and converting them to 1d with no strides.</p>\n\n<p>An ensemble of one 'basic' and one 'se' model ended up scoring 0.941 after a week of training. To round out the ensemble and boost the score into the 0.947 range I included some CNN based models that I trained on the side.  The CNNs mirrored the approach taken by many others in this challenge. I was hoping to include the best RNN and the best CNN model into a single one and fine tune the combined model to see how good it could get but ran out of time so they were all combined via ensemble.</p>\n\n<p>I posted a gist of just the cleaned up RNN models: <a href=\"https://gist.github.com/rwightman/893723e023ffccca4abf67f92cb76526\">https://gist.github.com/rwightman/893723e023ffccca4abf67f92cb76526</a></p>\n\n<p>I fed both of these with sequences that had been interpolated from the original stroke vectors to a consistent length somewhat arbirtarily chose to be 192 with x, y, t, and pen up inputs in the tensor normalized based on a partial pass of the dataset. </p>\n\n<p>I didn't have much (any) time to experiment with varations on these RNN models, the SE based one was still improving at the time of final submissions. Potential to match the pure CNN models? Perhaps...</p>",
  "messages": [
    {
      "id": 434040,
      "postDate": "2018-12-05T20:19:14.397Z",
      "content": "<p>With a busy Oct and Nov I jumped into this challenge very late. I really wanted to try something a little different so I started with an RNN based model. Inspired a bit by a Tensorflow example I saw, <a href=\"https://github.com/tensorflow/models/blob/master/tutorials/rnn/quickdraw/train_model.py\">https://github.com/tensorflow/models/blob/master/tutorials/rnn/quickdraw/train_model.py</a>, I created a 'basic' 1d CNN + LSTM model in PyTorch. I created a more complex one by incorporating SE blocks from the SEResNext network and converting them to 1d with no strides.</p>\n\n<p>An ensemble of one 'basic' and one 'se' model ended up scoring 0.941 after a week of training. To round out the ensemble and boost the score into the 0.947 range I included some CNN based models that I trained on the side.  The CNNs mirrored the approach taken by many others in this challenge. I was hoping to include the best RNN and the best CNN model into a single one and fine tune the combined model to see how good it could get but ran out of time so they were all combined via ensemble.</p>\n\n<p>I posted a gist of just the cleaned up RNN models: <a href=\"https://gist.github.com/rwightman/893723e023ffccca4abf67f92cb76526\">https://gist.github.com/rwightman/893723e023ffccca4abf67f92cb76526</a></p>\n\n<p>I fed both of these with sequences that had been interpolated from the original stroke vectors to a consistent length somewhat arbirtarily chose to be 192 with x, y, t, and pen up inputs in the tensor normalized based on a partial pass of the dataset. </p>\n\n<p>I didn't have much (any) time to experiment with varations on these RNN models, the SE based one was still improving at the time of final submissions. Potential to match the pure CNN models? Perhaps...</p>",
      "rawMarkdown": "With a busy Oct and Nov I jumped into this challenge very late. I really wanted to try something a little different so I started with an RNN based model. Inspired a bit by a Tensorflow example I saw, https://github.com/tensorflow/models/blob/master/tutorials/rnn/quickdraw/train_model.py, I created a 'basic' 1d CNN + LSTM model in PyTorch. I created a more complex one by incorporating SE blocks from the SEResNext network and converting them to 1d with no strides.\n\nAn ensemble of one 'basic' and one 'se' model ended up scoring 0.941 after a week of training. To round out the ensemble and boost the score into the 0.947 range I included some CNN based models that I trained on the side.  The CNNs mirrored the approach taken by many others in this challenge. I was hoping to include the best RNN and the best CNN model into a single one and fine tune the combined model to see how good it could get but ran out of time so they were all combined via ensemble.\n\nI posted a gist of just the cleaned up RNN models: https://gist.github.com/rwightman/893723e023ffccca4abf67f92cb76526\n\nI fed both of these with sequences that had been interpolated from the original stroke vectors to a consistent length somewhat arbirtarily chose to be 192 with x, y, t, and pen up inputs in the tensor normalized based on a partial pass of the dataset. \n\nI didn't have much (any) time to experiment with varations on these RNN models, the SE based one was still improving at the time of final submissions. Potential to match the pure CNN models? Perhaps...",
      "votes": 18
    },
    {
      "id": 439976,
      "postDate": "2018-12-16T18:56:05.133Z",
      "content": "<p>I ran through training of my 'basic sequence' model again from scratch with the fix (as pointed out by Heng) so that the reverse direction outputs are from the correct timestep.</p>\n\n<p>Single model scored .9415 private, .9427 public  ... so now better than my previous ensemble of two RNN models.</p>",
      "rawMarkdown": "I ran through training of my 'basic sequence' model again from scratch with the fix (as pointed out by Heng) so that the reverse direction outputs are from the correct timestep.\n\nSingle model scored .9415 private, .9427 public  ... so now better than my previous ensemble of two RNN models.",
      "votes": 1
    },
    {
      "id": 434355,
      "postDate": "2018-12-06T09:04:58.703Z",
      "content": "<p>\"I fed both of these with sequences that had been interpolated from the original stroke vectors to a consistent length somewhat arbirtarily chose to be 192 with x, y, t, and pen up inputs in the tensor normalized based on a partial pass of the dataset.\"</p>\n\n<p>do you have code for this? thanks!</p>",
      "rawMarkdown": "\"I fed both of these with sequences that had been interpolated from the original stroke vectors to a consistent length somewhat arbirtarily chose to be 192 with x, y, t, and pen up inputs in the tensor normalized based on a partial pass of the dataset.\"\n\ndo you have code for this? thanks!",
      "votes": 1
    },
    {
      "id": 434350,
      "postDate": "2018-12-06T08:49:12.163Z",
      "content": "<p>@ RossWightman</p>\n\n<p>This is the first time i used bidirectional lstm in  pytorch, so my understanding my be wrong. In your code:</p>\n\n<p><a href=\"https://gist.github.com/rwightman/893723e023ffccca4abf67f92cb76526\">https://gist.github.com/rwightman/893723e023ffccca4abf67f92cb76526</a>, for BasicSeqStrokeNet:</p>\n\n<pre><code>x, (h, c) = self.rnn(x)\n... \nlast_output = xo.gather(1, idx).squeeze(1)\n</code></pre>\n\n<p>which i think you are taking the last output of forward lstm and first output of backward lstm. But I think we should be taking last output of both forward and backward.</p>\n\n<p>(Note a simpler way is to use torch.cat ([h[-1], h[-2]])</p>\n\n<hr>\n\n<p>Now according to </p>\n\n<p><a href=\"https://towardsdatascience.com/understanding-bidirectional-rnn-in-pytorch-5bd25a5dd66\">https://towardsdatascience.com/understanding-bidirectional-rnn-in-pytorch-5bd25a5dd66</a></p>\n\n<p>\"We should take output[-1, :, :hidden_size] (normal RNN) and output[0, :, hidden_size:] (reverse RNN), concatenate them, and feed the result to the subsequent dense neural network.\"</p>\n\n<p>also from <a href=\"https://stackoverflow.com/questions/50856936/taking-the-last-state-from-bilstm-bigru-in-pytorch\">https://stackoverflow.com/questions/50856936/taking-the-last-state-from-bilstm-bigru-in-pytorch</a></p>\n\n<pre><code>last_forward = torch.gather(output_forward, 0, lengths - 1).squeeze(0)\nlast_backward = output_backward[0, :, :]\n</code></pre>",
      "rawMarkdown": "@ RossWightman\n\nThis is the first time i used bidirectional lstm in  pytorch, so my understanding my be wrong. In your code:\n\nhttps://gist.github.com/rwightman/893723e023ffccca4abf67f92cb76526, for BasicSeqStrokeNet:\n\n\n    x, (h, c) = self.rnn(x)\n    ... \n    last_output = xo.gather(1, idx).squeeze(1)\n\n\nwhich i think you are taking the last output of forward lstm and first output of backward lstm. But I think we should be taking last output of both forward and backward.\n\n(Note a simpler way is to use torch.cat ([h[-1], h[-2]])\n\n\n---\n\nNow according to \n\nhttps://towardsdatascience.com/understanding-bidirectional-rnn-in-pytorch-5bd25a5dd66\n\n\"We should take output[-1, :, :hidden_size] (normal RNN) and output[0, :, hidden_size:] (reverse RNN), concatenate them, and feed the result to the subsequent dense neural network.\"\n\nalso from https://stackoverflow.com/questions/50856936/taking-the-last-state-from-bilstm-bigru-in-pytorch\n\n\n    last_forward = torch.gather(output_forward, 0, lengths - 1).squeeze(0)\n    last_backward = output_backward[0, :, :]",
      "replies": [
        {
          "id": 434796,
          "postDate": "2018-12-07T00:47:55.183Z",
          "content": "<p>I was lead to believe, after reading some examples and docs, that the gather on output was necessary to grab the valid output at the end of each (possibly variable length) sequence as per the sequence lengths for each input in the packed/padded format. The output was already hidden size * 2 which happens as soon as you enable bidrectional for the LSTM so it's my understanding nothing is needed wrt to that for output.</p>",
          "rawMarkdown": "I was lead to believe, after reading some examples and docs, that the gather on output was necessary to grab the valid output at the end of each (possibly variable length) sequence as per the sequence lengths for each input in the packed/padded format. The output was already hidden size * 2 which happens as soon as you enable bidrectional for the LSTM so it's my understanding nothing is needed wrt to that for output."
        },
        {
          "id": 436673,
          "postDate": "2018-12-10T18:14:50.887Z",
          "content": "<p>The output is 2xhidden size. But the time is wrong. You are getting last of forward lstm ( hidden size) and first of backward lstm ( another hidden size)</p>",
          "rawMarkdown": "The output is 2xhidden size. But the time is wrong. You are getting last of forward lstm ( hidden size) and first of backward lstm ( another hidden size)"
        },
        {
          "id": 436675,
          "postDate": "2018-12-10T18:17:19.643Z",
          "content": "<p><a href=\"https://discuss.pytorch.org/t/bilstm-hidden-states-and-output-dont-match/3561\">https://discuss.pytorch.org/t/bilstm-hidden-states-and-output-dont-match/3561</a></p>\n\n<p><img src=\"https://discuss.pytorch.org/uploads/default/original/1X/b01a9e75313b8c651c7e3cb22c2161abbaa4751c.png\" alt=\"enter image description here\"></p>",
          "rawMarkdown": "https://discuss.pytorch.org/t/bilstm-hidden-states-and-output-dont-match/3561\n\n\n  ![enter image description here][1]\n\n\n  [1]: https://discuss.pytorch.org/uploads/default/original/1X/b01a9e75313b8c651c7e3cb22c2161abbaa4751c.png",
          "votes": 1
        },
        {
          "id": 437325,
          "postDate": "2018-12-11T18:04:25.903Z",
          "content": "<p>Okay, yeah I see now, so, with that in mind, and taking into account that I'm in batch first I have...</p>\n\n<pre><code>    idxf = (xl - 1).view(-1, 1).expand(xl.size(0), xo.size(2) // 2)\n    idxf = idxf.unsqueeze(1)\n    idxf = idxf.cuda()\n\n    max_seq_len = int(x_seq_len[0])\n    idxr = (max_seq_len - xl).view(-1, 1).expand(xl.size(0), xo.size(2) // 2)\n    idxr = idxr.unsqueeze(1)\n    idxr = idxr.cuda()\n\n    # Shape: (batch_size, rnn_hidden_dim)\n    last_output_forward = xo[:, :, :512].gather(1, idxf).squeeze(1)\n    last_output_reverse = xo[:, :, 512:].gather(1, idxr).squeeze(1)\n\n    last_output = torch.cat((last_output_forward, last_output_reverse), 1)\n</code></pre>\n\n<p>The one thing that I'm still not at all clear on is whether I can just use 0 for the reverse direction as in </p>\n\n<pre><code>    last_output_reverse = xo[:, 0, 512:]\n</code></pre>\n\n<p>I'm not at all clear on how the padding is handled in the reverse direction for bidir when I'm doing the pack/pad. Does padding end up on the left? so I'd have [pad, pad, pad, ...., 3, 2, 1] in the reverse or would it be [...., 3, 2, 1, pad, pad, pad] ... that would change what you need to pick off from the output...</p>\n\n<p>Or perhaps I could just cat((h[-2], h[-1])) as that is now matching my output. BUT, my seq lengths are uniform since I resampled them. I'm not sure if I can still use h as is if that's not the case....</p>\n\n<p><em>EDIT</em></p>\n\n<p>I disabled interpolation for inputs &lt; my max sequence to test the behaviour with variable sequence lengths. The last element in the reverse sequence is always at index 0 and no compensation for padding needed.  So, <code>last_output_reverse = xo[:, 0, 512:]</code> is indeed correct and my <code>last_output == torch.cat((h[-2], h[-1]))</code> now for both fixed and variable length sequences.</p>",
          "rawMarkdown": "Okay, yeah I see now, so, with that in mind, and taking into account that I'm in batch first I have...\n\n        idxf = (xl - 1).view(-1, 1).expand(xl.size(0), xo.size(2) // 2)\n        idxf = idxf.unsqueeze(1)\n        idxf = idxf.cuda()\n\n        max_seq_len = int(x_seq_len[0])\n        idxr = (max_seq_len - xl).view(-1, 1).expand(xl.size(0), xo.size(2) // 2)\n        idxr = idxr.unsqueeze(1)\n        idxr = idxr.cuda()\n\n        # Shape: (batch_size, rnn_hidden_dim)\n        last_output_forward = xo[:, :, :512].gather(1, idxf).squeeze(1)\n        last_output_reverse = xo[:, :, 512:].gather(1, idxr).squeeze(1)\n\n        last_output = torch.cat((last_output_forward, last_output_reverse), 1)\n\nThe one thing that I'm still not at all clear on is whether I can just use 0 for the reverse direction as in \n\n        last_output_reverse = xo[:, 0, 512:]\n\nI'm not at all clear on how the padding is handled in the reverse direction for bidir when I'm doing the pack/pad. Does padding end up on the left? so I'd have [pad, pad, pad, ...., 3, 2, 1] in the reverse or would it be [...., 3, 2, 1, pad, pad, pad] ... that would change what you need to pick off from the output...\n\nOr perhaps I could just cat((h[-2], h[-1])) as that is now matching my output. BUT, my seq lengths are uniform since I resampled them. I'm not sure if I can still use h as is if that's not the case....\n\n*EDIT*\n\nI disabled interpolation for inputs &lt; my max sequence to test the behaviour with variable sequence lengths. The last element in the reverse sequence is always at index 0 and no compensation for padding needed.  So, `last_output_reverse = xo[:, 0, 512:]` is indeed correct and my `last_output == torch.cat((h[-2], h[-1]))` now for both fixed and variable length sequences.\n\n"
        }
      ]
    },
    {
      "id": 434101,
      "postDate": "2018-12-05T22:10:49.487Z",
      "content": "<p>thanks for the post!</p>\n\n<p>\"An ensemble of one 'basic' and one 'se' model ended up scoring 0.941 after a week of training\"</p>\n\n<p>do you have LB score of each individual ones?</p>",
      "rawMarkdown": "thanks for the post!\n\n\"An ensemble of one 'basic' and one 'se' model ended up scoring 0.941 after a week of training\"\n\ndo you have LB score of each individual ones?",
      "replies": [
        {
          "id": 434109,
          "postDate": "2018-12-05T22:42:51.453Z",
          "content": "<p>In interest of saving submissions, I wasn't submitting single models once the local cv was reliable...</p>\n\n<p>I just submitted solo now to see..</p>\n\n<p>'basic seq' 0.93979 priv, 0.94029 pub\n'se seq' 0.93112, 0.93419</p>\n\n<p>So the basic model was much stronger. I feel that was largely due to it having more time to train. I created and started training the SE model several days later, it was still noteably improving at the point I had to stop. </p>",
          "rawMarkdown": "In interest of saving submissions, I wasn't submitting single models once the local cv was reliable...\n\nI just submitted solo now to see..\n\n'basic seq' 0.93979 priv, 0.94029 pub\n'se seq' 0.93112, 0.93419\n\nSo the basic model was much stronger. I feel that was largely due to it having more time to train. I created and started training the SE model several days later, it was still noteably improving at the point I had to stop. \n "
        }
      ]
    },
    {
      "id": 434095,
      "postDate": "2018-12-05T21:53:23.980Z",
      "content": "<p>Congrats and thanks for sharing your solution.</p>",
      "rawMarkdown": "Congrats and thanks for sharing your solution."
    }
  ],
  "comments": [
    {
      "id": 439976,
      "author_name": "RossWightman",
      "author_url": "",
      "post_date": "2018-12-16T18:56:05.133000",
      "content": "<p>I ran through training of my 'basic sequence' model again from scratch with the fix (as pointed out by Heng) so that the reverse direction outputs are from the correct timestep.</p>\n\n<p>Single model scored .9415 private, .9427 public  ... so now better than my previous ensemble of two RNN models.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 434355,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-12-06T09:04:58.703000",
      "content": "<p>\"I fed both of these with sequences that had been interpolated from the original stroke vectors to a consistent length somewhat arbirtarily chose to be 192 with x, y, t, and pen up inputs in the tensor normalized based on a partial pass of the dataset.\"</p>\n\n<p>do you have code for this? thanks!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 434350,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-12-06T08:49:12.163000",
      "content": "<p>@ RossWightman</p>\n\n<p>This is the first time i used bidirectional lstm in  pytorch, so my understanding my be wrong. In your code:</p>\n\n<p><a href=\"https://gist.github.com/rwightman/893723e023ffccca4abf67f92cb76526\">https://gist.github.com/rwightman/893723e023ffccca4abf67f92cb76526</a>, for BasicSeqStrokeNet:</p>\n\n<pre><code>x, (h, c) = self.rnn(x)\n... \nlast_output = xo.gather(1, idx).squeeze(1)\n</code></pre>\n\n<p>which i think you are taking the last output of forward lstm and first output of backward lstm. But I think we should be taking last output of both forward and backward.</p>\n\n<p>(Note a simpler way is to use torch.cat ([h[-1], h[-2]])</p>\n\n<hr>\n\n<p>Now according to </p>\n\n<p><a href=\"https://towardsdatascience.com/understanding-bidirectional-rnn-in-pytorch-5bd25a5dd66\">https://towardsdatascience.com/understanding-bidirectional-rnn-in-pytorch-5bd25a5dd66</a></p>\n\n<p>\"We should take output[-1, :, :hidden_size] (normal RNN) and output[0, :, hidden_size:] (reverse RNN), concatenate them, and feed the result to the subsequent dense neural network.\"</p>\n\n<p>also from <a href=\"https://stackoverflow.com/questions/50856936/taking-the-last-state-from-bilstm-bigru-in-pytorch\">https://stackoverflow.com/questions/50856936/taking-the-last-state-from-bilstm-bigru-in-pytorch</a></p>\n\n<pre><code>last_forward = torch.gather(output_forward, 0, lengths - 1).squeeze(0)\nlast_backward = output_backward[0, :, :]\n</code></pre>",
      "votes": 0,
      "replies": [
        {
          "id": 434796,
          "author_name": "RossWightman",
          "author_url": "",
          "post_date": "2018-12-07T00:47:55.183000",
          "content": "<p>I was lead to believe, after reading some examples and docs, that the gather on output was necessary to grab the valid output at the end of each (possibly variable length) sequence as per the sequence lengths for each input in the packed/padded format. The output was already hidden size * 2 which happens as soon as you enable bidrectional for the LSTM so it's my understanding nothing is needed wrt to that for output.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 436673,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2018-12-10T18:14:50.887000",
          "content": "<p>The output is 2xhidden size. But the time is wrong. You are getting last of forward lstm ( hidden size) and first of backward lstm ( another hidden size)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 436675,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2018-12-10T18:17:19.643000",
          "content": "<p><a href=\"https://discuss.pytorch.org/t/bilstm-hidden-states-and-output-dont-match/3561\">https://discuss.pytorch.org/t/bilstm-hidden-states-and-output-dont-match/3561</a></p>\n\n<p><img src=\"https://discuss.pytorch.org/uploads/default/original/1X/b01a9e75313b8c651c7e3cb22c2161abbaa4751c.png\" alt=\"enter image description here\"></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 437325,
          "author_name": "RossWightman",
          "author_url": "",
          "post_date": "2018-12-11T18:04:25.903000",
          "content": "<p>Okay, yeah I see now, so, with that in mind, and taking into account that I'm in batch first I have...</p>\n\n<pre><code>    idxf = (xl - 1).view(-1, 1).expand(xl.size(0), xo.size(2) // 2)\n    idxf = idxf.unsqueeze(1)\n    idxf = idxf.cuda()\n\n    max_seq_len = int(x_seq_len[0])\n    idxr = (max_seq_len - xl).view(-1, 1).expand(xl.size(0), xo.size(2) // 2)\n    idxr = idxr.unsqueeze(1)\n    idxr = idxr.cuda()\n\n    # Shape: (batch_size, rnn_hidden_dim)\n    last_output_forward = xo[:, :, :512].gather(1, idxf).squeeze(1)\n    last_output_reverse = xo[:, :, 512:].gather(1, idxr).squeeze(1)\n\n    last_output = torch.cat((last_output_forward, last_output_reverse), 1)\n</code></pre>\n\n<p>The one thing that I'm still not at all clear on is whether I can just use 0 for the reverse direction as in </p>\n\n<pre><code>    last_output_reverse = xo[:, 0, 512:]\n</code></pre>\n\n<p>I'm not at all clear on how the padding is handled in the reverse direction for bidir when I'm doing the pack/pad. Does padding end up on the left? so I'd have [pad, pad, pad, ...., 3, 2, 1] in the reverse or would it be [...., 3, 2, 1, pad, pad, pad] ... that would change what you need to pick off from the output...</p>\n\n<p>Or perhaps I could just cat((h[-2], h[-1])) as that is now matching my output. BUT, my seq lengths are uniform since I resampled them. I'm not sure if I can still use h as is if that's not the case....</p>\n\n<p><em>EDIT</em></p>\n\n<p>I disabled interpolation for inputs &lt; my max sequence to test the behaviour with variable sequence lengths. The last element in the reverse sequence is always at index 0 and no compensation for padding needed.  So, <code>last_output_reverse = xo[:, 0, 512:]</code> is indeed correct and my <code>last_output == torch.cat((h[-2], h[-1]))</code> now for both fixed and variable length sequences.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 434101,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-12-05T22:10:49.487000",
      "content": "<p>thanks for the post!</p>\n\n<p>\"An ensemble of one 'basic' and one 'se' model ended up scoring 0.941 after a week of training\"</p>\n\n<p>do you have LB score of each individual ones?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 434109,
          "author_name": "RossWightman",
          "author_url": "",
          "post_date": "2018-12-05T22:42:51.453000",
          "content": "<p>In interest of saving submissions, I wasn't submitting single models once the local cv was reliable...</p>\n\n<p>I just submitted solo now to see..</p>\n\n<p>'basic seq' 0.93979 priv, 0.94029 pub\n'se seq' 0.93112, 0.93419</p>\n\n<p>So the basic model was much stronger. I feel that was largely due to it having more time to train. I created and started training the SE model several days later, it was still noteably improving at the point I had to stop. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 434095,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2018-12-05T21:53:23.980000",
      "content": "<p>Congrats and thanks for sharing your solution.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "434040": "With a busy Oct and Nov I jumped into this challenge very late. I really wanted to try something a little different so I started with an RNN based model. Inspired a bit by a Tensorflow example I saw, https://github.com/tensorflow/models/blob/master/tutorials/rnn/quickdraw/train_model.py, I created a 'basic' 1d CNN + LSTM model in PyTorch. I created a more complex one by incorporating SE blocks from the SEResNext network and converting them to 1d with no strides.\n\nAn ensemble of one 'basic' and one 'se' model ended up scoring 0.941 after a week of training. To round out the ensemble and boost the score into the 0.947 range I included some CNN based models that I trained on the side.  The CNNs mirrored the approach taken by many others in this challenge. I was hoping to include the best RNN and the best CNN model into a single one and fine tune the combined model to see how good it could get but ran out of time so they were all combined via ensemble.\n\nI posted a gist of just the cleaned up RNN models: https://gist.github.com/rwightman/893723e023ffccca4abf67f92cb76526\n\nI fed both of these with sequences that had been interpolated from the original stroke vectors to a consistent length somewhat arbirtarily chose to be 192 with x, y, t, and pen up inputs in the tensor normalized based on a partial pass of the dataset. \n\nI didn't have much (any) time to experiment with varations on these RNN models, the SE based one was still improving at the time of final submissions. Potential to match the pure CNN models? Perhaps...",
    "439976": "I ran through training of my 'basic sequence' model again from scratch with the fix (as pointed out by Heng) so that the reverse direction outputs are from the correct timestep.\n\nSingle model scored .9415 private, .9427 public  ... so now better than my previous ensemble of two RNN models.",
    "434355": "\"I fed both of these with sequences that had been interpolated from the original stroke vectors to a consistent length somewhat arbirtarily chose to be 192 with x, y, t, and pen up inputs in the tensor normalized based on a partial pass of the dataset.\"\n\ndo you have code for this? thanks!",
    "434350": "@ RossWightman\n\nThis is the first time i used bidirectional lstm in  pytorch, so my understanding my be wrong. In your code:\n\nhttps://gist.github.com/rwightman/893723e023ffccca4abf67f92cb76526, for BasicSeqStrokeNet:\n\n\n    x, (h, c) = self.rnn(x)\n    ... \n    last_output = xo.gather(1, idx).squeeze(1)\n\n\nwhich i think you are taking the last output of forward lstm and first output of backward lstm. But I think we should be taking last output of both forward and backward.\n\n(Note a simpler way is to use torch.cat ([h[-1], h[-2]])\n\n\n---\n\nNow according to \n\nhttps://towardsdatascience.com/understanding-bidirectional-rnn-in-pytorch-5bd25a5dd66\n\n\"We should take output[-1, :, :hidden_size] (normal RNN) and output[0, :, hidden_size:] (reverse RNN), concatenate them, and feed the result to the subsequent dense neural network.\"\n\nalso from https://stackoverflow.com/questions/50856936/taking-the-last-state-from-bilstm-bigru-in-pytorch\n\n\n    last_forward = torch.gather(output_forward, 0, lengths - 1).squeeze(0)\n    last_backward = output_backward[0, :, :]",
    "434101": "thanks for the post!\n\n\"An ensemble of one 'basic' and one 'se' model ended up scoring 0.941 after a week of training\"\n\ndo you have LB score of each individual ones?",
    "434095": "Congrats and thanks for sharing your solution."
  }
}