{
  "id": 73967,
  "title": "8th place novel solution",
  "url": "/competitions/quickdraw-doodle-recognition/discussion/73967",
  "author_name": "Aleksey Nozdryn-Plotnicki",
  "post_date": "2018-12-07T05:57:00.176000",
  "votes": 34,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Please pardon the quality of the following, but I’ve rushed it out. I will plan to release a detailed blog post/paper at a later date. My approach is novel, complex, and a challenge to communicate, but I wanted to get this information out while everyone was still interested!</p>\n\n<h1>Key Insight</h1>\n\n<p>I leveraged a principal lesson in Deep Learning: Let the network learn the features rather than hand crafting them. The standard approach is to rasterize the drawings, crudely encode time as colour, and then pass to a pre-trained RGB ResNet. Instead, I implement a differentiable, trainable module to do this. Furthermore, I replace the standard nn.Conv2d(3, 64, kernel_size=7, stride=2, ...) and 2x2 MaxPool at the start of a ResNet with the same module.</p>\n\n<p>High Level Visual:\n<img src=\"http://alekseynp.com/images/quickdraw_diagram1.png\" alt=\"enter image description here\"></p>\n\n<p>Key benefits:</p>\n\n<ul>\n<li><p>Deep features are computed from the time component rather than some RGB hack</p></li>\n<li><p>Deep features also computed from stroke data, leveraging the native format before “discarding” it and switching to images. Access to connectedness information.</p></li>\n<li><p>Uses efficiency of convolutions on grids where a pure sequence or point cloud model fails to</p></li>\n</ul>\n\n<h1>Implementation</h1>\n\n<p>Enormous amount of details that I will cover later.</p>\n\n<p>Detailed visual #1:\n<img src=\"http://alekseynp.com/images/quickdraw_diagram2.png\" alt=\"enter image description here\"></p>\n\n<ul>\n<li><p>Begin with strokes as defined by a series of points.</p></li>\n<li><p>Difference the x’s and y’s to get dx, dy segment vectors. Average the ts.</p></li>\n<li><p>Now we have strokes as defined by a series of segments.</p></li>\n<li><p>Process strokes with a sequence module to generate 32 features.</p></li>\n<li><p>Unroll those features with a window 2 convolution to generate 64 features per original point.</p></li>\n</ul>\n\n<p>Detailed visual #2:\n<img src=\"http://alekseynp.com/images/quickdraw_diagram3.png\" alt=\"enter image description here\"></p>\n\n<ul>\n<li><p>Draw points into a 32x32x64 image as per the diagram.</p></li>\n<li><p>When points collide their feature vectors are averaged. I think I would have max-pooled if I could have implemented it.</p></li>\n<li><p>At this stage in the network a 32x32 image could be thought of as equivalent to having started a normal image ResNet at 256x256.</p></li>\n</ul>\n\n<h1>Sequence Module</h1>\n\n<pre><code>Conv1d(3,  32, kernel_size=3, stride=1, padding=1, dilation=1)\nBatchNorm1d(32)\nReLU(inplace=True)\nConv1d(32, 32, kernel_size=3, stride=1, padding=2, dilation=2)\nBatchNorm1d(32)\nReLU(inplace=True)\nConv1d(32, 32, kernel_size=3, stride=1, padding=4, dilation=4)\nBatchNorm1d(32)\nReLU(inplace=True)\nConv1d(32, 32, kernel_size=3, stride=1, padding=8, dilation=8)\nBatchNorm1d(32)\nReLU(inplace=True)\nConv1d(32, 64, kernel_size=2, stride=1, padding=(1,0))\n</code></pre>\n\n<h1>Rasterization Module</h1>\n\n<pre><code>from apex import amp\nimport torch\nfrom torch.autograd import Function\n\nclass PointsToImage(Function):\n   @staticmethod\n   @amp.float_function\n   def forward(ctx, i, v):\n       device = i.device\n       batch_size, _, num_input_points = i.size()\n       feature_size = v.size()[2]\n\n       batch_idx = torch.arange(batch_size, device=device).view(-1, 1).repeat(1, num_input_points).view(-1)\n       idx_full = torch.cat([batch_idx.unsqueeze(0), i.permute(1, 0, 2).contiguous().view(2, -1)], dim=0)\n\n       v_full = v.contiguous().view(batch_size * num_input_points, feature_size)\n       mat_sparse = torch.cuda.sparse.FloatTensor(idx_full, v_full)\n       mat_dense = mat_sparse.to_dense()\n\n       ones_full = torch.ones(v_full.size(), device=device)\n       mat_sparse_count = torch.sparse.FloatTensor(idx_full, ones_full)\n       mat_dense_count = mat_sparse_count.to_dense()\n\n       ctx.save_for_backward(idx_full, mat_dense_count)\n\n       return mat_dense / torch.clamp(mat_dense_count, 1, 1e4)\n\n   @staticmethod\n   @amp.float_function\n   def backward(ctx, grad_output):\n       idx_full, mat_dense_count = ctx.saved_tensors\n       grad_i = grad_v = None\n\n       batch_size, _, _, feature_size = grad_output.size()\n\n       if ctx.needs_input_grad[0]:\n           raise Exception(\"Indices aren't differentiable.\")\n       if ctx.needs_input_grad[1]:\n           grad = grad_output[idx_full[0], idx_full[1], idx_full[2]]\n           coef = mat_dense_count[idx_full[0], idx_full[1], idx_full[2]]\n           grad_v = grad / coef\n           grad_v = grad_v.view(batch_size, -1, feature_size)\n\n       if isinstance(grad_output, torch.cuda.FloatTensor):\n           return grad_i, grad_v\n       else:\n           return grad_i, grad_v.half()\n\npoints_to_image = PointsToImage.apply\n</code></pre>\n\n<p>Various other details:</p>\n\n<ul>\n<li><p>Packing strokes of varying length into tensors of fixed size in order to do 1D CNNs is a non-trivial thing to do and beyond the scope of this post</p></li>\n<li><p>Pytorch 0.4.1</p></li>\n<li><p>Used NVIDIA’s apex amp (<a href=\"https://github.com/NVIDIA/apex/tree/master/apex/amp\">https://github.com/NVIDIA/apex/tree/master/apex/amp</a>) to train exclusively at half precision</p></li>\n<li><p>Trained on all the data</p></li>\n<li><p>Raw not simplified</p></li>\n<li><p>LMDB for memory mapped data</p></li>\n<li><p>Adam optimizer</p></li>\n<li><p>Models typically took about 2.5-3 days to converge on a system with a 1080Ti and a Titan</p></li>\n<li><p>Used pre-trained imagenet weights</p></li>\n<li><p>Froze those weights and only trained my additional modules for the first 1k-5k iterations</p></li>\n<li><p>Held out 50k examples for validation during training</p></li>\n<li><p>Held out 1 million examples for blending</p></li>\n<li><p>Probably never completed even 2 complete passes through all the data during training. Convergence came first.</p></li>\n<li><p>Used gradient accumulation via multiple backwards calls in Pytorch to finish training at huge batch sizes</p></li>\n</ul>\n\n<h1>Results</h1>\n\n<p>SEResNeXt50 32x4d at core.\nLocal validation:</p>\n\n<ul>\n<li><p>Acc@1: 0.8648</p></li>\n<li><p>Mapk3: 0.9073</p></li>\n<li><p>CE Loss: 0.5092</p></li>\n<li><p>Public LB: 0.94781</p></li>\n<li><p>Private LB: 0.94915</p></li>\n</ul>\n\n<p>Best ensemble:</p>\n\n<p>Six best models as measured by local CE Loss.</p>\n\n<ul>\n<li><p>2 x SEResNeXt50 at core</p></li>\n<li><p>3 x SEResNeXt101 at core</p></li>\n<li><p>1 x ResNet34 at core</p></li>\n<li><p>Weighted arithmetic mean of probabilities. Weights = (1/loss)**24</p></li>\n</ul>\n\n<p>Local validation:</p>\n\n<ul>\n<li><p>Acc#1: 0.8683</p></li>\n<li><p>Mapk3: 0.9101</p></li>\n<li><p>CE Loss: 0.4937</p></li>\n<li><p>Public LB: 0.95142</p></li>\n<li><p>Private LB: 0.95101</p></li>\n</ul>\n\n<p>Very slight optimization bias in local validation score. I used my validation set to select “six best” and the 24 exponent.\nZero overfit to public LB. hence why I moved up from 12 to 8 in the shakeup.</p>\n\n<p>Other comments:</p>\n\n<ul>\n<li><p>No RNNs in my ensemble</p></li>\n<li><p>No Image CNNs in my ensemble</p></li>\n<li><p>Kaggle competitions are a hell of an environment to try to do something novel in. I spent weeks messing around with Deep Residual PointNet++ networks, but never surpassed public LB 0.916</p></li>\n</ul>\n\n<p>Finally:</p>\n\n<ul>\n<li><p>GG to everyone</p></li>\n<li><p>Super proud of Guanshuo Xu for solid solo performance and disciplined execution enabling a move from 4th to 2nd into the money in the Private LB shakeup!</p></li>\n</ul>",
  "messages": [
    {
      "id": 434902,
      "postDate": "2018-12-07T05:57:00.177Z",
      "content": "<p>Please pardon the quality of the following, but I’ve rushed it out. I will plan to release a detailed blog post/paper at a later date. My approach is novel, complex, and a challenge to communicate, but I wanted to get this information out while everyone was still interested!</p>\n\n<h1>Key Insight</h1>\n\n<p>I leveraged a principal lesson in Deep Learning: Let the network learn the features rather than hand crafting them. The standard approach is to rasterize the drawings, crudely encode time as colour, and then pass to a pre-trained RGB ResNet. Instead, I implement a differentiable, trainable module to do this. Furthermore, I replace the standard nn.Conv2d(3, 64, kernel_size=7, stride=2, ...) and 2x2 MaxPool at the start of a ResNet with the same module.</p>\n\n<p>High Level Visual:\n<img src=\"http://alekseynp.com/images/quickdraw_diagram1.png\" alt=\"enter image description here\"></p>\n\n<p>Key benefits:</p>\n\n<ul>\n<li><p>Deep features are computed from the time component rather than some RGB hack</p></li>\n<li><p>Deep features also computed from stroke data, leveraging the native format before “discarding” it and switching to images. Access to connectedness information.</p></li>\n<li><p>Uses efficiency of convolutions on grids where a pure sequence or point cloud model fails to</p></li>\n</ul>\n\n<h1>Implementation</h1>\n\n<p>Enormous amount of details that I will cover later.</p>\n\n<p>Detailed visual #1:\n<img src=\"http://alekseynp.com/images/quickdraw_diagram2.png\" alt=\"enter image description here\"></p>\n\n<ul>\n<li><p>Begin with strokes as defined by a series of points.</p></li>\n<li><p>Difference the x’s and y’s to get dx, dy segment vectors. Average the ts.</p></li>\n<li><p>Now we have strokes as defined by a series of segments.</p></li>\n<li><p>Process strokes with a sequence module to generate 32 features.</p></li>\n<li><p>Unroll those features with a window 2 convolution to generate 64 features per original point.</p></li>\n</ul>\n\n<p>Detailed visual #2:\n<img src=\"http://alekseynp.com/images/quickdraw_diagram3.png\" alt=\"enter image description here\"></p>\n\n<ul>\n<li><p>Draw points into a 32x32x64 image as per the diagram.</p></li>\n<li><p>When points collide their feature vectors are averaged. I think I would have max-pooled if I could have implemented it.</p></li>\n<li><p>At this stage in the network a 32x32 image could be thought of as equivalent to having started a normal image ResNet at 256x256.</p></li>\n</ul>\n\n<h1>Sequence Module</h1>\n\n<pre><code>Conv1d(3,  32, kernel_size=3, stride=1, padding=1, dilation=1)\nBatchNorm1d(32)\nReLU(inplace=True)\nConv1d(32, 32, kernel_size=3, stride=1, padding=2, dilation=2)\nBatchNorm1d(32)\nReLU(inplace=True)\nConv1d(32, 32, kernel_size=3, stride=1, padding=4, dilation=4)\nBatchNorm1d(32)\nReLU(inplace=True)\nConv1d(32, 32, kernel_size=3, stride=1, padding=8, dilation=8)\nBatchNorm1d(32)\nReLU(inplace=True)\nConv1d(32, 64, kernel_size=2, stride=1, padding=(1,0))\n</code></pre>\n\n<h1>Rasterization Module</h1>\n\n<pre><code>from apex import amp\nimport torch\nfrom torch.autograd import Function\n\nclass PointsToImage(Function):\n   @staticmethod\n   @amp.float_function\n   def forward(ctx, i, v):\n       device = i.device\n       batch_size, _, num_input_points = i.size()\n       feature_size = v.size()[2]\n\n       batch_idx = torch.arange(batch_size, device=device).view(-1, 1).repeat(1, num_input_points).view(-1)\n       idx_full = torch.cat([batch_idx.unsqueeze(0), i.permute(1, 0, 2).contiguous().view(2, -1)], dim=0)\n\n       v_full = v.contiguous().view(batch_size * num_input_points, feature_size)\n       mat_sparse = torch.cuda.sparse.FloatTensor(idx_full, v_full)\n       mat_dense = mat_sparse.to_dense()\n\n       ones_full = torch.ones(v_full.size(), device=device)\n       mat_sparse_count = torch.sparse.FloatTensor(idx_full, ones_full)\n       mat_dense_count = mat_sparse_count.to_dense()\n\n       ctx.save_for_backward(idx_full, mat_dense_count)\n\n       return mat_dense / torch.clamp(mat_dense_count, 1, 1e4)\n\n   @staticmethod\n   @amp.float_function\n   def backward(ctx, grad_output):\n       idx_full, mat_dense_count = ctx.saved_tensors\n       grad_i = grad_v = None\n\n       batch_size, _, _, feature_size = grad_output.size()\n\n       if ctx.needs_input_grad[0]:\n           raise Exception(\"Indices aren't differentiable.\")\n       if ctx.needs_input_grad[1]:\n           grad = grad_output[idx_full[0], idx_full[1], idx_full[2]]\n           coef = mat_dense_count[idx_full[0], idx_full[1], idx_full[2]]\n           grad_v = grad / coef\n           grad_v = grad_v.view(batch_size, -1, feature_size)\n\n       if isinstance(grad_output, torch.cuda.FloatTensor):\n           return grad_i, grad_v\n       else:\n           return grad_i, grad_v.half()\n\npoints_to_image = PointsToImage.apply\n</code></pre>\n\n<p>Various other details:</p>\n\n<ul>\n<li><p>Packing strokes of varying length into tensors of fixed size in order to do 1D CNNs is a non-trivial thing to do and beyond the scope of this post</p></li>\n<li><p>Pytorch 0.4.1</p></li>\n<li><p>Used NVIDIA’s apex amp (<a href=\"https://github.com/NVIDIA/apex/tree/master/apex/amp\">https://github.com/NVIDIA/apex/tree/master/apex/amp</a>) to train exclusively at half precision</p></li>\n<li><p>Trained on all the data</p></li>\n<li><p>Raw not simplified</p></li>\n<li><p>LMDB for memory mapped data</p></li>\n<li><p>Adam optimizer</p></li>\n<li><p>Models typically took about 2.5-3 days to converge on a system with a 1080Ti and a Titan</p></li>\n<li><p>Used pre-trained imagenet weights</p></li>\n<li><p>Froze those weights and only trained my additional modules for the first 1k-5k iterations</p></li>\n<li><p>Held out 50k examples for validation during training</p></li>\n<li><p>Held out 1 million examples for blending</p></li>\n<li><p>Probably never completed even 2 complete passes through all the data during training. Convergence came first.</p></li>\n<li><p>Used gradient accumulation via multiple backwards calls in Pytorch to finish training at huge batch sizes</p></li>\n</ul>\n\n<h1>Results</h1>\n\n<p>SEResNeXt50 32x4d at core.\nLocal validation:</p>\n\n<ul>\n<li><p>Acc@1: 0.8648</p></li>\n<li><p>Mapk3: 0.9073</p></li>\n<li><p>CE Loss: 0.5092</p></li>\n<li><p>Public LB: 0.94781</p></li>\n<li><p>Private LB: 0.94915</p></li>\n</ul>\n\n<p>Best ensemble:</p>\n\n<p>Six best models as measured by local CE Loss.</p>\n\n<ul>\n<li><p>2 x SEResNeXt50 at core</p></li>\n<li><p>3 x SEResNeXt101 at core</p></li>\n<li><p>1 x ResNet34 at core</p></li>\n<li><p>Weighted arithmetic mean of probabilities. Weights = (1/loss)**24</p></li>\n</ul>\n\n<p>Local validation:</p>\n\n<ul>\n<li><p>Acc#1: 0.8683</p></li>\n<li><p>Mapk3: 0.9101</p></li>\n<li><p>CE Loss: 0.4937</p></li>\n<li><p>Public LB: 0.95142</p></li>\n<li><p>Private LB: 0.95101</p></li>\n</ul>\n\n<p>Very slight optimization bias in local validation score. I used my validation set to select “six best” and the 24 exponent.\nZero overfit to public LB. hence why I moved up from 12 to 8 in the shakeup.</p>\n\n<p>Other comments:</p>\n\n<ul>\n<li><p>No RNNs in my ensemble</p></li>\n<li><p>No Image CNNs in my ensemble</p></li>\n<li><p>Kaggle competitions are a hell of an environment to try to do something novel in. I spent weeks messing around with Deep Residual PointNet++ networks, but never surpassed public LB 0.916</p></li>\n</ul>\n\n<p>Finally:</p>\n\n<ul>\n<li><p>GG to everyone</p></li>\n<li><p>Super proud of Guanshuo Xu for solid solo performance and disciplined execution enabling a move from 4th to 2nd into the money in the Private LB shakeup!</p></li>\n</ul>",
      "rawMarkdown": "Please pardon the quality of the following, but I’ve rushed it out. I will plan to release a detailed blog post/paper at a later date. My approach is novel, complex, and a challenge to communicate, but I wanted to get this information out while everyone was still interested!\n\n# Key Insight\nI leveraged a principal lesson in Deep Learning: Let the network learn the features rather than hand crafting them. The standard approach is to rasterize the drawings, crudely encode time as colour, and then pass to a pre-trained RGB ResNet. Instead, I implement a differentiable, trainable module to do this. Furthermore, I replace the standard nn.Conv2d(3, 64, kernel_size=7, stride=2, ...) and 2x2 MaxPool at the start of a ResNet with the same module.\n\nHigh Level Visual:\n![enter image description here][1]\n\nKey benefits:\n\n- Deep features are computed from the time component rather than some RGB hack\n\n- Deep features also computed from stroke data, leveraging the native format before “discarding” it and switching to images. Access to connectedness information.\n\n- Uses efficiency of convolutions on grids where a pure sequence or point cloud model fails to\n\n# Implementation\nEnormous amount of details that I will cover later.\n\nDetailed visual #1:\n![enter image description here][2]\n\n- Begin with strokes as defined by a series of points.\n\n- Difference the x’s and y’s to get dx, dy segment vectors. Average the ts.\n\n- Now we have strokes as defined by a series of segments.\n\n- Process strokes with a sequence module to generate 32 features.\n\n- Unroll those features with a window 2 convolution to generate 64 features per original point.\n\nDetailed visual #2:\n![enter image description here][3]\n\n- Draw points into a 32x32x64 image as per the diagram.\n\n- When points collide their feature vectors are averaged. I think I would have max-pooled if I could have implemented it.\n\n- At this stage in the network a 32x32 image could be thought of as equivalent to having started a normal image ResNet at 256x256.\n\n# Sequence Module\n\n    Conv1d(3,  32, kernel_size=3, stride=1, padding=1, dilation=1)\n    BatchNorm1d(32)\n    ReLU(inplace=True)\n    Conv1d(32, 32, kernel_size=3, stride=1, padding=2, dilation=2)\n    BatchNorm1d(32)\n    ReLU(inplace=True)\n    Conv1d(32, 32, kernel_size=3, stride=1, padding=4, dilation=4)\n    BatchNorm1d(32)\n    ReLU(inplace=True)\n    Conv1d(32, 32, kernel_size=3, stride=1, padding=8, dilation=8)\n    BatchNorm1d(32)\n    ReLU(inplace=True)\n    Conv1d(32, 64, kernel_size=2, stride=1, padding=(1,0))\n\n\n# Rasterization Module\n\n    from apex import amp\n    import torch\n    from torch.autograd import Function\n    \n    class PointsToImage(Function):\n       @staticmethod\n       @amp.float_function\n       def forward(ctx, i, v):\n           device = i.device\n           batch_size, _, num_input_points = i.size()\n           feature_size = v.size()[2]\n    \n           batch_idx = torch.arange(batch_size, device=device).view(-1, 1).repeat(1, num_input_points).view(-1)\n           idx_full = torch.cat([batch_idx.unsqueeze(0), i.permute(1, 0, 2).contiguous().view(2, -1)], dim=0)\n    \n           v_full = v.contiguous().view(batch_size * num_input_points, feature_size)\n           mat_sparse = torch.cuda.sparse.FloatTensor(idx_full, v_full)\n           mat_dense = mat_sparse.to_dense()\n    \n           ones_full = torch.ones(v_full.size(), device=device)\n           mat_sparse_count = torch.sparse.FloatTensor(idx_full, ones_full)\n           mat_dense_count = mat_sparse_count.to_dense()\n    \n           ctx.save_for_backward(idx_full, mat_dense_count)\n    \n           return mat_dense / torch.clamp(mat_dense_count, 1, 1e4)\n    \n       @staticmethod\n       @amp.float_function\n       def backward(ctx, grad_output):\n           idx_full, mat_dense_count = ctx.saved_tensors\n           grad_i = grad_v = None\n    \n           batch_size, _, _, feature_size = grad_output.size()\n    \n           if ctx.needs_input_grad[0]:\n               raise Exception(\"Indices aren't differentiable.\")\n           if ctx.needs_input_grad[1]:\n               grad = grad_output[idx_full[0], idx_full[1], idx_full[2]]\n               coef = mat_dense_count[idx_full[0], idx_full[1], idx_full[2]]\n               grad_v = grad / coef\n               grad_v = grad_v.view(batch_size, -1, feature_size)\n    \n           if isinstance(grad_output, torch.cuda.FloatTensor):\n               return grad_i, grad_v\n           else:\n               return grad_i, grad_v.half()\n    \n    points_to_image = PointsToImage.apply\n\n\nVarious other details:\n\n- Packing strokes of varying length into tensors of fixed size in order to do 1D CNNs is a non-trivial thing to do and beyond the scope of this post\n\n- Pytorch 0.4.1\n\n- Used NVIDIA’s apex amp (https://github.com/NVIDIA/apex/tree/master/apex/amp) to train exclusively at half precision\n\n- Trained on all the data\n\n- Raw not simplified\n\n- LMDB for memory mapped data\n\n- Adam optimizer\n\n- Models typically took about 2.5-3 days to converge on a system with a 1080Ti and a Titan\n\n- Used pre-trained imagenet weights\n\n- Froze those weights and only trained my additional modules for the first 1k-5k iterations\n\n- Held out 50k examples for validation during training\n\n- Held out 1 million examples for blending\n\n- Probably never completed even 2 complete passes through all the data during training. Convergence came first.\n\n- Used gradient accumulation via multiple backwards calls in Pytorch to finish training at huge batch sizes\n\n# Results\n\nSEResNeXt50 32x4d at core.\nLocal validation:\n\n- Acc@1: 0.8648\n\n- Mapk3: 0.9073\n\n- CE Loss: 0.5092\n\n- Public LB: 0.94781\n\n- Private LB: 0.94915\n\nBest ensemble:\n\nSix best models as measured by local CE Loss.\n\n- 2 x SEResNeXt50 at core\n\n- 3 x SEResNeXt101 at core\n\n- 1 x ResNet34 at core\n\n- Weighted arithmetic mean of probabilities. Weights = (1/loss)**24\n\nLocal validation:\n\n- Acc#1: 0.8683\n\n- Mapk3: 0.9101\n\n- CE Loss: 0.4937\n\n- Public LB: 0.95142\n\n- Private LB: 0.95101\n\nVery slight optimization bias in local validation score. I used my validation set to select “six best” and the 24 exponent.\nZero overfit to public LB. hence why I moved up from 12 to 8 in the shakeup.\n\nOther comments:\n\n- No RNNs in my ensemble\n\n- No Image CNNs in my ensemble\n\n- Kaggle competitions are a hell of an environment to try to do something novel in. I spent weeks messing around with Deep Residual PointNet++ networks, but never surpassed public LB 0.916\n\nFinally:\n\n- GG to everyone\n\n- Super proud of Guanshuo Xu for solid solo performance and disciplined execution enabling a move from 4th to 2nd into the money in the Private LB shakeup!\n\n  [1]: http://alekseynp.com/images/quickdraw_diagram1.png\n  [2]: http://alekseynp.com/images/quickdraw_diagram2.png\n  [3]: http://alekseynp.com/images/quickdraw_diagram3.png",
      "votes": 34
    },
    {
      "id": 436252,
      "postDate": "2018-12-10T01:39:01.940Z",
      "content": "<p>I just put code up here: <a href=\"https://github.com/alekseynp/kaggle-quickdraw\">https://github.com/alekseynp/kaggle-quickdraw</a></p>",
      "rawMarkdown": "I just put code up here: https://github.com/alekseynp/kaggle-quickdraw",
      "votes": 5,
      "replies": [
        {
          "id": 471437,
          "postDate": "2019-02-14T13:10:46.753Z",
          "content": "<p>Thanks for sharing code :) </p>",
          "rawMarkdown": "Thanks for sharing code :) "
        }
      ]
    },
    {
      "id": 436654,
      "postDate": "2018-12-10T17:52:41.997Z",
      "content": "<p>for resampling to 256 points, did you use \"def resample_to(drawing, n)\" from</p>\n\n<p><a href=\"https://github.com/alekseynp/kaggle-quickdraw/blob/master/data_set.py\">https://github.com/alekseynp/kaggle-quickdraw/blob/master/data_set.py</a></p>\n\n<p>Thanks!</p>\n\n<p>it seems that making the drawing equal lengths also make the data more consistent and improve rnn results.</p>",
      "rawMarkdown": "for resampling to 256 points, did you use \"def resample_to(drawing, n)\" from\n\nhttps://github.com/alekseynp/kaggle-quickdraw/blob/master/data_set.py\n\nThanks!\n\n\nit seems that making the drawing equal lengths also make the data more consistent and improve rnn results.",
      "votes": 1,
      "replies": [
        {
          "id": 436694,
          "postDate": "2018-12-10T18:54:32.913Z",
          "content": "<p>Yes I did.\nI don't know if it creates as much consistency as I would like. Within any one drawing the segments are close to the same length, but not necessarily between drawings. I didn't study the raw data close enough to know. How consistent are lengths? By forcing \"long\" drawings and \"short\" drawings all to have 256 points, segment lengths may actually vary more after <code>resample_to</code></p>",
          "rawMarkdown": "Yes I did.\nI don't know if it creates as much consistency as I would like. Within any one drawing the segments are close to the same length, but not necessarily between drawings. I didn't study the raw data close enough to know. How consistent are lengths? By forcing \"long\" drawings and \"short\" drawings all to have 256 points, segment lengths may actually vary more after `resample_to`"
        }
      ]
    },
    {
      "id": 436400,
      "postDate": "2018-12-10T08:29:59.297Z",
      "content": "<p>Hi Aleksey, It seems that you wrote Private LB: 0.94804 as  Private LB: 0.94915 by mistake. I run your test.py getting the result of priv 0.94804 pub 0.94780. Otherwise your submission had a much better performance in priv than pub.</p>",
      "rawMarkdown": "Hi Aleksey, It seems that you wrote Private LB: 0.94804 as  Private LB: 0.94915 by mistake. I run your test.py getting the result of priv 0.94804 pub 0.94780. Otherwise your submission had a much better performance in priv than pub.",
      "votes": 1,
      "replies": [
        {
          "id": 436691,
          "postDate": "2018-12-10T18:49:53.693Z",
          "content": "<p>Oops! Thanks</p>",
          "rawMarkdown": "Oops! Thanks"
        }
      ]
    },
    {
      "id": 435289,
      "postDate": "2018-12-07T19:59:09.463Z",
      "content": "<p>Absolutely wonderful approach!</p>",
      "rawMarkdown": "Absolutely wonderful approach!",
      "votes": 1
    },
    {
      "id": 435072,
      "postDate": "2018-12-07T13:04:13.670Z",
      "content": "<p>Thanks for sharing. it looks very interesting. look forward to your blog/paper. </p>",
      "rawMarkdown": "Thanks for sharing. it looks very interesting. look forward to your blog/paper. ",
      "votes": 1
    },
    {
      "id": 436275,
      "postDate": "2018-12-10T03:37:33.607Z",
      "content": "<p>Wow, you're a breath of a fresh air in Kaggle !</p>",
      "rawMarkdown": "Wow, you're a breath of a fresh air in Kaggle !",
      "votes": 2
    },
    {
      "id": 439370,
      "postDate": "2018-12-15T09:10:29.723Z",
      "content": "<p>amzing</p>",
      "rawMarkdown": "amzing"
    }
  ],
  "comments": [
    {
      "id": 436252,
      "author_name": "Aleksey Nozdryn-Plotnicki",
      "author_url": "",
      "post_date": "2018-12-10T01:39:01.940000",
      "content": "<p>I just put code up here: <a href=\"https://github.com/alekseynp/kaggle-quickdraw\">https://github.com/alekseynp/kaggle-quickdraw</a></p>",
      "votes": 5,
      "replies": [
        {
          "id": 471437,
          "author_name": "Dai Dao",
          "author_url": "",
          "post_date": "2019-02-14T13:10:46.753000",
          "content": "<p>Thanks for sharing code :) </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 436654,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-12-10T17:52:41.997000",
      "content": "<p>for resampling to 256 points, did you use \"def resample_to(drawing, n)\" from</p>\n\n<p><a href=\"https://github.com/alekseynp/kaggle-quickdraw/blob/master/data_set.py\">https://github.com/alekseynp/kaggle-quickdraw/blob/master/data_set.py</a></p>\n\n<p>Thanks!</p>\n\n<p>it seems that making the drawing equal lengths also make the data more consistent and improve rnn results.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 436694,
          "author_name": "Aleksey Nozdryn-Plotnicki",
          "author_url": "",
          "post_date": "2018-12-10T18:54:32.913000",
          "content": "<p>Yes I did.\nI don't know if it creates as much consistency as I would like. Within any one drawing the segments are close to the same length, but not necessarily between drawings. I didn't study the raw data close enough to know. How consistent are lengths? By forcing \"long\" drawings and \"short\" drawings all to have 256 points, segment lengths may actually vary more after <code>resample_to</code></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 436400,
      "author_name": "kalili",
      "author_url": "",
      "post_date": "2018-12-10T08:29:59.297000",
      "content": "<p>Hi Aleksey, It seems that you wrote Private LB: 0.94804 as  Private LB: 0.94915 by mistake. I run your test.py getting the result of priv 0.94804 pub 0.94780. Otherwise your submission had a much better performance in priv than pub.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 436691,
          "author_name": "Aleksey Nozdryn-Plotnicki",
          "author_url": "",
          "post_date": "2018-12-10T18:49:53.693000",
          "content": "<p>Oops! Thanks</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 435289,
      "author_name": "Dmytro Danevskyi",
      "author_url": "",
      "post_date": "2018-12-07T19:59:09.463000",
      "content": "<p>Absolutely wonderful approach!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 435072,
      "author_name": "yimacs",
      "author_url": "",
      "post_date": "2018-12-07T13:04:13.670000",
      "content": "<p>Thanks for sharing. it looks very interesting. look forward to your blog/paper. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 436275,
      "author_name": "kalili",
      "author_url": "",
      "post_date": "2018-12-10T03:37:33.607000",
      "content": "<p>Wow, you're a breath of a fresh air in Kaggle !</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 439370,
      "author_name": "hentai123",
      "author_url": "",
      "post_date": "2018-12-15T09:10:29.723000",
      "content": "<p>amzing</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "434902": "Please pardon the quality of the following, but I’ve rushed it out. I will plan to release a detailed blog post/paper at a later date. My approach is novel, complex, and a challenge to communicate, but I wanted to get this information out while everyone was still interested!\n\n# Key Insight\nI leveraged a principal lesson in Deep Learning: Let the network learn the features rather than hand crafting them. The standard approach is to rasterize the drawings, crudely encode time as colour, and then pass to a pre-trained RGB ResNet. Instead, I implement a differentiable, trainable module to do this. Furthermore, I replace the standard nn.Conv2d(3, 64, kernel_size=7, stride=2, ...) and 2x2 MaxPool at the start of a ResNet with the same module.\n\nHigh Level Visual:\n![enter image description here][1]\n\nKey benefits:\n\n- Deep features are computed from the time component rather than some RGB hack\n\n- Deep features also computed from stroke data, leveraging the native format before “discarding” it and switching to images. Access to connectedness information.\n\n- Uses efficiency of convolutions on grids where a pure sequence or point cloud model fails to\n\n# Implementation\nEnormous amount of details that I will cover later.\n\nDetailed visual #1:\n![enter image description here][2]\n\n- Begin with strokes as defined by a series of points.\n\n- Difference the x’s and y’s to get dx, dy segment vectors. Average the ts.\n\n- Now we have strokes as defined by a series of segments.\n\n- Process strokes with a sequence module to generate 32 features.\n\n- Unroll those features with a window 2 convolution to generate 64 features per original point.\n\nDetailed visual #2:\n![enter image description here][3]\n\n- Draw points into a 32x32x64 image as per the diagram.\n\n- When points collide their feature vectors are averaged. I think I would have max-pooled if I could have implemented it.\n\n- At this stage in the network a 32x32 image could be thought of as equivalent to having started a normal image ResNet at 256x256.\n\n# Sequence Module\n\n    Conv1d(3,  32, kernel_size=3, stride=1, padding=1, dilation=1)\n    BatchNorm1d(32)\n    ReLU(inplace=True)\n    Conv1d(32, 32, kernel_size=3, stride=1, padding=2, dilation=2)\n    BatchNorm1d(32)\n    ReLU(inplace=True)\n    Conv1d(32, 32, kernel_size=3, stride=1, padding=4, dilation=4)\n    BatchNorm1d(32)\n    ReLU(inplace=True)\n    Conv1d(32, 32, kernel_size=3, stride=1, padding=8, dilation=8)\n    BatchNorm1d(32)\n    ReLU(inplace=True)\n    Conv1d(32, 64, kernel_size=2, stride=1, padding=(1,0))\n\n\n# Rasterization Module\n\n    from apex import amp\n    import torch\n    from torch.autograd import Function\n    \n    class PointsToImage(Function):\n       @staticmethod\n       @amp.float_function\n       def forward(ctx, i, v):\n           device = i.device\n           batch_size, _, num_input_points = i.size()\n           feature_size = v.size()[2]\n    \n           batch_idx = torch.arange(batch_size, device=device).view(-1, 1).repeat(1, num_input_points).view(-1)\n           idx_full = torch.cat([batch_idx.unsqueeze(0), i.permute(1, 0, 2).contiguous().view(2, -1)], dim=0)\n    \n           v_full = v.contiguous().view(batch_size * num_input_points, feature_size)\n           mat_sparse = torch.cuda.sparse.FloatTensor(idx_full, v_full)\n           mat_dense = mat_sparse.to_dense()\n    \n           ones_full = torch.ones(v_full.size(), device=device)\n           mat_sparse_count = torch.sparse.FloatTensor(idx_full, ones_full)\n           mat_dense_count = mat_sparse_count.to_dense()\n    \n           ctx.save_for_backward(idx_full, mat_dense_count)\n    \n           return mat_dense / torch.clamp(mat_dense_count, 1, 1e4)\n    \n       @staticmethod\n       @amp.float_function\n       def backward(ctx, grad_output):\n           idx_full, mat_dense_count = ctx.saved_tensors\n           grad_i = grad_v = None\n    \n           batch_size, _, _, feature_size = grad_output.size()\n    \n           if ctx.needs_input_grad[0]:\n               raise Exception(\"Indices aren't differentiable.\")\n           if ctx.needs_input_grad[1]:\n               grad = grad_output[idx_full[0], idx_full[1], idx_full[2]]\n               coef = mat_dense_count[idx_full[0], idx_full[1], idx_full[2]]\n               grad_v = grad / coef\n               grad_v = grad_v.view(batch_size, -1, feature_size)\n    \n           if isinstance(grad_output, torch.cuda.FloatTensor):\n               return grad_i, grad_v\n           else:\n               return grad_i, grad_v.half()\n    \n    points_to_image = PointsToImage.apply\n\n\nVarious other details:\n\n- Packing strokes of varying length into tensors of fixed size in order to do 1D CNNs is a non-trivial thing to do and beyond the scope of this post\n\n- Pytorch 0.4.1\n\n- Used NVIDIA’s apex amp (https://github.com/NVIDIA/apex/tree/master/apex/amp) to train exclusively at half precision\n\n- Trained on all the data\n\n- Raw not simplified\n\n- LMDB for memory mapped data\n\n- Adam optimizer\n\n- Models typically took about 2.5-3 days to converge on a system with a 1080Ti and a Titan\n\n- Used pre-trained imagenet weights\n\n- Froze those weights and only trained my additional modules for the first 1k-5k iterations\n\n- Held out 50k examples for validation during training\n\n- Held out 1 million examples for blending\n\n- Probably never completed even 2 complete passes through all the data during training. Convergence came first.\n\n- Used gradient accumulation via multiple backwards calls in Pytorch to finish training at huge batch sizes\n\n# Results\n\nSEResNeXt50 32x4d at core.\nLocal validation:\n\n- Acc@1: 0.8648\n\n- Mapk3: 0.9073\n\n- CE Loss: 0.5092\n\n- Public LB: 0.94781\n\n- Private LB: 0.94915\n\nBest ensemble:\n\nSix best models as measured by local CE Loss.\n\n- 2 x SEResNeXt50 at core\n\n- 3 x SEResNeXt101 at core\n\n- 1 x ResNet34 at core\n\n- Weighted arithmetic mean of probabilities. Weights = (1/loss)**24\n\nLocal validation:\n\n- Acc#1: 0.8683\n\n- Mapk3: 0.9101\n\n- CE Loss: 0.4937\n\n- Public LB: 0.95142\n\n- Private LB: 0.95101\n\nVery slight optimization bias in local validation score. I used my validation set to select “six best” and the 24 exponent.\nZero overfit to public LB. hence why I moved up from 12 to 8 in the shakeup.\n\nOther comments:\n\n- No RNNs in my ensemble\n\n- No Image CNNs in my ensemble\n\n- Kaggle competitions are a hell of an environment to try to do something novel in. I spent weeks messing around with Deep Residual PointNet++ networks, but never surpassed public LB 0.916\n\nFinally:\n\n- GG to everyone\n\n- Super proud of Guanshuo Xu for solid solo performance and disciplined execution enabling a move from 4th to 2nd into the money in the Private LB shakeup!\n\n  [1]: http://alekseynp.com/images/quickdraw_diagram1.png\n  [2]: http://alekseynp.com/images/quickdraw_diagram2.png\n  [3]: http://alekseynp.com/images/quickdraw_diagram3.png",
    "436252": "I just put code up here: https://github.com/alekseynp/kaggle-quickdraw",
    "436654": "for resampling to 256 points, did you use \"def resample_to(drawing, n)\" from\n\nhttps://github.com/alekseynp/kaggle-quickdraw/blob/master/data_set.py\n\nThanks!\n\n\nit seems that making the drawing equal lengths also make the data more consistent and improve rnn results.",
    "436400": "Hi Aleksey, It seems that you wrote Private LB: 0.94804 as  Private LB: 0.94915 by mistake. I run your test.py getting the result of priv 0.94804 pub 0.94780. Otherwise your submission had a much better performance in priv than pub.",
    "435289": "Absolutely wonderful approach!",
    "435072": "Thanks for sharing. it looks very interesting. look forward to your blog/paper. ",
    "436275": "Wow, you're a breath of a fresh air in Kaggle !",
    "439370": "amzing"
  }
}