{
  "id": 70912,
  "title": "tricks for getting LB 0.945 and above",
  "url": "/competitions/quickdraw-doodle-recognition/discussion/70912",
  "author_name": "hengck23",
  "post_date": "2018-11-08T10:12:40.958000",
  "votes": 30,
  "comment_count": 20,
  "views": 0,
  "content": "<p>see attached:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/417458/10636/limit.png\" alt=\"enter image description here\"></p>",
  "messages": [
    {
      "id": 417458,
      "postDate": "2018-11-08T10:12:40.960Z",
      "content": "<p>see attached:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/417458/10636/limit.png\" alt=\"enter image description here\"></p>",
      "rawMarkdown": "see attached:\n\n\n   ![enter image description here][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/417458/10636/limit.png",
      "votes": 30
    },
    {
      "id": 417462,
      "postDate": "2018-11-08T10:17:41.420Z",
      "content": "<p>lesson learn:</p>\n\n<ol>\n<li><p>if you want to win kaggle, you have understand why your model does not perform well.</p></li>\n<li><p>for this challenge, there is considerable label noise. This is the \"real problem\" in this challenge. There is about 10% unrecognized label, and that is why my validation top1 is 85%. But about 50% of  unrecognized label are correct, giving my local LB (map@3) 90%.</p></li>\n<li><p>some of the unrecognized label are very wrong, and it will affect accuracy (e.g. the mislabeled \"rain\" above). We have to think of a way to learn against label noise, or clean them up by hand.</p></li>\n</ol>",
      "rawMarkdown": "lesson learn:\n\n1. if you want to win kaggle, you have understand why your model does not perform well.\n\n2. for this challenge, there is considerable label noise. This is the \"real problem\" in this challenge. There is about 10% unrecognized label, and that is why my validation top1 is 85%. But about 50% of  unrecognized label are correct, giving my local LB (map@3) 90%.\n\n3. some of the unrecognized label are very wrong, and it will affect accuracy (e.g. the mislabeled \"rain\" above). We have to think of a way to learn against label noise, or clean them up by hand.\n\n",
      "votes": 8,
      "replies": [
        {
          "id": 419461,
          "postDate": "2018-11-12T02:51:31.340Z",
          "content": "<p>Hi \nI trained the model using recognized data only. There are also gaps between scores of validation and submitted results(about 4%).</p>",
          "rawMarkdown": "Hi \nI trained the model using recognized data only. There are also gaps between scores of validation and submitted results(about 4%).",
          "votes": 1
        },
        {
          "id": 420641,
          "postDate": "2018-11-13T23:39:34.667Z",
          "content": "<p>@Heng CherKeng \nFor point 1, how to understand why my architecture fails or which element of the whole process fails (hyperparameter, architecture of model,train/test size split and so on)? Currently, I do not have any overfitting/underfitting problem, just want to improve the accuracy, so I tried various models, but all cannot beat ResNet50, can you share with us how to \"understand\" the model?</p>",
          "rawMarkdown": "@Heng CherKeng \nFor point 1, how to understand why my architecture fails or which element of the whole process fails (hyperparameter, architecture of model,train/test size split and so on)? Currently, I do not have any overfitting/underfitting problem, just want to improve the accuracy, so I tried various models, but all cannot beat ResNet50, can you share with us how to \"understand\" the model?"
        }
      ]
    },
    {
      "id": 424176,
      "postDate": "2018-11-19T17:50:07.940Z",
      "content": "<p>time encode CNN</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/424176/10692/time_encoded_cnn.png\" alt=\"enter image description here\"></p>",
      "rawMarkdown": "time encode CNN\n\n  ![enter image description here][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/424176/10692/time_encoded_cnn.png",
      "votes": 6
    },
    {
      "id": 424423,
      "postDate": "2018-11-20T05:20:12.013Z",
      "content": "<p>code to transfer timestamp from raw to simplified:</p>\n\n<p>df_r = df of raw csv</p>\n\n<p>df_r = df of simplified  csv</p>\n\n<p>for each stroke, the end points were make to have the same timestamp. then, the timestamp of in-between points of the simplified drawings were  then interpolated using distance of the point from start of stroke.</p>\n\n<pre><code>def do_one(param):\n\ndf_r, df_s, n = param\nassert(df_r.key_id[n] == df_s.key_id[n])\nprint('\\r\\t%d  %s'%(n, df_r.key_id[n]), end ='', flush=True)\n\ndrawing_s = eval(df_s.drawing[n])\ndrawing_r = eval(df_r.drawing[n])\nassert(len(drawing_s)==len(drawing_r))\n\ndrawing=[]\nfor i, (d_s,d_r) in enumerate(zip(drawing_s,drawing_r)):\n    x_s,y_s     = np.array(d_s)\n    x_r,y_r,t_r = np.array(d_r)\n\n    N_s = len(x_s)\n    N_r = len(x_r)\n    if (N_s&amp;gt;1) and (N_r&amp;gt;1):\n        d_s = ((x_s[1:]-x_s[:-1])**2 + (y_s[1:]-y_s[:-1])**2 )**0.5\n        d_s = np.insert(d_s, 0, 0)\n        distance = d_s.sum()+ EPS\n        t_s = (d_s.cumsum()/distance)*(t_r[-1]-t_r[0]) + t_r[0]\n\n    elif (N_s==1) and (N_r==1):\n        t_s = t_r\n    elif (N_s==1) and (N_r&amp;gt;1):\n        t_s = t_r[[0]]\n    elif (N_s&amp;gt;1) and (N_r==1):\n        t_s = t_r[[0]*N_s]\n\n    t_s = list(t_s.astype(np.int32))\n    x_s = list(x_s)\n    y_s = list(y_s)\n    drawing.append([x_s,y_s,t_s])\ndrawing= str(drawing)\nreturn drawing\n</code></pre>",
      "rawMarkdown": "code to transfer timestamp from raw to simplified:\n\ndf_r = df of raw csv\n\ndf_r = df of simplified  csv\n\nfor each stroke, the end points were make to have the same timestamp. then, the timestamp of in-between points of the simplified drawings were  then interpolated using distance of the point from start of stroke.\n\n\n    def do_one(param):\n\n    df_r, df_s, n = param\n    assert(df_r.key_id[n] == df_s.key_id[n])\n    print('\\r\\t%d  %s'%(n, df_r.key_id[n]), end ='', flush=True)\n\n    drawing_s = eval(df_s.drawing[n])\n    drawing_r = eval(df_r.drawing[n])\n    assert(len(drawing_s)==len(drawing_r))\n\n    drawing=[]\n    for i, (d_s,d_r) in enumerate(zip(drawing_s,drawing_r)):\n        x_s,y_s     = np.array(d_s)\n        x_r,y_r,t_r = np.array(d_r)\n\n        N_s = len(x_s)\n        N_r = len(x_r)\n        if (N_s&gt;1) and (N_r&gt;1):\n            d_s = ((x_s[1:]-x_s[:-1])**2 + (y_s[1:]-y_s[:-1])**2 )**0.5\n            d_s = np.insert(d_s, 0, 0)\n            distance = d_s.sum()+ EPS\n            t_s = (d_s.cumsum()/distance)*(t_r[-1]-t_r[0]) + t_r[0]\n\n        elif (N_s==1) and (N_r==1):\n            t_s = t_r\n        elif (N_s==1) and (N_r&gt;1):\n            t_s = t_r[[0]]\n        elif (N_s&gt;1) and (N_r==1):\n            t_s = t_r[[0]*N_s]\n\n        t_s = list(t_s.astype(np.int32))\n        x_s = list(x_s)\n        y_s = list(y_s)\n        drawing.append([x_s,y_s,t_s])\n    drawing= str(drawing)\n    return drawing\n",
      "votes": 4
    },
    {
      "id": 423544,
      "postDate": "2018-11-18T14:12:11.367Z",
      "content": "<p>here is the magic. large batch size is important!\nIf you can backprop end to end, results will even be better. Else due to memory constraint, freezing/partial-freezing cnn and rnn would be the next best solution.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/423544/10689/magic.png\" alt=\"enter image description here\"></p>",
      "rawMarkdown": "here is the magic. large batch size is important!\nIf you can backprop end to end, results will even be better. Else due to memory constraint, freezing/partial-freezing cnn and rnn would be the next best solution.\n\n\n  ![enter image description here][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/423544/10689/magic.png",
      "votes": 4
    },
    {
      "id": 417475,
      "postDate": "2018-11-08T10:44:57.383Z",
      "content": "<p>you can google for papers with the keyword:\n'deep learning', noisy label, etc ... there are many works that deal with this.</p>\n\n<hr>\n\n<p>on a side note:</p>\n\n<p>\"generalized orderless pooling performs implicit salient matching\"</p>\n\n<p><a href=\"https://arxiv.org/pdf/1705.00487.pdf\">https://arxiv.org/pdf/1705.00487.pdf</a></p>\n\n<p>this paper shows you which train images affect the decision of a test image:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/417475/10637/vnn.png\" alt=\"enter image description here\"></p>",
      "rawMarkdown": "you can google for papers with the keyword:\n'deep learning', noisy label, etc ... there are many works that deal with this.\n\n---\n\non a side note:\n\n\"generalized orderless pooling performs implicit salient matching\"\n\nhttps://arxiv.org/pdf/1705.00487.pdf\n\n\nthis paper shows you which train images affect the decision of a test image:\n\n\n  ![enter image description here][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/417475/10637/vnn.png",
      "votes": 1
    },
    {
      "id": 417470,
      "postDate": "2018-11-08T10:39:12.917Z",
      "content": "<p>these are probably enough to win:\n1.  good CNN\n2. LSTM/RNN (the order of stroke affect results, e.g. drawing bat before/after baseball) or ways to use timestamp in raw data\n3. handling noisy label\n4. handing multi-output /top-k friendly loss,etc</p>",
      "rawMarkdown": "these are probably enough to win:\n1.  good CNN\n2. LSTM/RNN (the order of stroke affect results, e.g. drawing bat before/after baseball) or ways to use timestamp in raw data\n3. handling noisy label\n4. handing multi-output /top-k friendly loss,etc",
      "replies": [
        {
          "id": 417681,
          "postDate": "2018-11-08T16:28:05.060Z",
          "content": "<p>what should be the score of CNN to be considered a good CNN?</p>",
          "rawMarkdown": "what should be the score of CNN to be considered a good CNN?"
        },
        {
          "id": 421030,
          "postDate": "2018-11-14T13:52:32.520Z",
          "content": "<p>Hi~Heng For your point 3.Should we delete the unrecognized image?</p>",
          "rawMarkdown": "Hi~Heng For your point 3.Should we delete the unrecognized image?"
        },
        {
          "id": 421034,
          "postDate": "2018-11-14T13:57:32.493Z",
          "content": "<p>remove those unrecognized image that you are sure are wrong.\nkeep those that are correct.</p>\n\n<p>about 50% of unrecognized image are correct.</p>\n\n<p>if you are unsure, you can leave them there.</p>\n\n<p>deep network can tolerate some noise</p>",
          "rawMarkdown": "remove those unrecognized image that you are sure are wrong.\nkeep those that are correct.\n\nabout 50% of unrecognized image are correct.\n \nif you are unsure, you can leave them there.\n\ndeep network can tolerate some noise",
          "votes": 1
        },
        {
          "id": 421057,
          "postDate": "2018-11-14T14:20:34.780Z",
          "content": "<p>It is not easy.</p>",
          "rawMarkdown": "It is not easy."
        },
        {
          "id": 421186,
          "postDate": "2018-11-14T17:25:42.133Z",
          "content": "<p>you can get LB 0.945 just by using CNN and ensemble. you don;t have to handle noise, etc</p>\n\n<p>But i suspect the final top LB score should be in the range of 0.960. This will require handling of noise and LSTM approach, etc</p>",
          "rawMarkdown": "you can get LB 0.945 just by using CNN and ensemble. you don;t have to handle noise, etc\n\nBut i suspect the final top LB score should be in the range of 0.960. This will require handling of noise and LSTM approach, etc",
          "votes": 1
        },
        {
          "id": 421812,
          "postDate": "2018-11-15T12:57:55.017Z",
          "content": "<p>Did you dispose of samples that are identified as noise by hand?</p>",
          "rawMarkdown": "Did you dispose of samples that are identified as noise by hand?"
        },
        {
          "id": 426133,
          "postDate": "2018-11-22T17:38:37.037Z",
          "content": "<p>Heng CherKeng. How use timestamp if for test we don't have it ?</p>",
          "rawMarkdown": "Heng CherKeng. How use timestamp if for test we don't have it ?"
        },
        {
          "id": 426141,
          "postDate": "2018-11-22T17:57:51.450Z",
          "content": "<p>You can get it from the raw</p>",
          "rawMarkdown": "You can get it from the raw"
        },
        {
          "id": 426143,
          "postDate": "2018-11-22T18:12:35.960Z",
          "content": "<p>Time of creation? How?</p>",
          "rawMarkdown": "Time of creation? How?"
        }
      ]
    },
    {
      "id": 423804,
      "postDate": "2018-11-19T04:52:45.880Z",
      "rawMarkdown": "",
      "votes": 7,
      "isDeleted": true
    },
    {
      "id": 426070,
      "postDate": "2018-11-22T14:52:22.593Z",
      "content": "<p>Thanks a lot.</p>",
      "rawMarkdown": "Thanks a lot."
    },
    {
      "id": 420468,
      "postDate": "2018-11-13T17:07:06.013Z",
      "content": "<p>thanks for your sharing</p>",
      "rawMarkdown": "thanks for your sharing"
    }
  ],
  "comments": [
    {
      "id": 417462,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-11-08T10:17:41.420000",
      "content": "<p>lesson learn:</p>\n\n<ol>\n<li><p>if you want to win kaggle, you have understand why your model does not perform well.</p></li>\n<li><p>for this challenge, there is considerable label noise. This is the \"real problem\" in this challenge. There is about 10% unrecognized label, and that is why my validation top1 is 85%. But about 50% of  unrecognized label are correct, giving my local LB (map@3) 90%.</p></li>\n<li><p>some of the unrecognized label are very wrong, and it will affect accuracy (e.g. the mislabeled \"rain\" above). We have to think of a way to learn against label noise, or clean them up by hand.</p></li>\n</ol>",
      "votes": 8,
      "replies": [
        {
          "id": 419461,
          "author_name": "June",
          "author_url": "",
          "post_date": "2018-11-12T02:51:31.340000",
          "content": "<p>Hi \nI trained the model using recognized data only. There are also gaps between scores of validation and submitted results(about 4%).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 420641,
          "author_name": "Joe Ho",
          "author_url": "",
          "post_date": "2018-11-13T23:39:34.667000",
          "content": "<p>@Heng CherKeng \nFor point 1, how to understand why my architecture fails or which element of the whole process fails (hyperparameter, architecture of model,train/test size split and so on)? Currently, I do not have any overfitting/underfitting problem, just want to improve the accuracy, so I tried various models, but all cannot beat ResNet50, can you share with us how to \"understand\" the model?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 424176,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-11-19T17:50:07.940000",
      "content": "<p>time encode CNN</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/424176/10692/time_encoded_cnn.png\" alt=\"enter image description here\"></p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 424423,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-11-20T05:20:12.013000",
      "content": "<p>code to transfer timestamp from raw to simplified:</p>\n\n<p>df_r = df of raw csv</p>\n\n<p>df_r = df of simplified  csv</p>\n\n<p>for each stroke, the end points were make to have the same timestamp. then, the timestamp of in-between points of the simplified drawings were  then interpolated using distance of the point from start of stroke.</p>\n\n<pre><code>def do_one(param):\n\ndf_r, df_s, n = param\nassert(df_r.key_id[n] == df_s.key_id[n])\nprint('\\r\\t%d  %s'%(n, df_r.key_id[n]), end ='', flush=True)\n\ndrawing_s = eval(df_s.drawing[n])\ndrawing_r = eval(df_r.drawing[n])\nassert(len(drawing_s)==len(drawing_r))\n\ndrawing=[]\nfor i, (d_s,d_r) in enumerate(zip(drawing_s,drawing_r)):\n    x_s,y_s     = np.array(d_s)\n    x_r,y_r,t_r = np.array(d_r)\n\n    N_s = len(x_s)\n    N_r = len(x_r)\n    if (N_s&amp;gt;1) and (N_r&amp;gt;1):\n        d_s = ((x_s[1:]-x_s[:-1])**2 + (y_s[1:]-y_s[:-1])**2 )**0.5\n        d_s = np.insert(d_s, 0, 0)\n        distance = d_s.sum()+ EPS\n        t_s = (d_s.cumsum()/distance)*(t_r[-1]-t_r[0]) + t_r[0]\n\n    elif (N_s==1) and (N_r==1):\n        t_s = t_r\n    elif (N_s==1) and (N_r&amp;gt;1):\n        t_s = t_r[[0]]\n    elif (N_s&amp;gt;1) and (N_r==1):\n        t_s = t_r[[0]*N_s]\n\n    t_s = list(t_s.astype(np.int32))\n    x_s = list(x_s)\n    y_s = list(y_s)\n    drawing.append([x_s,y_s,t_s])\ndrawing= str(drawing)\nreturn drawing\n</code></pre>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 423544,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-11-18T14:12:11.367000",
      "content": "<p>here is the magic. large batch size is important!\nIf you can backprop end to end, results will even be better. Else due to memory constraint, freezing/partial-freezing cnn and rnn would be the next best solution.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/423544/10689/magic.png\" alt=\"enter image description here\"></p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 417475,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-11-08T10:44:57.383000",
      "content": "<p>you can google for papers with the keyword:\n'deep learning', noisy label, etc ... there are many works that deal with this.</p>\n\n<hr>\n\n<p>on a side note:</p>\n\n<p>\"generalized orderless pooling performs implicit salient matching\"</p>\n\n<p><a href=\"https://arxiv.org/pdf/1705.00487.pdf\">https://arxiv.org/pdf/1705.00487.pdf</a></p>\n\n<p>this paper shows you which train images affect the decision of a test image:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/417475/10637/vnn.png\" alt=\"enter image description here\"></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 417470,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-11-08T10:39:12.917000",
      "content": "<p>these are probably enough to win:\n1.  good CNN\n2. LSTM/RNN (the order of stroke affect results, e.g. drawing bat before/after baseball) or ways to use timestamp in raw data\n3. handling noisy label\n4. handing multi-output /top-k friendly loss,etc</p>",
      "votes": 0,
      "replies": [
        {
          "id": 417681,
          "author_name": "Ankit Sati",
          "author_url": "",
          "post_date": "2018-11-08T16:28:05.060000",
          "content": "<p>what should be the score of CNN to be considered a good CNN?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 421030,
          "author_name": "TPloveYXT520",
          "author_url": "",
          "post_date": "2018-11-14T13:52:32.520000",
          "content": "<p>Hi~Heng For your point 3.Should we delete the unrecognized image?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 421034,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2018-11-14T13:57:32.493000",
          "content": "<p>remove those unrecognized image that you are sure are wrong.\nkeep those that are correct.</p>\n\n<p>about 50% of unrecognized image are correct.</p>\n\n<p>if you are unsure, you can leave them there.</p>\n\n<p>deep network can tolerate some noise</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 421057,
          "author_name": "Gary",
          "author_url": "",
          "post_date": "2018-11-14T14:20:34.780000",
          "content": "<p>It is not easy.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 421186,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2018-11-14T17:25:42.133000",
          "content": "<p>you can get LB 0.945 just by using CNN and ensemble. you don;t have to handle noise, etc</p>\n\n<p>But i suspect the final top LB score should be in the range of 0.960. This will require handling of noise and LSTM approach, etc</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 421812,
          "author_name": "Gary",
          "author_url": "",
          "post_date": "2018-11-15T12:57:55.017000",
          "content": "<p>Did you dispose of samples that are identified as noise by hand?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 426133,
          "author_name": "Leigh",
          "author_url": "",
          "post_date": "2018-11-22T17:38:37.037000",
          "content": "<p>Heng CherKeng. How use timestamp if for test we don't have it ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 426141,
          "author_name": "Strideradu",
          "author_url": "",
          "post_date": "2018-11-22T17:57:51.450000",
          "content": "<p>You can get it from the raw</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 426143,
          "author_name": "Leigh",
          "author_url": "",
          "post_date": "2018-11-22T18:12:35.960000",
          "content": "<p>Time of creation? How?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 423804,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-19T04:52:45.880000",
      "content": "",
      "votes": 7,
      "replies": []
    },
    {
      "id": 426070,
      "author_name": "Tim Wu",
      "author_url": "",
      "post_date": "2018-11-22T14:52:22.593000",
      "content": "<p>Thanks a lot.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 420468,
      "author_name": "Salaryman",
      "author_url": "",
      "post_date": "2018-11-13T17:07:06.013000",
      "content": "<p>thanks for your sharing</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "417458": "see attached:\n\n\n   ![enter image description here][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/417458/10636/limit.png",
    "417462": "lesson learn:\n\n1. if you want to win kaggle, you have understand why your model does not perform well.\n\n2. for this challenge, there is considerable label noise. This is the \"real problem\" in this challenge. There is about 10% unrecognized label, and that is why my validation top1 is 85%. But about 50% of  unrecognized label are correct, giving my local LB (map@3) 90%.\n\n3. some of the unrecognized label are very wrong, and it will affect accuracy (e.g. the mislabeled \"rain\" above). We have to think of a way to learn against label noise, or clean them up by hand.\n\n",
    "424176": "time encode CNN\n\n  ![enter image description here][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/424176/10692/time_encoded_cnn.png",
    "424423": "code to transfer timestamp from raw to simplified:\n\ndf_r = df of raw csv\n\ndf_r = df of simplified  csv\n\nfor each stroke, the end points were make to have the same timestamp. then, the timestamp of in-between points of the simplified drawings were  then interpolated using distance of the point from start of stroke.\n\n\n    def do_one(param):\n\n    df_r, df_s, n = param\n    assert(df_r.key_id[n] == df_s.key_id[n])\n    print('\\r\\t%d  %s'%(n, df_r.key_id[n]), end ='', flush=True)\n\n    drawing_s = eval(df_s.drawing[n])\n    drawing_r = eval(df_r.drawing[n])\n    assert(len(drawing_s)==len(drawing_r))\n\n    drawing=[]\n    for i, (d_s,d_r) in enumerate(zip(drawing_s,drawing_r)):\n        x_s,y_s     = np.array(d_s)\n        x_r,y_r,t_r = np.array(d_r)\n\n        N_s = len(x_s)\n        N_r = len(x_r)\n        if (N_s&gt;1) and (N_r&gt;1):\n            d_s = ((x_s[1:]-x_s[:-1])**2 + (y_s[1:]-y_s[:-1])**2 )**0.5\n            d_s = np.insert(d_s, 0, 0)\n            distance = d_s.sum()+ EPS\n            t_s = (d_s.cumsum()/distance)*(t_r[-1]-t_r[0]) + t_r[0]\n\n        elif (N_s==1) and (N_r==1):\n            t_s = t_r\n        elif (N_s==1) and (N_r&gt;1):\n            t_s = t_r[[0]]\n        elif (N_s&gt;1) and (N_r==1):\n            t_s = t_r[[0]*N_s]\n\n        t_s = list(t_s.astype(np.int32))\n        x_s = list(x_s)\n        y_s = list(y_s)\n        drawing.append([x_s,y_s,t_s])\n    drawing= str(drawing)\n    return drawing\n",
    "423544": "here is the magic. large batch size is important!\nIf you can backprop end to end, results will even be better. Else due to memory constraint, freezing/partial-freezing cnn and rnn would be the next best solution.\n\n\n  ![enter image description here][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/423544/10689/magic.png",
    "417475": "you can google for papers with the keyword:\n'deep learning', noisy label, etc ... there are many works that deal with this.\n\n---\n\non a side note:\n\n\"generalized orderless pooling performs implicit salient matching\"\n\nhttps://arxiv.org/pdf/1705.00487.pdf\n\n\nthis paper shows you which train images affect the decision of a test image:\n\n\n  ![enter image description here][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/417475/10637/vnn.png",
    "417470": "these are probably enough to win:\n1.  good CNN\n2. LSTM/RNN (the order of stroke affect results, e.g. drawing bat before/after baseball) or ways to use timestamp in raw data\n3. handling noisy label\n4. handing multi-output /top-k friendly loss,etc",
    "423804": "",
    "426070": "Thanks a lot.",
    "420468": "thanks for your sharing"
  }
}