{
  "id": 68006,
  "title": "new/interesting ideas to try",
  "url": "/competitions/quickdraw-doodle-recognition/discussion/68006",
  "author_name": "hengck23",
  "post_date": "2018-10-08T12:03:43.828000",
  "votes": 35,
  "comment_count": 33,
  "views": 0,
  "content": "<p>some ideas that i would like to try</p>\n\n<p><a href=\"https://arxiv.org/abs/1802.07595\">https://arxiv.org/abs/1802.07595</a></p>\n\n<p>Smooth Loss Functions for Deep Top-k Classification\n - Leonard Berrada, Andrew Zisserman, M. Pawan Kumar</p>\n\n<hr>\n\n<p><a href=\"https://arxiv.org/abs/1806.05594\">https://arxiv.org/abs/1806.05594</a></p>\n\n<p>Improving Consistency-Based Semi-Supervised Learning with Weight Averaging\n - Ben Athiwaratkun, Marc Finzi, Pavel Izmailov, Andrew Gordon Wilson</p>\n\n<hr>\n\n<p>to be udpated ...</p>\n\n<p>quickdraw GAN,  Reinforcement Learning Deep-Q Nets ....</p>\n\n<p>Beyond rnn,eg attention, temporal cnn</p>\n\n<p>deformable cnn, non-local mean, graph-cnn</p>",
  "messages": [
    {
      "id": 400492,
      "postDate": "2018-10-08T12:03:43.830Z",
      "content": "<p>some ideas that i would like to try</p>\n\n<p><a href=\"https://arxiv.org/abs/1802.07595\">https://arxiv.org/abs/1802.07595</a></p>\n\n<p>Smooth Loss Functions for Deep Top-k Classification\n - Leonard Berrada, Andrew Zisserman, M. Pawan Kumar</p>\n\n<hr>\n\n<p><a href=\"https://arxiv.org/abs/1806.05594\">https://arxiv.org/abs/1806.05594</a></p>\n\n<p>Improving Consistency-Based Semi-Supervised Learning with Weight Averaging\n - Ben Athiwaratkun, Marc Finzi, Pavel Izmailov, Andrew Gordon Wilson</p>\n\n<hr>\n\n<p>to be udpated ...</p>\n\n<p>quickdraw GAN,  Reinforcement Learning Deep-Q Nets ....</p>\n\n<p>Beyond rnn,eg attention, temporal cnn</p>\n\n<p>deformable cnn, non-local mean, graph-cnn</p>",
      "rawMarkdown": "some ideas that i would like to try\n\nhttps://arxiv.org/abs/1802.07595\n\nSmooth Loss Functions for Deep Top-k Classification\n - Leonard Berrada, Andrew Zisserman, M. Pawan Kumar\n\n---\nhttps://arxiv.org/abs/1806.05594\n\nImproving Consistency-Based Semi-Supervised Learning with Weight Averaging\n - Ben Athiwaratkun, Marc Finzi, Pavel Izmailov, Andrew Gordon Wilson\n\n\n---\n\nto be udpated ...\n\nquickdraw GAN,  Reinforcement Learning Deep-Q Nets ....\n\n\nBeyond rnn,eg attention, temporal cnn\n\ndeformable cnn, non-local mean, graph-cnn",
      "votes": 35
    },
    {
      "id": 410647,
      "postDate": "2018-10-26T11:57:38.250Z",
      "content": "<p>I have implemented the top-k loss described in <a href=\"http://openaccess.thecvf.com/content_cvpr_2016/papers/Lapin_Loss_Functions_for_CVPR_2016_paper.pdf\">http://openaccess.thecvf.com/content_cvpr_2016/papers/Lapin_Loss_Functions_for_CVPR_2016_paper.pdf</a>. And although my top3 accuracy went up from ~92% to ~95%, at the same time the top1 accuracy went down by a lot, making the MAP@3 go down significantly (from .91 to .61 or so). So don't bother optimizing for top3 accuracy. </p>",
      "rawMarkdown": "I have implemented the top-k loss described in http://openaccess.thecvf.com/content_cvpr_2016/papers/Lapin_Loss_Functions_for_CVPR_2016_paper.pdf. And although my top3 accuracy went up from ~92% to ~95%, at the same time the top1 accuracy went down by a lot, making the MAP@3 go down significantly (from .91 to .61 or so). So don't bother optimizing for top3 accuracy. ",
      "votes": 5,
      "replies": [
        {
          "id": 410661,
          "postDate": "2018-10-26T12:21:37.733Z",
          "content": "<p>did you implement the truncated top-k cross entropy or the smooth top-k Hinge ?  </p>",
          "rawMarkdown": "did you implement the truncated top-k cross entropy or the smooth top-k Hinge ?  "
        },
        {
          "id": 410714,
          "postDate": "2018-10-26T14:21:05.030Z",
          "content": "<p>the top-k cross entropy :)</p>",
          "rawMarkdown": "the top-k cross entropy :)"
        },
        {
          "id": 413520,
          "postDate": "2018-11-01T04:17:41.693Z",
          "content": "<p>kaggle loss is not exactly top-k. you have to modify a bit. e.g.</p>\n\n<p>kaggle_loss = top1+top2+top3</p>\n\n<p>also, if the different loss has different difficulties. the network may need to be larger.</p>",
          "rawMarkdown": "kaggle loss is not exactly top-k. you have to modify a bit. e.g.\n\nkaggle_loss = top1+top2+top3\n\nalso, if the different loss has different difficulties. the network may need to be larger."
        },
        {
          "id": 413676,
          "postDate": "2018-11-01T10:56:56.400Z",
          "content": "<p>Yeah, I figured. Also depends on to hat extent optimizing for top1 also co-optimizes the top3 and vice versa. Probably best to train a network for each and then train how to combine them after?</p>",
          "rawMarkdown": "Yeah, I figured. Also depends on to hat extent optimizing for top1 also co-optimizes the top3 and vice versa. Probably best to train a network for each and then train how to combine them after?"
        }
      ]
    },
    {
      "id": 401019,
      "postDate": "2018-10-09T09:29:59.370Z",
      "content": "<p>useful augmentation:</p>\n\n<ul>\n<li><p>stroke removal</p></li>\n<li><p>stroke simplification</p></li>\n<li><p>stroke deformation</p>\n\n<p><img src=\"https://media.springernature.com/original/springer-static/image/art%3A10.1007%2Fs11263-016-0932-3/MediaObjects/11263_2016_932_Fig2_HTML.gif\" alt=\"enter image description here\"></p>\n\n<p><img src=\"https://media.springernature.com/lw785/springer-static/image/art%3A10.1007%2Fs11263-016-0932-3/MediaObjects/11263_2016_932_Fig5_HTML.gif\" alt=\"enter image description here\"></p></li>\n</ul>\n\n<p><a href=\"http://sketchx.eecs.qmul.ac.uk/\">http://sketchx.eecs.qmul.ac.uk/</a></p>\n\n<p><a href=\"https://github.com/yuchuochuo1023/sketch-specific-data-augmentation\">https://github.com/yuchuochuo1023/sketch-specific-data-augmentation</a></p>\n\n<p><a href=\"https://www.eecs.qmul.ac.uk/~qian/Qian\">https://www.eecs.qmul.ac.uk/~qian/Qian</a>'s%20Materials/paper/IJCV_revised_version.pdf</p>",
      "rawMarkdown": "useful augmentation:\n\n-  stroke removal\n\n- stroke simplification\n\n- stroke deformation\n\n\n   ![enter image description here][1]\n\n\n    ![enter image description here][2]\n\n\n  [1]: https://media.springernature.com/original/springer-static/image/art%3A10.1007%2Fs11263-016-0932-3/MediaObjects/11263_2016_932_Fig2_HTML.gif\n  [2]: https://media.springernature.com/lw785/springer-static/image/art%3A10.1007%2Fs11263-016-0932-3/MediaObjects/11263_2016_932_Fig5_HTML.gif\n\n\nhttp://sketchx.eecs.qmul.ac.uk/\n\nhttps://github.com/yuchuochuo1023/sketch-specific-data-augmentation\n\nhttps://www.eecs.qmul.ac.uk/~qian/Qian's%20Materials/paper/IJCV_revised_version.pdf",
      "votes": 5,
      "replies": [
        {
          "id": 410380,
          "postDate": "2018-10-26T00:25:49.807Z",
          "content": "<p>Hi, Heng. The multi-scale fusion mentioned here, is each scale separately trained, or multiple scales simultaneously trained?</p>",
          "rawMarkdown": "Hi, Heng. The multi-scale fusion mentioned here, is each scale separately trained, or multiple scales simultaneously trained?"
        },
        {
          "id": 413670,
          "postDate": "2018-11-01T10:40:57.773Z",
          "content": "<p>I looked up this original paper, they used a far smaller dataset with only 80 sketches per category. I don't know augmentation is necessary for this amount of data?</p>",
          "rawMarkdown": "I looked up this original paper, they used a far smaller dataset with only 80 sketches per category. I don't know augmentation is necessary for this amount of data?"
        }
      ]
    },
    {
      "id": 400557,
      "postDate": "2018-10-08T14:43:01.540Z",
      "content": "<p><a href=\"http://openaccess.thecvf.com/content_cvpr_2018/CameraReady/2763.pdf\">http://openaccess.thecvf.com/content_cvpr_2018/CameraReady/2763.pdf</a></p>\n\n<p>SketchMate: Deep Hashing for Million-Scale Human Sketch Retrieval\n- Peng Xu</p>\n\n<ul>\n<li><p>joint stroke and image branch is interesting.</p></li>\n<li><p>with 2 separate models (stroke and image), distillation between them for semi-supervised learning is interesting</p>\n\n<p><img src=\"http://www.eecs.qmul.ac.uk/~kp306/Kaiyue%20Material/CVPR2018_HASH/thumbnail.png\" alt=\"enter image description here\"></p></li>\n</ul>",
      "rawMarkdown": "http://openaccess.thecvf.com/content_cvpr_2018/CameraReady/2763.pdf\n\n\nSketchMate: Deep Hashing for Million-Scale Human Sketch Retrieval\n- Peng Xu\n\n\n- joint stroke and image branch is interesting.\n\n- with 2 separate models (stroke and image), distillation between them for semi-supervised learning is interesting\n\n   ![enter image description here][1]\n\n\n  [1]: http://www.eecs.qmul.ac.uk/~kp306/Kaiyue%20Material/CVPR2018_HASH/thumbnail.png",
      "votes": 6,
      "replies": [
        {
          "id": 400575,
          "postDate": "2018-10-08T15:22:40.157Z",
          "content": "<p>oh Wow! Interesting idea!</p>",
          "rawMarkdown": "oh Wow! Interesting idea!"
        },
        {
          "id": 409871,
          "postDate": "2018-10-25T01:38:15.127Z",
          "content": "<p>Did you try this? I think CNN+RNN is a great idea.</p>",
          "rawMarkdown": "Did you try this? I think CNN+RNN is a great idea."
        },
        {
          "id": 410559,
          "postDate": "2018-10-26T09:00:27.130Z",
          "content": "<p>You could also use a third network that takes as inputs a sequence of images.</p>",
          "rawMarkdown": "You could also use a third network that takes as inputs a sequence of images.\n ",
          "votes": 1
        }
      ]
    },
    {
      "id": 402816,
      "postDate": "2018-10-12T11:49:45.693Z",
      "content": "<blockquote>\n  <p><a href=\"https://arxiv.org/abs/1802.07595\">https://arxiv.org/abs/1802.07595</a>\n  Smooth Loss Functions for Deep Top-k Classification - Leonard Berrada, Andrew Zisserman, M. Pawan Kumar</p>\n</blockquote>\n\n<p>There is a repository for this paper:\n<a href=\"https://github.com/oval-group/smooth-topk\">https://github.com/oval-group/smooth-topk</a></p>",
      "rawMarkdown": "&gt; https://arxiv.org/abs/1802.07595\n&gt; Smooth Loss Functions for Deep Top-k Classification - Leonard Berrada, Andrew Zisserman, M. Pawan Kumar\n\nThere is a repository for this paper:\nhttps://github.com/oval-group/smooth-topk",
      "votes": 3
    },
    {
      "id": 407763,
      "postDate": "2018-10-21T18:01:58.607Z",
      "content": "<p>RNN weakness?\nEspecially, rnn encode dx,dy and not the absolute x,y.</p>\n\n<p>It seems that the spatial information is not important?\nNote that i split the components of \"key\" and \"smiley face\" apart</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/inbox/113660/b5dba192c86735f135da305714a41bdc/rnn.png\" alt=\"enter image description here\"></p>",
      "rawMarkdown": "RNN weakness?\nEspecially, rnn encode dx,dy and not the absolute x,y.\n\n\nIt seems that the spatial information is not important?\nNote that i split the components of \"key\" and \"smiley face\" apart\n\n  ![enter image description here][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/inbox/113660/b5dba192c86735f135da305714a41bdc/rnn.png",
      "votes": 4,
      "replies": [
        {
          "id": 407979,
          "postDate": "2018-10-22T05:12:39.150Z",
          "content": "<p>The current Model behind this game is RNN? If so, seems CNN should be better for this game?</p>",
          "rawMarkdown": "The current Model behind this game is RNN? If so, seems CNN should be better for this game?",
          "votes": 1
        },
        {
          "id": 408764,
          "postDate": "2018-10-23T12:21:37.490Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 419691,
          "postDate": "2018-11-12T12:05:45.807Z",
          "content": "<p>I used dx, dy and it made no difference. </p>",
          "rawMarkdown": "I used dx, dy and it made no difference. "
        }
      ]
    },
    {
      "id": 420176,
      "postDate": "2018-11-13T07:48:31.223Z",
      "content": "<p><a href=\"https://arxiv.org/pdf/1708.02716.pdf\">https://arxiv.org/pdf/1708.02716.pdf</a></p>\n\n<p><a href=\"http://www.yugangjiang.info/publication/17MM-Sketch.pdf\">http://www.yugangjiang.info/publication/17MM-Sketch.pdf</a></p>\n\n<hr>\n\n<p><a href=\"https://ravika.github.io/publications.html\">https://ravika.github.io/publications.html</a></p>\n\n<p><a href=\"https://arxiv.org/pdf/1608.03369v1.pdf\">https://arxiv.org/pdf/1608.03369v1.pdf</a></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/420176/10664/cnnlstm1.png\" alt=\"enter image description here\"></p>\n\n<p>a pretty smart way to deal with incomplete sketch.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/420176/10665/cnnlstm2.png\" alt=\"enter image description here\">\n   <img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/420176/10666/cnnlstm3.png\" alt=\"enter image description here\"></p>",
      "rawMarkdown": "https://arxiv.org/pdf/1708.02716.pdf\n\nhttp://www.yugangjiang.info/publication/17MM-Sketch.pdf\n\n---\n\nhttps://ravika.github.io/publications.html\n\nhttps://arxiv.org/pdf/1608.03369v1.pdf\n\n\n   ![enter image description here][1]\n\na pretty smart way to deal with incomplete sketch.\n\n   ![enter image description here][2]\n   ![enter image description here][3]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/420176/10664/cnnlstm1.png\n  [2]: https://storage.googleapis.com/kaggle-forum-message-attachments/420176/10665/cnnlstm2.png\n  [3]: https://storage.googleapis.com/kaggle-forum-message-attachments/420176/10666/cnnlstm3.png",
      "votes": 1,
      "replies": [
        {
          "id": 420963,
          "postDate": "2018-11-14T12:02:38.237Z",
          "content": "<p>Thanks again, lot of useful info</p>",
          "rawMarkdown": "Thanks again, lot of useful info"
        }
      ]
    },
    {
      "id": 419430,
      "postDate": "2018-11-12T00:17:04.583Z",
      "content": "<p>\"Learning From Noisy Large-Scale Datasets With Minimal Supervision\" -Andreas Veit, cvpr 2018</p>\n\n<p><a href=\"https://www.youtube.com/watch?v=RAHiCtJhyDg\">https://www.youtube.com/watch?v=RAHiCtJhyDg</a></p>\n\n<p>inspired by this paper, one could do this:</p>\n\n<ol>\n<li><p>construct feature extractor.  Train both cross-entropy(image classification head) and sigmoid classifier (label cleaning head).</p></li>\n<li><p>cross entropy is trained on  all images. sigmoid is trained on recognised images only.</p></li>\n<li><p>loss can be like:  loss = (weight) * (cross entropy loss). weight = 1 if it is recognised. weight = sigmoid classifier probability if it is non-recognised </p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/419430/10657/architecture_1.jpg\" alt=\"enter image description here\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/419430/10662/label_cleaning1.png\" alt=\"enter image description here\"></p></li>\n</ol>",
      "rawMarkdown": "\"Learning From Noisy Large-Scale Datasets With Minimal Supervision\" -Andreas Veit, cvpr 2018\n\nhttps://www.youtube.com/watch?v=RAHiCtJhyDg\n\ninspired by this paper, one could do this:\n\n1. construct feature extractor.  Train both cross-entropy(image classification head) and sigmoid classifier (label cleaning head).\n\n2. cross entropy is trained on  all images. sigmoid is trained on recognised images only.\n\n\n3. loss can be like:  loss = (weight) * (cross entropy loss). weight = 1 if it is recognised. weight = sigmoid classifier probability if it is non-recognised \n\n\n  ![enter image description here][1]\n\n\n  ![enter image description here][2]\n\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/419430/10657/architecture_1.jpg\n  [2]: https://storage.googleapis.com/kaggle-forum-message-attachments/419430/10662/label_cleaning1.png",
      "votes": 1
    },
    {
      "id": 406544,
      "postDate": "2018-10-19T12:21:53.050Z",
      "content": "<p>Soon, new models pretrained on shapes:</p>\n\n<p><a href=\"https://openreview.net/pdf?id=Bygh9j09KX\">https://openreview.net/pdf?id=Bygh9j09KX</a></p>",
      "rawMarkdown": "Soon, new models pretrained on shapes:\n\nhttps://openreview.net/pdf?id=Bygh9j09KX\n",
      "votes": 1
    },
    {
      "id": 400934,
      "postDate": "2018-10-09T06:51:13.103Z",
      "content": "<p>Loss design:</p>\n\n<p><a href=\"https://liusi-group.com/pdf/ijcai_2018.pdf\">https://liusi-group.com/pdf/ijcai_2018.pdf</a></p>\n\n<p>'Ensemble Soft-Margin Softmax Loss for Image Classification' - Xiaobo Wang</p>\n\n<p>related:  cross entropy, center loss,A-softmax,Large-margin softmax ...</p>",
      "rawMarkdown": "Loss design:\n\nhttps://liusi-group.com/pdf/ijcai_2018.pdf\n\n'Ensemble Soft-Margin Softmax Loss for Image Classification' - Xiaobo Wang\n\n\nrelated:  cross entropy, center loss,A-softmax,Large-margin softmax ...",
      "votes": 1,
      "replies": [
        {
          "id": 409927,
          "postDate": "2018-10-25T04:44:13.293Z",
          "content": "<p>I tried l2-constrained loss, seem slightly worse than cross entropy loss. Also tried a simpler version AM-softmax loss, it's also slightly worse than cross entropy loss</p>\n\n<pre><code>AM-softmax loss\nfrom keras.models import Model\nfrom keras.layers import *\nimport keras.backend as K\nfrom keras.constraints import unit_norm\n\n\nx_in = Input(shape=(maxlen,))\nx_embedded = Embedding(len(chars)+2,\n                   word_size)(x_in)\nx = CuDNNGRU(word_size)(x_embedded)\nx = Lambda(lambda x: K.l2_normalize(x, 1))(x)\n\npred = Dense(num_train,\n         use_bias=False,\n         kernel_constraint=unit_norm())(x)\n\nencoder = Model(x_in, x) # 最终的目的是要得到一个编码器\nmodel = Model(x_in, pred) # 用分类问题做训练\n\ndef amsoftmax_loss(y_true, y_pred, scale=30, margin=0.35):\n y_pred = y_true * (y_pred - margin) + (1 - y_true) * y_pred\n y_pred *= scale\n return K.categorical_crossentropy(y_true, y_pred, from_logits=True)\n\nmodel.compile(loss=amsoftmax_loss,\n          optimizer='adam',\n          metrics=['accuracy'])\n</code></pre>\n\n<p>Reference: <a href=\"https://kexue.fm/archives/5743#%E5%9F%BA%E6%9C%AC%E5%AE%9E%E7%8E%B0\">https://kexue.fm/archives/5743#%E5%9F%BA%E6%9C%AC%E5%AE%9E%E7%8E%B0</a></p>",
          "rawMarkdown": "I tried l2-constrained loss, seem slightly worse than cross entropy loss. Also tried a simpler version AM-softmax loss, it's also slightly worse than cross entropy loss\n\n    AM-softmax loss\n    from keras.models import Model\n    from keras.layers import *\n    import keras.backend as K\n    from keras.constraints import unit_norm\n\n\n    x_in = Input(shape=(maxlen,))\n    x_embedded = Embedding(len(chars)+2,\n                       word_size)(x_in)\n    x = CuDNNGRU(word_size)(x_embedded)\n    x = Lambda(lambda x: K.l2_normalize(x, 1))(x)\n\n    pred = Dense(num_train,\n             use_bias=False,\n             kernel_constraint=unit_norm())(x)\n\n    encoder = Model(x_in, x) # 最终的目的是要得到一个编码器\n    model = Model(x_in, pred) # 用分类问题做训练\n\n    def amsoftmax_loss(y_true, y_pred, scale=30, margin=0.35):\n     y_pred = y_true * (y_pred - margin) + (1 - y_true) * y_pred\n     y_pred *= scale\n     return K.categorical_crossentropy(y_true, y_pred, from_logits=True)\n\n    model.compile(loss=amsoftmax_loss,\n              optimizer='adam',\n              metrics=['accuracy'])\n\n\nReference: https://kexue.fm/archives/5743#%E5%9F%BA%E6%9C%AC%E5%AE%9E%E7%8E%B0"
        },
        {
          "id": 410418,
          "postDate": "2018-10-26T03:01:19.660Z",
          "content": "<p>Do u modify the scale and margin param of amsoftmax ？</p>",
          "rawMarkdown": "Do u modify the scale and margin param of amsoftmax ？"
        },
        {
          "id": 410423,
          "postDate": "2018-10-26T03:15:20.710Z",
          "content": "<p>No,  I use default settings</p>",
          "rawMarkdown": "No,  I use default settings"
        },
        {
          "id": 410427,
          "postDate": "2018-10-26T03:27:11.200Z",
          "content": "<p>The default param is used for face verification task on CASIA-webface datasets which has much more ids than this competition. I think margin based softmax can perform better than the origin softmax only if we tune the params carefully. </p>",
          "rawMarkdown": "The default param is used for face verification task on CASIA-webface datasets which has much more ids than this competition. I think margin based softmax can perform better than the origin softmax only if we tune the params carefully. "
        },
        {
          "id": 410450,
          "postDate": "2018-10-26T04:53:43.943Z",
          "content": "<p>I see, I will update my result later...</p>\n\n<p>BTW, l2-constrained loss is a bit weird, I tried alpha from 30 to 40, it worse the result, but alpha is from 7 to 15, seems the result has been improved....</p>",
          "rawMarkdown": "I see, I will update my result later...\n\nBTW, l2-constrained loss is a bit weird, I tried alpha from 30 to 40, it worse the result, but alpha is from 7 to 15, seems the result has been improved...."
        }
      ]
    },
    {
      "id": 420212,
      "postDate": "2018-11-13T09:43:57.097Z",
      "content": "<p>one can treat the strokes as a video (i,e x,y time). Methods like video-based action recognition can then be apply.</p>",
      "rawMarkdown": "one can treat the strokes as a video (i,e x,y time). Methods like video-based action recognition can then be apply.",
      "replies": [
        {
          "id": 420962,
          "postDate": "2018-11-14T12:02:16.063Z",
          "content": "<p>I'd say there would be a lot of redundant information in the video. You don't add any information compared to using an image where the time dimension is color-coded, but you do require much larger datasamples. I don't think that's the best use of resources.</p>",
          "rawMarkdown": "I'd say there would be a lot of redundant information in the video. You don't add any information compared to using an image where the time dimension is color-coded, but you do require much larger datasamples. I don't think that's the best use of resources."
        }
      ]
    },
    {
      "id": 413540,
      "postDate": "2018-11-01T05:16:11.740Z",
      "content": "<p>It is not exactly top1+top2+top3.</p>\n\n<p>ton-3 = found1 + found2 + found3</p>\n\n<p>map@3 (this competition metric) = 1.0*found1 + 0.5*found2 + 0.33*found3</p>\n\n<p><a href=\"https://www.kaggle.com/wendykan/map-k-demo\">https://www.kaggle.com/wendykan/map-k-demo</a></p>",
      "rawMarkdown": "It is not exactly top1+top2+top3.\n\nton-3 = found1 + found2 + found3\n\nmap@3 (this competition metric) = 1.0*found1 + 0.5*found2 + 0.33*found3\n\nhttps://www.kaggle.com/wendykan/map-k-demo"
    },
    {
      "id": 409558,
      "postDate": "2018-10-24T13:40:32.140Z",
      "content": "<p>There is a lot of interesting work on sketch generation:\nSketch-pix2seq : <a href=\"https://arxiv.org/abs/1709.04121\">https://arxiv.org/abs/1709.04121</a>\n<a href=\"https://arxiv.org/abs/1704.03477\">https://arxiv.org/abs/1704.03477</a></p>\n\n<p>Based on that, one could develop an architecture for classification with VAE regularisation; which means that a VAE shares the same encoder used for classification.</p>\n\n<p>Alternatively a sketch generator conditioned on a given sketch could be used for test time augmentation. </p>",
      "rawMarkdown": "There is a lot of interesting work on sketch generation:\nSketch-pix2seq : https://arxiv.org/abs/1709.04121\nhttps://arxiv.org/abs/1704.03477\n\nBased on that, one could develop an architecture for classification with VAE regularisation; which means that a VAE shares the same encoder used for classification.\n\nAlternatively a sketch generator conditioned on a given sketch could be used for test time augmentation. \n"
    },
    {
      "id": 401743,
      "postDate": "2018-10-10T16:15:14.507Z",
      "content": "<p>good stroke-based lstm resources:</p>\n\n<p><a href=\"https://distill.pub/2016/handwriting/\">https://distill.pub/2016/handwriting/</a></p>",
      "rawMarkdown": "good stroke-based lstm resources:\n\nhttps://distill.pub/2016/handwriting/"
    },
    {
      "id": 400541,
      "postDate": "2018-10-08T13:46:21.787Z",
      "content": "<p>Weight Averaging sounds interesting, Here are some other paper that might be interesting as well. One is dealing with Hyperparamter optimization for large networks and Second is about using wideResents which supposed to be better and train faster. </p>\n\n<p><a href=\"https://arxiv.org/pdf/1803.09820.pdf\">https://arxiv.org/pdf/1803.09820.pdf</a>\n<a href=\"https://arxiv.org/pdf/1605.07146.pdf\">https://arxiv.org/pdf/1605.07146.pdf</a></p>",
      "rawMarkdown": "Weight Averaging sounds interesting, Here are some other paper that might be interesting as well. One is dealing with Hyperparamter optimization for large networks and Second is about using wideResents which supposed to be better and train faster. \n\nhttps://arxiv.org/pdf/1803.09820.pdf\nhttps://arxiv.org/pdf/1605.07146.pdf"
    }
  ],
  "comments": [
    {
      "id": 410647,
      "author_name": "Kees van Rooijen",
      "author_url": "",
      "post_date": "2018-10-26T11:57:38.250000",
      "content": "<p>I have implemented the top-k loss described in <a href=\"http://openaccess.thecvf.com/content_cvpr_2016/papers/Lapin_Loss_Functions_for_CVPR_2016_paper.pdf\">http://openaccess.thecvf.com/content_cvpr_2016/papers/Lapin_Loss_Functions_for_CVPR_2016_paper.pdf</a>. And although my top3 accuracy went up from ~92% to ~95%, at the same time the top1 accuracy went down by a lot, making the MAP@3 go down significantly (from .91 to .61 or so). So don't bother optimizing for top3 accuracy. </p>",
      "votes": 5,
      "replies": [
        {
          "id": 410661,
          "author_name": "jeandebleau",
          "author_url": "",
          "post_date": "2018-10-26T12:21:37.733000",
          "content": "<p>did you implement the truncated top-k cross entropy or the smooth top-k Hinge ?  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 410714,
          "author_name": "Kees van Rooijen",
          "author_url": "",
          "post_date": "2018-10-26T14:21:05.030000",
          "content": "<p>the top-k cross entropy :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 413520,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2018-11-01T04:17:41.693000",
          "content": "<p>kaggle loss is not exactly top-k. you have to modify a bit. e.g.</p>\n\n<p>kaggle_loss = top1+top2+top3</p>\n\n<p>also, if the different loss has different difficulties. the network may need to be larger.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 413676,
          "author_name": "Kees van Rooijen",
          "author_url": "",
          "post_date": "2018-11-01T10:56:56.400000",
          "content": "<p>Yeah, I figured. Also depends on to hat extent optimizing for top1 also co-optimizes the top3 and vice versa. Probably best to train a network for each and then train how to combine them after?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 401019,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-10-09T09:29:59.370000",
      "content": "<p>useful augmentation:</p>\n\n<ul>\n<li><p>stroke removal</p></li>\n<li><p>stroke simplification</p></li>\n<li><p>stroke deformation</p>\n\n<p><img src=\"https://media.springernature.com/original/springer-static/image/art%3A10.1007%2Fs11263-016-0932-3/MediaObjects/11263_2016_932_Fig2_HTML.gif\" alt=\"enter image description here\"></p>\n\n<p><img src=\"https://media.springernature.com/lw785/springer-static/image/art%3A10.1007%2Fs11263-016-0932-3/MediaObjects/11263_2016_932_Fig5_HTML.gif\" alt=\"enter image description here\"></p></li>\n</ul>\n\n<p><a href=\"http://sketchx.eecs.qmul.ac.uk/\">http://sketchx.eecs.qmul.ac.uk/</a></p>\n\n<p><a href=\"https://github.com/yuchuochuo1023/sketch-specific-data-augmentation\">https://github.com/yuchuochuo1023/sketch-specific-data-augmentation</a></p>\n\n<p><a href=\"https://www.eecs.qmul.ac.uk/~qian/Qian\">https://www.eecs.qmul.ac.uk/~qian/Qian</a>'s%20Materials/paper/IJCV_revised_version.pdf</p>",
      "votes": 5,
      "replies": [
        {
          "id": 410380,
          "author_name": "Gary",
          "author_url": "",
          "post_date": "2018-10-26T00:25:49.807000",
          "content": "<p>Hi, Heng. The multi-scale fusion mentioned here, is each scale separately trained, or multiple scales simultaneously trained?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 413670,
          "author_name": "HuyenNguyen",
          "author_url": "",
          "post_date": "2018-11-01T10:40:57.773000",
          "content": "<p>I looked up this original paper, they used a far smaller dataset with only 80 sketches per category. I don't know augmentation is necessary for this amount of data?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 400557,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-10-08T14:43:01.540000",
      "content": "<p><a href=\"http://openaccess.thecvf.com/content_cvpr_2018/CameraReady/2763.pdf\">http://openaccess.thecvf.com/content_cvpr_2018/CameraReady/2763.pdf</a></p>\n\n<p>SketchMate: Deep Hashing for Million-Scale Human Sketch Retrieval\n- Peng Xu</p>\n\n<ul>\n<li><p>joint stroke and image branch is interesting.</p></li>\n<li><p>with 2 separate models (stroke and image), distillation between them for semi-supervised learning is interesting</p>\n\n<p><img src=\"http://www.eecs.qmul.ac.uk/~kp306/Kaiyue%20Material/CVPR2018_HASH/thumbnail.png\" alt=\"enter image description here\"></p></li>\n</ul>",
      "votes": 6,
      "replies": [
        {
          "id": 400575,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2018-10-08T15:22:40.157000",
          "content": "<p>oh Wow! Interesting idea!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 409871,
          "author_name": "Gary",
          "author_url": "",
          "post_date": "2018-10-25T01:38:15.127000",
          "content": "<p>Did you try this? I think CNN+RNN is a great idea.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 410559,
          "author_name": "jeandebleau",
          "author_url": "",
          "post_date": "2018-10-26T09:00:27.130000",
          "content": "<p>You could also use a third network that takes as inputs a sequence of images.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 402816,
      "author_name": "Vladimir Larin",
      "author_url": "",
      "post_date": "2018-10-12T11:49:45.693000",
      "content": "<blockquote>\n  <p><a href=\"https://arxiv.org/abs/1802.07595\">https://arxiv.org/abs/1802.07595</a>\n  Smooth Loss Functions for Deep Top-k Classification - Leonard Berrada, Andrew Zisserman, M. Pawan Kumar</p>\n</blockquote>\n\n<p>There is a repository for this paper:\n<a href=\"https://github.com/oval-group/smooth-topk\">https://github.com/oval-group/smooth-topk</a></p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 407763,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-10-21T18:01:58.607000",
      "content": "<p>RNN weakness?\nEspecially, rnn encode dx,dy and not the absolute x,y.</p>\n\n<p>It seems that the spatial information is not important?\nNote that i split the components of \"key\" and \"smiley face\" apart</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/inbox/113660/b5dba192c86735f135da305714a41bdc/rnn.png\" alt=\"enter image description here\"></p>",
      "votes": 4,
      "replies": [
        {
          "id": 407979,
          "author_name": "Joe Ho",
          "author_url": "",
          "post_date": "2018-10-22T05:12:39.150000",
          "content": "<p>The current Model behind this game is RNN? If so, seems CNN should be better for this game?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 408764,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-10-23T12:21:37.490000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 419691,
          "author_name": "HuyenNguyen",
          "author_url": "",
          "post_date": "2018-11-12T12:05:45.807000",
          "content": "<p>I used dx, dy and it made no difference. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 420176,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-11-13T07:48:31.223000",
      "content": "<p><a href=\"https://arxiv.org/pdf/1708.02716.pdf\">https://arxiv.org/pdf/1708.02716.pdf</a></p>\n\n<p><a href=\"http://www.yugangjiang.info/publication/17MM-Sketch.pdf\">http://www.yugangjiang.info/publication/17MM-Sketch.pdf</a></p>\n\n<hr>\n\n<p><a href=\"https://ravika.github.io/publications.html\">https://ravika.github.io/publications.html</a></p>\n\n<p><a href=\"https://arxiv.org/pdf/1608.03369v1.pdf\">https://arxiv.org/pdf/1608.03369v1.pdf</a></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/420176/10664/cnnlstm1.png\" alt=\"enter image description here\"></p>\n\n<p>a pretty smart way to deal with incomplete sketch.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/420176/10665/cnnlstm2.png\" alt=\"enter image description here\">\n   <img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/420176/10666/cnnlstm3.png\" alt=\"enter image description here\"></p>",
      "votes": 1,
      "replies": [
        {
          "id": 420963,
          "author_name": "Kees van Rooijen",
          "author_url": "",
          "post_date": "2018-11-14T12:02:38.237000",
          "content": "<p>Thanks again, lot of useful info</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 419430,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-11-12T00:17:04.583000",
      "content": "<p>\"Learning From Noisy Large-Scale Datasets With Minimal Supervision\" -Andreas Veit, cvpr 2018</p>\n\n<p><a href=\"https://www.youtube.com/watch?v=RAHiCtJhyDg\">https://www.youtube.com/watch?v=RAHiCtJhyDg</a></p>\n\n<p>inspired by this paper, one could do this:</p>\n\n<ol>\n<li><p>construct feature extractor.  Train both cross-entropy(image classification head) and sigmoid classifier (label cleaning head).</p></li>\n<li><p>cross entropy is trained on  all images. sigmoid is trained on recognised images only.</p></li>\n<li><p>loss can be like:  loss = (weight) * (cross entropy loss). weight = 1 if it is recognised. weight = sigmoid classifier probability if it is non-recognised </p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/419430/10657/architecture_1.jpg\" alt=\"enter image description here\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/419430/10662/label_cleaning1.png\" alt=\"enter image description here\"></p></li>\n</ol>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 406544,
      "author_name": "jeandebleau",
      "author_url": "",
      "post_date": "2018-10-19T12:21:53.050000",
      "content": "<p>Soon, new models pretrained on shapes:</p>\n\n<p><a href=\"https://openreview.net/pdf?id=Bygh9j09KX\">https://openreview.net/pdf?id=Bygh9j09KX</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 400934,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-10-09T06:51:13.103000",
      "content": "<p>Loss design:</p>\n\n<p><a href=\"https://liusi-group.com/pdf/ijcai_2018.pdf\">https://liusi-group.com/pdf/ijcai_2018.pdf</a></p>\n\n<p>'Ensemble Soft-Margin Softmax Loss for Image Classification' - Xiaobo Wang</p>\n\n<p>related:  cross entropy, center loss,A-softmax,Large-margin softmax ...</p>",
      "votes": 1,
      "replies": [
        {
          "id": 409927,
          "author_name": "Joe Ho",
          "author_url": "",
          "post_date": "2018-10-25T04:44:13.293000",
          "content": "<p>I tried l2-constrained loss, seem slightly worse than cross entropy loss. Also tried a simpler version AM-softmax loss, it's also slightly worse than cross entropy loss</p>\n\n<pre><code>AM-softmax loss\nfrom keras.models import Model\nfrom keras.layers import *\nimport keras.backend as K\nfrom keras.constraints import unit_norm\n\n\nx_in = Input(shape=(maxlen,))\nx_embedded = Embedding(len(chars)+2,\n                   word_size)(x_in)\nx = CuDNNGRU(word_size)(x_embedded)\nx = Lambda(lambda x: K.l2_normalize(x, 1))(x)\n\npred = Dense(num_train,\n         use_bias=False,\n         kernel_constraint=unit_norm())(x)\n\nencoder = Model(x_in, x) # 最终的目的是要得到一个编码器\nmodel = Model(x_in, pred) # 用分类问题做训练\n\ndef amsoftmax_loss(y_true, y_pred, scale=30, margin=0.35):\n y_pred = y_true * (y_pred - margin) + (1 - y_true) * y_pred\n y_pred *= scale\n return K.categorical_crossentropy(y_true, y_pred, from_logits=True)\n\nmodel.compile(loss=amsoftmax_loss,\n          optimizer='adam',\n          metrics=['accuracy'])\n</code></pre>\n\n<p>Reference: <a href=\"https://kexue.fm/archives/5743#%E5%9F%BA%E6%9C%AC%E5%AE%9E%E7%8E%B0\">https://kexue.fm/archives/5743#%E5%9F%BA%E6%9C%AC%E5%AE%9E%E7%8E%B0</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 410418,
          "author_name": "SeuTao",
          "author_url": "",
          "post_date": "2018-10-26T03:01:19.660000",
          "content": "<p>Do u modify the scale and margin param of amsoftmax ？</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 410423,
          "author_name": "Joe Ho",
          "author_url": "",
          "post_date": "2018-10-26T03:15:20.710000",
          "content": "<p>No,  I use default settings</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 410427,
          "author_name": "SeuTao",
          "author_url": "",
          "post_date": "2018-10-26T03:27:11.200000",
          "content": "<p>The default param is used for face verification task on CASIA-webface datasets which has much more ids than this competition. I think margin based softmax can perform better than the origin softmax only if we tune the params carefully. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 410450,
          "author_name": "Joe Ho",
          "author_url": "",
          "post_date": "2018-10-26T04:53:43.943000",
          "content": "<p>I see, I will update my result later...</p>\n\n<p>BTW, l2-constrained loss is a bit weird, I tried alpha from 30 to 40, it worse the result, but alpha is from 7 to 15, seems the result has been improved....</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 420212,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-11-13T09:43:57.097000",
      "content": "<p>one can treat the strokes as a video (i,e x,y time). Methods like video-based action recognition can then be apply.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 420962,
          "author_name": "Kees van Rooijen",
          "author_url": "",
          "post_date": "2018-11-14T12:02:16.063000",
          "content": "<p>I'd say there would be a lot of redundant information in the video. You don't add any information compared to using an image where the time dimension is color-coded, but you do require much larger datasamples. I don't think that's the best use of resources.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 413540,
      "author_name": "Vladimir Larin",
      "author_url": "",
      "post_date": "2018-11-01T05:16:11.740000",
      "content": "<p>It is not exactly top1+top2+top3.</p>\n\n<p>ton-3 = found1 + found2 + found3</p>\n\n<p>map@3 (this competition metric) = 1.0*found1 + 0.5*found2 + 0.33*found3</p>\n\n<p><a href=\"https://www.kaggle.com/wendykan/map-k-demo\">https://www.kaggle.com/wendykan/map-k-demo</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 409558,
      "author_name": "jeandebleau",
      "author_url": "",
      "post_date": "2018-10-24T13:40:32.140000",
      "content": "<p>There is a lot of interesting work on sketch generation:\nSketch-pix2seq : <a href=\"https://arxiv.org/abs/1709.04121\">https://arxiv.org/abs/1709.04121</a>\n<a href=\"https://arxiv.org/abs/1704.03477\">https://arxiv.org/abs/1704.03477</a></p>\n\n<p>Based on that, one could develop an architecture for classification with VAE regularisation; which means that a VAE shares the same encoder used for classification.</p>\n\n<p>Alternatively a sketch generator conditioned on a given sketch could be used for test time augmentation. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 401743,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-10-10T16:15:14.507000",
      "content": "<p>good stroke-based lstm resources:</p>\n\n<p><a href=\"https://distill.pub/2016/handwriting/\">https://distill.pub/2016/handwriting/</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 400541,
      "author_name": "DrHB",
      "author_url": "",
      "post_date": "2018-10-08T13:46:21.787000",
      "content": "<p>Weight Averaging sounds interesting, Here are some other paper that might be interesting as well. One is dealing with Hyperparamter optimization for large networks and Second is about using wideResents which supposed to be better and train faster. </p>\n\n<p><a href=\"https://arxiv.org/pdf/1803.09820.pdf\">https://arxiv.org/pdf/1803.09820.pdf</a>\n<a href=\"https://arxiv.org/pdf/1605.07146.pdf\">https://arxiv.org/pdf/1605.07146.pdf</a></p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "400492": "some ideas that i would like to try\n\nhttps://arxiv.org/abs/1802.07595\n\nSmooth Loss Functions for Deep Top-k Classification\n - Leonard Berrada, Andrew Zisserman, M. Pawan Kumar\n\n---\nhttps://arxiv.org/abs/1806.05594\n\nImproving Consistency-Based Semi-Supervised Learning with Weight Averaging\n - Ben Athiwaratkun, Marc Finzi, Pavel Izmailov, Andrew Gordon Wilson\n\n\n---\n\nto be udpated ...\n\nquickdraw GAN,  Reinforcement Learning Deep-Q Nets ....\n\n\nBeyond rnn,eg attention, temporal cnn\n\ndeformable cnn, non-local mean, graph-cnn",
    "410647": "I have implemented the top-k loss described in http://openaccess.thecvf.com/content_cvpr_2016/papers/Lapin_Loss_Functions_for_CVPR_2016_paper.pdf. And although my top3 accuracy went up from ~92% to ~95%, at the same time the top1 accuracy went down by a lot, making the MAP@3 go down significantly (from .91 to .61 or so). So don't bother optimizing for top3 accuracy. ",
    "401019": "useful augmentation:\n\n-  stroke removal\n\n- stroke simplification\n\n- stroke deformation\n\n\n   ![enter image description here][1]\n\n\n    ![enter image description here][2]\n\n\n  [1]: https://media.springernature.com/original/springer-static/image/art%3A10.1007%2Fs11263-016-0932-3/MediaObjects/11263_2016_932_Fig2_HTML.gif\n  [2]: https://media.springernature.com/lw785/springer-static/image/art%3A10.1007%2Fs11263-016-0932-3/MediaObjects/11263_2016_932_Fig5_HTML.gif\n\n\nhttp://sketchx.eecs.qmul.ac.uk/\n\nhttps://github.com/yuchuochuo1023/sketch-specific-data-augmentation\n\nhttps://www.eecs.qmul.ac.uk/~qian/Qian's%20Materials/paper/IJCV_revised_version.pdf",
    "400557": "http://openaccess.thecvf.com/content_cvpr_2018/CameraReady/2763.pdf\n\n\nSketchMate: Deep Hashing for Million-Scale Human Sketch Retrieval\n- Peng Xu\n\n\n- joint stroke and image branch is interesting.\n\n- with 2 separate models (stroke and image), distillation between them for semi-supervised learning is interesting\n\n   ![enter image description here][1]\n\n\n  [1]: http://www.eecs.qmul.ac.uk/~kp306/Kaiyue%20Material/CVPR2018_HASH/thumbnail.png",
    "402816": "&gt; https://arxiv.org/abs/1802.07595\n&gt; Smooth Loss Functions for Deep Top-k Classification - Leonard Berrada, Andrew Zisserman, M. Pawan Kumar\n\nThere is a repository for this paper:\nhttps://github.com/oval-group/smooth-topk",
    "407763": "RNN weakness?\nEspecially, rnn encode dx,dy and not the absolute x,y.\n\n\nIt seems that the spatial information is not important?\nNote that i split the components of \"key\" and \"smiley face\" apart\n\n  ![enter image description here][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/inbox/113660/b5dba192c86735f135da305714a41bdc/rnn.png",
    "420176": "https://arxiv.org/pdf/1708.02716.pdf\n\nhttp://www.yugangjiang.info/publication/17MM-Sketch.pdf\n\n---\n\nhttps://ravika.github.io/publications.html\n\nhttps://arxiv.org/pdf/1608.03369v1.pdf\n\n\n   ![enter image description here][1]\n\na pretty smart way to deal with incomplete sketch.\n\n   ![enter image description here][2]\n   ![enter image description here][3]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/420176/10664/cnnlstm1.png\n  [2]: https://storage.googleapis.com/kaggle-forum-message-attachments/420176/10665/cnnlstm2.png\n  [3]: https://storage.googleapis.com/kaggle-forum-message-attachments/420176/10666/cnnlstm3.png",
    "419430": "\"Learning From Noisy Large-Scale Datasets With Minimal Supervision\" -Andreas Veit, cvpr 2018\n\nhttps://www.youtube.com/watch?v=RAHiCtJhyDg\n\ninspired by this paper, one could do this:\n\n1. construct feature extractor.  Train both cross-entropy(image classification head) and sigmoid classifier (label cleaning head).\n\n2. cross entropy is trained on  all images. sigmoid is trained on recognised images only.\n\n\n3. loss can be like:  loss = (weight) * (cross entropy loss). weight = 1 if it is recognised. weight = sigmoid classifier probability if it is non-recognised \n\n\n  ![enter image description here][1]\n\n\n  ![enter image description here][2]\n\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/419430/10657/architecture_1.jpg\n  [2]: https://storage.googleapis.com/kaggle-forum-message-attachments/419430/10662/label_cleaning1.png",
    "406544": "Soon, new models pretrained on shapes:\n\nhttps://openreview.net/pdf?id=Bygh9j09KX\n",
    "400934": "Loss design:\n\nhttps://liusi-group.com/pdf/ijcai_2018.pdf\n\n'Ensemble Soft-Margin Softmax Loss for Image Classification' - Xiaobo Wang\n\n\nrelated:  cross entropy, center loss,A-softmax,Large-margin softmax ...",
    "420212": "one can treat the strokes as a video (i,e x,y time). Methods like video-based action recognition can then be apply.",
    "413540": "It is not exactly top1+top2+top3.\n\nton-3 = found1 + found2 + found3\n\nmap@3 (this competition metric) = 1.0*found1 + 0.5*found2 + 0.33*found3\n\nhttps://www.kaggle.com/wendykan/map-k-demo",
    "409558": "There is a lot of interesting work on sketch generation:\nSketch-pix2seq : https://arxiv.org/abs/1709.04121\nhttps://arxiv.org/abs/1704.03477\n\nBased on that, one could develop an architecture for classification with VAE regularisation; which means that a VAE shares the same encoder used for classification.\n\nAlternatively a sketch generator conditioned on a given sketch could be used for test time augmentation. \n",
    "401743": "good stroke-based lstm resources:\n\nhttps://distill.pub/2016/handwriting/",
    "400541": "Weight Averaging sounds interesting, Here are some other paper that might be interesting as well. One is dealing with Hyperparamter optimization for large networks and Second is about using wideResents which supposed to be better and train faster. \n\nhttps://arxiv.org/pdf/1803.09820.pdf\nhttps://arxiv.org/pdf/1605.07146.pdf"
  }
}