{
  "id": 71313,
  "title": "how is the test data collected?",
  "url": "/competitions/quickdraw-doodle-recognition/discussion/71313",
  "author_name": "hengck23",
  "post_date": "2018-11-12T14:20:29.181000",
  "votes": 8,
  "comment_count": 7,
  "views": 0,
  "content": "<p>this is my guess, may or may not be correct</p>\n\n<ol>\n<li><p>there are unusual pause in the strokes in the test (just like the train) so i think test set is also from quickdraw game.</p></li>\n<li><p>it is mentioned on kaggle that \" ... perform well on a manually-labeled test set from a different distribution.\" But i don't think the test data is annotated from scratch ... there are just too many of them!</p></li>\n<li><p>Rather, a subset of data from quickdraw game is selected. Then a human annotator verifies the recognized drawing and attempt the hand label the non-recognized ones. The rubbish drawing are discarded.</p></li>\n</ol>\n\n<p>Hence if you use recognized train drawing only, you would get a lower LB score.\nIf you use all train drawing, you get a higher LB score.</p>\n\n<p>To further verify this:</p>\n\n<p>select a validation set. Spit it into recognized and non-recognized. For the  non-recognized, hand correct them. Monitor your validation metric for: </p>\n\n<ul>\n<li><p>all validation set</p></li>\n<li><p>recognized validation set</p></li>\n<li><p>non-recognized validation set</p></li>\n<li><p>hand corrected non-recognized  validation set</p></li>\n<li><p>discarded non-recognized  validation set</p></li>\n<li><p>hand corrected non-recognized +  recognized validation set</p></li>\n</ul>",
  "messages": [
    {
      "id": 419760,
      "postDate": "2018-11-12T14:20:29.180Z",
      "content": "<p>this is my guess, may or may not be correct</p>\n\n<ol>\n<li><p>there are unusual pause in the strokes in the test (just like the train) so i think test set is also from quickdraw game.</p></li>\n<li><p>it is mentioned on kaggle that \" ... perform well on a manually-labeled test set from a different distribution.\" But i don't think the test data is annotated from scratch ... there are just too many of them!</p></li>\n<li><p>Rather, a subset of data from quickdraw game is selected. Then a human annotator verifies the recognized drawing and attempt the hand label the non-recognized ones. The rubbish drawing are discarded.</p></li>\n</ol>\n\n<p>Hence if you use recognized train drawing only, you would get a lower LB score.\nIf you use all train drawing, you get a higher LB score.</p>\n\n<p>To further verify this:</p>\n\n<p>select a validation set. Spit it into recognized and non-recognized. For the  non-recognized, hand correct them. Monitor your validation metric for: </p>\n\n<ul>\n<li><p>all validation set</p></li>\n<li><p>recognized validation set</p></li>\n<li><p>non-recognized validation set</p></li>\n<li><p>hand corrected non-recognized  validation set</p></li>\n<li><p>discarded non-recognized  validation set</p></li>\n<li><p>hand corrected non-recognized +  recognized validation set</p></li>\n</ul>",
      "rawMarkdown": "this is my guess, may or may not be correct\n\n1. there are unusual pause in the strokes in the test (just like the train) so i think test set is also from quickdraw game.\n\n2. it is mentioned on kaggle that \" ... perform well on a manually-labeled test set from a different distribution.\" But i don't think the test data is annotated from scratch ... there are just too many of them!\n\n3. Rather, a subset of data from quickdraw game is selected. Then a human annotator verifies the recognized drawing and attempt the hand label the non-recognized ones. The rubbish drawing are discarded.\n\nHence if you use recognized train drawing only, you would get a lower LB score.\nIf you use all train drawing, you get a higher LB score.\n\nTo further verify this:\n\nselect a validation set. Spit it into recognized and non-recognized. For the  non-recognized, hand correct them. Monitor your validation metric for: \n\n- all validation set\n\n- recognized validation set\n\n- non-recognized validation set\n\n- hand corrected non-recognized  validation set\n\n- discarded non-recognized  validation set\n\n- hand corrected non-recognized +  recognized validation set\n\n",
      "votes": 8
    },
    {
      "id": 420527,
      "postDate": "2018-11-13T19:13:53.350Z",
      "content": "<p>If we only use recognized training we are actually forcing our model to simulate Google's model, a clearly wrong direction. Not to say Google's model is an RNN which differs more from human visual system than a CNN does. </p>",
      "rawMarkdown": "If we only use recognized training we are actually forcing our model to simulate Google's model, a clearly wrong direction. Not to say Google's model is an RNN which differs more from human visual system than a CNN does. ",
      "votes": 1
    },
    {
      "id": 420512,
      "postDate": "2018-11-13T18:47:24.487Z",
      "content": "<p>The unrecognized drawings might contain a sub-class that was ignored by google's model.\nFor example, most of the recognized training examples for 'zebra' are full-body poses, but a few of the 12.1% of non-recognized zebras are head-shots:\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/420512/10671/bad_zebras.png\" alt=\"unrecognized zebras\">\nSo a model trained only on recognized zebras might miss these two drawings from the test examples:\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/420512/10672/test_007024.png\" alt=\"zebra head\">\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/420512/10673/test_009603.png\" alt=\"\"></p>\n\n<p>(I just noticed a similar problem with 'camera'.  The recognized images are mostly front views and three-quarter views--no side views)</p>",
      "rawMarkdown": "The unrecognized drawings might contain a sub-class that was ignored by google's model.\nFor example, most of the recognized training examples for 'zebra' are full-body poses, but a few of the 12.1% of non-recognized zebras are head-shots:\n![unrecognized zebras][1]\nSo a model trained only on recognized zebras might miss these two drawings from the test examples:\n![zebra head][2]\n![][3]\n\n\n(I just noticed a similar problem with 'camera'.  The recognized images are mostly front views and three-quarter views--no side views)\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/420512/10671/bad_zebras.png\n  [2]: https://storage.googleapis.com/kaggle-forum-message-attachments/420512/10672/test_007024.png\n  [3]: https://storage.googleapis.com/kaggle-forum-message-attachments/420512/10673/test_009603.png",
      "votes": 2
    },
    {
      "id": 419782,
      "postDate": "2018-11-12T14:54:47.660Z",
      "content": "<p>Something to consider though is that any drawing that needs to be hand-labeled is so bad that it is just noise.  So even if you train on those, it is not likely there will ever be another one like that, so it just makes the model less accurate by having these outliers. But, your idea for testing all of this would be a great way to get the answer. I have not done any experiments with it yet.</p>",
      "rawMarkdown": "Something to consider though is that any drawing that needs to be hand-labeled is so bad that it is just noise.  So even if you train on those, it is not likely there will ever be another one like that, so it just makes the model less accurate by having these outliers. But, your idea for testing all of this would be a great way to get the answer. I have not done any experiments with it yet.",
      "replies": [
        {
          "id": 419790,
          "postDate": "2018-11-12T15:04:51.590Z",
          "content": "<p>\"Something to consider though is that any drawing that needs to be hand-labeled is so bad that it is just noise\"</p>\n\n<p>Not true. 50% of the non-recognized drawing are actually good drawing. It is just that the accuracy of the quickdraw game is  not accurate enough to recognized them.</p>\n\n<p>These 50% can be easily identified as human as valid drawing. I actually plot out the non-recognized drawing and verify by visual inspection and by experiment below:</p>\n\n<p>e.g:</p>\n\n<ul>\n<li><p>addition of xx valid drawing cause +zz improvement in LB.</p></li>\n<li><p>addition of xx noise label (random scribble) cause -zz drop in LB.</p></li>\n</ul>",
          "rawMarkdown": "\"Something to consider though is that any drawing that needs to be hand-labeled is so bad that it is just noise\"\n\nNot true. 50% of the non-recognized drawing are actually good drawing. It is just that the accuracy of the quickdraw game is  not accurate enough to recognized them.\n\nThese 50% can be easily identified as human as valid drawing. I actually plot out the non-recognized drawing and verify by visual inspection and by experiment below:\n\n\ne.g:\n\n- addition of xx valid drawing cause +zz improvement in LB.\n\n- addition of xx noise label (random scribble) cause -zz drop in LB.\n"
        },
        {
          "id": 419809,
          "postDate": "2018-11-12T15:32:18.890Z",
          "content": "<p>Yes, the results of that experiment would be very interesting.</p>",
          "rawMarkdown": "Yes, the results of that experiment would be very interesting."
        },
        {
          "id": 420261,
          "postDate": "2018-11-13T11:16:34.270Z",
          "content": "<p>but even a lot of the recognised images are rubbish and really should be considered noise: <a href=\"https://www.kaggle.com/gaborfodor/un-recognized-drawings/output\">https://www.kaggle.com/gaborfodor/un-recognized-drawings/output</a></p>",
          "rawMarkdown": "but even a lot of the recognised images are rubbish and really should be considered noise: https://www.kaggle.com/gaborfodor/un-recognized-drawings/output"
        }
      ]
    },
    {
      "id": 419797,
      "postDate": "2018-11-12T15:12:44.313Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 420527,
      "author_name": "Guanshuo Xu",
      "author_url": "",
      "post_date": "2018-11-13T19:13:53.350000",
      "content": "<p>If we only use recognized training we are actually forcing our model to simulate Google's model, a clearly wrong direction. Not to say Google's model is an RNN which differs more from human visual system than a CNN does. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 420512,
      "author_name": "raifer",
      "author_url": "",
      "post_date": "2018-11-13T18:47:24.487000",
      "content": "<p>The unrecognized drawings might contain a sub-class that was ignored by google's model.\nFor example, most of the recognized training examples for 'zebra' are full-body poses, but a few of the 12.1% of non-recognized zebras are head-shots:\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/420512/10671/bad_zebras.png\" alt=\"unrecognized zebras\">\nSo a model trained only on recognized zebras might miss these two drawings from the test examples:\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/420512/10672/test_007024.png\" alt=\"zebra head\">\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/420512/10673/test_009603.png\" alt=\"\"></p>\n\n<p>(I just noticed a similar problem with 'camera'.  The recognized images are mostly front views and three-quarter views--no side views)</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 419782,
      "author_name": "impulsecorp",
      "author_url": "",
      "post_date": "2018-11-12T14:54:47.660000",
      "content": "<p>Something to consider though is that any drawing that needs to be hand-labeled is so bad that it is just noise.  So even if you train on those, it is not likely there will ever be another one like that, so it just makes the model less accurate by having these outliers. But, your idea for testing all of this would be a great way to get the answer. I have not done any experiments with it yet.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 419790,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2018-11-12T15:04:51.590000",
          "content": "<p>\"Something to consider though is that any drawing that needs to be hand-labeled is so bad that it is just noise\"</p>\n\n<p>Not true. 50% of the non-recognized drawing are actually good drawing. It is just that the accuracy of the quickdraw game is  not accurate enough to recognized them.</p>\n\n<p>These 50% can be easily identified as human as valid drawing. I actually plot out the non-recognized drawing and verify by visual inspection and by experiment below:</p>\n\n<p>e.g:</p>\n\n<ul>\n<li><p>addition of xx valid drawing cause +zz improvement in LB.</p></li>\n<li><p>addition of xx noise label (random scribble) cause -zz drop in LB.</p></li>\n</ul>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 419809,
          "author_name": "impulsecorp",
          "author_url": "",
          "post_date": "2018-11-12T15:32:18.890000",
          "content": "<p>Yes, the results of that experiment would be very interesting.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 420261,
          "author_name": "HuyenNguyen",
          "author_url": "",
          "post_date": "2018-11-13T11:16:34.270000",
          "content": "<p>but even a lot of the recognised images are rubbish and really should be considered noise: <a href=\"https://www.kaggle.com/gaborfodor/un-recognized-drawings/output\">https://www.kaggle.com/gaborfodor/un-recognized-drawings/output</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 419797,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-12T15:12:44.313000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "419760": "this is my guess, may or may not be correct\n\n1. there are unusual pause in the strokes in the test (just like the train) so i think test set is also from quickdraw game.\n\n2. it is mentioned on kaggle that \" ... perform well on a manually-labeled test set from a different distribution.\" But i don't think the test data is annotated from scratch ... there are just too many of them!\n\n3. Rather, a subset of data from quickdraw game is selected. Then a human annotator verifies the recognized drawing and attempt the hand label the non-recognized ones. The rubbish drawing are discarded.\n\nHence if you use recognized train drawing only, you would get a lower LB score.\nIf you use all train drawing, you get a higher LB score.\n\nTo further verify this:\n\nselect a validation set. Spit it into recognized and non-recognized. For the  non-recognized, hand correct them. Monitor your validation metric for: \n\n- all validation set\n\n- recognized validation set\n\n- non-recognized validation set\n\n- hand corrected non-recognized  validation set\n\n- discarded non-recognized  validation set\n\n- hand corrected non-recognized +  recognized validation set\n\n",
    "420527": "If we only use recognized training we are actually forcing our model to simulate Google's model, a clearly wrong direction. Not to say Google's model is an RNN which differs more from human visual system than a CNN does. ",
    "420512": "The unrecognized drawings might contain a sub-class that was ignored by google's model.\nFor example, most of the recognized training examples for 'zebra' are full-body poses, but a few of the 12.1% of non-recognized zebras are head-shots:\n![unrecognized zebras][1]\nSo a model trained only on recognized zebras might miss these two drawings from the test examples:\n![zebra head][2]\n![][3]\n\n\n(I just noticed a similar problem with 'camera'.  The recognized images are mostly front views and three-quarter views--no side views)\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/420512/10671/bad_zebras.png\n  [2]: https://storage.googleapis.com/kaggle-forum-message-attachments/420512/10672/test_007024.png\n  [3]: https://storage.googleapis.com/kaggle-forum-message-attachments/420512/10673/test_009603.png",
    "419782": "Something to consider though is that any drawing that needs to be hand-labeled is so bad that it is just noise.  So even if you train on those, it is not likely there will ever be another one like that, so it just makes the model less accurate by having these outliers. But, your idea for testing all of this would be a great way to get the answer. I have not done any experiments with it yet.",
    "419797": ""
  }
}