{
  "id": 71662,
  "title": "About noisy samples",
  "url": "/competitions/quickdraw-doodle-recognition/discussion/71662",
  "author_name": "upup",
  "post_date": "2018-11-15T12:15:19.322000",
  "votes": 3,
  "comment_count": 2,
  "views": 0,
  "content": "<p>From this <a href=\"https://www.kaggle.com/gaborfodor/un-recognized-drawings/output\">kernel</a>(from beluga, thanks) we can see there are some \"noisy\" samples in both unrecognized and recognized samples. <br>\ne.g.(green -&gt; recognized, red -&gt; unrecognized) \n<img src=\"http://ww1.sinaimg.cn/large/0068zQzgly1fx9072qzaej30nz0g5tdr.jpg\" alt=\"some noise\">\nThese noisy samples are not so good for our models, what should we do? Should we design a binary classifier to find all these samples or drop them by hand?</p>",
  "messages": [
    {
      "id": 421783,
      "postDate": "2018-11-15T12:15:19.323Z",
      "content": "<p>From this <a href=\"https://www.kaggle.com/gaborfodor/un-recognized-drawings/output\">kernel</a>(from beluga, thanks) we can see there are some \"noisy\" samples in both unrecognized and recognized samples. <br>\ne.g.(green -&gt; recognized, red -&gt; unrecognized) \n<img src=\"http://ww1.sinaimg.cn/large/0068zQzgly1fx9072qzaej30nz0g5tdr.jpg\" alt=\"some noise\">\nThese noisy samples are not so good for our models, what should we do? Should we design a binary classifier to find all these samples or drop them by hand?</p>",
      "rawMarkdown": "From this [kernel](https://www.kaggle.com/gaborfodor/un-recognized-drawings/output)(from beluga, thanks) we can see there are some \"noisy\" samples in both unrecognized and recognized samples.    \ne.g.(green -&gt; recognized, red -&gt; unrecognized) \n![some noise][1]\nThese noisy samples are not so good for our models, what should we do? Should we design a binary classifier to find all these samples or drop them by hand?\n\n  [1]: http://ww1.sinaimg.cn/large/0068zQzgly1fx9072qzaej30nz0g5tdr.jpg",
      "votes": 3
    },
    {
      "id": 421931,
      "postDate": "2018-11-15T15:21:23.447Z",
      "content": "<p>There is a entropy calculator in public kernels that could classify.</p>",
      "rawMarkdown": "There is a entropy calculator in public kernels that could classify.",
      "votes": 2,
      "replies": [
        {
          "id": 421943,
          "postDate": "2018-11-15T15:33:02.297Z",
          "content": "<p>I got it, thanks!</p>",
          "rawMarkdown": "I got it, thanks!"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 421931,
      "author_name": "Zineng Tang",
      "author_url": "",
      "post_date": "2018-11-15T15:21:23.447000",
      "content": "<p>There is a entropy calculator in public kernels that could classify.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 421943,
          "author_name": "upup",
          "author_url": "",
          "post_date": "2018-11-15T15:33:02.297000",
          "content": "<p>I got it, thanks!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "421783": "From this [kernel](https://www.kaggle.com/gaborfodor/un-recognized-drawings/output)(from beluga, thanks) we can see there are some \"noisy\" samples in both unrecognized and recognized samples.    \ne.g.(green -&gt; recognized, red -&gt; unrecognized) \n![some noise][1]\nThese noisy samples are not so good for our models, what should we do? Should we design a binary classifier to find all these samples or drop them by hand?\n\n  [1]: http://ww1.sinaimg.cn/large/0068zQzgly1fx9072qzaej30nz0g5tdr.jpg",
    "421931": "There is a entropy calculator in public kernels that could classify."
  }
}