{
  "id": 72921,
  "title": " It's all about noise ",
  "url": "/competitions/quickdraw-doodle-recognition/discussion/72921",
  "author_name": "miguel perez",
  "post_date": "2018-11-28T11:07:05.793000",
  "votes": 8,
  "comment_count": 2,
  "views": 0,
  "content": "<p>This competition is shaped by two ingredients that highly condition best performing solutions; size of data and noise. </p>\n\n<p>Data <strong>size</strong> makes it closer to theoretical \"ideal\" situation, in which our universal aproximators can learn just anything necessary. (One of first obvious consequences of this is some of the usual go-to tools are no longer necessary and possibly harmful, like image augmentation or even pretrained weights)</p>\n\n<p>But in my opinion the elephant in the room here is <strong>noise</strong>. Specifically the kind of noise made of completely mislabeled images. I don't mean a kangaroo slightly resembling a squirrel. I mean a kangaroo that someone, for whatever reason, decided to draw as a tennis raquet. That kind of noise. So lets think about that noise :</p>\n\n<p>-<strong>In theory, as long as  noise is evenly distributed across classes a neural net with enough data can filter mislabeling noise (\"swap noise\") perfectly</strong>. \n-In practice, I'm finding it difficult with 30% of data that is already quite a lot. </p>\n\n<p>I suspect that part of the difficulty is due to another of the usually go-to tools we all use without much questioning; <strong>adaptative learning rates, cyclic training, or both</strong>. The issue is,  noise correction (of this kind) should be balanced, wrong labels correcting each other in a balanced way throughout training. But if learning rates vary then this balance can take many more epochs to happen. As this is just an intuition I will like to know  if anyone has a good insight about this relatioship.</p>\n\n<p>And, to make all this even more interesting, public leaderboard scores seem to account for less noise than expected. Will that hold for private LB? I wouldn't bet on it...</p>",
  "messages": [
    {
      "id": 429113,
      "postDate": "2018-11-28T11:07:05.793Z",
      "content": "<p>This competition is shaped by two ingredients that highly condition best performing solutions; size of data and noise. </p>\n\n<p>Data <strong>size</strong> makes it closer to theoretical \"ideal\" situation, in which our universal aproximators can learn just anything necessary. (One of first obvious consequences of this is some of the usual go-to tools are no longer necessary and possibly harmful, like image augmentation or even pretrained weights)</p>\n\n<p>But in my opinion the elephant in the room here is <strong>noise</strong>. Specifically the kind of noise made of completely mislabeled images. I don't mean a kangaroo slightly resembling a squirrel. I mean a kangaroo that someone, for whatever reason, decided to draw as a tennis raquet. That kind of noise. So lets think about that noise :</p>\n\n<p>-<strong>In theory, as long as  noise is evenly distributed across classes a neural net with enough data can filter mislabeling noise (\"swap noise\") perfectly</strong>. \n-In practice, I'm finding it difficult with 30% of data that is already quite a lot. </p>\n\n<p>I suspect that part of the difficulty is due to another of the usually go-to tools we all use without much questioning; <strong>adaptative learning rates, cyclic training, or both</strong>. The issue is,  noise correction (of this kind) should be balanced, wrong labels correcting each other in a balanced way throughout training. But if learning rates vary then this balance can take many more epochs to happen. As this is just an intuition I will like to know  if anyone has a good insight about this relatioship.</p>\n\n<p>And, to make all this even more interesting, public leaderboard scores seem to account for less noise than expected. Will that hold for private LB? I wouldn't bet on it...</p>",
      "rawMarkdown": "This competition is shaped by two ingredients that highly condition best performing solutions; size of data and noise. \n\nData **size** makes it closer to theoretical \"ideal\" situation, in which our universal aproximators can learn just anything necessary. (One of first obvious consequences of this is some of the usual go-to tools are no longer necessary and possibly harmful, like image augmentation or even pretrained weights)\n\nBut in my opinion the elephant in the room here is **noise**. Specifically the kind of noise made of completely mislabeled images. I don't mean a kangaroo slightly resembling a squirrel. I mean a kangaroo that someone, for whatever reason, decided to draw as a tennis raquet. That kind of noise. So lets think about that noise :\n\n-**In theory, as long as  noise is evenly distributed across classes a neural net with enough data can filter mislabeling noise (\"swap noise\") perfectly**. \n-In practice, I'm finding it difficult with 30% of data that is already quite a lot. \n\nI suspect that part of the difficulty is due to another of the usually go-to tools we all use without much questioning; **adaptative learning rates, cyclic training, or both**. The issue is,  noise correction (of this kind) should be balanced, wrong labels correcting each other in a balanced way throughout training. But if learning rates vary then this balance can take many more epochs to happen. As this is just an intuition I will like to know  if anyone has a good insight about this relatioship.\n\nAnd, to make all this even more interesting, public leaderboard scores seem to account for less noise than expected. Will that hold for private LB? I wouldn't bet on it...",
      "votes": 8
    },
    {
      "id": 431369,
      "postDate": "2018-12-02T05:45:45.643Z",
      "content": "<p>With the level of noise I can't help but believe that the LB scores for the private are going to be much worse than public.  A good part of that, of course, depends on how many noisy images are in the test set. <br>\nI played the game for about 30 minutes - which was 25 too long....   The last 15 minutes I was drawing to fool google - I was successful.  The key to my success - I think -- I drew almost the exact same crap every time.  In my theory, if I give you the same wrong answer every time, no neural net regardless of data size can determine the object I am not trying to draw.   I am very worried that my current 0.93 model is huge over fit.</p>\n\n<p>Maybe I am confused about the metric -- but it seems to me that the top LB score is around what I would expect for a well evaluated set of images like imagenet.  To expect that score to be correct  for quick draw images assumes that a very high percentage of drawings were done with honest intent rather than dishonest intent like my drawings (or that google salted the percentage of noise).</p>\n\n<p>The above comments don't stop me from a frantic attempt to get to 0.94, but that is because I have never practiced what I preach for 72 years - no sense starting now!</p>",
      "rawMarkdown": "With the level of noise I can't help but believe that the LB scores for the private are going to be much worse than public.  A good part of that, of course, depends on how many noisy images are in the test set.  \nI played the game for about 30 minutes - which was 25 too long....   The last 15 minutes I was drawing to fool google - I was successful.  The key to my success - I think -- I drew almost the exact same crap every time.  In my theory, if I give you the same wrong answer every time, no neural net regardless of data size can determine the object I am not trying to draw.   I am very worried that my current 0.93 model is huge over fit.\n\nMaybe I am confused about the metric -- but it seems to me that the top LB score is around what I would expect for a well evaluated set of images like imagenet.  To expect that score to be correct  for quick draw images assumes that a very high percentage of drawings were done with honest intent rather than dishonest intent like my drawings (or that google salted the percentage of noise).\n\nThe above comments don't stop me from a frantic attempt to get to 0.94, but that is because I have never practiced what I preach for 72 years - no sense starting now!",
      "votes": 5,
      "replies": [
        {
          "id": 431439,
          "postDate": "2018-12-02T09:19:13.890Z",
          "content": "<p>My take is that the private test data will be as noisy as the training data, if not even noisier. So I felt there is no need to remove them from my training data and let the NN deal with it as it wishes. Perhaps I am fooling myself. Either way I will know in 3 days.</p>",
          "rawMarkdown": "My take is that the private test data will be as noisy as the training data, if not even noisier. So I felt there is no need to remove them from my training data and let the NN deal with it as it wishes. Perhaps I am fooling myself. Either way I will know in 3 days."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 431369,
      "author_name": "PC Jimmmy",
      "author_url": "",
      "post_date": "2018-12-02T05:45:45.643000",
      "content": "<p>With the level of noise I can't help but believe that the LB scores for the private are going to be much worse than public.  A good part of that, of course, depends on how many noisy images are in the test set. <br>\nI played the game for about 30 minutes - which was 25 too long....   The last 15 minutes I was drawing to fool google - I was successful.  The key to my success - I think -- I drew almost the exact same crap every time.  In my theory, if I give you the same wrong answer every time, no neural net regardless of data size can determine the object I am not trying to draw.   I am very worried that my current 0.93 model is huge over fit.</p>\n\n<p>Maybe I am confused about the metric -- but it seems to me that the top LB score is around what I would expect for a well evaluated set of images like imagenet.  To expect that score to be correct  for quick draw images assumes that a very high percentage of drawings were done with honest intent rather than dishonest intent like my drawings (or that google salted the percentage of noise).</p>\n\n<p>The above comments don't stop me from a frantic attempt to get to 0.94, but that is because I have never practiced what I preach for 72 years - no sense starting now!</p>",
      "votes": 5,
      "replies": [
        {
          "id": 431439,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-12-02T09:19:13.890000",
          "content": "<p>My take is that the private test data will be as noisy as the training data, if not even noisier. So I felt there is no need to remove them from my training data and let the NN deal with it as it wishes. Perhaps I am fooling myself. Either way I will know in 3 days.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "429113": "This competition is shaped by two ingredients that highly condition best performing solutions; size of data and noise. \n\nData **size** makes it closer to theoretical \"ideal\" situation, in which our universal aproximators can learn just anything necessary. (One of first obvious consequences of this is some of the usual go-to tools are no longer necessary and possibly harmful, like image augmentation or even pretrained weights)\n\nBut in my opinion the elephant in the room here is **noise**. Specifically the kind of noise made of completely mislabeled images. I don't mean a kangaroo slightly resembling a squirrel. I mean a kangaroo that someone, for whatever reason, decided to draw as a tennis raquet. That kind of noise. So lets think about that noise :\n\n-**In theory, as long as  noise is evenly distributed across classes a neural net with enough data can filter mislabeling noise (\"swap noise\") perfectly**. \n-In practice, I'm finding it difficult with 30% of data that is already quite a lot. \n\nI suspect that part of the difficulty is due to another of the usually go-to tools we all use without much questioning; **adaptative learning rates, cyclic training, or both**. The issue is,  noise correction (of this kind) should be balanced, wrong labels correcting each other in a balanced way throughout training. But if learning rates vary then this balance can take many more epochs to happen. As this is just an intuition I will like to know  if anyone has a good insight about this relatioship.\n\nAnd, to make all this even more interesting, public leaderboard scores seem to account for less noise than expected. Will that hold for private LB? I wouldn't bet on it...",
    "431369": "With the level of noise I can't help but believe that the LB scores for the private are going to be much worse than public.  A good part of that, of course, depends on how many noisy images are in the test set.  \nI played the game for about 30 minutes - which was 25 too long....   The last 15 minutes I was drawing to fool google - I was successful.  The key to my success - I think -- I drew almost the exact same crap every time.  In my theory, if I give you the same wrong answer every time, no neural net regardless of data size can determine the object I am not trying to draw.   I am very worried that my current 0.93 model is huge over fit.\n\nMaybe I am confused about the metric -- but it seems to me that the top LB score is around what I would expect for a well evaluated set of images like imagenet.  To expect that score to be correct  for quick draw images assumes that a very high percentage of drawings were done with honest intent rather than dishonest intent like my drawings (or that google salted the percentage of noise).\n\nThe above comments don't stop me from a frantic attempt to get to 0.94, but that is because I have never practiced what I preach for 72 years - no sense starting now!"
  }
}