{
  "id": 72199,
  "title": "An elementary method of data clean? Ingenious or nonsense?",
  "url": "/competitions/quickdraw-doodle-recognition/discussion/72199",
  "author_name": "Dilapsky Lee",
  "post_date": "2018-11-21T08:18:08.953000",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>In this competition, the dataset can be divided into two parts: successful recognition and failed recognition.</p>\n\n<p>Only pictures with label ‘successful recognition’ counts. If we need to clean the data and reduce the noise, we can delete all data with ‘failed recognition’. However, in binary classification, we have these four conditions:</p>\n\n<p>(1) True positive: Drawer draw a correct doodle, and google mark it as ‘successful recognition’.\n(2) True negative: Drawer draw a wrong doodle, and google mark it as ‘failed recognition’.\n(3) False positive: Drawer draw a wrong doodle, but google mark it as ‘successful recognition’.\n(4) False negative: Drawer draw a correct doodle, but google mark it as ‘failed recognition’.</p>\n\n<p>The noise data contains 2 and 3, but if we delete all ‘failed recognition’ data, we delete 2 and 4. It’s hard to say that new data will perform better than original data.</p>\n\n<p>However, we can assume that the proportion of noise in ‘failed recognition’ is much larger than that in ‘successful recognition’ data.</p>\n\n<p>Here, I have an idea of using ‘average doodle’, for example: we have 500 doodles of train-simplified cats, \nFirst, we transfer the (x,y) form data to numpy-picture data. \nSecond, we add all pixels up and calculate the average value, and we will get an average numpy-picture data as ‘average doodle’ based on 500 doodles. \nThen, we calculate the Euclidean distance between each doodle picture and ‘average doodle’ picture.\nAt last, we can delete all data whose distance is exceed than a threshold.</p>\n\n<p>The principle of this method is simple: If your doodle has more difference than other’s doodles, you doodle is wrong.</p>\n\n<p>Perhaps you will ask: how can you prove that your method is effective? In other words, will your new dataset perform better than original data?</p>\n\n<p>We have reached the assumption that ‘the proportion of noise in ‘failed recognition’ is much larger than that in ‘successful recognition’ data.’ Thus, if we can prove that the proportion of ‘False’ in dataframe reduce, we can prove that.</p>\n\n<p>In fact, I have established a kernel to show my method, and the result is:</p>\n\n<p>success num: 261</p>\n\n<p>fail num: 79</p>\n\n<p>You can find the kernel <a href=\"https://www.kaggle.com/dilapsky/an-elementary-method-of-data-clean?scriptVersionId=7602883\">Here</a>\nand thanks beluga for his public kernel ‘GreyScale Mobilenet’<a href=\"https://www.kaggle.com/gaborfodor/greyscale-mobilenet-lb-0-892\">Here</a>, because I use part of his code in it</p>\n\n<p>However, this is only my trail of data clean. Perhaps it will be useful, perhaps it is only nonsense. Well, what do you think of this method? Welcome to publish your opinion!</p>",
  "messages": [
    {
      "id": 425269,
      "postDate": "2018-11-21T11:27:36.160Z",
      "content": "<p>before any data cleaning work, you should verify its effect.</p>\n\n<p>(1) how much data are noisy? you should random sub sample and do manual counting and get some statistics.</p>\n\n<p>e.g. 5% of all data are noisy.</p>\n\n<hr>\n\n<p>(2) if i can clean some label, how much improvement?</p>\n\n<p>e.g. if i hand correct 1%, improvement is xx%?</p>\n\n<p>e.g. if i hand correct 2%, improvement is xx%?</p>\n\n<hr>\n\n<p>without these basic experiment results, it is difficult to judge how much to clean and what to clean.</p>\n\n<p>you can also just take only the recognized images and flip the label synthetically to do experiments to find the relationship of noise vs accuracy.</p>",
      "rawMarkdown": "before any data cleaning work, you should verify its effect.\n\n(1) how much data are noisy? you should random sub sample and do manual counting and get some statistics.\n \ne.g. 5% of all data are noisy.\n\n---\n\n(2) if i can clean some label, how much improvement?\n     \ne.g. if i hand correct 1%, improvement is xx%?\n\ne.g. if i hand correct 2%, improvement is xx%?\n\n\n---\n\nwithout these basic experiment results, it is difficult to judge how much to clean and what to clean.\n\n\nyou can also just take only the recognized images and flip the label synthetically to do experiments to find the relationship of noise vs accuracy.\n",
      "votes": 4
    },
    {
      "id": 425164,
      "postDate": "2018-11-21T08:18:08.953Z",
      "content": "<p>In this competition, the dataset can be divided into two parts: successful recognition and failed recognition.</p>\n\n<p>Only pictures with label ‘successful recognition’ counts. If we need to clean the data and reduce the noise, we can delete all data with ‘failed recognition’. However, in binary classification, we have these four conditions:</p>\n\n<p>(1) True positive: Drawer draw a correct doodle, and google mark it as ‘successful recognition’.\n(2) True negative: Drawer draw a wrong doodle, and google mark it as ‘failed recognition’.\n(3) False positive: Drawer draw a wrong doodle, but google mark it as ‘successful recognition’.\n(4) False negative: Drawer draw a correct doodle, but google mark it as ‘failed recognition’.</p>\n\n<p>The noise data contains 2 and 3, but if we delete all ‘failed recognition’ data, we delete 2 and 4. It’s hard to say that new data will perform better than original data.</p>\n\n<p>However, we can assume that the proportion of noise in ‘failed recognition’ is much larger than that in ‘successful recognition’ data.</p>\n\n<p>Here, I have an idea of using ‘average doodle’, for example: we have 500 doodles of train-simplified cats, \nFirst, we transfer the (x,y) form data to numpy-picture data. \nSecond, we add all pixels up and calculate the average value, and we will get an average numpy-picture data as ‘average doodle’ based on 500 doodles. \nThen, we calculate the Euclidean distance between each doodle picture and ‘average doodle’ picture.\nAt last, we can delete all data whose distance is exceed than a threshold.</p>\n\n<p>The principle of this method is simple: If your doodle has more difference than other’s doodles, you doodle is wrong.</p>\n\n<p>Perhaps you will ask: how can you prove that your method is effective? In other words, will your new dataset perform better than original data?</p>\n\n<p>We have reached the assumption that ‘the proportion of noise in ‘failed recognition’ is much larger than that in ‘successful recognition’ data.’ Thus, if we can prove that the proportion of ‘False’ in dataframe reduce, we can prove that.</p>\n\n<p>In fact, I have established a kernel to show my method, and the result is:</p>\n\n<p>success num: 261</p>\n\n<p>fail num: 79</p>\n\n<p>You can find the kernel <a href=\"https://www.kaggle.com/dilapsky/an-elementary-method-of-data-clean?scriptVersionId=7602883\">Here</a>\nand thanks beluga for his public kernel ‘GreyScale Mobilenet’<a href=\"https://www.kaggle.com/gaborfodor/greyscale-mobilenet-lb-0-892\">Here</a>, because I use part of his code in it</p>\n\n<p>However, this is only my trail of data clean. Perhaps it will be useful, perhaps it is only nonsense. Well, what do you think of this method? Welcome to publish your opinion!</p>",
      "rawMarkdown": "\nIn this competition, the dataset can be divided into two parts: successful recognition and failed recognition.\n\nOnly pictures with label ‘successful recognition’ counts. If we need to clean the data and reduce the noise, we can delete all data with ‘failed recognition’. However, in binary classification, we have these four conditions:\n\n(1)\tTrue positive: Drawer draw a correct doodle, and google mark it as ‘successful recognition’.\n(2)\tTrue negative: Drawer draw a wrong doodle, and google mark it as ‘failed recognition’.\n(3)\tFalse positive: Drawer draw a wrong doodle, but google mark it as ‘successful recognition’.\n(4)\tFalse negative: Drawer draw a correct doodle, but google mark it as ‘failed recognition’.\n\nThe noise data contains 2 and 3, but if we delete all ‘failed recognition’ data, we delete 2 and 4. It’s hard to say that new data will perform better than original data.\n\nHowever, we can assume that the proportion of noise in ‘failed recognition’ is much larger than that in ‘successful recognition’ data.\n\nHere, I have an idea of using ‘average doodle’, for example: we have 500 doodles of train-simplified cats, \nFirst, we transfer the (x,y) form data to numpy-picture data. \nSecond, we add all pixels up and calculate the average value, and we will get an average numpy-picture data as ‘average doodle’ based on 500 doodles. \nThen, we calculate the Euclidean distance between each doodle picture and ‘average doodle’ picture.\nAt last, we can delete all data whose distance is exceed than a threshold.\n\nThe principle of this method is simple: If your doodle has more difference than other’s doodles, you doodle is wrong.\n\nPerhaps you will ask: how can you prove that your method is effective? In other words, will your new dataset perform better than original data?\n\nWe have reached the assumption that ‘the proportion of noise in ‘failed recognition’ is much larger than that in ‘successful recognition’ data.’ Thus, if we can prove that the proportion of ‘False’ in dataframe reduce, we can prove that.\n\nIn fact, I have established a kernel to show my method, and the result is:\n\nsuccess num: 261\n\nfail num: 79\n\nYou can find the kernel [Here][1]\nand thanks beluga for his public kernel ‘GreyScale Mobilenet’[Here][2], because I use part of his code in it\n\nHowever, this is only my trail of data clean. Perhaps it will be useful, perhaps it is only nonsense. Well, what do you think of this method? Welcome to publish your opinion!\n\n\n  [1]: https://www.kaggle.com/dilapsky/an-elementary-method-of-data-clean?scriptVersionId=7602883\n  [2]: https://www.kaggle.com/gaborfodor/greyscale-mobilenet-lb-0-892",
      "votes": 1
    },
    {
      "id": 426465,
      "postDate": "2018-11-23T09:46:29.423Z",
      "content": "<p>I tried to do data augmentation today, but the dataset is so large that my data augmentation program crushed......Perhaps the best way is using the raw dataset...........</p>",
      "rawMarkdown": "I tried to do data augmentation today, but the dataset is so large that my data augmentation program crushed......Perhaps the best way is using the raw dataset..........."
    },
    {
      "id": 425331,
      "postDate": "2018-11-21T13:04:47.400Z",
      "content": "<p>How about just excluding images that are too crappy? </p>",
      "rawMarkdown": "How about just excluding images that are too crappy? "
    },
    {
      "id": 425266,
      "postDate": "2018-11-21T11:23:03.757Z",
      "content": "<p>Recognize noise label correctly can improve model performance significantly.</p>\n\n<p>But normal strategy should be as follows:</p>\n\n<p>Use a CNN classifier to distinguish the noise sample and correct sample with thresholding operation on unrecognized image will be more precise than just use the distance between sample image and average image.</p>",
      "rawMarkdown": "Recognize noise label correctly can improve model performance significantly.\n\nBut normal strategy should be as follows:\n \nUse a CNN classifier to distinguish the noise sample and correct sample with thresholding operation on unrecognized image will be more precise than just use the distance between sample image and average image.\n",
      "replies": [
        {
          "id": 425746,
          "postDate": "2018-11-22T03:31:37.997Z",
          "content": "<p>Using a CNN classifier to distinguish the noise sample and correct sample.\nBut our job is using a CNN classifier to distinguish the type of sample.\nThen they will be the same thing, right?\nOr, BTW, are there any pre-trained model of  distinguishing the noise sample?</p>",
          "rawMarkdown": " Using a CNN classifier to distinguish the noise sample and correct sample.\nBut our job is using a CNN classifier to distinguish the type of sample.\nThen they will be the same thing, right?\nOr, BTW, are there any pre-trained model of  distinguishing the noise sample?"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 425269,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-11-21T11:27:36.160000",
      "content": "<p>before any data cleaning work, you should verify its effect.</p>\n\n<p>(1) how much data are noisy? you should random sub sample and do manual counting and get some statistics.</p>\n\n<p>e.g. 5% of all data are noisy.</p>\n\n<hr>\n\n<p>(2) if i can clean some label, how much improvement?</p>\n\n<p>e.g. if i hand correct 1%, improvement is xx%?</p>\n\n<p>e.g. if i hand correct 2%, improvement is xx%?</p>\n\n<hr>\n\n<p>without these basic experiment results, it is difficult to judge how much to clean and what to clean.</p>\n\n<p>you can also just take only the recognized images and flip the label synthetically to do experiments to find the relationship of noise vs accuracy.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 426465,
      "author_name": "Dilapsky Lee",
      "author_url": "",
      "post_date": "2018-11-23T09:46:29.423000",
      "content": "<p>I tried to do data augmentation today, but the dataset is so large that my data augmentation program crushed......Perhaps the best way is using the raw dataset...........</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 425331,
      "author_name": "HuyenNguyen",
      "author_url": "",
      "post_date": "2018-11-21T13:04:47.400000",
      "content": "<p>How about just excluding images that are too crappy? </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 425266,
      "author_name": "wh1te",
      "author_url": "",
      "post_date": "2018-11-21T11:23:03.757000",
      "content": "<p>Recognize noise label correctly can improve model performance significantly.</p>\n\n<p>But normal strategy should be as follows:</p>\n\n<p>Use a CNN classifier to distinguish the noise sample and correct sample with thresholding operation on unrecognized image will be more precise than just use the distance between sample image and average image.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 425746,
          "author_name": "Dilapsky Lee",
          "author_url": "",
          "post_date": "2018-11-22T03:31:37.997000",
          "content": "<p>Using a CNN classifier to distinguish the noise sample and correct sample.\nBut our job is using a CNN classifier to distinguish the type of sample.\nThen they will be the same thing, right?\nOr, BTW, are there any pre-trained model of  distinguishing the noise sample?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "425269": "before any data cleaning work, you should verify its effect.\n\n(1) how much data are noisy? you should random sub sample and do manual counting and get some statistics.\n \ne.g. 5% of all data are noisy.\n\n---\n\n(2) if i can clean some label, how much improvement?\n     \ne.g. if i hand correct 1%, improvement is xx%?\n\ne.g. if i hand correct 2%, improvement is xx%?\n\n\n---\n\nwithout these basic experiment results, it is difficult to judge how much to clean and what to clean.\n\n\nyou can also just take only the recognized images and flip the label synthetically to do experiments to find the relationship of noise vs accuracy.\n",
    "425164": "\nIn this competition, the dataset can be divided into two parts: successful recognition and failed recognition.\n\nOnly pictures with label ‘successful recognition’ counts. If we need to clean the data and reduce the noise, we can delete all data with ‘failed recognition’. However, in binary classification, we have these four conditions:\n\n(1)\tTrue positive: Drawer draw a correct doodle, and google mark it as ‘successful recognition’.\n(2)\tTrue negative: Drawer draw a wrong doodle, and google mark it as ‘failed recognition’.\n(3)\tFalse positive: Drawer draw a wrong doodle, but google mark it as ‘successful recognition’.\n(4)\tFalse negative: Drawer draw a correct doodle, but google mark it as ‘failed recognition’.\n\nThe noise data contains 2 and 3, but if we delete all ‘failed recognition’ data, we delete 2 and 4. It’s hard to say that new data will perform better than original data.\n\nHowever, we can assume that the proportion of noise in ‘failed recognition’ is much larger than that in ‘successful recognition’ data.\n\nHere, I have an idea of using ‘average doodle’, for example: we have 500 doodles of train-simplified cats, \nFirst, we transfer the (x,y) form data to numpy-picture data. \nSecond, we add all pixels up and calculate the average value, and we will get an average numpy-picture data as ‘average doodle’ based on 500 doodles. \nThen, we calculate the Euclidean distance between each doodle picture and ‘average doodle’ picture.\nAt last, we can delete all data whose distance is exceed than a threshold.\n\nThe principle of this method is simple: If your doodle has more difference than other’s doodles, you doodle is wrong.\n\nPerhaps you will ask: how can you prove that your method is effective? In other words, will your new dataset perform better than original data?\n\nWe have reached the assumption that ‘the proportion of noise in ‘failed recognition’ is much larger than that in ‘successful recognition’ data.’ Thus, if we can prove that the proportion of ‘False’ in dataframe reduce, we can prove that.\n\nIn fact, I have established a kernel to show my method, and the result is:\n\nsuccess num: 261\n\nfail num: 79\n\nYou can find the kernel [Here][1]\nand thanks beluga for his public kernel ‘GreyScale Mobilenet’[Here][2], because I use part of his code in it\n\nHowever, this is only my trail of data clean. Perhaps it will be useful, perhaps it is only nonsense. Well, what do you think of this method? Welcome to publish your opinion!\n\n\n  [1]: https://www.kaggle.com/dilapsky/an-elementary-method-of-data-clean?scriptVersionId=7602883\n  [2]: https://www.kaggle.com/gaborfodor/greyscale-mobilenet-lb-0-892",
    "426465": "I tried to do data augmentation today, but the dataset is so large that my data augmentation program crushed......Perhaps the best way is using the raw dataset...........",
    "425331": "How about just excluding images that are too crappy? ",
    "425266": "Recognize noise label correctly can improve model performance significantly.\n\nBut normal strategy should be as follows:\n \nUse a CNN classifier to distinguish the noise sample and correct sample with thresholding operation on unrecognized image will be more precise than just use the distance between sample image and average image.\n"
  }
}