{
  "id": 70540,
  "title": "cake v birthday_cake",
  "url": "/competitions/quickdraw-doodle-recognition/discussion/70540",
  "author_name": "raifer",
  "post_date": "2018-11-05T04:27:11.440000",
  "votes": 7,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Has anyone computed something like a confusion matrix? I just noticed that most of the training cases in 'cake' look like birthday cakes. I suspect those classes are nearly statistically the same.</p>",
  "messages": [
    {
      "id": 415414,
      "postDate": "2018-11-05T04:27:11.440Z",
      "content": "<p>Has anyone computed something like a confusion matrix? I just noticed that most of the training cases in 'cake' look like birthday cakes. I suspect those classes are nearly statistically the same.</p>",
      "rawMarkdown": "Has anyone computed something like a confusion matrix? I just noticed that most of the training cases in 'cake' look like birthday cakes. I suspect those classes are nearly statistically the same.",
      "votes": 7
    },
    {
      "id": 423448,
      "postDate": "2018-11-18T09:08:11.517Z",
      "content": "<p>Based on my current model, the following are the categories which are confused the most (confusion &gt;20%):</p>\n\n<pre><code>bicycle vs. motorbike: 0.320\nbirthday cake vs. cake: 0.554\nbucket vs. paint can: 0.221\nbus vs. school bus: 0.461\ncello vs. violin: 0.309\ncell phone vs. telephone: 0.210\ncoffee cup vs. cup: 0.218\ncoffee cup vs. mug: 0.334\ncup vs. mug: 0.259\nface vs. smiley face: 0.213\ngolf club vs. hockey stick: 0.245\nguitar vs. violin: 0.303\nhexagon vs. octagon: 0.347\nhurricane vs. tornado: 0.379\njacket vs. sweater: 0.250\npickup truck vs. truck: 0.319\nradio vs. stereo: 0.302\ntrombone vs. trumpet: 0.247\n</code></pre>",
      "rawMarkdown": "Based on my current model, the following are the categories which are confused the most (confusion &gt;20%):\n\n    bicycle vs. motorbike: 0.320\n    birthday cake vs. cake: 0.554\n    bucket vs. paint can: 0.221\n    bus vs. school bus: 0.461\n    cello vs. violin: 0.309\n    cell phone vs. telephone: 0.210\n    coffee cup vs. cup: 0.218\n    coffee cup vs. mug: 0.334\n    cup vs. mug: 0.259\n    face vs. smiley face: 0.213\n    golf club vs. hockey stick: 0.245\n    guitar vs. violin: 0.303\n    hexagon vs. octagon: 0.347\n    hurricane vs. tornado: 0.379\n    jacket vs. sweater: 0.250\n    pickup truck vs. truck: 0.319\n    radio vs. stereo: 0.302\n    trombone vs. trumpet: 0.247",
      "votes": 4,
      "replies": [
        {
          "id": 423707,
          "postDate": "2018-11-18T22:46:43.700Z",
          "content": "<p>Just wondering how do you calculate the confusion?</p>",
          "rawMarkdown": "Just wondering how do you calculate the confusion?"
        },
        {
          "id": 424289,
          "postDate": "2018-11-19T21:03:03.843Z",
          "content": "<p>In PyTorch, something like the following:</p>\n\n<pre><code>confusion = np.zeros((num_categories, num_categories), dtype=np.float32)\n\nfor images, categories in val_data_loader:\n    predictions = F.softmax(model(images), dim=1)\n    _, prediction_categories = predictions.topk(3, dim=1, sorted=True)\n\n    # only the first prediction (out of 3) is considered for the confusion matrix\n    for bpc, bc in zip(prediction_categories[:, 0], categories):\n        confusion[bpc, bc] += 1\n\n# normalize the confusion matrix to get percentages\nfor c in range(confusion.shape[0]):\n    category_count = confusion[c, :].sum()\n    if category_count != 0:\n        confusion[c, :] /= category_count\n</code></pre>\n\n<p>There might be other approaches for calculating the confusion matrix (like considering all 3 predictions) but the code above suited my needs.</p>",
          "rawMarkdown": "In PyTorch, something like the following:\n\n    confusion = np.zeros((num_categories, num_categories), dtype=np.float32)\n\n    for images, categories in val_data_loader:\n        predictions = F.softmax(model(images), dim=1)\n        _, prediction_categories = predictions.topk(3, dim=1, sorted=True)\n\n        # only the first prediction (out of 3) is considered for the confusion matrix\n        for bpc, bc in zip(prediction_categories[:, 0], categories):\n            confusion[bpc, bc] += 1\n\n    # normalize the confusion matrix to get percentages\n    for c in range(confusion.shape[0]):\n        category_count = confusion[c, :].sum()\n        if category_count != 0:\n            confusion[c, :] /= category_count\n\nThere might be other approaches for calculating the confusion matrix (like considering all 3 predictions) but the code above suited my needs.",
          "votes": 2
        }
      ]
    },
    {
      "id": 427069,
      "postDate": "2018-11-24T13:46:07.540Z",
      "content": "<p>Everyone draws a birthday cake when asked to 'draw a cake'. It seems impossible that any program could correctly recognize any birthday cake as a cake. I suspect that a few of the categories in this competition are artificially biased to QuickDraw's view of the world.\nGiven the non-independent categories, maybe the original app could have been specific: \n\"please draw a cake (not a birthday cake)\"\n\"please draw a bus (not a school bus)\"</p>",
      "rawMarkdown": "Everyone draws a birthday cake when asked to 'draw a cake'. It seems impossible that any program could correctly recognize any birthday cake as a cake. I suspect that a few of the categories in this competition are artificially biased to QuickDraw's view of the world.\nGiven the non-independent categories, maybe the original app could have been specific: \n\"please draw a cake (not a birthday cake)\"\n\"please draw a bus (not a school bus)\"",
      "votes": 1
    },
    {
      "id": 417080,
      "postDate": "2018-11-07T17:54:25.947Z",
      "content": "<p>check the octagon/hexagon classes. The confusion can be solved by some specialized method</p>",
      "rawMarkdown": "check the octagon/hexagon classes. The confusion can be solved by some specialized method",
      "votes": 1
    },
    {
      "id": 416271,
      "postDate": "2018-11-06T14:00:38.317Z",
      "content": "<p>base on my validation results,  if i have a public LB of say 0.945, it means my top 1 and top 3 accuracy could be about 0.90 and 0.98?</p>\n\n<p>this means that if i can reorder about 10% of my submission, i can improve my LB substantially.  </p>\n\n<p>i just have to think of way to pick up the most probably error images and think of ways to correct them.</p>\n\n<p>maybe we can build specialized predictor for the error case like cup/coffee cup etc</p>\n\n<hr>\n\n<p>alternatively, i can think of way to improve the LB via detecting the missing 2% (false negative)</p>",
      "rawMarkdown": "base on my validation results,  if i have a public LB of say 0.945, it means my top 1 and top 3 accuracy could be about 0.90 and 0.98?\n\nthis means that if i can reorder about 10% of my submission, i can improve my LB substantially.  \n\ni just have to think of way to pick up the most probably error images and think of ways to correct them.\n\nmaybe we can build specialized predictor for the error case like cup/coffee cup etc\n\n---\n\nalternatively, i can think of way to improve the LB via detecting the missing 2% (false negative)",
      "votes": 2,
      "replies": [
        {
          "id": 417332,
          "postDate": "2018-11-08T05:24:29.003Z",
          "content": "<p>Hi, I have the same idea. train another specialized classifier for those error-like-group classes. But there is almost no improvement. Can you tell me how much improvement can you gain? Thx</p>",
          "rawMarkdown": "Hi, I have the same idea. train another specialized classifier for those error-like-group classes. But there is almost no improvement. Can you tell me how much improvement can you gain? Thx",
          "votes": 1
        }
      ]
    },
    {
      "id": 415817,
      "postDate": "2018-11-05T18:47:05.787Z",
      "content": "<p>there are similar problems with cup/coffee cup, bus/school bus, hurricane/tornado</p>",
      "rawMarkdown": "there are similar problems with cup/coffee cup, bus/school bus, hurricane/tornado"
    },
    {
      "id": 416307,
      "postDate": "2018-11-06T14:24:11.520Z",
      "content": "<p>there seems to be a difference for the confusion case?</p>\n\n<p>check the mean no. of stroke at:</p>\n\n<p><a href=\"http://rstudio-pubs-static.s3.amazonaws.com/292508_8ef4c9ec5f76421e92803fccba9765df.html\">http://rstudio-pubs-static.s3.amazonaws.com/292508_8ef4c9ec5f76421e92803fccba9765df.html</a></p>",
      "rawMarkdown": "there seems to be a difference for the confusion case?\n\ncheck the mean no. of stroke at:\n\nhttp://rstudio-pubs-static.s3.amazonaws.com/292508_8ef4c9ec5f76421e92803fccba9765df.html"
    },
    {
      "id": 416136,
      "postDate": "2018-11-06T08:44:54.560Z",
      "content": "<p>May be you should check other column to ensure which label should be assigned to :)</p>",
      "rawMarkdown": "May be you should check other column to ensure which label should be assigned to :)"
    },
    {
      "id": 416021,
      "postDate": "2018-11-06T03:44:49.543Z",
      "content": "<p>Looking at google's failure rates, I would suggest the following (when in doubt): \"school bus\" over \"bus\", \"birthday cake\" over \"cake\", but then \"cup\" over \"coffee cup\"</p>",
      "rawMarkdown": "Looking at google's failure rates, I would suggest the following (when in doubt): \"school bus\" over \"bus\", \"birthday cake\" over \"cake\", but then \"cup\" over \"coffee cup\"",
      "replies": [
        {
          "id": 416772,
          "postDate": "2018-11-07T09:00:25.580Z",
          "content": "<p>this is correct. there are 330 items per test class. for example, square, sun, paperclip, star are the most confident classes. their item count is 330.</p>\n\n<p>cup and birthday cake are under populated, and incidentally, coffee mug and cake are over-populated .</p>\n\n<p>we need a ranking-based predictor for these cases!</p>\n\n<p>most predicted coffee_cup are \"coffee_cup, mug, cup\", you can flip the order to try your luck :)</p>\n\n<hr>\n\n<p>a submission of cup only  ('cup cup cup') has LB score of 0.002. kaggle score is truncation.\nso number of cups in public LB set is:</p>\n\n<p>0.002*112199 = 224.398</p>\n\n<p>to</p>\n\n<p>0.00299999*112199 = 336.59587801000004</p>",
          "rawMarkdown": "this is correct. there are 330 items per test class. for example, square, sun, paperclip, star are the most confident classes. their item count is 330.\n\ncup and birthday cake are under populated, and incidentally, coffee mug and cake are over-populated .\n\nwe need a ranking-based predictor for these cases!\n\nmost predicted coffee_cup are \"coffee_cup, mug, cup\", you can flip the order to try your luck :)\n\n---\na submission of cup only  ('cup cup cup') has LB score of 0.002. kaggle score is truncation.\nso number of cups in public LB set is:\n\n\n0.002*112199 = 224.398\n\n\nto\n\n0.00299999*112199 = 336.59587801000004\n",
          "votes": 5
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 423448,
      "author_name": "omallo",
      "author_url": "",
      "post_date": "2018-11-18T09:08:11.517000",
      "content": "<p>Based on my current model, the following are the categories which are confused the most (confusion &gt;20%):</p>\n\n<pre><code>bicycle vs. motorbike: 0.320\nbirthday cake vs. cake: 0.554\nbucket vs. paint can: 0.221\nbus vs. school bus: 0.461\ncello vs. violin: 0.309\ncell phone vs. telephone: 0.210\ncoffee cup vs. cup: 0.218\ncoffee cup vs. mug: 0.334\ncup vs. mug: 0.259\nface vs. smiley face: 0.213\ngolf club vs. hockey stick: 0.245\nguitar vs. violin: 0.303\nhexagon vs. octagon: 0.347\nhurricane vs. tornado: 0.379\njacket vs. sweater: 0.250\npickup truck vs. truck: 0.319\nradio vs. stereo: 0.302\ntrombone vs. trumpet: 0.247\n</code></pre>",
      "votes": 4,
      "replies": [
        {
          "id": 423707,
          "author_name": "Strideradu",
          "author_url": "",
          "post_date": "2018-11-18T22:46:43.700000",
          "content": "<p>Just wondering how do you calculate the confusion?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 424289,
          "author_name": "omallo",
          "author_url": "",
          "post_date": "2018-11-19T21:03:03.843000",
          "content": "<p>In PyTorch, something like the following:</p>\n\n<pre><code>confusion = np.zeros((num_categories, num_categories), dtype=np.float32)\n\nfor images, categories in val_data_loader:\n    predictions = F.softmax(model(images), dim=1)\n    _, prediction_categories = predictions.topk(3, dim=1, sorted=True)\n\n    # only the first prediction (out of 3) is considered for the confusion matrix\n    for bpc, bc in zip(prediction_categories[:, 0], categories):\n        confusion[bpc, bc] += 1\n\n# normalize the confusion matrix to get percentages\nfor c in range(confusion.shape[0]):\n    category_count = confusion[c, :].sum()\n    if category_count != 0:\n        confusion[c, :] /= category_count\n</code></pre>\n\n<p>There might be other approaches for calculating the confusion matrix (like considering all 3 predictions) but the code above suited my needs.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 427069,
      "author_name": "raifer",
      "author_url": "",
      "post_date": "2018-11-24T13:46:07.540000",
      "content": "<p>Everyone draws a birthday cake when asked to 'draw a cake'. It seems impossible that any program could correctly recognize any birthday cake as a cake. I suspect that a few of the categories in this competition are artificially biased to QuickDraw's view of the world.\nGiven the non-independent categories, maybe the original app could have been specific: \n\"please draw a cake (not a birthday cake)\"\n\"please draw a bus (not a school bus)\"</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 417080,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-11-07T17:54:25.947000",
      "content": "<p>check the octagon/hexagon classes. The confusion can be solved by some specialized method</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 416271,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-11-06T14:00:38.317000",
      "content": "<p>base on my validation results,  if i have a public LB of say 0.945, it means my top 1 and top 3 accuracy could be about 0.90 and 0.98?</p>\n\n<p>this means that if i can reorder about 10% of my submission, i can improve my LB substantially.  </p>\n\n<p>i just have to think of way to pick up the most probably error images and think of ways to correct them.</p>\n\n<p>maybe we can build specialized predictor for the error case like cup/coffee cup etc</p>\n\n<hr>\n\n<p>alternatively, i can think of way to improve the LB via detecting the missing 2% (false negative)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 417332,
          "author_name": "good good study",
          "author_url": "",
          "post_date": "2018-11-08T05:24:29.003000",
          "content": "<p>Hi, I have the same idea. train another specialized classifier for those error-like-group classes. But there is almost no improvement. Can you tell me how much improvement can you gain? Thx</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 415817,
      "author_name": "Miha Skalic",
      "author_url": "",
      "post_date": "2018-11-05T18:47:05.787000",
      "content": "<p>there are similar problems with cup/coffee cup, bus/school bus, hurricane/tornado</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 416307,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-11-06T14:24:11.520000",
      "content": "<p>there seems to be a difference for the confusion case?</p>\n\n<p>check the mean no. of stroke at:</p>\n\n<p><a href=\"http://rstudio-pubs-static.s3.amazonaws.com/292508_8ef4c9ec5f76421e92803fccba9765df.html\">http://rstudio-pubs-static.s3.amazonaws.com/292508_8ef4c9ec5f76421e92803fccba9765df.html</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 416136,
      "author_name": "wh1te",
      "author_url": "",
      "post_date": "2018-11-06T08:44:54.560000",
      "content": "<p>May be you should check other column to ensure which label should be assigned to :)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 416021,
      "author_name": "raifer",
      "author_url": "",
      "post_date": "2018-11-06T03:44:49.543000",
      "content": "<p>Looking at google's failure rates, I would suggest the following (when in doubt): \"school bus\" over \"bus\", \"birthday cake\" over \"cake\", but then \"cup\" over \"coffee cup\"</p>",
      "votes": 0,
      "replies": [
        {
          "id": 416772,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2018-11-07T09:00:25.580000",
          "content": "<p>this is correct. there are 330 items per test class. for example, square, sun, paperclip, star are the most confident classes. their item count is 330.</p>\n\n<p>cup and birthday cake are under populated, and incidentally, coffee mug and cake are over-populated .</p>\n\n<p>we need a ranking-based predictor for these cases!</p>\n\n<p>most predicted coffee_cup are \"coffee_cup, mug, cup\", you can flip the order to try your luck :)</p>\n\n<hr>\n\n<p>a submission of cup only  ('cup cup cup') has LB score of 0.002. kaggle score is truncation.\nso number of cups in public LB set is:</p>\n\n<p>0.002*112199 = 224.398</p>\n\n<p>to</p>\n\n<p>0.00299999*112199 = 336.59587801000004</p>",
          "votes": 5,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "415414": "Has anyone computed something like a confusion matrix? I just noticed that most of the training cases in 'cake' look like birthday cakes. I suspect those classes are nearly statistically the same.",
    "423448": "Based on my current model, the following are the categories which are confused the most (confusion &gt;20%):\n\n    bicycle vs. motorbike: 0.320\n    birthday cake vs. cake: 0.554\n    bucket vs. paint can: 0.221\n    bus vs. school bus: 0.461\n    cello vs. violin: 0.309\n    cell phone vs. telephone: 0.210\n    coffee cup vs. cup: 0.218\n    coffee cup vs. mug: 0.334\n    cup vs. mug: 0.259\n    face vs. smiley face: 0.213\n    golf club vs. hockey stick: 0.245\n    guitar vs. violin: 0.303\n    hexagon vs. octagon: 0.347\n    hurricane vs. tornado: 0.379\n    jacket vs. sweater: 0.250\n    pickup truck vs. truck: 0.319\n    radio vs. stereo: 0.302\n    trombone vs. trumpet: 0.247",
    "427069": "Everyone draws a birthday cake when asked to 'draw a cake'. It seems impossible that any program could correctly recognize any birthday cake as a cake. I suspect that a few of the categories in this competition are artificially biased to QuickDraw's view of the world.\nGiven the non-independent categories, maybe the original app could have been specific: \n\"please draw a cake (not a birthday cake)\"\n\"please draw a bus (not a school bus)\"",
    "417080": "check the octagon/hexagon classes. The confusion can be solved by some specialized method",
    "416271": "base on my validation results,  if i have a public LB of say 0.945, it means my top 1 and top 3 accuracy could be about 0.90 and 0.98?\n\nthis means that if i can reorder about 10% of my submission, i can improve my LB substantially.  \n\ni just have to think of way to pick up the most probably error images and think of ways to correct them.\n\nmaybe we can build specialized predictor for the error case like cup/coffee cup etc\n\n---\n\nalternatively, i can think of way to improve the LB via detecting the missing 2% (false negative)",
    "415817": "there are similar problems with cup/coffee cup, bus/school bus, hurricane/tornado",
    "416307": "there seems to be a difference for the confusion case?\n\ncheck the mean no. of stroke at:\n\nhttp://rstudio-pubs-static.s3.amazonaws.com/292508_8ef4c9ec5f76421e92803fccba9765df.html",
    "416136": "May be you should check other column to ensure which label should be assigned to :)",
    "416021": "Looking at google's failure rates, I would suggest the following (when in doubt): \"school bus\" over \"bus\", \"birthday cake\" over \"cake\", but then \"cup\" over \"coffee cup\""
  }
}