{
  "id": 73803,
  "title": "Predictions balancing: source code",
  "url": "/competitions/quickdraw-doodle-recognition/discussion/73803",
  "author_name": "Pavel Ostyakov",
  "post_date": "2018-12-05T16:52:34.264000",
  "votes": 76,
  "comment_count": 18,
  "views": 0,
  "content": "<p>Hey guys,</p>\n\n<p>I've just published my code for balancing predictions: <a href=\"https://github.com/PavelOstyakov/predictions_balancing\">https://github.com/PavelOstyakov/predictions_balancing</a></p>\n\n<p>It gives ~0.7% boost for any submission. Feel free to use it in future competitions :)</p>\n\n<p><img src=\"https://sd.keepcalm-o-matic.co.uk/i-w600/keep-calm-and-balance-your-predictions.jpg\" alt=\"enter image description here\"></p>",
  "messages": [
    {
      "id": 433924,
      "postDate": "2018-12-05T16:52:34.263Z",
      "content": "<p>Hey guys,</p>\n\n<p>I've just published my code for balancing predictions: <a href=\"https://github.com/PavelOstyakov/predictions_balancing\">https://github.com/PavelOstyakov/predictions_balancing</a></p>\n\n<p>It gives ~0.7% boost for any submission. Feel free to use it in future competitions :)</p>\n\n<p><img src=\"https://sd.keepcalm-o-matic.co.uk/i-w600/keep-calm-and-balance-your-predictions.jpg\" alt=\"enter image description here\"></p>",
      "rawMarkdown": "Hey guys,\n\nI've just published my code for balancing predictions: https://github.com/PavelOstyakov/predictions_balancing\n\nIt gives ~0.7% boost for any submission. Feel free to use it in future competitions :)\n\n![enter image description here][1]\n\n\n  [1]: https://sd.keepcalm-o-matic.co.uk/i-w600/keep-calm-and-balance-your-predictions.jpg",
      "votes": 76
    },
    {
      "id": 434108,
      "postDate": "2018-12-05T22:41:43.257Z",
      "content": "<p>thanks for the code!</p>\n\n<p>any idea what keyword i should use for goggle search for similar methods? </p>",
      "rawMarkdown": "thanks for the code!\n\nany idea what keyword i should use for goggle search for similar methods? ",
      "votes": 3
    },
    {
      "id": 434505,
      "postDate": "2018-12-06T14:14:34.537Z",
      "content": "<p>Hi,\nI am new to competitive data science and Kaggle. Could anyone explain what balancing of predictions mean?</p>",
      "rawMarkdown": "Hi,\nI am new to competitive data science and Kaggle. Could anyone explain what balancing of predictions mean?",
      "votes": 4,
      "replies": [
        {
          "id": 435494,
          "postDate": "2018-12-08T05:34:20.063Z",
          "content": "<p>from <a href=\"https://www.kaggle.com/c/quickdraw-doodle-recognition/discussion/73738\">https://www.kaggle.com/c/quickdraw-doodle-recognition/discussion/73738</a></p>\n\n<blockquote>\n  <p>The algorithm behind postprocessing is the following: for the most popular class decrease all the  probabilities iteratively by the same small value until it is no longer the most popular, repeat this procedure until all classes become equal. This technique was also used in one of the previous competitions (see github link for the code).</p>\n</blockquote>\n\n<p>This is the implementation of the explanation above. \nA variable alpha is the one he mentioned as \"the same small value\". In this case, they use 0.001 as initial alpha and iteratively halve the value when the score no longer improves until it gets 0.0001.</p>\n\n<p>Without balancing, for example if you directly use the prediction of your model, one class might be presented a few times more than minor classes. This results in worse score if the distribution of your prediction much differs from the class distribution of test set. To counter this, you can balance the probability of your prediction so that all classes are equally presented.</p>\n\n<p>This kind of balancing works when you know the distribution of test set as a prior knowledge which is the case for this competition. According to inversion, the claases were actually not evenly balanced but I think it's still fairly balanced and that's why it worked very well.</p>",
          "rawMarkdown": "from https://www.kaggle.com/c/quickdraw-doodle-recognition/discussion/73738\n\n&gt; The algorithm behind postprocessing is the following: for the most popular class decrease all the  probabilities iteratively by the same small value until it is no longer the most popular, repeat this procedure until all classes become equal. This technique was also used in one of the previous competitions (see github link for the code).\n\nThis is the implementation of the explanation above. \nA variable alpha is the one he mentioned as \"the same small value\". In this case, they use 0.001 as initial alpha and iteratively halve the value when the score no longer improves until it gets 0.0001.\n\nWithout balancing, for example if you directly use the prediction of your model, one class might be presented a few times more than minor classes. This results in worse score if the distribution of your prediction much differs from the class distribution of test set. To counter this, you can balance the probability of your prediction so that all classes are equally presented.\n\nThis kind of balancing works when you know the distribution of test set as a prior knowledge which is the case for this competition. According to inversion, the claases were actually not evenly balanced but I think it's still fairly balanced and that's why it worked very well.\n",
          "votes": 6
        },
        {
          "id": 435646,
          "postDate": "2018-12-08T13:04:20.973Z",
          "content": "<p>Thanks for the nice explanation. By \"all classes become equal\" does it mean that we know all classes are balanced (i.e. equal probabilities)? So if in our prior knowledge the classes are not balanced, we also tune the output distributions to match this unbalanced distribution? </p>",
          "rawMarkdown": "Thanks for the nice explanation. By \"all classes become equal\" does it mean that we know all classes are balanced (i.e. equal probabilities)? So if in our prior knowledge the classes are not balanced, we also tune the output distributions to match this unbalanced distribution? ",
          "votes": 1
        },
        {
          "id": 435906,
          "postDate": "2018-12-09T03:27:36.423Z",
          "content": "<p>Yes. Modifying a few lines such as torch.ones() which address equally balanced label distribution to something else should achive what you want.</p>",
          "rawMarkdown": "Yes. Modifying a few lines such as torch.ones() which address equally balanced label distribution to something else should achive what you want.",
          "votes": 2
        },
        {
          "id": 436499,
          "postDate": "2018-12-10T12:33:02.967Z",
          "content": "<p>Thanks a lot for that great explanation!!</p>",
          "rawMarkdown": "Thanks a lot for that great explanation!!\n\n",
          "votes": 1
        },
        {
          "id": 437915,
          "postDate": "2018-12-12T18:48:46.837Z",
          "content": "<p>thanks for the explanation, this is extremely helpful</p>",
          "rawMarkdown": "thanks for the explanation, this is extremely helpful"
        }
      ]
    },
    {
      "id": 434158,
      "postDate": "2018-12-06T01:24:19.117Z",
      "content": "<p>Awesome!! After applying this algorithm, my lb score boosts from <strong>Private LB 0.94974</strong> to <strong>0.95467</strong>.\nThanks for the code! It will be very useful for the future competitions.</p>",
      "rawMarkdown": "Awesome!! After applying this algorithm, my lb score boosts from **Private LB 0.94974** to **0.95467**.\nThanks for the code! It will be very useful for the future competitions.",
      "votes": 2,
      "replies": [
        {
          "id": 434357,
          "postDate": "2018-12-06T09:13:29.817Z",
          "content": "<p>You are welcome!</p>",
          "rawMarkdown": "You are welcome!"
        }
      ]
    },
    {
      "id": 434114,
      "postDate": "2018-12-05T22:50:53.877Z",
      "content": "<p>here is probably one related paper:\n\"Probabilistic n-Choose-k Models for Classification and Ranking\"</p>\n\n<p><a href=\"http://www.cs.toronto.edu/~kswersky/wp-content/uploads/pnck.pdf\">http://www.cs.toronto.edu/~kswersky/wp-content/uploads/pnck.pdf</a></p>\n\n<p>\"Instead, we focus on work related to the main novelty in this paper,\nthe explicit modeling of structure on label counts. That is, given that we have prior knowledge of\nlabel count structure, or are modeling a domain that exhibits such structure, the question is how can\nthe structure be leveraged to improve a model.\"</p>",
      "rawMarkdown": "here is probably one related paper:\n\"Probabilistic n-Choose-k Models for Classification and Ranking\"\n\nhttp://www.cs.toronto.edu/~kswersky/wp-content/uploads/pnck.pdf\n\n\"Instead, we focus on work related to the main novelty in this paper,\nthe explicit modeling of structure on label counts. That is, given that we have prior knowledge of\nlabel count structure, or are modeling a domain that exhibits such structure, the question is how can\nthe structure be leveraged to improve a model.\"",
      "votes": 2,
      "replies": [
        {
          "id": 434359,
          "postDate": "2018-12-06T09:13:54.677Z",
          "content": "<p>Thanks for sharing </p>",
          "rawMarkdown": "Thanks for sharing "
        }
      ]
    },
    {
      "id": 438270,
      "postDate": "2018-12-13T11:49:21.193Z",
      "content": "<p>GOOD JOB!</p>",
      "rawMarkdown": "GOOD JOB!"
    },
    {
      "id": 435233,
      "postDate": "2018-12-07T18:08:51.103Z",
      "content": "<p>For what it's worth, classes in the test set were not evenly balanced.  But, obviously, they were not highly skewed either. :-)</p>",
      "rawMarkdown": "For what it's worth, classes in the test set were not evenly balanced.  But, obviously, they were not highly skewed either. :-)",
      "replies": [
        {
          "id": 435236,
          "postDate": "2018-12-07T18:12:20.763Z",
          "content": "<p>That's interesting! Could I ask you to provide a distribution of classes in the test set? I may try to improve my script for this competition :)</p>",
          "rawMarkdown": "That's interesting! Could I ask you to provide a distribution of classes in the test set? I may try to improve my script for this competition :)"
        }
      ]
    },
    {
      "id": 434869,
      "postDate": "2018-12-07T04:47:04.233Z",
      "content": "<p>Amazing and extremely simple. I'm really wondering how it works? Thanks for your explanation in advance.</p>",
      "rawMarkdown": "Amazing and extremely simple. I'm really wondering how it works? Thanks for your explanation in advance."
    },
    {
      "id": 434523,
      "postDate": "2018-12-06T14:33:06.130Z",
      "content": "<p>good job!</p>",
      "rawMarkdown": "good job!"
    },
    {
      "id": 434194,
      "postDate": "2018-12-06T03:22:54.950Z",
      "rawMarkdown": "",
      "votes": 8,
      "isDeleted": true
    },
    {
      "id": 434199,
      "postDate": "2018-12-06T03:31:02.783Z",
      "content": "<p>Many Thanks Pavel!</p>",
      "rawMarkdown": "Many Thanks Pavel!",
      "votes": 7
    }
  ],
  "comments": [
    {
      "id": 434108,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-12-05T22:41:43.257000",
      "content": "<p>thanks for the code!</p>\n\n<p>any idea what keyword i should use for goggle search for similar methods? </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 434505,
      "author_name": "Sreyan Ghosh",
      "author_url": "",
      "post_date": "2018-12-06T14:14:34.537000",
      "content": "<p>Hi,\nI am new to competitive data science and Kaggle. Could anyone explain what balancing of predictions mean?</p>",
      "votes": 4,
      "replies": [
        {
          "id": 435494,
          "author_name": "Appian",
          "author_url": "",
          "post_date": "2018-12-08T05:34:20.063000",
          "content": "<p>from <a href=\"https://www.kaggle.com/c/quickdraw-doodle-recognition/discussion/73738\">https://www.kaggle.com/c/quickdraw-doodle-recognition/discussion/73738</a></p>\n\n<blockquote>\n  <p>The algorithm behind postprocessing is the following: for the most popular class decrease all the  probabilities iteratively by the same small value until it is no longer the most popular, repeat this procedure until all classes become equal. This technique was also used in one of the previous competitions (see github link for the code).</p>\n</blockquote>\n\n<p>This is the implementation of the explanation above. \nA variable alpha is the one he mentioned as \"the same small value\". In this case, they use 0.001 as initial alpha and iteratively halve the value when the score no longer improves until it gets 0.0001.</p>\n\n<p>Without balancing, for example if you directly use the prediction of your model, one class might be presented a few times more than minor classes. This results in worse score if the distribution of your prediction much differs from the class distribution of test set. To counter this, you can balance the probability of your prediction so that all classes are equally presented.</p>\n\n<p>This kind of balancing works when you know the distribution of test set as a prior knowledge which is the case for this competition. According to inversion, the claases were actually not evenly balanced but I think it's still fairly balanced and that's why it worked very well.</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 435646,
          "author_name": "Shaohua Li",
          "author_url": "",
          "post_date": "2018-12-08T13:04:20.973000",
          "content": "<p>Thanks for the nice explanation. By \"all classes become equal\" does it mean that we know all classes are balanced (i.e. equal probabilities)? So if in our prior knowledge the classes are not balanced, we also tune the output distributions to match this unbalanced distribution? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 435906,
          "author_name": "Appian",
          "author_url": "",
          "post_date": "2018-12-09T03:27:36.423000",
          "content": "<p>Yes. Modifying a few lines such as torch.ones() which address equally balanced label distribution to something else should achive what you want.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 436499,
          "author_name": "Sreyan Ghosh",
          "author_url": "",
          "post_date": "2018-12-10T12:33:02.967000",
          "content": "<p>Thanks a lot for that great explanation!!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 437915,
          "author_name": "Stan Chen",
          "author_url": "",
          "post_date": "2018-12-12T18:48:46.837000",
          "content": "<p>thanks for the explanation, this is extremely helpful</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 434158,
      "author_name": "pudae",
      "author_url": "",
      "post_date": "2018-12-06T01:24:19.117000",
      "content": "<p>Awesome!! After applying this algorithm, my lb score boosts from <strong>Private LB 0.94974</strong> to <strong>0.95467</strong>.\nThanks for the code! It will be very useful for the future competitions.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 434357,
          "author_name": "Pavel Ostyakov",
          "author_url": "",
          "post_date": "2018-12-06T09:13:29.817000",
          "content": "<p>You are welcome!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 434114,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-12-05T22:50:53.877000",
      "content": "<p>here is probably one related paper:\n\"Probabilistic n-Choose-k Models for Classification and Ranking\"</p>\n\n<p><a href=\"http://www.cs.toronto.edu/~kswersky/wp-content/uploads/pnck.pdf\">http://www.cs.toronto.edu/~kswersky/wp-content/uploads/pnck.pdf</a></p>\n\n<p>\"Instead, we focus on work related to the main novelty in this paper,\nthe explicit modeling of structure on label counts. That is, given that we have prior knowledge of\nlabel count structure, or are modeling a domain that exhibits such structure, the question is how can\nthe structure be leveraged to improve a model.\"</p>",
      "votes": 2,
      "replies": [
        {
          "id": 434359,
          "author_name": "Pavel Ostyakov",
          "author_url": "",
          "post_date": "2018-12-06T09:13:54.677000",
          "content": "<p>Thanks for sharing </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 438270,
      "author_name": "Bear Guan",
      "author_url": "",
      "post_date": "2018-12-13T11:49:21.193000",
      "content": "<p>GOOD JOB!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 435233,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "2018-12-07T18:08:51.103000",
      "content": "<p>For what it's worth, classes in the test set were not evenly balanced.  But, obviously, they were not highly skewed either. :-)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 435236,
          "author_name": "Pavel Ostyakov",
          "author_url": "",
          "post_date": "2018-12-07T18:12:20.763000",
          "content": "<p>That's interesting! Could I ask you to provide a distribution of classes in the test set? I may try to improve my script for this competition :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 434869,
      "author_name": "Shaohua Li",
      "author_url": "",
      "post_date": "2018-12-07T04:47:04.233000",
      "content": "<p>Amazing and extremely simple. I'm really wondering how it works? Thanks for your explanation in advance.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 434523,
      "author_name": "天大狂徒",
      "author_url": "",
      "post_date": "2018-12-06T14:33:06.130000",
      "content": "<p>good job!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 434194,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-06T03:22:54.950000",
      "content": "",
      "votes": 8,
      "replies": []
    },
    {
      "id": 434199,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-06T03:31:02.783000",
      "content": "<p>Many Thanks Pavel!</p>",
      "votes": 7,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "433924": "Hey guys,\n\nI've just published my code for balancing predictions: https://github.com/PavelOstyakov/predictions_balancing\n\nIt gives ~0.7% boost for any submission. Feel free to use it in future competitions :)\n\n![enter image description here][1]\n\n\n  [1]: https://sd.keepcalm-o-matic.co.uk/i-w600/keep-calm-and-balance-your-predictions.jpg",
    "434108": "thanks for the code!\n\nany idea what keyword i should use for goggle search for similar methods? ",
    "434505": "Hi,\nI am new to competitive data science and Kaggle. Could anyone explain what balancing of predictions mean?",
    "434158": "Awesome!! After applying this algorithm, my lb score boosts from **Private LB 0.94974** to **0.95467**.\nThanks for the code! It will be very useful for the future competitions.",
    "434114": "here is probably one related paper:\n\"Probabilistic n-Choose-k Models for Classification and Ranking\"\n\nhttp://www.cs.toronto.edu/~kswersky/wp-content/uploads/pnck.pdf\n\n\"Instead, we focus on work related to the main novelty in this paper,\nthe explicit modeling of structure on label counts. That is, given that we have prior knowledge of\nlabel count structure, or are modeling a domain that exhibits such structure, the question is how can\nthe structure be leveraged to improve a model.\"",
    "438270": "GOOD JOB!",
    "435233": "For what it's worth, classes in the test set were not evenly balanced.  But, obviously, they were not highly skewed either. :-)",
    "434869": "Amazing and extremely simple. I'm really wondering how it works? Thanks for your explanation in advance.",
    "434523": "good job!",
    "434194": "",
    "434199": "Many Thanks Pavel!"
  }
}