{
  "id": 376472,
  "title": "A tiny improvement of WeightedRandomSampler",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/376472",
  "author_name": "Chenglu",
  "post_date": "2023-01-06T13:45:14.188000",
  "votes": 24,
  "comment_count": 15,
  "views": 0,
  "content": "<p>If you are using <code>WeightedRandomSampler</code> in this game, you may have some interest in this.</p>\n<p>A simple idea, <code>WeightedRandomSampler</code> is randomly sampling negative samples repeatedly, it means that some samples may have been sampled many times but some others may not be sampled. So I write this simple sampler to address this problem: <a href=\"https://github.com/louis-she/exhaustive-weighted-random-sampler\" target=\"_blank\">https://github.com/louis-she/exhaustive-weighted-random-sampler</a>, the goal here is not to waste any samples even if there is a lot of them!</p>\n<p>A demo notebook can be found here: <a href=\"https://www.kaggle.com/code/snaker/exhaustiveweightedrandomsampler/notebook\" target=\"_blank\">https://www.kaggle.com/code/snaker/exhaustiveweightedrandomsampler/notebook</a></p>",
  "messages": [
    {
      "id": 2088617,
      "postDate": "2023-01-06T13:45:14.190Z",
      "content": "<p>If you are using <code>WeightedRandomSampler</code> in this game, you may have some interest in this.</p>\n<p>A simple idea, <code>WeightedRandomSampler</code> is randomly sampling negative samples repeatedly, it means that some samples may have been sampled many times but some others may not be sampled. So I write this simple sampler to address this problem: <a href=\"https://github.com/louis-she/exhaustive-weighted-random-sampler\" target=\"_blank\">https://github.com/louis-she/exhaustive-weighted-random-sampler</a>, the goal here is not to waste any samples even if there is a lot of them!</p>\n<p>A demo notebook can be found here: <a href=\"https://www.kaggle.com/code/snaker/exhaustiveweightedrandomsampler/notebook\" target=\"_blank\">https://www.kaggle.com/code/snaker/exhaustiveweightedrandomsampler/notebook</a></p>",
      "rawMarkdown": "If you are using `WeightedRandomSampler` in this game, you may have some interest in this.\n\nA simple idea, `WeightedRandomSampler` is randomly sampling negative samples repeatedly, it means that some samples may have been sampled many times but some others may not be sampled. So I write this simple sampler to address this problem: https://github.com/louis-she/exhaustive-weighted-random-sampler, the goal here is not to waste any samples even if there is a lot of them!\n\nA demo notebook can be found here: https://www.kaggle.com/code/snaker/exhaustiveweightedrandomsampler/notebook",
      "votes": 24
    },
    {
      "id": 2114022,
      "postDate": "2023-01-24T16:54:34.687Z",
      "content": "<p>I've always taken the approach of:</p>\n<ul>\n<li>…at the start of every epoch</li>\n<li>…oversample all the positive cases using the epoch number as the random seed and make that your dataset for the epoch</li>\n<li>…iterate over all that dataset in the standard pytorch way</li>\n</ul>\n<p>This means every epoch, every negative case gets used, and which positive cases get oversampled changes each epoch.</p>\n<p>I can't see why using this exhaustive weighted random sampler is any better than this, but would be keen to hear what you think <a href=\"https://www.kaggle.com/snaker\" target=\"_blank\">@snaker</a>? </p>",
      "rawMarkdown": "I've always taken the approach of:\n- ...at the start of every epoch\n- ...oversample all the positive cases using the epoch number as the random seed and make that your dataset for the epoch\n- ...iterate over all that dataset in the standard pytorch way\n\nThis means every epoch, every negative case gets used, and which positive cases get oversampled changes each epoch.\n\nI can't see why using this exhaustive weighted random sampler is any better than this, but would be keen to hear what you think @snaker? ",
      "replies": [
        {
          "id": 2114493,
          "postDate": "2023-01-25T03:08:23.687Z",
          "content": "<p>It depends on how to define one \"epoch\". If the dataset is large and is very imbalanced ( like this one ), I prefer to use <code>WeightedRandomSampler</code> with <code>num_samples</code> much smaller than the whole dataset, usually it's <code>number of positive samples</code> * N. </p>",
          "rawMarkdown": "It depends on how to define one \"epoch\". If the dataset is large and is very imbalanced ( like this one ), I prefer to use `WeightedRandomSampler` with `num_samples` much smaller than the whole dataset, usually it's `number of positive samples` * N. ",
          "votes": 1,
          "replies": [
            {
              "id": 2115344,
              "postDate": "2023-01-25T17:26:09.150Z",
              "content": "<p>Ok, that makes sense, but it's effectively the same strategy but just a different way of setting training time?<br>\nAlso, I may be wrong, but I think your method doesn't ensure a uniform oversampling of the minority class, because <code>replace=True</code>?</p>\n<p>I think my approach from above might have some slight advantages, but not 100% sure.</p>",
              "rawMarkdown": "Ok, that makes sense, but it's effectively the same strategy but just a different way of setting training time?\nAlso, I may be wrong, but I think your method doesn't ensure a uniform oversampling of the minority class, because `replace=True`?\n\nI think my approach from above might have some slight advantages, but not 100% sure."
            },
            {
              "id": 2115836,
              "postDate": "2023-01-26T01:40:55.700Z",
              "content": "<p>\"uniform oversampling of the minority class\"  -  that's true, <code>WeightedRandomSampler</code> can not ensure equal sample every minority class, maybe some class sample 5 time while others sample 2 or 3 times.</p>",
              "rawMarkdown": "\"uniform oversampling of the minority class\"  -  that's true, `WeightedRandomSampler` can not ensure equal sample every minority class, maybe some class sample 5 time while others sample 2 or 3 times."
            }
          ]
        }
      ]
    },
    {
      "id": 2108919,
      "postDate": "2023-01-21T00:11:36.990Z",
      "content": "<p>Thanks for sharing. But when I use ExhaustiveWeightedRandomSampler, my CV improved and my LB didn't improve. It confused me a lot</p>",
      "rawMarkdown": "Thanks for sharing. But when I use ExhaustiveWeightedRandomSampler, my CV improved and my LB didn't improve. It confused me a lot",
      "replies": [
        {
          "id": 2108950,
          "postDate": "2023-01-21T01:08:17.900Z",
          "content": "<p>LB is not stable… My CV boosted from 0.43 to 0.51, but LB is still the same…</p>",
          "rawMarkdown": "LB is not stable... My CV boosted from 0.43 to 0.51, but LB is still the same...",
          "votes": 1,
          "replies": [
            {
              "id": 2108984,
              "postDate": "2023-01-21T02:48:22.570Z",
              "content": "<p>So shall we trust cv or lb? haha</p>",
              "rawMarkdown": "So shall we trust cv or lb? haha"
            },
            {
              "id": 2109004,
              "postDate": "2023-01-21T03:37:39.307Z",
              "content": "<p>In my opinion, both, but give them a weight</p>",
              "rawMarkdown": "In my opinion, both, but give them a weight"
            },
            {
              "id": 2109076,
              "postDate": "2023-01-21T05:24:17.463Z",
              "content": "<p>use this playground code to see what really happens.<br>\nsee how oversampling changes the decision space</p>\n<p><a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/372628\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/372628</a></p>",
              "rawMarkdown": "use this playground code to see what really happens.\nsee how oversampling changes the decision space\n\nhttps://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/372628\n",
              "votes": 2
            },
            {
              "id": 2109085,
              "postDate": "2023-01-21T05:33:41.850Z",
              "content": "<p>why choose the best?<br>\nyou can choose the best k-models from different weights in loss, etc</p>\n<p>[1]Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time</p>\n<p>use larger learning rate to learn a coarse model.<br>\nthen finetune using different seed (+ different hyperparameters, rate, weighing, etc)</p>\n<p>ensemble all using the method  in [1], i.e. just average the weights of the best k-models in greedy way.</p>",
              "rawMarkdown": "why choose the best?\nyou can choose the best k-models from different weights in loss, etc\n\n[1]Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time\n\nuse larger learning rate to learn a coarse model.\nthen finetune using different seed (+ different hyperparameters, rate, weighing, etc)\n\nensemble all using the method  in [1], i.e. just average the weights of the best k-models in greedy way.\n"
            },
            {
              "id": 2109268,
              "postDate": "2023-01-21T08:37:27.767Z",
              "content": "<p>Really cool visualizations, thanks heng</p>",
              "rawMarkdown": "Really cool visualizations, thanks heng"
            },
            {
              "id": 2114855,
              "postDate": "2023-01-25T09:48:55.643Z",
              "content": "<p>Yes, but surely that is only a one fold boost, not the entire CV..   😲</p>",
              "rawMarkdown": "Yes, but surely that is only a one fold boost, not the entire CV..   😲"
            }
          ]
        }
      ]
    },
    {
      "id": 2090365,
      "postDate": "2023-01-07T08:52:44.363Z",
      "content": "<p>thanks for sharing.  down sample or up sample, should improve the LB.</p>",
      "rawMarkdown": "thanks for sharing.  down sample or up sample, should improve the LB.",
      "replies": [
        {
          "id": 2090449,
          "postDate": "2023-01-07T10:46:02.627Z",
          "content": "<p>When you find way to recognize (or show to model) cancer it will definitely help. For me challange is to find such way. Some images are classified very very well … some of them … no. <br>\nWhen I see cancer … model see it as well but when I don't (even label is cancer) model has the same problem. We are both poor in cancer recognition. 😂</p>",
          "rawMarkdown": "When you find way to recognize (or show to model) cancer it will definitely help. For me challange is to find such way. Some images are classified very very well ... some of them ... no. \nWhen I see cancer ... model see it as well but when I don't (even label is cancer) model has the same problem. We are both poor in cancer recognition. 😂",
          "votes": 2,
          "replies": [
            {
              "id": 2094924,
              "postDate": "2023-01-11T05:19:04.537Z",
              "content": "<p>that is true. It is impossible for no specialist to deal with abnormal data or wrong labelled data.</p>",
              "rawMarkdown": "that is true. It is impossible for no specialist to deal with abnormal data or wrong labelled data."
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2114022,
      "author_name": "James Howard",
      "author_url": "",
      "post_date": "2023-01-24T16:54:34.687000",
      "content": "<p>I've always taken the approach of:</p>\n<ul>\n<li>…at the start of every epoch</li>\n<li>…oversample all the positive cases using the epoch number as the random seed and make that your dataset for the epoch</li>\n<li>…iterate over all that dataset in the standard pytorch way</li>\n</ul>\n<p>This means every epoch, every negative case gets used, and which positive cases get oversampled changes each epoch.</p>\n<p>I can't see why using this exhaustive weighted random sampler is any better than this, but would be keen to hear what you think <a href=\"https://www.kaggle.com/snaker\" target=\"_blank\">@snaker</a>? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2114493,
          "author_name": "Chenglu",
          "author_url": "",
          "post_date": "2023-01-25T03:08:23.687000",
          "content": "<p>It depends on how to define one \"epoch\". If the dataset is large and is very imbalanced ( like this one ), I prefer to use <code>WeightedRandomSampler</code> with <code>num_samples</code> much smaller than the whole dataset, usually it's <code>number of positive samples</code> * N. </p>",
          "votes": 1,
          "replies": [
            {
              "id": 2115344,
              "author_name": "James Howard",
              "author_url": "",
              "post_date": "2023-01-25T17:26:09.150000",
              "content": "<p>Ok, that makes sense, but it's effectively the same strategy but just a different way of setting training time?<br>\nAlso, I may be wrong, but I think your method doesn't ensure a uniform oversampling of the minority class, because <code>replace=True</code>?</p>\n<p>I think my approach from above might have some slight advantages, but not 100% sure.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2115836,
              "author_name": "Chenglu",
              "author_url": "",
              "post_date": "2023-01-26T01:40:55.700000",
              "content": "<p>\"uniform oversampling of the minority class\"  -  that's true, <code>WeightedRandomSampler</code> can not ensure equal sample every minority class, maybe some class sample 5 time while others sample 2 or 3 times.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2108919,
      "author_name": "Leimeng46",
      "author_url": "",
      "post_date": "2023-01-21T00:11:36.990000",
      "content": "<p>Thanks for sharing. But when I use ExhaustiveWeightedRandomSampler, my CV improved and my LB didn't improve. It confused me a lot</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2108950,
          "author_name": "Chenglu",
          "author_url": "",
          "post_date": "2023-01-21T01:08:17.900000",
          "content": "<p>LB is not stable… My CV boosted from 0.43 to 0.51, but LB is still the same…</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2108984,
              "author_name": "Leimeng46",
              "author_url": "",
              "post_date": "2023-01-21T02:48:22.570000",
              "content": "<p>So shall we trust cv or lb? haha</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2109004,
              "author_name": "Chenglu",
              "author_url": "",
              "post_date": "2023-01-21T03:37:39.307000",
              "content": "<p>In my opinion, both, but give them a weight</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2109076,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-01-21T05:24:17.463000",
              "content": "<p>use this playground code to see what really happens.<br>\nsee how oversampling changes the decision space</p>\n<p><a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/372628\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/372628</a></p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2109085,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-01-21T05:33:41.850000",
              "content": "<p>why choose the best?<br>\nyou can choose the best k-models from different weights in loss, etc</p>\n<p>[1]Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time</p>\n<p>use larger learning rate to learn a coarse model.<br>\nthen finetune using different seed (+ different hyperparameters, rate, weighing, etc)</p>\n<p>ensemble all using the method  in [1], i.e. just average the weights of the best k-models in greedy way.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2109268,
              "author_name": "Chenglu",
              "author_url": "",
              "post_date": "2023-01-21T08:37:27.767000",
              "content": "<p>Really cool visualizations, thanks heng</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2114855,
              "author_name": "Ivan Aerlic",
              "author_url": "",
              "post_date": "2023-01-25T09:48:55.643000",
              "content": "<p>Yes, but surely that is only a one fold boost, not the entire CV..   😲</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2090365,
      "author_name": "dragon zhang",
      "author_url": "",
      "post_date": "2023-01-07T08:52:44.363000",
      "content": "<p>thanks for sharing.  down sample or up sample, should improve the LB.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2090449,
          "author_name": "Remek Kinas",
          "author_url": "",
          "post_date": "2023-01-07T10:46:02.627000",
          "content": "<p>When you find way to recognize (or show to model) cancer it will definitely help. For me challange is to find such way. Some images are classified very very well … some of them … no. <br>\nWhen I see cancer … model see it as well but when I don't (even label is cancer) model has the same problem. We are both poor in cancer recognition. 😂</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2094924,
              "author_name": "dragon zhang",
              "author_url": "",
              "post_date": "2023-01-11T05:19:04.537000",
              "content": "<p>that is true. It is impossible for no specialist to deal with abnormal data or wrong labelled data.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2088617": "If you are using `WeightedRandomSampler` in this game, you may have some interest in this.\n\nA simple idea, `WeightedRandomSampler` is randomly sampling negative samples repeatedly, it means that some samples may have been sampled many times but some others may not be sampled. So I write this simple sampler to address this problem: https://github.com/louis-she/exhaustive-weighted-random-sampler, the goal here is not to waste any samples even if there is a lot of them!\n\nA demo notebook can be found here: https://www.kaggle.com/code/snaker/exhaustiveweightedrandomsampler/notebook",
    "2114022": "I've always taken the approach of:\n- ...at the start of every epoch\n- ...oversample all the positive cases using the epoch number as the random seed and make that your dataset for the epoch\n- ...iterate over all that dataset in the standard pytorch way\n\nThis means every epoch, every negative case gets used, and which positive cases get oversampled changes each epoch.\n\nI can't see why using this exhaustive weighted random sampler is any better than this, but would be keen to hear what you think @snaker? ",
    "2108919": "Thanks for sharing. But when I use ExhaustiveWeightedRandomSampler, my CV improved and my LB didn't improve. It confused me a lot",
    "2090365": "thanks for sharing.  down sample or up sample, should improve the LB."
  }
}