{
  "id": 373628,
  "title": "Focal or BCE loss?",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/373628",
  "author_name": "nicehzj",
  "post_date": "2022-12-22T13:13:44.429000",
  "votes": 5,
  "comment_count": 19,
  "views": 0,
  "content": "<p>To deal with the unbalanced dataset, I'm now using BCE loss with positive weighted.<br>\nAnd I also saw some ideas using Focal lose to solve this.<br>\nAny performance difference on these two? Which one is better on CV problem?</p>",
  "messages": [
    {
      "id": 2072853,
      "postDate": "2022-12-22T13:13:44.430Z",
      "content": "<p>To deal with the unbalanced dataset, I'm now using BCE loss with positive weighted.<br>\nAnd I also saw some ideas using Focal lose to solve this.<br>\nAny performance difference on these two? Which one is better on CV problem?</p>",
      "rawMarkdown": "To deal with the unbalanced dataset, I'm now using BCE loss with positive weighted.\nAnd I also saw some ideas using Focal lose to solve this.\nAny performance difference on these two? Which one is better on CV problem?",
      "votes": 5
    },
    {
      "id": 2073328,
      "postDate": "2022-12-22T21:52:23.863Z",
      "content": "<p>BCELosswithlogits + weight works better then Focal Loss for me. I use custom sampler to oversample positive cases too.<br>\nHowever I am not TOP (my single model is no more then 0.4) on LB :) so I can be totally wrong :) </p>",
      "rawMarkdown": "BCELosswithlogits + weight works better then Focal Loss for me. I use custom sampler to oversample positive cases too.\nHowever I am not TOP (my single model is no more then 0.4) on LB :) so I can be totally wrong :) ",
      "votes": 3,
      "replies": [
        {
          "id": 2073947,
          "postDate": "2022-12-23T15:10:04.890Z",
          "content": "<p>Same here. I tried both and it seems that BCELosswithlogits with a different weight for positive examples is better here than focal loss (don't focus on my current score, I only submitted a dummy model to see if my submission notebook work). Maybe Focal loss is better when there are more than 2 classes or I have the wrong parameters but BCE gave me better results.</p>",
          "rawMarkdown": "Same here. I tried both and it seems that BCELosswithlogits with a different weight for positive examples is better here than focal loss (don't focus on my current score, I only submitted a dummy model to see if my submission notebook work). Maybe Focal loss is better when there are more than 2 classes or I have the wrong parameters but BCE gave me better results.",
          "votes": 2,
          "replies": [
            {
              "id": 2073963,
              "postDate": "2022-12-23T15:38:39.097Z",
              "content": "<p>This could be reallly helpful… Because there are too many things to try, I need to try something new, many many thanks!</p>",
              "rawMarkdown": "This could be reallly helpful... Because there are too many things to try, I need to try something new, many many thanks!"
            },
            {
              "id": 2073978,
              "postDate": "2022-12-23T15:54:17.413Z",
              "content": "<p>Thank you for your comment.<br>\nAnother great step for me was upsample data. I use modified BalanceSampler from <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>.</p>\n<pre><code>class BalanceSampler(torch.utils.data.Sampler):\n\n    def __init__(self, dataset, ratio = 3):\n        self.r = ratio-1\n        self.dataset = dataset\n        self.pos_index = np.where(dataset.df.cancer&gt;0)[0]\n        self.neg_index = np.where(dataset.df.cancer==0)[0]\n\n        self.length = self.r * int(np.floor(len(self.neg_index)/self.r)) \n        self.ds_len =  self.length + (self.length // self.r) \n\n    def __iter__(self):\n        pos_index = self.pos_index.copy()\n        neg_index = self.neg_index.copy()\n        np.random.shuffle(pos_index)\n        np.random.shuffle(neg_index)\n\n        neg_index = neg_index[:self.length].reshape(-1,self.r)\n        #pos_index = np.random.choice(pos_index, self.length//self.r).reshape(-1,1)\n        pos_index = np.tile(pos_index, (len(neg_index) // len(pos_index)) + 1)[:len(neg_index)].reshape(-1,1)\n\n        index = np.concatenate([pos_index,neg_index],-1).reshape(-1)\n        return iter(index)\n\n    def __len__(self):\n        return self.ds_len\n</code></pre>\n<p>Then in Pytorch DataLoader (for train) you can upsample positive cases.  For validation I use SequentialSampler (default one).</p>",
              "rawMarkdown": "Thank you for your comment.\nAnother great step for me was upsample data. I use modified BalanceSampler from @hengck23.\n\n```\nclass BalanceSampler(torch.utils.data.Sampler):\n\n    def __init__(self, dataset, ratio = 3):\n        self.r = ratio-1\n        self.dataset = dataset\n        self.pos_index = np.where(dataset.df.cancer>0)[0]\n        self.neg_index = np.where(dataset.df.cancer==0)[0]\n\n        self.length = self.r * int(np.floor(len(self.neg_index)/self.r)) \n        self.ds_len =  self.length + (self.length // self.r) \n\n    def __iter__(self):\n        pos_index = self.pos_index.copy()\n        neg_index = self.neg_index.copy()\n        np.random.shuffle(pos_index)\n        np.random.shuffle(neg_index)\n\n        neg_index = neg_index[:self.length].reshape(-1,self.r)\n        #pos_index = np.random.choice(pos_index, self.length//self.r).reshape(-1,1)\n        pos_index = np.tile(pos_index, (len(neg_index) // len(pos_index)) + 1)[:len(neg_index)].reshape(-1,1)\n\n        index = np.concatenate([pos_index,neg_index],-1).reshape(-1)\n        return iter(index)\n\n    def __len__(self):\n        return self.ds_len\n```\n\nThen in Pytorch DataLoader (for train) you can upsample positive cases.  For validation I use SequentialSampler (default one).",
              "votes": 9
            },
            {
              "id": 2074055,
              "postDate": "2022-12-23T17:17:48.917Z",
              "content": "<p>For upsample, I use RandomOverSampler from imblearn package.<br>\nIt can randomly choose and copy the positive data until the pos/neg rate you want.<br>\nMaybe you can have a try. just twos lines added after you get the csv.<br>\n<code>oversample = RandomOverSampler(sampling_strategy=0.2)</code><br>\n<code>train_UP_X, train_UP_Y = oversample.fit_resample(train[[col for col in train.columns if col!='cancer']], train['cancer'])</code></p>",
              "rawMarkdown": "For upsample, I use RandomOverSampler from imblearn package.\nIt can randomly choose and copy the positive data until the pos/neg rate you want.\nMaybe you can have a try. just twos lines added after you get the csv.\n`oversample = RandomOverSampler(sampling_strategy=0.2)`\n`train_UP_X, train_UP_Y = oversample.fit_resample(train[[col for col in train.columns if col!='cancer']], train['cancer'])`"
            },
            {
              "id": 2074117,
              "postDate": "2022-12-23T18:09:37.717Z",
              "content": "<p>I totally agree, I think there is a right balance between the amount of oversampling and the reweighting of the positive class. Personally for oversampling I just duplicated all positives examples IDs n times from the train CSV (around 3 to 8 max in different tests). <br>\nThis may not be the best but given the few positive examples we have, this ensure that a subset of positives samples is not oversampled compared to other positive samples and does not lead the model to learn to only recognize a few positives examples.</p>",
              "rawMarkdown": "I totally agree, I think there is a right balance between the amount of oversampling and the reweighting of the positive class. Personally for oversampling I just duplicated all positives examples IDs n times from the train CSV (around 3 to 8 max in different tests). \nThis may not be the best but given the few positive examples we have, this ensure that a subset of positives samples is not oversampled compared to other positive samples and does not lead the model to learn to only recognize a few positives examples."
            },
            {
              "id": 2074127,
              "postDate": "2022-12-23T18:24:02.910Z",
              "content": "<p>I thnik DataLoader Sampler has advantage over DS oversampling - you can balance for each batch and provide model more stable way for learning. </p>\n<p>Nice disucssion 👍</p>",
              "rawMarkdown": "I thnik DataLoader Sampler has advantage over DS oversampling - you can balance for each batch and provide model more stable way for learning. \n\nNice disucssion 👍"
            },
            {
              "id": 2074245,
              "postDate": "2022-12-23T22:07:10.860Z",
              "content": "<p>I didn't think about it. It could probably give a smoother pf1 value during training instead of the unstable curves I am currently getting.</p>\n<p>Thanks for the discussion, was a pleasure 🤝</p>",
              "rawMarkdown": "I didn't think about it. It could probably give a smoother pf1 value during training instead of the unstable curves I am currently getting.\n\nThanks for the discussion, was a pleasure 🤝"
            },
            {
              "id": 2074573,
              "postDate": "2022-12-24T10:40:14.227Z",
              "content": "<p>Hi Remek. Beginner here. Really appreciate your efforts so others life is a bit easy on kaggle and for your helpful nature.<br>\nJust a quick question. Does oversampling/undersamplig minority/majority classes improve model performance in general? I have read many say that it has zero to negative effect on the predictive capability since it distorts the original distribution of data. Kindly advice.</p>",
              "rawMarkdown": "Hi Remek. Beginner here. Really appreciate your efforts so others life is a bit easy on kaggle and for your helpful nature.\nJust a quick question. Does oversampling/undersamplig minority/majority classes improve model performance in general? I have read many say that it has zero to negative effect on the predictive capability since it distorts the original distribution of data. Kindly advice."
            },
            {
              "id": 2086707,
              "postDate": "2023-01-05T00:36:47.460Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a>, would you mind attaching an example code of how to apply the sampler ? <br>\nThanks in advance. </p>",
              "rawMarkdown": "Hi @remekkinas, would you mind attaching an example code of how to apply the sampler ? \nThanks in advance. "
            },
            {
              "id": 2087090,
              "postDate": "2023-01-05T10:17:48.013Z",
              "content": "<p>Sure. The best way to understand Sampler and Batch Sampler is to print output of batch. This code (above) presents Sampler (not batch sampler - this is different one). My recomendation is always to print output in experimentation notebook (usually I have two notebooks - one for quick experimentation and second as a final one).</p>\n<p>Example:</p>\n<ul>\n<li>ratio = 4</li>\n</ul>\n<pre><code>train_dataset = TrainDataset(train_folds, \n                                 transform=get_transforms(data='train'))\ntrain_loader = DataLoader(train_dataset, \n                              batch_size= 16, \n                              shuffle=False, \n                              sampler=BalanceSampler(train_dataset, ratio = 8),\n                              num_workers=8, pin_memory=False, drop_last=True,)\n</code></pre>",
              "rawMarkdown": "Sure. The best way to understand Sampler and Batch Sampler is to print output of batch. This code (above) presents Sampler (not batch sampler - this is different one). My recomendation is always to print output in experimentation notebook (usually I have two notebooks - one for quick experimentation and second as a final one).\n\nExample:\n- ratio = 4\n\n```\ntrain_dataset = TrainDataset(train_folds, \n                                 transform=get_transforms(data='train'))\ntrain_loader = DataLoader(train_dataset, \n                              batch_size= 16, \n                              shuffle=False, \n                              sampler=BalanceSampler(train_dataset, ratio = 8),\n                              num_workers=8, pin_memory=False, drop_last=True,)\n```\n\n\n",
              "votes": 1
            },
            {
              "id": 2087095,
              "postDate": "2023-01-05T10:21:37.630Z",
              "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> !</p>",
              "rawMarkdown": "Thanks a lot @remekkinas !"
            },
            {
              "id": 2087162,
              "postDate": "2023-01-05T11:27:52.007Z",
              "content": "<p>I print almost everything. I plot almost everytning. To be sure what is going - what looks input and out from my network. This is very important. I spent a lot of time frustrating with score … and then it appeared that … ploting everytning gave me an answer in seconds. This is very important to have controll. </p>",
              "rawMarkdown": "I print almost everything. I plot almost everytning. To be sure what is going - what looks input and out from my network. This is very important. I spent a lot of time frustrating with score ... and then it appeared that ... ploting everytning gave me an answer in seconds. This is very important to have controll. ",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2073220,
      "postDate": "2022-12-22T19:09:20.767Z",
      "content": "<p>I cant tell you in particular which is working better in this competition. </p>\n<p>You can also try different things which work well with class imbalance e.g.:</p>\n<p>Loss e.g.:<br>\n-Bi-Tempered Logistic Loss</p>\n<p>…</p>\n<p>Other methods e.g.:<br>\n-Label smoothing<br>\n-Minoririty augmentation<br>\n-Oversampling<br>\n-Form it into an AnomalyDetection use case</p>\n<p>…</p>\n<p>There are a lot of different options available.</p>",
      "rawMarkdown": "I cant tell you in particular which is working better in this competition. \n\nYou can also try different things which work well with class imbalance e.g.:\n\nLoss e.g.:\n-Bi-Tempered Logistic Loss\n\n…\n\nOther methods e.g.:\n-Label smoothing\n-Minoririty augmentation\n-Oversampling\n-Form it into an AnomalyDetection use case\n\n …\n\nThere are a lot of different options available.",
      "votes": 1,
      "replies": [
        {
          "id": 2073320,
          "postDate": "2022-12-22T21:30:51.470Z",
          "content": "<p>Thanks for your answers and help! I will try some of these methods~👍</p>",
          "rawMarkdown": "Thanks for your answers and help! I will try some of these methods~👍"
        }
      ]
    },
    {
      "id": 2077875,
      "postDate": "2022-12-27T23:36:59.563Z",
      "content": "<p>Congratulations on your progress so far with the competition.</p>\n<p>The best answer for all question of this type is: Try and see what works best..<br>\nFrom my experience, I find that I usually get better results when using a combination of label smoothing and sometimes also focal-loss. <br>\nGood luck!</p>\n<p>The Devastator.</p>",
      "rawMarkdown": "Congratulations on your progress so far with the competition.\n\nThe best answer for all question of this type is: Try and see what works best..\nFrom my experience, I find that I usually get better results when using a combination of label smoothing and sometimes also focal-loss. \nGood luck!\n\nThe Devastator.\n",
      "votes": 2,
      "replies": [
        {
          "id": 2077889,
          "postDate": "2022-12-27T23:46:56.047Z",
          "content": "<p>Thanks a lot! I'm now trying upsampling and also BCE with weight. <br>\nLooks okay but need some tuning about the upsampling ratio and BCE weight</p>",
          "rawMarkdown": "Thanks a lot! I'm now trying upsampling and also BCE with weight. \nLooks okay but need some tuning about the upsampling ratio and BCE weight"
        }
      ]
    },
    {
      "id": 2083840,
      "postDate": "2023-01-02T22:55:31.860Z",
      "content": "<p>Hi!<br>\nCongratulations on your progress!</p>\n<p>Which dataset do you use? ROI 1024 or just no ROI 512. I found that on 512 no ROI dataset, I could get 0.24LB by using ensemble + weighted loss + upsample. Here is my code.</p>\n<p><a href=\"https://www.kaggle.com/code/fanyang99/train-no-roi-512-lb0-24\" target=\"_blank\">https://www.kaggle.com/code/fanyang99/train-no-roi-512-lb0-24</a></p>\n<p>Hope can get some suggestion. Should I go on to next dataset like ROI to do test? Or just try to improve methods on this dataset? Do you think 0.24 is already good on 512 data?</p>\n<p>Thank you so much!</p>",
      "rawMarkdown": "Hi!\nCongratulations on your progress!\n\nWhich dataset do you use? ROI 1024 or just no ROI 512. I found that on 512 no ROI dataset, I could get 0.24LB by using ensemble + weighted loss + upsample. Here is my code.\n\nhttps://www.kaggle.com/code/fanyang99/train-no-roi-512-lb0-24\n\nHope can get some suggestion. Should I go on to next dataset like ROI to do test? Or just try to improve methods on this dataset? Do you think 0.24 is already good on 512 data?\n\nThank you so much!"
    },
    {
      "id": 2075867,
      "postDate": "2022-12-26T01:21:35.560Z",
      "content": "<p>Hi,<br>\nFor me, I'm working with 50/50 balanced training dataset and standard BCE loss (so with no weighting). LB always higher than CV… Maybe I've found the right train/val division (by chance). However, if I were you, I'd suggest you to concentrate more on hard examples and data augmentation.</p>\n<p>Happy kaggling!</p>",
      "rawMarkdown": "Hi,\nFor me, I'm working with 50/50 balanced training dataset and standard BCE loss (so with no weighting). LB always higher than CV... Maybe I've found the right train/val division (by chance). However, if I were you, I'd suggest you to concentrate more on hard examples and data augmentation.\n\nHappy kaggling!"
    }
  ],
  "comments": [
    {
      "id": 2073328,
      "author_name": "Remek Kinas",
      "author_url": "",
      "post_date": "2022-12-22T21:52:23.863000",
      "content": "<p>BCELosswithlogits + weight works better then Focal Loss for me. I use custom sampler to oversample positive cases too.<br>\nHowever I am not TOP (my single model is no more then 0.4) on LB :) so I can be totally wrong :) </p>",
      "votes": 3,
      "replies": [
        {
          "id": 2073947,
          "author_name": "Natyu",
          "author_url": "",
          "post_date": "2022-12-23T15:10:04.890000",
          "content": "<p>Same here. I tried both and it seems that BCELosswithlogits with a different weight for positive examples is better here than focal loss (don't focus on my current score, I only submitted a dummy model to see if my submission notebook work). Maybe Focal loss is better when there are more than 2 classes or I have the wrong parameters but BCE gave me better results.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2073963,
              "author_name": "nicehzj",
              "author_url": "",
              "post_date": "2022-12-23T15:38:39.097000",
              "content": "<p>This could be reallly helpful… Because there are too many things to try, I need to try something new, many many thanks!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2073978,
              "author_name": "Remek Kinas",
              "author_url": "",
              "post_date": "2022-12-23T15:54:17.413000",
              "content": "<p>Thank you for your comment.<br>\nAnother great step for me was upsample data. I use modified BalanceSampler from <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>.</p>\n<pre><code>class BalanceSampler(torch.utils.data.Sampler):\n\n    def __init__(self, dataset, ratio = 3):\n        self.r = ratio-1\n        self.dataset = dataset\n        self.pos_index = np.where(dataset.df.cancer&gt;0)[0]\n        self.neg_index = np.where(dataset.df.cancer==0)[0]\n\n        self.length = self.r * int(np.floor(len(self.neg_index)/self.r)) \n        self.ds_len =  self.length + (self.length // self.r) \n\n    def __iter__(self):\n        pos_index = self.pos_index.copy()\n        neg_index = self.neg_index.copy()\n        np.random.shuffle(pos_index)\n        np.random.shuffle(neg_index)\n\n        neg_index = neg_index[:self.length].reshape(-1,self.r)\n        #pos_index = np.random.choice(pos_index, self.length//self.r).reshape(-1,1)\n        pos_index = np.tile(pos_index, (len(neg_index) // len(pos_index)) + 1)[:len(neg_index)].reshape(-1,1)\n\n        index = np.concatenate([pos_index,neg_index],-1).reshape(-1)\n        return iter(index)\n\n    def __len__(self):\n        return self.ds_len\n</code></pre>\n<p>Then in Pytorch DataLoader (for train) you can upsample positive cases.  For validation I use SequentialSampler (default one).</p>",
              "votes": 9,
              "replies": []
            },
            {
              "id": 2074055,
              "author_name": "nicehzj",
              "author_url": "",
              "post_date": "2022-12-23T17:17:48.917000",
              "content": "<p>For upsample, I use RandomOverSampler from imblearn package.<br>\nIt can randomly choose and copy the positive data until the pos/neg rate you want.<br>\nMaybe you can have a try. just twos lines added after you get the csv.<br>\n<code>oversample = RandomOverSampler(sampling_strategy=0.2)</code><br>\n<code>train_UP_X, train_UP_Y = oversample.fit_resample(train[[col for col in train.columns if col!='cancer']], train['cancer'])</code></p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2074117,
              "author_name": "Natyu",
              "author_url": "",
              "post_date": "2022-12-23T18:09:37.717000",
              "content": "<p>I totally agree, I think there is a right balance between the amount of oversampling and the reweighting of the positive class. Personally for oversampling I just duplicated all positives examples IDs n times from the train CSV (around 3 to 8 max in different tests). <br>\nThis may not be the best but given the few positive examples we have, this ensure that a subset of positives samples is not oversampled compared to other positive samples and does not lead the model to learn to only recognize a few positives examples.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2074127,
              "author_name": "Remek Kinas",
              "author_url": "",
              "post_date": "2022-12-23T18:24:02.910000",
              "content": "<p>I thnik DataLoader Sampler has advantage over DS oversampling - you can balance for each batch and provide model more stable way for learning. </p>\n<p>Nice disucssion 👍</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2074245,
              "author_name": "Natyu",
              "author_url": "",
              "post_date": "2022-12-23T22:07:10.860000",
              "content": "<p>I didn't think about it. It could probably give a smoother pf1 value during training instead of the unstable curves I am currently getting.</p>\n<p>Thanks for the discussion, was a pleasure 🤝</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2074573,
              "author_name": "Sandy",
              "author_url": "",
              "post_date": "2022-12-24T10:40:14.227000",
              "content": "<p>Hi Remek. Beginner here. Really appreciate your efforts so others life is a bit easy on kaggle and for your helpful nature.<br>\nJust a quick question. Does oversampling/undersamplig minority/majority classes improve model performance in general? I have read many say that it has zero to negative effect on the predictive capability since it distorts the original distribution of data. Kindly advice.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2086707,
              "author_name": "Francisco Javier Gallego",
              "author_url": "",
              "post_date": "2023-01-05T00:36:47.460000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a>, would you mind attaching an example code of how to apply the sampler ? <br>\nThanks in advance. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2087090,
              "author_name": "Remek Kinas",
              "author_url": "",
              "post_date": "2023-01-05T10:17:48.013000",
              "content": "<p>Sure. The best way to understand Sampler and Batch Sampler is to print output of batch. This code (above) presents Sampler (not batch sampler - this is different one). My recomendation is always to print output in experimentation notebook (usually I have two notebooks - one for quick experimentation and second as a final one).</p>\n<p>Example:</p>\n<ul>\n<li>ratio = 4</li>\n</ul>\n<pre><code>train_dataset = TrainDataset(train_folds, \n                                 transform=get_transforms(data='train'))\ntrain_loader = DataLoader(train_dataset, \n                              batch_size= 16, \n                              shuffle=False, \n                              sampler=BalanceSampler(train_dataset, ratio = 8),\n                              num_workers=8, pin_memory=False, drop_last=True,)\n</code></pre>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2087095,
              "author_name": "Francisco Javier Gallego",
              "author_url": "",
              "post_date": "2023-01-05T10:21:37.630000",
              "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> !</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2087162,
              "author_name": "Remek Kinas",
              "author_url": "",
              "post_date": "2023-01-05T11:27:52.007000",
              "content": "<p>I print almost everything. I plot almost everytning. To be sure what is going - what looks input and out from my network. This is very important. I spent a lot of time frustrating with score … and then it appeared that … ploting everytning gave me an answer in seconds. This is very important to have controll. </p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2073220,
      "author_name": "Ali Abdin",
      "author_url": "",
      "post_date": "2022-12-22T19:09:20.767000",
      "content": "<p>I cant tell you in particular which is working better in this competition. </p>\n<p>You can also try different things which work well with class imbalance e.g.:</p>\n<p>Loss e.g.:<br>\n-Bi-Tempered Logistic Loss</p>\n<p>…</p>\n<p>Other methods e.g.:<br>\n-Label smoothing<br>\n-Minoririty augmentation<br>\n-Oversampling<br>\n-Form it into an AnomalyDetection use case</p>\n<p>…</p>\n<p>There are a lot of different options available.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2073320,
          "author_name": "nicehzj",
          "author_url": "",
          "post_date": "2022-12-22T21:30:51.470000",
          "content": "<p>Thanks for your answers and help! I will try some of these methods~👍</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2077875,
      "author_name": "The Devastator",
      "author_url": "",
      "post_date": "2022-12-27T23:36:59.563000",
      "content": "<p>Congratulations on your progress so far with the competition.</p>\n<p>The best answer for all question of this type is: Try and see what works best..<br>\nFrom my experience, I find that I usually get better results when using a combination of label smoothing and sometimes also focal-loss. <br>\nGood luck!</p>\n<p>The Devastator.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2077889,
          "author_name": "nicehzj",
          "author_url": "",
          "post_date": "2022-12-27T23:46:56.047000",
          "content": "<p>Thanks a lot! I'm now trying upsampling and also BCE with weight. <br>\nLooks okay but need some tuning about the upsampling ratio and BCE weight</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2083840,
      "author_name": "ChenxiangSun@NJU",
      "author_url": "",
      "post_date": "2023-01-02T22:55:31.860000",
      "content": "<p>Hi!<br>\nCongratulations on your progress!</p>\n<p>Which dataset do you use? ROI 1024 or just no ROI 512. I found that on 512 no ROI dataset, I could get 0.24LB by using ensemble + weighted loss + upsample. Here is my code.</p>\n<p><a href=\"https://www.kaggle.com/code/fanyang99/train-no-roi-512-lb0-24\" target=\"_blank\">https://www.kaggle.com/code/fanyang99/train-no-roi-512-lb0-24</a></p>\n<p>Hope can get some suggestion. Should I go on to next dataset like ROI to do test? Or just try to improve methods on this dataset? Do you think 0.24 is already good on 512 data?</p>\n<p>Thank you so much!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2075867,
      "author_name": "Giovanni Cavallin",
      "author_url": "",
      "post_date": "2022-12-26T01:21:35.560000",
      "content": "<p>Hi,<br>\nFor me, I'm working with 50/50 balanced training dataset and standard BCE loss (so with no weighting). LB always higher than CV… Maybe I've found the right train/val division (by chance). However, if I were you, I'd suggest you to concentrate more on hard examples and data augmentation.</p>\n<p>Happy kaggling!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2072853": "To deal with the unbalanced dataset, I'm now using BCE loss with positive weighted.\nAnd I also saw some ideas using Focal lose to solve this.\nAny performance difference on these two? Which one is better on CV problem?",
    "2073328": "BCELosswithlogits + weight works better then Focal Loss for me. I use custom sampler to oversample positive cases too.\nHowever I am not TOP (my single model is no more then 0.4) on LB :) so I can be totally wrong :) ",
    "2073220": "I cant tell you in particular which is working better in this competition. \n\nYou can also try different things which work well with class imbalance e.g.:\n\nLoss e.g.:\n-Bi-Tempered Logistic Loss\n\n…\n\nOther methods e.g.:\n-Label smoothing\n-Minoririty augmentation\n-Oversampling\n-Form it into an AnomalyDetection use case\n\n …\n\nThere are a lot of different options available.",
    "2077875": "Congratulations on your progress so far with the competition.\n\nThe best answer for all question of this type is: Try and see what works best..\nFrom my experience, I find that I usually get better results when using a combination of label smoothing and sometimes also focal-loss. \nGood luck!\n\nThe Devastator.\n",
    "2083840": "Hi!\nCongratulations on your progress!\n\nWhich dataset do you use? ROI 1024 or just no ROI 512. I found that on 512 no ROI dataset, I could get 0.24LB by using ensemble + weighted loss + upsample. Here is my code.\n\nhttps://www.kaggle.com/code/fanyang99/train-no-roi-512-lb0-24\n\nHope can get some suggestion. Should I go on to next dataset like ROI to do test? Or just try to improve methods on this dataset? Do you think 0.24 is already good on 512 data?\n\nThank you so much!",
    "2075867": "Hi,\nFor me, I'm working with 50/50 balanced training dataset and standard BCE loss (so with no weighting). LB always higher than CV... Maybe I've found the right train/val division (by chance). However, if I were you, I'd suggest you to concentrate more on hard examples and data augmentation.\n\nHappy kaggling!"
  }
}