{
  "id": 430684,
  "title": "26th Place Solution",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430684",
  "author_name": "Shashwat Raman",
  "post_date": "2023-08-10T18:42:21.703000",
  "votes": 23,
  "comment_count": 8,
  "views": 0,
  "content": "<p>First of all I would like to thank the organisers and Kaggle for this great competition. It was an amazing learning experience, I got to learn a lot of things. Congratulations to all the winners!</p>\n<h1>Summary</h1>\n<ul>\n<li>Data: Ash Color Images and Soft Labels (Average of individual labels)</li>\n<li>Image Size: 512</li>\n<li>Cross Validation: Split the Train Set to 5 folds after sorting by time.</li>\n<li>Ensemble of 3 Models</li>\n<li>Encoders: efficientnet-b5, regnety_120, seresnextaa101d_32x8d</li>\n<li>Decoder: Unet</li>\n</ul>\n<h2>Data</h2>\n<p>I trained on false color images and used soft labels. This gave me a big boost in both cv and lb score as compared to when training on the pixel masks. I think this was one of the most important tricks for this competiton. It was very surprising for me to see that the model was able to capture the uncertainty depicted by the probabilities of the pixels. As individual labels were not available for the valid set, so I didnt include it in training. I used it as hold set.</p>\n<h2>Cross Validation</h2>\n<p>This was a bit tricky due to the presence of many duplicates in the training set. So, to build a robust local validation, I first sorted the training set by date and time which was available in train.json and then split the sorted data into 5 folds without shuffling. After this, I clipped the valid folds during training like this:</p>\n<pre><code> self.fold == :\n    df = df[:-]\n self.fold == :\n    df = df[:]\n:\n    df = df[:-]\n</code></pre>\n<p>The cv score was not at all correlated with the leaderboard score, but still I trusted my cv in the end to select my submissions.</p>\n<h2>Models</h2>\n<p>I tried many encoders but used efficientnet-b5, regnety_120 and seresnextaa101d_32x8d in the end as they gave the best cv score. I only used Unet for all the encoders.</p>\n<h2>Training</h2>\n<ul>\n<li>Epochs: 30</li>\n<li>Learning Rate: Between 3e-4 to 3e-3 (Different for different Encoders)</li>\n<li>Optimizer: Adam</li>\n<li>Scheduler: CosineAnnealingLR (I saved the epoch with the best valid fold score)</li>\n<li>Loss: BCE Loss (As I was training with soft labels)</li>\n<li>Augmentations: RandomResizedCrop, ShiftScaleRotate (HFlip, VFlip and RandomRotate90 didn't work for me also, but I had no idea that this can be due to some problem in the masks)</li>\n</ul>\n<h3>Trick to reduce training time</h3>\n<p>To reduce training time, I tried to remove some images during training. Removing all the non-contrail images reduced the local cv score. Then I predicted the train set with one of my best model and made a column with the highest pixel probability predicted by the model for the images. Then I removed the images which had max pixel probability lower than 0.05 and had no contrail.</p>\n<p><code>df = df[~((df['max_prob_pred'] &lt; 0.05) &amp; (df['contrail'] == 0))]</code></p>\n<p>Using this I was able to remove about 4000 images from training, reduce training time a lot and it didn't effect the cv score at all.</p>\n<h2>Ensemble</h2>\n<p>I used my cv folds to get the ensemble weights for my final models.<br>\nWeighted ensembling: 0.3172[eb5] + 0.2928[regty120] + 0.39[sr101].</p>\n<p>For final prediction, first I ensembled the different models separately for each fold. Then I converted the 5 fold predictions to binary predictions using the thresholds I got using my valid sets. The thresholds were 0.49, 0.47, 0.47, 0.48, 0.46 for the folds from 0 to 4. Then I took the average of the 5 folds and used 0.3 as the final threshold (which I got using the valid hold set). </p>\n<p>My teammate's account was deleted just after I teamed up with him. I don't know if he did anything wrong or not. Whatever be the case, I wish him best of luck for his future.<br>\nI really hope I won't be removed from the leaderboard after verification. Let's see what happens ☺️</p>",
  "messages": [
    {
      "id": 2384030,
      "postDate": "2023-08-10T18:42:21.703Z",
      "content": "<p>First of all I would like to thank the organisers and Kaggle for this great competition. It was an amazing learning experience, I got to learn a lot of things. Congratulations to all the winners!</p>\n<h1>Summary</h1>\n<ul>\n<li>Data: Ash Color Images and Soft Labels (Average of individual labels)</li>\n<li>Image Size: 512</li>\n<li>Cross Validation: Split the Train Set to 5 folds after sorting by time.</li>\n<li>Ensemble of 3 Models</li>\n<li>Encoders: efficientnet-b5, regnety_120, seresnextaa101d_32x8d</li>\n<li>Decoder: Unet</li>\n</ul>\n<h2>Data</h2>\n<p>I trained on false color images and used soft labels. This gave me a big boost in both cv and lb score as compared to when training on the pixel masks. I think this was one of the most important tricks for this competiton. It was very surprising for me to see that the model was able to capture the uncertainty depicted by the probabilities of the pixels. As individual labels were not available for the valid set, so I didnt include it in training. I used it as hold set.</p>\n<h2>Cross Validation</h2>\n<p>This was a bit tricky due to the presence of many duplicates in the training set. So, to build a robust local validation, I first sorted the training set by date and time which was available in train.json and then split the sorted data into 5 folds without shuffling. After this, I clipped the valid folds during training like this:</p>\n<pre><code> self.fold == :\n    df = df[:-]\n self.fold == :\n    df = df[:]\n:\n    df = df[:-]\n</code></pre>\n<p>The cv score was not at all correlated with the leaderboard score, but still I trusted my cv in the end to select my submissions.</p>\n<h2>Models</h2>\n<p>I tried many encoders but used efficientnet-b5, regnety_120 and seresnextaa101d_32x8d in the end as they gave the best cv score. I only used Unet for all the encoders.</p>\n<h2>Training</h2>\n<ul>\n<li>Epochs: 30</li>\n<li>Learning Rate: Between 3e-4 to 3e-3 (Different for different Encoders)</li>\n<li>Optimizer: Adam</li>\n<li>Scheduler: CosineAnnealingLR (I saved the epoch with the best valid fold score)</li>\n<li>Loss: BCE Loss (As I was training with soft labels)</li>\n<li>Augmentations: RandomResizedCrop, ShiftScaleRotate (HFlip, VFlip and RandomRotate90 didn't work for me also, but I had no idea that this can be due to some problem in the masks)</li>\n</ul>\n<h3>Trick to reduce training time</h3>\n<p>To reduce training time, I tried to remove some images during training. Removing all the non-contrail images reduced the local cv score. Then I predicted the train set with one of my best model and made a column with the highest pixel probability predicted by the model for the images. Then I removed the images which had max pixel probability lower than 0.05 and had no contrail.</p>\n<p><code>df = df[~((df['max_prob_pred'] &lt; 0.05) &amp; (df['contrail'] == 0))]</code></p>\n<p>Using this I was able to remove about 4000 images from training, reduce training time a lot and it didn't effect the cv score at all.</p>\n<h2>Ensemble</h2>\n<p>I used my cv folds to get the ensemble weights for my final models.<br>\nWeighted ensembling: 0.3172[eb5] + 0.2928[regty120] + 0.39[sr101].</p>\n<p>For final prediction, first I ensembled the different models separately for each fold. Then I converted the 5 fold predictions to binary predictions using the thresholds I got using my valid sets. The thresholds were 0.49, 0.47, 0.47, 0.48, 0.46 for the folds from 0 to 4. Then I took the average of the 5 folds and used 0.3 as the final threshold (which I got using the valid hold set). </p>\n<p>My teammate's account was deleted just after I teamed up with him. I don't know if he did anything wrong or not. Whatever be the case, I wish him best of luck for his future.<br>\nI really hope I won't be removed from the leaderboard after verification. Let's see what happens ☺️</p>",
      "rawMarkdown": "First of all I would like to thank the organisers and Kaggle for this great competition. It was an amazing learning experience, I got to learn a lot of things. Congratulations to all the winners!\n\n# Summary\n\n- Data: Ash Color Images and Soft Labels (Average of individual labels)\n- Image Size: 512\n- Cross Validation: Split the Train Set to 5 folds after sorting by time.\n- Ensemble of 3 Models\n- Encoders: efficientnet-b5, regnety_120, seresnextaa101d_32x8d\n- Decoder: Unet\n\n## Data\n\nI trained on false color images and used soft labels. This gave me a big boost in both cv and lb score as compared to when training on the pixel masks. I think this was one of the most important tricks for this competiton. It was very surprising for me to see that the model was able to capture the uncertainty depicted by the probabilities of the pixels. As individual labels were not available for the valid set, so I didnt include it in training. I used it as hold set.\n\n## Cross Validation\n\nThis was a bit tricky due to the presence of many duplicates in the training set. So, to build a robust local validation, I first sorted the training set by date and time which was available in train.json and then split the sorted data into 5 folds without shuffling. After this, I clipped the valid folds during training like this:\n\n```python\nif self.fold == 0:\n    df = df[:-2000]\nelif self.fold == 4:\n    df = df[2000:]\nelse:\n    df = df[1000:-1000]\n```\n\nThe cv score was not at all correlated with the leaderboard score, but still I trusted my cv in the end to select my submissions.\n\n## Models\n\nI tried many encoders but used efficientnet-b5, regnety_120 and seresnextaa101d_32x8d in the end as they gave the best cv score. I only used Unet for all the encoders.\n\n## Training\n\n- Epochs: 30\n- Learning Rate: Between 3e-4 to 3e-3 (Different for different Encoders)\n- Optimizer: Adam\n- Scheduler: CosineAnnealingLR (I saved the epoch with the best valid fold score)\n- Loss: BCE Loss (As I was training with soft labels)\n- Augmentations: RandomResizedCrop, ShiftScaleRotate (HFlip, VFlip and RandomRotate90 didn't work for me also, but I had no idea that this can be due to some problem in the masks)\n\n### Trick to reduce training time\nTo reduce training time, I tried to remove some images during training. Removing all the non-contrail images reduced the local cv score. Then I predicted the train set with one of my best model and made a column with the highest pixel probability predicted by the model for the images. Then I removed the images which had max pixel probability lower than 0.05 and had no contrail.\n\n`df = df[~((df['max_prob_pred'] < 0.05) & (df['contrail'] == 0))]`\n\nUsing this I was able to remove about 4000 images from training, reduce training time a lot and it didn't effect the cv score at all.\n\n## Ensemble\n\nI used my cv folds to get the ensemble weights for my final models.\nWeighted ensembling: 0.3172[eb5] + 0.2928[regty120] + 0.39[sr101].\n\nFor final prediction, first I ensembled the different models separately for each fold. Then I converted the 5 fold predictions to binary predictions using the thresholds I got using my valid sets. The thresholds were 0.49, 0.47, 0.47, 0.48, 0.46 for the folds from 0 to 4. Then I took the average of the 5 folds and used 0.3 as the final threshold (which I got using the valid hold set). \n\n\nMy teammate's account was deleted just after I teamed up with him. I don't know if he did anything wrong or not. Whatever be the case, I wish him best of luck for his future.\nI really hope I won't be removed from the leaderboard after verification. Let's see what happens ☺️",
      "votes": 22
    },
    {
      "id": 2387031,
      "postDate": "2023-08-12T10:58:17.557Z",
      "content": "<p>Congrats Shashwat,<br>\nWell deserved. The solution is clean and very intuitive. Hope we can team up in the future.</p>",
      "rawMarkdown": "Congrats Shashwat,\nWell deserved. The solution is clean and very intuitive. Hope we can team up in the future.",
      "votes": 1,
      "replies": [
        {
          "id": 2387482,
          "postDate": "2023-08-12T17:35:36.973Z",
          "content": "<p>Thank you very much <a href=\"https://www.kaggle.com/thejavanka\" target=\"_blank\">@thejavanka</a></p>",
          "rawMarkdown": "Thank you very much @thejavanka"
        }
      ]
    },
    {
      "id": 2385339,
      "postDate": "2023-08-11T09:21:45.687Z",
      "content": "<p>Congrats!<br>\nCool trick with removing 4000 images for extra speed up.<br>\nWhen you talk about weights ensembling do you mean that ensemble prediction masks outputs with different weights?<br>\nDid you plan to release your solution?</p>",
      "rawMarkdown": "Congrats!\nCool trick with removing 4000 images for extra speed up.\nWhen you talk about weights ensembling do you mean that ensemble prediction masks outputs with different weights?\nDid you plan to release your solution?",
      "votes": 1,
      "replies": [
        {
          "id": 2386207,
          "postDate": "2023-08-11T19:21:52.973Z",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/kuddai\" target=\"_blank\">@kuddai</a> <br>\nYes I mean that only. I will publish my code in a few days.</p>",
          "rawMarkdown": "Thank you @kuddai \nYes I mean that only. I will publish my code in a few days."
        }
      ]
    },
    {
      "id": 2384147,
      "postDate": "2023-08-10T21:02:08.910Z",
      "content": "<p>Very nice and clean solution!</p>\n<p>How are you able to do so much training and testing on your CV? Do you have your own GPU or did Kaggle's suffice? I ask because running your train notebook took me an hour every time and ran out of GPU hours.</p>\n<p>Could you also elaborate on your cross-validation as well? Did you clip that way because of the duplicates?</p>",
      "rawMarkdown": "Very nice and clean solution!\n\nHow are you able to do so much training and testing on your CV? Do you have your own GPU or did Kaggle's suffice? I ask because running your train notebook took me an hour every time and ran out of GPU hours.\n\nCould you also elaborate on your cross-validation as well? Did you clip that way because of the duplicates?",
      "votes": 1,
      "replies": [
        {
          "id": 2384777,
          "postDate": "2023-08-11T04:34:16.250Z",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/echen333us\" target=\"_blank\">@echen333us</a>,<br>\nI used JavisLabs' A5000 GPU for all my experiments. It worked really well for me. Kaggle is also great but I find the 30 hour GPU limit very low. <br>\nYes. There were many duplicates present in the training set. A satellite image from the same location after a few hours can be very similar to the ones before. So sorting, splitting and clipping like this ensured that training and validation are far from each other with respect to time.</p>",
          "rawMarkdown": "Thank you @echen333us,\nI used JavisLabs' A5000 GPU for all my experiments. It worked really well for me. Kaggle is also great but I find the 30 hour GPU limit very low. \nYes. There were many duplicates present in the training set. A satellite image from the same location after a few hours can be very similar to the ones before. So sorting, splitting and clipping like this ensured that training and validation are far from each other with respect to time.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2384039,
      "postDate": "2023-08-10T18:55:08.747Z",
      "content": "<p>Congratulations for 26th place! </p>\n<p>Fingers crossed that you won't get removed from the leaderboard due to bad team-up! I like that you have cared about the image similarities in the data when setting up your validation folds. We experimented with it as well and were able to find a very good correlation between CV and LB by grouping similar images into same folds. If you are interested in it we published a notebook about it. </p>\n<p>I also like the trick of removing empty masks… have to try it out with our balanced CV set as well to check whether our score would still have been almost the same as well.</p>\n<p>Do you remember how much boost you got by using soft labels with bce loss?</p>",
      "rawMarkdown": "Congratulations for 26th place! \n\nFingers crossed that you won't get removed from the leaderboard due to bad team-up! I like that you have cared about the image similarities in the data when setting up your validation folds. We experimented with it as well and were able to find a very good correlation between CV and LB by grouping similar images into same folds. If you are interested in it we published a notebook about it. \n\nI also like the trick of removing empty masks... have to try it out with our balanced CV set as well to check whether our score would still have been almost the same as well.\n\nDo you remember how much boost you got by using soft labels with bce loss?",
      "votes": 2,
      "replies": [
        {
          "id": 2384748,
          "postDate": "2023-08-11T04:20:41.337Z",
          "content": "<p>Thank you very much <a href=\"https://www.kaggle.com/allunia\" target=\"_blank\">@allunia</a> <br>\nCongratulations for 52nd Place!<br>\nSure, I will take a look at the notebook. Thanks!<br>\nAfter using soft labels, the score boost was around 0.02-0.03, if I remember correctly!</p>",
          "rawMarkdown": "Thank you very much @allunia \nCongratulations for 52nd Place!\nSure, I will take a look at the notebook. Thanks!\nAfter using soft labels, the score boost was around 0.02-0.03, if I remember correctly!",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2387031,
      "author_name": "Kenni",
      "author_url": "",
      "post_date": "2023-08-12T10:58:17.557000",
      "content": "<p>Congrats Shashwat,<br>\nWell deserved. The solution is clean and very intuitive. Hope we can team up in the future.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2387482,
          "author_name": "Shashwat Raman",
          "author_url": "",
          "post_date": "2023-08-12T17:35:36.973000",
          "content": "<p>Thank you very much <a href=\"https://www.kaggle.com/thejavanka\" target=\"_blank\">@thejavanka</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2385339,
      "author_name": "iadduk",
      "author_url": "",
      "post_date": "2023-08-11T09:21:45.687000",
      "content": "<p>Congrats!<br>\nCool trick with removing 4000 images for extra speed up.<br>\nWhen you talk about weights ensembling do you mean that ensemble prediction masks outputs with different weights?<br>\nDid you plan to release your solution?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2386207,
          "author_name": "Shashwat Raman",
          "author_url": "",
          "post_date": "2023-08-11T19:21:52.973000",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/kuddai\" target=\"_blank\">@kuddai</a> <br>\nYes I mean that only. I will publish my code in a few days.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2384147,
      "author_name": "Eddy",
      "author_url": "",
      "post_date": "2023-08-10T21:02:08.910000",
      "content": "<p>Very nice and clean solution!</p>\n<p>How are you able to do so much training and testing on your CV? Do you have your own GPU or did Kaggle's suffice? I ask because running your train notebook took me an hour every time and ran out of GPU hours.</p>\n<p>Could you also elaborate on your cross-validation as well? Did you clip that way because of the duplicates?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2384777,
          "author_name": "Shashwat Raman",
          "author_url": "",
          "post_date": "2023-08-11T04:34:16.250000",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/echen333us\" target=\"_blank\">@echen333us</a>,<br>\nI used JavisLabs' A5000 GPU for all my experiments. It worked really well for me. Kaggle is also great but I find the 30 hour GPU limit very low. <br>\nYes. There were many duplicates present in the training set. A satellite image from the same location after a few hours can be very similar to the ones before. So sorting, splitting and clipping like this ensured that training and validation are far from each other with respect to time.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2384039,
      "author_name": "Laura Fink",
      "author_url": "",
      "post_date": "2023-08-10T18:55:08.747000",
      "content": "<p>Congratulations for 26th place! </p>\n<p>Fingers crossed that you won't get removed from the leaderboard due to bad team-up! I like that you have cared about the image similarities in the data when setting up your validation folds. We experimented with it as well and were able to find a very good correlation between CV and LB by grouping similar images into same folds. If you are interested in it we published a notebook about it. </p>\n<p>I also like the trick of removing empty masks… have to try it out with our balanced CV set as well to check whether our score would still have been almost the same as well.</p>\n<p>Do you remember how much boost you got by using soft labels with bce loss?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2384748,
          "author_name": "Shashwat Raman",
          "author_url": "",
          "post_date": "2023-08-11T04:20:41.337000",
          "content": "<p>Thank you very much <a href=\"https://www.kaggle.com/allunia\" target=\"_blank\">@allunia</a> <br>\nCongratulations for 52nd Place!<br>\nSure, I will take a look at the notebook. Thanks!<br>\nAfter using soft labels, the score boost was around 0.02-0.03, if I remember correctly!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2384030": "First of all I would like to thank the organisers and Kaggle for this great competition. It was an amazing learning experience, I got to learn a lot of things. Congratulations to all the winners!\n\n# Summary\n\n- Data: Ash Color Images and Soft Labels (Average of individual labels)\n- Image Size: 512\n- Cross Validation: Split the Train Set to 5 folds after sorting by time.\n- Ensemble of 3 Models\n- Encoders: efficientnet-b5, regnety_120, seresnextaa101d_32x8d\n- Decoder: Unet\n\n## Data\n\nI trained on false color images and used soft labels. This gave me a big boost in both cv and lb score as compared to when training on the pixel masks. I think this was one of the most important tricks for this competiton. It was very surprising for me to see that the model was able to capture the uncertainty depicted by the probabilities of the pixels. As individual labels were not available for the valid set, so I didnt include it in training. I used it as hold set.\n\n## Cross Validation\n\nThis was a bit tricky due to the presence of many duplicates in the training set. So, to build a robust local validation, I first sorted the training set by date and time which was available in train.json and then split the sorted data into 5 folds without shuffling. After this, I clipped the valid folds during training like this:\n\n```python\nif self.fold == 0:\n    df = df[:-2000]\nelif self.fold == 4:\n    df = df[2000:]\nelse:\n    df = df[1000:-1000]\n```\n\nThe cv score was not at all correlated with the leaderboard score, but still I trusted my cv in the end to select my submissions.\n\n## Models\n\nI tried many encoders but used efficientnet-b5, regnety_120 and seresnextaa101d_32x8d in the end as they gave the best cv score. I only used Unet for all the encoders.\n\n## Training\n\n- Epochs: 30\n- Learning Rate: Between 3e-4 to 3e-3 (Different for different Encoders)\n- Optimizer: Adam\n- Scheduler: CosineAnnealingLR (I saved the epoch with the best valid fold score)\n- Loss: BCE Loss (As I was training with soft labels)\n- Augmentations: RandomResizedCrop, ShiftScaleRotate (HFlip, VFlip and RandomRotate90 didn't work for me also, but I had no idea that this can be due to some problem in the masks)\n\n### Trick to reduce training time\nTo reduce training time, I tried to remove some images during training. Removing all the non-contrail images reduced the local cv score. Then I predicted the train set with one of my best model and made a column with the highest pixel probability predicted by the model for the images. Then I removed the images which had max pixel probability lower than 0.05 and had no contrail.\n\n`df = df[~((df['max_prob_pred'] < 0.05) & (df['contrail'] == 0))]`\n\nUsing this I was able to remove about 4000 images from training, reduce training time a lot and it didn't effect the cv score at all.\n\n## Ensemble\n\nI used my cv folds to get the ensemble weights for my final models.\nWeighted ensembling: 0.3172[eb5] + 0.2928[regty120] + 0.39[sr101].\n\nFor final prediction, first I ensembled the different models separately for each fold. Then I converted the 5 fold predictions to binary predictions using the thresholds I got using my valid sets. The thresholds were 0.49, 0.47, 0.47, 0.48, 0.46 for the folds from 0 to 4. Then I took the average of the 5 folds and used 0.3 as the final threshold (which I got using the valid hold set). \n\n\nMy teammate's account was deleted just after I teamed up with him. I don't know if he did anything wrong or not. Whatever be the case, I wish him best of luck for his future.\nI really hope I won't be removed from the leaderboard after verification. Let's see what happens ☺️",
    "2387031": "Congrats Shashwat,\nWell deserved. The solution is clean and very intuitive. Hope we can team up in the future.",
    "2385339": "Congrats!\nCool trick with removing 4000 images for extra speed up.\nWhen you talk about weights ensembling do you mean that ensemble prediction masks outputs with different weights?\nDid you plan to release your solution?",
    "2384147": "Very nice and clean solution!\n\nHow are you able to do so much training and testing on your CV? Do you have your own GPU or did Kaggle's suffice? I ask because running your train notebook took me an hour every time and ran out of GPU hours.\n\nCould you also elaborate on your cross-validation as well? Did you clip that way because of the duplicates?",
    "2384039": "Congratulations for 26th place! \n\nFingers crossed that you won't get removed from the leaderboard due to bad team-up! I like that you have cared about the image similarities in the data when setting up your validation folds. We experimented with it as well and were able to find a very good correlation between CV and LB by grouping similar images into same folds. If you are interested in it we published a notebook about it. \n\nI also like the trick of removing empty masks... have to try it out with our balanced CV set as well to check whether our score would still have been almost the same as well.\n\nDo you remember how much boost you got by using soft labels with bce loss?"
  }
}