{
  "id": 435176,
  "title": "73rd Place Solution",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/435176",
  "author_name": "Raki",
  "post_date": "2023-08-28T08:51:14.001000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Thanks to the organizers and Kaggle for hosting, I really liked working on it. Also thanks to many Kagglers who shared great ideas with the community!</p>\n<p>I did a larger project write-up on my <a href=\"https://github.com/rachiki/Finding-A-Great-ML-Job/tree/main/Code%20and%20Projects/Google%20Research%20-%20Identify%20Contrails%20in%20Satellite%20Images\" target=\"_blank\">github</a> and only put a summary here.</p>\n<h2>Overview</h2>\n<ul>\n<li>Architecture: Ensemble of 12 DL models, mostly UNet+ResNeSt </li>\n<li>Input: Ash Color Scheme</li>\n<li>Target Labels: Soft labels averaged from annotators</li>\n<li>Loss Function: Custom Dice loss</li>\n<li>Time-Dependent Data: Utilized in the form of input sequences, but not effectively implemented in the architecture</li>\n<li>Data Augmentation: Almost none (due to the 0.5 pixel shift, which I did not find).</li>\n</ul>\n<h2>Custom Dice Loss</h2>\n<p>The standard dice loss performed better than WBCE for me on hard labels, but struggled with soft labels. <br>\nI looked at the dice loss calculation more closely and noticed that it does not have local minima at the perfect solutions in most cases. This is described in more detail <a href=\"https://www.kaggle.com/code/raki21/adjusted-dice-loss-for-soft-labels\" target=\"_blank\">in this notebook</a>. I found that soft labels work really well with an adjusted dice loss.<br>\nThese are my val dice score results with optimal thresholds for different losses and labels on a ResNeSt26d+UNet baseline.</p>\n<pre><code> + dice:  . at .\n + WBCE:  . at .\n\n + dice:  . at .\n + WBCE:  . at .\n + cdice: . at .\n</code></pre>\n<h2>Data Augmentation</h2>\n<p>I only used rotations up to 10 degrees and 0.1 scale adjustment. I rationalized the problems with data augmentation away, thinking that the model gains geographical information like:<br>\n      Contrail distribution =&gt; Flight routes =&gt; Location =&gt; Information about contrail formation.<br>\nThis was not the main, as it turned out. I should check these kind of problems in more detail in the future!</p>\n<h2>Project Timeline</h2>\n<p>I joined the competition during its final week, which led to a very condensed and focused period of development:</p>\n<ul>\n<li>Day 1-2: Data overview and exploratory analysis</li>\n<li>Day 3-6: Model pipeline and optimization</li>\n<li>Day 7: Model ensembling and final submission</li>\n</ul>\n<h2>Models</h2>\n<p>Framework Setup<br>\nBuilt using Segmentation Models PyTorch, I tested out many of the standard backbones and some timm models, like coatv3 and all decoders.<br>\nThe only worthwhile decoders for me where UNet/++ and DeepLabV3/+.<br>\nAs encoders I stayed mostly with ResNeSt and trained some EfficientNet. Most used model was ResNeSt200e.</p>\n<h2>Upscale</h2>\n<p>I did not have enough compute for training 512x512 or greater models, so I trained 256 and 384 only.</p>\n<h2>Ensembling and Submission</h2>\n<p>Looked through all my checkpoints and combined 12 models in the end. If I had more good checkpoints I think it would have scaled even further with model number. Climbed from around 150th place to 70th in the ranking with ensembling. I was surprised it worked so good.</p>\n<h2>Conclusion</h2>\n<p>Despite the time constraints, I was able to achieve a respectable ranking. The key takeaway was the need to investigate issues in data augmentation deeply.</p>\n<h3>Things I couldn't explore in depth in the given time</h3>\n<p>Better utilization of time-dependent data, larger models and bigger upscale.<br>\nExploration of pseudo-labeling and post-processing.</p>",
  "messages": [
    {
      "id": 2412400,
      "postDate": "2023-08-28T08:51:14Z",
      "content": "<p>Thanks to the organizers and Kaggle for hosting, I really liked working on it. Also thanks to many Kagglers who shared great ideas with the community!</p>\n<p>I did a larger project write-up on my <a href=\"https://github.com/rachiki/Finding-A-Great-ML-Job/tree/main/Code%20and%20Projects/Google%20Research%20-%20Identify%20Contrails%20in%20Satellite%20Images\" target=\"_blank\">github</a> and only put a summary here.</p>\n<h2>Overview</h2>\n<ul>\n<li>Architecture: Ensemble of 12 DL models, mostly UNet+ResNeSt </li>\n<li>Input: Ash Color Scheme</li>\n<li>Target Labels: Soft labels averaged from annotators</li>\n<li>Loss Function: Custom Dice loss</li>\n<li>Time-Dependent Data: Utilized in the form of input sequences, but not effectively implemented in the architecture</li>\n<li>Data Augmentation: Almost none (due to the 0.5 pixel shift, which I did not find).</li>\n</ul>\n<h2>Custom Dice Loss</h2>\n<p>The standard dice loss performed better than WBCE for me on hard labels, but struggled with soft labels. <br>\nI looked at the dice loss calculation more closely and noticed that it does not have local minima at the perfect solutions in most cases. This is described in more detail <a href=\"https://www.kaggle.com/code/raki21/adjusted-dice-loss-for-soft-labels\" target=\"_blank\">in this notebook</a>. I found that soft labels work really well with an adjusted dice loss.<br>\nThese are my val dice score results with optimal thresholds for different losses and labels on a ResNeSt26d+UNet baseline.</p>\n<pre><code> + dice:  . at .\n + WBCE:  . at .\n\n + dice:  . at .\n + WBCE:  . at .\n + cdice: . at .\n</code></pre>\n<h2>Data Augmentation</h2>\n<p>I only used rotations up to 10 degrees and 0.1 scale adjustment. I rationalized the problems with data augmentation away, thinking that the model gains geographical information like:<br>\n      Contrail distribution =&gt; Flight routes =&gt; Location =&gt; Information about contrail formation.<br>\nThis was not the main, as it turned out. I should check these kind of problems in more detail in the future!</p>\n<h2>Project Timeline</h2>\n<p>I joined the competition during its final week, which led to a very condensed and focused period of development:</p>\n<ul>\n<li>Day 1-2: Data overview and exploratory analysis</li>\n<li>Day 3-6: Model pipeline and optimization</li>\n<li>Day 7: Model ensembling and final submission</li>\n</ul>\n<h2>Models</h2>\n<p>Framework Setup<br>\nBuilt using Segmentation Models PyTorch, I tested out many of the standard backbones and some timm models, like coatv3 and all decoders.<br>\nThe only worthwhile decoders for me where UNet/++ and DeepLabV3/+.<br>\nAs encoders I stayed mostly with ResNeSt and trained some EfficientNet. Most used model was ResNeSt200e.</p>\n<h2>Upscale</h2>\n<p>I did not have enough compute for training 512x512 or greater models, so I trained 256 and 384 only.</p>\n<h2>Ensembling and Submission</h2>\n<p>Looked through all my checkpoints and combined 12 models in the end. If I had more good checkpoints I think it would have scaled even further with model number. Climbed from around 150th place to 70th in the ranking with ensembling. I was surprised it worked so good.</p>\n<h2>Conclusion</h2>\n<p>Despite the time constraints, I was able to achieve a respectable ranking. The key takeaway was the need to investigate issues in data augmentation deeply.</p>\n<h3>Things I couldn't explore in depth in the given time</h3>\n<p>Better utilization of time-dependent data, larger models and bigger upscale.<br>\nExploration of pseudo-labeling and post-processing.</p>",
      "rawMarkdown": "Thanks to the organizers and Kaggle for hosting, I really liked working on it. Also thanks to many Kagglers who shared great ideas with the community!\n\nI did a larger project write-up on my [github](https://github.com/rachiki/Finding-A-Great-ML-Job/tree/main/Code%20and%20Projects/Google%20Research%20-%20Identify%20Contrails%20in%20Satellite%20Images) and only put a summary here.\n\n## Overview\n- Architecture: Ensemble of 12 DL models, mostly UNet+ResNeSt \n- Input: Ash Color Scheme\n- Target Labels: Soft labels averaged from annotators\n- Loss Function: Custom Dice loss\n- Time-Dependent Data: Utilized in the form of input sequences, but not effectively implemented in the architecture\n- Data Augmentation: Almost none (due to the 0.5 pixel shift, which I did not find).\n\n## Custom Dice Loss\nThe standard dice loss performed better than WBCE for me on hard labels, but struggled with soft labels. \nI looked at the dice loss calculation more closely and noticed that it does not have local minima at the perfect solutions in most cases. This is described in more detail [in this notebook](https://www.kaggle.com/code/raki21/adjusted-dice-loss-for-soft-labels). I found that soft labels work really well with an adjusted dice loss.\nThese are my val dice score results with optimal thresholds for different losses and labels on a ResNeSt26d+UNet baseline.\n\n    Hard + dice:  0.621 at 0.17\n    Hard + WBCE:  0.618 at 0.85\n\n    Soft + dice:  0.629 at 0.999\n    Soft + WBCE:  0.621 at 0.9\n    Soft + cdice: 0.633 at 0.62\n\n\n## Data Augmentation\nI only used rotations up to 10 degrees and 0.1 scale adjustment. I rationalized the problems with data augmentation away, thinking that the model gains geographical information like:\n      Contrail distribution => Flight routes => Location => Information about contrail formation.\nThis was not the main, as it turned out. I should check these kind of problems in more detail in the future!\n\n## Project Timeline\nI joined the competition during its final week, which led to a very condensed and focused period of development:\n- Day 1-2: Data overview and exploratory analysis\n- Day 3-6: Model pipeline and optimization\n- Day 7: Model ensembling and final submission\n\n## Models\nFramework Setup\nBuilt using Segmentation Models PyTorch, I tested out many of the standard backbones and some timm models, like coatv3 and all decoders.\nThe only worthwhile decoders for me where UNet/++ and DeepLabV3/+.\nAs encoders I stayed mostly with ResNeSt and trained some EfficientNet. Most used model was ResNeSt200e.\n\n## Upscale\nI did not have enough compute for training 512x512 or greater models, so I trained 256 and 384 only.\n\n## Ensembling and Submission\nLooked through all my checkpoints and combined 12 models in the end. If I had more good checkpoints I think it would have scaled even further with model number. Climbed from around 150th place to 70th in the ranking with ensembling. I was surprised it worked so good.\n\n## Conclusion\nDespite the time constraints, I was able to achieve a respectable ranking. The key takeaway was the need to investigate issues in data augmentation deeply.\n\n### Things I couldn't explore in depth in the given time\nBetter utilization of time-dependent data, larger models and bigger upscale.\nExploration of pseudo-labeling and post-processing.",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2412400": "Thanks to the organizers and Kaggle for hosting, I really liked working on it. Also thanks to many Kagglers who shared great ideas with the community!\n\nI did a larger project write-up on my [github](https://github.com/rachiki/Finding-A-Great-ML-Job/tree/main/Code%20and%20Projects/Google%20Research%20-%20Identify%20Contrails%20in%20Satellite%20Images) and only put a summary here.\n\n## Overview\n- Architecture: Ensemble of 12 DL models, mostly UNet+ResNeSt \n- Input: Ash Color Scheme\n- Target Labels: Soft labels averaged from annotators\n- Loss Function: Custom Dice loss\n- Time-Dependent Data: Utilized in the form of input sequences, but not effectively implemented in the architecture\n- Data Augmentation: Almost none (due to the 0.5 pixel shift, which I did not find).\n\n## Custom Dice Loss\nThe standard dice loss performed better than WBCE for me on hard labels, but struggled with soft labels. \nI looked at the dice loss calculation more closely and noticed that it does not have local minima at the perfect solutions in most cases. This is described in more detail [in this notebook](https://www.kaggle.com/code/raki21/adjusted-dice-loss-for-soft-labels). I found that soft labels work really well with an adjusted dice loss.\nThese are my val dice score results with optimal thresholds for different losses and labels on a ResNeSt26d+UNet baseline.\n\n    Hard + dice:  0.621 at 0.17\n    Hard + WBCE:  0.618 at 0.85\n\n    Soft + dice:  0.629 at 0.999\n    Soft + WBCE:  0.621 at 0.9\n    Soft + cdice: 0.633 at 0.62\n\n\n## Data Augmentation\nI only used rotations up to 10 degrees and 0.1 scale adjustment. I rationalized the problems with data augmentation away, thinking that the model gains geographical information like:\n      Contrail distribution => Flight routes => Location => Information about contrail formation.\nThis was not the main, as it turned out. I should check these kind of problems in more detail in the future!\n\n## Project Timeline\nI joined the competition during its final week, which led to a very condensed and focused period of development:\n- Day 1-2: Data overview and exploratory analysis\n- Day 3-6: Model pipeline and optimization\n- Day 7: Model ensembling and final submission\n\n## Models\nFramework Setup\nBuilt using Segmentation Models PyTorch, I tested out many of the standard backbones and some timm models, like coatv3 and all decoders.\nThe only worthwhile decoders for me where UNet/++ and DeepLabV3/+.\nAs encoders I stayed mostly with ResNeSt and trained some EfficientNet. Most used model was ResNeSt200e.\n\n## Upscale\nI did not have enough compute for training 512x512 or greater models, so I trained 256 and 384 only.\n\n## Ensembling and Submission\nLooked through all my checkpoints and combined 12 models in the end. If I had more good checkpoints I think it would have scaled even further with model number. Climbed from around 150th place to 70th in the ranking with ensembling. I was surprised it worked so good.\n\n## Conclusion\nDespite the time constraints, I was able to achieve a respectable ranking. The key takeaway was the need to investigate issues in data augmentation deeply.\n\n### Things I couldn't explore in depth in the given time\nBetter utilization of time-dependent data, larger models and bigger upscale.\nExploration of pseudo-labeling and post-processing."
  }
}