{
  "id": 430543,
  "title": "8th place solution (OneFormer works)",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430543",
  "author_name": "ts",
  "post_date": "2023-08-10T08:13:42.577000",
  "votes": 32,
  "comment_count": 2,
  "views": 0,
  "content": "<h1><strong>Summary</strong></h1>\n<ul>\n<li>Models:&nbsp;<strong>OneFormer</strong>, effnet, MaxViT, resnetrs, nfnetf5</li>\n<li>Train on <strong>soft labels and pseudo label</strong></li>\n<li>Optimizing Ensemble Weights</li>\n<li>Loss: BCE</li>\n<li>Only random image cropping without TTA and without mask shift</li>\n</ul>\n<h1><strong>Introduction</strong></h1>\n<p>Our team would like to thank the organizers and Kaggle for making this competition possible. Also, I want to express my gratitude to my outstanding teammates&nbsp;<a href=\"https://www.kaggle.com/cnumber\" target=\"_blank\">@cnumber</a> and&nbsp;<a href=\"https://www.kaggle.com/yukiokumura1\" target=\"_blank\">@yukiokumura1</a> for their incredible contribution toward our final result.</p>\n<h1><strong>Details</strong></h1>\n<p>Our main strategy is the adoption and optimization of various models and the use of soft and pseudo labels.</p>\n<ul>\n<li><strong>Models:</strong></li>\n</ul>\n<p>In the beginning of the competition, we mainly experimented with a combination of smp and timm, but from the middle of the competition, we also started using OneFormer.</p>\n<p>OneFormer demonstrated the best results. <br>\n<strong>private lb score of single oneformer (dinat-l) : 0.70204</strong><br>\n<strong>With CV score of holdout dice, effnet b7: 0.671, oneformer(dinat-l): 0.693</strong>.</p>\n<p>Model resolutions are below: (inference settings)&nbsp;</p>\n<table>\n<thead>\n<tr>\n<th>OneFormer</th>\n<th>effnetb7, b8</th>\n<th>MaxViT-Tiny, Base</th>\n<th>resnetrs</th>\n<th>nfnetf5</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1024</td>\n<td>640</td>\n<td>512</td>\n<td>512</td>\n<td>640</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li><strong>Data: Soft and Pseudo Labels:</strong></li>\n</ul>\n<p>Using soft labels (the average of all annotator labels) during training.  We also introduced pseudo labels for some models. When both are used together, training is performed randomly on images with the 0th through 7th time frame, with soft label used only for the 4th time frame and pseudo label for all other cases.</p>\n<ul>\n<li><strong>Optimizing Ensemble Weights:</strong></li>\n</ul>\n<p>To maximized the performance of ensembling, we incorporated optimization using Optuna within the CV environment. This facilitated the efficient verification of the best model combinations.</p>\n<h1>Be open to consideration</h1>\n<ul>\n<li><strong>Interplay Between a Strong Backbone and Pseudo Labels:</strong></li>\n</ul>\n<p>The impact of using pseudo labels<strong>, when trained with pseudo label of when trained with pseudo labels *, effnet b7: 0.685, oneformer(dinat-l): 0.695</strong> *The average of each model trained in holdout is used as pseudo label. </p>\n<p>However, the enhancements were relatively modest with OneFormer.<br>\nOne hypothesis suggests that pseudo labels might bring about a knowledge distillation effect. The impact of this might reach a limit when the backbone is of a certain strength.</p>",
  "messages": [
    {
      "id": 2383181,
      "postDate": "2023-08-10T08:13:42.577Z",
      "content": "<h1><strong>Summary</strong></h1>\n<ul>\n<li>Models:&nbsp;<strong>OneFormer</strong>, effnet, MaxViT, resnetrs, nfnetf5</li>\n<li>Train on <strong>soft labels and pseudo label</strong></li>\n<li>Optimizing Ensemble Weights</li>\n<li>Loss: BCE</li>\n<li>Only random image cropping without TTA and without mask shift</li>\n</ul>\n<h1><strong>Introduction</strong></h1>\n<p>Our team would like to thank the organizers and Kaggle for making this competition possible. Also, I want to express my gratitude to my outstanding teammates&nbsp;<a href=\"https://www.kaggle.com/cnumber\" target=\"_blank\">@cnumber</a> and&nbsp;<a href=\"https://www.kaggle.com/yukiokumura1\" target=\"_blank\">@yukiokumura1</a> for their incredible contribution toward our final result.</p>\n<h1><strong>Details</strong></h1>\n<p>Our main strategy is the adoption and optimization of various models and the use of soft and pseudo labels.</p>\n<ul>\n<li><strong>Models:</strong></li>\n</ul>\n<p>In the beginning of the competition, we mainly experimented with a combination of smp and timm, but from the middle of the competition, we also started using OneFormer.</p>\n<p>OneFormer demonstrated the best results. <br>\n<strong>private lb score of single oneformer (dinat-l) : 0.70204</strong><br>\n<strong>With CV score of holdout dice, effnet b7: 0.671, oneformer(dinat-l): 0.693</strong>.</p>\n<p>Model resolutions are below: (inference settings)&nbsp;</p>\n<table>\n<thead>\n<tr>\n<th>OneFormer</th>\n<th>effnetb7, b8</th>\n<th>MaxViT-Tiny, Base</th>\n<th>resnetrs</th>\n<th>nfnetf5</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1024</td>\n<td>640</td>\n<td>512</td>\n<td>512</td>\n<td>640</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li><strong>Data: Soft and Pseudo Labels:</strong></li>\n</ul>\n<p>Using soft labels (the average of all annotator labels) during training.  We also introduced pseudo labels for some models. When both are used together, training is performed randomly on images with the 0th through 7th time frame, with soft label used only for the 4th time frame and pseudo label for all other cases.</p>\n<ul>\n<li><strong>Optimizing Ensemble Weights:</strong></li>\n</ul>\n<p>To maximized the performance of ensembling, we incorporated optimization using Optuna within the CV environment. This facilitated the efficient verification of the best model combinations.</p>\n<h1>Be open to consideration</h1>\n<ul>\n<li><strong>Interplay Between a Strong Backbone and Pseudo Labels:</strong></li>\n</ul>\n<p>The impact of using pseudo labels<strong>, when trained with pseudo label of when trained with pseudo labels *, effnet b7: 0.685, oneformer(dinat-l): 0.695</strong> *The average of each model trained in holdout is used as pseudo label. </p>\n<p>However, the enhancements were relatively modest with OneFormer.<br>\nOne hypothesis suggests that pseudo labels might bring about a knowledge distillation effect. The impact of this might reach a limit when the backbone is of a certain strength.</p>",
      "rawMarkdown": "# **Summary**\n\n- Models: **OneFormer**, effnet, MaxViT, resnetrs, nfnetf5\n- Train on **soft labels and pseudo label**\n- Optimizing Ensemble Weights\n- Loss: BCE\n- Only random image cropping without TTA and without mask shift\n\n# **Introduction**\n\nOur team would like to thank the organizers and Kaggle for making this competition possible. Also, I want to express my gratitude to my outstanding teammates [@cnumber](https://www.kaggle.com/cnumber) and [@yukiokumura1](https://www.kaggle.com/yukiokumura1) for their incredible contribution toward our final result.\n\n# **Details**\n\nOur main strategy is the adoption and optimization of various models and the use of soft and pseudo labels.\n\n- **Models:**\n\nIn the beginning of the competition, we mainly experimented with a combination of smp and timm, but from the middle of the competition, we also started using OneFormer.\n\nOneFormer demonstrated the best results. \n**private lb score of single oneformer (dinat-l) : 0.70204**\n**With CV score of holdout dice, effnet b7: 0.671, oneformer(dinat-l): 0.693**.\n\nModel resolutions are below: (inference settings) \n\n| OneFormer | effnetb7, b8 | MaxViT-Tiny, Base | resnetrs | nfnetf5 |\n| --- | --- | --- | --- | --- |\n| 1024 | 640 | 512 | 512 | 640 |\n- **Data: Soft and Pseudo Labels:**\n\nUsing soft labels (the average of all annotator labels) during training.  We also introduced pseudo labels for some models. When both are used together, training is performed randomly on images with the 0th through 7th time frame, with soft label used only for the 4th time frame and pseudo label for all other cases.\n\n- **Optimizing Ensemble Weights:**\n\nTo maximized the performance of ensembling, we incorporated optimization using Optuna within the CV environment. This facilitated the efficient verification of the best model combinations.\n\n# Be open to consideration\n\n- **Interplay Between a Strong Backbone and Pseudo Labels:**\n\nThe impact of using pseudo labels**, when trained with pseudo label of when trained with pseudo labels *, effnet b7: 0.685, oneformer(dinat-l): 0.695** *The average of each model trained in holdout is used as pseudo label. \n\nHowever, the enhancements were relatively modest with OneFormer.\nOne hypothesis suggests that pseudo labels might bring about a knowledge distillation effect. The impact of this might reach a limit when the backbone is of a certain strength.",
      "votes": 32
    },
    {
      "id": 2383261,
      "postDate": "2023-08-10T09:19:42.407Z",
      "content": "<p>Congratulations!</p>\n<blockquote>\n  <p>To maximized the performance of ensembling, we incorporated optimization using Optuna within the CV environment. This facilitated the efficient verification of the best model combinations.</p>\n</blockquote>\n<p>Could you describe more about:</p>\n<ol>\n<li>Optuna ensemble optimization? How did you perform this optimization (what was your objective function, what search method did you use)? </li>\n<li>OneFormer training process - which repo did you use for training? How many epochs did you train it? Which optimizer, lr scheduler? It is not often (for me) to see OneFormer fine tuning process during Kaggle competition (certainly I could be wrong).</li>\n</ol>\n<p>Thank you!</p>",
      "rawMarkdown": "Congratulations!\n>To maximized the performance of ensembling, we incorporated optimization using Optuna within the CV environment. This facilitated the efficient verification of the best model combinations.\n\nCould you describe more about:\n1. Optuna ensemble optimization? How did you perform this optimization (what was your objective function, what search method did you use)? \n2. OneFormer training process - which repo did you use for training? How many epochs did you train it? Which optimizer, lr scheduler? It is not often (for me) to see OneFormer fine tuning process during Kaggle competition (certainly I could be wrong).\n\nThank you!\n",
      "votes": 1,
      "replies": [
        {
          "id": 2383275,
          "postDate": "2023-08-10T09:31:52.907Z",
          "content": "<p>Thank you for commenting!</p>\n<blockquote>\n  <ol>\n  <li>Optuna ensemble optimization? How did you perform this optimization (what was your objective function, what search method did you use)?</li>\n  </ol>\n</blockquote>\n<p>We conducted optimization calculations by adjusting the weights and thresholds for each model to maximize the global dice.</p>\n<blockquote>\n  <ol>\n  <li>OneFormer training process - which repo did you use for training? How many epochs did you train it? Which optimizer, lr scheduler? It is not often (for me) to see OneFormer fine tuning process during Kaggle competition (certainly I could be wrong).</li>\n  </ol>\n</blockquote>\n<p>We referred to the official repository and made some custom modifications to it. By the way, my teammate, c-number, pulled an all-nighter to make this happen.</p>",
          "rawMarkdown": "Thank you for commenting!\n\n>1. Optuna ensemble optimization? How did you perform this optimization (what was your objective function, what search method did you use)?\n\nWe conducted optimization calculations by adjusting the weights and thresholds for each model to maximize the global dice.\n\n>2. OneFormer training process - which repo did you use for training? How many epochs did you train it? Which optimizer, lr scheduler? It is not often (for me) to see OneFormer fine tuning process during Kaggle competition (certainly I could be wrong).\n\nWe referred to the official repository and made some custom modifications to it. By the way, my teammate, c-number, pulled an all-nighter to make this happen.",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2383261,
      "author_name": "Remek Kinas",
      "author_url": "",
      "post_date": "2023-08-10T09:19:42.407000",
      "content": "<p>Congratulations!</p>\n<blockquote>\n  <p>To maximized the performance of ensembling, we incorporated optimization using Optuna within the CV environment. This facilitated the efficient verification of the best model combinations.</p>\n</blockquote>\n<p>Could you describe more about:</p>\n<ol>\n<li>Optuna ensemble optimization? How did you perform this optimization (what was your objective function, what search method did you use)? </li>\n<li>OneFormer training process - which repo did you use for training? How many epochs did you train it? Which optimizer, lr scheduler? It is not often (for me) to see OneFormer fine tuning process during Kaggle competition (certainly I could be wrong).</li>\n</ol>\n<p>Thank you!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2383275,
          "author_name": "ts",
          "author_url": "",
          "post_date": "2023-08-10T09:31:52.907000",
          "content": "<p>Thank you for commenting!</p>\n<blockquote>\n  <ol>\n  <li>Optuna ensemble optimization? How did you perform this optimization (what was your objective function, what search method did you use)?</li>\n  </ol>\n</blockquote>\n<p>We conducted optimization calculations by adjusting the weights and thresholds for each model to maximize the global dice.</p>\n<blockquote>\n  <ol>\n  <li>OneFormer training process - which repo did you use for training? How many epochs did you train it? Which optimizer, lr scheduler? It is not often (for me) to see OneFormer fine tuning process during Kaggle competition (certainly I could be wrong).</li>\n  </ol>\n</blockquote>\n<p>We referred to the official repository and made some custom modifications to it. By the way, my teammate, c-number, pulled an all-nighter to make this happen.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2383181": "# **Summary**\n\n- Models: **OneFormer**, effnet, MaxViT, resnetrs, nfnetf5\n- Train on **soft labels and pseudo label**\n- Optimizing Ensemble Weights\n- Loss: BCE\n- Only random image cropping without TTA and without mask shift\n\n# **Introduction**\n\nOur team would like to thank the organizers and Kaggle for making this competition possible. Also, I want to express my gratitude to my outstanding teammates [@cnumber](https://www.kaggle.com/cnumber) and [@yukiokumura1](https://www.kaggle.com/yukiokumura1) for their incredible contribution toward our final result.\n\n# **Details**\n\nOur main strategy is the adoption and optimization of various models and the use of soft and pseudo labels.\n\n- **Models:**\n\nIn the beginning of the competition, we mainly experimented with a combination of smp and timm, but from the middle of the competition, we also started using OneFormer.\n\nOneFormer demonstrated the best results. \n**private lb score of single oneformer (dinat-l) : 0.70204**\n**With CV score of holdout dice, effnet b7: 0.671, oneformer(dinat-l): 0.693**.\n\nModel resolutions are below: (inference settings) \n\n| OneFormer | effnetb7, b8 | MaxViT-Tiny, Base | resnetrs | nfnetf5 |\n| --- | --- | --- | --- | --- |\n| 1024 | 640 | 512 | 512 | 640 |\n- **Data: Soft and Pseudo Labels:**\n\nUsing soft labels (the average of all annotator labels) during training.  We also introduced pseudo labels for some models. When both are used together, training is performed randomly on images with the 0th through 7th time frame, with soft label used only for the 4th time frame and pseudo label for all other cases.\n\n- **Optimizing Ensemble Weights:**\n\nTo maximized the performance of ensembling, we incorporated optimization using Optuna within the CV environment. This facilitated the efficient verification of the best model combinations.\n\n# Be open to consideration\n\n- **Interplay Between a Strong Backbone and Pseudo Labels:**\n\nThe impact of using pseudo labels**, when trained with pseudo label of when trained with pseudo labels *, effnet b7: 0.685, oneformer(dinat-l): 0.695** *The average of each model trained in holdout is used as pseudo label. \n\nHowever, the enhancements were relatively modest with OneFormer.\nOne hypothesis suggests that pseudo labels might bring about a knowledge distillation effect. The impact of this might reach a limit when the backbone is of a certain strength.",
    "2383261": "Congratulations!\n>To maximized the performance of ensembling, we incorporated optimization using Optuna within the CV environment. This facilitated the efficient verification of the best model combinations.\n\nCould you describe more about:\n1. Optuna ensemble optimization? How did you perform this optimization (what was your objective function, what search method did you use)? \n2. OneFormer training process - which repo did you use for training? How many epochs did you train it? Which optimizer, lr scheduler? It is not often (for me) to see OneFormer fine tuning process during Kaggle competition (certainly I could be wrong).\n\nThank you!\n"
  }
}