{
  "id": 430685,
  "title": "3rd Place Solution: 2.5D U-Net",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430685",
  "author_name": "knshnb",
  "post_date": "2023-08-10T18:50:04.361000",
  "votes": 48,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Thanks to the host and Kaggle staff for holding the competition, and congratulations to the winners! I also appreciate my teammates ( <a href=\"https://www.kaggle.com/charmq\" target=\"_blank\">@charmq</a> and <a href=\"https://www.kaggle.com/yoichi7yamakawa\" target=\"_blank\">@yoichi7yamakawa</a>) a lot.</p>\n<h2>Overview</h2>\n<p>Our solution is an ensemble of three pipelines by each member. Here, let me mainly explain my pipeline, which was used as a main solution.<br>\nMy solution is based on a simple 2.5D U-Net (described below). I used 512x512 (256x256 for experimental phases) ash color images of the <a href=\"https://www.kaggle.com/code/inversion/visualizing-contrails\" target=\"_blank\">official notebook</a> as input and predicted a mean value of human individual masks. I trained a model for 25 epochs by AdamW with a cosine annealing scheduler with warmup. The loss function I used was <code>(-dice_coefficient + binary_cross_entropy) / 2</code>.</p>\n<h2>Validation</h2>\n<p>At the early stage of the competition, I was doing 5-fold cross validation. The trends of out-of-fold scores and validation scores differed a little possibly because of the difference in positive-pixel ratios.<br>\nAfter the training get computationally heavy, I decided to check only the validation score of the training of the entire training data. The validation score seemed correlated with a public LB with a small noise.</p>\n<h2>Architecture</h2>\n<p>Our main approach was so-called 2.5D: making 3D input 2D by stacking frames to a batch dimension and input to 2D backbones.<br>\nI first tried 2.5D U-Net architecture with 3D convolutions after the whole U-Net and got a small gain compared to 2D models (around +0.01 in validation).<br>\nThen I tried to move 3D convolutions to the middle of U-Net's skip connection layers, which have richer information of each downsampled feature. We used frames 2, 3, and 4 (0-indexed). In 3D convolution, we reduced the frame dimension from 3 to 1 by stacking two convolutions with kernel_size=2 and padding=0. With this, we got a big gain (around +0.02 in validation).<br>\nThe pseudo-code is as follows (depends heavily on segmentation_models.pytorch library):</p>\n<pre><code> (torch.nn.Sequential):\n     ():\n        ().__init__(\n            torch.nn.Conv3d(in_channels, out_channels, kernel_size, padding=padding, padding_mode=),\n            torch.nn.BatchNorm3d(out_channels),\n            torch.nn.LeakyReLU(),\n        )\n\n (torch.nn.Module):\n     ():\n        ().__init__()\n        self.n_frames = \n        self.backbone = smp.Unet(...)\n        conv3ds = [\n            torch.nn.Sequential(\n                Conv3dBlock(ch, ch, (, , ), (, , )), Conv3dBlock(ch, ch, (, , ), (, , ))\n            )\n             ch  self.backbone.encoder.out_channels[:]\n        ]\n        self.conv3ds = torch.nn.ModuleList(conv3ds)\n\n     () -&gt; torch.Tensor:\n        total_batch, ch, H, W = feature.shape\n        feat_3d = feature.reshape(total_batch // self.n_frames, self.n_frames, ch, H, W).transpose(, )\n         conv3d_block(feat_3d).squeeze()\n\n     () -&gt; torch.Tensor:\n        n_batch, in_ch, n_frame, H, W = x.shape\n        x = x.transpose(, ).reshape(n_batch * n_frame, in_ch, H, W)\n\n        self.backbone.check_input_shape(x)\n\n        features = self.backbone.encoder(x)\n        features[:] = [self._to2d(conv3d, feature)  conv3d, feature  (self.conv3ds, features[:])]\n        decoder_output = self.backbone.decoder(*features)\n\n        masks = self.backbone.segmentation_head(decoder_output)\n         masks\n</code></pre>\n<h2>Data augmentation</h2>\n<p>While full flip and rotation augmentation did not work because of the pixel-shift issues pointed out by <a href=\"https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430618\" target=\"_blank\">1st place solution</a> and <a href=\"https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430479\" target=\"_blank\">9th place solution</a>, applying them by a small ratio enhanced the performance a little. The whole augmentation I used was as follows:</p>\n<pre><code> albumentations  A\n\naugments = [\n    A.HorizontalFlip(p=),\n    A.VerticalFlip(p=),\n    A.RandomRotate90(p=),\n    A.ShiftScaleRotate(, , , p=),\n    A.RandomResizedCrop(, , scale=(, ), ratio=(, ), p=),\n]\n</code></pre>\n<h2>Pseudo label</h2>\n<p>We discretized by 0.25 the models' predictions for 2, 3, 5, 6, and 7 frames and used them as pseudo labels. Since there may be a distribution shift from the original training data, we used pseudo labels for pretraining and the original training data for finetuning.</p>\n<h2>Threshold</h2>\n<p>Since we used validation data for training some models, optimizing the threshold by validation data was not easy. We adopted a percentile threshold. We confirmed by validation data that the optimal percentile was almost equal to the ratio of positive pixels (=0.18%). Therefore, we identified the percentile in test data of the best threshold of some models in validation data by LB probing using submission time. It was about 0.16% and I hope it is a correct ratio.</p>\n<h2>Other tips</h2>\n<ul>\n<li>Setting grad_checkpointing saved more than half of memory usage.</li>\n<li>Increasing the number of decoder channels of U-Net slightly enhanced the performance.</li>\n<li>The mean of the batch-wise dice coefficient does not correspond to the global dice coefficient and is also unstable with small batch sizes. To mitigate this, I heuristically added 700000 to the numerator and 1000000 to the denominator.</li>\n</ul>\n<h2>Final submission</h2>\n<p>Our best single model was 0.706/0.71770/0.71629 (validation/private/public) by maxvit_large. This could still win 3rd place!<br>\nBy ensembling 18 2.5D models with different backbones (maxvit_large_tf_512, tf_efficientnet_l2, resnest269e, maxvit_base_tf_512, maxvit_xlarge_tf_512, tf_efficientnetv2_xl) and slightly different setups (use pseudo label, include validation data for training, finetune lr, etc), we achieved 0.72233 of private LB. Also, adding <a href=\"https://www.kaggle.com/yoichi7yamakawa\" target=\"_blank\">@yoichi7yamakawa</a> and <a href=\"https://www.kaggle.com/charmq\" target=\"_blank\">@charmq</a>'s models increased the private score to 0.72305, which could have won 2nd place by 0.00001!<br>\nUnfortunately, we couldn't select that submission. On the last day, we added a model trained with full flip and rotation augmentation and test-time augmentation, which didn't change the validation score a lot. This lowered the generalization to private data probably because of the pixel-shift issue (or just a random fluctuation).<br>\nAnyway, being unable to find out the pixel-shift issue was the reason for the loss, and I learned a lot.</p>\n<h2>What did not work</h2>\n<ul>\n<li>Double U-Net architecture (training was unstable…)</li>\n<li>Increasing image size by finetune</li>\n<li>Adding conv2d layers between conv3d layers</li>\n</ul>\n<h2>Yiemon773 Part</h2>\n<p>Results of my 2d models' ensemble<br>\nPrivate LB: 0.709+  w/ pseudo labels<br>\nPrivate LB: 0.705+  w/o pseudo labels</p>\n<h3>Preprocess</h3>\n<p>I changed the normalization process from <code>normalize_range</code> to <code>normalize_mean_std</code>. The mean and std of each band are calculated using the training data.</p>\n<h3>backbone</h3>\n<p>resnest269e, maxvit_base, maxvit_large, efficientnetv2_l</p>\n<h3>Loss function</h3>\n<p>Weighted mean of following losses: Hard-label Dice, Hard-label BCE, Soft-label Dice, Soft-label BCE</p>\n<h3>What did not work</h3>\n<ul>\n<li>Many augmentations    <ul>\n<li>Mixup</li>\n<li>frame shuffle</li>\n<li>channel shuffle</li>\n<li>…</li></ul></li>\n<li>Using bands other than 11, 13, 14, and 15 </li>\n</ul>\n<h2>charmq Part</h2>\n<h3>Architecture</h3>\n<p>2.5d models which are mentioned above and 2d models. 2.5d models performed better than 2d models.</p>\n<h3>backbone</h3>\n<p>resnest269e, resnetrs420, maxvit_base, maxvit_large, efficientnetv2_l</p>\n<h3>image size</h3>\n<p>2.5d models and 2d maxvit models were trained with (512, 512). (1024, 1024) was better than (512, 512) with 2d resnest269 but didn't work with other models.</p>\n<h3>Loss function</h3>\n<p>Weighted mean of following losses: Hard-label (pixel mask) Dice, Hard-label (1-(pixel mask))Dice, Soft-label (mean of individual mask) Dice, and Hard-label (min and max of individual mask) Dice</p>\n<h3>What did not work</h3>\n<ul>\n<li>augmentation cropping the area around the positive pixels</li>\n<li>UPerNet with convnext and swin transformer (mmsegmentation)</li>\n</ul>\n<h3>Findings</h3>\n<ul>\n<li>large architectures such as resnest269e, resnetrs420 and maxvit works well</li>\n<li>the training of maxvit with amp is unstable<ul>\n<li>we turned off amp with maxvit</li>\n<li>gradient checkpointing contributes to reduce memory </li></ul></li>\n<li>efficientnet_l2 works well but sometimes output nan<ul>\n<li>we simply replaced the nan output to 0</li></ul></li>\n<li>ensemble boosts score<ul>\n<li>e.g. 0.68 model + 0.68 model -&gt; 0.695</li>\n<li>even ensemble of 5 fold model boosts public score</li></ul></li>\n</ul>\n<h2>Acknowledgment</h2>\n<p>We would like to appreciate Preferred Networks, Inc for allowing us to use computational resources.</p>",
  "messages": [
    {
      "id": 2384036,
      "postDate": "2023-08-10T18:50:04.360Z",
      "content": "<p>Thanks to the host and Kaggle staff for holding the competition, and congratulations to the winners! I also appreciate my teammates ( <a href=\"https://www.kaggle.com/charmq\" target=\"_blank\">@charmq</a> and <a href=\"https://www.kaggle.com/yoichi7yamakawa\" target=\"_blank\">@yoichi7yamakawa</a>) a lot.</p>\n<h2>Overview</h2>\n<p>Our solution is an ensemble of three pipelines by each member. Here, let me mainly explain my pipeline, which was used as a main solution.<br>\nMy solution is based on a simple 2.5D U-Net (described below). I used 512x512 (256x256 for experimental phases) ash color images of the <a href=\"https://www.kaggle.com/code/inversion/visualizing-contrails\" target=\"_blank\">official notebook</a> as input and predicted a mean value of human individual masks. I trained a model for 25 epochs by AdamW with a cosine annealing scheduler with warmup. The loss function I used was <code>(-dice_coefficient + binary_cross_entropy) / 2</code>.</p>\n<h2>Validation</h2>\n<p>At the early stage of the competition, I was doing 5-fold cross validation. The trends of out-of-fold scores and validation scores differed a little possibly because of the difference in positive-pixel ratios.<br>\nAfter the training get computationally heavy, I decided to check only the validation score of the training of the entire training data. The validation score seemed correlated with a public LB with a small noise.</p>\n<h2>Architecture</h2>\n<p>Our main approach was so-called 2.5D: making 3D input 2D by stacking frames to a batch dimension and input to 2D backbones.<br>\nI first tried 2.5D U-Net architecture with 3D convolutions after the whole U-Net and got a small gain compared to 2D models (around +0.01 in validation).<br>\nThen I tried to move 3D convolutions to the middle of U-Net's skip connection layers, which have richer information of each downsampled feature. We used frames 2, 3, and 4 (0-indexed). In 3D convolution, we reduced the frame dimension from 3 to 1 by stacking two convolutions with kernel_size=2 and padding=0. With this, we got a big gain (around +0.02 in validation).<br>\nThe pseudo-code is as follows (depends heavily on segmentation_models.pytorch library):</p>\n<pre><code> (torch.nn.Sequential):\n     ():\n        ().__init__(\n            torch.nn.Conv3d(in_channels, out_channels, kernel_size, padding=padding, padding_mode=),\n            torch.nn.BatchNorm3d(out_channels),\n            torch.nn.LeakyReLU(),\n        )\n\n (torch.nn.Module):\n     ():\n        ().__init__()\n        self.n_frames = \n        self.backbone = smp.Unet(...)\n        conv3ds = [\n            torch.nn.Sequential(\n                Conv3dBlock(ch, ch, (, , ), (, , )), Conv3dBlock(ch, ch, (, , ), (, , ))\n            )\n             ch  self.backbone.encoder.out_channels[:]\n        ]\n        self.conv3ds = torch.nn.ModuleList(conv3ds)\n\n     () -&gt; torch.Tensor:\n        total_batch, ch, H, W = feature.shape\n        feat_3d = feature.reshape(total_batch // self.n_frames, self.n_frames, ch, H, W).transpose(, )\n         conv3d_block(feat_3d).squeeze()\n\n     () -&gt; torch.Tensor:\n        n_batch, in_ch, n_frame, H, W = x.shape\n        x = x.transpose(, ).reshape(n_batch * n_frame, in_ch, H, W)\n\n        self.backbone.check_input_shape(x)\n\n        features = self.backbone.encoder(x)\n        features[:] = [self._to2d(conv3d, feature)  conv3d, feature  (self.conv3ds, features[:])]\n        decoder_output = self.backbone.decoder(*features)\n\n        masks = self.backbone.segmentation_head(decoder_output)\n         masks\n</code></pre>\n<h2>Data augmentation</h2>\n<p>While full flip and rotation augmentation did not work because of the pixel-shift issues pointed out by <a href=\"https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430618\" target=\"_blank\">1st place solution</a> and <a href=\"https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430479\" target=\"_blank\">9th place solution</a>, applying them by a small ratio enhanced the performance a little. The whole augmentation I used was as follows:</p>\n<pre><code> albumentations  A\n\naugments = [\n    A.HorizontalFlip(p=),\n    A.VerticalFlip(p=),\n    A.RandomRotate90(p=),\n    A.ShiftScaleRotate(, , , p=),\n    A.RandomResizedCrop(, , scale=(, ), ratio=(, ), p=),\n]\n</code></pre>\n<h2>Pseudo label</h2>\n<p>We discretized by 0.25 the models' predictions for 2, 3, 5, 6, and 7 frames and used them as pseudo labels. Since there may be a distribution shift from the original training data, we used pseudo labels for pretraining and the original training data for finetuning.</p>\n<h2>Threshold</h2>\n<p>Since we used validation data for training some models, optimizing the threshold by validation data was not easy. We adopted a percentile threshold. We confirmed by validation data that the optimal percentile was almost equal to the ratio of positive pixels (=0.18%). Therefore, we identified the percentile in test data of the best threshold of some models in validation data by LB probing using submission time. It was about 0.16% and I hope it is a correct ratio.</p>\n<h2>Other tips</h2>\n<ul>\n<li>Setting grad_checkpointing saved more than half of memory usage.</li>\n<li>Increasing the number of decoder channels of U-Net slightly enhanced the performance.</li>\n<li>The mean of the batch-wise dice coefficient does not correspond to the global dice coefficient and is also unstable with small batch sizes. To mitigate this, I heuristically added 700000 to the numerator and 1000000 to the denominator.</li>\n</ul>\n<h2>Final submission</h2>\n<p>Our best single model was 0.706/0.71770/0.71629 (validation/private/public) by maxvit_large. This could still win 3rd place!<br>\nBy ensembling 18 2.5D models with different backbones (maxvit_large_tf_512, tf_efficientnet_l2, resnest269e, maxvit_base_tf_512, maxvit_xlarge_tf_512, tf_efficientnetv2_xl) and slightly different setups (use pseudo label, include validation data for training, finetune lr, etc), we achieved 0.72233 of private LB. Also, adding <a href=\"https://www.kaggle.com/yoichi7yamakawa\" target=\"_blank\">@yoichi7yamakawa</a> and <a href=\"https://www.kaggle.com/charmq\" target=\"_blank\">@charmq</a>'s models increased the private score to 0.72305, which could have won 2nd place by 0.00001!<br>\nUnfortunately, we couldn't select that submission. On the last day, we added a model trained with full flip and rotation augmentation and test-time augmentation, which didn't change the validation score a lot. This lowered the generalization to private data probably because of the pixel-shift issue (or just a random fluctuation).<br>\nAnyway, being unable to find out the pixel-shift issue was the reason for the loss, and I learned a lot.</p>\n<h2>What did not work</h2>\n<ul>\n<li>Double U-Net architecture (training was unstable…)</li>\n<li>Increasing image size by finetune</li>\n<li>Adding conv2d layers between conv3d layers</li>\n</ul>\n<h2>Yiemon773 Part</h2>\n<p>Results of my 2d models' ensemble<br>\nPrivate LB: 0.709+  w/ pseudo labels<br>\nPrivate LB: 0.705+  w/o pseudo labels</p>\n<h3>Preprocess</h3>\n<p>I changed the normalization process from <code>normalize_range</code> to <code>normalize_mean_std</code>. The mean and std of each band are calculated using the training data.</p>\n<h3>backbone</h3>\n<p>resnest269e, maxvit_base, maxvit_large, efficientnetv2_l</p>\n<h3>Loss function</h3>\n<p>Weighted mean of following losses: Hard-label Dice, Hard-label BCE, Soft-label Dice, Soft-label BCE</p>\n<h3>What did not work</h3>\n<ul>\n<li>Many augmentations    <ul>\n<li>Mixup</li>\n<li>frame shuffle</li>\n<li>channel shuffle</li>\n<li>…</li></ul></li>\n<li>Using bands other than 11, 13, 14, and 15 </li>\n</ul>\n<h2>charmq Part</h2>\n<h3>Architecture</h3>\n<p>2.5d models which are mentioned above and 2d models. 2.5d models performed better than 2d models.</p>\n<h3>backbone</h3>\n<p>resnest269e, resnetrs420, maxvit_base, maxvit_large, efficientnetv2_l</p>\n<h3>image size</h3>\n<p>2.5d models and 2d maxvit models were trained with (512, 512). (1024, 1024) was better than (512, 512) with 2d resnest269 but didn't work with other models.</p>\n<h3>Loss function</h3>\n<p>Weighted mean of following losses: Hard-label (pixel mask) Dice, Hard-label (1-(pixel mask))Dice, Soft-label (mean of individual mask) Dice, and Hard-label (min and max of individual mask) Dice</p>\n<h3>What did not work</h3>\n<ul>\n<li>augmentation cropping the area around the positive pixels</li>\n<li>UPerNet with convnext and swin transformer (mmsegmentation)</li>\n</ul>\n<h3>Findings</h3>\n<ul>\n<li>large architectures such as resnest269e, resnetrs420 and maxvit works well</li>\n<li>the training of maxvit with amp is unstable<ul>\n<li>we turned off amp with maxvit</li>\n<li>gradient checkpointing contributes to reduce memory </li></ul></li>\n<li>efficientnet_l2 works well but sometimes output nan<ul>\n<li>we simply replaced the nan output to 0</li></ul></li>\n<li>ensemble boosts score<ul>\n<li>e.g. 0.68 model + 0.68 model -&gt; 0.695</li>\n<li>even ensemble of 5 fold model boosts public score</li></ul></li>\n</ul>\n<h2>Acknowledgment</h2>\n<p>We would like to appreciate Preferred Networks, Inc for allowing us to use computational resources.</p>",
      "rawMarkdown": "Thanks to the host and Kaggle staff for holding the competition, and congratulations to the winners! I also appreciate my teammates ( @charmq and @yoichi7yamakawa) a lot.\n\n## Overview\nOur solution is an ensemble of three pipelines by each member. Here, let me mainly explain my pipeline, which was used as a main solution.\nMy solution is based on a simple 2.5D U-Net (described below). I used 512x512 (256x256 for experimental phases) ash color images of the [official notebook](https://www.kaggle.com/code/inversion/visualizing-contrails) as input and predicted a mean value of human individual masks. I trained a model for 25 epochs by AdamW with a cosine annealing scheduler with warmup. The loss function I used was `(-dice_coefficient + binary_cross_entropy) / 2`.\n\n## Validation\nAt the early stage of the competition, I was doing 5-fold cross validation. The trends of out-of-fold scores and validation scores differed a little possibly because of the difference in positive-pixel ratios.\nAfter the training get computationally heavy, I decided to check only the validation score of the training of the entire training data. The validation score seemed correlated with a public LB with a small noise.\n\n## Architecture\nOur main approach was so-called 2.5D: making 3D input 2D by stacking frames to a batch dimension and input to 2D backbones.\nI first tried 2.5D U-Net architecture with 3D convolutions after the whole U-Net and got a small gain compared to 2D models (around +0.01 in validation).\nThen I tried to move 3D convolutions to the middle of U-Net's skip connection layers, which have richer information of each downsampled feature. We used frames 2, 3, and 4 (0-indexed). In 3D convolution, we reduced the frame dimension from 3 to 1 by stacking two convolutions with kernel_size=2 and padding=0. With this, we got a big gain (around +0.02 in validation).\nThe pseudo-code is as follows (depends heavily on segmentation_models.pytorch library):\n```python\nclass Conv3dBlock(torch.nn.Sequential):\n    def __init__(\n        self, in_channels: int, out_channels: int, kernel_size: tuple[int, int, int], padding: tuple[int, int, int]\n    ):\n        super().__init__(\n            torch.nn.Conv3d(in_channels, out_channels, kernel_size, padding=padding, padding_mode=\"replicate\"),\n            torch.nn.BatchNorm3d(out_channels),\n            torch.nn.LeakyReLU(),\n        )\n\nclass Segmentor25d(torch.nn.Module):\n    def __init__(self, ...):\n        super().__init__()\n        self.n_frames = 3\n        self.backbone = smp.Unet(...)\n        conv3ds = [\n            torch.nn.Sequential(\n                Conv3dBlock(ch, ch, (2, 3, 3), (0, 1, 1)), Conv3dBlock(ch, ch, (2, 3, 3), (0, 1, 1))\n            )\n            for ch in self.backbone.encoder.out_channels[1:]\n        ]\n        self.conv3ds = torch.nn.ModuleList(conv3ds)\n\n    def _to2d(self, conv3d_block: torch.nn.Module, feature: torch.Tensor) -> torch.Tensor:\n        total_batch, ch, H, W = feature.shape\n        feat_3d = feature.reshape(total_batch // self.n_frames, self.n_frames, ch, H, W).transpose(1, 2)\n        return conv3d_block(feat_3d).squeeze(2)\n\n    def forward(self, x: torch.Tensor) -> torch.Tensor:\n        n_batch, in_ch, n_frame, H, W = x.shape\n        x = x.transpose(1, 2).reshape(n_batch * n_frame, in_ch, H, W)\n\n        self.backbone.check_input_shape(x)\n\n        features = self.backbone.encoder(x)\n        features[1:] = [self._to2d(conv3d, feature) for conv3d, feature in zip(self.conv3ds, features[1:])]\n        decoder_output = self.backbone.decoder(*features)\n\n        masks = self.backbone.segmentation_head(decoder_output)\n        return masks\n\n```\n\n## Data augmentation\nWhile full flip and rotation augmentation did not work because of the pixel-shift issues pointed out by [1st place solution](https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430618) and [9th place solution](https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430479), applying them by a small ratio enhanced the performance a little. The whole augmentation I used was as follows:\n```python\nimport albumentations as A\n\naugments = [\n    A.HorizontalFlip(p=0.1),\n    A.VerticalFlip(p=0.1),\n    A.RandomRotate90(p=0.4),\n    A.ShiftScaleRotate(0.05, 0.1, 15, p=0.3),\n    A.RandomResizedCrop(512, 512, scale=(0.75, 1.0), ratio=(0.9, 1.1111111111111), p=0.7),\n]\n```\n\n## Pseudo label\nWe discretized by 0.25 the models' predictions for 2, 3, 5, 6, and 7 frames and used them as pseudo labels. Since there may be a distribution shift from the original training data, we used pseudo labels for pretraining and the original training data for finetuning.\n\n## Threshold\nSince we used validation data for training some models, optimizing the threshold by validation data was not easy. We adopted a percentile threshold. We confirmed by validation data that the optimal percentile was almost equal to the ratio of positive pixels (=0.18%). Therefore, we identified the percentile in test data of the best threshold of some models in validation data by LB probing using submission time. It was about 0.16% and I hope it is a correct ratio.\n\n## Other tips\n- Setting grad_checkpointing saved more than half of memory usage.\n- Increasing the number of decoder channels of U-Net slightly enhanced the performance.\n- The mean of the batch-wise dice coefficient does not correspond to the global dice coefficient and is also unstable with small batch sizes. To mitigate this, I heuristically added 700000 to the numerator and 1000000 to the denominator.\n\n## Final submission\nOur best single model was 0.706/0.71770/0.71629 (validation/private/public) by maxvit_large. This could still win 3rd place!\nBy ensembling 18 2.5D models with different backbones (maxvit_large_tf_512, tf_efficientnet_l2, resnest269e, maxvit_base_tf_512, maxvit_xlarge_tf_512, tf_efficientnetv2_xl) and slightly different setups (use pseudo label, include validation data for training, finetune lr, etc), we achieved 0.72233 of private LB. Also, adding @yoichi7yamakawa and @charmq's models increased the private score to 0.72305, which could have won 2nd place by 0.00001!\nUnfortunately, we couldn't select that submission. On the last day, we added a model trained with full flip and rotation augmentation and test-time augmentation, which didn't change the validation score a lot. This lowered the generalization to private data probably because of the pixel-shift issue (or just a random fluctuation).\nAnyway, being unable to find out the pixel-shift issue was the reason for the loss, and I learned a lot.\n\n## What did not work\n- Double U-Net architecture (training was unstable...)\n- Increasing image size by finetune\n- Adding conv2d layers between conv3d layers\n\n## Yiemon773 Part\nResults of my 2d models' ensemble\nPrivate LB: 0.709+  w/ pseudo labels\nPrivate LB: 0.705+  w/o pseudo labels\n\n### Preprocess \nI changed the normalization process from `normalize_range` to `normalize_mean_std`. The mean and std of each band are calculated using the training data.\n\n### backbone\nresnest269e, maxvit_base, maxvit_large, efficientnetv2_l\n\n### Loss function\nWeighted mean of following losses: Hard-label Dice, Hard-label BCE, Soft-label Dice, Soft-label BCE\n\n### What did not work\n- Many augmentations\t\n  - Mixup\n  - frame shuffle\n  - channel shuffle\n  - ...\n- Using bands other than 11, 13, 14, and 15 \n\n## charmq Part\n### Architecture\n2.5d models which are mentioned above and 2d models. 2.5d models performed better than 2d models.\n\n### backbone\nresnest269e, resnetrs420, maxvit_base, maxvit_large, efficientnetv2_l\n\n### image size\n2.5d models and 2d maxvit models were trained with (512, 512). (1024, 1024) was better than (512, 512) with 2d resnest269 but didn't work with other models.\n\n### Loss function\nWeighted mean of following losses: Hard-label (pixel mask) Dice, Hard-label (1-(pixel mask))Dice, Soft-label (mean of individual mask) Dice, and Hard-label (min and max of individual mask) Dice\n\n### What did not work\n- augmentation cropping the area around the positive pixels\n- UPerNet with convnext and swin transformer (mmsegmentation)\n\n### Findings\n- large architectures such as resnest269e, resnetrs420 and maxvit works well\n- the training of maxvit with amp is unstable\n    - we turned off amp with maxvit\n    - gradient checkpointing contributes to reduce memory \n- efficientnet_l2 works well but sometimes output nan\n    - we simply replaced the nan output to 0\n- ensemble boosts score\n    - e.g. 0.68 model + 0.68 model -> 0.695\n    - even ensemble of 5 fold model boosts public score\n\n## Acknowledgment\nWe would like to appreciate Preferred Networks, Inc for allowing us to use computational resources.",
      "votes": 48
    },
    {
      "id": 2384131,
      "postDate": "2023-08-10T20:29:37.863Z",
      "content": "<p>Amazing use of the temporal information. Thank you for sharing and congratulations on 3rd place!</p>\n<p>How do you set gradient checkpointing?</p>",
      "rawMarkdown": "Amazing use of the temporal information. Thank you for sharing and congratulations on 3rd place!\n\nHow do you set gradient checkpointing?",
      "votes": 1,
      "replies": [
        {
          "id": 2386864,
          "postDate": "2023-08-12T07:39:30.473Z",
          "content": "<p>Thanks! If you are using segmentation_models.pytorch, you can set gradient checkpointing for most timm models by the following code:</p>\n<pre><code>backbone = smp(...)\nbackbone()\n</code></pre>",
          "rawMarkdown": "Thanks! If you are using segmentation_models.pytorch, you can set gradient checkpointing for most timm models by the following code:\n```\nbackbone = smp.Unet(...)\nbackbone.encoder.model.set_grad_checkpointing()\n```",
          "votes": 4
        }
      ]
    },
    {
      "id": 2384129,
      "postDate": "2023-08-10T20:23:17.353Z",
      "content": "<p>Thank you for the write up and congrats! <a href=\"https://www.kaggle.com/knshnb\" target=\"_blank\">@knshnb</a> <a href=\"https://www.kaggle.com/charmq\" target=\"_blank\">@charmq</a> <a href=\"https://www.kaggle.com/yoichi7yamakawa\" target=\"_blank\">@yoichi7yamakawa</a> </p>",
      "rawMarkdown": "Thank you for the write up and congrats! @knshnb @charmq @yoichi7yamakawa ",
      "votes": 1,
      "replies": [
        {
          "id": 2386865,
          "postDate": "2023-08-12T07:39:44Z",
          "content": "<p>Thank you!!!</p>",
          "rawMarkdown": "Thank you!!!",
          "votes": 1
        }
      ]
    },
    {
      "id": 2407857,
      "postDate": "2023-08-25T09:33:44.223Z",
      "content": "<p>We published our source code!</p>\n<p><a href=\"https://github.com/knshnb/kaggle-contrails-3rd-place\" target=\"_blank\">https://github.com/knshnb/kaggle-contrails-3rd-place</a><br>\n<a href=\"https://github.com/yoichi-yamakawa/kaggle-contrail-3rd-place-solution\" target=\"_blank\">https://github.com/yoichi-yamakawa/kaggle-contrail-3rd-place-solution</a><br>\n<a href=\"https://github.com/tyamaguchi17/contrails_charm_public\" target=\"_blank\">https://github.com/tyamaguchi17/contrails_charm_public</a></p>",
      "rawMarkdown": "We published our source code!\n\nhttps://github.com/knshnb/kaggle-contrails-3rd-place\nhttps://github.com/yoichi-yamakawa/kaggle-contrail-3rd-place-solution\nhttps://github.com/tyamaguchi17/contrails_charm_public"
    },
    {
      "id": 2390176,
      "postDate": "2023-08-14T12:28:12.287Z",
      "content": "<p>Congratulations and thanks a lot for sharing! 😀<br>\nWhat's the advantage of choosing the optimal threshold using the percentile vs optimizing the threshold value directly?</p>",
      "rawMarkdown": "Congratulations and thanks a lot for sharing! 😀\nWhat's the advantage of choosing the optimal threshold using the percentile vs optimizing the threshold value directly?",
      "replies": [
        {
          "id": 2392288,
          "postDate": "2023-08-15T15:18:19.137Z",
          "content": "<p>Thanks!<br>\nChoosing the optimal threshold value directly for validation data was fine, but we couldn't do that because we trained several models also on validation data. That's why we identified the best percentile using models that didn't use validation data for training.</p>",
          "rawMarkdown": "Thanks!\nChoosing the optimal threshold value directly for validation data was fine, but we couldn't do that because we trained several models also on validation data. That's why we identified the best percentile using models that didn't use validation data for training.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2385004,
      "postDate": "2023-08-11T06:24:19.807Z",
      "content": "<p>May I ask what tool you use for hyperparameter search? Is it NNI</p>",
      "rawMarkdown": "May I ask what tool you use for hyperparameter search? Is it NNI",
      "replies": [
        {
          "id": 2386868,
          "postDate": "2023-08-12T07:41:35.080Z",
          "content": "<p>I didn't use any hyperparameter optimization tool for this competition. I have used Optuna for the previous one: <a href=\"https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/320192\" target=\"_blank\">https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/320192</a>.</p>",
          "rawMarkdown": "I didn't use any hyperparameter optimization tool for this competition. I have used Optuna for the previous one: https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/320192.",
          "replies": [
            {
              "id": 2387982,
              "postDate": "2023-08-13T05:15:21.170Z",
              "content": "<p>So how did you determine the backbone and various parameters to use in this competition? It's really impressive</p>",
              "rawMarkdown": "So how did you determine the backbone and various parameters to use in this competition? It's really impressive"
            },
            {
              "id": 2388397,
              "postDate": "2023-08-13T11:10:28.827Z",
              "content": "<p>We just manually tried several parameters. There might still be some room for tuning, but we thought it was more worthwhile to spend our time and computational resources on other experiments.</p>",
              "rawMarkdown": "We just manually tried several parameters. There might still be some room for tuning, but we thought it was more worthwhile to spend our time and computational resources on other experiments."
            },
            {
              "id": 2413545,
              "postDate": "2023-08-29T02:05:35.130Z",
              "content": "<p>Thank you. I have been researching automl recently and have found that people rarely use it in the DL field, and the development of related tools is not as complete. I wonder if it is the reason why DL is too resource intensive? ML doesn't consume as much computing power, so it seems that Optuna is more used for ML. I don't know if my understanding is correct? I would greatly appreciate it if you could correct me</p>",
              "rawMarkdown": "Thank you. I have been researching automl recently and have found that people rarely use it in the DL field, and the development of related tools is not as complete. I wonder if it is the reason why DL is too resource intensive? ML doesn't consume as much computing power, so it seems that Optuna is more used for ML. I don't know if my understanding is correct? I would greatly appreciate it if you could correct me"
            },
            {
              "id": 2418504,
              "postDate": "2023-09-01T10:19:36.413Z",
              "content": "<p>Of course, it depends on the situation, but I think hyperparameter optimization of DL is too computationally heavy in many cases.</p>",
              "rawMarkdown": "Of course, it depends on the situation, but I think hyperparameter optimization of DL is too computationally heavy in many cases."
            }
          ]
        }
      ]
    },
    {
      "id": 2413538,
      "postDate": "2023-08-29T01:48:06.093Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2384131,
      "author_name": "Yousef Rabi",
      "author_url": "",
      "post_date": "2023-08-10T20:29:37.863000",
      "content": "<p>Amazing use of the temporal information. Thank you for sharing and congratulations on 3rd place!</p>\n<p>How do you set gradient checkpointing?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2386864,
          "author_name": "knshnb",
          "author_url": "",
          "post_date": "2023-08-12T07:39:30.473000",
          "content": "<p>Thanks! If you are using segmentation_models.pytorch, you can set gradient checkpointing for most timm models by the following code:</p>\n<pre><code>backbone = smp(...)\nbackbone()\n</code></pre>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 2384129,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "2023-08-10T20:23:17.353000",
      "content": "<p>Thank you for the write up and congrats! <a href=\"https://www.kaggle.com/knshnb\" target=\"_blank\">@knshnb</a> <a href=\"https://www.kaggle.com/charmq\" target=\"_blank\">@charmq</a> <a href=\"https://www.kaggle.com/yoichi7yamakawa\" target=\"_blank\">@yoichi7yamakawa</a> </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2386865,
          "author_name": "knshnb",
          "author_url": "",
          "post_date": "2023-08-12T07:39:44",
          "content": "<p>Thank you!!!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2407857,
      "author_name": "knshnb",
      "author_url": "",
      "post_date": "2023-08-25T09:33:44.223000",
      "content": "<p>We published our source code!</p>\n<p><a href=\"https://github.com/knshnb/kaggle-contrails-3rd-place\" target=\"_blank\">https://github.com/knshnb/kaggle-contrails-3rd-place</a><br>\n<a href=\"https://github.com/yoichi-yamakawa/kaggle-contrail-3rd-place-solution\" target=\"_blank\">https://github.com/yoichi-yamakawa/kaggle-contrail-3rd-place-solution</a><br>\n<a href=\"https://github.com/tyamaguchi17/contrails_charm_public\" target=\"_blank\">https://github.com/tyamaguchi17/contrails_charm_public</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2390176,
      "author_name": "delai50",
      "author_url": "",
      "post_date": "2023-08-14T12:28:12.287000",
      "content": "<p>Congratulations and thanks a lot for sharing! 😀<br>\nWhat's the advantage of choosing the optimal threshold using the percentile vs optimizing the threshold value directly?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2392288,
          "author_name": "knshnb",
          "author_url": "",
          "post_date": "2023-08-15T15:18:19.137000",
          "content": "<p>Thanks!<br>\nChoosing the optimal threshold value directly for validation data was fine, but we couldn't do that because we trained several models also on validation data. That's why we identified the best percentile using models that didn't use validation data for training.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2385004,
      "author_name": "kongweihao",
      "author_url": "",
      "post_date": "2023-08-11T06:24:19.807000",
      "content": "<p>May I ask what tool you use for hyperparameter search? Is it NNI</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2386868,
          "author_name": "knshnb",
          "author_url": "",
          "post_date": "2023-08-12T07:41:35.080000",
          "content": "<p>I didn't use any hyperparameter optimization tool for this competition. I have used Optuna for the previous one: <a href=\"https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/320192\" target=\"_blank\">https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/320192</a>.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2387982,
              "author_name": "kongweihao",
              "author_url": "",
              "post_date": "2023-08-13T05:15:21.170000",
              "content": "<p>So how did you determine the backbone and various parameters to use in this competition? It's really impressive</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2388397,
              "author_name": "knshnb",
              "author_url": "",
              "post_date": "2023-08-13T11:10:28.827000",
              "content": "<p>We just manually tried several parameters. There might still be some room for tuning, but we thought it was more worthwhile to spend our time and computational resources on other experiments.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2413545,
              "author_name": "kongweihao",
              "author_url": "",
              "post_date": "2023-08-29T02:05:35.130000",
              "content": "<p>Thank you. I have been researching automl recently and have found that people rarely use it in the DL field, and the development of related tools is not as complete. I wonder if it is the reason why DL is too resource intensive? ML doesn't consume as much computing power, so it seems that Optuna is more used for ML. I don't know if my understanding is correct? I would greatly appreciate it if you could correct me</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2418504,
              "author_name": "knshnb",
              "author_url": "",
              "post_date": "2023-09-01T10:19:36.413000",
              "content": "<p>Of course, it depends on the situation, but I think hyperparameter optimization of DL is too computationally heavy in many cases.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2413538,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-08-29T01:48:06.093000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2384036": "Thanks to the host and Kaggle staff for holding the competition, and congratulations to the winners! I also appreciate my teammates ( @charmq and @yoichi7yamakawa) a lot.\n\n## Overview\nOur solution is an ensemble of three pipelines by each member. Here, let me mainly explain my pipeline, which was used as a main solution.\nMy solution is based on a simple 2.5D U-Net (described below). I used 512x512 (256x256 for experimental phases) ash color images of the [official notebook](https://www.kaggle.com/code/inversion/visualizing-contrails) as input and predicted a mean value of human individual masks. I trained a model for 25 epochs by AdamW with a cosine annealing scheduler with warmup. The loss function I used was `(-dice_coefficient + binary_cross_entropy) / 2`.\n\n## Validation\nAt the early stage of the competition, I was doing 5-fold cross validation. The trends of out-of-fold scores and validation scores differed a little possibly because of the difference in positive-pixel ratios.\nAfter the training get computationally heavy, I decided to check only the validation score of the training of the entire training data. The validation score seemed correlated with a public LB with a small noise.\n\n## Architecture\nOur main approach was so-called 2.5D: making 3D input 2D by stacking frames to a batch dimension and input to 2D backbones.\nI first tried 2.5D U-Net architecture with 3D convolutions after the whole U-Net and got a small gain compared to 2D models (around +0.01 in validation).\nThen I tried to move 3D convolutions to the middle of U-Net's skip connection layers, which have richer information of each downsampled feature. We used frames 2, 3, and 4 (0-indexed). In 3D convolution, we reduced the frame dimension from 3 to 1 by stacking two convolutions with kernel_size=2 and padding=0. With this, we got a big gain (around +0.02 in validation).\nThe pseudo-code is as follows (depends heavily on segmentation_models.pytorch library):\n```python\nclass Conv3dBlock(torch.nn.Sequential):\n    def __init__(\n        self, in_channels: int, out_channels: int, kernel_size: tuple[int, int, int], padding: tuple[int, int, int]\n    ):\n        super().__init__(\n            torch.nn.Conv3d(in_channels, out_channels, kernel_size, padding=padding, padding_mode=\"replicate\"),\n            torch.nn.BatchNorm3d(out_channels),\n            torch.nn.LeakyReLU(),\n        )\n\nclass Segmentor25d(torch.nn.Module):\n    def __init__(self, ...):\n        super().__init__()\n        self.n_frames = 3\n        self.backbone = smp.Unet(...)\n        conv3ds = [\n            torch.nn.Sequential(\n                Conv3dBlock(ch, ch, (2, 3, 3), (0, 1, 1)), Conv3dBlock(ch, ch, (2, 3, 3), (0, 1, 1))\n            )\n            for ch in self.backbone.encoder.out_channels[1:]\n        ]\n        self.conv3ds = torch.nn.ModuleList(conv3ds)\n\n    def _to2d(self, conv3d_block: torch.nn.Module, feature: torch.Tensor) -> torch.Tensor:\n        total_batch, ch, H, W = feature.shape\n        feat_3d = feature.reshape(total_batch // self.n_frames, self.n_frames, ch, H, W).transpose(1, 2)\n        return conv3d_block(feat_3d).squeeze(2)\n\n    def forward(self, x: torch.Tensor) -> torch.Tensor:\n        n_batch, in_ch, n_frame, H, W = x.shape\n        x = x.transpose(1, 2).reshape(n_batch * n_frame, in_ch, H, W)\n\n        self.backbone.check_input_shape(x)\n\n        features = self.backbone.encoder(x)\n        features[1:] = [self._to2d(conv3d, feature) for conv3d, feature in zip(self.conv3ds, features[1:])]\n        decoder_output = self.backbone.decoder(*features)\n\n        masks = self.backbone.segmentation_head(decoder_output)\n        return masks\n\n```\n\n## Data augmentation\nWhile full flip and rotation augmentation did not work because of the pixel-shift issues pointed out by [1st place solution](https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430618) and [9th place solution](https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430479), applying them by a small ratio enhanced the performance a little. The whole augmentation I used was as follows:\n```python\nimport albumentations as A\n\naugments = [\n    A.HorizontalFlip(p=0.1),\n    A.VerticalFlip(p=0.1),\n    A.RandomRotate90(p=0.4),\n    A.ShiftScaleRotate(0.05, 0.1, 15, p=0.3),\n    A.RandomResizedCrop(512, 512, scale=(0.75, 1.0), ratio=(0.9, 1.1111111111111), p=0.7),\n]\n```\n\n## Pseudo label\nWe discretized by 0.25 the models' predictions for 2, 3, 5, 6, and 7 frames and used them as pseudo labels. Since there may be a distribution shift from the original training data, we used pseudo labels for pretraining and the original training data for finetuning.\n\n## Threshold\nSince we used validation data for training some models, optimizing the threshold by validation data was not easy. We adopted a percentile threshold. We confirmed by validation data that the optimal percentile was almost equal to the ratio of positive pixels (=0.18%). Therefore, we identified the percentile in test data of the best threshold of some models in validation data by LB probing using submission time. It was about 0.16% and I hope it is a correct ratio.\n\n## Other tips\n- Setting grad_checkpointing saved more than half of memory usage.\n- Increasing the number of decoder channels of U-Net slightly enhanced the performance.\n- The mean of the batch-wise dice coefficient does not correspond to the global dice coefficient and is also unstable with small batch sizes. To mitigate this, I heuristically added 700000 to the numerator and 1000000 to the denominator.\n\n## Final submission\nOur best single model was 0.706/0.71770/0.71629 (validation/private/public) by maxvit_large. This could still win 3rd place!\nBy ensembling 18 2.5D models with different backbones (maxvit_large_tf_512, tf_efficientnet_l2, resnest269e, maxvit_base_tf_512, maxvit_xlarge_tf_512, tf_efficientnetv2_xl) and slightly different setups (use pseudo label, include validation data for training, finetune lr, etc), we achieved 0.72233 of private LB. Also, adding @yoichi7yamakawa and @charmq's models increased the private score to 0.72305, which could have won 2nd place by 0.00001!\nUnfortunately, we couldn't select that submission. On the last day, we added a model trained with full flip and rotation augmentation and test-time augmentation, which didn't change the validation score a lot. This lowered the generalization to private data probably because of the pixel-shift issue (or just a random fluctuation).\nAnyway, being unable to find out the pixel-shift issue was the reason for the loss, and I learned a lot.\n\n## What did not work\n- Double U-Net architecture (training was unstable...)\n- Increasing image size by finetune\n- Adding conv2d layers between conv3d layers\n\n## Yiemon773 Part\nResults of my 2d models' ensemble\nPrivate LB: 0.709+  w/ pseudo labels\nPrivate LB: 0.705+  w/o pseudo labels\n\n### Preprocess \nI changed the normalization process from `normalize_range` to `normalize_mean_std`. The mean and std of each band are calculated using the training data.\n\n### backbone\nresnest269e, maxvit_base, maxvit_large, efficientnetv2_l\n\n### Loss function\nWeighted mean of following losses: Hard-label Dice, Hard-label BCE, Soft-label Dice, Soft-label BCE\n\n### What did not work\n- Many augmentations\t\n  - Mixup\n  - frame shuffle\n  - channel shuffle\n  - ...\n- Using bands other than 11, 13, 14, and 15 \n\n## charmq Part\n### Architecture\n2.5d models which are mentioned above and 2d models. 2.5d models performed better than 2d models.\n\n### backbone\nresnest269e, resnetrs420, maxvit_base, maxvit_large, efficientnetv2_l\n\n### image size\n2.5d models and 2d maxvit models were trained with (512, 512). (1024, 1024) was better than (512, 512) with 2d resnest269 but didn't work with other models.\n\n### Loss function\nWeighted mean of following losses: Hard-label (pixel mask) Dice, Hard-label (1-(pixel mask))Dice, Soft-label (mean of individual mask) Dice, and Hard-label (min and max of individual mask) Dice\n\n### What did not work\n- augmentation cropping the area around the positive pixels\n- UPerNet with convnext and swin transformer (mmsegmentation)\n\n### Findings\n- large architectures such as resnest269e, resnetrs420 and maxvit works well\n- the training of maxvit with amp is unstable\n    - we turned off amp with maxvit\n    - gradient checkpointing contributes to reduce memory \n- efficientnet_l2 works well but sometimes output nan\n    - we simply replaced the nan output to 0\n- ensemble boosts score\n    - e.g. 0.68 model + 0.68 model -> 0.695\n    - even ensemble of 5 fold model boosts public score\n\n## Acknowledgment\nWe would like to appreciate Preferred Networks, Inc for allowing us to use computational resources.",
    "2384131": "Amazing use of the temporal information. Thank you for sharing and congratulations on 3rd place!\n\nHow do you set gradient checkpointing?",
    "2384129": "Thank you for the write up and congrats! @knshnb @charmq @yoichi7yamakawa ",
    "2407857": "We published our source code!\n\nhttps://github.com/knshnb/kaggle-contrails-3rd-place\nhttps://github.com/yoichi-yamakawa/kaggle-contrail-3rd-place-solution\nhttps://github.com/tyamaguchi17/contrails_charm_public",
    "2390176": "Congratulations and thanks a lot for sharing! 😀\nWhat's the advantage of choosing the optimal threshold using the percentile vs optimizing the threshold value directly?",
    "2385004": "May I ask what tool you use for hyperparameter search? Is it NNI",
    "2413538": ""
  }
}