{
  "id": 434086,
  "title": "105th Place Solution for the Google Research - Identify Contrails to Reduce Global Warming  Competition",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/434086",
  "author_name": "Jenny Ding",
  "post_date": "2023-08-23T22:03:04.507000",
  "votes": 2,
  "comment_count": 0,
  "views": 0,
  "content": "<h1>1. Context section</h1>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/overview\" target=\"_blank\">Business context</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/data\" target=\"_blank\">Data context</a></li>\n</ul>\n<h1>2. Overview of the Approach</h1>\n<p>In this competition, the motivation is to contribute to the improvement of contrail prediction models, which are used to predict the formation of contrails generated by aircraft engine exhaust. Contrails are clouds of ice crystals with a line shape. They form when aircraft fly through super humid regions in the atmosphere and contribute greatly to global warming by trapping the heat in the atmosphere. Our goal is to validate the models that use satellite imagery as input data and then provide airlines with more accurate ways to mitigate the problem of climate change by avoiding the formation of contrails.</p>\n<p>In this study, we addressed the contrail detection challenge using geostationary satellite imagery from the GOES-16 ABI. Our approach involved data preprocessing for standardization and ash RGB image creation, contrail annotation with IoU-based object detection, and model training with EfficientNet and Dice coefficient loss in a K-Fold cross-validation setup. A distinctive feature was threshold selection for image masking, with a 0.52 threshold derived from the fifth image's Dice coefficient distribution. Validation included an error metric assessing labeling accuracy. Importantly, the Dice coefficient distributions from training on the fifth image alone and additional-image training closely aligned, confirming the chosen threshold's effectiveness.</p>\n<h1>3. Details of the submission</h1>\n<h2>3.1 Data Exploration</h2>\n<p>The data used in this competition are geostationary satellite images, sourced from the GOES-16 Advanced Baseline Imager (ABI). The images are provided as a sequence of 10-minute interval snapshots, where each sequence consists of eight images, with labeling applied only to the fifth image. Each dataset entry is represented by a unique record_id and contains precisely one labeled frame. The training dataset offers both individual label annotations and aggregated sound truth annotations, while the validation data only contains the latter.</p>\n<h2>3.2 Data Preprocessing</h2>\n<h3>3.2.1 Image Standardization</h3>\n<p>Data preprocessing plays a crucial role in this competition. The provided data consists of eight different spectral bands, and the pixel values are not in the standard 0-255 range. To make the satellite data suitable for our models, we first normalize and transform them using the following two steps:</p>\n<ul>\n<li>Map the pixel values of each spectral band to the standard range of 0 to 255 to create visualizable images;</li>\n<li>Create an ash RGB image by combining these eight bands, forming a three-channel color image.</li>\n</ul>\n<h3>3.2.2 Contrail Annotation</h3>\n<p>Next, we use the Intersection over Union (IoU) method to locate and annotate the contrails in the images. Our steps are as follows:</p>\n<ul>\n<li>Utilize object detection models to detect contrails within the ash RGB images. Outline them using bounding boxes;</li>\n<li>For potentially overlapping boxes, calculate the IoU to determine the degree of overlap between the boxes;</li>\n<li>Based on the calculation, merge smaller boxes with significant overlap into larger bounding boxes to enhance annotation accuracy and readability.</li>\n</ul>\n<h2>3.3 Model Training</h2>\n<h3>3.3.1 Random Crop and Image Selection</h3>\n<p>Random cropping is implemented as a form of data augmentation. We select only the fifth image from each sequence for training, as it contains the binary mask information (0 to 1). </p>\n<p>We applied a total of 25 different methods of cropping and the four labeled green are the most effective ones: <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2497266%2F782ddf910435c2bca436c491cfae2869%2Fcropping%20methods.png?generation=1692827919871994&amp;alt=media\" alt=\"cropping methods\"></p>\n<p>The following are the training results of each method: <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2497266%2Ff6aba1d97ceee02ef0296aebd4ee3e82%2Fcropping%20results.png?generation=1692827970836162&amp;alt=media\" alt=\"cropping results\"></p>\n<p>Some of the blank data is due to the fact that during training, we initially didn’t head in the right direction, so the experiment didn’t continue on those methods. In the end, we opted for the seventh method: Randomized square cropping with at least one complete contrail and the size of the cropped image is always greater than 128 pixels. The cropped images could be proportionally scaled as needed. For instance, we could expand them to 512 pixels by random cropping images greater than 256 pixels and then resizing them. Likewise, if we wanted to increase the image size to 1024 pixels, we followed a similar process by random cropping images larger than or equal to 512 pixels and resizing them to 1024 pixels. </p>\n<h3>3.3.2 EfficientNet with K-Fold Validation and Dice Coefficient</h3>\n<p>The model is trained using the EfficientNet architecture. The K-Fold cross-validation is employed to assess the performance of the model across different subsets of the training dataset and to ensure the robustness and generalization of the model. For loss calculation, the Dice coefficient function is used, which is useful for segmentation tasks.</p>\n<h3>3.3.3 Threshold Selection for Masking</h3>\n<p>After training the model with all the fifth images from each sequence, the next step is to apply the trained model to the rest of the images in each sequence (the first to fourth and sixth to eighth) to create the masks. The most creative part of our solution is to determine an appropriate threshold (or confidence) within the range from 0 to 1 for masking. This threshold is based on the distribution of the fifth image of each sequence. The following are the steps in detail:</p>\n<ul>\n<li>Apply the trained EfficientNet models to the first to fourth and sixth to eighth images in each sequence;</li>\n<li>Generate binary probability (ranging from 0 to 1) based on the model’s predictions;</li>\n<li>Determine an optimal threshold within the 0-1 range by analyzing the distribution of the fifth image;</li>\n<li>Apply the selected threshold to the rest of the images (the first to fourth and sixth to eighth).</li>\n</ul>\n<p>This approach leverages the power of our EfficientNet models and ensures that the models can better identify and retain the relevant images. </p>\n<p>The training results are illustrated in the following two figures. The left figure displays the distribution of Dice coefficients obtained when training exclusively with the fifth image of each sequence. Conversely, the right figure showcases the distribution of Dice coefficients when training with additional images, specifically the first to fourth and sixth to eighth images within each sequence. Remarkably, these two distributions exhibit a remarkable similarity, enabling us to determine a suitable threshold based on the distribution derived from the fifth image.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2497266%2F38af041ab03c025ae9fb897aaee999c3%2FDice%20Coefficient%20Distribution%20.png?generation=1692828079636644&amp;alt=media\" alt=\"dice coefficient distribution\"><br>\nTo assess the accuracy of our approach, we introduced an error metric. If the fifth image was not correctly labeled, it was counted as an error. Subsequently, we calculated the proportion of results that contained errors. Our statistical analysis revealed that a Dice coefficient exceeding 0.5275080198049544 (approximately equal to 0.52 in our case) reliably indicated utility. Therefore, the threshold was set at 0.52, ensuring that the chosen threshold effectively distinguishes valuable information in the images. </p>\n<h1>4. Sources</h1>\n<ul>\n<li><a href=\"https://ai.googleblog.com/2019/05/efficientnet-improving-accuracy-and.html\" target=\"_blank\">EfficientNet: Improving Accuracy and Efficiency through AutoML and Model Scaling</a></li>\n<li><a href=\"https://eumetrain.org/sites/default/files/2020-05/RGB_recipes.pdf\" target=\"_blank\">Compilation of RGB Recipes</a></li>\n</ul>",
  "messages": [
    {
      "id": 2405456,
      "postDate": "2023-08-23T22:03:04.507Z",
      "content": "<h1>1. Context section</h1>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/overview\" target=\"_blank\">Business context</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/data\" target=\"_blank\">Data context</a></li>\n</ul>\n<h1>2. Overview of the Approach</h1>\n<p>In this competition, the motivation is to contribute to the improvement of contrail prediction models, which are used to predict the formation of contrails generated by aircraft engine exhaust. Contrails are clouds of ice crystals with a line shape. They form when aircraft fly through super humid regions in the atmosphere and contribute greatly to global warming by trapping the heat in the atmosphere. Our goal is to validate the models that use satellite imagery as input data and then provide airlines with more accurate ways to mitigate the problem of climate change by avoiding the formation of contrails.</p>\n<p>In this study, we addressed the contrail detection challenge using geostationary satellite imagery from the GOES-16 ABI. Our approach involved data preprocessing for standardization and ash RGB image creation, contrail annotation with IoU-based object detection, and model training with EfficientNet and Dice coefficient loss in a K-Fold cross-validation setup. A distinctive feature was threshold selection for image masking, with a 0.52 threshold derived from the fifth image's Dice coefficient distribution. Validation included an error metric assessing labeling accuracy. Importantly, the Dice coefficient distributions from training on the fifth image alone and additional-image training closely aligned, confirming the chosen threshold's effectiveness.</p>\n<h1>3. Details of the submission</h1>\n<h2>3.1 Data Exploration</h2>\n<p>The data used in this competition are geostationary satellite images, sourced from the GOES-16 Advanced Baseline Imager (ABI). The images are provided as a sequence of 10-minute interval snapshots, where each sequence consists of eight images, with labeling applied only to the fifth image. Each dataset entry is represented by a unique record_id and contains precisely one labeled frame. The training dataset offers both individual label annotations and aggregated sound truth annotations, while the validation data only contains the latter.</p>\n<h2>3.2 Data Preprocessing</h2>\n<h3>3.2.1 Image Standardization</h3>\n<p>Data preprocessing plays a crucial role in this competition. The provided data consists of eight different spectral bands, and the pixel values are not in the standard 0-255 range. To make the satellite data suitable for our models, we first normalize and transform them using the following two steps:</p>\n<ul>\n<li>Map the pixel values of each spectral band to the standard range of 0 to 255 to create visualizable images;</li>\n<li>Create an ash RGB image by combining these eight bands, forming a three-channel color image.</li>\n</ul>\n<h3>3.2.2 Contrail Annotation</h3>\n<p>Next, we use the Intersection over Union (IoU) method to locate and annotate the contrails in the images. Our steps are as follows:</p>\n<ul>\n<li>Utilize object detection models to detect contrails within the ash RGB images. Outline them using bounding boxes;</li>\n<li>For potentially overlapping boxes, calculate the IoU to determine the degree of overlap between the boxes;</li>\n<li>Based on the calculation, merge smaller boxes with significant overlap into larger bounding boxes to enhance annotation accuracy and readability.</li>\n</ul>\n<h2>3.3 Model Training</h2>\n<h3>3.3.1 Random Crop and Image Selection</h3>\n<p>Random cropping is implemented as a form of data augmentation. We select only the fifth image from each sequence for training, as it contains the binary mask information (0 to 1). </p>\n<p>We applied a total of 25 different methods of cropping and the four labeled green are the most effective ones: <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2497266%2F782ddf910435c2bca436c491cfae2869%2Fcropping%20methods.png?generation=1692827919871994&amp;alt=media\" alt=\"cropping methods\"></p>\n<p>The following are the training results of each method: <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2497266%2Ff6aba1d97ceee02ef0296aebd4ee3e82%2Fcropping%20results.png?generation=1692827970836162&amp;alt=media\" alt=\"cropping results\"></p>\n<p>Some of the blank data is due to the fact that during training, we initially didn’t head in the right direction, so the experiment didn’t continue on those methods. In the end, we opted for the seventh method: Randomized square cropping with at least one complete contrail and the size of the cropped image is always greater than 128 pixels. The cropped images could be proportionally scaled as needed. For instance, we could expand them to 512 pixels by random cropping images greater than 256 pixels and then resizing them. Likewise, if we wanted to increase the image size to 1024 pixels, we followed a similar process by random cropping images larger than or equal to 512 pixels and resizing them to 1024 pixels. </p>\n<h3>3.3.2 EfficientNet with K-Fold Validation and Dice Coefficient</h3>\n<p>The model is trained using the EfficientNet architecture. The K-Fold cross-validation is employed to assess the performance of the model across different subsets of the training dataset and to ensure the robustness and generalization of the model. For loss calculation, the Dice coefficient function is used, which is useful for segmentation tasks.</p>\n<h3>3.3.3 Threshold Selection for Masking</h3>\n<p>After training the model with all the fifth images from each sequence, the next step is to apply the trained model to the rest of the images in each sequence (the first to fourth and sixth to eighth) to create the masks. The most creative part of our solution is to determine an appropriate threshold (or confidence) within the range from 0 to 1 for masking. This threshold is based on the distribution of the fifth image of each sequence. The following are the steps in detail:</p>\n<ul>\n<li>Apply the trained EfficientNet models to the first to fourth and sixth to eighth images in each sequence;</li>\n<li>Generate binary probability (ranging from 0 to 1) based on the model’s predictions;</li>\n<li>Determine an optimal threshold within the 0-1 range by analyzing the distribution of the fifth image;</li>\n<li>Apply the selected threshold to the rest of the images (the first to fourth and sixth to eighth).</li>\n</ul>\n<p>This approach leverages the power of our EfficientNet models and ensures that the models can better identify and retain the relevant images. </p>\n<p>The training results are illustrated in the following two figures. The left figure displays the distribution of Dice coefficients obtained when training exclusively with the fifth image of each sequence. Conversely, the right figure showcases the distribution of Dice coefficients when training with additional images, specifically the first to fourth and sixth to eighth images within each sequence. Remarkably, these two distributions exhibit a remarkable similarity, enabling us to determine a suitable threshold based on the distribution derived from the fifth image.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2497266%2F38af041ab03c025ae9fb897aaee999c3%2FDice%20Coefficient%20Distribution%20.png?generation=1692828079636644&amp;alt=media\" alt=\"dice coefficient distribution\"><br>\nTo assess the accuracy of our approach, we introduced an error metric. If the fifth image was not correctly labeled, it was counted as an error. Subsequently, we calculated the proportion of results that contained errors. Our statistical analysis revealed that a Dice coefficient exceeding 0.5275080198049544 (approximately equal to 0.52 in our case) reliably indicated utility. Therefore, the threshold was set at 0.52, ensuring that the chosen threshold effectively distinguishes valuable information in the images. </p>\n<h1>4. Sources</h1>\n<ul>\n<li><a href=\"https://ai.googleblog.com/2019/05/efficientnet-improving-accuracy-and.html\" target=\"_blank\">EfficientNet: Improving Accuracy and Efficiency through AutoML and Model Scaling</a></li>\n<li><a href=\"https://eumetrain.org/sites/default/files/2020-05/RGB_recipes.pdf\" target=\"_blank\">Compilation of RGB Recipes</a></li>\n</ul>",
      "rawMarkdown": "# 1. Context section\n- [Business context](https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/overview)\n- [Data context](https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/data)\n\n# 2. Overview of the Approach\nIn this competition, the motivation is to contribute to the improvement of contrail prediction models, which are used to predict the formation of contrails generated by aircraft engine exhaust. Contrails are clouds of ice crystals with a line shape. They form when aircraft fly through super humid regions in the atmosphere and contribute greatly to global warming by trapping the heat in the atmosphere. Our goal is to validate the models that use satellite imagery as input data and then provide airlines with more accurate ways to mitigate the problem of climate change by avoiding the formation of contrails.\n\nIn this study, we addressed the contrail detection challenge using geostationary satellite imagery from the GOES-16 ABI. Our approach involved data preprocessing for standardization and ash RGB image creation, contrail annotation with IoU-based object detection, and model training with EfficientNet and Dice coefficient loss in a K-Fold cross-validation setup. A distinctive feature was threshold selection for image masking, with a 0.52 threshold derived from the fifth image's Dice coefficient distribution. Validation included an error metric assessing labeling accuracy. Importantly, the Dice coefficient distributions from training on the fifth image alone and additional-image training closely aligned, confirming the chosen threshold's effectiveness.\n\n# 3. Details of the submission\n## 3.1 Data Exploration\nThe data used in this competition are geostationary satellite images, sourced from the GOES-16 Advanced Baseline Imager (ABI). The images are provided as a sequence of 10-minute interval snapshots, where each sequence consists of eight images, with labeling applied only to the fifth image. Each dataset entry is represented by a unique record_id and contains precisely one labeled frame. The training dataset offers both individual label annotations and aggregated sound truth annotations, while the validation data only contains the latter.\n\n## 3.2 Data Preprocessing\n### 3.2.1 Image Standardization\nData preprocessing plays a crucial role in this competition. The provided data consists of eight different spectral bands, and the pixel values are not in the standard 0-255 range. To make the satellite data suitable for our models, we first normalize and transform them using the following two steps:\n- Map the pixel values of each spectral band to the standard range of 0 to 255 to create visualizable images;\n- Create an ash RGB image by combining these eight bands, forming a three-channel color image.\n\n### 3.2.2 Contrail Annotation\nNext, we use the Intersection over Union (IoU) method to locate and annotate the contrails in the images. Our steps are as follows:\n- Utilize object detection models to detect contrails within the ash RGB images. Outline them using bounding boxes;\n- For potentially overlapping boxes, calculate the IoU to determine the degree of overlap between the boxes;\n- Based on the calculation, merge smaller boxes with significant overlap into larger bounding boxes to enhance annotation accuracy and readability.\n\n## 3.3 Model Training\n### 3.3.1 Random Crop and Image Selection\nRandom cropping is implemented as a form of data augmentation. We select only the fifth image from each sequence for training, as it contains the binary mask information (0 to 1). \n\nWe applied a total of 25 different methods of cropping and the four labeled green are the most effective ones: \n![cropping methods](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2497266%2F782ddf910435c2bca436c491cfae2869%2Fcropping%20methods.png?generation=1692827919871994&alt=media)\n\nThe following are the training results of each method: \n![cropping results](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2497266%2Ff6aba1d97ceee02ef0296aebd4ee3e82%2Fcropping%20results.png?generation=1692827970836162&alt=media)\n\nSome of the blank data is due to the fact that during training, we initially didn’t head in the right direction, so the experiment didn’t continue on those methods. In the end, we opted for the seventh method: Randomized square cropping with at least one complete contrail and the size of the cropped image is always greater than 128 pixels. The cropped images could be proportionally scaled as needed. For instance, we could expand them to 512 pixels by random cropping images greater than 256 pixels and then resizing them. Likewise, if we wanted to increase the image size to 1024 pixels, we followed a similar process by random cropping images larger than or equal to 512 pixels and resizing them to 1024 pixels. \n\n### 3.3.2 EfficientNet with K-Fold Validation and Dice Coefficient\nThe model is trained using the EfficientNet architecture. The K-Fold cross-validation is employed to assess the performance of the model across different subsets of the training dataset and to ensure the robustness and generalization of the model. For loss calculation, the Dice coefficient function is used, which is useful for segmentation tasks.\n\n### 3.3.3 Threshold Selection for Masking\nAfter training the model with all the fifth images from each sequence, the next step is to apply the trained model to the rest of the images in each sequence (the first to fourth and sixth to eighth) to create the masks. The most creative part of our solution is to determine an appropriate threshold (or confidence) within the range from 0 to 1 for masking. This threshold is based on the distribution of the fifth image of each sequence. The following are the steps in detail:\n- Apply the trained EfficientNet models to the first to fourth and sixth to eighth images in each sequence;\n- Generate binary probability (ranging from 0 to 1) based on the model’s predictions;\n- Determine an optimal threshold within the 0-1 range by analyzing the distribution of the fifth image;\n- Apply the selected threshold to the rest of the images (the first to fourth and sixth to eighth).\n\nThis approach leverages the power of our EfficientNet models and ensures that the models can better identify and retain the relevant images. \n\nThe training results are illustrated in the following two figures. The left figure displays the distribution of Dice coefficients obtained when training exclusively with the fifth image of each sequence. Conversely, the right figure showcases the distribution of Dice coefficients when training with additional images, specifically the first to fourth and sixth to eighth images within each sequence. Remarkably, these two distributions exhibit a remarkable similarity, enabling us to determine a suitable threshold based on the distribution derived from the fifth image.\n![dice coefficient distribution](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2497266%2F38af041ab03c025ae9fb897aaee999c3%2FDice%20Coefficient%20Distribution%20.png?generation=1692828079636644&alt=media)\nTo assess the accuracy of our approach, we introduced an error metric. If the fifth image was not correctly labeled, it was counted as an error. Subsequently, we calculated the proportion of results that contained errors. Our statistical analysis revealed that a Dice coefficient exceeding 0.5275080198049544 (approximately equal to 0.52 in our case) reliably indicated utility. Therefore, the threshold was set at 0.52, ensuring that the chosen threshold effectively distinguishes valuable information in the images. \n\n# 4. Sources\n- [EfficientNet: Improving Accuracy and Efficiency through AutoML and Model Scaling](https://ai.googleblog.com/2019/05/efficientnet-improving-accuracy-and.html)\n- [Compilation of RGB Recipes](https://eumetrain.org/sites/default/files/2020-05/RGB_recipes.pdf)",
      "votes": 2
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2405456": "# 1. Context section\n- [Business context](https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/overview)\n- [Data context](https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/data)\n\n# 2. Overview of the Approach\nIn this competition, the motivation is to contribute to the improvement of contrail prediction models, which are used to predict the formation of contrails generated by aircraft engine exhaust. Contrails are clouds of ice crystals with a line shape. They form when aircraft fly through super humid regions in the atmosphere and contribute greatly to global warming by trapping the heat in the atmosphere. Our goal is to validate the models that use satellite imagery as input data and then provide airlines with more accurate ways to mitigate the problem of climate change by avoiding the formation of contrails.\n\nIn this study, we addressed the contrail detection challenge using geostationary satellite imagery from the GOES-16 ABI. Our approach involved data preprocessing for standardization and ash RGB image creation, contrail annotation with IoU-based object detection, and model training with EfficientNet and Dice coefficient loss in a K-Fold cross-validation setup. A distinctive feature was threshold selection for image masking, with a 0.52 threshold derived from the fifth image's Dice coefficient distribution. Validation included an error metric assessing labeling accuracy. Importantly, the Dice coefficient distributions from training on the fifth image alone and additional-image training closely aligned, confirming the chosen threshold's effectiveness.\n\n# 3. Details of the submission\n## 3.1 Data Exploration\nThe data used in this competition are geostationary satellite images, sourced from the GOES-16 Advanced Baseline Imager (ABI). The images are provided as a sequence of 10-minute interval snapshots, where each sequence consists of eight images, with labeling applied only to the fifth image. Each dataset entry is represented by a unique record_id and contains precisely one labeled frame. The training dataset offers both individual label annotations and aggregated sound truth annotations, while the validation data only contains the latter.\n\n## 3.2 Data Preprocessing\n### 3.2.1 Image Standardization\nData preprocessing plays a crucial role in this competition. The provided data consists of eight different spectral bands, and the pixel values are not in the standard 0-255 range. To make the satellite data suitable for our models, we first normalize and transform them using the following two steps:\n- Map the pixel values of each spectral band to the standard range of 0 to 255 to create visualizable images;\n- Create an ash RGB image by combining these eight bands, forming a three-channel color image.\n\n### 3.2.2 Contrail Annotation\nNext, we use the Intersection over Union (IoU) method to locate and annotate the contrails in the images. Our steps are as follows:\n- Utilize object detection models to detect contrails within the ash RGB images. Outline them using bounding boxes;\n- For potentially overlapping boxes, calculate the IoU to determine the degree of overlap between the boxes;\n- Based on the calculation, merge smaller boxes with significant overlap into larger bounding boxes to enhance annotation accuracy and readability.\n\n## 3.3 Model Training\n### 3.3.1 Random Crop and Image Selection\nRandom cropping is implemented as a form of data augmentation. We select only the fifth image from each sequence for training, as it contains the binary mask information (0 to 1). \n\nWe applied a total of 25 different methods of cropping and the four labeled green are the most effective ones: \n![cropping methods](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2497266%2F782ddf910435c2bca436c491cfae2869%2Fcropping%20methods.png?generation=1692827919871994&alt=media)\n\nThe following are the training results of each method: \n![cropping results](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2497266%2Ff6aba1d97ceee02ef0296aebd4ee3e82%2Fcropping%20results.png?generation=1692827970836162&alt=media)\n\nSome of the blank data is due to the fact that during training, we initially didn’t head in the right direction, so the experiment didn’t continue on those methods. In the end, we opted for the seventh method: Randomized square cropping with at least one complete contrail and the size of the cropped image is always greater than 128 pixels. The cropped images could be proportionally scaled as needed. For instance, we could expand them to 512 pixels by random cropping images greater than 256 pixels and then resizing them. Likewise, if we wanted to increase the image size to 1024 pixels, we followed a similar process by random cropping images larger than or equal to 512 pixels and resizing them to 1024 pixels. \n\n### 3.3.2 EfficientNet with K-Fold Validation and Dice Coefficient\nThe model is trained using the EfficientNet architecture. The K-Fold cross-validation is employed to assess the performance of the model across different subsets of the training dataset and to ensure the robustness and generalization of the model. For loss calculation, the Dice coefficient function is used, which is useful for segmentation tasks.\n\n### 3.3.3 Threshold Selection for Masking\nAfter training the model with all the fifth images from each sequence, the next step is to apply the trained model to the rest of the images in each sequence (the first to fourth and sixth to eighth) to create the masks. The most creative part of our solution is to determine an appropriate threshold (or confidence) within the range from 0 to 1 for masking. This threshold is based on the distribution of the fifth image of each sequence. The following are the steps in detail:\n- Apply the trained EfficientNet models to the first to fourth and sixth to eighth images in each sequence;\n- Generate binary probability (ranging from 0 to 1) based on the model’s predictions;\n- Determine an optimal threshold within the 0-1 range by analyzing the distribution of the fifth image;\n- Apply the selected threshold to the rest of the images (the first to fourth and sixth to eighth).\n\nThis approach leverages the power of our EfficientNet models and ensures that the models can better identify and retain the relevant images. \n\nThe training results are illustrated in the following two figures. The left figure displays the distribution of Dice coefficients obtained when training exclusively with the fifth image of each sequence. Conversely, the right figure showcases the distribution of Dice coefficients when training with additional images, specifically the first to fourth and sixth to eighth images within each sequence. Remarkably, these two distributions exhibit a remarkable similarity, enabling us to determine a suitable threshold based on the distribution derived from the fifth image.\n![dice coefficient distribution](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2497266%2F38af041ab03c025ae9fb897aaee999c3%2FDice%20Coefficient%20Distribution%20.png?generation=1692828079636644&alt=media)\nTo assess the accuracy of our approach, we introduced an error metric. If the fifth image was not correctly labeled, it was counted as an error. Subsequently, we calculated the proportion of results that contained errors. Our statistical analysis revealed that a Dice coefficient exceeding 0.5275080198049544 (approximately equal to 0.52 in our case) reliably indicated utility. Therefore, the threshold was set at 0.52, ensuring that the chosen threshold effectively distinguishes valuable information in the images. \n\n# 4. Sources\n- [EfficientNet: Improving Accuracy and Efficiency through AutoML and Model Scaling](https://ai.googleblog.com/2019/05/efficientnet-improving-accuracy-and.html)\n- [Compilation of RGB Recipes](https://eumetrain.org/sites/default/files/2020-05/RGB_recipes.pdf)"
  }
}