{
  "id": 390975,
  "title": "18th place solution",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/390975",
  "author_name": "Naoki Kato",
  "post_date": "2023-02-28T02:28:46.903000",
  "votes": 22,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Thanks to Kaggle, the hosts and competitors for this meaningful competition.</p>\n<p>In the following, I want to provide a brief summary of my solution.</p>\n<h2>Overview</h2>\n<p>Similar to many public codes, my pipeline is as follows.</p>\n<ol>\n<li>Detect a breast area for each image and crop that area</li>\n<li>Predict cancer image-wise using various backbones</li>\n<li>Aggregate image-wise predictions and apply thresholding to get final prediction for each target</li>\n</ol>\n<h2>Preprocessing</h2>\n<p>My preprocessing depends on many public codes. I am grateful to the authors of those codes.</p>\n<p>Sigmoid/linear windowing is applied based on <code>VOILUTFunction</code>, <code>WindowCenter</code> and <code>WindowWidth</code> in dicom data. After windowing, images are processed with min-max scaling and treated as 8-bit images.</p>\n<h2>Breast detector</h2>\n<p>I annotated breast bounding boxes for about 1000 images. In addition to those labels, I also used labels provided by <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> (about 500 images) in <a href=\"https://www.kaggle.com/code/remekkinas/breast-cancer-roi-brest-extractor/notebook\" target=\"_blank\">this code</a> to train a single YOLOv5n6 with the input size of 1024. mAP_0.5:0.95 of a validation split is 0.952.</p>\n<p>Given the detections, affine transformation is applied to obtain fixed size cropped images.</p>\n<p>At that time, expanding the bboxes so that the aspect ratio and the size of the bbox relative to the original images did not change too much improved somewhat of local cv.</p>\n<p>However, the submission with the highest private LB did not use a detector…</p>\n<h2>Cancer model</h2>\n<p>I used <code>tf_efficientnet_b3.ns_jft_in1k</code>, <code>tf_efficientnetv2_s.in21k</code>, <code>eca_nfnet_l1</code> and <code>dm_nfnet_f0</code> in timm library for ensembling in my final submisison. Each model is trained with different input size (800×1200 to 1024×1536), learning rate and training epochs.</p>\n<p>GeM pooling with fixed p is used as global pooling.</p>\n<h2>Augmentations</h2>\n<p>I used the following list of augmentations implemented with imgaug.</p>\n<pre><code>iaa.Sometimes(\n    ,\n    iaa.Affine(\n        rotate=(-, ),\n        shear=(-, ),\n        scale={: scale_x, : scale_y},\n        translate_percent={: shift_x, : shift_y},\n    ),\n),\niaa.Resize({: image_size[], : image_size[]},\n           interpolation=cv2.INTER_LINEAR),\niaa.Fliplr(),\niaa.Flipud(  row[] ==   ),\niaa.Sequential([\n    iaa.Sometimes(, iaa.SomeOf(, [\n        iaa.GaussianBlur(sigma=(, )),\n        iaa.AdditiveGaussianNoise(scale=(, )),\n    ])),\n    iaa.Sometimes(, iaa.Multiply((, ))),\n    iaa.Sometimes(, iaa.LinearContrast((, ))),\n], random_order=),\n</code></pre>\n<h2>Training settings</h2>\n<p>I applied bce loss for positive samples and focal loss for negative samples, since I thought that the correct identification of hard negative cases would contribute to the improvement of pF1.</p>\n<p>As an auxiliary loss, focal loss to classify invasive is adopted.</p>\n<p>To mitigate overfitting, I employed exponential moving average (ema) of model weights with warmup where the decay is 1 + t / 10 + t and t is training iteration.</p>\n<p>The final model weights are obtained by averaging normaly trained weights and ema weights that had the highest pF1 on a validation set.</p>\n<p>Other training settings are as follows:</p>\n<ul>\n<li>Optimizer: AdamW with weight decay 1e-3</li>\n<li>Scheduler: OneCycleLR</li>\n<li>6~8 epochs of training depending on backbones</li>\n</ul>\n<h2>Post-processing</h2>\n<ul>\n<li>Flip test is used as TTA</li>\n<li>Image-wise predictions were aggregated by LP pooling, where p is determined based on the validation score</li>\n<li>Ensemble is performed by simple averaging</li>\n<li>Referring to <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243463\" target=\"_blank\">2nd place solution of BirdCLEF 2021</a>, percentile based thresholding is adopted</li>\n</ul>\n<h2>Submissions</h2>\n<p>My final submission and the best submission are as follows. Simpler method performed better on private data in my case.</p>\n<table>\n<thead>\n<tr>\n<th>Submission</th>\n<th>Backbones</th>\n<th>W/ detector</th>\n<th>Agg. method</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Final submission</td>\n<td>tf_efficientnet_b3.ns_jft_in1k, tf_efficientnetv2_s.in21k, eca_nfnet_l1, dm_nfnet_f0</td>\n<td>✔︎</td>\n<td>LP pooling</td>\n<td>0.65</td>\n<td>0.48</td>\n</tr>\n<tr>\n<td>Best private LB</td>\n<td>tf_efficientnet_b3.ns_jft_in1k, tf_efficientnetv2_s.in21k, 2x nfnet_l0</td>\n<td>×</td>\n<td>mean</td>\n<td>0.56</td>\n<td>0.50</td>\n</tr>\n</tbody>\n</table>\n<p>Thank you for reading.</p>",
  "messages": [
    {
      "id": 2162122,
      "postDate": "2023-02-28T02:28:46.903Z",
      "content": "<p>Thanks to Kaggle, the hosts and competitors for this meaningful competition.</p>\n<p>In the following, I want to provide a brief summary of my solution.</p>\n<h2>Overview</h2>\n<p>Similar to many public codes, my pipeline is as follows.</p>\n<ol>\n<li>Detect a breast area for each image and crop that area</li>\n<li>Predict cancer image-wise using various backbones</li>\n<li>Aggregate image-wise predictions and apply thresholding to get final prediction for each target</li>\n</ol>\n<h2>Preprocessing</h2>\n<p>My preprocessing depends on many public codes. I am grateful to the authors of those codes.</p>\n<p>Sigmoid/linear windowing is applied based on <code>VOILUTFunction</code>, <code>WindowCenter</code> and <code>WindowWidth</code> in dicom data. After windowing, images are processed with min-max scaling and treated as 8-bit images.</p>\n<h2>Breast detector</h2>\n<p>I annotated breast bounding boxes for about 1000 images. In addition to those labels, I also used labels provided by <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> (about 500 images) in <a href=\"https://www.kaggle.com/code/remekkinas/breast-cancer-roi-brest-extractor/notebook\" target=\"_blank\">this code</a> to train a single YOLOv5n6 with the input size of 1024. mAP_0.5:0.95 of a validation split is 0.952.</p>\n<p>Given the detections, affine transformation is applied to obtain fixed size cropped images.</p>\n<p>At that time, expanding the bboxes so that the aspect ratio and the size of the bbox relative to the original images did not change too much improved somewhat of local cv.</p>\n<p>However, the submission with the highest private LB did not use a detector…</p>\n<h2>Cancer model</h2>\n<p>I used <code>tf_efficientnet_b3.ns_jft_in1k</code>, <code>tf_efficientnetv2_s.in21k</code>, <code>eca_nfnet_l1</code> and <code>dm_nfnet_f0</code> in timm library for ensembling in my final submisison. Each model is trained with different input size (800×1200 to 1024×1536), learning rate and training epochs.</p>\n<p>GeM pooling with fixed p is used as global pooling.</p>\n<h2>Augmentations</h2>\n<p>I used the following list of augmentations implemented with imgaug.</p>\n<pre><code>iaa.Sometimes(\n    ,\n    iaa.Affine(\n        rotate=(-, ),\n        shear=(-, ),\n        scale={: scale_x, : scale_y},\n        translate_percent={: shift_x, : shift_y},\n    ),\n),\niaa.Resize({: image_size[], : image_size[]},\n           interpolation=cv2.INTER_LINEAR),\niaa.Fliplr(),\niaa.Flipud(  row[] ==   ),\niaa.Sequential([\n    iaa.Sometimes(, iaa.SomeOf(, [\n        iaa.GaussianBlur(sigma=(, )),\n        iaa.AdditiveGaussianNoise(scale=(, )),\n    ])),\n    iaa.Sometimes(, iaa.Multiply((, ))),\n    iaa.Sometimes(, iaa.LinearContrast((, ))),\n], random_order=),\n</code></pre>\n<h2>Training settings</h2>\n<p>I applied bce loss for positive samples and focal loss for negative samples, since I thought that the correct identification of hard negative cases would contribute to the improvement of pF1.</p>\n<p>As an auxiliary loss, focal loss to classify invasive is adopted.</p>\n<p>To mitigate overfitting, I employed exponential moving average (ema) of model weights with warmup where the decay is 1 + t / 10 + t and t is training iteration.</p>\n<p>The final model weights are obtained by averaging normaly trained weights and ema weights that had the highest pF1 on a validation set.</p>\n<p>Other training settings are as follows:</p>\n<ul>\n<li>Optimizer: AdamW with weight decay 1e-3</li>\n<li>Scheduler: OneCycleLR</li>\n<li>6~8 epochs of training depending on backbones</li>\n</ul>\n<h2>Post-processing</h2>\n<ul>\n<li>Flip test is used as TTA</li>\n<li>Image-wise predictions were aggregated by LP pooling, where p is determined based on the validation score</li>\n<li>Ensemble is performed by simple averaging</li>\n<li>Referring to <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243463\" target=\"_blank\">2nd place solution of BirdCLEF 2021</a>, percentile based thresholding is adopted</li>\n</ul>\n<h2>Submissions</h2>\n<p>My final submission and the best submission are as follows. Simpler method performed better on private data in my case.</p>\n<table>\n<thead>\n<tr>\n<th>Submission</th>\n<th>Backbones</th>\n<th>W/ detector</th>\n<th>Agg. method</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Final submission</td>\n<td>tf_efficientnet_b3.ns_jft_in1k, tf_efficientnetv2_s.in21k, eca_nfnet_l1, dm_nfnet_f0</td>\n<td>✔︎</td>\n<td>LP pooling</td>\n<td>0.65</td>\n<td>0.48</td>\n</tr>\n<tr>\n<td>Best private LB</td>\n<td>tf_efficientnet_b3.ns_jft_in1k, tf_efficientnetv2_s.in21k, 2x nfnet_l0</td>\n<td>×</td>\n<td>mean</td>\n<td>0.56</td>\n<td>0.50</td>\n</tr>\n</tbody>\n</table>\n<p>Thank you for reading.</p>",
      "rawMarkdown": "Thanks to Kaggle, the hosts and competitors for this meaningful competition.\n\nIn the following, I want to provide a brief summary of my solution.\n\n## Overview\n\nSimilar to many public codes, my pipeline is as follows.\n\n1. Detect a breast area for each image and crop that area\n2. Predict cancer image-wise using various backbones\n3. Aggregate image-wise predictions and apply thresholding to get final prediction for each target\n\n## Preprocessing\n\nMy preprocessing depends on many public codes. I am grateful to the authors of those codes.\n\nSigmoid/linear windowing is applied based on `VOILUTFunction`, `WindowCenter` and `WindowWidth` in dicom data. After windowing, images are processed with min-max scaling and treated as 8-bit images.\n\n## Breast detector\n\nI annotated breast bounding boxes for about 1000 images. In addition to those labels, I also used labels provided by @remekkinas (about 500 images) in [this code](https://www.kaggle.com/code/remekkinas/breast-cancer-roi-brest-extractor/notebook) to train a single YOLOv5n6 with the input size of 1024. mAP_0.5:0.95 of a validation split is 0.952.\n\nGiven the detections, affine transformation is applied to obtain fixed size cropped images.\n\nAt that time, expanding the bboxes so that the aspect ratio and the size of the bbox relative to the original images did not change too much improved somewhat of local cv.\n\nHowever, the submission with the highest private LB did not use a detector…\n\n## Cancer model\n\nI used `tf_efficientnet_b3.ns_jft_in1k`, `tf_efficientnetv2_s.in21k`, `eca_nfnet_l1` and `dm_nfnet_f0` in timm library for ensembling in my final submisison. Each model is trained with different input size (800×1200 to 1024×1536), learning rate and training epochs.\n\nGeM pooling with fixed p is used as global pooling.\n\n## Augmentations\n\nI used the following list of augmentations implemented with imgaug.\n\n```python\niaa.Sometimes(\n    0.5,\n    iaa.Affine(\n        rotate=(-15, 15),\n        shear=(-3, 3),\n        scale={'x': scale_x, 'y': scale_y},\n        translate_percent={'x': shift_x, 'y': shift_y},\n    ),\n),\niaa.Resize({\"width\": image_size[0], \"height\": image_size[1]},\n           interpolation=cv2.INTER_LINEAR),\niaa.Fliplr(0.5),\niaa.Flipud(0.1 if row['view'] == 'CC' else 0.01),\niaa.Sequential([\n    iaa.Sometimes(0.1, iaa.SomeOf(1, [\n        iaa.GaussianBlur(sigma=(0, 1.5)),\n        iaa.AdditiveGaussianNoise(scale=(1.0, 4.0)),\n    ])),\n    iaa.Sometimes(0.3, iaa.Multiply((0.95, 1.05))),\n    iaa.Sometimes(0.3, iaa.LinearContrast((0.90, 1.10))),\n], random_order=True),\n```\n\n## Training settings\n\nI applied bce loss for positive samples and focal loss for negative samples, since I thought that the correct identification of hard negative cases would contribute to the improvement of pF1.\n\nAs an auxiliary loss, focal loss to classify invasive is adopted.\n\nTo mitigate overfitting, I employed exponential moving average (ema) of model weights with warmup where the decay is 1 + t / 10 + t and t is training iteration.\n\nThe final model weights are obtained by averaging normaly trained weights and ema weights that had the highest pF1 on a validation set.\n\nOther training settings are as follows:\n\n- Optimizer: AdamW with weight decay 1e-3\n- Scheduler: OneCycleLR\n- 6~8 epochs of training depending on backbones\n\n## Post-processing\n\n- Flip test is used as TTA\n- Image-wise predictions were aggregated by LP pooling, where p is determined based on the validation score\n- Ensemble is performed by simple averaging\n- Referring to [2nd place solution of BirdCLEF 2021](https://www.kaggle.com/competitions/birdclef-2021/discussion/243463), percentile based thresholding is adopted\n\n## Submissions\n\nMy final submission and the best submission are as follows. Simpler method performed better on private data in my case.\n\n| Submission | Backbones | W/ detector | Agg. method | Public LB | Private LB |\n| --- | --- | --- | --- | --- | --- |\n| Final submission | tf_efficientnet_b3.ns_jft_in1k, tf_efficientnetv2_s.in21k, eca_nfnet_l1, dm_nfnet_f0 | ✔︎ | LP pooling | 0.65 | 0.48 |\n| Best private LB | tf_efficientnet_b3.ns_jft_in1k, tf_efficientnetv2_s.in21k, 2x nfnet_l0 | × | mean | 0.56 | 0.50 |\n\n\nThank you for reading.",
      "votes": 22
    },
    {
      "id": 2162434,
      "postDate": "2023-02-28T07:51:15.960Z",
      "content": "<p>Great work! Smart to also include the best private solution and score as it will help further work, should be a write-up standard 👍</p>",
      "rawMarkdown": "Great work! Smart to also include the best private solution and score as it will help further work, should be a write-up standard 👍",
      "votes": 1,
      "replies": [
        {
          "id": 2162519,
          "postDate": "2023-02-28T09:04:35.450Z",
          "content": "<p>Thanks! I'm happy to hear that.</p>",
          "rawMarkdown": "Thanks! I'm happy to hear that."
        }
      ]
    },
    {
      "id": 2162372,
      "postDate": "2023-02-28T06:49:35.143Z",
      "content": "<p>Well done! I see OneCycleLR. Any others you experimented with?</p>",
      "rawMarkdown": "Well done! I see OneCycleLR. Any others you experimented with?",
      "replies": [
        {
          "id": 2162530,
          "postDate": "2023-02-28T09:15:14.357Z",
          "content": "<p>Thanks! I only experimented this shceduler with deferent pct_start. 0.2 is used in my solution, but there wasn't a big difference.</p>",
          "rawMarkdown": "Thanks! I only experimented this shceduler with deferent pct_start. 0.2 is used in my solution, but there wasn't a big difference.",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2162434,
      "author_name": "Kirderf",
      "author_url": "",
      "post_date": "2023-02-28T07:51:15.960000",
      "content": "<p>Great work! Smart to also include the best private solution and score as it will help further work, should be a write-up standard 👍</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2162519,
          "author_name": "Naoki Kato",
          "author_url": "",
          "post_date": "2023-02-28T09:04:35.450000",
          "content": "<p>Thanks! I'm happy to hear that.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2162372,
      "author_name": "Antti Isosalo",
      "author_url": "",
      "post_date": "2023-02-28T06:49:35.143000",
      "content": "<p>Well done! I see OneCycleLR. Any others you experimented with?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2162530,
          "author_name": "Naoki Kato",
          "author_url": "",
          "post_date": "2023-02-28T09:15:14.357000",
          "content": "<p>Thanks! I only experimented this shceduler with deferent pct_start. 0.2 is used in my solution, but there wasn't a big difference.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2162122": "Thanks to Kaggle, the hosts and competitors for this meaningful competition.\n\nIn the following, I want to provide a brief summary of my solution.\n\n## Overview\n\nSimilar to many public codes, my pipeline is as follows.\n\n1. Detect a breast area for each image and crop that area\n2. Predict cancer image-wise using various backbones\n3. Aggregate image-wise predictions and apply thresholding to get final prediction for each target\n\n## Preprocessing\n\nMy preprocessing depends on many public codes. I am grateful to the authors of those codes.\n\nSigmoid/linear windowing is applied based on `VOILUTFunction`, `WindowCenter` and `WindowWidth` in dicom data. After windowing, images are processed with min-max scaling and treated as 8-bit images.\n\n## Breast detector\n\nI annotated breast bounding boxes for about 1000 images. In addition to those labels, I also used labels provided by @remekkinas (about 500 images) in [this code](https://www.kaggle.com/code/remekkinas/breast-cancer-roi-brest-extractor/notebook) to train a single YOLOv5n6 with the input size of 1024. mAP_0.5:0.95 of a validation split is 0.952.\n\nGiven the detections, affine transformation is applied to obtain fixed size cropped images.\n\nAt that time, expanding the bboxes so that the aspect ratio and the size of the bbox relative to the original images did not change too much improved somewhat of local cv.\n\nHowever, the submission with the highest private LB did not use a detector…\n\n## Cancer model\n\nI used `tf_efficientnet_b3.ns_jft_in1k`, `tf_efficientnetv2_s.in21k`, `eca_nfnet_l1` and `dm_nfnet_f0` in timm library for ensembling in my final submisison. Each model is trained with different input size (800×1200 to 1024×1536), learning rate and training epochs.\n\nGeM pooling with fixed p is used as global pooling.\n\n## Augmentations\n\nI used the following list of augmentations implemented with imgaug.\n\n```python\niaa.Sometimes(\n    0.5,\n    iaa.Affine(\n        rotate=(-15, 15),\n        shear=(-3, 3),\n        scale={'x': scale_x, 'y': scale_y},\n        translate_percent={'x': shift_x, 'y': shift_y},\n    ),\n),\niaa.Resize({\"width\": image_size[0], \"height\": image_size[1]},\n           interpolation=cv2.INTER_LINEAR),\niaa.Fliplr(0.5),\niaa.Flipud(0.1 if row['view'] == 'CC' else 0.01),\niaa.Sequential([\n    iaa.Sometimes(0.1, iaa.SomeOf(1, [\n        iaa.GaussianBlur(sigma=(0, 1.5)),\n        iaa.AdditiveGaussianNoise(scale=(1.0, 4.0)),\n    ])),\n    iaa.Sometimes(0.3, iaa.Multiply((0.95, 1.05))),\n    iaa.Sometimes(0.3, iaa.LinearContrast((0.90, 1.10))),\n], random_order=True),\n```\n\n## Training settings\n\nI applied bce loss for positive samples and focal loss for negative samples, since I thought that the correct identification of hard negative cases would contribute to the improvement of pF1.\n\nAs an auxiliary loss, focal loss to classify invasive is adopted.\n\nTo mitigate overfitting, I employed exponential moving average (ema) of model weights with warmup where the decay is 1 + t / 10 + t and t is training iteration.\n\nThe final model weights are obtained by averaging normaly trained weights and ema weights that had the highest pF1 on a validation set.\n\nOther training settings are as follows:\n\n- Optimizer: AdamW with weight decay 1e-3\n- Scheduler: OneCycleLR\n- 6~8 epochs of training depending on backbones\n\n## Post-processing\n\n- Flip test is used as TTA\n- Image-wise predictions were aggregated by LP pooling, where p is determined based on the validation score\n- Ensemble is performed by simple averaging\n- Referring to [2nd place solution of BirdCLEF 2021](https://www.kaggle.com/competitions/birdclef-2021/discussion/243463), percentile based thresholding is adopted\n\n## Submissions\n\nMy final submission and the best submission are as follows. Simpler method performed better on private data in my case.\n\n| Submission | Backbones | W/ detector | Agg. method | Public LB | Private LB |\n| --- | --- | --- | --- | --- | --- |\n| Final submission | tf_efficientnet_b3.ns_jft_in1k, tf_efficientnetv2_s.in21k, eca_nfnet_l1, dm_nfnet_f0 | ✔︎ | LP pooling | 0.65 | 0.48 |\n| Best private LB | tf_efficientnet_b3.ns_jft_in1k, tf_efficientnetv2_s.in21k, 2x nfnet_l0 | × | mean | 0.56 | 0.50 |\n\n\nThank you for reading.",
    "2162434": "Great work! Smart to also include the best private solution and score as it will help further work, should be a write-up standard 👍",
    "2162372": "Well done! I see OneCycleLR. Any others you experimented with?"
  }
}