{
  "id": 169425,
  "title": "16th place solution (all you need to know)",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/169425",
  "author_name": "Dracarys",
  "post_date": "2020-07-23T19:46:14.233000",
  "votes": 14,
  "comment_count": 1,
  "views": 0,
  "content": "<p>First of all, I want to thank kaggle and the organizers for hosting such interesting competition.\nAnd i like to thank my teammates <a href=\"/mpware\">@mpware</a> , <a href=\"/tikutiku\">@tikutiku</a>,  <a href=\"/phoenix9032\">@phoenix9032</a>  and <a href=\"/virajbagal\">@virajbagal</a>  for such an interesting competition journey. And congratulations to all the winners!\nIt's been a great competition, and my team has spend a lot of time in this competition and finally glad to share that all the hard work paid off.</p>\n\n<h1>What worked:</h1>\n\n<h2>Image Preprocessing</h2>\n\n<ul>\n<li>We experiment with a lot of tiling strategies including publicly shared by <a href=\"/iafoss\">@iafoss</a> and <a href=\"/akensert\">@akensert</a> and end up using both in our training pipeline. (25x256x256) worked best for us.</li>\n<li>We also randomly replaced white tiles (tiles with color avg &gt; 240 ) with other tiles and it gives boost in our CV.</li>\n<li><p>Removing duplicates (or adding them to only train set), and removing few noisy images as mentioned <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/151323\">here</a> helped a little.</p>\n\n<h2>Augmentations</h2></li>\n<li><p>Our models are trained on a wide range of augmentations. Few of them are shown below: \n``` \ntransforms_train = albumentations.Compose([\nalbumentations.HorizontalFlip(p=0.3),\nalbumentations.VerticalFlip(p=0.3),\nalbumentations.Transpose(p=0.3),\nalbumentations.RandomGridShuffle(),\nalbumentations.OneOf([\n    albumentations.GaussianBlur(blur_limit=1),\n    albumentations.MotionBlur(),\n], p=0.01),\n])</p>\n\n<h2>use this after initial training for 40 epochs</h2></li>\n</ul>\n\n<p>transforms_train_hard = albumentations.Compose([\n    albumentations.HorizontalFlip(p=0.5),\n    albumentations.VerticalFlip(p=0.5),\n    albumentations.Transpose(p=0.3),\n    albumentations.ShiftScaleRotate(shift_limit=0.1, scale_limit=0.2, rotate_limit=5,\n                                    border_mode=cv2.BORDER_CONSTANT, p=0.1),\n    albumentations.RandomGridShuffle(),\n])\n```</p>\n\n<h2>Ensemble</h2>\n\n<p>Our final solution is a simple average of the following models.\nFor Karolinska:\n<code>\n1. ResNet34\n2. EfficientNet-B0\n3. EfficientNet-B1\n</code>\nFor Radboud:\n<code>\n1. EfficientNet-B0\n2. ResNet34\n3. EfficientNet-B1\n4. UNet with EfficientNet-B1 backbone, segmentation head + classification head\n</code></p>\n\n<h1>The short story about UNet models (MPWARE speaking):</h1>\n\n<p>Main part of the team was working on CNN approaches without masks so I decided to focus only on models involving masks to try to add more diversity to our ensemble. \nThe idea was that segmentation should help the model to discover what does matter for our target.\nAs masks were different between Radboud (0 to 5) and Karolinska (0 to 2), Unet models per provider were trained separately. Input was a 25x128x128 tile-based image from half medium resolution (mask built accordingly). \nAround 100 WSI had missing masks so it was a chance to create, with some other additional WSI, an hold-out balanced dataset to follow correlation between CV, Hold-Out and LB.\nIf all correlate then our models could be considered stable and safe for private dataset. After a few trainings we succeeded to have CV and Hold-Out correlated and we noticed that LB was always better.\nWe also noticed that CV were quite bad on Karolinska and quite good on Radboud. After some investigations and attempts to improve Karolinska masks, it became obvious that segmentation approach will not work for Karolinska. However, it had potential to bring boost on Radboud so more models were trained by applying random density (between 0.2 and 0.9) tile selection to cover most possible tiles configuration. This procedure provided a nice boost +0.01 but required around 128 epochs. 4 folds of such UNet/radboud models were integrated in the ensemble (using classification head only to average). Finally, it seems to have contributed to stabilize our ensemble.</p>\n\n<h1>What did not worked:</h1>\n\n<ol>\n<li>Removing gray backgrounds from radboud images.</li>\n<li>Training models with Multiple Sample Dropouts.</li>\n<li>Using H&amp;E Normalizations and augmentations.</li>\n<li>Relabelling bad labels.</li>\n<li>Models with heavy base like EfficientNet-B3,B4,B5, etc.</li>\n<li>RegNet y and x various variants 080 till 32.</li>\n<li>Adding data_provider as feature.</li>\n<li>Adding extra head to predict <code>isup_grade</code> from <code>gleason_scores</code>.</li>\n<li>ArcFaceloss</li>\n<li>Multitask learning with more weight on the Gleason Score base loss</li>\n</ol>",
  "messages": [
    {
      "id": 942528,
      "postDate": "2020-07-23T19:46:14.233Z",
      "content": "<p>First of all, I want to thank kaggle and the organizers for hosting such interesting competition.\nAnd i like to thank my teammates <a href=\"/mpware\">@mpware</a> , <a href=\"/tikutiku\">@tikutiku</a>,  <a href=\"/phoenix9032\">@phoenix9032</a>  and <a href=\"/virajbagal\">@virajbagal</a>  for such an interesting competition journey. And congratulations to all the winners!\nIt's been a great competition, and my team has spend a lot of time in this competition and finally glad to share that all the hard work paid off.</p>\n\n<h1>What worked:</h1>\n\n<h2>Image Preprocessing</h2>\n\n<ul>\n<li>We experiment with a lot of tiling strategies including publicly shared by <a href=\"/iafoss\">@iafoss</a> and <a href=\"/akensert\">@akensert</a> and end up using both in our training pipeline. (25x256x256) worked best for us.</li>\n<li>We also randomly replaced white tiles (tiles with color avg &gt; 240 ) with other tiles and it gives boost in our CV.</li>\n<li><p>Removing duplicates (or adding them to only train set), and removing few noisy images as mentioned <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/151323\">here</a> helped a little.</p>\n\n<h2>Augmentations</h2></li>\n<li><p>Our models are trained on a wide range of augmentations. Few of them are shown below: \n``` \ntransforms_train = albumentations.Compose([\nalbumentations.HorizontalFlip(p=0.3),\nalbumentations.VerticalFlip(p=0.3),\nalbumentations.Transpose(p=0.3),\nalbumentations.RandomGridShuffle(),\nalbumentations.OneOf([\n    albumentations.GaussianBlur(blur_limit=1),\n    albumentations.MotionBlur(),\n], p=0.01),\n])</p>\n\n<h2>use this after initial training for 40 epochs</h2></li>\n</ul>\n\n<p>transforms_train_hard = albumentations.Compose([\n    albumentations.HorizontalFlip(p=0.5),\n    albumentations.VerticalFlip(p=0.5),\n    albumentations.Transpose(p=0.3),\n    albumentations.ShiftScaleRotate(shift_limit=0.1, scale_limit=0.2, rotate_limit=5,\n                                    border_mode=cv2.BORDER_CONSTANT, p=0.1),\n    albumentations.RandomGridShuffle(),\n])\n```</p>\n\n<h2>Ensemble</h2>\n\n<p>Our final solution is a simple average of the following models.\nFor Karolinska:\n<code>\n1. ResNet34\n2. EfficientNet-B0\n3. EfficientNet-B1\n</code>\nFor Radboud:\n<code>\n1. EfficientNet-B0\n2. ResNet34\n3. EfficientNet-B1\n4. UNet with EfficientNet-B1 backbone, segmentation head + classification head\n</code></p>\n\n<h1>The short story about UNet models (MPWARE speaking):</h1>\n\n<p>Main part of the team was working on CNN approaches without masks so I decided to focus only on models involving masks to try to add more diversity to our ensemble. \nThe idea was that segmentation should help the model to discover what does matter for our target.\nAs masks were different between Radboud (0 to 5) and Karolinska (0 to 2), Unet models per provider were trained separately. Input was a 25x128x128 tile-based image from half medium resolution (mask built accordingly). \nAround 100 WSI had missing masks so it was a chance to create, with some other additional WSI, an hold-out balanced dataset to follow correlation between CV, Hold-Out and LB.\nIf all correlate then our models could be considered stable and safe for private dataset. After a few trainings we succeeded to have CV and Hold-Out correlated and we noticed that LB was always better.\nWe also noticed that CV were quite bad on Karolinska and quite good on Radboud. After some investigations and attempts to improve Karolinska masks, it became obvious that segmentation approach will not work for Karolinska. However, it had potential to bring boost on Radboud so more models were trained by applying random density (between 0.2 and 0.9) tile selection to cover most possible tiles configuration. This procedure provided a nice boost +0.01 but required around 128 epochs. 4 folds of such UNet/radboud models were integrated in the ensemble (using classification head only to average). Finally, it seems to have contributed to stabilize our ensemble.</p>\n\n<h1>What did not worked:</h1>\n\n<ol>\n<li>Removing gray backgrounds from radboud images.</li>\n<li>Training models with Multiple Sample Dropouts.</li>\n<li>Using H&amp;E Normalizations and augmentations.</li>\n<li>Relabelling bad labels.</li>\n<li>Models with heavy base like EfficientNet-B3,B4,B5, etc.</li>\n<li>RegNet y and x various variants 080 till 32.</li>\n<li>Adding data_provider as feature.</li>\n<li>Adding extra head to predict <code>isup_grade</code> from <code>gleason_scores</code>.</li>\n<li>ArcFaceloss</li>\n<li>Multitask learning with more weight on the Gleason Score base loss</li>\n</ol>",
      "rawMarkdown": "First of all, I want to thank kaggle and the organizers for hosting such interesting competition.\nAnd i like to thank my teammates @mpware , @tikutiku,  @phoenix9032  and @virajbagal  for such an interesting competition journey. And congratulations to all the winners!\nIt's been a great competition, and my team has spend a lot of time in this competition and finally glad to share that all the hard work paid off.\n# What worked:\n## Image Preprocessing\n* We experiment with a lot of tiling strategies including publicly shared by @iafoss and @akensert and end up using both in our training pipeline. (25x256x256) worked best for us.\n* We also randomly replaced white tiles (tiles with color avg &gt; 240 ) with other tiles and it gives boost in our CV.\n* Removing duplicates (or adding them to only train set), and removing few noisy images as mentioned [here](https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/151323) helped a little.\n## Augmentations\n* Our models are trained on a wide range of augmentations. Few of them are shown below: \n``` \ntransforms_train = albumentations.Compose([\n    albumentations.HorizontalFlip(p=0.3),\n    albumentations.VerticalFlip(p=0.3),\n    albumentations.Transpose(p=0.3),\n    albumentations.RandomGridShuffle(),\n    albumentations.OneOf([\n        albumentations.GaussianBlur(blur_limit=1),\n        albumentations.MotionBlur(),\n    ], p=0.01),\n])\n##use this after initial training for 40 epochs\ntransforms_train_hard = albumentations.Compose([\n    albumentations.HorizontalFlip(p=0.5),\n    albumentations.VerticalFlip(p=0.5),\n    albumentations.Transpose(p=0.3),\n    albumentations.ShiftScaleRotate(shift_limit=0.1, scale_limit=0.2, rotate_limit=5,\n                                    border_mode=cv2.BORDER_CONSTANT, p=0.1),\n    albumentations.RandomGridShuffle(),\n])\n```\n## Ensemble\nOur final solution is a simple average of the following models.\nFor Karolinska:\n```\n1. ResNet34\n2. EfficientNet-B0\n3. EfficientNet-B1\n```\nFor Radboud:\n```\n1. EfficientNet-B0\n2. ResNet34\n3. EfficientNet-B1\n4. UNet with EfficientNet-B1 backbone, segmentation head + classification head\n```\n# The short story about UNet models (MPWARE speaking):\nMain part of the team was working on CNN approaches without masks so I decided to focus only on models involving masks to try to add more diversity to our ensemble. \nThe idea was that segmentation should help the model to discover what does matter for our target.\nAs masks were different between Radboud (0 to 5) and Karolinska (0 to 2), Unet models per provider were trained separately. Input was a 25x128x128 tile-based image from half medium resolution (mask built accordingly). \nAround 100 WSI had missing masks so it was a chance to create, with some other additional WSI, an hold-out balanced dataset to follow correlation between CV, Hold-Out and LB.\nIf all correlate then our models could be considered stable and safe for private dataset. After a few trainings we succeeded to have CV and Hold-Out correlated and we noticed that LB was always better.\nWe also noticed that CV were quite bad on Karolinska and quite good on Radboud. After some investigations and attempts to improve Karolinska masks, it became obvious that segmentation approach will not work for Karolinska. However, it had potential to bring boost on Radboud so more models were trained by applying random density (between 0.2 and 0.9) tile selection to cover most possible tiles configuration. This procedure provided a nice boost +0.01 but required around 128 epochs. 4 folds of such UNet/radboud models were integrated in the ensemble (using classification head only to average). Finally, it seems to have contributed to stabilize our ensemble.\n\n# What did not worked:\n1. Removing gray backgrounds from radboud images.\n2. Training models with Multiple Sample Dropouts.\n3. Using H&amp;E Normalizations and augmentations.\n4. Relabelling bad labels.\n5. Models with heavy base like EfficientNet-B3,B4,B5, etc.\n6. RegNet y and x various variants 080 till 32.\n7. Adding data_provider as feature.\n8. Adding extra head to predict `isup_grade` from `gleason_scores`.\n9. ArcFaceloss\n10. Multitask learning with more weight on the Gleason Score base loss",
      "votes": 14
    },
    {
      "id": 942576,
      "postDate": "2020-07-23T20:41:23.650Z",
      "content": "<p>Just to add that both CV and HoldOut on UNet/Radboud was within 0.83-0.86 range.</p>",
      "rawMarkdown": "Just to add that both CV and HoldOut on UNet/Radboud was within 0.83-0.86 range.",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 942576,
      "author_name": "MPWARE",
      "author_url": "",
      "post_date": "2020-07-23T20:41:23.650000",
      "content": "<p>Just to add that both CV and HoldOut on UNet/Radboud was within 0.83-0.86 range.</p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "942528": "First of all, I want to thank kaggle and the organizers for hosting such interesting competition.\nAnd i like to thank my teammates @mpware , @tikutiku,  @phoenix9032  and @virajbagal  for such an interesting competition journey. And congratulations to all the winners!\nIt's been a great competition, and my team has spend a lot of time in this competition and finally glad to share that all the hard work paid off.\n# What worked:\n## Image Preprocessing\n* We experiment with a lot of tiling strategies including publicly shared by @iafoss and @akensert and end up using both in our training pipeline. (25x256x256) worked best for us.\n* We also randomly replaced white tiles (tiles with color avg &gt; 240 ) with other tiles and it gives boost in our CV.\n* Removing duplicates (or adding them to only train set), and removing few noisy images as mentioned [here](https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/151323) helped a little.\n## Augmentations\n* Our models are trained on a wide range of augmentations. Few of them are shown below: \n``` \ntransforms_train = albumentations.Compose([\n    albumentations.HorizontalFlip(p=0.3),\n    albumentations.VerticalFlip(p=0.3),\n    albumentations.Transpose(p=0.3),\n    albumentations.RandomGridShuffle(),\n    albumentations.OneOf([\n        albumentations.GaussianBlur(blur_limit=1),\n        albumentations.MotionBlur(),\n    ], p=0.01),\n])\n##use this after initial training for 40 epochs\ntransforms_train_hard = albumentations.Compose([\n    albumentations.HorizontalFlip(p=0.5),\n    albumentations.VerticalFlip(p=0.5),\n    albumentations.Transpose(p=0.3),\n    albumentations.ShiftScaleRotate(shift_limit=0.1, scale_limit=0.2, rotate_limit=5,\n                                    border_mode=cv2.BORDER_CONSTANT, p=0.1),\n    albumentations.RandomGridShuffle(),\n])\n```\n## Ensemble\nOur final solution is a simple average of the following models.\nFor Karolinska:\n```\n1. ResNet34\n2. EfficientNet-B0\n3. EfficientNet-B1\n```\nFor Radboud:\n```\n1. EfficientNet-B0\n2. ResNet34\n3. EfficientNet-B1\n4. UNet with EfficientNet-B1 backbone, segmentation head + classification head\n```\n# The short story about UNet models (MPWARE speaking):\nMain part of the team was working on CNN approaches without masks so I decided to focus only on models involving masks to try to add more diversity to our ensemble. \nThe idea was that segmentation should help the model to discover what does matter for our target.\nAs masks were different between Radboud (0 to 5) and Karolinska (0 to 2), Unet models per provider were trained separately. Input was a 25x128x128 tile-based image from half medium resolution (mask built accordingly). \nAround 100 WSI had missing masks so it was a chance to create, with some other additional WSI, an hold-out balanced dataset to follow correlation between CV, Hold-Out and LB.\nIf all correlate then our models could be considered stable and safe for private dataset. After a few trainings we succeeded to have CV and Hold-Out correlated and we noticed that LB was always better.\nWe also noticed that CV were quite bad on Karolinska and quite good on Radboud. After some investigations and attempts to improve Karolinska masks, it became obvious that segmentation approach will not work for Karolinska. However, it had potential to bring boost on Radboud so more models were trained by applying random density (between 0.2 and 0.9) tile selection to cover most possible tiles configuration. This procedure provided a nice boost +0.01 but required around 128 epochs. 4 folds of such UNet/radboud models were integrated in the ensemble (using classification head only to average). Finally, it seems to have contributed to stabilize our ensemble.\n\n# What did not worked:\n1. Removing gray backgrounds from radboud images.\n2. Training models with Multiple Sample Dropouts.\n3. Using H&amp;E Normalizations and augmentations.\n4. Relabelling bad labels.\n5. Models with heavy base like EfficientNet-B3,B4,B5, etc.\n6. RegNet y and x various variants 080 till 32.\n7. Adding data_provider as feature.\n8. Adding extra head to predict `isup_grade` from `gleason_scores`.\n9. ArcFaceloss\n10. Multitask learning with more weight on the Gleason Score base loss",
    "942576": "Just to add that both CV and HoldOut on UNet/Radboud was within 0.83-0.86 range."
  }
}