{
  "id": 358130,
  "title": "7th place solution -> 2-step MIL based strategy",
  "url": "/competitions/mayo-clinic-strip-ai/discussion/358130",
  "author_name": "tdiceman",
  "post_date": "2022-10-06T18:05:09.251000",
  "votes": 15,
  "comment_count": 9,
  "views": 0,
  "content": "<p>It was a nice starting experience for me in this competition and big scoop of beginner's luck to go with. Nevertheless, learnt a lot and final methodology with all learning implemented worked best, which is satisfying. Thanks to all who shared hints, tips and suggestions during this competitions. Here is a summary of the methodology used:</p>\n<p><br></p>\n<h2>Preprocessing - Tile Creation and Selection</h2>\n<p>Tried a lot of things at the start but finally settled with tiles created from image downsized by a factor of 6 to balance signal loss and manageable number of tiles. Created and selected tiles by:</p>\n<ul>\n<li>Creating (224,224,3) shape tiles and applying simple threshold on image array unique values and their counts - for removing mostly monochrome and background tiles. </li>\n<li>Reduced tile number further by running tiles through EfficientNetB0 (pretrained, ImageNet weights), scoring them and choosing top 50-60% tiles for each WSI (created another problem in tile number per image, addressed towards the end).</li>\n<li>Added augmented data for LAA class by creating extra tiles simply by padding image and shifting tile location in image by tile_size/2 in both directions. </li>\n</ul>\n<p><br></p>\n<h2>MIL - Pseudo Labelled Tiles based Feature Extractor Training, Feature Aggregation and Classifier</h2>\n<p>Followed the tile feature embedding and then aggregation at slide level for further classification method. There are two main steps:</p>\n<h3>1. Feature Extractor Training</h3>\n<p><a href=\"https://www.kaggle.com/code/icemantd/feature-cluster-tile-psuedo-label-train-mayo\" target=\"_blank\">This notebook</a> has detail implementation of this part. The main points are:</p>\n<ul>\n<li>Trained EfficientnetB0 starting with ImageNet weights, first epoch used all tiles with slide level labels.</li>\n<li>Added random hue augmentation to tiles (details towards the end).</li>\n<li>Defined a feature extractor which is the whole network upto one layer (shape (1,1280)) before the last FC classification layer.</li>\n<li>Used tile features from extractor to apply cluster (kmeans) distance based pseudo labels to tiles for next iteration.</li>\n</ul>\n<p>The logic used here is to attempt to provide the feature extractor the ability to not only give high output for the two classes but also a low output for out of class or non-critical tiles. Hence I used 'Other' label tiles with target [0.0] and LAA and CE with [1,0] and [0,1] respectively, with sigmoid activation and BCE loss. Every training epoch, features are created, clustering performed, and tiles scored on minimum distance from the negative class - i.e. top 20% highest min distance from LAA+Other feature clusters for CE tiles are labelled [0,1] and bottom 10% lowest min distance CE tiles are labelled [0,0] - and vice-versa for LAA tile pseudo labelling. The hope was this discriminates between the two classes, as well as between critical and non-critical tiles. The main clustering idea is inspired by <a href=\"https://arxiv.org/pdf/2206.08861.pdf\" target=\"_blank\">this</a> paper.</p>\n<ul>\n<li>Used random sampling of CE class to balance classes and used a simple metric for training evaluation at slide level - maxpooling AUC scores, i.e. taking AUC scores for maximum predictions of both classes by any tiles in the slide. 4-fold CV AUC scores of 0.6-0.65.</li>\n</ul>\n<h3>2. Feature Aggregation, Classification and Inference</h3>\n<p>Obtained tile level features from feature extractor and adapted <a href=\"https://arxiv.org/pdf/2011.08939.pdf\" target=\"_blank\">this</a> paper's methodology for feature aggregation at slide level, based on attention to tile with max output logit. </p>\n<ul>\n<li>Built a simple random forest classifier using aggregated features at slide level and max predictions at slide level (any tile) for both classes as extra features.</li>\n<li>Used only top 25 dark tiles from any image for training classifier and for inference.</li>\n<li>Inference: Take top 25 dark tiles from image upto 2.2 Gb in size, create features from extractor and max logits from classifier head, aggregate and predict classifier probabilities (<a href=\"https://www.kaggle.com/code/icemantd/new-predict-mil-logit-features-mayo-strip-ai/notebook\" target=\"_blank\">notebook</a>)</li>\n</ul>\n<p><br></p>\n<h2>Some important observations</h2>\n<ul>\n<li>Found out that tile selection was very important, especially if you are using smaller tiles. Observed that at one point, larger images were giving better log_loss scores but there are more smaller images to infer (at least in training dataset). So adjused the top% of tiles selected from larger images, especially because LAA class had large images and some images dominated the tile numbers and resulted in overfitting. This definitely helped in controlling penalty on bad log_loss over large number of images. Example of log loss spread over training data vs number of tiles per slide image (number of tiles per image represents original number created and represents size of image and not number of tiles selected for training):</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8536316%2Fbe12d4336f614efd1bd81a02d54edb67%2Flog_loss_tilenum_imbalance.png?generation=1665076862015109&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>Observed from <a href=\"https://arxiv.org/pdf/1902.06543.pdf\" target=\"_blank\">this</a> paper that adjusting hue of the image, even a little bit, resulted in better chance of predicting unseen strains better. So I implemented a random hue augmentation based on stretching/compressing the hue channel of each tile randomly by a factor betweem (0.85, 1.15). For example</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8536316%2F155fc81cbdc429a4c96bba1cc9c5c0cf%2Fhue_aug.png?generation=1665077255141305&amp;alt=media\" alt=\"![\"></p>\n<ul>\n<li>Due to the 2-step complexity, I could not come up with a robust end-to-end CV pipeline. Only spent time on feature extractor CV and evaluation using various data splits (60-40, 90-10) but ran into various issues with validation metric robustness. This part was the weakest for me and that's why I consider myself lucky to still land on something that kind of worked.</li>\n</ul>\n<p>Hopefully some of you can point out where things could have been better or what was unecessary and even incorrectly implemented. Thanks.</p>",
  "messages": [
    {
      "id": 1975327,
      "postDate": "2022-10-06T18:05:09.253Z",
      "content": "<p>It was a nice starting experience for me in this competition and big scoop of beginner's luck to go with. Nevertheless, learnt a lot and final methodology with all learning implemented worked best, which is satisfying. Thanks to all who shared hints, tips and suggestions during this competitions. Here is a summary of the methodology used:</p>\n<p><br></p>\n<h2>Preprocessing - Tile Creation and Selection</h2>\n<p>Tried a lot of things at the start but finally settled with tiles created from image downsized by a factor of 6 to balance signal loss and manageable number of tiles. Created and selected tiles by:</p>\n<ul>\n<li>Creating (224,224,3) shape tiles and applying simple threshold on image array unique values and their counts - for removing mostly monochrome and background tiles. </li>\n<li>Reduced tile number further by running tiles through EfficientNetB0 (pretrained, ImageNet weights), scoring them and choosing top 50-60% tiles for each WSI (created another problem in tile number per image, addressed towards the end).</li>\n<li>Added augmented data for LAA class by creating extra tiles simply by padding image and shifting tile location in image by tile_size/2 in both directions. </li>\n</ul>\n<p><br></p>\n<h2>MIL - Pseudo Labelled Tiles based Feature Extractor Training, Feature Aggregation and Classifier</h2>\n<p>Followed the tile feature embedding and then aggregation at slide level for further classification method. There are two main steps:</p>\n<h3>1. Feature Extractor Training</h3>\n<p><a href=\"https://www.kaggle.com/code/icemantd/feature-cluster-tile-psuedo-label-train-mayo\" target=\"_blank\">This notebook</a> has detail implementation of this part. The main points are:</p>\n<ul>\n<li>Trained EfficientnetB0 starting with ImageNet weights, first epoch used all tiles with slide level labels.</li>\n<li>Added random hue augmentation to tiles (details towards the end).</li>\n<li>Defined a feature extractor which is the whole network upto one layer (shape (1,1280)) before the last FC classification layer.</li>\n<li>Used tile features from extractor to apply cluster (kmeans) distance based pseudo labels to tiles for next iteration.</li>\n</ul>\n<p>The logic used here is to attempt to provide the feature extractor the ability to not only give high output for the two classes but also a low output for out of class or non-critical tiles. Hence I used 'Other' label tiles with target [0.0] and LAA and CE with [1,0] and [0,1] respectively, with sigmoid activation and BCE loss. Every training epoch, features are created, clustering performed, and tiles scored on minimum distance from the negative class - i.e. top 20% highest min distance from LAA+Other feature clusters for CE tiles are labelled [0,1] and bottom 10% lowest min distance CE tiles are labelled [0,0] - and vice-versa for LAA tile pseudo labelling. The hope was this discriminates between the two classes, as well as between critical and non-critical tiles. The main clustering idea is inspired by <a href=\"https://arxiv.org/pdf/2206.08861.pdf\" target=\"_blank\">this</a> paper.</p>\n<ul>\n<li>Used random sampling of CE class to balance classes and used a simple metric for training evaluation at slide level - maxpooling AUC scores, i.e. taking AUC scores for maximum predictions of both classes by any tiles in the slide. 4-fold CV AUC scores of 0.6-0.65.</li>\n</ul>\n<h3>2. Feature Aggregation, Classification and Inference</h3>\n<p>Obtained tile level features from feature extractor and adapted <a href=\"https://arxiv.org/pdf/2011.08939.pdf\" target=\"_blank\">this</a> paper's methodology for feature aggregation at slide level, based on attention to tile with max output logit. </p>\n<ul>\n<li>Built a simple random forest classifier using aggregated features at slide level and max predictions at slide level (any tile) for both classes as extra features.</li>\n<li>Used only top 25 dark tiles from any image for training classifier and for inference.</li>\n<li>Inference: Take top 25 dark tiles from image upto 2.2 Gb in size, create features from extractor and max logits from classifier head, aggregate and predict classifier probabilities (<a href=\"https://www.kaggle.com/code/icemantd/new-predict-mil-logit-features-mayo-strip-ai/notebook\" target=\"_blank\">notebook</a>)</li>\n</ul>\n<p><br></p>\n<h2>Some important observations</h2>\n<ul>\n<li>Found out that tile selection was very important, especially if you are using smaller tiles. Observed that at one point, larger images were giving better log_loss scores but there are more smaller images to infer (at least in training dataset). So adjused the top% of tiles selected from larger images, especially because LAA class had large images and some images dominated the tile numbers and resulted in overfitting. This definitely helped in controlling penalty on bad log_loss over large number of images. Example of log loss spread over training data vs number of tiles per slide image (number of tiles per image represents original number created and represents size of image and not number of tiles selected for training):</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8536316%2Fbe12d4336f614efd1bd81a02d54edb67%2Flog_loss_tilenum_imbalance.png?generation=1665076862015109&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>Observed from <a href=\"https://arxiv.org/pdf/1902.06543.pdf\" target=\"_blank\">this</a> paper that adjusting hue of the image, even a little bit, resulted in better chance of predicting unseen strains better. So I implemented a random hue augmentation based on stretching/compressing the hue channel of each tile randomly by a factor betweem (0.85, 1.15). For example</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8536316%2F155fc81cbdc429a4c96bba1cc9c5c0cf%2Fhue_aug.png?generation=1665077255141305&amp;alt=media\" alt=\"![\"></p>\n<ul>\n<li>Due to the 2-step complexity, I could not come up with a robust end-to-end CV pipeline. Only spent time on feature extractor CV and evaluation using various data splits (60-40, 90-10) but ran into various issues with validation metric robustness. This part was the weakest for me and that's why I consider myself lucky to still land on something that kind of worked.</li>\n</ul>\n<p>Hopefully some of you can point out where things could have been better or what was unecessary and even incorrectly implemented. Thanks.</p>",
      "rawMarkdown": "It was a nice starting experience for me in this competition and big scoop of beginner's luck to go with. Nevertheless, learnt a lot and final methodology with all learning implemented worked best, which is satisfying. Thanks to all who shared hints, tips and suggestions during this competitions. Here is a summary of the methodology used:\n\n</br>\n## Preprocessing - Tile Creation and Selection\n\nTried a lot of things at the start but finally settled with tiles created from image downsized by a factor of 6 to balance signal loss and manageable number of tiles. Created and selected tiles by:\n\n- Creating (224,224,3) shape tiles and applying simple threshold on image array unique values and their counts - for removing mostly monochrome and background tiles. \n- Reduced tile number further by running tiles through EfficientNetB0 (pretrained, ImageNet weights), scoring them and choosing top 50-60% tiles for each WSI (created another problem in tile number per image, addressed towards the end).\n- Added augmented data for LAA class by creating extra tiles simply by padding image and shifting tile location in image by tile_size/2 in both directions. \n\n</br>\n## MIL - Pseudo Labelled Tiles based Feature Extractor Training, Feature Aggregation and Classifier\n\nFollowed the tile feature embedding and then aggregation at slide level for further classification method. There are two main steps:\n\n### 1. Feature Extractor Training\n[This notebook](https://www.kaggle.com/code/icemantd/feature-cluster-tile-psuedo-label-train-mayo) has detail implementation of this part. The main points are:\n- Trained EfficientnetB0 starting with ImageNet weights, first epoch used all tiles with slide level labels.\n- Added random hue augmentation to tiles (details towards the end).\n- Defined a feature extractor which is the whole network upto one layer (shape (1,1280)) before the last FC classification layer.\n- Used tile features from extractor to apply cluster (kmeans) distance based pseudo labels to tiles for next iteration.\n\nThe logic used here is to attempt to provide the feature extractor the ability to not only give high output for the two classes but also a low output for out of class or non-critical tiles. Hence I used 'Other' label tiles with target [0.0] and LAA and CE with [1,0] and [0,1] respectively, with sigmoid activation and BCE loss. Every training epoch, features are created, clustering performed, and tiles scored on minimum distance from the negative class - i.e. top 20% highest min distance from LAA+Other feature clusters for CE tiles are labelled [0,1] and bottom 10% lowest min distance CE tiles are labelled [0,0] - and vice-versa for LAA tile pseudo labelling. The hope was this discriminates between the two classes, as well as between critical and non-critical tiles. The main clustering idea is inspired by [this](https://arxiv.org/pdf/2206.08861.pdf) paper.\n\n- Used random sampling of CE class to balance classes and used a simple metric for training evaluation at slide level - maxpooling AUC scores, i.e. taking AUC scores for maximum predictions of both classes by any tiles in the slide. 4-fold CV AUC scores of 0.6-0.65.\n\n### 2. Feature Aggregation, Classification and Inference\nObtained tile level features from feature extractor and adapted [this](https://arxiv.org/pdf/2011.08939.pdf) paper's methodology for feature aggregation at slide level, based on attention to tile with max output logit. \n\n- Built a simple random forest classifier using aggregated features at slide level and max predictions at slide level (any tile) for both classes as extra features.\n- Used only top 25 dark tiles from any image for training classifier and for inference.\n- Inference: Take top 25 dark tiles from image upto 2.2 Gb in size, create features from extractor and max logits from classifier head, aggregate and predict classifier probabilities ([notebook](https://www.kaggle.com/code/icemantd/new-predict-mil-logit-features-mayo-strip-ai/notebook))\n\n</br>\n## Some important observations\n- Found out that tile selection was very important, especially if you are using smaller tiles. Observed that at one point, larger images were giving better log_loss scores but there are more smaller images to infer (at least in training dataset). So adjused the top% of tiles selected from larger images, especially because LAA class had large images and some images dominated the tile numbers and resulted in overfitting. This definitely helped in controlling penalty on bad log_loss over large number of images. Example of log loss spread over training data vs number of tiles per slide image (number of tiles per image represents original number created and represents size of image and not number of tiles selected for training):\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8536316%2Fbe12d4336f614efd1bd81a02d54edb67%2Flog_loss_tilenum_imbalance.png?generation=1665076862015109&alt=media)\n\n- Observed from [this](https://arxiv.org/pdf/1902.06543.pdf) paper that adjusting hue of the image, even a little bit, resulted in better chance of predicting unseen strains better. So I implemented a random hue augmentation based on stretching/compressing the hue channel of each tile randomly by a factor betweem (0.85, 1.15). For example\n\n![![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8536316%2F155fc81cbdc429a4c96bba1cc9c5c0cf%2Fhue_aug.png?generation=1665077255141305&alt=media)\n\n- Due to the 2-step complexity, I could not come up with a robust end-to-end CV pipeline. Only spent time on feature extractor CV and evaluation using various data splits (60-40, 90-10) but ran into various issues with validation metric robustness. This part was the weakest for me and that's why I consider myself lucky to still land on something that kind of worked.\n\nHopefully some of you can point out where things could have been better or what was unecessary and even incorrectly implemented. Thanks.",
      "votes": 15
    },
    {
      "id": 2048812,
      "postDate": "2022-11-29T17:35:12.990Z",
      "content": "<p>This is a really practical way!</p>",
      "rawMarkdown": "This is a really practical way!",
      "votes": 1,
      "replies": [
        {
          "id": 2048982,
          "postDate": "2022-11-29T20:34:35.707Z",
          "content": "<p><a href=\"https://www.kaggle.com/deepkim\" target=\"_blank\">@deepkim</a> thanks :)</p>",
          "rawMarkdown": "@deepkim thanks :)"
        }
      ]
    },
    {
      "id": 1978839,
      "postDate": "2022-10-09T03:35:02.087Z",
      "content": "<p>Congratulations! Good job 👍</p>",
      "rawMarkdown": "Congratulations! Good job 👍",
      "votes": 1,
      "replies": [
        {
          "id": 1979649,
          "postDate": "2022-10-09T15:45:49.370Z",
          "content": "<p><a href=\"https://www.kaggle.com/willcramptonn\" target=\"_blank\">@willcramptonn</a> thanks</p>",
          "rawMarkdown": "@willcramptonn thanks"
        }
      ]
    },
    {
      "id": 1978676,
      "postDate": "2022-10-08T22:45:18.870Z",
      "content": "<p>Wow!! So much work goes into successful project. Congrats👍  and thanks for sharing.</p>",
      "rawMarkdown": "Wow!! So much work goes into successful project. Congrats👍  and thanks for sharing.",
      "votes": 1,
      "replies": [
        {
          "id": 1978748,
          "postDate": "2022-10-09T01:55:36.803Z",
          "content": "<p><a href=\"https://www.kaggle.com/cid007\" target=\"_blank\">@cid007</a> thanks so much for your encouraging comments  :)</p>",
          "rawMarkdown": "@cid007 thanks so much for your encouraging comments  :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 1975700,
      "postDate": "2022-10-07T01:03:51.170Z",
      "content": "<p>Congratulations, looks like a huge amount of work here and so a well deserved placing. Lots of innovative ideas like the clustering, use of \"other\" images, and 2-step process.</p>\n<p>IIUC, you downsampled the whole slides by 6-times, and then pulled tiles of 224x224 after dropping monochrome tiles and tiles that scored in the bottom 50th percentile on a pre-trained EfficientNetB0. But then even after that, based on the first figure, some images still had 100+ tiles, some even with 500+ tiles. Is that correct? I'm surprised there were that many tiles available after the pre-processing.</p>\n<p>Also, would you mind posting details of how you randomly adjusted the hue? Based on the notebook you posted I presume that with Tensorflow - is that correct? I used TF too but couldn't figure out how to implement a hue augmentation/adjustment. I suspect that hue adjustment was key because based on my testing I believe the hidden test set contained other center_ids than the n=11 in the training dataset, so it's probable the MSB staining methodology at new sites was different than the training set.</p>\n<p>Anyway, some really impressive work here</p>",
      "rawMarkdown": "Congratulations, looks like a huge amount of work here and so a well deserved placing. Lots of innovative ideas like the clustering, use of \"other\" images, and 2-step process.\n\nIIUC, you downsampled the whole slides by 6-times, and then pulled tiles of 224x224 after dropping monochrome tiles and tiles that scored in the bottom 50th percentile on a pre-trained EfficientNetB0. But then even after that, based on the first figure, some images still had 100+ tiles, some even with 500+ tiles. Is that correct? I'm surprised there were that many tiles available after the pre-processing.\n\nAlso, would you mind posting details of how you randomly adjusted the hue? Based on the notebook you posted I presume that with Tensorflow - is that correct? I used TF too but couldn't figure out how to implement a hue augmentation/adjustment. I suspect that hue adjustment was key because based on my testing I believe the hidden test set contained other center_ids than the n=11 in the training dataset, so it's probable the MSB staining methodology at new sites was different than the training set.\n\nAnyway, some really impressive work here\n\n",
      "votes": 1,
      "replies": [
        {
          "id": 1975718,
          "postDate": "2022-10-07T01:47:06.673Z",
          "content": "<p><a href=\"https://www.kaggle.com/josephmarturano\" target=\"_blank\">@josephmarturano</a> thanks a lot for your comments. </p>\n<p>To clarify, I always had all tiles and its data in a csv at image level, I just modified another csv that read certain selected tiles. So the plots you see use data for images that originally had up to 600 tiles, but actually in the two cases I select much lower number for training. This was a way to also represent image size data, such that I know that originally high tile number images and low tile number images were showing clear difference when I correct this imbalance. Hope that clears up that point - after two selection process and reduction of tiles for larger images, max tiles per image we about 60-70 in the training set for LAA class (edit: edited the main text, thanks for pointing out the confusion).</p>\n<p>I have to update the notebook shared, it was posted before I made all these final changes. But here is the function for hue augmentation:</p>\n<pre><code>def hue_augument(img):\n\n    hsv = cv2.cvtColor(img, cv2.COLOR_BGR2HSV)\n    f = np.random.uniform(low=0.85, high=1.15, size=1)[0]\n    hsv[:,:,0] = hsv[:,:,0]*f\n    img_aug = cv2.cvtColor(hsv, cv2.COLOR_HSV2BGR)\n\n    return img_aug\n</code></pre>\n<p>You're correct about unseen data strains, and with that issue in mind I found the mentioned article and so I implemented this. I did not have time to verify if without this the prediction methodology would suffer significantly. Perhaps something to try further.</p>",
          "rawMarkdown": "@josephmarturano thanks a lot for your comments. \n\nTo clarify, I always had all tiles and its data in a csv at image level, I just modified another csv that read certain selected tiles. So the plots you see use data for images that originally had up to 600 tiles, but actually in the two cases I select much lower number for training. This was a way to also represent image size data, such that I know that originally high tile number images and low tile number images were showing clear difference when I correct this imbalance. Hope that clears up that point - after two selection process and reduction of tiles for larger images, max tiles per image we about 60-70 in the training set for LAA class (edit: edited the main text, thanks for pointing out the confusion).\n\nI have to update the notebook shared, it was posted before I made all these final changes. But here is the function for hue augmentation:\n\n\n```\ndef hue_augument(img):\n\n    hsv = cv2.cvtColor(img, cv2.COLOR_BGR2HSV)\n    f = np.random.uniform(low=0.85, high=1.15, size=1)[0]\n    hsv[:,:,0] = hsv[:,:,0]*f\n    img_aug = cv2.cvtColor(hsv, cv2.COLOR_HSV2BGR)\n\n    return img_aug\n```\n\nYou're correct about unseen data strains, and with that issue in mind I found the mentioned article and so I implemented this. I did not have time to verify if without this the prediction methodology would suffer significantly. Perhaps something to try further.",
          "votes": 1
        },
        {
          "id": 1976449,
          "postDate": "2022-10-07T11:12:11.583Z",
          "content": "<p>Thanks that helps clarify that <code>num_tiles</code> in those plots are effectively a proxy for image size. </p>\n<p>And thanks for the <code>hue_augment</code> function, I actually have an upcoming work project where I suspect that function will be useful. There is <a href=\"https://github.com/tensorflow/tensorflow/blob/359c3cdfc5fabac82b3c70b3b6de2b0a8c16874f/tensorflow/python/ops/image_ops_impl.py#L2610-L2656\" target=\"_blank\">Tensorflow's random_hue</a> but I like your function better because it gives more flexibility on the underlying distribution (e.g., random, normal, etc.).</p>",
          "rawMarkdown": "Thanks that helps clarify that `num_tiles` in those plots are effectively a proxy for image size. \n\nAnd thanks for the `hue_augment` function, I actually have an upcoming work project where I suspect that function will be useful. There is [Tensorflow's random_hue](https://github.com/tensorflow/tensorflow/blob/359c3cdfc5fabac82b3c70b3b6de2b0a8c16874f/tensorflow/python/ops/image_ops_impl.py#L2610-L2656) but I like your function better because it gives more flexibility on the underlying distribution (e.g., random, normal, etc.).",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2048812,
      "author_name": "kaggler",
      "author_url": "",
      "post_date": "2022-11-29T17:35:12.990000",
      "content": "<p>This is a really practical way!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2048982,
          "author_name": "tdiceman",
          "author_url": "",
          "post_date": "2022-11-29T20:34:35.707000",
          "content": "<p><a href=\"https://www.kaggle.com/deepkim\" target=\"_blank\">@deepkim</a> thanks :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1978839,
      "author_name": "Will",
      "author_url": "",
      "post_date": "2022-10-09T03:35:02.087000",
      "content": "<p>Congratulations! Good job 👍</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1979649,
          "author_name": "tdiceman",
          "author_url": "",
          "post_date": "2022-10-09T15:45:49.370000",
          "content": "<p><a href=\"https://www.kaggle.com/willcramptonn\" target=\"_blank\">@willcramptonn</a> thanks</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1978676,
      "author_name": "Chirag Desai",
      "author_url": "",
      "post_date": "2022-10-08T22:45:18.870000",
      "content": "<p>Wow!! So much work goes into successful project. Congrats👍  and thanks for sharing.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1978748,
          "author_name": "tdiceman",
          "author_url": "",
          "post_date": "2022-10-09T01:55:36.803000",
          "content": "<p><a href=\"https://www.kaggle.com/cid007\" target=\"_blank\">@cid007</a> thanks so much for your encouraging comments  :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1975700,
      "author_name": "Joe Marturano",
      "author_url": "",
      "post_date": "2022-10-07T01:03:51.170000",
      "content": "<p>Congratulations, looks like a huge amount of work here and so a well deserved placing. Lots of innovative ideas like the clustering, use of \"other\" images, and 2-step process.</p>\n<p>IIUC, you downsampled the whole slides by 6-times, and then pulled tiles of 224x224 after dropping monochrome tiles and tiles that scored in the bottom 50th percentile on a pre-trained EfficientNetB0. But then even after that, based on the first figure, some images still had 100+ tiles, some even with 500+ tiles. Is that correct? I'm surprised there were that many tiles available after the pre-processing.</p>\n<p>Also, would you mind posting details of how you randomly adjusted the hue? Based on the notebook you posted I presume that with Tensorflow - is that correct? I used TF too but couldn't figure out how to implement a hue augmentation/adjustment. I suspect that hue adjustment was key because based on my testing I believe the hidden test set contained other center_ids than the n=11 in the training dataset, so it's probable the MSB staining methodology at new sites was different than the training set.</p>\n<p>Anyway, some really impressive work here</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1975718,
          "author_name": "tdiceman",
          "author_url": "",
          "post_date": "2022-10-07T01:47:06.673000",
          "content": "<p><a href=\"https://www.kaggle.com/josephmarturano\" target=\"_blank\">@josephmarturano</a> thanks a lot for your comments. </p>\n<p>To clarify, I always had all tiles and its data in a csv at image level, I just modified another csv that read certain selected tiles. So the plots you see use data for images that originally had up to 600 tiles, but actually in the two cases I select much lower number for training. This was a way to also represent image size data, such that I know that originally high tile number images and low tile number images were showing clear difference when I correct this imbalance. Hope that clears up that point - after two selection process and reduction of tiles for larger images, max tiles per image we about 60-70 in the training set for LAA class (edit: edited the main text, thanks for pointing out the confusion).</p>\n<p>I have to update the notebook shared, it was posted before I made all these final changes. But here is the function for hue augmentation:</p>\n<pre><code>def hue_augument(img):\n\n    hsv = cv2.cvtColor(img, cv2.COLOR_BGR2HSV)\n    f = np.random.uniform(low=0.85, high=1.15, size=1)[0]\n    hsv[:,:,0] = hsv[:,:,0]*f\n    img_aug = cv2.cvtColor(hsv, cv2.COLOR_HSV2BGR)\n\n    return img_aug\n</code></pre>\n<p>You're correct about unseen data strains, and with that issue in mind I found the mentioned article and so I implemented this. I did not have time to verify if without this the prediction methodology would suffer significantly. Perhaps something to try further.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1976449,
          "author_name": "Joe Marturano",
          "author_url": "",
          "post_date": "2022-10-07T11:12:11.583000",
          "content": "<p>Thanks that helps clarify that <code>num_tiles</code> in those plots are effectively a proxy for image size. </p>\n<p>And thanks for the <code>hue_augment</code> function, I actually have an upcoming work project where I suspect that function will be useful. There is <a href=\"https://github.com/tensorflow/tensorflow/blob/359c3cdfc5fabac82b3c70b3b6de2b0a8c16874f/tensorflow/python/ops/image_ops_impl.py#L2610-L2656\" target=\"_blank\">Tensorflow's random_hue</a> but I like your function better because it gives more flexibility on the underlying distribution (e.g., random, normal, etc.).</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1975327": "It was a nice starting experience for me in this competition and big scoop of beginner's luck to go with. Nevertheless, learnt a lot and final methodology with all learning implemented worked best, which is satisfying. Thanks to all who shared hints, tips and suggestions during this competitions. Here is a summary of the methodology used:\n\n</br>\n## Preprocessing - Tile Creation and Selection\n\nTried a lot of things at the start but finally settled with tiles created from image downsized by a factor of 6 to balance signal loss and manageable number of tiles. Created and selected tiles by:\n\n- Creating (224,224,3) shape tiles and applying simple threshold on image array unique values and their counts - for removing mostly monochrome and background tiles. \n- Reduced tile number further by running tiles through EfficientNetB0 (pretrained, ImageNet weights), scoring them and choosing top 50-60% tiles for each WSI (created another problem in tile number per image, addressed towards the end).\n- Added augmented data for LAA class by creating extra tiles simply by padding image and shifting tile location in image by tile_size/2 in both directions. \n\n</br>\n## MIL - Pseudo Labelled Tiles based Feature Extractor Training, Feature Aggregation and Classifier\n\nFollowed the tile feature embedding and then aggregation at slide level for further classification method. There are two main steps:\n\n### 1. Feature Extractor Training\n[This notebook](https://www.kaggle.com/code/icemantd/feature-cluster-tile-psuedo-label-train-mayo) has detail implementation of this part. The main points are:\n- Trained EfficientnetB0 starting with ImageNet weights, first epoch used all tiles with slide level labels.\n- Added random hue augmentation to tiles (details towards the end).\n- Defined a feature extractor which is the whole network upto one layer (shape (1,1280)) before the last FC classification layer.\n- Used tile features from extractor to apply cluster (kmeans) distance based pseudo labels to tiles for next iteration.\n\nThe logic used here is to attempt to provide the feature extractor the ability to not only give high output for the two classes but also a low output for out of class or non-critical tiles. Hence I used 'Other' label tiles with target [0.0] and LAA and CE with [1,0] and [0,1] respectively, with sigmoid activation and BCE loss. Every training epoch, features are created, clustering performed, and tiles scored on minimum distance from the negative class - i.e. top 20% highest min distance from LAA+Other feature clusters for CE tiles are labelled [0,1] and bottom 10% lowest min distance CE tiles are labelled [0,0] - and vice-versa for LAA tile pseudo labelling. The hope was this discriminates between the two classes, as well as between critical and non-critical tiles. The main clustering idea is inspired by [this](https://arxiv.org/pdf/2206.08861.pdf) paper.\n\n- Used random sampling of CE class to balance classes and used a simple metric for training evaluation at slide level - maxpooling AUC scores, i.e. taking AUC scores for maximum predictions of both classes by any tiles in the slide. 4-fold CV AUC scores of 0.6-0.65.\n\n### 2. Feature Aggregation, Classification and Inference\nObtained tile level features from feature extractor and adapted [this](https://arxiv.org/pdf/2011.08939.pdf) paper's methodology for feature aggregation at slide level, based on attention to tile with max output logit. \n\n- Built a simple random forest classifier using aggregated features at slide level and max predictions at slide level (any tile) for both classes as extra features.\n- Used only top 25 dark tiles from any image for training classifier and for inference.\n- Inference: Take top 25 dark tiles from image upto 2.2 Gb in size, create features from extractor and max logits from classifier head, aggregate and predict classifier probabilities ([notebook](https://www.kaggle.com/code/icemantd/new-predict-mil-logit-features-mayo-strip-ai/notebook))\n\n</br>\n## Some important observations\n- Found out that tile selection was very important, especially if you are using smaller tiles. Observed that at one point, larger images were giving better log_loss scores but there are more smaller images to infer (at least in training dataset). So adjused the top% of tiles selected from larger images, especially because LAA class had large images and some images dominated the tile numbers and resulted in overfitting. This definitely helped in controlling penalty on bad log_loss over large number of images. Example of log loss spread over training data vs number of tiles per slide image (number of tiles per image represents original number created and represents size of image and not number of tiles selected for training):\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8536316%2Fbe12d4336f614efd1bd81a02d54edb67%2Flog_loss_tilenum_imbalance.png?generation=1665076862015109&alt=media)\n\n- Observed from [this](https://arxiv.org/pdf/1902.06543.pdf) paper that adjusting hue of the image, even a little bit, resulted in better chance of predicting unseen strains better. So I implemented a random hue augmentation based on stretching/compressing the hue channel of each tile randomly by a factor betweem (0.85, 1.15). For example\n\n![![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8536316%2F155fc81cbdc429a4c96bba1cc9c5c0cf%2Fhue_aug.png?generation=1665077255141305&alt=media)\n\n- Due to the 2-step complexity, I could not come up with a robust end-to-end CV pipeline. Only spent time on feature extractor CV and evaluation using various data splits (60-40, 90-10) but ran into various issues with validation metric robustness. This part was the weakest for me and that's why I consider myself lucky to still land on something that kind of worked.\n\nHopefully some of you can point out where things could have been better or what was unecessary and even incorrectly implemented. Thanks.",
    "2048812": "This is a really practical way!",
    "1978839": "Congratulations! Good job 👍",
    "1978676": "Wow!! So much work goes into successful project. Congrats👍  and thanks for sharing.",
    "1975700": "Congratulations, looks like a huge amount of work here and so a well deserved placing. Lots of innovative ideas like the clustering, use of \"other\" images, and 2-step process.\n\nIIUC, you downsampled the whole slides by 6-times, and then pulled tiles of 224x224 after dropping monochrome tiles and tiles that scored in the bottom 50th percentile on a pre-trained EfficientNetB0. But then even after that, based on the first figure, some images still had 100+ tiles, some even with 500+ tiles. Is that correct? I'm surprised there were that many tiles available after the pre-processing.\n\nAlso, would you mind posting details of how you randomly adjusted the hue? Based on the notebook you posted I presume that with Tensorflow - is that correct? I used TF too but couldn't figure out how to implement a hue augmentation/adjustment. I suspect that hue adjustment was key because based on my testing I believe the hidden test set contained other center_ids than the n=11 in the training dataset, so it's probable the MSB staining methodology at new sites was different than the training set.\n\nAnyway, some really impressive work here\n\n"
  }
}