{
  "id": 358089,
  "title": "2nd place solution - EfficientNetB0+augmentations+smart tiling",
  "url": "/competitions/mayo-clinic-strip-ai/discussion/358089",
  "author_name": "Ilia los",
  "post_date": "2022-10-06T15:10:02.482000",
  "votes": 25,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Thanks to Mayo Clinic for preparing this competition and to Kaggle for hosting it. And I'm happy that it was hosted as a code competition, because from my opinion it helps a lot to vanish the gap between kaggle competitions and real life business problems. Also I would like to congratulate guys from the first place because their solution seems much more complex and interesting than mine. <br>\nAt the same time sorry to see so low scores on the top comparing with random solution. I'm afraid that this competition might not help guys from Mayo to solve this very important problem and this is really sad.<br>\nLooking at the private leaderboard it seems like it was all about proper validation and total ignoring public scores (and a bit of luck of course). So let me quickly describe my solution.</p>\n<p><strong>Data preparation</strong><br>\nFor training I used all train images except 5 too-blurry. And also all other image with label 'Other'. I mixed them with LAA images in order to increase the amount of samples for that class, because in general I just wanted to distinguish CE from all other types.<br>\nProcess step by step:</p>\n<ol>\n<li>Resize images by factor 24. I also tried 8, 16, 32, 48, but 24 showed the best final metric;</li>\n<li>Use line-to-line pixelwise difference to calc for each 28x28 block does it have blood or background;</li>\n<li>For each 224x224 block with stride 28 check does it have blood or background;</li>\n<li>Deduplicate tiles: remove all blood tiles that intersect by more than half;</li>\n<li>Get random 20 if we have more tiles than that;</li>\n<li>No color normalization. I've tried several approaches, but all of them worked worse than strong color jitter augmentation.</li>\n</ol>\n<p><strong>Validation process</strong><br>\nFirst of all I quickly trained a small model on the same data to classify tiles by center_id. This models easily showed above-random performance (sorry, I lost exact metrics) and it was a signal for me that model potentially might overfits on similar clinics. Also <a href=\"https://www.kaggle.com/competitions/mayo-clinic-strip-ai/discussion/347095\" target=\"_blank\">here</a> competition host mentioned that data was collected from 18 institutes, but in train data we have only 11. So either test has 7 other clinics or it was about data from other folder somehow. Anyway I decided to use cross-validation based on clinics. In order to handle \"small-size\" clinics (small test size might leads to unstable metric and potential overfit) I joined them into groups. So each group has at least 90 examples.<br>\n<code>(11,), (4,), (7,), (1, 5,), (10, 3), (6, 2, 8, 9,)</code></p>\n<p><strong>Model training</strong><br>\nDuring my experiments I had a lot of issues with overfitting and stability. So I tried to choose as simple approach as possible.<br>\nI used efficientNetB0 pretrained on imageNet as a backbone and just one FC layer as a head. Also backbone was frozen during the whole training. Model predicts only one probability \"is it CE type\". So basically I trained only 513 parameters. <br>\nIn order to stratify training dataset I used upsampling and also I found it useful to upsample it 4 more times (probably, it only helps with bigger batch size). Also I multiplied test dataset 20 times: if some image has 20 tiles then all of them will be used in test once, if less then some tiles from that image could be presented more than once.<br>\nOther bullet points:</p>\n<ul>\n<li>Ensemble of 6 models as a final model (one model from each CV fold);</li>\n<li>Bigger batch helps: 64 was my final choice;</li>\n<li>Learning rate is important;</li>\n<li>Learning process was still a bit unstable so I trained 3 models for each CV fold and got the best one;</li>\n<li>BCELoss as a loss function;</li>\n<li>Competition metric with (0.5, 0.5) weights as a main and the only metric;</li>\n<li>If image has no tiles (for example, too small or too-blurry) then replace their final score with 0.5.</li>\n</ul>\n<p><strong>Data augmentations</strong><br>\nFor my approach it was the most important part so I decided to write about it in a separate block.<br>\nI used 3 types of augmentations during the training:</p>\n<ol>\n<li>Random flips - obvious;</li>\n<li>Random adjust sharpness - by some reason helped a lot, but blur doesn't at the same time;</li>\n<li>Strong color jitter (brightness=0.2, saturation=0.5, hue=0.5) - the most important one, because helped to fix color difference problem.<br>\nAnd also used 2 and 3 during the test. I used seeds everywhere in order to make the process reproducible. </li>\n</ol>\n<p><strong>Results</strong><br>\nCompetition target metric on my CV: 0.6373838583629989<br>\nPublic LB metric: 0.70822<br>\nPrivate LB metric: 0.66421</p>\n<p><strong>What didn't work for me</strong></p>\n<ul>\n<li>MIL;</li>\n<li>More complex models;</li>\n<li>Color normalization;</li>\n<li>Final predictions clipping;</li>\n<li>Grayscale images;</li>\n<li>Gram matrices from VGG19 (see style transfer model).</li>\n</ul>\n<p>Code is available on my GitHub now: <a href=\"https://github.com/IlyaLos/mayo-clinic-strip-ai\" target=\"_blank\">https://github.com/IlyaLos/mayo-clinic-strip-ai</a><br>\nNotebooks (just run one by one): <br>\n<a href=\"https://www.kaggle.com/code/ilyalos/2nd-place-solution-tiles-generation\" target=\"_blank\">https://www.kaggle.com/code/ilyalos/2nd-place-solution-tiles-generation</a><br>\n<a href=\"https://www.kaggle.com/code/ilyalos/2nd-place-solution-train\" target=\"_blank\">https://www.kaggle.com/code/ilyalos/2nd-place-solution-train</a><br>\n<a href=\"https://www.kaggle.com/code/ilyalos/2nd-place-solution-inference\" target=\"_blank\">https://www.kaggle.com/code/ilyalos/2nd-place-solution-inference</a></p>",
  "messages": [
    {
      "id": 1974992,
      "postDate": "2022-10-06T15:10:02.483Z",
      "content": "<p>Thanks to Mayo Clinic for preparing this competition and to Kaggle for hosting it. And I'm happy that it was hosted as a code competition, because from my opinion it helps a lot to vanish the gap between kaggle competitions and real life business problems. Also I would like to congratulate guys from the first place because their solution seems much more complex and interesting than mine. <br>\nAt the same time sorry to see so low scores on the top comparing with random solution. I'm afraid that this competition might not help guys from Mayo to solve this very important problem and this is really sad.<br>\nLooking at the private leaderboard it seems like it was all about proper validation and total ignoring public scores (and a bit of luck of course). So let me quickly describe my solution.</p>\n<p><strong>Data preparation</strong><br>\nFor training I used all train images except 5 too-blurry. And also all other image with label 'Other'. I mixed them with LAA images in order to increase the amount of samples for that class, because in general I just wanted to distinguish CE from all other types.<br>\nProcess step by step:</p>\n<ol>\n<li>Resize images by factor 24. I also tried 8, 16, 32, 48, but 24 showed the best final metric;</li>\n<li>Use line-to-line pixelwise difference to calc for each 28x28 block does it have blood or background;</li>\n<li>For each 224x224 block with stride 28 check does it have blood or background;</li>\n<li>Deduplicate tiles: remove all blood tiles that intersect by more than half;</li>\n<li>Get random 20 if we have more tiles than that;</li>\n<li>No color normalization. I've tried several approaches, but all of them worked worse than strong color jitter augmentation.</li>\n</ol>\n<p><strong>Validation process</strong><br>\nFirst of all I quickly trained a small model on the same data to classify tiles by center_id. This models easily showed above-random performance (sorry, I lost exact metrics) and it was a signal for me that model potentially might overfits on similar clinics. Also <a href=\"https://www.kaggle.com/competitions/mayo-clinic-strip-ai/discussion/347095\" target=\"_blank\">here</a> competition host mentioned that data was collected from 18 institutes, but in train data we have only 11. So either test has 7 other clinics or it was about data from other folder somehow. Anyway I decided to use cross-validation based on clinics. In order to handle \"small-size\" clinics (small test size might leads to unstable metric and potential overfit) I joined them into groups. So each group has at least 90 examples.<br>\n<code>(11,), (4,), (7,), (1, 5,), (10, 3), (6, 2, 8, 9,)</code></p>\n<p><strong>Model training</strong><br>\nDuring my experiments I had a lot of issues with overfitting and stability. So I tried to choose as simple approach as possible.<br>\nI used efficientNetB0 pretrained on imageNet as a backbone and just one FC layer as a head. Also backbone was frozen during the whole training. Model predicts only one probability \"is it CE type\". So basically I trained only 513 parameters. <br>\nIn order to stratify training dataset I used upsampling and also I found it useful to upsample it 4 more times (probably, it only helps with bigger batch size). Also I multiplied test dataset 20 times: if some image has 20 tiles then all of them will be used in test once, if less then some tiles from that image could be presented more than once.<br>\nOther bullet points:</p>\n<ul>\n<li>Ensemble of 6 models as a final model (one model from each CV fold);</li>\n<li>Bigger batch helps: 64 was my final choice;</li>\n<li>Learning rate is important;</li>\n<li>Learning process was still a bit unstable so I trained 3 models for each CV fold and got the best one;</li>\n<li>BCELoss as a loss function;</li>\n<li>Competition metric with (0.5, 0.5) weights as a main and the only metric;</li>\n<li>If image has no tiles (for example, too small or too-blurry) then replace their final score with 0.5.</li>\n</ul>\n<p><strong>Data augmentations</strong><br>\nFor my approach it was the most important part so I decided to write about it in a separate block.<br>\nI used 3 types of augmentations during the training:</p>\n<ol>\n<li>Random flips - obvious;</li>\n<li>Random adjust sharpness - by some reason helped a lot, but blur doesn't at the same time;</li>\n<li>Strong color jitter (brightness=0.2, saturation=0.5, hue=0.5) - the most important one, because helped to fix color difference problem.<br>\nAnd also used 2 and 3 during the test. I used seeds everywhere in order to make the process reproducible. </li>\n</ol>\n<p><strong>Results</strong><br>\nCompetition target metric on my CV: 0.6373838583629989<br>\nPublic LB metric: 0.70822<br>\nPrivate LB metric: 0.66421</p>\n<p><strong>What didn't work for me</strong></p>\n<ul>\n<li>MIL;</li>\n<li>More complex models;</li>\n<li>Color normalization;</li>\n<li>Final predictions clipping;</li>\n<li>Grayscale images;</li>\n<li>Gram matrices from VGG19 (see style transfer model).</li>\n</ul>\n<p>Code is available on my GitHub now: <a href=\"https://github.com/IlyaLos/mayo-clinic-strip-ai\" target=\"_blank\">https://github.com/IlyaLos/mayo-clinic-strip-ai</a><br>\nNotebooks (just run one by one): <br>\n<a href=\"https://www.kaggle.com/code/ilyalos/2nd-place-solution-tiles-generation\" target=\"_blank\">https://www.kaggle.com/code/ilyalos/2nd-place-solution-tiles-generation</a><br>\n<a href=\"https://www.kaggle.com/code/ilyalos/2nd-place-solution-train\" target=\"_blank\">https://www.kaggle.com/code/ilyalos/2nd-place-solution-train</a><br>\n<a href=\"https://www.kaggle.com/code/ilyalos/2nd-place-solution-inference\" target=\"_blank\">https://www.kaggle.com/code/ilyalos/2nd-place-solution-inference</a></p>",
      "rawMarkdown": "Thanks to Mayo Clinic for preparing this competition and to Kaggle for hosting it. And I'm happy that it was hosted as a code competition, because from my opinion it helps a lot to vanish the gap between kaggle competitions and real life business problems. Also I would like to congratulate guys from the first place because their solution seems much more complex and interesting than mine. \nAt the same time sorry to see so low scores on the top comparing with random solution. I'm afraid that this competition might not help guys from Mayo to solve this very important problem and this is really sad.\nLooking at the private leaderboard it seems like it was all about proper validation and total ignoring public scores (and a bit of luck of course). So let me quickly describe my solution.\n\n\n**Data preparation**\nFor training I used all train images except 5 too-blurry. And also all other image with label 'Other'. I mixed them with LAA images in order to increase the amount of samples for that class, because in general I just wanted to distinguish CE from all other types.\nProcess step by step:\n1. Resize images by factor 24. I also tried 8, 16, 32, 48, but 24 showed the best final metric;\n2. Use line-to-line pixelwise difference to calc for each 28x28 block does it have blood or background;\n3. For each 224x224 block with stride 28 check does it have blood or background;\n4. Deduplicate tiles: remove all blood tiles that intersect by more than half;\n5. Get random 20 if we have more tiles than that;\n6. No color normalization. I've tried several approaches, but all of them worked worse than strong color jitter augmentation.\n\n**Validation process**\nFirst of all I quickly trained a small model on the same data to classify tiles by center_id. This models easily showed above-random performance (sorry, I lost exact metrics) and it was a signal for me that model potentially might overfits on similar clinics. Also [here](https://www.kaggle.com/competitions/mayo-clinic-strip-ai/discussion/347095) competition host mentioned that data was collected from 18 institutes, but in train data we have only 11. So either test has 7 other clinics or it was about data from other folder somehow. Anyway I decided to use cross-validation based on clinics. In order to handle \"small-size\" clinics (small test size might leads to unstable metric and potential overfit) I joined them into groups. So each group has at least 90 examples.\n`(11,), (4,), (7,), (1, 5,), (10, 3), (6, 2, 8, 9,)`\n\n**Model training**\nDuring my experiments I had a lot of issues with overfitting and stability. So I tried to choose as simple approach as possible.\nI used efficientNetB0 pretrained on imageNet as a backbone and just one FC layer as a head. Also backbone was frozen during the whole training. Model predicts only one probability \"is it CE type\". So basically I trained only 513 parameters. \nIn order to stratify training dataset I used upsampling and also I found it useful to upsample it 4 more times (probably, it only helps with bigger batch size). Also I multiplied test dataset 20 times: if some image has 20 tiles then all of them will be used in test once, if less then some tiles from that image could be presented more than once.\nOther bullet points:\n- Ensemble of 6 models as a final model (one model from each CV fold);\n- Bigger batch helps: 64 was my final choice;\n- Learning rate is important;\n- Learning process was still a bit unstable so I trained 3 models for each CV fold and got the best one;\n- BCELoss as a loss function;\n- Competition metric with (0.5, 0.5) weights as a main and the only metric;\n- If image has no tiles (for example, too small or too-blurry) then replace their final score with 0.5.\n\n**Data augmentations**\nFor my approach it was the most important part so I decided to write about it in a separate block.\nI used 3 types of augmentations during the training:\n1. Random flips - obvious;\n2. Random adjust sharpness - by some reason helped a lot, but blur doesn't at the same time;\n3. Strong color jitter (brightness=0.2, saturation=0.5, hue=0.5) - the most important one, because helped to fix color difference problem.\nAnd also used 2 and 3 during the test. I used seeds everywhere in order to make the process reproducible. \n\n**Results**\nCompetition target metric on my CV: 0.6373838583629989\nPublic LB metric: 0.70822\nPrivate LB metric: 0.66421\n\n**What didn't work for me**\n- MIL;\n- More complex models;\n- Color normalization;\n- Final predictions clipping;\n- Grayscale images;\n- Gram matrices from VGG19 (see style transfer model).\n\nCode is available on my GitHub now: https://github.com/IlyaLos/mayo-clinic-strip-ai\nNotebooks (just run one by one): \nhttps://www.kaggle.com/code/ilyalos/2nd-place-solution-tiles-generation\nhttps://www.kaggle.com/code/ilyalos/2nd-place-solution-train\nhttps://www.kaggle.com/code/ilyalos/2nd-place-solution-inference",
      "votes": 25
    },
    {
      "id": 2006392,
      "postDate": "2022-10-27T15:06:20.343Z",
      "content": "<p>Hi Ilya Ios,</p>\n<p>Congratulations for your medal/2nd place.  <br>\nThanks for sharing what worked and what didn't work for you.  And mostly for sharing your GitHub code/link. </p>\n<p>Best regards, <br>\nMarília</p>",
      "rawMarkdown": "Hi Ilya Ios,\n\nCongratulations for your medal/2nd place.  \nThanks for sharing what worked and what didn't work for you.  And mostly for sharing your GitHub code/link. \n\nBest regards, \nMarília",
      "votes": 1,
      "replies": [
        {
          "id": 2012764,
          "postDate": "2022-11-01T12:03:15.707Z",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/mpwolke\" target=\"_blank\">@mpwolke</a> !<br>\nHope you found some of my ideas useful.</p>",
          "rawMarkdown": "Thank you @mpwolke !\nHope you found some of my ideas useful.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1977373,
      "postDate": "2022-10-08T01:56:32.370Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/ilyalos\" target=\"_blank\">@ilyalos</a> , congrats for you solo gold. Did you average the score of all the patches of each tissues to get the final probability?</p>",
      "rawMarkdown": "Hi @ilyalos , congrats for you solo gold. Did you average the score of all the patches of each tissues to get the final probability?",
      "votes": 1,
      "replies": [
        {
          "id": 1977578,
          "postDate": "2022-10-08T05:30:57.887Z",
          "content": "<p>thank you, <a href=\"https://www.kaggle.com/forcewithme\" target=\"_blank\">@forcewithme</a> <br>\nYes, I had 6-folds CV and used 6-models ensemble for the final prediction. <br>\nMy logic about it: each model was validated on its own set of centers, so they are different somehow and combining them should help to cover all cases and probably even increase the quality.<br>\nTraining process was unstable so I was not able to train one final model on full data without validation.</p>",
          "rawMarkdown": "thank you, @forcewithme \nYes, I had 6-folds CV and used 6-models ensemble for the final prediction. \nMy logic about it: each model was validated on its own set of centers, so they are different somehow and combining them should help to cover all cases and probably even increase the quality.\nTraining process was unstable so I was not able to train one final model on full data without validation.",
          "votes": 1
        },
        {
          "id": 1977621,
          "postDate": "2022-10-08T06:08:36.583Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/ilyalos\" target=\"_blank\">@ilyalos</a> . Thank you for your answer. I mean, you get 20 patches for each tissues(wsl), then how did you get the probability for each wsl? I guess you average the score of the 20 patches of each wsl?</p>",
          "rawMarkdown": "Hi @ilyalos . Thank you for your answer. I mean, you get 20 patches for each tissues(wsl), then how did you get the probability for each wsl? I guess you average the score of the 20 patches of each wsl?",
          "votes": 1
        },
        {
          "id": 1977914,
          "postDate": "2022-10-08T10:36:57.893Z",
          "content": "<p>oh, sorry, I missed your point. Yes, you're right, I just got an average of all probabilities.</p>",
          "rawMarkdown": "oh, sorry, I missed your point. Yes, you're right, I just got an average of all probabilities."
        }
      ]
    },
    {
      "id": 2011219,
      "postDate": "2022-10-31T12:32:38.593Z",
      "content": "<p>Thanks for your great sharing. I do not know what is Final predictions clipping. Does this refer to ensemble modeling? </p>",
      "rawMarkdown": "Thanks for your great sharing. I do not know what is Final predictions clipping. Does this refer to ensemble modeling? ",
      "replies": [
        {
          "id": 2012762,
          "postDate": "2022-11-01T12:01:49.463Z",
          "content": "<p>In general if you have a model that predicts some probability then you might want to clip that probability in some ranges. For example, if probability is from range [0.4, 0.6] you might want to clip it to 0.5, because your models isn't sure about such samples and it might be better to go with the middle probability for them. And the same for higher/lower probabilities: [0.8, 1.0] -&gt; 1.0 and [0.0, 0.2] -&gt; 0.0 (all ranges here are just examples).<br>\nIn that competition it might works because of really small differences in scores from random solution (with all 0.5). So for some teams clipping uncertain predictions to 0.5 gave boost in score. </p>",
          "rawMarkdown": "In general if you have a model that predicts some probability then you might want to clip that probability in some ranges. For example, if probability is from range [0.4, 0.6] you might want to clip it to 0.5, because your models isn't sure about such samples and it might be better to go with the middle probability for them. And the same for higher/lower probabilities: [0.8, 1.0] -> 1.0 and [0.0, 0.2] -> 0.0 (all ranges here are just examples).\nIn that competition it might works because of really small differences in scores from random solution (with all 0.5). So for some teams clipping uncertain predictions to 0.5 gave boost in score. ",
          "votes": 2
        },
        {
          "id": 2013006,
          "postDate": "2022-11-01T15:21:23.650Z",
          "content": "<p>Thanks a lot! I understand it.</p>",
          "rawMarkdown": "Thanks a lot! I understand it.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1975066,
      "postDate": "2022-10-06T15:41:39.913Z",
      "content": "<p>Great work <a href=\"https://www.kaggle.com/ilyalos\" target=\"_blank\">@ilyalos</a>, congrats for your position. I also used EffnetB0, but I am amazed that frozen backbone worked so well here. Thanks for the details.</p>",
      "rawMarkdown": "Great work @ilyalos, congrats for your position. I also used EffnetB0, but I am amazed that frozen backbone worked so well here. Thanks for the details.",
      "replies": [
        {
          "id": 1975091,
          "postDate": "2022-10-06T15:48:43.070Z",
          "content": "<p>Thanks, <a href=\"https://www.kaggle.com/icemantd\" target=\"_blank\">@icemantd</a> and congrats with your 7th place too :)<br>\nyes, without freezing (and even with unfreezing only few last layers) learning process was really unstable in my case. I even tried to unfreeze backbone after some warmup for head only, but it showed worse performance as well.</p>",
          "rawMarkdown": "Thanks, @icemantd and congrats with your 7th place too :)\nyes, without freezing (and even with unfreezing only few last layers) learning process was really unstable in my case. I even tried to unfreeze backbone after some warmup for head only, but it showed worse performance as well.",
          "votes": 2
        },
        {
          "id": 1975103,
          "postDate": "2022-10-06T15:53:13.373Z",
          "content": "<p>Really good to know these details. And thanks :)</p>",
          "rawMarkdown": "Really good to know these details. And thanks :)",
          "votes": 1
        },
        {
          "id": 1976614,
          "postDate": "2022-10-07T13:08:01.857Z",
          "content": "<p>Very weird that imagenet features extracted from images of sharks, frogs, spiders, etc can yield good performance on whole slide images of biological tissue. I don't understand how this works. </p>\n<p>And thanks <a href=\"https://www.kaggle.com/ilyalos\" target=\"_blank\">@ilyalos</a> for sharing details and your cleanly written code!</p>",
          "rawMarkdown": "Very weird that imagenet features extracted from images of sharks, frogs, spiders, etc can yield good performance on whole slide images of biological tissue. I don't understand how this works. \n\nAnd thanks @ilyalos for sharing details and your cleanly written code!",
          "votes": 1
        },
        {
          "id": 1976650,
          "postDate": "2022-10-07T13:35:49.193Z",
          "content": "<p>In general these features might provide a lot of useful information for various number of different tasks. But you're right and blood clot images don't look like the best area anyway. Of course, with some biological pretrained model it suppose to work much better, but I was not able to find any such model. I tried several options, but they all showed even worse results than imageNet. And at the same time I was not able to train or even fine-tune the backbone myself without dramatic overfitting, because we don't have enough signal.</p>",
          "rawMarkdown": "In general these features might provide a lot of useful information for various number of different tasks. But you're right and blood clot images don't look like the best area anyway. Of course, with some biological pretrained model it suppose to work much better, but I was not able to find any such model. I tried several options, but they all showed even worse results than imageNet. And at the same time I was not able to train or even fine-tune the backbone myself without dramatic overfitting, because we don't have enough signal.",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2006392,
      "author_name": "Marília Prata",
      "author_url": "",
      "post_date": "2022-10-27T15:06:20.343000",
      "content": "<p>Hi Ilya Ios,</p>\n<p>Congratulations for your medal/2nd place.  <br>\nThanks for sharing what worked and what didn't work for you.  And mostly for sharing your GitHub code/link. </p>\n<p>Best regards, <br>\nMarília</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2012764,
          "author_name": "Ilia los",
          "author_url": "",
          "post_date": "2022-11-01T12:03:15.707000",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/mpwolke\" target=\"_blank\">@mpwolke</a> !<br>\nHope you found some of my ideas useful.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1977373,
      "author_name": "ForcewithMe",
      "author_url": "",
      "post_date": "2022-10-08T01:56:32.370000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/ilyalos\" target=\"_blank\">@ilyalos</a> , congrats for you solo gold. Did you average the score of all the patches of each tissues to get the final probability?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1977578,
          "author_name": "Ilia los",
          "author_url": "",
          "post_date": "2022-10-08T05:30:57.887000",
          "content": "<p>thank you, <a href=\"https://www.kaggle.com/forcewithme\" target=\"_blank\">@forcewithme</a> <br>\nYes, I had 6-folds CV and used 6-models ensemble for the final prediction. <br>\nMy logic about it: each model was validated on its own set of centers, so they are different somehow and combining them should help to cover all cases and probably even increase the quality.<br>\nTraining process was unstable so I was not able to train one final model on full data without validation.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1977621,
          "author_name": "ForcewithMe",
          "author_url": "",
          "post_date": "2022-10-08T06:08:36.583000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/ilyalos\" target=\"_blank\">@ilyalos</a> . Thank you for your answer. I mean, you get 20 patches for each tissues(wsl), then how did you get the probability for each wsl? I guess you average the score of the 20 patches of each wsl?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1977914,
          "author_name": "Ilia los",
          "author_url": "",
          "post_date": "2022-10-08T10:36:57.893000",
          "content": "<p>oh, sorry, I missed your point. Yes, you're right, I just got an average of all probabilities.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2011219,
      "author_name": "gray98",
      "author_url": "",
      "post_date": "2022-10-31T12:32:38.593000",
      "content": "<p>Thanks for your great sharing. I do not know what is Final predictions clipping. Does this refer to ensemble modeling? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2012762,
          "author_name": "Ilia los",
          "author_url": "",
          "post_date": "2022-11-01T12:01:49.463000",
          "content": "<p>In general if you have a model that predicts some probability then you might want to clip that probability in some ranges. For example, if probability is from range [0.4, 0.6] you might want to clip it to 0.5, because your models isn't sure about such samples and it might be better to go with the middle probability for them. And the same for higher/lower probabilities: [0.8, 1.0] -&gt; 1.0 and [0.0, 0.2] -&gt; 0.0 (all ranges here are just examples).<br>\nIn that competition it might works because of really small differences in scores from random solution (with all 0.5). So for some teams clipping uncertain predictions to 0.5 gave boost in score. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2013006,
          "author_name": "gray98",
          "author_url": "",
          "post_date": "2022-11-01T15:21:23.650000",
          "content": "<p>Thanks a lot! I understand it.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1975066,
      "author_name": "tdiceman",
      "author_url": "",
      "post_date": "2022-10-06T15:41:39.913000",
      "content": "<p>Great work <a href=\"https://www.kaggle.com/ilyalos\" target=\"_blank\">@ilyalos</a>, congrats for your position. I also used EffnetB0, but I am amazed that frozen backbone worked so well here. Thanks for the details.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1975091,
          "author_name": "Ilia los",
          "author_url": "",
          "post_date": "2022-10-06T15:48:43.070000",
          "content": "<p>Thanks, <a href=\"https://www.kaggle.com/icemantd\" target=\"_blank\">@icemantd</a> and congrats with your 7th place too :)<br>\nyes, without freezing (and even with unfreezing only few last layers) learning process was really unstable in my case. I even tried to unfreeze backbone after some warmup for head only, but it showed worse performance as well.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1975103,
          "author_name": "tdiceman",
          "author_url": "",
          "post_date": "2022-10-06T15:53:13.373000",
          "content": "<p>Really good to know these details. And thanks :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1976614,
          "author_name": "cosmosaa",
          "author_url": "",
          "post_date": "2022-10-07T13:08:01.857000",
          "content": "<p>Very weird that imagenet features extracted from images of sharks, frogs, spiders, etc can yield good performance on whole slide images of biological tissue. I don't understand how this works. </p>\n<p>And thanks <a href=\"https://www.kaggle.com/ilyalos\" target=\"_blank\">@ilyalos</a> for sharing details and your cleanly written code!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1976650,
          "author_name": "Ilia los",
          "author_url": "",
          "post_date": "2022-10-07T13:35:49.193000",
          "content": "<p>In general these features might provide a lot of useful information for various number of different tasks. But you're right and blood clot images don't look like the best area anyway. Of course, with some biological pretrained model it suppose to work much better, but I was not able to find any such model. I tried several options, but they all showed even worse results than imageNet. And at the same time I was not able to train or even fine-tune the backbone myself without dramatic overfitting, because we don't have enough signal.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1974992": "Thanks to Mayo Clinic for preparing this competition and to Kaggle for hosting it. And I'm happy that it was hosted as a code competition, because from my opinion it helps a lot to vanish the gap between kaggle competitions and real life business problems. Also I would like to congratulate guys from the first place because their solution seems much more complex and interesting than mine. \nAt the same time sorry to see so low scores on the top comparing with random solution. I'm afraid that this competition might not help guys from Mayo to solve this very important problem and this is really sad.\nLooking at the private leaderboard it seems like it was all about proper validation and total ignoring public scores (and a bit of luck of course). So let me quickly describe my solution.\n\n\n**Data preparation**\nFor training I used all train images except 5 too-blurry. And also all other image with label 'Other'. I mixed them with LAA images in order to increase the amount of samples for that class, because in general I just wanted to distinguish CE from all other types.\nProcess step by step:\n1. Resize images by factor 24. I also tried 8, 16, 32, 48, but 24 showed the best final metric;\n2. Use line-to-line pixelwise difference to calc for each 28x28 block does it have blood or background;\n3. For each 224x224 block with stride 28 check does it have blood or background;\n4. Deduplicate tiles: remove all blood tiles that intersect by more than half;\n5. Get random 20 if we have more tiles than that;\n6. No color normalization. I've tried several approaches, but all of them worked worse than strong color jitter augmentation.\n\n**Validation process**\nFirst of all I quickly trained a small model on the same data to classify tiles by center_id. This models easily showed above-random performance (sorry, I lost exact metrics) and it was a signal for me that model potentially might overfits on similar clinics. Also [here](https://www.kaggle.com/competitions/mayo-clinic-strip-ai/discussion/347095) competition host mentioned that data was collected from 18 institutes, but in train data we have only 11. So either test has 7 other clinics or it was about data from other folder somehow. Anyway I decided to use cross-validation based on clinics. In order to handle \"small-size\" clinics (small test size might leads to unstable metric and potential overfit) I joined them into groups. So each group has at least 90 examples.\n`(11,), (4,), (7,), (1, 5,), (10, 3), (6, 2, 8, 9,)`\n\n**Model training**\nDuring my experiments I had a lot of issues with overfitting and stability. So I tried to choose as simple approach as possible.\nI used efficientNetB0 pretrained on imageNet as a backbone and just one FC layer as a head. Also backbone was frozen during the whole training. Model predicts only one probability \"is it CE type\". So basically I trained only 513 parameters. \nIn order to stratify training dataset I used upsampling and also I found it useful to upsample it 4 more times (probably, it only helps with bigger batch size). Also I multiplied test dataset 20 times: if some image has 20 tiles then all of them will be used in test once, if less then some tiles from that image could be presented more than once.\nOther bullet points:\n- Ensemble of 6 models as a final model (one model from each CV fold);\n- Bigger batch helps: 64 was my final choice;\n- Learning rate is important;\n- Learning process was still a bit unstable so I trained 3 models for each CV fold and got the best one;\n- BCELoss as a loss function;\n- Competition metric with (0.5, 0.5) weights as a main and the only metric;\n- If image has no tiles (for example, too small or too-blurry) then replace their final score with 0.5.\n\n**Data augmentations**\nFor my approach it was the most important part so I decided to write about it in a separate block.\nI used 3 types of augmentations during the training:\n1. Random flips - obvious;\n2. Random adjust sharpness - by some reason helped a lot, but blur doesn't at the same time;\n3. Strong color jitter (brightness=0.2, saturation=0.5, hue=0.5) - the most important one, because helped to fix color difference problem.\nAnd also used 2 and 3 during the test. I used seeds everywhere in order to make the process reproducible. \n\n**Results**\nCompetition target metric on my CV: 0.6373838583629989\nPublic LB metric: 0.70822\nPrivate LB metric: 0.66421\n\n**What didn't work for me**\n- MIL;\n- More complex models;\n- Color normalization;\n- Final predictions clipping;\n- Grayscale images;\n- Gram matrices from VGG19 (see style transfer model).\n\nCode is available on my GitHub now: https://github.com/IlyaLos/mayo-clinic-strip-ai\nNotebooks (just run one by one): \nhttps://www.kaggle.com/code/ilyalos/2nd-place-solution-tiles-generation\nhttps://www.kaggle.com/code/ilyalos/2nd-place-solution-train\nhttps://www.kaggle.com/code/ilyalos/2nd-place-solution-inference",
    "2006392": "Hi Ilya Ios,\n\nCongratulations for your medal/2nd place.  \nThanks for sharing what worked and what didn't work for you.  And mostly for sharing your GitHub code/link. \n\nBest regards, \nMarília",
    "1977373": "Hi @ilyalos , congrats for you solo gold. Did you average the score of all the patches of each tissues to get the final probability?",
    "2011219": "Thanks for your great sharing. I do not know what is Final predictions clipping. Does this refer to ensemble modeling? ",
    "1975066": "Great work @ilyalos, congrats for your position. I also used EffnetB0, but I am amazed that frozen backbone worked so well here. Thanks for the details."
  }
}