{
  "id": 357892,
  "title": "1st place solution",
  "url": "/competitions/mayo-clinic-strip-ai/discussion/357892",
  "author_name": "khyeh",
  "post_date": "2022-10-06T02:57:55.069000",
  "votes": 67,
  "comment_count": 39,
  "views": 0,
  "content": "<p>Thanks to Mayo Clinic and Kaggle for this competition. I enjoyed playing around with it. Also thanks to my teammate <a href=\"https://www.kaggle.com/evilpsycho42\" target=\"_blank\">@evilpsycho42</a> for all the support. I will try to illustrate my solution here even though I decided to stop working on it a month ago…</p>\n<h3>Data</h3>\n<ul>\n<li>Tiling and pick the top 16 darkest tiles</li>\n</ul>\n<h3>Single model</h3>\n<ul>\n<li>backbone: <code>swin_large_patch4_window12_384</code> + customized head</li>\n<li>classification head: customized: replace average pooling with attention pooling</li>\n</ul>\n<h3>Loss\\Metric</h3>\n<ul>\n<li><strong>implement loss\\metric following competition metric</strong></li>\n</ul>\n<h3>CV Strategy</h3>\n<ul>\n<li>5-fold Stratified Grouped KFold<ul>\n<li>Stratified by class and grouped by <code>patientid</code></li></ul></li>\n</ul>\n<h3>What Works</h3>\n<ul>\n<li><strong>attention pooling: 5-fold cv average improves from 0.69-&gt;0.66</strong></li>\n<li><strong>moco-v3 pretraining: 5-fold cv deviation improves from 0.30-&gt;0.15</strong></li>\n<li>ensemble: 5-fold cv average improves from 0.662-&gt;0.658 (very small improvement)</li>\n</ul>\n<h3>What Doesn't Work</h3>\n<ul>\n<li>More tiles</li>\n<li>Different preprocessing<ul>\n<li>top 16 highest pixel deviation instead of darkness</li>\n<li>stain normalization</li></ul></li>\n</ul>\n<h3>Final Solution</h3>\n<ul>\n<li>Ensemble of <code>swin_large_patch4_window12_384</code> and <code>coat_lite_medium</code></li>\n<li>Some luck 🙏</li>\n</ul>\n<h3>Some lessons learned from other competitions and applied here:</h3>\n<ul>\n<li>Since public LB contains very few samples, we have to do a correct CV, which is what we could do and rely on.</li>\n<li>I tuned models not only to improve the average CV score but also to reduce the deviation of the 5-fold CV, so the model could perform more stable in an unseen test set.</li>\n<li>Implement the right loss and metric: I found a lot of public kernels using logloss as train loss and even evaluation directly instead of implementing competition metric and using it as a loss function for modeling. <ul>\n<li>Logloss as evaluation behaves very differently from the competition metric.</li>\n<li>Logloss as training loss cause worse competition metric in my validation.</li></ul></li>\n<li>Try our best and hope for the best. <ul>\n<li>We try our best to do CV correctly, to improve CV score while reducing uncertainty.</li>\n<li>We hope for the best, since the dataset is not large enough to tell the difference between the last few digits, there will always be shakeup as expected. Keep an optimistic mind and move forward (Less painful for a bad shakeup)</li></ul></li>\n</ul>\n<p>My best CV with the highest mean and lowest deviations also achieve the best private scores.<br>\n<img src=\"https://imgur.com/gallery/2iIA7wO\" alt=\"final selection and best private lb\"></p>\n<p>Submission Kernel:<br>\n<a href=\"https://www.kaggle.com/code/khyeh0719/mayo-submission\" target=\"_blank\">https://www.kaggle.com/code/khyeh0719/mayo-submission</a> </p>",
  "messages": [
    {
      "id": 1974009,
      "postDate": "2022-10-06T02:57:55.070Z",
      "content": "<p>Thanks to Mayo Clinic and Kaggle for this competition. I enjoyed playing around with it. Also thanks to my teammate <a href=\"https://www.kaggle.com/evilpsycho42\" target=\"_blank\">@evilpsycho42</a> for all the support. I will try to illustrate my solution here even though I decided to stop working on it a month ago…</p>\n<h3>Data</h3>\n<ul>\n<li>Tiling and pick the top 16 darkest tiles</li>\n</ul>\n<h3>Single model</h3>\n<ul>\n<li>backbone: <code>swin_large_patch4_window12_384</code> + customized head</li>\n<li>classification head: customized: replace average pooling with attention pooling</li>\n</ul>\n<h3>Loss\\Metric</h3>\n<ul>\n<li><strong>implement loss\\metric following competition metric</strong></li>\n</ul>\n<h3>CV Strategy</h3>\n<ul>\n<li>5-fold Stratified Grouped KFold<ul>\n<li>Stratified by class and grouped by <code>patientid</code></li></ul></li>\n</ul>\n<h3>What Works</h3>\n<ul>\n<li><strong>attention pooling: 5-fold cv average improves from 0.69-&gt;0.66</strong></li>\n<li><strong>moco-v3 pretraining: 5-fold cv deviation improves from 0.30-&gt;0.15</strong></li>\n<li>ensemble: 5-fold cv average improves from 0.662-&gt;0.658 (very small improvement)</li>\n</ul>\n<h3>What Doesn't Work</h3>\n<ul>\n<li>More tiles</li>\n<li>Different preprocessing<ul>\n<li>top 16 highest pixel deviation instead of darkness</li>\n<li>stain normalization</li></ul></li>\n</ul>\n<h3>Final Solution</h3>\n<ul>\n<li>Ensemble of <code>swin_large_patch4_window12_384</code> and <code>coat_lite_medium</code></li>\n<li>Some luck 🙏</li>\n</ul>\n<h3>Some lessons learned from other competitions and applied here:</h3>\n<ul>\n<li>Since public LB contains very few samples, we have to do a correct CV, which is what we could do and rely on.</li>\n<li>I tuned models not only to improve the average CV score but also to reduce the deviation of the 5-fold CV, so the model could perform more stable in an unseen test set.</li>\n<li>Implement the right loss and metric: I found a lot of public kernels using logloss as train loss and even evaluation directly instead of implementing competition metric and using it as a loss function for modeling. <ul>\n<li>Logloss as evaluation behaves very differently from the competition metric.</li>\n<li>Logloss as training loss cause worse competition metric in my validation.</li></ul></li>\n<li>Try our best and hope for the best. <ul>\n<li>We try our best to do CV correctly, to improve CV score while reducing uncertainty.</li>\n<li>We hope for the best, since the dataset is not large enough to tell the difference between the last few digits, there will always be shakeup as expected. Keep an optimistic mind and move forward (Less painful for a bad shakeup)</li></ul></li>\n</ul>\n<p>My best CV with the highest mean and lowest deviations also achieve the best private scores.<br>\n<img src=\"https://imgur.com/gallery/2iIA7wO\" alt=\"final selection and best private lb\"></p>\n<p>Submission Kernel:<br>\n<a href=\"https://www.kaggle.com/code/khyeh0719/mayo-submission\" target=\"_blank\">https://www.kaggle.com/code/khyeh0719/mayo-submission</a> </p>",
      "rawMarkdown": "Thanks to Mayo Clinic and Kaggle for this competition. I enjoyed playing around with it. Also thanks to my teammate @evilpsycho42 for all the support. I will try to illustrate my solution here even though I decided to stop working on it a month ago...\n\n### Data\n- Tiling and pick the top 16 darkest tiles\n\n### Single model\n- backbone: ```swin_large_patch4_window12_384``` + customized head\n- classification head: customized: replace average pooling with attention pooling\n\n### Loss\\Metric\n- **implement loss\\metric following competition metric**\n\n### CV Strategy\n- 5-fold Stratified Grouped KFold\n    - Stratified by class and grouped by ```patientid```\n\n### What Works\n- **attention pooling: 5-fold cv average improves from 0.69->0.66**\n- **moco-v3 pretraining: 5-fold cv deviation improves from 0.30->0.15**\n- ensemble: 5-fold cv average improves from 0.662->0.658 (very small improvement)\n\n### What Doesn't Work\n- More tiles\n- Different preprocessing\n  - top 16 highest pixel deviation instead of darkness\n  - stain normalization\n\n### Final Solution\n- Ensemble of ```swin_large_patch4_window12_384``` and ```coat_lite_medium```\n- Some luck 🙏\n\n### Some lessons learned from other competitions and applied here:\n- Since public LB contains very few samples, we have to do a correct CV, which is what we could do and rely on.\n- I tuned models not only to improve the average CV score but also to reduce the deviation of the 5-fold CV, so the model could perform more stable in an unseen test set.\n- Implement the right loss and metric: I found a lot of public kernels using logloss as train loss and even evaluation directly instead of implementing competition metric and using it as a loss function for modeling. \n  - Logloss as evaluation behaves very differently from the competition metric.\n  - Logloss as training loss cause worse competition metric in my validation.\n- Try our best and hope for the best. \n  - We try our best to do CV correctly, to improve CV score while reducing uncertainty.\n  - We hope for the best, since the dataset is not large enough to tell the difference between the last few digits, there will always be shakeup as expected. Keep an optimistic mind and move forward (Less painful for a bad shakeup)\n\n\nMy best CV with the highest mean and lowest deviations also achieve the best private scores.\n![final selection and best private lb](https://imgur.com/gallery/2iIA7wO)\n\nSubmission Kernel:\nhttps://www.kaggle.com/code/khyeh0719/mayo-submission ",
      "votes": 67
    },
    {
      "id": 1981992,
      "postDate": "2022-10-11T07:14:06.970Z",
      "content": "<p>Congratulation for winning , great work , keep going :).</p>",
      "rawMarkdown": "Congratulation for winning , great work , keep going :).",
      "votes": 1
    },
    {
      "id": 1981745,
      "postDate": "2022-10-11T03:24:34.007Z",
      "content": "<p>Awesome work! Very impressive - initially I thought it's going to be close to random guess…</p>\n<p>One question: what is the input image size and magnification level? Did you see any improvements with different input size / magnification level?</p>",
      "rawMarkdown": "Awesome work! Very impressive - initially I thought it's going to be close to random guess...\n\nOne question: what is the input image size and magnification level? Did you see any improvements with different input size / magnification level?",
      "votes": 1,
      "replies": [
        {
          "id": 1982006,
          "postDate": "2022-10-11T07:24:47.697Z",
          "content": "<p><a href=\"https://www.kaggle.com/naotous\" target=\"_blank\">@naotous</a> <br>\nI used 384 as the image size since it is the image size required by the transformer-based backbone (SWIN large 384)<br>\nI also tried BEIT 512, which is another transformer-based backbone that takes 512 as input image size, however, I did not get better results with that.</p>\n<ul>\n<li>The BEIT model is larger and easier to get overfitting might be the reason for not giving me improvement.</li>\n<li>I did not try the hybrid SWIN transformer by adding extra convolutions between images and the SWIN backbone to support larger images and to verify your question.</li>\n</ul>",
          "rawMarkdown": "@naotous \nI used 384 as the image size since it is the image size required by the transformer-based backbone (SWIN large 384)\nI also tried BEIT 512, which is another transformer-based backbone that takes 512 as input image size, however, I did not get better results with that.\n- The BEIT model is larger and easier to get overfitting might be the reason for not giving me improvement.\n- I did not try the hybrid SWIN transformer by adding extra convolutions between images and the SWIN backbone to support larger images and to verify your question.",
          "votes": 1
        },
        {
          "id": 1983304,
          "postDate": "2022-10-12T00:47:15.060Z",
          "content": "<p>Thank you for your insights! Very interesting.<br>\nIn addition to potential overfitting issue, I think it's also the patch embedding kernel size is the key difference (Swin 4px vs BEiT 16px?).</p>",
          "rawMarkdown": "Thank you for your insights! Very interesting.\nIn addition to potential overfitting issue, I think it's also the patch embedding kernel size is the key difference (Swin 4px vs BEiT 16px?).",
          "votes": 1
        }
      ]
    },
    {
      "id": 1980237,
      "postDate": "2022-10-10T03:30:40.183Z",
      "content": "<p>Congrats and thank you for sharing your solution!</p>",
      "rawMarkdown": "Congrats and thank you for sharing your solution!",
      "votes": 1
    },
    {
      "id": 1980024,
      "postDate": "2022-10-10T00:22:11.480Z",
      "content": "<p>Congrats on winning and great solution!</p>",
      "rawMarkdown": "Congrats on winning and great solution!",
      "votes": 1
    },
    {
      "id": 1979333,
      "postDate": "2022-10-09T10:37:52.383Z",
      "content": "<p>Congratulations! </p>\n<p>I would like to ask about the hyperparameters when pre-learning swin-L with mocoV3 and points to note when using self-supervised learning.</p>",
      "rawMarkdown": "Congratulations! \n\nI would like to ask about the hyperparameters when pre-learning swin-L with mocoV3 and points to note when using self-supervised learning.",
      "votes": 1,
      "replies": [
        {
          "id": 1979397,
          "postDate": "2022-10-09T11:56:32.077Z",
          "content": "<p><a href=\"https://www.kaggle.com/abebe9849\" target=\"_blank\">@abebe9849</a> my hyperparameters: </p>\n<pre><code>lr = 0.0001 \nbatch-size = 4 \nepochs = 20\n</code></pre>\n<p>I modified the repo from <a href=\"https://github.com/CupidJay/MoCov3-pytorch\" target=\"_blank\">https://github.com/CupidJay/MoCov3-pytorch</a> to support <em>timm</em> created models (at least swin transformer)</p>",
          "rawMarkdown": "@abebe9849 my hyperparameters: \n```\nlr = 0.0001 \nbatch-size = 4 \nepochs = 20\n```\nI modified the repo from https://github.com/CupidJay/MoCov3-pytorch to support *timm* created models (at least swin transformer)",
          "votes": 1
        },
        {
          "id": 1980150,
          "postDate": "2022-10-10T02:33:27.363Z",
          "content": "<p>I was surprised that contrastive learning worked well with such a small batch size.<br>\nDid you shorten the epoch experimentally? Since I usually use DINO, I have never been able to acquire effective features in 20 epochs.</p>",
          "rawMarkdown": "I was surprised that contrastive learning worked well with such a small batch size.\nDid you shorten the epoch experimentally? Since I usually use DINO, I have never been able to acquire effective features in 20 epochs."
        },
        {
          "id": 1980209,
          "postDate": "2022-10-10T03:19:48.813Z",
          "content": "<p><a href=\"https://www.kaggle.com/abebe9849\" target=\"_blank\">@abebe9849</a> Yes, I tried 10\\20\\40 epochs, and 20 epochs of moco pretraining worked for me the best in reducing cv variation.</p>",
          "rawMarkdown": "@abebe9849 Yes, I tried 10\\20\\40 epochs, and 20 epochs of moco pretraining worked for me the best in reducing cv variation.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1978634,
      "postDate": "2022-10-08T20:42:48.580Z",
      "content": "<p>Great work congrats! <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a> </p>",
      "rawMarkdown": "Great work congrats! @khyeh0719 ",
      "votes": 1
    },
    {
      "id": 1978614,
      "postDate": "2022-10-08T20:06:42.903Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a> 🎉</p>",
      "rawMarkdown": "Congratulations @khyeh0719 🎉",
      "votes": 1
    },
    {
      "id": 1978442,
      "postDate": "2022-10-08T18:02:51.993Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a> </p>",
      "rawMarkdown": "Congratulations @khyeh0719 ",
      "votes": 1
    },
    {
      "id": 1977834,
      "postDate": "2022-10-08T09:57:49.960Z",
      "content": "<p>Congratulation <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a> </p>",
      "rawMarkdown": "Congratulation @khyeh0719 ",
      "votes": 1
    },
    {
      "id": 1977286,
      "postDate": "2022-10-07T23:21:48.197Z",
      "content": "<p>congratulations !! Although tell me what were your main challenges with the data!</p>",
      "rawMarkdown": "congratulations !! Although tell me what were your main challenges with the data!",
      "votes": 1,
      "replies": [
        {
          "id": 1978202,
          "postDate": "2022-10-08T15:41:55.567Z",
          "content": "<p>the challenge with the data is that I need to do all the preprocessing for the train da dataset on the Kaggle before moving into my rented server…</p>",
          "rawMarkdown": "the challenge with the data is that I need to do all the preprocessing for the train da dataset on the Kaggle before moving into my rented server..."
        }
      ]
    },
    {
      "id": 1975786,
      "postDate": "2022-10-07T02:54:50.543Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a> and <a href=\"https://www.kaggle.com/evilpsycho42\" target=\"_blank\">@evilpsycho42</a>.</p>",
      "rawMarkdown": "Congratulations @khyeh0719 and @evilpsycho42.",
      "votes": 1
    },
    {
      "id": 1975647,
      "postDate": "2022-10-06T23:36:04.023Z",
      "content": "<p>Thank you for sharing your solution and mostly for sharing your mayo-submission Kaggle Notebook.<br>\nCongratulations for your 1st Place Khyeh0719!</p>",
      "rawMarkdown": "Thank you for sharing your solution and mostly for sharing your mayo-submission Kaggle Notebook.\nCongratulations for your 1st Place Khyeh0719!",
      "votes": 1
    },
    {
      "id": 1975633,
      "postDate": "2022-10-06T23:21:05.073Z",
      "content": "<p>Wow! Congratulations <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a> Very well deserved!🎉</p>",
      "rawMarkdown": "Wow! Congratulations @khyeh0719 Very well deserved!🎉",
      "votes": 1
    },
    {
      "id": 1974918,
      "postDate": "2022-10-06T14:29:27.893Z",
      "content": "<p>Congrats! Great way to overcome low signals. </p>",
      "rawMarkdown": "Congrats! Great way to overcome low signals. ",
      "votes": 1
    },
    {
      "id": 1974761,
      "postDate": "2022-10-06T12:44:29.927Z",
      "content": "<p>Congrats!!! Great job</p>",
      "rawMarkdown": "Congrats!!! Great job",
      "votes": 1
    },
    {
      "id": 1974436,
      "postDate": "2022-10-06T08:49:27.130Z",
      "content": "<p>Congratz on first place ! <br>\nI could not make tile models competitive, looks like I should've pushed a bit more.</p>",
      "rawMarkdown": "Congratz on first place ! \nI could not make tile models competitive, looks like I should've pushed a bit more.",
      "votes": 1,
      "replies": [
        {
          "id": 1974458,
          "postDate": "2022-10-06T09:05:06.137Z",
          "content": "<p>You got good results in the end still 😁</p>",
          "rawMarkdown": "You got good results in the end still 😁",
          "votes": 1
        }
      ]
    },
    {
      "id": 1974237,
      "postDate": "2022-10-06T06:24:23.837Z",
      "content": "<p>congratulations and thanks for the accurate description of your experience. </p>",
      "rawMarkdown": "congratulations and thanks for the accurate description of your experience. ",
      "votes": 1
    },
    {
      "id": 1974100,
      "postDate": "2022-10-06T04:31:51.950Z",
      "content": "<p>Congrats! I think you really predicted that the test is slightly different than the training set. I also wanted to work on attention pooling but I didn't have enough time :/</p>",
      "rawMarkdown": "Congrats! I think you really predicted that the test is slightly different than the training set. I also wanted to work on attention pooling but I didn't have enough time :/",
      "votes": 1
    },
    {
      "id": 1974050,
      "postDate": "2022-10-06T03:51:25.387Z",
      "content": "<p>congratulations.</p>",
      "rawMarkdown": "congratulations.",
      "votes": 1
    },
    {
      "id": 1975766,
      "postDate": "2022-10-07T02:25:04.407Z",
      "content": "<p>Great work <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a> and <a href=\"https://www.kaggle.com/evilpsycho42\" target=\"_blank\">@evilpsycho42</a> ! Glad to see you both win a competition!! </p>",
      "rawMarkdown": "Great work @khyeh0719 and @evilpsycho42 ! Glad to see you both win a competition!! ",
      "votes": 2
    },
    {
      "id": 1975013,
      "postDate": "2022-10-06T15:22:52.257Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a>. Lot of good points here, thanks.</p>\n<p>Could you, as short as you like, tell how you used the tiles? As far as I can understand you trained with tile level but with slide level labels with an end-to-end model that includes tile feature aggregation to slide level before classification. Is that right?</p>",
      "rawMarkdown": "Congrats @khyeh0719. Lot of good points here, thanks.\n\nCould you, as short as you like, tell how you used the tiles? As far as I can understand you trained with tile level but with slide level labels with an end-to-end model that includes tile feature aggregation to slide level before classification. Is that right?",
      "votes": 2,
      "replies": [
        {
          "id": 1975135,
          "postDate": "2022-10-06T16:06:48Z",
          "content": "<p>I trained in slide level as well by using the attention module across all the tiles for the classification head.</p>",
          "rawMarkdown": "I trained in slide level as well by using the attention module across all the tiles for the classification head."
        },
        {
          "id": 1975231,
          "postDate": "2022-10-06T16:40:50.633Z",
          "content": "<p>Thanks for clarifying. You did this end to end, right? I mean tiles go as input and slide level prediction as output, with attention aggregation before classification head. I'm sorry, I will go through your code, its just that I cannot read pytorch code well right now :)</p>\n<p>Thanks for taking the time.</p>",
          "rawMarkdown": "Thanks for clarifying. You did this end to end, right? I mean tiles go as input and slide level prediction as output, with attention aggregation before classification head. I'm sorry, I will go through your code, its just that I cannot read pytorch code well right now :)\n\nThanks for taking the time.",
          "votes": 1
        },
        {
          "id": 1975655,
          "postDate": "2022-10-06T23:47:05.157Z",
          "content": "<p>yes, you are right, it is end to end. </p>",
          "rawMarkdown": "yes, you are right, it is end to end. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1974066,
      "postDate": "2022-10-06T04:09:19.423Z",
      "content": "<p>Great job, KH!   </p>",
      "rawMarkdown": "Great job, KH!   ",
      "votes": 2
    },
    {
      "id": 2249565,
      "postDate": "2023-05-07T21:18:59.913Z",
      "content": "<p>Congratulations and Thank you for sharing your experience!</p>",
      "rawMarkdown": "Congratulations and Thank you for sharing your experience!"
    },
    {
      "id": 2226421,
      "postDate": "2023-04-18T22:38:48.757Z",
      "content": "<p>Hi! I am curious what data you used for pretraining?</p>",
      "rawMarkdown": "Hi! I am curious what data you used for pretraining?"
    },
    {
      "id": 1987238,
      "postDate": "2022-10-14T15:34:00.647Z",
      "content": "<p>Congrats team.</p>",
      "rawMarkdown": "Congrats team."
    },
    {
      "id": 1977806,
      "postDate": "2022-10-08T09:27:57.257Z",
      "content": "<p>Just to have an idea of the practical applicability, did you evaluate  metrics like accuracy, false positives, negatives, etc (on your folds of course)? Thanks. </p>",
      "rawMarkdown": "Just to have an idea of the practical applicability, did you evaluate  metrics like accuracy, false positives, negatives, etc (on your folds of course)? Thanks. ",
      "replies": [
        {
          "id": 1978207,
          "postDate": "2022-10-08T15:46:03.663Z",
          "content": "<p>I did check non-weighted logloss and accuracy, but still only focused on the competition metric. For practical applicability,  my suggestion is to have different weights according to real-world data distribution. AUC curves might be also a more intuitive way to compare the trade-off between false positives and negatives for different models.</p>",
          "rawMarkdown": "I did check non-weighted logloss and accuracy, but still only focused on the competition metric. For practical applicability,  my suggestion is to have different weights according to real-world data distribution. AUC curves might be also a more intuitive way to compare the trade-off between false positives and negatives for different models."
        }
      ]
    },
    {
      "id": 1974103,
      "postDate": "2022-10-06T04:35:03.497Z",
      "content": "<p>What stain normalization techniques did you try?</p>",
      "rawMarkdown": "What stain normalization techniques did you try?",
      "replies": [
        {
          "id": 1974146,
          "postDate": "2022-10-06T05:01:43.237Z",
          "content": "<p><a href=\"https://github.com/EIDOSLAB/torchstain\" target=\"_blank\">https://github.com/EIDOSLAB/torchstain</a> <br>\nAssign a target stained image and convert other images to have similar stain color distribution.</p>",
          "rawMarkdown": "https://github.com/EIDOSLAB/torchstain \nAssign a target stained image and convert other images to have similar stain color distribution.",
          "votes": 4
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1981992,
      "author_name": "Dr. Alaa Temimy",
      "author_url": "",
      "post_date": "2022-10-11T07:14:06.970000",
      "content": "<p>Congratulation for winning , great work , keep going :).</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1981745,
      "author_name": "Naoto Usuyama",
      "author_url": "",
      "post_date": "2022-10-11T03:24:34.007000",
      "content": "<p>Awesome work! Very impressive - initially I thought it's going to be close to random guess…</p>\n<p>One question: what is the input image size and magnification level? Did you see any improvements with different input size / magnification level?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1982006,
          "author_name": "khyeh",
          "author_url": "",
          "post_date": "2022-10-11T07:24:47.697000",
          "content": "<p><a href=\"https://www.kaggle.com/naotous\" target=\"_blank\">@naotous</a> <br>\nI used 384 as the image size since it is the image size required by the transformer-based backbone (SWIN large 384)<br>\nI also tried BEIT 512, which is another transformer-based backbone that takes 512 as input image size, however, I did not get better results with that.</p>\n<ul>\n<li>The BEIT model is larger and easier to get overfitting might be the reason for not giving me improvement.</li>\n<li>I did not try the hybrid SWIN transformer by adding extra convolutions between images and the SWIN backbone to support larger images and to verify your question.</li>\n</ul>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1983304,
          "author_name": "Naoto Usuyama",
          "author_url": "",
          "post_date": "2022-10-12T00:47:15.060000",
          "content": "<p>Thank you for your insights! Very interesting.<br>\nIn addition to potential overfitting issue, I think it's also the patch embedding kernel size is the key difference (Swin 4px vs BEiT 16px?).</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1980237,
      "author_name": "Gangin Park",
      "author_url": "",
      "post_date": "2022-10-10T03:30:40.183000",
      "content": "<p>Congrats and thank you for sharing your solution!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1980024,
      "author_name": "Ahmed Ibrahim",
      "author_url": "",
      "post_date": "2022-10-10T00:22:11.480000",
      "content": "<p>Congrats on winning and great solution!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1979333,
      "author_name": "patriot",
      "author_url": "",
      "post_date": "2022-10-09T10:37:52.383000",
      "content": "<p>Congratulations! </p>\n<p>I would like to ask about the hyperparameters when pre-learning swin-L with mocoV3 and points to note when using self-supervised learning.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1979397,
          "author_name": "khyeh",
          "author_url": "",
          "post_date": "2022-10-09T11:56:32.077000",
          "content": "<p><a href=\"https://www.kaggle.com/abebe9849\" target=\"_blank\">@abebe9849</a> my hyperparameters: </p>\n<pre><code>lr = 0.0001 \nbatch-size = 4 \nepochs = 20\n</code></pre>\n<p>I modified the repo from <a href=\"https://github.com/CupidJay/MoCov3-pytorch\" target=\"_blank\">https://github.com/CupidJay/MoCov3-pytorch</a> to support <em>timm</em> created models (at least swin transformer)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1980150,
          "author_name": "patriot",
          "author_url": "",
          "post_date": "2022-10-10T02:33:27.363000",
          "content": "<p>I was surprised that contrastive learning worked well with such a small batch size.<br>\nDid you shorten the epoch experimentally? Since I usually use DINO, I have never been able to acquire effective features in 20 epochs.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1980209,
          "author_name": "khyeh",
          "author_url": "",
          "post_date": "2022-10-10T03:19:48.813000",
          "content": "<p><a href=\"https://www.kaggle.com/abebe9849\" target=\"_blank\">@abebe9849</a> Yes, I tried 10\\20\\40 epochs, and 20 epochs of moco pretraining worked for me the best in reducing cv variation.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1978634,
      "author_name": "Felicius Brando",
      "author_url": "",
      "post_date": "2022-10-08T20:42:48.580000",
      "content": "<p>Great work congrats! <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1978614,
      "author_name": "Firulyra",
      "author_url": "",
      "post_date": "2022-10-08T20:06:42.903000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a> 🎉</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1978442,
      "author_name": "Bocetufy",
      "author_url": "",
      "post_date": "2022-10-08T18:02:51.993000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1977834,
      "author_name": "Harsh Vats",
      "author_url": "",
      "post_date": "2022-10-08T09:57:49.960000",
      "content": "<p>Congratulation <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1977286,
      "author_name": "Muhammad Ammar Jamshed",
      "author_url": "",
      "post_date": "2022-10-07T23:21:48.197000",
      "content": "<p>congratulations !! Although tell me what were your main challenges with the data!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1978202,
          "author_name": "khyeh",
          "author_url": "",
          "post_date": "2022-10-08T15:41:55.567000",
          "content": "<p>the challenge with the data is that I need to do all the preprocessing for the train da dataset on the Kaggle before moving into my rented server…</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1975786,
      "author_name": "djagatiya",
      "author_url": "",
      "post_date": "2022-10-07T02:54:50.543000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a> and <a href=\"https://www.kaggle.com/evilpsycho42\" target=\"_blank\">@evilpsycho42</a>.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1975647,
      "author_name": "Marília Prata",
      "author_url": "",
      "post_date": "2022-10-06T23:36:04.023000",
      "content": "<p>Thank you for sharing your solution and mostly for sharing your mayo-submission Kaggle Notebook.<br>\nCongratulations for your 1st Place Khyeh0719!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1975633,
      "author_name": "Godsent Abode",
      "author_url": "",
      "post_date": "2022-10-06T23:21:05.073000",
      "content": "<p>Wow! Congratulations <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a> Very well deserved!🎉</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1974918,
      "author_name": "Francesco Uccelli",
      "author_url": "",
      "post_date": "2022-10-06T14:29:27.893000",
      "content": "<p>Congrats! Great way to overcome low signals. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1974761,
      "author_name": "Pablo Larrosa",
      "author_url": "",
      "post_date": "2022-10-06T12:44:29.927000",
      "content": "<p>Congrats!!! Great job</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1974436,
      "author_name": "Theo Viel",
      "author_url": "",
      "post_date": "2022-10-06T08:49:27.130000",
      "content": "<p>Congratz on first place ! <br>\nI could not make tile models competitive, looks like I should've pushed a bit more.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1974458,
          "author_name": "khyeh",
          "author_url": "",
          "post_date": "2022-10-06T09:05:06.137000",
          "content": "<p>You got good results in the end still 😁</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1974237,
      "author_name": "MITEL-UNIUD",
      "author_url": "",
      "post_date": "2022-10-06T06:24:23.837000",
      "content": "<p>congratulations and thanks for the accurate description of your experience. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1974100,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2022-10-06T04:31:51.950000",
      "content": "<p>Congrats! I think you really predicted that the test is slightly different than the training set. I also wanted to work on attention pooling but I didn't have enough time :/</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1974050,
      "author_name": "XXXXlf",
      "author_url": "",
      "post_date": "2022-10-06T03:51:25.387000",
      "content": "<p>congratulations.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1975766,
      "author_name": "Trushant Kalyanpur",
      "author_url": "",
      "post_date": "2022-10-07T02:25:04.407000",
      "content": "<p>Great work <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a> and <a href=\"https://www.kaggle.com/evilpsycho42\" target=\"_blank\">@evilpsycho42</a> ! Glad to see you both win a competition!! </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1975013,
      "author_name": "tdiceman",
      "author_url": "",
      "post_date": "2022-10-06T15:22:52.257000",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a>. Lot of good points here, thanks.</p>\n<p>Could you, as short as you like, tell how you used the tiles? As far as I can understand you trained with tile level but with slide level labels with an end-to-end model that includes tile feature aggregation to slide level before classification. Is that right?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1975135,
          "author_name": "khyeh",
          "author_url": "",
          "post_date": "2022-10-06T16:06:48",
          "content": "<p>I trained in slide level as well by using the attention module across all the tiles for the classification head.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1975231,
          "author_name": "tdiceman",
          "author_url": "",
          "post_date": "2022-10-06T16:40:50.633000",
          "content": "<p>Thanks for clarifying. You did this end to end, right? I mean tiles go as input and slide level prediction as output, with attention aggregation before classification head. I'm sorry, I will go through your code, its just that I cannot read pytorch code well right now :)</p>\n<p>Thanks for taking the time.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1975655,
          "author_name": "khyeh",
          "author_url": "",
          "post_date": "2022-10-06T23:47:05.157000",
          "content": "<p>yes, you are right, it is end to end. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1974066,
      "author_name": "KKY",
      "author_url": "",
      "post_date": "2022-10-06T04:09:19.423000",
      "content": "<p>Great job, KH!   </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2249565,
      "author_name": "Areej Malkawi",
      "author_url": "",
      "post_date": "2023-05-07T21:18:59.913000",
      "content": "<p>Congratulations and Thank you for sharing your experience!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2226421,
      "author_name": "Mara Pl",
      "author_url": "",
      "post_date": "2023-04-18T22:38:48.757000",
      "content": "<p>Hi! I am curious what data you used for pretraining?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1987238,
      "author_name": "Giri Kunche",
      "author_url": "",
      "post_date": "2022-10-14T15:34:00.647000",
      "content": "<p>Congrats team.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1977806,
      "author_name": "MITEL-UNIUD",
      "author_url": "",
      "post_date": "2022-10-08T09:27:57.257000",
      "content": "<p>Just to have an idea of the practical applicability, did you evaluate  metrics like accuracy, false positives, negatives, etc (on your folds of course)? Thanks. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1978207,
          "author_name": "khyeh",
          "author_url": "",
          "post_date": "2022-10-08T15:46:03.663000",
          "content": "<p>I did check non-weighted logloss and accuracy, but still only focused on the competition metric. For practical applicability,  my suggestion is to have different weights according to real-world data distribution. AUC curves might be also a more intuitive way to compare the trade-off between false positives and negatives for different models.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1974103,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-10-06T04:35:03.497000",
      "content": "<p>What stain normalization techniques did you try?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1974146,
          "author_name": "khyeh",
          "author_url": "",
          "post_date": "2022-10-06T05:01:43.237000",
          "content": "<p><a href=\"https://github.com/EIDOSLAB/torchstain\" target=\"_blank\">https://github.com/EIDOSLAB/torchstain</a> <br>\nAssign a target stained image and convert other images to have similar stain color distribution.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1974009": "Thanks to Mayo Clinic and Kaggle for this competition. I enjoyed playing around with it. Also thanks to my teammate @evilpsycho42 for all the support. I will try to illustrate my solution here even though I decided to stop working on it a month ago...\n\n### Data\n- Tiling and pick the top 16 darkest tiles\n\n### Single model\n- backbone: ```swin_large_patch4_window12_384``` + customized head\n- classification head: customized: replace average pooling with attention pooling\n\n### Loss\\Metric\n- **implement loss\\metric following competition metric**\n\n### CV Strategy\n- 5-fold Stratified Grouped KFold\n    - Stratified by class and grouped by ```patientid```\n\n### What Works\n- **attention pooling: 5-fold cv average improves from 0.69->0.66**\n- **moco-v3 pretraining: 5-fold cv deviation improves from 0.30->0.15**\n- ensemble: 5-fold cv average improves from 0.662->0.658 (very small improvement)\n\n### What Doesn't Work\n- More tiles\n- Different preprocessing\n  - top 16 highest pixel deviation instead of darkness\n  - stain normalization\n\n### Final Solution\n- Ensemble of ```swin_large_patch4_window12_384``` and ```coat_lite_medium```\n- Some luck 🙏\n\n### Some lessons learned from other competitions and applied here:\n- Since public LB contains very few samples, we have to do a correct CV, which is what we could do and rely on.\n- I tuned models not only to improve the average CV score but also to reduce the deviation of the 5-fold CV, so the model could perform more stable in an unseen test set.\n- Implement the right loss and metric: I found a lot of public kernels using logloss as train loss and even evaluation directly instead of implementing competition metric and using it as a loss function for modeling. \n  - Logloss as evaluation behaves very differently from the competition metric.\n  - Logloss as training loss cause worse competition metric in my validation.\n- Try our best and hope for the best. \n  - We try our best to do CV correctly, to improve CV score while reducing uncertainty.\n  - We hope for the best, since the dataset is not large enough to tell the difference between the last few digits, there will always be shakeup as expected. Keep an optimistic mind and move forward (Less painful for a bad shakeup)\n\n\nMy best CV with the highest mean and lowest deviations also achieve the best private scores.\n![final selection and best private lb](https://imgur.com/gallery/2iIA7wO)\n\nSubmission Kernel:\nhttps://www.kaggle.com/code/khyeh0719/mayo-submission ",
    "1981992": "Congratulation for winning , great work , keep going :).",
    "1981745": "Awesome work! Very impressive - initially I thought it's going to be close to random guess...\n\nOne question: what is the input image size and magnification level? Did you see any improvements with different input size / magnification level?",
    "1980237": "Congrats and thank you for sharing your solution!",
    "1980024": "Congrats on winning and great solution!",
    "1979333": "Congratulations! \n\nI would like to ask about the hyperparameters when pre-learning swin-L with mocoV3 and points to note when using self-supervised learning.",
    "1978634": "Great work congrats! @khyeh0719 ",
    "1978614": "Congratulations @khyeh0719 🎉",
    "1978442": "Congratulations @khyeh0719 ",
    "1977834": "Congratulation @khyeh0719 ",
    "1977286": "congratulations !! Although tell me what were your main challenges with the data!",
    "1975786": "Congratulations @khyeh0719 and @evilpsycho42.",
    "1975647": "Thank you for sharing your solution and mostly for sharing your mayo-submission Kaggle Notebook.\nCongratulations for your 1st Place Khyeh0719!",
    "1975633": "Wow! Congratulations @khyeh0719 Very well deserved!🎉",
    "1974918": "Congrats! Great way to overcome low signals. ",
    "1974761": "Congrats!!! Great job",
    "1974436": "Congratz on first place ! \nI could not make tile models competitive, looks like I should've pushed a bit more.",
    "1974237": "congratulations and thanks for the accurate description of your experience. ",
    "1974100": "Congrats! I think you really predicted that the test is slightly different than the training set. I also wanted to work on attention pooling but I didn't have enough time :/",
    "1974050": "congratulations.",
    "1975766": "Great work @khyeh0719 and @evilpsycho42 ! Glad to see you both win a competition!! ",
    "1975013": "Congrats @khyeh0719. Lot of good points here, thanks.\n\nCould you, as short as you like, tell how you used the tiles? As far as I can understand you trained with tile level but with slide level labels with an end-to-end model that includes tile feature aggregation to slide level before classification. Is that right?",
    "1974066": "Great job, KH!   ",
    "2249565": "Congratulations and Thank you for sharing your experience!",
    "2226421": "Hi! I am curious what data you used for pretraining?",
    "1987238": "Congrats team.",
    "1977806": "Just to have an idea of the practical applicability, did you evaluate  metrics like accuracy, false positives, negatives, etc (on your folds of course)? Thanks. ",
    "1974103": "What stain normalization techniques did you try?"
  }
}