{
  "id": 169114,
  "title": "8Th place solution",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/169114",
  "author_name": "Arnaud Roussel",
  "post_date": "2020-07-23T01:03:53.382000",
  "votes": 37,
  "comment_count": 19,
  "views": 0,
  "content": "<p>Despite some evidence of randomness I'd like to share the ideas we used:\n* 10 model ensemble based on local CV and decent LB.\n* Different predictions between models (regression, bins and ordinal regression)\n* Some of it used bags of tiles and other stack the tiles in squares\n* Efficient nets (I trained only b0 but partner had a few b4\n* My models were trained in two steps. First a model with an attention layer is made (was shared by me in some thread). Then this attention layer and model is reused to predict weights for tiles. Then a model is retrained with a lower number of tiles (9 or 16). I have some 9 tiles models that were both fast and were going at 0.90CV+. On top of it it allowed us to inspect a larger amount of tiles during inference (128 tiles) and just select the best 9 or 16.\n* My partner used in his model a NetVlad layer which maybe hell talk about in this thread.\n* Ensembling with mean + round was better than majority voting for LB (and is what we used) but actually our best solution uses majority voting (which we didnt select).\n* We also built a CV without duplicates and without \"suspicious slides\".</p>\n\n<p>In the last weeks after making the team and how obvious it seemed the shake up would be big I started to mistrust LB and try to bring diversity to the ensemble. As long as a model was at LB &gt; 0.88 that was good enough if the CV was among the best ones.</p>\n\n<p>Learned a lot during this competition. Thanks to organizer.</p>\n\n<p>Note: We also have a solution at 0.936 that we didn't select :(</p>",
  "messages": [
    {
      "id": 940466,
      "postDate": "2020-07-23T01:03:53.383Z",
      "content": "<p>Despite some evidence of randomness I'd like to share the ideas we used:\n* 10 model ensemble based on local CV and decent LB.\n* Different predictions between models (regression, bins and ordinal regression)\n* Some of it used bags of tiles and other stack the tiles in squares\n* Efficient nets (I trained only b0 but partner had a few b4\n* My models were trained in two steps. First a model with an attention layer is made (was shared by me in some thread). Then this attention layer and model is reused to predict weights for tiles. Then a model is retrained with a lower number of tiles (9 or 16). I have some 9 tiles models that were both fast and were going at 0.90CV+. On top of it it allowed us to inspect a larger amount of tiles during inference (128 tiles) and just select the best 9 or 16.\n* My partner used in his model a NetVlad layer which maybe hell talk about in this thread.\n* Ensembling with mean + round was better than majority voting for LB (and is what we used) but actually our best solution uses majority voting (which we didnt select).\n* We also built a CV without duplicates and without \"suspicious slides\".</p>\n\n<p>In the last weeks after making the team and how obvious it seemed the shake up would be big I started to mistrust LB and try to bring diversity to the ensemble. As long as a model was at LB &gt; 0.88 that was good enough if the CV was among the best ones.</p>\n\n<p>Learned a lot during this competition. Thanks to organizer.</p>\n\n<p>Note: We also have a solution at 0.936 that we didn't select :(</p>",
      "rawMarkdown": "Despite some evidence of randomness I'd like to share the ideas we used:\n* 10 model ensemble based on local CV and decent LB.\n* Different predictions between models (regression, bins and ordinal regression)\n* Some of it used bags of tiles and other stack the tiles in squares\n* Efficient nets (I trained only b0 but partner had a few b4\n* My models were trained in two steps. First a model with an attention layer is made (was shared by me in some thread). Then this attention layer and model is reused to predict weights for tiles. Then a model is retrained with a lower number of tiles (9 or 16). I have some 9 tiles models that were both fast and were going at 0.90CV+. On top of it it allowed us to inspect a larger amount of tiles during inference (128 tiles) and just select the best 9 or 16.\n* My partner used in his model a NetVlad layer which maybe hell talk about in this thread.\n* Ensembling with mean + round was better than majority voting for LB (and is what we used) but actually our best solution uses majority voting (which we didnt select).\n* We also built a CV without duplicates and without \"suspicious slides\".\n\nIn the last weeks after making the team and how obvious it seemed the shake up would be big I started to mistrust LB and try to bring diversity to the ensemble. As long as a model was at LB &gt; 0.88 that was good enough if the CV was among the best ones.\n\nLearned a lot during this competition. Thanks to organizer.\n\nNote: We also have a solution at 0.936 that we didn't select :(",
      "votes": 37
    },
    {
      "id": 941332,
      "postDate": "2020-07-23T07:08:19.053Z",
      "content": "<p>Thanks every one who shared their ideas in this competition !\nHere is my code:  <a href=\"https://github.com/ChienYiChi/kaggle-panda-challenge\">solution code for PANDA </a>\nFeel free to ask me about any questions.\nWe are lucky to get into gold zone in the private LB, and this is my second time to survive in big shake up. Last time is the <a href=\"https://www.kaggle.com/c/severstal-steel-defect-detection\">Severstal: Steel Defect Detection</a>. </p>\n\n<h3>Best model of Private LeaderBoard</h3>\n\n<strong>1st stage  model</strong>: efficientnet-b0 with attention layer, regression type, 64 tiles for train\n\n<strong>2nd stage model</strong>:  efficientnet-b0 with NetVLAD layer, regression type, top 16 tiles for train\n\n<strong>Score</strong>:\n\n<ul>\n<li><p>private lb:0.920 ; public lb: 0.881 ; local kappa: 0.8952 ; karolinska kappa: 0.8976 ; radboud: 0.8704</p>\n\n<strong>Detail</strong>:</li>\n<li><p>thank my teammate Arnaud , I did the same thing as he did,  train a Efficientnet-b0 with Attention layer to generate weight for each tiles . The number of tiles used to train is 64, and then I select the top 16 tiles as the training tiles for the 2nd stage model.</p></li>\n<li>train the 2nd stage to predict the final score. I use efficientnet-b0 with a NetVLAD layer to generate the final feature for final prediction. The NetVLAD outperform adaptive max pooling and average pooling layer. </li>\n</ul>\n\n<h3>Best model of Public LeaderBoard</h3>\n\n<strong>Model</strong>: efficientnet-b4 with NetVLAD layer, regression type, 36 tiles selected by Blue Ratio Selection.\n\n<strong>Score</strong>:\n\n<ul>\n<li>private lb: 0.913 ; public lb: 0.902 ; local kappa: 0.8826 ; karolinska kappa: 0.9047 ; radboud: 0.8335 \n<strong>Detail</strong>:</li>\n</ul>\n\n<p>I use Blue Ratio Selection for selecting the top 36 tiles  and to accelerate the selection function, firstly I select 128 tiles based on sum of tile pixels and then I select 16 tiles based on them with blue ratio value. </p>\n\n<p>In inference, all model use 8 TTA same as public kernel of lafoss. </p>\n\n<p>As Arnaud described above , our final submission ensemble 10 model.</p>",
      "rawMarkdown": "Thanks every one who shared their ideas in this competition !\nHere is my code:  [solution code for PANDA ](https://github.com/ChienYiChi/kaggle-panda-challenge)\nFeel free to ask me about any questions.\nWe are lucky to get into gold zone in the private LB, and this is my second time to survive in big shake up. Last time is the [Severstal: Steel Defect Detection](https://www.kaggle.com/c/severstal-steel-defect-detection). \n\n### Best model of Private LeaderBoard\n#### **1st stage  model**: efficientnet-b0 with attention layer, regression type, 64 tiles for train \n#### **2nd stage model**:  efficientnet-b0 with NetVLAD layer, regression type, top 16 tiles for train\n#### **Score**: \n- private lb:0.920 ; public lb: 0.881 ; local kappa: 0.8952 ; karolinska kappa: 0.8976 ; radboud: 0.8704\n#### **Detail**:\n- thank my teammate Arnaud , I did the same thing as he did,  train a Efficientnet-b0 with Attention layer to generate weight for each tiles . The number of tiles used to train is 64, and then I select the top 16 tiles as the training tiles for the 2nd stage model.\n- train the 2nd stage to predict the final score. I use efficientnet-b0 with a NetVLAD layer to generate the final feature for final prediction. The NetVLAD outperform adaptive max pooling and average pooling layer. \n\n### Best model of Public LeaderBoard \n#### **Model**: efficientnet-b4 with NetVLAD layer, regression type, 36 tiles selected by Blue Ratio Selection.\n#### **Score**:\n- private lb: 0.913 ; public lb: 0.902 ; local kappa: 0.8826 ; karolinska kappa: 0.9047 ; radboud: 0.8335 \n#### **Detail**:\nI use Blue Ratio Selection for selecting the top 36 tiles  and to accelerate the selection function, firstly I select 128 tiles based on sum of tile pixels and then I select 16 tiles based on them with blue ratio value. \n\nIn inference, all model use 8 TTA same as public kernel of lafoss. \n\nAs Arnaud described above , our final submission ensemble 10 model.\n",
      "votes": 5
    },
    {
      "id": 940475,
      "postDate": "2020-07-23T01:13:53.627Z",
      "content": "<p>Additionally my code is available on github (in a messy state :)) push/pulling was the easiest way for me when using GCP to train bigger models.\nhere: <a href=\"https://github.com/arroqc/pandacancer_kaggle\">https://github.com/arroqc/pandacancer_kaggle</a></p>",
      "rawMarkdown": "Additionally my code is available on github (in a messy state :)) push/pulling was the easiest way for me when using GCP to train bigger models.\nhere: https://github.com/arroqc/pandacancer_kaggle",
      "votes": 4,
      "replies": [
        {
          "id": 940481,
          "postDate": "2020-07-23T01:22:01.143Z",
          "content": "<p>Congratulations <a href=\"/arroqc\">@arroqc</a>  . Your code is still private :) </p>",
          "rawMarkdown": "Congratulations @arroqc  . Your code is still private :) ",
          "votes": 1
        },
        {
          "id": 940482,
          "postDate": "2020-07-23T01:23:37.907Z",
          "content": "<p>Sorry about that, fixed.</p>",
          "rawMarkdown": "Sorry about that, fixed.",
          "votes": 1
        },
        {
          "id": 940563,
          "postDate": "2020-07-23T02:57:46.637Z",
          "content": "<p>congrats </p>",
          "rawMarkdown": "congrats "
        }
      ]
    },
    {
      "id": 941756,
      "postDate": "2020-07-23T11:52:21.803Z",
      "content": "<p>Thanks guys for sharing your githubs, I'll work on those right away. Congratulations on your medal ;)</p>",
      "rawMarkdown": "Thanks guys for sharing your githubs, I'll work on those right away. Congratulations on your medal ;)",
      "votes": 1
    },
    {
      "id": 940506,
      "postDate": "2020-07-23T01:54:00.997Z",
      "content": "<p><a href=\"/arroqc\">@arroqc</a> Congratulations for you gold, looking forward to work with you in future.</p>",
      "rawMarkdown": "@arroqc Congratulations for you gold, looking forward to work with you in future.",
      "votes": 1
    },
    {
      "id": 948158,
      "postDate": "2020-07-27T17:43:24.413Z",
      "content": "<p><a href=\"/arroqc\">@arroqc</a> \nCongrats &amp; thanks for sharing your smart solution!</p>\n\n<p>&gt; We also built a CV without duplicates and without \"suspicious slides\".</p>\n\n<p>I'm curious about this part. How did you achieve this?</p>",
      "rawMarkdown": "@arroqc \nCongrats &amp; thanks for sharing your smart solution!\n\n&gt; We also built a CV without duplicates and without \"suspicious slides\".\n\nI'm curious about this part. How did you achieve this?",
      "replies": [
        {
          "id": 948412,
          "postDate": "2020-07-27T23:04:48.223Z",
          "content": "<p><a href=\"/yukkyo\">@yukkyo</a> We just used some kernel someone shared that uses imagehash. For duplicates we used a lower threshold than suggested and got quite a few false positives but we don't think it matters much since all it did was to force those slides to be in the same splits. For suspicious someone released a list at some point of weird cases like blank slides, no mask, pen marks etc. I assumed these were likely to be not representative of the test set so I removed them from the validation splits.</p>\n\n<p>That being said it is not clear this strategy really paid off. It just seemed like a good idea at the time amd we didnt really have a way of testing because we believed the LB to be a poor predictor of the private.</p>",
          "rawMarkdown": "@yukkyo We just used some kernel someone shared that uses imagehash. For duplicates we used a lower threshold than suggested and got quite a few false positives but we don't think it matters much since all it did was to force those slides to be in the same splits. For suspicious someone released a list at some point of weird cases like blank slides, no mask, pen marks etc. I assumed these were likely to be not representative of the test set so I removed them from the validation splits.\n\nThat being said it is not clear this strategy really paid off. It just seemed like a good idea at the time amd we didnt really have a way of testing because we believed the LB to be a poor predictor of the private.",
          "votes": 1
        },
        {
          "id": 948468,
          "postDate": "2020-07-28T01:33:43.327Z",
          "content": "<p><a href=\"/arroqc\">@arroqc</a> Thanks for sharing the details! It's very helpful.</p>",
          "rawMarkdown": "@arroqc Thanks for sharing the details! It's very helpful."
        }
      ]
    },
    {
      "id": 945866,
      "postDate": "2020-07-26T07:30:31.527Z",
      "content": "<p>Congratulation with 8th place and thanks for sharing  your ideas!</p>",
      "rawMarkdown": "Congratulation with 8th place and thanks for sharing  your ideas!"
    },
    {
      "id": 945830,
      "postDate": "2020-07-26T06:54:01.607Z",
      "content": "<p>Hi <a href=\"/arroqc\">@arroqc</a>,</p>\n\n<p>Your attention approach is interesting. Do I use batch size of 1 (1 bag of n_tiles from same slide) to train the attention model?</p>",
      "rawMarkdown": "Hi @arroqc,\n\nYour attention approach is interesting. Do I use batch size of 1 (1 bag of n_tiles from same slide) to train the attention model?",
      "replies": [
        {
          "id": 945936,
          "postDate": "2020-07-26T08:22:49.243Z",
          "content": "<p>Batch size should be greater than 4 </p>",
          "rawMarkdown": "Batch size should be greater than 4 ",
          "votes": 1
        },
        {
          "id": 946870,
          "postDate": "2020-07-26T23:04:56.970Z",
          "content": "<p>No, use as large a batch as you can while training at least 32 tiles per case. I also add a lot of randomness (random tile selection weighted by \"darkness\". augmentated datasets etc.) to help the attention train to find the best tiles.</p>",
          "rawMarkdown": "No, use as large a batch as you can while training at least 32 tiles per case. I also add a lot of randomness (random tile selection weighted by \"darkness\". augmentated datasets etc.) to help the attention train to find the best tiles.",
          "votes": 1
        }
      ]
    },
    {
      "id": 944105,
      "postDate": "2020-07-24T20:45:47.657Z",
      "content": "<p>Congratulations and thank you for sharing your code.</p>",
      "rawMarkdown": "Congratulations and thank you for sharing your code."
    },
    {
      "id": 943171,
      "postDate": "2020-07-24T07:22:32.630Z",
      "content": "<p>Congrats <a href=\"/arroqc\">@arroqc</a> and your team, you contributed a lot by sharing your ideas during the 3 past months</p>",
      "rawMarkdown": "Congrats @arroqc and your team, you contributed a lot by sharing your ideas during the 3 past months"
    },
    {
      "id": 940623,
      "postDate": "2020-07-23T03:57:22.947Z",
      "content": "<p>Congrats <a href=\"/arroqc\">@arroqc</a> </p>",
      "rawMarkdown": "Congrats @arroqc "
    },
    {
      "id": 940565,
      "postDate": "2020-07-23T03:01:04.717Z",
      "content": "<p>Congrats on 8th place and gold medal <a href=\"/arroqc\">@arroqc</a> and thanks for sharing details solution.</p>",
      "rawMarkdown": "Congrats on 8th place and gold medal @arroqc and thanks for sharing details solution."
    },
    {
      "id": 947913,
      "postDate": "2020-07-27T14:53:34.680Z",
      "content": "<p>Cool...Thank you for sharing 👍 </p>",
      "rawMarkdown": "Cool...Thank you for sharing 👍 "
    }
  ],
  "comments": [
    {
      "id": 941332,
      "author_name": "ChienYiChi",
      "author_url": "",
      "post_date": "2020-07-23T07:08:19.053000",
      "content": "<p>Thanks every one who shared their ideas in this competition !\nHere is my code:  <a href=\"https://github.com/ChienYiChi/kaggle-panda-challenge\">solution code for PANDA </a>\nFeel free to ask me about any questions.\nWe are lucky to get into gold zone in the private LB, and this is my second time to survive in big shake up. Last time is the <a href=\"https://www.kaggle.com/c/severstal-steel-defect-detection\">Severstal: Steel Defect Detection</a>. </p>\n\n<h3>Best model of Private LeaderBoard</h3>\n\n<strong>1st stage  model</strong>: efficientnet-b0 with attention layer, regression type, 64 tiles for train\n\n<strong>2nd stage model</strong>:  efficientnet-b0 with NetVLAD layer, regression type, top 16 tiles for train\n\n<strong>Score</strong>:\n\n<ul>\n<li><p>private lb:0.920 ; public lb: 0.881 ; local kappa: 0.8952 ; karolinska kappa: 0.8976 ; radboud: 0.8704</p>\n\n<strong>Detail</strong>:</li>\n<li><p>thank my teammate Arnaud , I did the same thing as he did,  train a Efficientnet-b0 with Attention layer to generate weight for each tiles . The number of tiles used to train is 64, and then I select the top 16 tiles as the training tiles for the 2nd stage model.</p></li>\n<li>train the 2nd stage to predict the final score. I use efficientnet-b0 with a NetVLAD layer to generate the final feature for final prediction. The NetVLAD outperform adaptive max pooling and average pooling layer. </li>\n</ul>\n\n<h3>Best model of Public LeaderBoard</h3>\n\n<strong>Model</strong>: efficientnet-b4 with NetVLAD layer, regression type, 36 tiles selected by Blue Ratio Selection.\n\n<strong>Score</strong>:\n\n<ul>\n<li>private lb: 0.913 ; public lb: 0.902 ; local kappa: 0.8826 ; karolinska kappa: 0.9047 ; radboud: 0.8335 \n<strong>Detail</strong>:</li>\n</ul>\n\n<p>I use Blue Ratio Selection for selecting the top 36 tiles  and to accelerate the selection function, firstly I select 128 tiles based on sum of tile pixels and then I select 16 tiles based on them with blue ratio value. </p>\n\n<p>In inference, all model use 8 TTA same as public kernel of lafoss. </p>\n\n<p>As Arnaud described above , our final submission ensemble 10 model.</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 940475,
      "author_name": "Arnaud Roussel",
      "author_url": "",
      "post_date": "2020-07-23T01:13:53.627000",
      "content": "<p>Additionally my code is available on github (in a messy state :)) push/pulling was the easiest way for me when using GCP to train bigger models.\nhere: <a href=\"https://github.com/arroqc/pandacancer_kaggle\">https://github.com/arroqc/pandacancer_kaggle</a></p>",
      "votes": 4,
      "replies": [
        {
          "id": 940481,
          "author_name": "Nirjhar Roy",
          "author_url": "",
          "post_date": "2020-07-23T01:22:01.143000",
          "content": "<p>Congratulations <a href=\"/arroqc\">@arroqc</a>  . Your code is still private :) </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 940482,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2020-07-23T01:23:37.907000",
          "content": "<p>Sorry about that, fixed.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 940563,
          "author_name": "yuvaramsingh",
          "author_url": "",
          "post_date": "2020-07-23T02:57:46.637000",
          "content": "<p>congrats </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 941756,
      "author_name": "Benjamin Dubreu",
      "author_url": "",
      "post_date": "2020-07-23T11:52:21.803000",
      "content": "<p>Thanks guys for sharing your githubs, I'll work on those right away. Congratulations on your medal ;)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 940506,
      "author_name": "Dracarys",
      "author_url": "",
      "post_date": "2020-07-23T01:54:00.997000",
      "content": "<p><a href=\"/arroqc\">@arroqc</a> Congratulations for you gold, looking forward to work with you in future.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 948158,
      "author_name": "fam_taro",
      "author_url": "",
      "post_date": "2020-07-27T17:43:24.413000",
      "content": "<p><a href=\"/arroqc\">@arroqc</a> \nCongrats &amp; thanks for sharing your smart solution!</p>\n\n<p>&gt; We also built a CV without duplicates and without \"suspicious slides\".</p>\n\n<p>I'm curious about this part. How did you achieve this?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 948412,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2020-07-27T23:04:48.223000",
          "content": "<p><a href=\"/yukkyo\">@yukkyo</a> We just used some kernel someone shared that uses imagehash. For duplicates we used a lower threshold than suggested and got quite a few false positives but we don't think it matters much since all it did was to force those slides to be in the same splits. For suspicious someone released a list at some point of weird cases like blank slides, no mask, pen marks etc. I assumed these were likely to be not representative of the test set so I removed them from the validation splits.</p>\n\n<p>That being said it is not clear this strategy really paid off. It just seemed like a good idea at the time amd we didnt really have a way of testing because we believed the LB to be a poor predictor of the private.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 948468,
          "author_name": "fam_taro",
          "author_url": "",
          "post_date": "2020-07-28T01:33:43.327000",
          "content": "<p><a href=\"/arroqc\">@arroqc</a> Thanks for sharing the details! It's very helpful.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 945866,
      "author_name": "Artem Shibaev",
      "author_url": "",
      "post_date": "2020-07-26T07:30:31.527000",
      "content": "<p>Congratulation with 8th place and thanks for sharing  your ideas!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 945830,
      "author_name": "Yousef Rabi",
      "author_url": "",
      "post_date": "2020-07-26T06:54:01.607000",
      "content": "<p>Hi <a href=\"/arroqc\">@arroqc</a>,</p>\n\n<p>Your attention approach is interesting. Do I use batch size of 1 (1 bag of n_tiles from same slide) to train the attention model?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 945936,
          "author_name": "ChienYiChi",
          "author_url": "",
          "post_date": "2020-07-26T08:22:49.243000",
          "content": "<p>Batch size should be greater than 4 </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 946870,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2020-07-26T23:04:56.970000",
          "content": "<p>No, use as large a batch as you can while training at least 32 tiles per case. I also add a lot of randomness (random tile selection weighted by \"darkness\". augmentated datasets etc.) to help the attention train to find the best tiles.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 944105,
      "author_name": "Yousef Rabi",
      "author_url": "",
      "post_date": "2020-07-24T20:45:47.657000",
      "content": "<p>Congratulations and thank you for sharing your code.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 943171,
      "author_name": "Alex",
      "author_url": "",
      "post_date": "2020-07-24T07:22:32.630000",
      "content": "<p>Congrats <a href=\"/arroqc\">@arroqc</a> and your team, you contributed a lot by sharing your ideas during the 3 past months</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 940623,
      "author_name": "Kurian Benoy",
      "author_url": "",
      "post_date": "2020-07-23T03:57:22.947000",
      "content": "<p>Congrats <a href=\"/arroqc\">@arroqc</a> </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 940565,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2020-07-23T03:01:04.717000",
      "content": "<p>Congrats on 8th place and gold medal <a href=\"/arroqc\">@arroqc</a> and thanks for sharing details solution.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 947913,
      "author_name": "Jaseem C K",
      "author_url": "",
      "post_date": "2020-07-27T14:53:34.680000",
      "content": "<p>Cool...Thank you for sharing 👍 </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "940466": "Despite some evidence of randomness I'd like to share the ideas we used:\n* 10 model ensemble based on local CV and decent LB.\n* Different predictions between models (regression, bins and ordinal regression)\n* Some of it used bags of tiles and other stack the tiles in squares\n* Efficient nets (I trained only b0 but partner had a few b4\n* My models were trained in two steps. First a model with an attention layer is made (was shared by me in some thread). Then this attention layer and model is reused to predict weights for tiles. Then a model is retrained with a lower number of tiles (9 or 16). I have some 9 tiles models that were both fast and were going at 0.90CV+. On top of it it allowed us to inspect a larger amount of tiles during inference (128 tiles) and just select the best 9 or 16.\n* My partner used in his model a NetVlad layer which maybe hell talk about in this thread.\n* Ensembling with mean + round was better than majority voting for LB (and is what we used) but actually our best solution uses majority voting (which we didnt select).\n* We also built a CV without duplicates and without \"suspicious slides\".\n\nIn the last weeks after making the team and how obvious it seemed the shake up would be big I started to mistrust LB and try to bring diversity to the ensemble. As long as a model was at LB &gt; 0.88 that was good enough if the CV was among the best ones.\n\nLearned a lot during this competition. Thanks to organizer.\n\nNote: We also have a solution at 0.936 that we didn't select :(",
    "941332": "Thanks every one who shared their ideas in this competition !\nHere is my code:  [solution code for PANDA ](https://github.com/ChienYiChi/kaggle-panda-challenge)\nFeel free to ask me about any questions.\nWe are lucky to get into gold zone in the private LB, and this is my second time to survive in big shake up. Last time is the [Severstal: Steel Defect Detection](https://www.kaggle.com/c/severstal-steel-defect-detection). \n\n### Best model of Private LeaderBoard\n#### **1st stage  model**: efficientnet-b0 with attention layer, regression type, 64 tiles for train \n#### **2nd stage model**:  efficientnet-b0 with NetVLAD layer, regression type, top 16 tiles for train\n#### **Score**: \n- private lb:0.920 ; public lb: 0.881 ; local kappa: 0.8952 ; karolinska kappa: 0.8976 ; radboud: 0.8704\n#### **Detail**:\n- thank my teammate Arnaud , I did the same thing as he did,  train a Efficientnet-b0 with Attention layer to generate weight for each tiles . The number of tiles used to train is 64, and then I select the top 16 tiles as the training tiles for the 2nd stage model.\n- train the 2nd stage to predict the final score. I use efficientnet-b0 with a NetVLAD layer to generate the final feature for final prediction. The NetVLAD outperform adaptive max pooling and average pooling layer. \n\n### Best model of Public LeaderBoard \n#### **Model**: efficientnet-b4 with NetVLAD layer, regression type, 36 tiles selected by Blue Ratio Selection.\n#### **Score**:\n- private lb: 0.913 ; public lb: 0.902 ; local kappa: 0.8826 ; karolinska kappa: 0.9047 ; radboud: 0.8335 \n#### **Detail**:\nI use Blue Ratio Selection for selecting the top 36 tiles  and to accelerate the selection function, firstly I select 128 tiles based on sum of tile pixels and then I select 16 tiles based on them with blue ratio value. \n\nIn inference, all model use 8 TTA same as public kernel of lafoss. \n\nAs Arnaud described above , our final submission ensemble 10 model.\n",
    "940475": "Additionally my code is available on github (in a messy state :)) push/pulling was the easiest way for me when using GCP to train bigger models.\nhere: https://github.com/arroqc/pandacancer_kaggle",
    "941756": "Thanks guys for sharing your githubs, I'll work on those right away. Congratulations on your medal ;)",
    "940506": "@arroqc Congratulations for you gold, looking forward to work with you in future.",
    "948158": "@arroqc \nCongrats &amp; thanks for sharing your smart solution!\n\n&gt; We also built a CV without duplicates and without \"suspicious slides\".\n\nI'm curious about this part. How did you achieve this?",
    "945866": "Congratulation with 8th place and thanks for sharing  your ideas!",
    "945830": "Hi @arroqc,\n\nYour attention approach is interesting. Do I use batch size of 1 (1 bag of n_tiles from same slide) to train the attention model?",
    "944105": "Congratulations and thank you for sharing your code.",
    "943171": "Congrats @arroqc and your team, you contributed a lot by sharing your ideas during the 3 past months",
    "940623": "Congrats @arroqc ",
    "940565": "Congrats on 8th place and gold medal @arroqc and thanks for sharing details solution.",
    "947913": "Cool...Thank you for sharing 👍 "
  }
}