{
  "id": 169111,
  "title": "65th solution",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/169111",
  "author_name": "Shiyuan Zeng",
  "post_date": "2020-07-23T00:33:21.770000",
  "votes": 6,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Thanks for organizer! Congratulations to winners! I create this topic just want to sharing my solution(get 0.920 private lb). I know my solution will not so great as top solutions, but here I want to record something and hope you can give some advice.😄 </p>\n\n<ul>\n<li><p>Image Pre-process</p>\n\n<ul><li>Filter polluted tiles.(precompute and save the useful(pass filtering) tile ids of each image)\n<ul><li>change each tile from rgb to hsv, and observe some polluted tiles you will find the thresholds to filter polluted tiles :D</li>\n<li>how check each tile(already hsv)? Actually if you use thresholds above to fiter on tile-level, you will waste many useful tiles which have tissue information about half of tile-size. So I random sample 20 squares(length of side = tile-size//10), if the the number of squares, not be filted by the thresholds above, can reach 20xpass-ratio, I will treat this tile is useful and will also use it.</li>\n<li>If the number of tiles from a iamge is &lt; 36(I choose 36 tiles, size is 224x224). I use np.random.choice to supplement until reaching 36.</li></ul></li>\n<li>Patch Image(game changer 1)\n<ul><li>Like what <a href=\"https://www.kaggle.com/iafoss/panda-16x128x128-tiles\">PANDA 16x128x128 tiles</a> did. Thanks <a href=\"/iafoss\">@iafoss</a> !!!</li></ul></li></ul></li>\n<li><p>Models</p>\n\n<ul><li>efn-b0x3(I will call them b0-1, b0-2, b0-3 below) + efn-b4</li></ul></li>\n<li>Optimizer\n<ul><li>AdamW(default parameters)</li></ul></li>\n<li>Scheduler\n<ul><li>OneCycle(epochs=30, steps-per-epoch=int(np.ceil(len(train_dl)/acc-grad-step))\n<ul><li>about pct-start, I try 0.1, 1/30, 2/30, and 0.1 works best.</li></ul></li></ul></li>\n<li>Loss(it's the game changer 2)\n<ul><li>BCEWithLogitsLoss(thanks <a href=\"/haqishen\">@haqishen</a> !!!), details in [Train EfficientNet-B0 w/ 36 tiles_256 [LB0.87]] (<a href=\"https://www.kaggle.com/haqishen/train-efficientnet-b0-w-36-tiles-256-lb0-87\"></a><a href=\"https://www.kaggle.com/haqishen/train-efficientnet-b0-w-36-tiles-256-lb0-87\">https://www.kaggle.com/haqishen/train-efficientnet-b0-w-36-tiles-256-lb0-87</a>)</li></ul></li>\n<li>Train Details\n<ul><li>efn-b0(use tiles)\n<ul><li>batch-size=4, use gradient accumulation(also 4), so the batch-size can be treated as 4x4(not actually 4x4 of course)</li></ul></li>\n<li>efn-b4(use full-image)\n<ul><li>batch-size=16, resize lv1 tiff to 448x448</li></ul></li>\n<li>apex amp in kaggle is too hard to be installed well. So I use gradscaler from torcu.cuda.amp(seems torch version should &gt;= 1.5, I forgot the specific version).</li></ul></li>\n<li>Ensemble\n<ul><li>0.35 x (b0-1 + b0-2) + 0.25 x b0-3 + 0.05 x b4</li></ul></li>\n<li>Models Details\n<ul><li>b0-1(local score: 0.885) and b0-2(local score: 0.873) are the 2 highest local score folds of my 5-stratified-folds, and if mix other folds(local score all &lt; 0.87) will hurt lb, so I only remain the 2 highest.</li>\n<li>b0-3(local score: 0.872) was trained by re-sampling data. And I only keep 5% as valid-set in order the model can see more data.</li>\n<li>b4(local score: 0.794). It was trained by full-image(resize lv1 tiff to 448x448, data is same like the best local score fold of 5-stratified-folds). The reason why I train a full-image model is I can't find some efficient ways to make model know the context relationship by using tiles, so I train a model.</li></ul></li>\n<li>TTA\n<ul><li>only use different patch-mode, same like [Train EfficientNet-B0 w/ 36 tiles_256 [LB0.87]] (<a href=\"https://www.kaggle.com/haqishen/train-efficientnet-b0-w-36-tiles-256-lb0-87\"></a><a href=\"https://www.kaggle.com/haqishen/train-efficientnet-b0-w-36-tiles-256-lb0-87\">https://www.kaggle.com/haqishen/train-efficientnet-b0-w-36-tiles-256-lb0-87</a>)</li></ul></li>\n</ul>\n\n<hr>\n\n<p>As for bce to label. First, pred.sigmoid(), and then check how many value &gt; 0.5. And I will treat the pred is what value. pseudo code: x, = np.where(pred[i]&gt;0.5), label[i] = len(x)\nWhy I supplement a b0 and b4 is because I listen the iWildCam 1st solution from megvii. They got iWildCam 2020 1st and shared they solution. Thanks megvii !!!\nNever dream about get 0.920 and shake to 70( main reason maybe the test-set is too small for qwk and I was lucky enough :) )\nWhat's more, thanks <a href=\"/lopuhin\">@lopuhin</a> , I use <a href=\"https://www.kaggle.com/lopuhin/panda-2020-level-1-2\">PANDA: Level 1 and 2 images</a> on colab which created by him :D\nFeel free to discuss and post you point about my solution. Thanks for reading :D</p>",
  "messages": [
    {
      "id": 940448,
      "postDate": "2020-07-23T00:33:21.770Z",
      "content": "<p>Thanks for organizer! Congratulations to winners! I create this topic just want to sharing my solution(get 0.920 private lb). I know my solution will not so great as top solutions, but here I want to record something and hope you can give some advice.😄 </p>\n\n<ul>\n<li><p>Image Pre-process</p>\n\n<ul><li>Filter polluted tiles.(precompute and save the useful(pass filtering) tile ids of each image)\n<ul><li>change each tile from rgb to hsv, and observe some polluted tiles you will find the thresholds to filter polluted tiles :D</li>\n<li>how check each tile(already hsv)? Actually if you use thresholds above to fiter on tile-level, you will waste many useful tiles which have tissue information about half of tile-size. So I random sample 20 squares(length of side = tile-size//10), if the the number of squares, not be filted by the thresholds above, can reach 20xpass-ratio, I will treat this tile is useful and will also use it.</li>\n<li>If the number of tiles from a iamge is &lt; 36(I choose 36 tiles, size is 224x224). I use np.random.choice to supplement until reaching 36.</li></ul></li>\n<li>Patch Image(game changer 1)\n<ul><li>Like what <a href=\"https://www.kaggle.com/iafoss/panda-16x128x128-tiles\">PANDA 16x128x128 tiles</a> did. Thanks <a href=\"/iafoss\">@iafoss</a> !!!</li></ul></li></ul></li>\n<li><p>Models</p>\n\n<ul><li>efn-b0x3(I will call them b0-1, b0-2, b0-3 below) + efn-b4</li></ul></li>\n<li>Optimizer\n<ul><li>AdamW(default parameters)</li></ul></li>\n<li>Scheduler\n<ul><li>OneCycle(epochs=30, steps-per-epoch=int(np.ceil(len(train_dl)/acc-grad-step))\n<ul><li>about pct-start, I try 0.1, 1/30, 2/30, and 0.1 works best.</li></ul></li></ul></li>\n<li>Loss(it's the game changer 2)\n<ul><li>BCEWithLogitsLoss(thanks <a href=\"/haqishen\">@haqishen</a> !!!), details in [Train EfficientNet-B0 w/ 36 tiles_256 [LB0.87]] (<a href=\"https://www.kaggle.com/haqishen/train-efficientnet-b0-w-36-tiles-256-lb0-87\"></a><a href=\"https://www.kaggle.com/haqishen/train-efficientnet-b0-w-36-tiles-256-lb0-87\">https://www.kaggle.com/haqishen/train-efficientnet-b0-w-36-tiles-256-lb0-87</a>)</li></ul></li>\n<li>Train Details\n<ul><li>efn-b0(use tiles)\n<ul><li>batch-size=4, use gradient accumulation(also 4), so the batch-size can be treated as 4x4(not actually 4x4 of course)</li></ul></li>\n<li>efn-b4(use full-image)\n<ul><li>batch-size=16, resize lv1 tiff to 448x448</li></ul></li>\n<li>apex amp in kaggle is too hard to be installed well. So I use gradscaler from torcu.cuda.amp(seems torch version should &gt;= 1.5, I forgot the specific version).</li></ul></li>\n<li>Ensemble\n<ul><li>0.35 x (b0-1 + b0-2) + 0.25 x b0-3 + 0.05 x b4</li></ul></li>\n<li>Models Details\n<ul><li>b0-1(local score: 0.885) and b0-2(local score: 0.873) are the 2 highest local score folds of my 5-stratified-folds, and if mix other folds(local score all &lt; 0.87) will hurt lb, so I only remain the 2 highest.</li>\n<li>b0-3(local score: 0.872) was trained by re-sampling data. And I only keep 5% as valid-set in order the model can see more data.</li>\n<li>b4(local score: 0.794). It was trained by full-image(resize lv1 tiff to 448x448, data is same like the best local score fold of 5-stratified-folds). The reason why I train a full-image model is I can't find some efficient ways to make model know the context relationship by using tiles, so I train a model.</li></ul></li>\n<li>TTA\n<ul><li>only use different patch-mode, same like [Train EfficientNet-B0 w/ 36 tiles_256 [LB0.87]] (<a href=\"https://www.kaggle.com/haqishen/train-efficientnet-b0-w-36-tiles-256-lb0-87\"></a><a href=\"https://www.kaggle.com/haqishen/train-efficientnet-b0-w-36-tiles-256-lb0-87\">https://www.kaggle.com/haqishen/train-efficientnet-b0-w-36-tiles-256-lb0-87</a>)</li></ul></li>\n</ul>\n\n<hr>\n\n<p>As for bce to label. First, pred.sigmoid(), and then check how many value &gt; 0.5. And I will treat the pred is what value. pseudo code: x, = np.where(pred[i]&gt;0.5), label[i] = len(x)\nWhy I supplement a b0 and b4 is because I listen the iWildCam 1st solution from megvii. They got iWildCam 2020 1st and shared they solution. Thanks megvii !!!\nNever dream about get 0.920 and shake to 70( main reason maybe the test-set is too small for qwk and I was lucky enough :) )\nWhat's more, thanks <a href=\"/lopuhin\">@lopuhin</a> , I use <a href=\"https://www.kaggle.com/lopuhin/panda-2020-level-1-2\">PANDA: Level 1 and 2 images</a> on colab which created by him :D\nFeel free to discuss and post you point about my solution. Thanks for reading :D</p>",
      "rawMarkdown": "Thanks for organizer! Congratulations to winners! I create this topic just want to sharing my solution(get 0.920 private lb). I know my solution will not so great as top solutions, but here I want to record something and hope you can give some advice.😄 \n\n- Image Pre-process\n    - Filter polluted tiles.(precompute and save the useful(pass filtering) tile ids of each image)\n        - change each tile from rgb to hsv, and observe some polluted tiles you will find the thresholds to filter polluted tiles :D\n        - how check each tile(already hsv)? Actually if you use thresholds above to fiter on tile-level, you will waste many useful tiles which have tissue information about half of tile-size. So I random sample 20 squares(length of side = tile-size//10), if the the number of squares, not be filted by the thresholds above, can reach 20xpass-ratio, I will treat this tile is useful and will also use it.\n        - If the number of tiles from a iamge is &lt; 36(I choose 36 tiles, size is 224x224). I use np.random.choice to supplement until reaching 36.\n    - Patch Image(game changer 1)\n        - Like what [PANDA 16x128x128 tiles](https://www.kaggle.com/iafoss/panda-16x128x128-tiles) did. Thanks @iafoss !!!\n\n- Models\n    - efn-b0x3(I will call them b0-1, b0-2, b0-3 below) + efn-b4\n- Optimizer\n    - AdamW(default parameters)\n- Scheduler\n    - OneCycle(epochs=30, steps-per-epoch=int(np.ceil(len(train_dl)/acc-grad-step))\n        - about pct-start, I try 0.1, 1/30, 2/30, and 0.1 works best.\n- Loss(it's the game changer 2)\n    - BCEWithLogitsLoss(thanks @haqishen !!!), details in [Train EfficientNet-B0 w/ 36 tiles_256 [LB0.87]] (https://www.kaggle.com/haqishen/train-efficientnet-b0-w-36-tiles-256-lb0-87)\n- Train Details\n    - efn-b0(use tiles)\n        - batch-size=4, use gradient accumulation(also 4), so the batch-size can be treated as 4x4(not actually 4x4 of course)\n    - efn-b4(use full-image)\n        - batch-size=16, resize lv1 tiff to 448x448\n    - apex amp in kaggle is too hard to be installed well. So I use gradscaler from torcu.cuda.amp(seems torch version should &gt;= 1.5, I forgot the specific version).\n- Ensemble\n    - 0.35 x (b0-1 + b0-2) + 0.25 x b0-3 + 0.05 x b4\n- Models Details\n    - b0-1(local score: 0.885) and b0-2(local score: 0.873) are the 2 highest local score folds of my 5-stratified-folds, and if mix other folds(local score all &lt; 0.87) will hurt lb, so I only remain the 2 highest.\n    - b0-3(local score: 0.872) was trained by re-sampling data. And I only keep 5% as valid-set in order the model can see more data.\n    - b4(local score: 0.794). It was trained by full-image(resize lv1 tiff to 448x448, data is same like the best local score fold of 5-stratified-folds). The reason why I train a full-image model is I can't find some efficient ways to make model know the context relationship by using tiles, so I train a model.\n- TTA\n    - only use different patch-mode, same like [Train EfficientNet-B0 w/ 36 tiles_256 [LB0.87]] (https://www.kaggle.com/haqishen/train-efficientnet-b0-w-36-tiles-256-lb0-87)\n\n***\nAs for bce to label. First, pred.sigmoid(), and then check how many value &gt; 0.5. And I will treat the pred is what value. pseudo code: x, = np.where(pred[i]&gt;0.5), label[i] = len(x)\nWhy I supplement a b0 and b4 is because I listen the iWildCam 1st solution from megvii. They got iWildCam 2020 1st and shared they solution. Thanks megvii !!!\nNever dream about get 0.920 and shake to 70( main reason maybe the test-set is too small for qwk and I was lucky enough :) )\nWhat's more, thanks @lopuhin , I use [PANDA: Level 1 and 2 images](https://www.kaggle.com/lopuhin/panda-2020-level-1-2) on colab which created by him :D\nFeel free to discuss and post you point about my solution. Thanks for reading :D",
      "votes": 6
    },
    {
      "id": 942241,
      "postDate": "2020-07-23T16:50:06.667Z",
      "content": "<p>Congratulations <a href=\"/cnzengshiyuan\">@cnzengshiyuan</a>  ! Good approach and a nice leaderboard position ! </p>",
      "rawMarkdown": "Congratulations @cnzengshiyuan  ! Good approach and a nice leaderboard position ! ",
      "votes": 1,
      "replies": [
        {
          "id": 942697,
          "postDate": "2020-07-24T00:05:43.170Z",
          "content": "<p>Thank you ! I also read your team's solution, the experiments and results are quite robust :D\nI'm ready to meet your team in next cv competition✋</p>",
          "rawMarkdown": "Thank you ! I also read your team's solution, the experiments and results are quite robust :D\nI'm ready to meet your team in next cv competition✋"
        }
      ]
    },
    {
      "id": 941733,
      "postDate": "2020-07-23T11:33:41.303Z",
      "content": "<p>Congratulations  <a href=\"/cnzengshiyuan\">@cnzengshiyuan</a>  and thanks for sharing. But since I'm a beginner, I don't understand two things: what thresholds will you talk about at the pre-processing stage and how did you choose the weights for the ensemble?\nCould you please clarify?</p>",
      "rawMarkdown": "Congratulations  @cnzengshiyuan  and thanks for sharing. But since I'm a beginner, I don't understand two things: what thresholds will you talk about at the pre-processing stage and how did you choose the weights for the ensemble?\nCould you please clarify?",
      "votes": 1,
      "replies": [
        {
          "id": 942689,
          "postDate": "2020-07-23T23:59:41.377Z",
          "content": "<p>Of course :D\n1. about thresholds\nthe thresholds are the thresholds about h,s,v channel respectively to filter the polluted tiles(main color of pollutions are green and blue). And if we turn a image from rgb space to hsv space, we can filter the blue and green just focus on h channel. And with auxiliary s,v channel, I filter a big part of pollution :D\n2. about ensemble weights\nBecause test data is so small, I turn to trust public lb. So the weights finally I choose are after trying some combinations.(not the best combination, after ending, I try some different weights and get better private lb result but worse public lb result). But I don't try many experiments for limited submissions during competition.</p>\n\n<p>Hope it can help you :D</p>",
          "rawMarkdown": "Of course :D\n1. about thresholds\nthe thresholds are the thresholds about h,s,v channel respectively to filter the polluted tiles(main color of pollutions are green and blue). And if we turn a image from rgb space to hsv space, we can filter the blue and green just focus on h channel. And with auxiliary s,v channel, I filter a big part of pollution :D\n2. about ensemble weights\nBecause test data is so small, I turn to trust public lb. So the weights finally I choose are after trying some combinations.(not the best combination, after ending, I try some different weights and get better private lb result but worse public lb result). But I don't try many experiments for limited submissions during competition.\n\nHope it can help you :D",
          "votes": 1
        },
        {
          "id": 943266,
          "postDate": "2020-07-24T08:43:11.350Z",
          "content": "<p>Many thanks! It was very helpful for me :D\nDid you use the lowest image quality or intermediate?</p>",
          "rawMarkdown": "Many thanks! It was very helpful for me :D\nDid you use the lowest image quality or intermediate?",
          "votes": 1
        },
        {
          "id": 943512,
          "postDate": "2020-07-24T12:10:55.880Z",
          "content": "<p>I use intermediate image. And as I use <a href=\"https://www.kaggle.com/lopuhin/panda-2020-level-1-2\">this dataset</a>, I often call it lv1 image 😄 </p>",
          "rawMarkdown": "I use intermediate image. And as I use [this dataset](https://www.kaggle.com/lopuhin/panda-2020-level-1-2), I often call it lv1 image 😄 ",
          "votes": 1
        },
        {
          "id": 943851,
          "postDate": "2020-07-24T16:13:40.800Z",
          "content": "<p>Understood, thanks :D</p>",
          "rawMarkdown": "Understood, thanks :D",
          "votes": 1
        }
      ]
    },
    {
      "id": 940450,
      "postDate": "2020-07-23T00:35:44.450Z",
      "content": "<p>Congrats and thanks for sharing your approach.</p>",
      "rawMarkdown": "Congrats and thanks for sharing your approach.",
      "votes": 1,
      "replies": [
        {
          "id": 940463,
          "postDate": "2020-07-23T00:55:11.180Z",
          "content": "<p>You are welcome :D</p>",
          "rawMarkdown": "You are welcome :D",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 942241,
      "author_name": "Vlad Vaduva",
      "author_url": "",
      "post_date": "2020-07-23T16:50:06.667000",
      "content": "<p>Congratulations <a href=\"/cnzengshiyuan\">@cnzengshiyuan</a>  ! Good approach and a nice leaderboard position ! </p>",
      "votes": 1,
      "replies": [
        {
          "id": 942697,
          "author_name": "Shiyuan Zeng",
          "author_url": "",
          "post_date": "2020-07-24T00:05:43.170000",
          "content": "<p>Thank you ! I also read your team's solution, the experiments and results are quite robust :D\nI'm ready to meet your team in next cv competition✋</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 941733,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-23T11:33:41.303000",
      "content": "<p>Congratulations  <a href=\"/cnzengshiyuan\">@cnzengshiyuan</a>  and thanks for sharing. But since I'm a beginner, I don't understand two things: what thresholds will you talk about at the pre-processing stage and how did you choose the weights for the ensemble?\nCould you please clarify?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 942689,
          "author_name": "Shiyuan Zeng",
          "author_url": "",
          "post_date": "2020-07-23T23:59:41.377000",
          "content": "<p>Of course :D\n1. about thresholds\nthe thresholds are the thresholds about h,s,v channel respectively to filter the polluted tiles(main color of pollutions are green and blue). And if we turn a image from rgb space to hsv space, we can filter the blue and green just focus on h channel. And with auxiliary s,v channel, I filter a big part of pollution :D\n2. about ensemble weights\nBecause test data is so small, I turn to trust public lb. So the weights finally I choose are after trying some combinations.(not the best combination, after ending, I try some different weights and get better private lb result but worse public lb result). But I don't try many experiments for limited submissions during competition.</p>\n\n<p>Hope it can help you :D</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 943266,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-24T08:43:11.350000",
          "content": "<p>Many thanks! It was very helpful for me :D\nDid you use the lowest image quality or intermediate?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 943512,
          "author_name": "Shiyuan Zeng",
          "author_url": "",
          "post_date": "2020-07-24T12:10:55.880000",
          "content": "<p>I use intermediate image. And as I use <a href=\"https://www.kaggle.com/lopuhin/panda-2020-level-1-2\">this dataset</a>, I often call it lv1 image 😄 </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 943851,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-24T16:13:40.800000",
          "content": "<p>Understood, thanks :D</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 940450,
      "author_name": "Karan",
      "author_url": "",
      "post_date": "2020-07-23T00:35:44.450000",
      "content": "<p>Congrats and thanks for sharing your approach.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 940463,
          "author_name": "Shiyuan Zeng",
          "author_url": "",
          "post_date": "2020-07-23T00:55:11.180000",
          "content": "<p>You are welcome :D</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "940448": "Thanks for organizer! Congratulations to winners! I create this topic just want to sharing my solution(get 0.920 private lb). I know my solution will not so great as top solutions, but here I want to record something and hope you can give some advice.😄 \n\n- Image Pre-process\n    - Filter polluted tiles.(precompute and save the useful(pass filtering) tile ids of each image)\n        - change each tile from rgb to hsv, and observe some polluted tiles you will find the thresholds to filter polluted tiles :D\n        - how check each tile(already hsv)? Actually if you use thresholds above to fiter on tile-level, you will waste many useful tiles which have tissue information about half of tile-size. So I random sample 20 squares(length of side = tile-size//10), if the the number of squares, not be filted by the thresholds above, can reach 20xpass-ratio, I will treat this tile is useful and will also use it.\n        - If the number of tiles from a iamge is &lt; 36(I choose 36 tiles, size is 224x224). I use np.random.choice to supplement until reaching 36.\n    - Patch Image(game changer 1)\n        - Like what [PANDA 16x128x128 tiles](https://www.kaggle.com/iafoss/panda-16x128x128-tiles) did. Thanks @iafoss !!!\n\n- Models\n    - efn-b0x3(I will call them b0-1, b0-2, b0-3 below) + efn-b4\n- Optimizer\n    - AdamW(default parameters)\n- Scheduler\n    - OneCycle(epochs=30, steps-per-epoch=int(np.ceil(len(train_dl)/acc-grad-step))\n        - about pct-start, I try 0.1, 1/30, 2/30, and 0.1 works best.\n- Loss(it's the game changer 2)\n    - BCEWithLogitsLoss(thanks @haqishen !!!), details in [Train EfficientNet-B0 w/ 36 tiles_256 [LB0.87]] (https://www.kaggle.com/haqishen/train-efficientnet-b0-w-36-tiles-256-lb0-87)\n- Train Details\n    - efn-b0(use tiles)\n        - batch-size=4, use gradient accumulation(also 4), so the batch-size can be treated as 4x4(not actually 4x4 of course)\n    - efn-b4(use full-image)\n        - batch-size=16, resize lv1 tiff to 448x448\n    - apex amp in kaggle is too hard to be installed well. So I use gradscaler from torcu.cuda.amp(seems torch version should &gt;= 1.5, I forgot the specific version).\n- Ensemble\n    - 0.35 x (b0-1 + b0-2) + 0.25 x b0-3 + 0.05 x b4\n- Models Details\n    - b0-1(local score: 0.885) and b0-2(local score: 0.873) are the 2 highest local score folds of my 5-stratified-folds, and if mix other folds(local score all &lt; 0.87) will hurt lb, so I only remain the 2 highest.\n    - b0-3(local score: 0.872) was trained by re-sampling data. And I only keep 5% as valid-set in order the model can see more data.\n    - b4(local score: 0.794). It was trained by full-image(resize lv1 tiff to 448x448, data is same like the best local score fold of 5-stratified-folds). The reason why I train a full-image model is I can't find some efficient ways to make model know the context relationship by using tiles, so I train a model.\n- TTA\n    - only use different patch-mode, same like [Train EfficientNet-B0 w/ 36 tiles_256 [LB0.87]] (https://www.kaggle.com/haqishen/train-efficientnet-b0-w-36-tiles-256-lb0-87)\n\n***\nAs for bce to label. First, pred.sigmoid(), and then check how many value &gt; 0.5. And I will treat the pred is what value. pseudo code: x, = np.where(pred[i]&gt;0.5), label[i] = len(x)\nWhy I supplement a b0 and b4 is because I listen the iWildCam 1st solution from megvii. They got iWildCam 2020 1st and shared they solution. Thanks megvii !!!\nNever dream about get 0.920 and shake to 70( main reason maybe the test-set is too small for qwk and I was lucky enough :) )\nWhat's more, thanks @lopuhin , I use [PANDA: Level 1 and 2 images](https://www.kaggle.com/lopuhin/panda-2020-level-1-2) on colab which created by him :D\nFeel free to discuss and post you point about my solution. Thanks for reading :D",
    "942241": "Congratulations @cnzengshiyuan  ! Good approach and a nice leaderboard position ! ",
    "941733": "Congratulations  @cnzengshiyuan  and thanks for sharing. But since I'm a beginner, I don't understand two things: what thresholds will you talk about at the pre-processing stage and how did you choose the weights for the ensemble?\nCould you please clarify?",
    "940450": "Congrats and thanks for sharing your approach."
  }
}