{
  "id": 148894,
  "title": "Best single model",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/148894",
  "author_name": "Iafoss",
  "post_date": "2020-05-06T03:41:08.624000",
  "votes": 98,
  "comment_count": 208,
  "views": 0,
  "content": "<p>I'm starting a common competition thread: <strong>What is your current best single model?</strong></p>\n\n<p>My one:\n<strong>0.886 single fold CV, 0.90 LB</strong>\n<code>\n[[615  75  19   7   3   0]\n [ 48 497  96  11   2   0]\n [  4  86 143  84  17   1]\n [ 10  13  41 136  91  15]\n [  2  13  17  44 156  79]\n [  2   6   9  13  78 196]]\n</code>\n- ResNext50 based like in <a href=\"https://www.kaggle.com/iafoss/panda-concat-tile-pooling-starter-0-79-lb\">my kernel</a> but with a number of additional tricks. \n- Tiles from a tiff layer of intermediate resolution based on <a href=\"https://www.kaggle.com/iafoss/panda-16x128x128-tiles\">this kernel</a>.\n- No segmentation</p>\n\n<p><strong>Additional observations:</strong>\n- Low res tiff layer (12x128x128 tiles) can give up to:\n4 fold CV of 0.843, ~0.80 LB (I'd expect that LB may be not stable, and I may face with a dilemma: should I trust more LB or CV... But I didn't find any leaks so far).\n<code>\n[[2290  424  106   41   11    1]\n [ 463 1563  455  116   17    2]\n [  60  396  547  261   67   10]\n [  26   81  213  397  395  114]\n [  29   65  128  182  450  391]\n [  13   31   42   91  270  768]]\n</code>\nOn low res tiff layer the same fold I ran 0.886 model (best one above) gives:\n<code>\n0.853 CV\n[[580 108  18   9   4   0]\n [122 414  98  15   4   1]\n [ 14  89 154  67  11   0]\n [  5  17  71 106  86  21]\n [  6  17  35  49 124  80]\n [  2   7  13  30  74 178]]\n</code>\nSo <strong>going to large images doesn't give a substantial CV boost</strong> so far, while LB is quite different.</p>\n\n<ul>\n<li>Classification+Segmentation aux gave me only a tiny boost when I made a direct comparison:\n0.842 -&gt; 0.843 (4 fold CV on low res tiff layer)</li>\n</ul>\n\n<p>One also may  check the related topic on <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/145296\">CV vs LB match</a></p>",
  "messages": [
    {
      "id": 835094,
      "postDate": "2020-05-06T03:41:08.623Z",
      "content": "<p>I'm starting a common competition thread: <strong>What is your current best single model?</strong></p>\n\n<p>My one:\n<strong>0.886 single fold CV, 0.90 LB</strong>\n<code>\n[[615  75  19   7   3   0]\n [ 48 497  96  11   2   0]\n [  4  86 143  84  17   1]\n [ 10  13  41 136  91  15]\n [  2  13  17  44 156  79]\n [  2   6   9  13  78 196]]\n</code>\n- ResNext50 based like in <a href=\"https://www.kaggle.com/iafoss/panda-concat-tile-pooling-starter-0-79-lb\">my kernel</a> but with a number of additional tricks. \n- Tiles from a tiff layer of intermediate resolution based on <a href=\"https://www.kaggle.com/iafoss/panda-16x128x128-tiles\">this kernel</a>.\n- No segmentation</p>\n\n<p><strong>Additional observations:</strong>\n- Low res tiff layer (12x128x128 tiles) can give up to:\n4 fold CV of 0.843, ~0.80 LB (I'd expect that LB may be not stable, and I may face with a dilemma: should I trust more LB or CV... But I didn't find any leaks so far).\n<code>\n[[2290  424  106   41   11    1]\n [ 463 1563  455  116   17    2]\n [  60  396  547  261   67   10]\n [  26   81  213  397  395  114]\n [  29   65  128  182  450  391]\n [  13   31   42   91  270  768]]\n</code>\nOn low res tiff layer the same fold I ran 0.886 model (best one above) gives:\n<code>\n0.853 CV\n[[580 108  18   9   4   0]\n [122 414  98  15   4   1]\n [ 14  89 154  67  11   0]\n [  5  17  71 106  86  21]\n [  6  17  35  49 124  80]\n [  2   7  13  30  74 178]]\n</code>\nSo <strong>going to large images doesn't give a substantial CV boost</strong> so far, while LB is quite different.</p>\n\n<ul>\n<li>Classification+Segmentation aux gave me only a tiny boost when I made a direct comparison:\n0.842 -&gt; 0.843 (4 fold CV on low res tiff layer)</li>\n</ul>\n\n<p>One also may  check the related topic on <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/145296\">CV vs LB match</a></p>",
      "rawMarkdown": "I'm starting a common competition thread: **What is your current best single model?**\n\nMy one:\n**0.886 single fold CV, 0.90 LB**\n```\n[[615  75  19   7   3   0]\n [ 48 497  96  11   2   0]\n [  4  86 143  84  17   1]\n [ 10  13  41 136  91  15]\n [  2  13  17  44 156  79]\n [  2   6   9  13  78 196]]\n```\n- ResNext50 based like in [my kernel](https://www.kaggle.com/iafoss/panda-concat-tile-pooling-starter-0-79-lb) but with a number of additional tricks. \n- Tiles from a tiff layer of intermediate resolution based on [this kernel](https://www.kaggle.com/iafoss/panda-16x128x128-tiles).\n- No segmentation\n\n\n**Additional observations:**\n- Low res tiff layer (12x128x128 tiles) can give up to:\n4 fold CV of 0.843, ~0.80 LB (I'd expect that LB may be not stable, and I may face with a dilemma: should I trust more LB or CV... But I didn't find any leaks so far).\n```\n[[2290  424  106   41   11    1]\n [ 463 1563  455  116   17    2]\n [  60  396  547  261   67   10]\n [  26   81  213  397  395  114]\n [  29   65  128  182  450  391]\n [  13   31   42   91  270  768]]\n```\nOn low res tiff layer the same fold I ran 0.886 model (best one above) gives:\n```\n0.853 CV\n[[580 108  18   9   4   0]\n [122 414  98  15   4   1]\n [ 14  89 154  67  11   0]\n [  5  17  71 106  86  21]\n [  6  17  35  49 124  80]\n [  2   7  13  30  74 178]]\n```\nSo **going to large images doesn't give a substantial CV boost** so far, while LB is quite different.\n\n- Classification+Segmentation aux gave me only a tiny boost when I made a direct comparison:\n0.842 -&gt; 0.843 (4 fold CV on low res tiff layer)\n\nOne also may  check the related topic on [CV vs LB match](https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/145296)",
      "votes": 95
    },
    {
      "id": 836153,
      "postDate": "2020-05-06T19:03:12.863Z",
      "content": "<p>Great work. My hypothesis is that overfitting more easily occurs at lower magnifications. In my own experience, I have higher CV using level 1 vs level 0 (0.91 vs, 0.88), but LB is essentially the same (0.87). If you're using the same patch size for levels 1 and 2, then the patches for level 1 will have a much higher proportion of tissue vs. background. That might be one reason for CV-LB discrepancy. </p>",
      "rawMarkdown": "Great work. My hypothesis is that overfitting more easily occurs at lower magnifications. In my own experience, I have higher CV using level 1 vs level 0 (0.91 vs, 0.88), but LB is essentially the same (0.87). If you're using the same patch size for levels 1 and 2, then the patches for level 1 will have a much higher proportion of tissue vs. background. That might be one reason for CV-LB discrepancy. ",
      "votes": 10,
      "replies": [
        {
          "id": 836170,
          "postDate": "2020-05-06T19:26:32.747Z",
          "content": "<p>Thanks, quite interesting observation. I use different tile size, but the the fraction of background, indeed, may be different for my current setups. So the hypothesis may be that the average tissue area in train and test sets are different that creates a gap for particular setups.</p>",
          "rawMarkdown": "Thanks, quite interesting observation. I use different tile size, but the the fraction of background, indeed, may be different for my current setups. So the hypothesis may be that the average tissue area in train and test sets are different that creates a gap for particular setups.",
          "votes": 1
        },
        {
          "id": 836177,
          "postDate": "2020-05-06T19:38:04.700Z",
          "content": "<p>One thing that might be interesting to try, is to choose the tiles in a stochastic fashion, which should result in a more robust model. However, then the tile generation would have to be done during training</p>",
          "rawMarkdown": "One thing that might be interesting to try, is to choose the tiles in a stochastic fashion, which should result in a more robust model. However, then the tile generation would have to be done during training",
          "votes": 1
        },
        {
          "id": 836247,
          "postDate": "2020-05-06T20:44:31.277Z",
          "content": "<p>You could sample the tiles weighted by number of tissue pixels. That will add some stochasticity to the training process without interfering too much with learning by including many non-informative tiles. </p>",
          "rawMarkdown": "You could sample the tiles weighted by number of tissue pixels. That will add some stochasticity to the training process without interfering too much with learning by including many non-informative tiles. ",
          "votes": 2
        },
        {
          "id": 837966,
          "postDate": "2020-05-08T07:37:32.540Z",
          "content": "<p><a href=\"/vaillant\">@vaillant</a> \"You could sample the tiles weighted by number of tissue pixels\" I believe that's what lafoss' approach does already. Well, more precisely, it uses a proxy for this: it tries to avoid white pixels as much as possible (by sorting the tiles according to the sum of their pixel values and taking the first N).</p>",
          "rawMarkdown": "@vaillant \"You could sample the tiles weighted by number of tissue pixels\" I believe that's what lafoss' approach does already. Well, more precisely, it uses a proxy for this: it tries to avoid white pixels as much as possible (by sorting the tiles according to the sum of their pixel values and taking the first N)."
        },
        {
          "id": 838462,
          "postDate": "2020-05-08T15:37:48.287Z",
          "content": "<p>I was referring to a stochastic sampling during training where each tile is sampled with probability proportional to the amount of tissue pixels for more regularization vs the deterministic approach. </p>",
          "rawMarkdown": "I was referring to a stochastic sampling during training where each tile is sampled with probability proportional to the amount of tissue pixels for more regularization vs the deterministic approach. ",
          "votes": 3
        },
        {
          "id": 838641,
          "postDate": "2020-05-08T17:34:15.987Z",
          "content": "<p>owh ok I hadn't understood, thanks for the clarification. Quite the nice trick indeed !</p>",
          "rawMarkdown": "owh ok I hadn't understood, thanks for the clarification. Quite the nice trick indeed !"
        },
        {
          "id": 915890,
          "postDate": "2020-07-05T07:18:10.837Z",
          "content": "<p><a href=\"/vaillant\">@vaillant</a> first of all congrats for being at top.\nTO your post stochastic sampling weighted by Tissues ,\n1) How would u decided the weight of tiles\n2)if  few set of tiles are  weighted less compared to other tile but carrying more cancerous region in them which should be decider for isup grade in terms of amount of abnormal cells region in overall wsi slide then  wont we miss classify that wsi to lower isup grade ?</p>",
          "rawMarkdown": "@vaillant first of all congrats for being at top.\nTO your post stochastic sampling weighted by Tissues ,\n1) How would u decided the weight of tiles\n2)if  few set of tiles are  weighted less compared to other tile but carrying more cancerous region in them which should be decider for isup grade in terms of amount of abnormal cells region in overall wsi slide then  wont we miss classify that wsi to lower isup grade ?"
        },
        {
          "id": 916762,
          "postDate": "2020-07-06T02:22:38.490Z",
          "content": "<p>I don't fully understand your post, but I never tried doing this. I don't really think it will make a difference. </p>\n\n<p>If you want to try it:\n1. Take all the tiles in an image and compute mean pixel value (as Iafoss does).\n2. Normalize these values so that they sum to 1. \n3. Sample each tile based on the normalized values using np.random.choice or something similar.</p>",
          "rawMarkdown": "I don't fully understand your post, but I never tried doing this. I don't really think it will make a difference. \n\nIf you want to try it:\n1. Take all the tiles in an image and compute mean pixel value (as Iafoss does).\n2. Normalize these values so that they sum to 1. \n3. Sample each tile based on the normalized values using np.random.choice or something similar."
        }
      ]
    },
    {
      "id": 842237,
      "postDate": "2020-05-11T09:14:51.657Z",
      "content": "<p>model: se-resnet50\nfold: single fold \ncv-qwk: 0.894\nlb-qwk: 0.87</p>",
      "rawMarkdown": "model: se-resnet50\nfold: single fold \ncv-qwk: 0.894\nlb-qwk: 0.87",
      "votes": 7,
      "replies": [
        {
          "id": 847231,
          "postDate": "2020-05-14T08:42:21.723Z",
          "content": "<p>hi, great results, classification or regression?</p>",
          "rawMarkdown": "hi, great results, classification or regression?"
        },
        {
          "id": 847502,
          "postDate": "2020-05-14T12:52:24.627Z",
          "content": "<p>This is the result of regression model. </p>",
          "rawMarkdown": "This is the result of regression model. ",
          "votes": 1
        },
        {
          "id": 848149,
          "postDate": "2020-05-14T19:11:07.493Z",
          "content": "<p>May I know how many epoches you used to get this kappa score? At that epoch what is the training kappa value?</p>",
          "rawMarkdown": "May I know how many epoches you used to get this kappa score? At that epoch what is the training kappa value?"
        },
        {
          "id": 848233,
          "postDate": "2020-05-14T20:07:15.743Z",
          "content": "<p><a href=\"/hirune924\">@hirune924</a>, how you go about to do a single fold ?</p>",
          "rawMarkdown": "@hirune924, how you go about to do a single fold ?"
        },
        {
          "id": 848418,
          "postDate": "2020-05-15T00:46:08.093Z",
          "content": "<p><a href=\"/bethewinner\">@bethewinner</a> I spent 60 epochs training and I'm not monitoring QWK during train. Instead, the RMSE is 0.51.\n<a href=\"/thestoneca\">@thestoneca</a> you can see the code here(<a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/145296#828631\">https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/145296#828631</a>  ).</p>",
          "rawMarkdown": "@bethewinner I spent 60 epochs training and I'm not monitoring QWK during train. Instead, the RMSE is 0.51.\n@thestoneca you can see the code here(https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/145296#828631  ).",
          "votes": 1
        },
        {
          "id": 848423,
          "postDate": "2020-05-15T00:50:35.307Z",
          "content": "<p><a href=\"/hirune924\">@hirune924</a> thanks but I dont know if I am missing something.... </p>\n\n<p>df = pd.read_csv(os.path.join(data_dir,'train.csv'))\nskf = StratifiedKFold(n_splits=5, shuffle = True, random_state = 2020)\nfor fold, (train_index, val_index) in enumerate(skf.split(df.values, df['isup_grade'])):\n    df.loc[val_index, 'fold'] = int(fold)</p>\n\n<p>here you are creating 5 folds, and StratifiedKFold the least you can create is 2.</p>\n\n<p>Then my question here is how do you create one fold..... </p>",
          "rawMarkdown": "@hirune924 thanks but I dont know if I am missing something.... \n\ndf = pd.read_csv(os.path.join(data_dir,'train.csv'))\nskf = StratifiedKFold(n_splits=5, shuffle = True, random_state = 2020)\nfor fold, (train_index, val_index) in enumerate(skf.split(df.values, df['isup_grade'])):\n    df.loc[val_index, 'fold'] = int(fold)\n\nhere you are creating 5 folds, and StratifiedKFold the least you can create is 2.\n\nThen my question here is how do you create one fold..... \n\n"
        },
        {
          "id": 848427,
          "postDate": "2020-05-15T00:56:02.053Z",
          "content": "<p><a href=\"/thestoneca\">@thestoneca</a> I think it's called a single fold when you use only one fold of what you split as a 5fold. It's a single fold divided into 80:20</p>",
          "rawMarkdown": "@thestoneca I think it's called a single fold when you use only one fold of what you split as a 5fold. It's a single fold divided into 80:20"
        },
        {
          "id": 848438,
          "postDate": "2020-05-15T01:10:56.407Z",
          "content": "<p><a href=\"/hirune924\">@hirune924</a> meaning that of the 5 folds create you only you one? sorry, language barrier.....</p>",
          "rawMarkdown": "@hirune924 meaning that of the 5 folds create you only you one? sorry, language barrier....."
        },
        {
          "id": 848442,
          "postDate": "2020-05-15T01:17:40.877Z",
          "content": "<p><a href=\"/thestoneca\">@thestoneca</a> I just split all data up so that train:val=80:20</p>",
          "rawMarkdown": "@thestoneca I just split all data up so that train:val=80:20"
        },
        {
          "id": 849040,
          "postDate": "2020-05-15T13:02:42.113Z",
          "content": "<p><a href=\"/hirune924\">@hirune924</a>, how do you handle the images? like tiles and feed them 1 by 1? thanks.</p>",
          "rawMarkdown": "@hirune924, how do you handle the images? like tiles and feed them 1 by 1? thanks."
        },
        {
          "id": 915887,
          "postDate": "2020-07-05T07:14:04.010Z",
          "content": "<p><a href=\"/hirune924\">@hirune924</a>  i tried enough of regression in beginning but was not getting better than 0.84 .\nWhat could be conributing to the score.\n1) Tiling method ?public or your own\n2) Some customization in loss used for regression ?</p>",
          "rawMarkdown": "@hirune924  i tried enough of regression in beginning but was not getting better than 0.84 .\nWhat could be conributing to the score.\n1) Tiling method ?public or your own\n2) Some customization in loss used for regression ?"
        }
      ]
    },
    {
      "id": 840058,
      "postDate": "2020-05-09T18:38:53.680Z",
      "content": "<p>A little update: resnet34 based architecture, 0.88 single fold val -&gt; 0.85 lb. 4 fold cv 0.87 -&gt; lb 0.87. It seems that the choice of architecture (or backbone) may not be the most important</p>",
      "rawMarkdown": "A little update: resnet34 based architecture, 0.88 single fold val -&gt; 0.85 lb. 4 fold cv 0.87 -&gt; lb 0.87. It seems that the choice of architecture (or backbone) may not be the most important",
      "votes": 5,
      "replies": [
        {
          "id": 841207,
          "postDate": "2020-05-10T17:09:36.167Z",
          "content": "<p>Hello, great results. Are you using tiles as Iafoss?</p>",
          "rawMarkdown": "Hello, great results. Are you using tiles as Iafoss?"
        },
        {
          "id": 841362,
          "postDate": "2020-05-10T18:51:08.023Z",
          "content": "<p>Yes</p>",
          "rawMarkdown": "Yes",
          "votes": 2
        },
        {
          "id": 841411,
          "postDate": "2020-05-10T19:27:26.400Z",
          "content": "<p>Hi <a href=\"/shujun717\">@shujun717</a> ! getting CV and LB on par is quite curious ^^ Any idea how you achieved that feat ? </p>",
          "rawMarkdown": "Hi @shujun717 ! getting CV and LB on par is quite curious ^^ Any idea how you achieved that feat ? "
        },
        {
          "id": 841444,
          "postDate": "2020-05-10T19:46:14.713Z",
          "content": "<p>You should ask <a href=\"/iafoss\">@iafoss</a> that. His public kernel had val score of 0.77~0.78 yet the lb score is 0.79, which was quite a surprise to me. My latest 0.88 submission had a 4 fold cv score of 0.874, but iafoss achieved 0.90 with a single model with val score of 0.88. In any case, the most immediate thing that comes to my mind is that I used a very simple backbone ~ resnet34, since I have tried efficientnets which did not work well for me</p>",
          "rawMarkdown": "You should ask @iafoss that. His public kernel had val score of 0.77~0.78 yet the lb score is 0.79, which was quite a surprise to me. My latest 0.88 submission had a 4 fold cv score of 0.874, but iafoss achieved 0.90 with a single model with val score of 0.88. In any case, the most immediate thing that comes to my mind is that I used a very simple backbone ~ resnet34, since I have tried efficientnets which did not work well for me",
          "votes": 1
        },
        {
          "id": 841484,
          "postDate": "2020-05-10T20:12:43.850Z",
          "content": "<p>My impression so far is that LB mostly correlates with karolinska CV. However, the training data seems to be noisy( and one needs to find a sweet spot between LB and CV(</p>",
          "rawMarkdown": "My impression so far is that LB mostly correlates with karolinska CV. However, the training data seems to be noisy( and one needs to find a sweet spot between LB and CV(",
          "votes": 4
        },
        {
          "id": 842939,
          "postDate": "2020-05-11T17:57:39.877Z",
          "content": "<p><a href=\"/shujun717\">@shujun717</a> sorry, are u using regression or classification?</p>",
          "rawMarkdown": "@shujun717 sorry, are u using regression or classification?\n"
        }
      ]
    },
    {
      "id": 836585,
      "postDate": "2020-05-07T04:47:37.040Z",
      "content": "<p><a href=\"/iafoss\">@iafoss</a> what is the breakdown of QWK between the 2 institutions? I noticed that when my CV went from 0.88 to 0.91 with no change in LB, there was only CV improvement for Radboud data - Karolinska data performance was the same. </p>",
      "rawMarkdown": "@iafoss what is the breakdown of QWK between the 2 institutions? I noticed that when my CV went from 0.88 to 0.91 with no change in LB, there was only CV improvement for Radboud data - Karolinska data performance was the same. ",
      "votes": 6,
      "replies": [
        {
          "id": 837125,
          "postDate": "2020-05-07T14:38:10.977Z",
          "content": "<p>in my case, the score for karolinska subset was way worse than the one for radboud too</p>",
          "rawMarkdown": "in my case, the score for karolinska subset was way worse than the one for radboud too",
          "votes": 1
        },
        {
          "id": 837223,
          "postDate": "2020-05-07T16:08:37.267Z",
          "content": "<p>It looks like this:\n<code>\nkarolinska 0.909\n[[433  41   4   2   1   0]\n [ 34 382  55   1   1   0]\n [  0  49  71  36   6   0]\n [  0   5  20  40  19   1]\n [  0   6   6  12  63  23]\n [  0   1   0   3  17  44]]\nradboud 0.8403027122576369\n[[182  34  15   5   2   0]\n [ 14 115  41  10   1   0]\n [  4  37  72  48  11   1]\n [ 10   8  21  96  72  14]\n [  2   7  11  32  93  56]\n [  2   5   9  10  61 152]]\n</code>\nMeanwhile my low res models look like this:\n<code>\nkarolinska 0.805\n[[1609  254   45    6   10    1]\n [ 300 1220  237   39   15    3]\n [  42  274  227   97   24    4]\n [   9   32   92   86   73   25]\n [  20   53   46   62  174  126]\n [  11    5   14   24   52  145]]\nradboud 0.829\n[[745 122  37  27  14   3]\n [100 405 217  65  13   2]\n [ 17 117 304 181  47   7]\n [ 16  44 129 267 346 107]\n [ 14  24  47  98 273 308]\n [  7  19  19  58 219 642]]\n</code>\nSo, it seems that <strong>public LB may account only for karolinska data</strong></p>",
          "rawMarkdown": "It looks like this:\n```\nkarolinska 0.909\n[[433  41   4   2   1   0]\n [ 34 382  55   1   1   0]\n [  0  49  71  36   6   0]\n [  0   5  20  40  19   1]\n [  0   6   6  12  63  23]\n [  0   1   0   3  17  44]]\nradboud 0.8403027122576369\n[[182  34  15   5   2   0]\n [ 14 115  41  10   1   0]\n [  4  37  72  48  11   1]\n [ 10   8  21  96  72  14]\n [  2   7  11  32  93  56]\n [  2   5   9  10  61 152]]\n```\nMeanwhile my low res models look like this:\n```\nkarolinska 0.805\n[[1609  254   45    6   10    1]\n [ 300 1220  237   39   15    3]\n [  42  274  227   97   24    4]\n [   9   32   92   86   73   25]\n [  20   53   46   62  174  126]\n [  11    5   14   24   52  145]]\nradboud 0.829\n[[745 122  37  27  14   3]\n [100 405 217  65  13   2]\n [ 17 117 304 181  47   7]\n [ 16  44 129 267 346 107]\n [ 14  24  47  98 273 308]\n [  7  19  19  58 219 642]]\n```\nSo, it seems that **public LB may account only for karolinska data**",
          "votes": 6
        },
        {
          "id": 837244,
          "postDate": "2020-05-07T16:28:51.007Z",
          "content": "<p>That is my suspicion as well. </p>\n\n<p>For LB 0.87, I have this for CV:</p>\n\n<p><code>\nKarolinska \n0.88638\nRadboud\n0.90636\nOverall\n0.90936\n</code></p>\n\n<p>The question is whether the entire test set is skewed towards Karolinska or if the public test set was randomly (or intentionally) sampled that way. </p>",
          "rawMarkdown": "That is my suspicion as well. \n \nFor LB 0.87, I have this for CV:\n\n```\nKarolinska \n0.88638\nRadboud\n0.90636\nOverall\n0.90936\n```\n\nThe question is whether the entire test set is skewed towards Karolinska or if the public test set was randomly (or intentionally) sampled that way. ",
          "votes": 3
        },
        {
          "id": 837267,
          "postDate": "2020-05-07T16:43:28.473Z",
          "content": "<p>It would be a quite bad move if organizers split test into public/private based on the provider( like public= karolinska and private=radboud</p>",
          "rawMarkdown": "It would be a quite bad move if organizers split test into public/private based on the provider( like public= karolinska and private=radboud",
          "votes": 1
        },
        {
          "id": 837316,
          "postDate": "2020-05-07T17:29:07.930Z",
          "content": "<p>in Bengali Private dataset was containing graphemes which were not present in Public Test .... So I will be not surprised if private is divided based on institution ...</p>",
          "rawMarkdown": "in Bengali Private dataset was containing graphemes which were not present in Public Test .... So I will be not surprised if private is divided based on institution ...",
          "votes": 2
        },
        {
          "id": 837381,
          "postDate": "2020-05-07T18:18:42.593Z",
          "content": "<p>It doesn't really make sense for this competition because then we can just optimize for one institution's data, and the resulting model would not be as generalizable. Note that there are no new institutions in the test set. There seems to be some evidence that at least the public test set is skewed towards Karolinska. If the entire test set has a similar distribution to the training set (roughly 50/50), that would mean the private test set is skewed towards Radboud. </p>",
          "rawMarkdown": "It doesn't really make sense for this competition because then we can just optimize for one institution's data, and the resulting model would not be as generalizable. Note that there are no new institutions in the test set. There seems to be some evidence that at least the public test set is skewed towards Karolinska. If the entire test set has a similar distribution to the training set (roughly 50/50), that would mean the private test set is skewed towards Radboud. ",
          "votes": 4
        },
        {
          "id": 838212,
          "postDate": "2020-05-08T11:37:26.827Z",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> and you reminded that shake up the competition, a nightmare 😭 \nI hope it wouldn't be the same for this comp too. 🤕 </p>",
          "rawMarkdown": "@drhabib and you reminded that shake up the competition, a nightmare 😭 \nI hope it wouldn't be the same for this comp too. 🤕 ",
          "votes": 2
        },
        {
          "id": 838426,
          "postDate": "2020-05-08T15:04:52.010Z",
          "content": "<p>Great discussions. Both public and private test sets contain images from both institutions. As someone already mentioned, from the patient/medical perspective the goal is to build a model that can work on datasets from multiple institutions/labs. That's also the aim of the <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/overview/miccai-2020\">paper/workshop</a>.</p>\n\n<p>Please take the data description into account regarding the labels:</p>\n\n<blockquote>\n  <p>The labels are imperfect. This is a challenging area of pathology and even experts in the field with years of experience do not always agree on how to interpret a slide. This will make training models more difficult, but increases the potential medical value of having a strong model to provide consistent ratings. All of the private test set images and most of the public test set images were graded by multiple pathologists, but this was not feasible for the training set. You can find additional details about how consistently the pathologist's labels matched <a href=\"https://zenodo.org/record/3715938#.XrVzk6gzZPa\">here</a>.</p>\n</blockquote>\n\n<p>Also, the <a href=\"https://zenodo.org/record/3715938#.XrVzk6gzZPa\">challenge document</a> contains a lot of background info on how the data was collected.</p>",
          "rawMarkdown": "Great discussions. Both public and private test sets contain images from both institutions. As someone already mentioned, from the patient/medical perspective the goal is to build a model that can work on datasets from multiple institutions/labs. That's also the aim of the [paper/workshop](https://www.kaggle.com/c/prostate-cancer-grade-assessment/overview/miccai-2020).\n\nPlease take the data description into account regarding the labels:\n\n&gt; The labels are imperfect. This is a challenging area of pathology and even experts in the field with years of experience do not always agree on how to interpret a slide. This will make training models more difficult, but increases the potential medical value of having a strong model to provide consistent ratings. All of the private test set images and most of the public test set images were graded by multiple pathologists, but this was not feasible for the training set. You can find additional details about how consistently the pathologist's labels matched [here](https://zenodo.org/record/3715938#.XrVzk6gzZPa).\n\nAlso, the [challenge document](https://zenodo.org/record/3715938#.XrVzk6gzZPa) contains a lot of background info on how the data was collected.",
          "votes": 20
        },
        {
          "id": 838447,
          "postDate": "2020-05-08T15:29:03.710Z",
          "content": "<p>Thanks for the reference.</p>\n\n<p>I went quickly thru challenge document and highlighted important information just in case if you dont want to read 13 pages =) </p>\n\n<p><strong>TL DR:</strong>\n```\nTraining set: +/- 11.000 cases\nPublic test set: +/- 400 cases (with expert gradings)\nPrivate test set: +/- 400 cases (with expert gradings)</p>\n\n<p>```</p>\n\n<p><strong>d) Mention further important characteristics of the training, validation and test cases (e.g. class distribution in\nclassification tasks chosen according to real-world distribution vs. equal class distribution) and justify the choice.</strong></p>\n\n<p><code>\nCases were sampled based on the Gleason grade group.\n</code></p>\n\n<p>The test set was graded independently by three pathologists who are subspecialized in\nuropathology. A final consensus score was determined in three rounds. The non-expert students that annotated the training set were all medical students with prior experience in annotating pathology cases.</p>\n\n<p><strong>Radboudumc data</strong></p>\n\n<p>The training set contains label noise. This label noise is introduced due to several reasons, including inconclusive pathologist reports, annotation errors, errors in the original diagnosis, disagreement between pathologists. To test the level of label noise, we let students annotate the test set with the same protocol as the training set. The labels, as determined by the students, were then compared to the consensus labels set by the experts. On grade group, the accuracy was <strong>0.720 (quadratic weighted kappa 0.853)</strong>. These values indicate a high agreement, but show the presence of label errors. Given the nature of Gleason grading and the problems of rater disagreement, handling this label noise is part of the challenge.</p>\n\n<p>The test set was graded by three experts in consensus. We determined this as the best possible gold standard for this grading task. Still, due to the subjective nature of Gleason grading, some errors can still be present. </p>\n\n<p><strong>Karolinska data</strong></p>\n\n<p>All cases were retrieved from the STHLM3 study with participants from the Stockholm county, Sweden, during the years 2013-2015.</p>\n\n<p>Each file represents a single case/slide and has one grade. Each slide typically consists of two sections from the same biopsy, but there is occasionally only one. In the case of cancer in the slide, one of the sections has a pen mark adjacent to the tissue where cancer is present.</p>\n\n<p>Similarly to the cases from Radboudumc, the Karolinska cases contain label noise due to the subjective nature of the Gleason grading system</p>",
          "rawMarkdown": "Thanks for the reference.\n\nI went quickly thru challenge document and highlighted important information just in case if you dont want to read 13 pages =) \n\n**TL DR:**\n```\nTraining set: +/- 11.000 cases\nPublic test set: +/- 400 cases (with expert gradings)\nPrivate test set: +/- 400 cases (with expert gradings)\n\n```\n\n**d) Mention further important characteristics of the training, validation and test cases (e.g. class distribution in\nclassification tasks chosen according to real-world distribution vs. equal class distribution) and justify the choice.**\n\n```\nCases were sampled based on the Gleason grade group.\n```\n\nThe test set was graded independently by three pathologists who are subspecialized in\nuropathology. A final consensus score was determined in three rounds. The non-expert students that annotated the training set were all medical students with prior experience in annotating pathology cases.\n\n**Radboudumc data**\n\nThe training set contains label noise. This label noise is introduced due to several reasons, including inconclusive pathologist reports, annotation errors, errors in the original diagnosis, disagreement between pathologists. To test the level of label noise, we let students annotate the test set with the same protocol as the training set. The labels, as determined by the students, were then compared to the consensus labels set by the experts. On grade group, the accuracy was **0.720 (quadratic weighted kappa 0.853)**. These values indicate a high agreement, but show the presence of label errors. Given the nature of Gleason grading and the problems of rater disagreement, handling this label noise is part of the challenge.\n\nThe test set was graded by three experts in consensus. We determined this as the best possible gold standard for this grading task. Still, due to the subjective nature of Gleason grading, some errors can still be present. \n\n**Karolinska data**\n\nAll cases were retrieved from the STHLM3 study with participants from the Stockholm county, Sweden, during the years 2013-2015.\n\nEach file represents a single case/slide and has one grade. Each slide typically consists of two sections from the same biopsy, but there is occasionally only one. In the case of cancer in the slide, one of the sections has a pen mark adjacent to the tissue where cancer is present.\n\nSimilarly to the cases from Radboudumc, the Karolinska cases contain label noise due to the subjective nature of the Gleason grading system",
          "votes": 18
        },
        {
          "id": 838480,
          "postDate": "2020-05-08T15:46:34.173Z",
          "content": "<p><a href=\"/wouterbulten\">@wouterbulten</a> Thank you for clarification of the structure of the test set and additional reference material on the data preparation.\n<a href=\"/drhabib\">@drhabib</a> , really good summary, thanks.</p>",
          "rawMarkdown": "@wouterbulten Thank you for clarification of the structure of the test set and additional reference material on the data preparation.\n@drhabib , really good summary, thanks.",
          "votes": 2
        },
        {
          "id": 838645,
          "postDate": "2020-05-08T17:38:34.303Z",
          "content": "<p><a href=\"/wouterbulten\">@wouterbulten</a> <a href=\"/drhabib\">@drhabib</a> </p>\n\n<p>\"Each file represents a single case/slide and has one grade. Each slide typically consists of two sections from the same biopsy, but there is occasionally only one. In the case of cancer in the slide, one of the sections has a pen mark adjacent to the tissue where cancer is present.\"</p>\n\n<p>Can anyone clarify what the situation is with pen marks in the test set? </p>",
          "rawMarkdown": "@wouterbulten @drhabib \n\n\"Each file represents a single case/slide and has one grade. Each slide typically consists of two sections from the same biopsy, but there is occasionally only one. In the case of cancer in the slide, one of the sections has a pen mark adjacent to the tissue where cancer is present.\"\n\nCan anyone clarify what the situation is with pen marks in the test set? ",
          "votes": 1
        },
        {
          "id": 838659,
          "postDate": "2020-05-08T17:44:32.630Z",
          "content": "<p><a href=\"/fergusoci\">@fergusoci</a> : I believe I read somewhere (probably in the data tab section) that there is no pen marks in the test set.</p>",
          "rawMarkdown": "@fergusoci : I believe I read somewhere (probably in the data tab section) that there is no pen marks in the test set.",
          "votes": 5
        },
        {
          "id": 838668,
          "postDate": "2020-05-08T17:54:13.543Z",
          "content": "<p>Ah, thank you, <a href=\"/bdubreu\">@bdubreu</a> </p>",
          "rawMarkdown": "Ah, thank you, @bdubreu ",
          "votes": 2
        },
        {
          "id": 838916,
          "postDate": "2020-05-08T22:27:36.687Z",
          "content": "<p>Regarding <code>\"All of the private test set images and most of the public test set images were graded by multiple pathologists, but this was not feasible for the training set\"</code> and <code>\"The labels, as determined by the students, were then compared to the consensus labels set by the experts. On grade group, the accuracy was 0.720 (quadratic weighted kappa 0.853)\"</code>, It explains why CV and LB can have an opposit  trend, even for individual components: <code>[CV 0.886 (karolinska 0.909, radboud 0.840), LB 0.90]</code> vs. <code>[CV 0.892 (karolinska 0.923, radboud 0.842), LB 0.89]</code>. \nIt seems that we will need to deal with untrustable CV and noisy LB, which can be easily overfitted(</p>",
          "rawMarkdown": "Regarding `\"All of the private test set images and most of the public test set images were graded by multiple pathologists, but this was not feasible for the training set\"` and `\"The labels, as determined by the students, were then compared to the consensus labels set by the experts. On grade group, the accuracy was 0.720 (quadratic weighted kappa 0.853)\"`, It explains why CV and LB can have an opposit  trend, even for individual components: `[CV 0.886 (karolinska 0.909, radboud 0.840), LB 0.90]` vs. `[CV 0.892 (karolinska 0.923, radboud 0.842), LB 0.89]`. \nIt seems that we will need to deal with untrustable CV and noisy LB, which can be easily overfitted(",
          "votes": 2
        },
        {
          "id": 838923,
          "postDate": "2020-05-08T22:45:41.703Z",
          "content": "<p>Time to start fine tuning <code>random_seeds</code> =) </p>",
          "rawMarkdown": "Time to start fine tuning `random_seeds` =) ",
          "votes": 4
        },
        {
          "id": 915894,
          "postDate": "2020-07-05T07:23:06.250Z",
          "content": "<p><a href=\"/wouterbulten\">@wouterbulten</a> \nis there a possiblity that kernels public score are calculated using different portion set of  test data every time the submission happens,otherwise it is surprising that many  are having difference of .01 to 0.03 in their score every time they submit using even though same models</p>",
          "rawMarkdown": "@wouterbulten \nis there a possiblity that kernels public score are calculated using different portion set of  test data every time the submission happens,otherwise it is surprising that many  are having difference of .01 to 0.03 in their score every time they submit using even though same models"
        }
      ]
    },
    {
      "id": 856434,
      "postDate": "2020-05-21T17:50:08.520Z",
      "content": "<p>I have a lot of difference between local CV and LB\nBest single fold 0.8090CV 0.85LB\nSame model with 5 fold ensemble average (0.8012 average local CV) 0.86LB</p>\n\n<p>I did the tile selection a bit different than you and used to have more \"bright\" tiles. Currently running a model with your tile selection to see what it does I'm seeing an increase in local CV but will be interesting to see how it translates to LB.</p>\n\n<p>Haven't yet output the per provider results. Will do later or for another model run.</p>",
      "rawMarkdown": "I have a lot of difference between local CV and LB\nBest single fold 0.8090CV 0.85LB\nSame model with 5 fold ensemble average (0.8012 average local CV) 0.86LB\n\nI did the tile selection a bit different than you and used to have more \"bright\" tiles. Currently running a model with your tile selection to see what it does I'm seeing an increase in local CV but will be interesting to see how it translates to LB.\n\nHaven't yet output the per provider results. Will do later or for another model run.",
      "votes": 3
    },
    {
      "id": 854879,
      "postDate": "2020-05-20T12:06:37.153Z",
      "content": "<p>model: resnext50\nfold: single fold\ncv-qwk : 0.877\nLb-qwk : 0.89</p>",
      "rawMarkdown": "model: resnext50\nfold: single fold\ncv-qwk : 0.877\nLb-qwk : 0.89",
      "votes": 4
    },
    {
      "id": 904856,
      "postDate": "2020-06-28T03:03:07.190Z",
      "content": "<p>resnet34, single fold, single model 0.91 LB</p>",
      "rawMarkdown": "resnet34, single fold, single model 0.91 LB",
      "votes": 3,
      "replies": [
        {
          "id": 904953,
          "postDate": "2020-06-28T05:34:03.997Z",
          "content": "<p>Quite impressive! What kind of dataset are you using? Iafoss's version, Qishen Ha's version, your own?</p>",
          "rawMarkdown": "Quite impressive! What kind of dataset are you using? Iafoss's version, Qishen Ha's version, your own?"
        },
        {
          "id": 915873,
          "postDate": "2020-07-05T07:05:58.937Z",
          "content": "<p>@zhang well done...\ni wont ask much details. Just the public kernels tiling approach + res34 is enough for you to get the score or you had to adopt different tiling and loss approach.</p>",
          "rawMarkdown": "@zhang well done...\ni wont ask much details. Just the public kernels tiling approach + res34 is enough for you to get the score or you had to adopt different tiling and loss approach.",
          "votes": 2
        }
      ]
    },
    {
      "id": 857651,
      "postDate": "2020-05-22T19:41:28.943Z",
      "content": "<p><code>\nmodel: resnet34\nfold:     1\ncv:       0.884\nkarolinska: 0.8750\nradbound:  0.857\n</code></p>",
      "rawMarkdown": "```\nmodel: resnet34\nfold:     1\ncv:       0.884\nkarolinska: 0.8750\nradbound:  0.857\n```\n\n",
      "votes": 2
    },
    {
      "id": 885139,
      "postDate": "2020-06-14T00:08:02.427Z",
      "content": "<p>Can you please share what is the maximum number of tiles from intermediate layer with which you managed to train network?</p>",
      "rawMarkdown": "Can you please share what is the maximum number of tiles from intermediate layer with which you managed to train network?",
      "votes": 1,
      "replies": [
        {
          "id": 885151,
          "postDate": "2020-06-14T00:34:21.097Z",
          "content": "<p>If u check <a href=\"https://www.kaggle.com/haqishen/panda-inference-w-36-tiles-256\">this kernel</a>, u will see some setup that is working with intermediate res layer.</p>",
          "rawMarkdown": "If u check [this kernel](https://www.kaggle.com/haqishen/panda-inference-w-36-tiles-256), u will see some setup that is working with intermediate res layer."
        },
        {
          "id": 885162,
          "postDate": "2020-06-14T01:04:14.367Z",
          "content": "<p>I know there we managed like 36 tiles of 256 size. I just wanted to confirm about maximum number of patches possible through the network &gt;= 256*256. </p>",
          "rawMarkdown": "I know there we managed like 36 tiles of 256 size. I just wanted to confirm about maximum number of patches possible through the network &gt;= 256*256. "
        }
      ]
    },
    {
      "id": 874494,
      "postDate": "2020-06-05T03:44:57.837Z",
      "content": "<p>Actually I could boost single fold single model performance to 0.91 LB. Unfortunately, based on my previous subs, ensemble doesn't give much boost, and my expectation is getting only 0.01- improvement from it at single model performance of 0.90-0.91.</p>",
      "rawMarkdown": "Actually I could boost single fold single model performance to 0.91 LB. Unfortunately, based on my previous subs, ensemble doesn't give much boost, and my expectation is getting only 0.01- improvement from it at single model performance of 0.90-0.91.",
      "votes": 1,
      "replies": [
        {
          "id": 875115,
          "postDate": "2020-06-05T14:27:23.570Z",
          "content": "<p>Impressive... Does that single model at 0.91 involve better preprocessing, or different modeling tricks ? </p>",
          "rawMarkdown": "Impressive... Does that single model at 0.91 involve better preprocessing, or different modeling tricks ? "
        },
        {
          "id": 875313,
          "postDate": "2020-06-05T16:42:13.297Z",
          "content": "<p>Different tiles from ones I considered before.</p>",
          "rawMarkdown": "Different tiles from ones I considered before.",
          "votes": 5
        }
      ]
    },
    {
      "id": 860380,
      "postDate": "2020-05-25T09:00:59.837Z",
      "content": "<p>I have LB 0.84 with efficientnet-b4 (CV 0.82). Efficientnet-b0 gave me LB 0.8 (CV 0.79). </p>",
      "rawMarkdown": "I have LB 0.84 with efficientnet-b4 (CV 0.82). Efficientnet-b0 gave me LB 0.8 (CV 0.79). \n",
      "votes": 1,
      "replies": [
        {
          "id": 860597,
          "postDate": "2020-05-25T13:02:19.127Z",
          "content": "<p>Do you use the orginial image sizes from the EfficientNet paper which are optimized for its compound method? Hence, 224x224 for EffNetB0 and 380x380 for EffNetB4, or do you just use 128x128 from iafoss's tiles?</p>",
          "rawMarkdown": "Do you use the orginial image sizes from the EfficientNet paper which are optimized for its compound method? Hence, 224x224 for EffNetB0 and 380x380 for EffNetB4, or do you just use 128x128 from iafoss's tiles?"
        },
        {
          "id": 860792,
          "postDate": "2020-05-25T15:38:06.970Z",
          "content": "<p>I used 224x224 for both b0 and b4</p>",
          "rawMarkdown": "I used 224x224 for both b0 and b4",
          "votes": 1
        },
        {
          "id": 860799,
          "postDate": "2020-05-25T15:41:36.783Z",
          "content": "<p>Alright, thanks! I had the same result for B0 - I shall try it with B4:)</p>",
          "rawMarkdown": "Alright, thanks! I had the same result for B0 - I shall try it with B4:)"
        },
        {
          "id": 860816,
          "postDate": "2020-05-25T15:55:55.820Z",
          "content": "<p>Yep, there's a tradeoff between the input size and how many tiles you can fit in a batch, obviously. I kept it at 224 because otherwise I couldn't get enough tiles to fit into memory during training and it hurt the results</p>",
          "rawMarkdown": "Yep, there's a tradeoff between the input size and how many tiles you can fit in a batch, obviously. I kept it at 224 because otherwise I couldn't get enough tiles to fit into memory during training and it hurt the results",
          "votes": 1
        },
        {
          "id": 860826,
          "postDate": "2020-05-25T16:02:53.120Z",
          "content": "<p>Yes, I came across the same problem. May I ask if you again used 12 tiles?\nFurthermore, how many epochs did it take you for convergence with B4?</p>",
          "rawMarkdown": "Yes, I came across the same problem. May I ask if you again used 12 tiles?\nFurthermore, how many epochs did it take you for convergence with B4?"
        },
        {
          "id": 860965,
          "postDate": "2020-05-25T18:25:27.797Z",
          "content": "<p>I use batches of 24 (split across 3 x 1080s). I got decent convergence after about 20-30 epochs, but ran for 50 in the end.</p>",
          "rawMarkdown": "I use batches of 24 (split across 3 x 1080s). I got decent convergence after about 20-30 epochs, but ran for 50 in the end."
        },
        {
          "id": 861005,
          "postDate": "2020-05-25T19:00:19.503Z",
          "content": "<p>Okay that makes sense, as the max BS I could reach in a Kaggle Notebook is 8. Thus, only using a single GPU.</p>",
          "rawMarkdown": "Okay that makes sense, as the max BS I could reach in a Kaggle Notebook is 8. Thus, only using a single GPU."
        },
        {
          "id": 894708,
          "postDate": "2020-06-20T17:20:43.720Z",
          "content": "<p><a href=\"/fergusoci\">@fergusoci</a> \nhow many tiles you are able to use for a an image.\nWHat is your current CV vs LB.. \nI find hard to  get CV and LB match until i was 0.88 . </p>",
          "rawMarkdown": "@fergusoci \nhow many tiles you are able to use for a an image.\nWHat is your current CV vs LB.. \nI find hard to  get CV and LB match until i was 0.88 . \n"
        },
        {
          "id": 894736,
          "postDate": "2020-06-20T17:53:59.730Z",
          "content": "<p>When I wrote this post I was using a different tiling approach (from <a href=\"https://developer.ibm.com/articles/an-automatic-method-to-identify-tissues-from-big-whole-slide-images-pt1/\">here</a>). Now I'm using 36 x 256 tiles, with a tiling approach more similar to iafoss's. CV and LB are pretty well aligned, not much between them.</p>",
          "rawMarkdown": "When I wrote this post I was using a different tiling approach (from [here](https://developer.ibm.com/articles/an-automatic-method-to-identify-tissues-from-big-whole-slide-images-pt1/)). Now I'm using 36 x 256 tiles, with a tiling approach more similar to iafoss's. CV and LB are pretty well aligned, not much between them.",
          "votes": 2
        },
        {
          "id": 915883,
          "postDate": "2020-07-05T07:11:00.203Z",
          "content": "<p><a href=\"/fergusoci\">@fergusoci</a>  with effnet how much best cv loss are u able to get. I am not getting loss better than 0.230/0.240 </p>",
          "rawMarkdown": "@fergusoci  with effnet how much best cv loss are u able to get. I am not getting loss better than 0.230/0.240 "
        },
        {
          "id": 916313,
          "postDate": "2020-07-05T14:41:33.907Z",
          "content": "<p>I presume you are referring to BCE loss? I'm finding that QWK is only loosely correlated to the loss. I'm getting losses in the region of 0.19-0.2 with effb0 if I include all tiff files, 0.22-0.23 if I exclude duplicates. But qwk is roughly similar (around 0.89x) in both cases.</p>",
          "rawMarkdown": "I presume you are referring to BCE loss? I'm finding that QWK is only loosely correlated to the loss. I'm getting losses in the region of 0.19-0.2 with effb0 if I include all tiff files, 0.22-0.23 if I exclude duplicates. But qwk is roughly similar (around 0.89x) in both cases."
        }
      ]
    },
    {
      "id": 857585,
      "postDate": "2020-05-22T18:12:26.123Z",
      "content": "<p>Any idea for improvement is welcome 😃 💪 </p>\n\n<p>Framework:</p>\n\n<ul>\n<li>Tensorflow, Keras </li>\n</ul>\n\n<p>Single fold:</p>\n\n<pre><code>X_train, X_val = train_test_split(train, test_size=.2, stratify=train['isup_grade'], random_state=SEED)\n</code></pre>\n\n<p>Pre processing: </p>\n\n<ul>\n<li>4X4 tiled images (384, 384, 3)</li>\n<li>Augmentations (Hor/Ver Flip, ShiftScaleRotate)</li>\n</ul>\n\n<p>Config:</p>\n\n<pre><code>LR: 1e-3 \nBS: 16\nEpoch: 40\n</code></pre>\n\n<p>Model: Seresnext50 backbone <br> \nClassification: </p>\n\n<pre><code>loss='categorical_crossentropy'\noptimizer=optimizers.Adam(lr=LR)\nmetrics=[qw_kappa_score]\n</code></pre>\n\n<p>Callback= ReduceLROnPlateau <br></p>\n\n<pre><code>CV: 0.78\nLB: 0.79\n</code></pre>",
      "rawMarkdown": "Any idea for improvement is welcome 😃 💪 \n\nFramework:\n\n- Tensorflow, Keras \n\nSingle fold:\n\n    X_train, X_val = train_test_split(train, test_size=.2, stratify=train['isup_grade'], random_state=SEED)\n\nPre processing: \n\n- 4X4 tiled images (384, 384, 3)\n- Augmentations (Hor/Ver Flip, ShiftScaleRotate)\n\nConfig:\n\n    LR: 1e-3 \n    BS: 16\n    Epoch: 40\n\nModel: Seresnext50 backbone <br> \nClassification: \n    \n    loss='categorical_crossentropy'\n    optimizer=optimizers.Adam(lr=LR)\n    metrics=[qw_kappa_score]\n    \nCallback= ReduceLROnPlateau <br>\n\n    CV: 0.78\n    LB: 0.79",
      "votes": 1
    },
    {
      "id": 857095,
      "postDate": "2020-05-22T09:45:22.517Z",
      "content": "<p>singlefold\nLB：0.82\nCV：0.88    Karolinska ：0.8387   Radboud：0.8806</p>\n\n<p>It's funny!  Is LB all about Karolinska？：）</p>\n\n<p>Update:LB:0.84\nCV:0.87   Karolinska ：0.8640   Radboud：0.8477</p>",
      "rawMarkdown": "singlefold\nLB：0.82\nCV：0.88    Karolinska ：0.8387   Radboud：0.8806\n\nIt's funny!  Is LB all about Karolinska？：）\n\nUpdate:LB:0.84\nCV:0.87   Karolinska ：0.8640   Radboud：0.8477\n",
      "votes": 1
    },
    {
      "id": 852885,
      "postDate": "2020-05-18T18:55:11.367Z",
      "content": "<p>model: resnext50\nfold: single fold\ncv-qwk : 0.67\nLb-qwk : 0.60</p>\n\n<p>loss fn : CategoricalCrossentropy\noptimizer : adam\ninput : tile concatenated into 512X512X3</p>\n\n<p>it will be helpful if someone can share ideas on how i can improve my model next. </p>",
      "rawMarkdown": "model: resnext50\nfold: single fold\ncv-qwk : 0.67\nLb-qwk : 0.60\n\nloss fn : CategoricalCrossentropy\noptimizer : adam\ninput : tile concatenated into 512X512X3\n\nit will be helpful if someone can share ideas on how i can improve my model next. \n",
      "votes": 1,
      "replies": [
        {
          "id": 854116,
          "postDate": "2020-05-19T19:08:49.537Z",
          "content": "<p>try this approach <a href=\"https://www.kaggle.com/iafoss/panda-concat-tile-pooling-starter-0-79-lb\">https://www.kaggle.com/iafoss/panda-concat-tile-pooling-starter-0-79-lb</a></p>",
          "rawMarkdown": "try this approach https://www.kaggle.com/iafoss/panda-concat-tile-pooling-starter-0-79-lb",
          "votes": 1
        }
      ]
    },
    {
      "id": 849615,
      "postDate": "2020-05-15T23:32:54.663Z",
      "content": "<p>How long does it take you to train a 5CV 20 epoch ? I feel my biggest pain right now is that testing ideas take like 9Hours on my GTX1080 ti (for 16x128x128) so it feels very slow. I guess I can always use a cloud solution.</p>\n\n<p>Also any tip for faster iterations ?</p>",
      "rawMarkdown": "How long does it take you to train a 5CV 20 epoch ? I feel my biggest pain right now is that testing ideas take like 9Hours on my GTX1080 ti (for 16x128x128) so it feels very slow. I guess I can always use a cloud solution.\n\nAlso any tip for faster iterations ?",
      "votes": 1,
      "replies": [
        {
          "id": 849672,
          "postDate": "2020-05-16T01:01:14.010Z",
          "content": "<p>I think you should do your tests with only one fold and then when you have a good solution train 5 folds! </p>",
          "rawMarkdown": "I think you should do your tests with only one fold and then when you have a good solution train 5 folds! ",
          "votes": 3
        },
        {
          "id": 849692,
          "postDate": "2020-05-16T01:55:06.540Z",
          "content": "<p>Yeah that is kinda what I defacto ended up doing: Stopping after one fold if I didn't see any interesting improvement.</p>",
          "rawMarkdown": "Yeah that is kinda what I defacto ended up doing: Stopping after one fold if I didn't see any interesting improvement."
        }
      ]
    },
    {
      "id": 845825,
      "postDate": "2020-05-13T12:52:30.010Z",
      "content": "<p>May I ask how many tiles you use to get 0.90 LB? \nI use 112x112x64 but it doesn't come up to higher score.</p>",
      "rawMarkdown": "May I ask how many tiles you use to get 0.90 LB? \nI use 112x112x64 but it doesn't come up to higher score.",
      "votes": 1,
      "replies": [
        {
          "id": 846213,
          "postDate": "2020-05-13T16:15:30.510Z",
          "content": "<p>For the low res layer 12x128x128 seems to be near the most optimal setup. For intermediate resolution I have changed both: the size and the number of tiles. When you select your setup u should keep in mind that <code>N*sz*sz ~ tissue area</code>. Also, smaller tiles would eliminate white space (allow to use less tiles), but may degrade the model performance because it may be difficult to say what is the kind of tissue is shown on a given small piece of an image without seeing the surrounding. \nAnd for intermediate resolution it is not only about the tile setup, as <a href=\"/oscarrangel\">@oscarrangel</a> asked below. Just straight use of intermediate tiles with my method may not work because of small bs limited by GPU RAM, and other trick(s) should be added. My first trial on intermediate res also didn't give better results than low res, while the second trial got 0.90 LB. So just keep exploring different tricks, and u may find something even better than I use.</p>",
          "rawMarkdown": "For the low res layer 12x128x128 seems to be near the most optimal setup. For intermediate resolution I have changed both: the size and the number of tiles. When you select your setup u should keep in mind that `N*sz*sz ~ tissue area`. Also, smaller tiles would eliminate white space (allow to use less tiles), but may degrade the model performance because it may be difficult to say what is the kind of tissue is shown on a given small piece of an image without seeing the surrounding. \nAnd for intermediate resolution it is not only about the tile setup, as @oscarrangel asked below. Just straight use of intermediate tiles with my method may not work because of small bs limited by GPU RAM, and other trick(s) should be added. My first trial on intermediate res also didn't give better results than low res, while the second trial got 0.90 LB. So just keep exploring different tricks, and u may find something even better than I use.",
          "votes": 6
        },
        {
          "id": 846334,
          "postDate": "2020-05-13T17:33:25.870Z",
          "content": "<p>I could not make it work, debugging it, I found the GPU error lack of memory happens on the forward pass in the train met method... so I end up buying an Nvidia RTX with 24 Ggs, so now I will have about 40 gigs of GPU memory with data-parallel.  waiting to arrive. and there is a problem with the tqdm on kaggle.... <a href=\"https://www.kaggle.com/product-feedback/150817\">https://www.kaggle.com/product-feedback/150817</a></p>",
          "rawMarkdown": "I could not make it work, debugging it, I found the GPU error lack of memory happens on the forward pass in the train met method... so I end up buying an Nvidia RTX with 24 Ggs, so now I will have about 40 gigs of GPU memory with data-parallel.  waiting to arrive. and there is a problem with the tqdm on kaggle.... https://www.kaggle.com/product-feedback/150817"
        },
        {
          "id": 846478,
          "postDate": "2020-05-13T19:35:25.153Z",
          "content": "<p>I'm not sure about batchsize, i train with batch size 6 and i can get decent results. But probably higher bs could help boost a little bit the performance!</p>\n\n<p>As this paper says, mini batch size lower than 32 might be the way to go, it usually achieves better results: <a href=\"https://arxiv.org/pdf/1804.07612.pdf\">https://arxiv.org/pdf/1804.07612.pdf</a></p>\n\n<p>But I wish i had more memory.. haha!  </p>",
          "rawMarkdown": "I'm not sure about batchsize, i train with batch size 6 and i can get decent results. But probably higher bs could help boost a little bit the performance!\n\nAs this paper says, mini batch size lower than 32 might be the way to go, it usually achieves better results: https://arxiv.org/pdf/1804.07612.pdf\n\nBut I wish i had more memory.. haha!  ",
          "votes": 2
        },
        {
          "id": 846524,
          "postDate": "2020-05-13T20:15:45.973Z",
          "content": "<p>FWIW it's possible to get 0.86 LB on one 2080 Ti in about 9 hours of training. But yes memory is a huge problem.</p>",
          "rawMarkdown": "FWIW it's possible to get 0.86 LB on one 2080 Ti in about 9 hours of training. But yes memory is a huge problem.",
          "votes": 1
        },
        {
          "id": 847288,
          "postDate": "2020-05-14T09:22:18.307Z",
          "content": "<p>9 hours of training is a lot! Is it one model only or all folds?</p>",
          "rawMarkdown": "9 hours of training is a lot! Is it one model only or all folds?"
        },
        {
          "id": 847337,
          "postDate": "2020-05-14T10:22:40.210Z",
          "content": "<p>It's one model (resnet50), yes not quite as fast as I'd like to.</p>",
          "rawMarkdown": "It's one model (resnet50), yes not quite as fast as I'd like to.",
          "votes": 2
        },
        {
          "id": 847595,
          "postDate": "2020-05-14T14:10:23.503Z",
          "content": "<p>There might better solution for low batch training. Recently in <a href=\"https://arxiv.org/pdf/2004.02967.pdf\">this</a> paper they propose new normalization-activation layer <code>Evo-Norm-S0</code>. Not only it performs better in different task but it also extremely stable across many batch sizes (32 bs is almost same as 4096). In figure 1 you can see nice summarization of performance. </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2Fdff4da383bfcef9078554a8038553f2b%2FScreen%20Shot%202020-05-14%20at%2010.07.53%20AM.png?generation=1589465309983305&amp;alt=media\" alt=\"\"></p>\n\n<p>Pytorch code - <a href=\"https://github.com/digantamisra98/EvoNorm\">https://github.com/digantamisra98/EvoNorm</a>\nTensorflow code - <a href=\"https://github.com/sayakpaul/EvoNorms-in-TensorFlow-2\">https://github.com/sayakpaul/EvoNorms-in-TensorFlow-2</a></p>",
          "rawMarkdown": "There might better solution for low batch training. Recently in [this](https://arxiv.org/pdf/2004.02967.pdf) paper they propose new normalization-activation layer `Evo-Norm-S0`. Not only it performs better in different task but it also extremely stable across many batch sizes (32 bs is almost same as 4096). In figure 1 you can see nice summarization of performance. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2Fdff4da383bfcef9078554a8038553f2b%2FScreen%20Shot%202020-05-14%20at%2010.07.53%20AM.png?generation=1589465309983305&amp;alt=media)\n\nPytorch code - https://github.com/digantamisra98/EvoNorm\nTensorflow code - https://github.com/sayakpaul/EvoNorms-in-TensorFlow-2",
          "votes": 5
        },
        {
          "id": 847622,
          "postDate": "2020-05-14T14:26:33.130Z",
          "content": "<p>Another thing that might be interesting is inplace activated batchnorm, which can help with a big issue of the competition - memory, although it has not worked for me... Have you tried anything like that?</p>",
          "rawMarkdown": "Another thing that might be interesting is inplace activated batchnorm, which can help with a big issue of the competition - memory, although it has not worked for me... Have you tried anything like that?",
          "votes": 1
        },
        {
          "id": 847638,
          "postDate": "2020-05-14T14:35:14.343Z",
          "content": "<p>I always used inplace... =) </p>",
          "rawMarkdown": "I always used inplace... =) "
        },
        {
          "id": 847662,
          "postDate": "2020-05-14T14:44:14.363Z",
          "content": "<p>Im not quite sure what is inplace activated batchnorm? is it like an complete other batchnorm?</p>",
          "rawMarkdown": "Im not quite sure what is inplace activated batchnorm? is it like an complete other batchnorm?"
        },
        {
          "id": 868454,
          "postDate": "2020-05-31T08:48:15.430Z",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> Sorry to disturb you. If I wanna replace BN with EvoNorm, how can I do it?\n<code>\ndef dfs(net):\n    for x in net.children():\n        if isinstance(x, nn.BatchNorm2d):\n            print(x, x.num_features)\n        else:\n            dfs(x)\n</code>\nI write a dfs to print all the BatchNorm2d in backbone, but I can't change the type of it by using ’x=EvoNrom‘. Is python exists a way to use quotative-vaiable in for-loop?(like the c++ code below)\n<code>\nfor ( int&amp; i = 0; i &amp;lt; N; i ++ ) { ...... }\n</code>\nOr I should copy the whole codes of my backbone and then modify it by hand?(it's so terrible ;_;)</p>",
          "rawMarkdown": "@drhabib Sorry to disturb you. If I wanna replace BN with EvoNorm, how can I do it?\n```\ndef dfs(net):\n    for x in net.children():\n        if isinstance(x, nn.BatchNorm2d):\n            print(x, x.num_features)\n        else:\n            dfs(x)\n```\nI write a dfs to print all the BatchNorm2d in backbone, but I can't change the type of it by using ’x=EvoNrom‘. Is python exists a way to use quotative-vaiable in for-loop?(like the c++ code below)\n```\nfor ( int&amp; i = 0; i &lt; N; i ++ ) { ...... }\n```\nOr I should copy the whole codes of my backbone and then modify it by hand?(it's so terrible ;_;)",
          "votes": 1
        },
        {
          "id": 868980,
          "postDate": "2020-05-31T15:59:06.013Z",
          "content": "<p>You can consider the following code I tried for GN:\n<code>\ndef to_GN(model):\n    for child_name, child in model.named_children():\n        if isinstance(child, nn.BatchNorm2d):\n            setattr(model, child_name, nn.GroupNorm(32,child.num_features))\n        else:\n            to_GN(child)\n</code></p>\n\n<p>Though, such replacement quite destroys the pretrained weights and didn't work well for me.</p>",
          "rawMarkdown": "You can consider the following code I tried for GN:\n```\ndef to_GN(model):\n    for child_name, child in model.named_children():\n        if isinstance(child, nn.BatchNorm2d):\n            setattr(model, child_name, nn.GroupNorm(32,child.num_features))\n        else:\n            to_GN(child)\n```\n\nThough, such replacement quite destroys the pretrained weights and didn't work well for me.",
          "votes": 3
        },
        {
          "id": 870821,
          "postDate": "2020-06-02T00:35:15.930Z",
          "content": "<p>Thank you so much!💯  I considered to trun BN to EvoNorm is my batch-size becomes samller and samller when I add more tiles. But you are right, abandon the pretrained weights also a big problem☹️ </p>",
          "rawMarkdown": "Thank you so much!💯  I considered to trun BN to EvoNorm is my batch-size becomes samller and samller when I add more tiles. But you are right, abandon the pretrained weights also a big problem☹️ "
        },
        {
          "id": 870880,
          "postDate": "2020-06-02T02:24:45.877Z",
          "content": "<p>While your overall batch size gets smaller, you still send a lot of tiles through the CNN if you use the tile strategy so potentially say 4 x 32 is still good enough for batchnorm in the CNN. While the tiles are correlated it probably isn't that bad and discarding pretrained weight is probably not worth it. On the other hand, for the FC part maybe I'll try something that is less dependent on batch-size.</p>",
          "rawMarkdown": "While your overall batch size gets smaller, you still send a lot of tiles through the CNN if you use the tile strategy so potentially say 4 x 32 is still good enough for batchnorm in the CNN. While the tiles are correlated it probably isn't that bad and discarding pretrained weight is probably not worth it. On the other hand, for the FC part maybe I'll try something that is less dependent on batch-size.",
          "votes": 1
        },
        {
          "id": 871001,
          "postDate": "2020-06-02T04:49:25.793Z",
          "content": "<p>Your words make sense!👍  Maybe I should try iafoss's way to feed the image into model instead of seaming tiles into a large image.</p>",
          "rawMarkdown": "Your words make sense!👍  Maybe I should try iafoss's way to feed the image into model instead of seaming tiles into a large image."
        },
        {
          "id": 884344,
          "postDate": "2020-06-13T09:52:54.303Z",
          "content": "<p>So, did any of you guys managed to make gradient accumulation work ? I managed to do grad_acc, but then, the small batch size screws the batch normalizations and results are not good in the end.\nTo that effect, I tried replacing batchNorm layers with groupNorm and EvoNorm, but both failed... Any clue on how one can effectively find a way to solve the batch-size problem ? </p>",
          "rawMarkdown": "So, did any of you guys managed to make gradient accumulation work ? I managed to do grad_acc, but then, the small batch size screws the batch normalizations and results are not good in the end.\nTo that effect, I tried replacing batchNorm layers with groupNorm and EvoNorm, but both failed... Any clue on how one can effectively find a way to solve the batch-size problem ? ",
          "votes": 1
        },
        {
          "id": 885130,
          "postDate": "2020-06-13T23:43:33.077Z",
          "content": "<p>As a side note I use pytorch lightning to make grad accumulation work easily but of course it doesn't solve the batch norm issue...</p>",
          "rawMarkdown": "As a side note I use pytorch lightning to make grad accumulation work easily but of course it doesn't solve the batch norm issue..."
        },
        {
          "id": 887910,
          "postDate": "2020-06-16T01:52:34.367Z",
          "content": "<p><a href=\"/bdubreu\">@bdubreu</a> did it not work for you? That's a little surprising actually. When I used resnext50, I had to use a batch size of 4 and gradient accumulation every 12 batches, but my model turned out ok anyway (my current best 0.90 run). </p>",
          "rawMarkdown": "@bdubreu did it not work for you? That's a little surprising actually. When I used resnext50, I had to use a batch size of 4 and gradient accumulation every 12 batches, but my model turned out ok anyway (my current best 0.90 run). ",
          "votes": 1
        },
        {
          "id": 887952,
          "postDate": "2020-06-16T03:06:24.047Z",
          "content": "<p>I think what Roussel means is that accumulate gradient is just a train trick and not the same with really enlarge batch-size?\nJust a guess😜 </p>",
          "rawMarkdown": "I think what Roussel means is that accumulate gradient is just a train trick and not the same with really enlarge batch-size?\nJust a guess😜 "
        },
        {
          "id": 887964,
          "postDate": "2020-06-16T03:19:36.333Z",
          "content": "<p><a href=\"/cnzengshiyuan\">@cnzengshiyuan</a>  It's an intermediary solution. You get good gradients from it but it doesn't solve the batch norm issue. I still use it though.</p>",
          "rawMarkdown": "@cnzengshiyuan  It's an intermediary solution. You get good gradients from it but it doesn't solve the batch norm issue. I still use it though.",
          "votes": 1
        },
        {
          "id": 888069,
          "postDate": "2020-06-16T05:27:06.533Z",
          "content": "<p><a href=\"/arroqc\">@arroqc</a> Thanks for your reply, I just started to use it😄 </p>",
          "rawMarkdown": "@arroqc Thanks for your reply, I just started to use it😄 "
        },
        {
          "id": 888189,
          "postDate": "2020-06-16T07:13:41.463Z",
          "content": "<p><a href=\"/shujun717\">@shujun717</a> thanks for telling me this. I was trying to replicate my partner's results (she has access to a 24gb GPU). So, having only 6gigs, I tried bs8 (she was doing 32) and tried everything I could to replicate her results. Never managed to. That being said, I wanted to \"fake\" a batch size of 32, so I never tried to do gradient accumulation for more than 4 batches. Maybe doing more will help, I will try that and report here !</p>",
          "rawMarkdown": "@shujun717 thanks for telling me this. I was trying to replicate my partner's results (she has access to a 24gb GPU). So, having only 6gigs, I tried bs8 (she was doing 32) and tried everything I could to replicate her results. Never managed to. That being said, I wanted to \"fake\" a batch size of 32, so I never tried to do gradient accumulation for more than 4 batches. Maybe doing more will help, I will try that and report here !"
        },
        {
          "id": 890579,
          "postDate": "2020-06-17T15:24:06.320Z",
          "content": "<p>New result for me with a CV at 0.9009 and LB at 0.89 (TTA) with single fold (and not even sure if this is the best fold as this is my first try with new architecture/process).</p>",
          "rawMarkdown": "New result for me with a CV at 0.9009 and LB at 0.89 (TTA) with single fold (and not even sure if this is the best fold as this is my first try with new architecture/process)."
        },
        {
          "id": 890808,
          "postDate": "2020-06-17T17:56:03.520Z",
          "content": "<p><a href=\"/arroqc\">@arroqc</a> Do you know how much of the lift was down to architecture versus process? </p>\n\n<p>I've been playing around trying to get a feel for how much each of these could contribute. Thinking that how the tiles are constructed is key. I was originally selecting the top tiles based on the approach \n <a href=\"https://developer.ibm.com/articles/an-automatic-method-to-identify-tissues-from-big-whole-slide-images-pt1/\">here</a>. Switching to the (much simpler!) <a href=\"/iafoss\">@iafoss</a> method led to LB 0.83 -&gt; 0.87 with the same Efficientnet B0 architecture. Now experimenting with other approaches to tiling; it feels like using full resolution (or half resolution) would be useful, but I'm running out of space on the GPUs and gradient accumulation doesn't seem to be playing ball! Can't get the batch norm working...</p>",
          "rawMarkdown": "@arroqc Do you know how much of the lift was down to architecture versus process? \n\nI've been playing around trying to get a feel for how much each of these could contribute. Thinking that how the tiles are constructed is key. I was originally selecting the top tiles based on the approach \n [here](https://developer.ibm.com/articles/an-automatic-method-to-identify-tissues-from-big-whole-slide-images-pt1/). Switching to the (much simpler!) @iafoss method led to LB 0.83 -&gt; 0.87 with the same Efficientnet B0 architecture. Now experimenting with other approaches to tiling; it feels like using full resolution (or half resolution) would be useful, but I'm running out of space on the GPUs and gradient accumulation doesn't seem to be playing ball! Can't get the batch norm working..."
        },
        {
          "id": 890851,
          "postDate": "2020-06-17T18:14:18.963Z",
          "content": "<p>Process is definitely most of the lift. Let's just say that my tile selection is much better than it used to be (and that this result actually uses a lower number of tiles than I used to). I am now working with it to see how far I can push this idea. Like most people I still struggle with compute time and batch norms when using something like 36 tiles.</p>\n\n<p>Also, I'm getting better results with stitched tiles than bag of tiles. The main reason is Karolinska is better with stitched (0.91CV). I am still investigating why but maybe it has to do with the higher number of benign slides with karolinska where your network can more quickly classify it by seeing all tiles at the same time. Radboud is exactly the same for both methods for me though (0.87CV).</p>",
          "rawMarkdown": "Process is definitely most of the lift. Let's just say that my tile selection is much better than it used to be (and that this result actually uses a lower number of tiles than I used to). I am now working with it to see how far I can push this idea. Like most people I still struggle with compute time and batch norms when using something like 36 tiles.\n\nAlso, I'm getting better results with stitched tiles than bag of tiles. The main reason is Karolinska is better with stitched (0.91CV). I am still investigating why but maybe it has to do with the higher number of benign slides with karolinska where your network can more quickly classify it by seeing all tiles at the same time. Radboud is exactly the same for both methods for me though (0.87CV).",
          "votes": 1
        },
        {
          "id": 891209,
          "postDate": "2020-06-18T02:41:45.187Z",
          "content": "<p>what is \"stitched tiles\" ?</p>",
          "rawMarkdown": "what is \"stitched tiles\" ?"
        },
        {
          "id": 891231,
          "postDate": "2020-06-18T03:22:40.893Z",
          "content": "<p>Did you try to maintain sequence of patches in a way? Or something similar to that?</p>",
          "rawMarkdown": "Did you try to maintain sequence of patches in a way? Or something similar to that?"
        },
        {
          "id": 891461,
          "postDate": "2020-06-18T07:36:23.890Z",
          "content": "<p>I've tried gradient accumulation without success. \nThanks <a href=\"/arroqc\">@arroqc</a> for you feedbacks. I think stitched tiles means concatenated tiles ?</p>",
          "rawMarkdown": "I've tried gradient accumulation without success. \nThanks @arroqc for you feedbacks. I think stitched tiles means concatenated tiles ?",
          "votes": 1
        },
        {
          "id": 891849,
          "postDate": "2020-06-18T13:45:37.817Z",
          "content": "<p>Yes by stitched I mean concatenated in a square.</p>",
          "rawMarkdown": "Yes by stitched I mean concatenated in a square.",
          "votes": 1
        },
        {
          "id": 891906,
          "postDate": "2020-06-18T14:24:13.990Z",
          "content": "<p><a href=\"/arroqc\">@arroqc</a> congrats on the lift ! does your better tiles involves using the masks ? All my experiments trying to use the mask for tile selection seem to fail. I haven't made a segmentation model to mask the test files and pick the tiles accordingly though, but I think most people haven't done that anyways...</p>",
          "rawMarkdown": "@arroqc congrats on the lift ! does your better tiles involves using the masks ? All my experiments trying to use the mask for tile selection seem to fail. I haven't made a segmentation model to mask the test files and pick the tiles accordingly though, but I think most people haven't done that anyways..."
        },
        {
          "id": 891914,
          "postDate": "2020-06-18T14:29:35.597Z",
          "content": "<p>No I don't use the masks as it seems to be very noisy and also differences between the two providers...</p>",
          "rawMarkdown": "No I don't use the masks as it seems to be very noisy and also differences between the two providers...",
          "votes": 2
        },
        {
          "id": 893390,
          "postDate": "2020-06-19T15:13:19.737Z",
          "content": "<p><a href=\"/arroqc\">@arroqc</a> did you get this result using regression or classification?</p>",
          "rawMarkdown": "\n@arroqc did you get this result using regression or classification?"
        },
        {
          "id": 893406,
          "postDate": "2020-06-19T15:25:53.723Z",
          "content": "<p>This is with the bin + binary cross entropy method similar to the current best kernel. But I haven't tried with other losses so I don't know how it compares.</p>",
          "rawMarkdown": "This is with the bin + binary cross entropy method similar to the current best kernel. But I haven't tried with other losses so I don't know how it compares.",
          "votes": 1
        },
        {
          "id": 902952,
          "postDate": "2020-06-26T13:37:28.107Z",
          "content": "<p><a href=\"/shujun717\">@shujun717</a> does your gradient accumulation use reduction=sum, or reduction=mean ? \nI tried numerous setups with a step every N batches, to no avail. I think my implementation is not correct (since it worked for you). When you do a step, do you divide the accumulated gradients by the number of batches you do between two steps ? </p>\n\n<p>Sorry for annoying you with the specifics, but I'm trying a whole bunch of things (even with GroupNorm, EvoNorm and the like) and nothing seems to work ^^ </p>",
          "rawMarkdown": "@shujun717 does your gradient accumulation use reduction=sum, or reduction=mean ? \nI tried numerous setups with a step every N batches, to no avail. I think my implementation is not correct (since it worked for you). When you do a step, do you divide the accumulated gradients by the number of batches you do between two steps ? \n\nSorry for annoying you with the specifics, but I'm trying a whole bunch of things (even with GroupNorm, EvoNorm and the like) and nothing seems to work ^^ "
        },
        {
          "id": 903050,
          "postDate": "2020-06-26T14:43:12.823Z",
          "content": "<p>I used reduction=\"none\" then torch.mean and divide by <code>gradient_accumulation_steps</code>. And \n<code>python\nif step%gradient_accumulation_steps==0:\n  optimizer.step()\n  optimizer.zero_grad()\n</code>\nThis should not be different from reduction=sum, or reduction=mean, provided that you divide your loss accordingly. I only used reduction=\"none\" to I can scale the losses individually. I don't freeze bn or anything either </p>",
          "rawMarkdown": "I used reduction=\"none\" then torch.mean and divide by ```gradient_accumulation_steps```. And \n```python\nif step%gradient_accumulation_steps==0:\n  optimizer.step()\n  optimizer.zero_grad()\n```\nThis should not be different from reduction=sum, or reduction=mean, provided that you divide your loss accordingly. I only used reduction=\"none\" to I can scale the losses individually. I don't freeze bn or anything either \n"
        },
        {
          "id": 903076,
          "postDate": "2020-06-26T14:58:41.810Z",
          "content": "<p>Then I will have to take a look at everybody else's code (including yours) at the end of comp' because I don't see what I'm missing here ^^ Thanks for taking the time to reply though ! </p>",
          "rawMarkdown": "Then I will have to take a look at everybody else's code (including yours) at the end of comp' because I don't see what I'm missing here ^^ Thanks for taking the time to reply though ! "
        },
        {
          "id": 903157,
          "postDate": "2020-06-26T16:00:43.140Z",
          "content": "<p>I can take a look at your training loop if you want. It might just be some simple mistake I feel, since it worked for me the first time I tried it </p>",
          "rawMarkdown": "I can take a look at your training loop if you want. It might just be some simple mistake I feel, since it worked for me the first time I tried it "
        },
        {
          "id": 903659,
          "postDate": "2020-06-27T03:12:09.453Z",
          "content": "<p>I have the same problem. I have implemented the model update process as follows (refer to <a href=\"https://discuss.pytorch.org/t/why-do-we-need-to-set-the-gradients-manually-to-zero-in-pytorch/4903/20\">https://discuss.pytorch.org/t/why-do-we-need-to-set-the-gradients-manually-to-zero-in-pytorch/4903/20</a> ):\n```\ndef update_model(self, model, loss, optimizer):\n        if self.mode == \"train\":\n            need_gradient_step = (\n                self._accumulation_counter + 1) % self.accumulation_steps == 0\n            model.zero_grad()\n            if self.use_amp and amp_enable:\n                delay_unscale = not need_gradient_step\n                with amp.scale_loss(loss, optimizer, delay_unscale=delay_unscale) as scaled_loss:\n                    scaled_loss.backward()\n            else:\n                loss.backward()</p>\n\n<pre><code>        if need_gradient_step:\n            optimizer.step()\n            self._accumulation_counter = 0\n</code></pre>\n\n<p>```\nIs this a mistake?</p>",
          "rawMarkdown": "I have the same problem. I have implemented the model update process as follows (refer to https://discuss.pytorch.org/t/why-do-we-need-to-set-the-gradients-manually-to-zero-in-pytorch/4903/20 ):\n```\ndef update_model(self, model, loss, optimizer):\n        if self.mode == \"train\":\n            need_gradient_step = (\n                self._accumulation_counter + 1) % self.accumulation_steps == 0\n            model.zero_grad()\n            if self.use_amp and amp_enable:\n                delay_unscale = not need_gradient_step\n                with amp.scale_loss(loss, optimizer, delay_unscale=delay_unscale) as scaled_loss:\n                    scaled_loss.backward()\n            else:\n                loss.backward()\n\n            if need_gradient_step:\n                optimizer.step()\n                self._accumulation_counter = 0\n```\nIs this a mistake?"
        },
        {
          "id": 916978,
          "postDate": "2020-07-06T06:50:24.253Z",
          "content": "<p>This looks fine to me, not sure why you are setting _accumulation_counter to 0 tho</p>",
          "rawMarkdown": "This looks fine to me, not sure why you are setting _accumulation_counter to 0 tho"
        },
        {
          "id": 917072,
          "postDate": "2020-07-06T08:05:29.830Z",
          "content": "<p>Thank you reply.\n <code>accumelation_counter</code> is incremented with the training loop.\nI keep debugging...</p>",
          "rawMarkdown": "Thank you reply.\n `accumelation_counter` is incremented with the training loop.\nI keep debugging..."
        },
        {
          "id": 917115,
          "postDate": "2020-07-06T08:51:46.963Z",
          "content": "<p>What specific bug/issue do you have with this code?</p>",
          "rawMarkdown": "What specific bug/issue do you have with this code?"
        },
        {
          "id": 917143,
          "postDate": "2020-07-06T09:11:15Z",
          "content": "<p>Hi <a href=\"/shujun717\">@shujun717</a> ! for some reason, I never saw your comment above ! Thank you so much for the offer to review my code. But as this is a competition I believe I should be able to handle this on my own. I might come back to you <em>after</em> the competition if I don't see what I did wrong by looking at everyone's code though, if that's ok for you ;)\nGood luck for the end, I hope you stay top10 ! </p>",
          "rawMarkdown": "Hi @shujun717 ! for some reason, I never saw your comment above ! Thank you so much for the offer to review my code. But as this is a competition I believe I should be able to handle this on my own. I might come back to you _after_ the competition if I don't see what I did wrong by looking at everyone's code though, if that's ok for you ;)\nGood luck for the end, I hope you stay top10 ! "
        }
      ]
    },
    {
      "id": 842703,
      "postDate": "2020-05-11T15:25:53.827Z",
      "content": "<p><a href=\"/iafoss\">@iafoss</a> how many epochs do you train per fold? In my experiments, my model seems to need more than 30 epochs per fold to converge</p>",
      "rawMarkdown": "@iafoss how many epochs do you train per fold? In my experiments, my model seems to need more than 30 epochs per fold to converge",
      "votes": 1,
      "replies": [
        {
          "id": 842708,
          "postDate": "2020-05-11T15:29:09.190Z",
          "content": "<p>It's correct, the convergence is not very fast, and I train for more than 30 epochs.</p>",
          "rawMarkdown": "It's correct, the convergence is not very fast, and I train for more than 30 epochs.",
          "votes": 2
        },
        {
          "id": 842712,
          "postDate": "2020-05-11T15:35:10.067Z",
          "content": "<p>I see. Thanks</p>",
          "rawMarkdown": "I see. Thanks"
        }
      ]
    },
    {
      "id": 835690,
      "postDate": "2020-05-06T12:52:41.957Z",
      "content": "<p><a href=\"/iafoss\">@iafoss</a> are you using masks?</p>",
      "rawMarkdown": "@iafoss are you using masks?",
      "votes": 1,
      "replies": [
        {
          "id": 835950,
          "postDate": "2020-05-06T16:03:41.897Z",
          "content": "<p>In my experiment with mask aux, as I pointed out above, I got only a slight improvment. Meanwhile, the training time increased quite a bit. So, for larger resolution I didn't use it.</p>",
          "rawMarkdown": "In my experiment with mask aux, as I pointed out above, I got only a slight improvment. Meanwhile, the training time increased quite a bit. So, for larger resolution I didn't use it."
        },
        {
          "id": 836930,
          "postDate": "2020-05-07T11:28:16.763Z",
          "content": "<p><a href=\"/iafoss\">@iafoss</a> okay thanks!</p>",
          "rawMarkdown": "@iafoss okay thanks!"
        }
      ]
    },
    {
      "id": 855587,
      "postDate": "2020-05-21T03:43:56.537Z",
      "content": "<p>model: single fold seresnext50\nCV: 0.9\nLB: 0.88</p>",
      "rawMarkdown": "model: single fold seresnext50\nCV: 0.9\nLB: 0.88",
      "votes": 2
    },
    {
      "id": 835514,
      "postDate": "2020-05-06T10:13:58.833Z",
      "content": "<p>Nice result <a href=\"/iafoss\">@iafoss</a> . If you don't mind, your tricks are related to model architecture or data augmentation/preprocessing ?</p>",
      "rawMarkdown": "Nice result @iafoss . If you don't mind, your tricks are related to model architecture or data augmentation/preprocessing ?",
      "votes": 2,
      "replies": [
        {
          "id": 835943,
          "postDate": "2020-05-06T16:00:10.557Z",
          "content": "<p>I'd say they are related to both you mentioned and and other things as well.</p>",
          "rawMarkdown": "I'd say they are related to both you mentioned and and other things as well.",
          "votes": 2
        }
      ]
    },
    {
      "id": 842831,
      "postDate": "2020-05-11T16:40:05.743Z",
      "content": "<p>Hi there all,</p>\n\n<p>is any one know how to translate this Pytorch code into fastai ?</p>\n\n<p>~~~\nfor i, (inputs, labels) in enumerate(training_set):\n    predictions = model(inputs)                     # Forward pass\n    loss = loss_function(predictions, labels)       # Compute loss function\n    loss = loss / accumulation_steps                # Normalize our loss (if averaged)\n    loss.backward()                                 # Backward pass\n    if (i+1) % accumulation_steps == 0:             # Wait for several backward steps\n        optimizer.step()                            # Now we can do an optimizer step\n        model.zero_grad()                           # Reset gradients tensors\n        if (i+1) % evaluation_steps == 0:           # Evaluate the model when we...\n            evaluate_model() \n~~~</p>",
      "rawMarkdown": "Hi there all,\n\nis any one know how to translate this Pytorch code into fastai ?\n\n~~~\nfor i, (inputs, labels) in enumerate(training_set):\n    predictions = model(inputs)                     # Forward pass\n    loss = loss_function(predictions, labels)       # Compute loss function\n    loss = loss / accumulation_steps                # Normalize our loss (if averaged)\n    loss.backward()                                 # Backward pass\n    if (i+1) % accumulation_steps == 0:             # Wait for several backward steps\n        optimizer.step()                            # Now we can do an optimizer step\n        model.zero_grad()                           # Reset gradients tensors\n        if (i+1) % evaluation_steps == 0:           # Evaluate the model when we...\n            evaluate_model() \n~~~",
      "replies": [
        {
          "id": 842898,
          "postDate": "2020-05-11T17:28:51.523Z",
          "content": "<p>You always can use your favorite search engine and type \"fastai gradient accumulation callback\". Specifically now fast.ai has AccumulateScheduler. Also I wrote my own one a while back in this kernel\n<a href=\"https://www.kaggle.com/iafoss/hypercolumns-pneumothorax-fastai-0-831-lb\">https://www.kaggle.com/iafoss/hypercolumns-pneumothorax-fastai-0-831-lb</a></p>",
          "rawMarkdown": "You always can use your favorite search engine and type \"fastai gradient accumulation callback\". Specifically now fast.ai has AccumulateScheduler. Also I wrote my own one a while back in this kernel\nhttps://www.kaggle.com/iafoss/hypercolumns-pneumothorax-fastai-0-831-lb",
          "votes": 1
        },
        {
          "id": 842966,
          "postDate": "2020-05-11T18:19:08.093Z",
          "content": "<p>Or why not just use pure pytorch?</p>",
          "rawMarkdown": "Or why not just use pure pytorch?"
        },
        {
          "id": 842988,
          "postDate": "2020-05-11T18:33:03.227Z",
          "content": "<p>@lafoss, I did search and posted the question on Fastai forum, but everyone send m me to fastai2, I have tried to upgrade your project to fastai2 but it does not works.... I guess I don't know much about fastai in order to upgrade it....</p>\n\n<p>Thanks for the post, trying to follow your advise on GPU optimization.... so I have have gradients update not when I use bz=8, but simulate bz=32</p>",
          "rawMarkdown": "@lafoss, I did search and posted the question on Fastai forum, but everyone send m me to fastai2, I have tried to upgrade your project to fastai2 but it does not works.... I guess I don't know much about fastai in order to upgrade it....\n\nThanks for the post, trying to follow your advise on GPU optimization.... so I have have gradients update not when I use bz=8, but simulate bz=32\n"
        },
        {
          "id": 842991,
          "postDate": "2020-05-11T18:35:25.507Z",
          "content": "<p><a href=\"/shujun717\">@shujun717</a> , because then you will have to get into the fit_one_cycle() code which I dont w want to mess with it, I know it can be implemented with a callback call, I am learning, maybe there is a way, but I dont knnow it... \nMaybe someone else can answer your questions better.</p>",
          "rawMarkdown": "@shujun717 , because then you will have to get into the fit_one_cycle() code which I dont w want to mess with it, I know it can be implemented with a callback call, I am learning, maybe there is a way, but I dont knnow it... \nMaybe someone else can answer your questions better."
        },
        {
          "id": 843005,
          "postDate": "2020-05-11T18:45:53.627Z",
          "content": "<p><a href=\"https://forums.fast.ai/t/gradient-accumulation/70968/5?u=orangelmx\">https://forums.fast.ai/t/gradient-accumulation/70968/5?u=orangelmx</a></p>\n\n<p><a href=\"https://forums.fast.ai/t/how-can-you-train-your-model-on-large-batches-when-your-gpu-can-t-hold-more-than-a-few-samples/70895/13?u=orangelmx\">https://forums.fast.ai/t/how-can-you-train-your-model-on-large-batches-when-your-gpu-can-t-hold-more-than-a-few-samples/70895/13?u=orangelmx</a></p>",
          "rawMarkdown": "https://forums.fast.ai/t/gradient-accumulation/70968/5?u=orangelmx\n\nhttps://forums.fast.ai/t/how-can-you-train-your-model-on-large-batches-when-your-gpu-can-t-hold-more-than-a-few-samples/70895/13?u=orangelmx"
        },
        {
          "id": 845911,
          "postDate": "2020-05-13T13:41:24.973Z",
          "content": "<p><a href=\"/oscarrangel\">@oscarrangel</a>  Pytorch has native <code>one cycle</code> <a href=\"https://pytorch.org/docs/stable/_modules/torch/optim/lr_scheduler.html#OneCycleLR\">https://pytorch.org/docs/stable/_modules/torch/optim/lr_scheduler.html#OneCycleLR</a> </p>",
          "rawMarkdown": "@oscarrangel  Pytorch has native `one cycle ` https://pytorch.org/docs/stable/_modules/torch/optim/lr_scheduler.html#OneCycleLR ",
          "votes": 2
        },
        {
          "id": 849539,
          "postDate": "2020-05-15T21:30:17.730Z",
          "content": "<p>Do you know if there is a <code>find_lr</code> for Pytorch? I am also using the native <code>OneCycleLR</code>, but struggle for <code>find_lr</code>. I implemented one from <code>https://gist.github.com/NegatioN/07edc229f9d668b2b366528d94500f49</code> but the result is weird... not sure it is due to the model or due to the bug (which I do not see it... yet :( )</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F101755%2Fdfdcb92184b38ed68d392e793479d53e%2F2020-05-15_232839.png?generation=1589578214713980&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Do you know if there is a `find_lr` for Pytorch? I am also using the native `OneCycleLR`, but struggle for `find_lr`. I implemented one from `https://gist.github.com/NegatioN/07edc229f9d668b2b366528d94500f49` but the result is weird... not sure it is due to the model or due to the bug (which I do not see it... yet :( )\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F101755%2Fdfdcb92184b38ed68d392e793479d53e%2F2020-05-15_232839.png?generation=1589578214713980&amp;alt=media)\n"
        },
        {
          "id": 849572,
          "postDate": "2020-05-15T22:53:59.910Z",
          "content": "<p>It seams that the initial value of smoothed loss may be not accounted correctly (also, if you start from untrained model, the loss should initially drop).</p>",
          "rawMarkdown": "It seams that the initial value of smoothed loss may be not accounted correctly (also, if you start from untrained model, the loss should initially drop).",
          "votes": 1
        }
      ]
    },
    {
      "id": 926069,
      "postDate": "2020-07-12T13:24:25.847Z",
      "content": "<p>Hi @lafoss, I was wondering if can help me how to load my model from candence pretrainedmodels, I am still learning.\n~~~\nimport sys\nfrom pathlib import Path</p>\n\n<p>sys.path.append('../input/pytorch-pretrained-models/repository/pretrained-models.pytorch-master')\nimport pretrainedmodels</p>\n\n<p>m = pretrainedmodels.se_resnext50(pretrained='imagenet')\nchildren = list(m.children())\nhead = nn.Sequential(nn.AdaptiveAvgPool2d(1), Flatten(), \n                                  nn.Linear(children[-1].in_features,200))\nmodel = nn.Sequential(nn.Sequential(*children[:-2]), head)\n~~~\nThen tries to download but as we know we dont have internet, I have the model in the directory, but I dont know how to load it,</p>\n\n<p>thanks for the help.</p>",
      "rawMarkdown": "Hi @lafoss, I was wondering if can help me how to load my model from candence pretrainedmodels, I am still learning.\n~~~\nimport sys\nfrom pathlib import Path\n\nsys.path.append('../input/pytorch-pretrained-models/repository/pretrained-models.pytorch-master')\nimport pretrainedmodels\n\nm = pretrainedmodels.se_resnext50(pretrained='imagenet')\nchildren = list(m.children())\nhead = nn.Sequential(nn.AdaptiveAvgPool2d(1), Flatten(), \n                                  nn.Linear(children[-1].in_features,200))\nmodel = nn.Sequential(nn.Sequential(*children[:-2]), head)\n~~~\nThen tries to download but as we know we dont have internet, I have the model in the directory, but I dont know how to load it,\n\nthanks for the help.",
      "votes": -4,
      "replies": [
        {
          "id": 926193,
          "postDate": "2020-07-12T15:12:31.767Z",
          "content": "<p>You don't need to load pretrained weights in inference kernel, just load your model. Try to look how to build the model without pretrained weights.</p>",
          "rawMarkdown": "You don't need to load pretrained weights in inference kernel, just load your model. Try to look how to build the model without pretrained weights.",
          "votes": 3
        },
        {
          "id": 926243,
          "postDate": "2020-07-12T15:40:09.673Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 926292,
          "postDate": "2020-07-12T16:21:44.063Z",
          "content": "<p>yes</p>",
          "rawMarkdown": "yes",
          "votes": 1
        },
        {
          "id": 926330,
          "postDate": "2020-07-12T16:44:23.760Z",
          "content": "<p>I am getting this error,  is looking for the model directory first... \n[Errno 2] No such file or directory: 'models/../input/pandas-models/1_july9_model.pth'</p>\n\n<p>thanks for the help!!!</p>",
          "rawMarkdown": "I am getting this error,  is looking for the model directory first... \n[Errno 2] No such file or directory: 'models/../input/pandas-models/1_july9_model.pth'\n\nthanks for the help!!!"
        },
        {
          "id": 926373,
          "postDate": "2020-07-12T17:09:15.350Z",
          "content": "<p>????\n~~~\nPath('./models').mkdir(exist_ok=True, parents=True)</p>\n\n<p>!cp '../input/panda-models/july9_model.pth' './models/july9_model'</p>\n\n<p>FileNotFoundError: [Errno 2] No such file or directory: 'models/july9_model.pth'\n~~~</p>",
          "rawMarkdown": "????\n~~~\nPath('./models').mkdir(exist_ok=True, parents=True)\n\n!cp '../input/panda-models/july9_model.pth' './models/july9_model'\n\nFileNotFoundError: [Errno 2] No such file or directory: 'models/july9_model.pth'\n~~~"
        },
        {
          "id": 926406,
          "postDate": "2020-07-12T17:37:52.107Z",
          "content": "<p>I was able to solve it changing the name of the model to 'saved'</p>\n\n<p>~~~</p>\n\n<h1>learn.load('saved')</h1>\n\n<p>~~~</p>",
          "rawMarkdown": "I was able to solve it changing the name of the model to 'saved'\n\n~~~\n# learn.load('saved')\n~~~",
          "votes": 1
        }
      ]
    },
    {
      "id": 840802,
      "postDate": "2020-05-10T11:51:20.340Z",
      "content": "<p>densenet121 architechture.</p>",
      "rawMarkdown": "densenet121 architechture.",
      "votes": -1
    },
    {
      "id": 928008,
      "postDate": "2020-07-13T17:25:37.090Z",
      "content": "<p>Hi there my frind @lafoss I was wondering if you can give me a help one more time before this ends, I havent be able to submit anymore</p>\n\n<p>the error is that  \"csv file not found.... \" ????? when I run it on draft everything works fine! as I do the commit and creates the cvs file for the submition,** Submission CSV Not Found**</p>\n\n<p>I was wondering if you see something wrong with the code?</p>\n\n<p>~~~\nTEST_PATH = '/kaggle/input/prostate-cancer-grade-assessment/test_images'</p>\n\n<p>if os.path.exists(TEST_PATH):\n    dls = dBlock.dataloaders(df_test, bs=8)</p>\n\n<pre><code>learn = Learner(dls, get_model())\nlearn.load('saved') \n\ntest_dl = dls.test_dl(df_test)\n_,_, preds = learn.get_preds(dl=test_dl, with_decoded=True)\n\ndf_test[\"isup_grade\"] = preds\nsub = df_test[[\"image_id\",\"isup_grade\"]]\nsub.to_csv('submission.csv', index=False) \n</code></pre>\n\n<p>else:\n    df_train =  df_train.loc[:5]\n    dls = dBlock.dataloaders(df_train, bs=8)</p>\n\n<pre><code>learn = Learner(dls, get_model())\nlearn.load('saved') \n\ntrain_dl = dls.test_dl(df_train)\n_,_, preds = learn.get_preds(dl=train_dl, with_decoded=True)\n\ndf_train[\"isup_grade\"] = preds\nsub = df_train[[\"image_id\",\"isup_grade\"]]\nsub.to_csv('submission.csv', index=Fals\n</code></pre>\n\n<p>~~~\nThanks a lot for all your help, I have learn a lot on this competition thanks to you.</p>",
      "rawMarkdown": "Hi there my frind @lafoss I was wondering if you can give me a help one more time before this ends, I havent be able to submit anymore\n\nthe error is that  \"csv file not found.... \" ????? when I run it on draft everything works fine! as I do the commit and creates the cvs file for the submition,** Submission CSV Not Found**\n\nI was wondering if you see something wrong with the code?\n\n~~~\nTEST_PATH = '/kaggle/input/prostate-cancer-grade-assessment/test_images'\n\nif os.path.exists(TEST_PATH):\n    dls = dBlock.dataloaders(df_test, bs=8)\n\n    learn = Learner(dls, get_model())\n    learn.load('saved') \n\n    test_dl = dls.test_dl(df_test)\n    _,_, preds = learn.get_preds(dl=test_dl, with_decoded=True)\n\n    df_test[\"isup_grade\"] = preds\n    sub = df_test[[\"image_id\",\"isup_grade\"]]\n    sub.to_csv('submission.csv', index=False) \nelse:\n    df_train =  df_train.loc[:5]\n    dls = dBlock.dataloaders(df_train, bs=8)\n\n    learn = Learner(dls, get_model())\n    learn.load('saved') \n\n    train_dl = dls.test_dl(df_train)\n    _,_, preds = learn.get_preds(dl=train_dl, with_decoded=True)\n\n    df_train[\"isup_grade\"] = preds\n    sub = df_train[[\"image_id\",\"isup_grade\"]]\n    sub.to_csv('submission.csv', index=Fals\n~~~\nThanks a lot for all your help, I have learn a lot on this competition thanks to you.",
      "replies": [
        {
          "id": 929580,
          "postDate": "2020-07-14T18:58:53.787Z",
          "content": "<p>I'd think that df_test is not opened or something similar. That is happening is u get some error if the condition in the if statement is True, and the execution of the cell is stopped before u write csv.</p>",
          "rawMarkdown": "I'd think that df_test is not opened or something similar. That is happening is u get some error if the condition in the if statement is True, and the execution of the cell is stopped before u write csv.",
          "votes": 1
        }
      ]
    },
    {
      "id": 882507,
      "postDate": "2020-06-11T21:35:07.237Z",
      "content": "<p><a href=\"/iafoss\">@iafoss</a> can you tell me about your best CV and LB with lowest res tiles?</p>",
      "rawMarkdown": "@iafoss can you tell me about your best CV and LB with lowest res tiles?",
      "replies": [
        {
          "id": 882557,
          "postDate": "2020-06-11T23:13:47.337Z",
          "content": "<p>As written above, on the lowest res 4 fold CV of 0.843 gives ~0.80 LB, and the max LB I got for these images is 0.82. </p>",
          "rawMarkdown": "As written above, on the lowest res 4 fold CV of 0.843 gives ~0.80 LB, and the max LB I got for these images is 0.82. "
        }
      ]
    },
    {
      "id": 867598,
      "postDate": "2020-05-30T12:51:08.773Z",
      "content": "<p>Method: Based on lafoss tiling\nModel: ResNet18\nFold: Stratified\nLeaderboard: 0.84\nValidation: \n<code>\nQWP:0.8556\n[[487  68  16   1   3   0]\n [ 96 331  81   8   6   1]\n [  8  84 128  34  14   1]\n [  1  20  47  65  87  25]\n [  6  17  20  33  77  96]\n [  4   3   9  15  76 136]]\n</code></p>",
      "rawMarkdown": "Method: Based on lafoss tiling\nModel: ResNet18\nFold: Stratified\nLeaderboard: 0.84\nValidation: \n```\nQWP:0.8556\n[[487  68  16   1   3   0]\n [ 96 331  81   8   6   1]\n [  8  84 128  34  14   1]\n [  1  20  47  65  87  25]\n [  6  17  20  33  77  96]\n [  4   3   9  15  76 136]]\n```"
    },
    {
      "id": 861402,
      "postDate": "2020-05-26T03:12:39.043Z",
      "content": "<p><a href=\"/iafoss\">@iafoss</a> is your 0.90 run from a single fold or multiple folds ensembled?</p>",
      "rawMarkdown": "@iafoss is your 0.90 run from a single fold or multiple folds ensembled?",
      "replies": [
        {
          "id": 861466,
          "postDate": "2020-05-26T04:15:04.253Z",
          "content": "<p>As I wrote above, I got 0.90 by a single fold model. Though, 4 folds also give 0.90.</p>",
          "rawMarkdown": "As I wrote above, I got 0.90 by a single fold model. Though, 4 folds also give 0.90.",
          "votes": 1
        },
        {
          "id": 861485,
          "postDate": "2020-05-26T04:31:12.717Z",
          "content": "<p>Right just checking. My 0.90 run was also single fold</p>",
          "rawMarkdown": "Right just checking. My 0.90 run was also single fold",
          "votes": 1
        }
      ]
    },
    {
      "id": 859954,
      "postDate": "2020-05-24T22:08:57.803Z",
      "content": "<p>single fold seresnext50\nCV: .810\nLB: .85</p>",
      "rawMarkdown": "single fold seresnext50\nCV: .810\nLB: .85"
    },
    {
      "id": 854529,
      "postDate": "2020-05-20T05:01:20.027Z",
      "content": "<p>model - Resnext50\nfold - single fold\nCV-QWK : 0.85\nLB-QWK: 0.81</p>",
      "rawMarkdown": "model - Resnext50\nfold - single fold\nCV-QWK : 0.85\nLB-QWK: 0.81",
      "replies": [
        {
          "id": 855368,
          "postDate": "2020-05-20T21:31:22.840Z",
          "content": "<p>Just a few questions, did u use the lowest res layer, and did u check the score based on institution? </p>",
          "rawMarkdown": "Just a few questions, did u use the lowest res layer, and did u check the score based on institution? "
        },
        {
          "id": 855557,
          "postDate": "2020-05-21T03:17:00.737Z",
          "content": "<p><a href=\"/iafoss\">@iafoss</a> i used level 1 images and directly applied your tile method without any rescaling.\nI did not check the CV score based on the institution, but i will check them and post here.</p>",
          "rawMarkdown": "@iafoss i used level 1 images and directly applied your tile method without any rescaling.\nI did not check the CV score based on the institution, but i will check them and post here.",
          "votes": 1
        },
        {
          "id": 855559,
          "postDate": "2020-05-21T03:17:51.800Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 855691,
          "postDate": "2020-05-21T05:47:18.753Z",
          "content": "<p>Update <a href=\"/iafoss\">@iafoss</a> \nmodel - Resnext50\nfold - single fold\nImage level - level 1\nLB-QWK: 0.81\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1554871%2F82d3f9e996df7869781291b6debc43eb%2Foverall.png?generation=1590039929674473&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1554871%2Fdb787fe740b7933959868278c8c3add3%2Fkarolinska.png?generation=1590039930098666&amp;alt=media\" alt=\"\">   <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1554871%2Fc9a6cb6aa1112e3245911113405716f4%2Fradboud.png?generation=1590039931743737&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Update @iafoss \nmodel - Resnext50\nfold - single fold\nImage level - level 1\nLB-QWK: 0.81\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1554871%2F82d3f9e996df7869781291b6debc43eb%2Foverall.png?generation=1590039929674473&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1554871%2Fdb787fe740b7933959868278c8c3add3%2Fkarolinska.png?generation=1590039930098666&amp;alt=media)   ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1554871%2Fc9a6cb6aa1112e3245911113405716f4%2Fradboud.png?generation=1590039931743737&amp;alt=media)\n\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 842197,
      "postDate": "2020-05-11T08:52:39.740Z",
      "content": "<p>Question: X-fold CV -&gt; means it was not single model, right?</p>",
      "rawMarkdown": "Question: X-fold CV -&gt; means it was not single model, right?",
      "replies": [
        {
          "id": 842687,
          "postDate": "2020-05-11T15:01:34.557Z",
          "content": "<p>I'd still consider it as a single model.</p>",
          "rawMarkdown": "I'd still consider it as a single model.",
          "votes": 2
        }
      ]
    },
    {
      "id": 842102,
      "postDate": "2020-05-11T07:39:47.183Z",
      "content": "<p>I train 4-5 models in the same time using built and designated platform - cnvrg.io.\nIt also enables me to get the most accurate model automatically using their conditionals feature.\nI use their community version</p>",
      "rawMarkdown": "I train 4-5 models in the same time using built and designated platform - cnvrg.io.\nIt also enables me to get the most accurate model automatically using their conditionals feature.\nI use their community version"
    },
    {
      "id": 841695,
      "postDate": "2020-05-11T00:59:05.427Z",
      "content": "<p>Impressive!</p>",
      "rawMarkdown": "Impressive!"
    },
    {
      "id": 840268,
      "postDate": "2020-05-09T22:09:37.680Z",
      "content": "<p>Good to Learn </p>",
      "rawMarkdown": "Good to Learn "
    },
    {
      "id": 839789,
      "postDate": "2020-05-09T15:39:38.067Z",
      "content": "<p>Thank you for this. I am new to Kaggle and this is going to be my first competition. \nIf I understand correctly are you not using the 'gleason_scores' for training? \nAre you only using the 'isup_grade' for training?</p>",
      "rawMarkdown": "Thank you for this. I am new to Kaggle and this is going to be my first competition. \nIf I understand correctly are you not using the 'gleason_scores' for training? \nAre you only using the 'isup_grade' for training?",
      "replies": [
        {
          "id": 839819,
          "postDate": "2020-05-09T16:01:19.550Z",
          "content": "<p>I just wrote that I'm not using masks atm because they slow down the training quite a lot, and the boost is insignificant. Other things may be used.</p>",
          "rawMarkdown": "I just wrote that I'm not using masks atm because they slow down the training quite a lot, and the boost is insignificant. Other things may be used.",
          "votes": 1
        },
        {
          "id": 840360,
          "postDate": "2020-05-10T00:28:48.540Z",
          "content": "<p><a href=\"/iafoss\">@iafoss</a> what video card are you using? can we use a different video card that is not Nvidia for training? the Nvidias are too expensive...... I have two 2080 TI with 11 gig ea each one, the highes I I can train is bs 8 with 8 blocks of images of 300x300 each and takes forever to train, and besidese because I am using bz 8 it really doesn't  give good convergence.... I get better results with 32x128x128 up to 0.84 but with 300 and 8 batch size I never go over 77's...... </p>",
          "rawMarkdown": "@iafoss what video card are you using? can we use a different video card that is not Nvidia for training? the Nvidias are too expensive...... I have two 2080 TI with 11 gig ea each one, the highes I I can train is bs 8 with 8 blocks of images of 300x300 each and takes forever to train, and besidese because I am using bz 8 it really doesn't  give good convergence.... I get better results with 32x128x128 up to 0.84 but with 300 and 8 batch size I never go over 77's...... "
        },
        {
          "id": 840396,
          "postDate": "2020-05-10T01:26:24.567Z",
          "content": "<p>I think at the moment only Nvidia GPUs could be used for training since others do not support CUDA. One could also do something with TPUs, but I'm not sure that u can buy them easily. I also have 2x2080Ti. I think you should consider optimizing the way how u use GPUs. To begin with, 300x300 is quite bad resolution to use. You can consider the tile sizes multiple of 32, based on the models u use. There are also other things that u should find out. You can investigate why your models do not converge and try to improve it. It's a competition.  </p>",
          "rawMarkdown": "I think at the moment only Nvidia GPUs could be used for training since others do not support CUDA. One could also do something with TPUs, but I'm not sure that u can buy them easily. I also have 2x2080Ti. I think you should consider optimizing the way how u use GPUs. To begin with, 300x300 is quite bad resolution to use. You can consider the tile sizes multiple of 32, based on the models u use. There are also other things that u should find out. You can investigate why your models do not converge and try to improve it. It's a competition.  "
        },
        {
          "id": 840426,
          "postDate": "2020-05-10T02:45:35.740Z",
          "content": "<p><a href=\"/iafoss\">@iafoss</a> hey I really appreciate your time and posts, I am learning lost from them, so a rule of thumb in image classification use always multiple of 32, got it.\nBy any chance, you have something that will help me understand how to optimize my GPUs? \n\" You can investigate why your models do not converge and try to improve it. It's a competition.\" traying amigo, trying there is so much to learn..... what course dod you recommend?</p>",
          "rawMarkdown": "@iafoss hey I really appreciate your time and posts, I am learning lost from them, so a rule of thumb in image classification use always multiple of 32, got it.\nBy any chance, you have something that will help me understand how to optimize my GPUs? \n\" You can investigate why your models do not converge and try to improve it. It's a competition.\" traying amigo, trying there is so much to learn..... what course dod you recommend?"
        },
        {
          "id": 840428,
          "postDate": "2020-05-10T02:51:50.737Z",
          "content": "<p>In the past I took many deep learning online courses, but the one that really has changed the way I think about deep learning is fast.ai course by Jeremy I took about 2 years ago.</p>",
          "rawMarkdown": "In the past I took many deep learning online courses, but the one that really has changed the way I think about deep learning is fast.ai course by Jeremy I took about 2 years ago.",
          "votes": 3
        },
        {
          "id": 843115,
          "postDate": "2020-05-11T21:03:22.943Z",
          "content": "<p>FAST.AI course is available only as video and notebooks? </p>",
          "rawMarkdown": "FAST.AI course is available only as video and notebooks? "
        },
        {
          "id": 843135,
          "postDate": "2020-05-11T21:19:34.247Z",
          "content": "<p>Yes. Presumably one could take it in person as well, but it will cost $ $ $, and you will need to stay at the place for several month. Also there was an option to watch the lectures during the course in real time.</p>",
          "rawMarkdown": "Yes. Presumably one could take it in person as well, but it will cost $ $ $, and you will need to stay at the place for several month. Also there was an option to watch the lectures during the course in real time."
        }
      ]
    },
    {
      "id": 839187,
      "postDate": "2020-05-09T06:28:06.307Z",
      "content": "<p>👍 </p>",
      "rawMarkdown": "👍 "
    },
    {
      "id": 839117,
      "postDate": "2020-05-09T05:04:18.533Z",
      "content": "<p>I am working on it</p>",
      "rawMarkdown": "I am working on it"
    },
    {
      "id": 838973,
      "postDate": "2020-05-09T00:56:49.847Z",
      "content": "<p><a href=\"/iafoss\">@iafoss</a> Hi there,</p>\n\n<p>1.- I was wondering where do you train, on your personal computer? or kaggle?\n2.- What is number of tiles per image and what dimension (sz)?\n3.- What bs are you using?\n4.- and for how many epochs ?</p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "@iafoss Hi there,\n\n1.- I was wondering where do you train, on your personal computer? or kaggle?\n2.- What is number of tiles per image and what dimension (sz)?\n3.- What bs are you using?\n4.- and for how many epochs ?\n\nThanks!",
      "replies": [
        {
          "id": 838995,
          "postDate": "2020-05-09T01:43:59.140Z",
          "content": "<p>I do almost all work on my personal computer because kaggle has introduced quite limited GPU quota. Based on my previous experience, at least 100-150 kaggle GPU hours per week are needed to fully participate in a typical deep learning competition. Also GPU kernels have very weak CPUs, so the real speed drops down quite a lot (like at kaggle my public kernel takes ~10 min per epoch, while on my PC with 2 GPUs - ~1 min).  Given these limitations, I don't think that it is quite feasible to train something with intermediate resolution using kaggle: training my intermediate resolution model took ~9 hours locally, at kaggle it would be ~90 hours. It's quite sad that computational resources limit participants now.  A year ago, when kaggle didn't impose any GPU limits, I got 1 gold and 1 silver medal just using kaggle GPUs.</p>\n\n<p>Regarding other questions, I would like not to provide too many details on my intermediate resolution models. For low resolution results, I used 12x128x128 setup, as in my public kernel, but trained for longer and used other tricks.</p>",
          "rawMarkdown": "I do almost all work on my personal computer because kaggle has introduced quite limited GPU quota. Based on my previous experience, at least 100-150 kaggle GPU hours per week are needed to fully participate in a typical deep learning competition. Also GPU kernels have very weak CPUs, so the real speed drops down quite a lot (like at kaggle my public kernel takes ~10 min per epoch, while on my PC with 2 GPUs - ~1 min).  Given these limitations, I don't think that it is quite feasible to train something with intermediate resolution using kaggle: training my intermediate resolution model took ~9 hours locally, at kaggle it would be ~90 hours. It's quite sad that computational resources limit participants now.  A year ago, when kaggle didn't impose any GPU limits, I got 1 gold and 1 silver medal just using kaggle GPUs.\n\nRegarding other questions, I would like not to provide too many details on my intermediate resolution models. For low resolution results, I used 12x128x128 setup, as in my public kernel, but trained for longer and used other tricks.",
          "votes": 5
        },
        {
          "id": 839050,
          "postDate": "2020-05-09T03:05:39.987Z",
          "content": "<p>Thank you <a href=\"/iafoss\">@iafoss</a> learning from you a lot, I am still trying to understand so anything.... hehe I guess it takes years, I only been doing DeepLearning for a year or so.</p>\n\n<p>But are you saving model with </p>\n\n<p>learn.model = torch.nn.DataParallel(learn.model) </p>\n\n<p>because I think there is a problem training with multiple GPUs and doing inference on Kaggle that has one GPU, right ?</p>",
          "rawMarkdown": "Thank you @iafoss learning from you a lot, I am still trying to understand so anything.... hehe I guess it takes years, I only been doing DeepLearning for a year or so.\n\nBut are you saving model with \n\nlearn.model = torch.nn.DataParallel(learn.model) \n\n because I think there is a problem training with multiple GPUs and doing inference on Kaggle that has one GPU, right ?"
        },
        {
          "id": 839057,
          "postDate": "2020-05-09T03:20:32.817Z",
          "content": "<p>There are no problems, just use <code>.module</code> when save: <code>torch.save(learn.model.module.state_dict(),os.path.join(OUT,f'{fname}_{fold}.pth'))</code>. Also, standard fast.ai save function has another issue: it saves gradients, etc., so the resulting files are 3-4 times larger than files with just weights (at least it was the case in one of fast.ai versions I used about a half year ago). \nThere are also other ways to load models saved with nn.DataParallel, like creating nn.DataParallel model at kaggle and remapping to a single GPU or manual modification of the names in the state dict, but saving all models without nn.DataParallel just makes your life simpler.</p>",
          "rawMarkdown": "There are no problems, just use `.module` when save: `torch.save(learn.model.module.state_dict(),os.path.join(OUT,f'{fname}_{fold}.pth'))`. Also, standard fast.ai save function has another issue: it saves gradients, etc., so the resulting files are 3-4 times larger than files with just weights (at least it was the case in one of fast.ai versions I used about a half year ago). \nThere are also other ways to load models saved with nn.DataParallel, like creating nn.DataParallel model at kaggle and remapping to a single GPU or manual modification of the names in the state dict, but saving all models without nn.DataParallel just makes your life simpler.",
          "votes": 2
        },
        {
          "id": 839062,
          "postDate": "2020-05-09T03:26:35.570Z",
          "content": "<p>Excellent explanation, here is a little bit more on the last point if you train model locally on multiple gpu and want to load in local kernel you can do following: </p>\n\n<p>```</p>\n\n<h1>here ['model'] refer to single dictionary</h1>\n\n<p>state_dict = torch.load('your_model.pth')['model']\nmodel.load_state_dict(state_dict)\n```</p>\n\n<p>In the case of excellent <a href=\"/iafoss\">@iafoss</a> inference kernel it should approx. look like this (copy, pasted, need to test) :</p>\n\n<p>```</p>\n\n<h1>if you trained locally using nn.Dataparallel and want to load the model in kaggle inference kernel</h1>\n\n<p>models = []\nfor path in MODELS:\n    state_dict = torch.load(path,map_location=torch.device('cpu'))['model']\n    model = Model()\n    model.load_state_dict(state_dict)\n    model.float()\n    model.eval()\n    model.cuda()\n    models.append(model)\n```</p>",
          "rawMarkdown": "Excellent explanation, here is a little bit more on the last point if you train model locally on multiple gpu and want to load in local kernel you can do following: \n\n```\n#here ['model'] refer to single dictionary \nstate_dict = torch.load('your_model.pth')['model']\nmodel.load_state_dict(state_dict)\n```\n\nIn the case of excellent @iafoss inference kernel it should approx. look like this (copy, pasted, need to test) :\n\n```\n#if you trained locally using nn.Dataparallel and want to load the model in kaggle inference kernel\nmodels = []\nfor path in MODELS:\n    state_dict = torch.load(path,map_location=torch.device('cpu'))['model']\n    model = Model()\n    model.load_state_dict(state_dict)\n    model.float()\n    model.eval()\n    model.cuda()\n    models.append(model)\n```\n",
          "votes": 2
        },
        {
          "id": 839171,
          "postDate": "2020-05-09T06:03:58.200Z",
          "content": "<p>Why do you do model.float() here? </p>",
          "rawMarkdown": "Why do you do model.float() here? "
        },
        {
          "id": 839686,
          "postDate": "2020-05-09T14:32:28.313Z",
          "content": "<p><a href=\"/iafoss\">@iafoss</a> thanks a lot for the code sniped will save me lost of headaches for this and other competitions.</p>",
          "rawMarkdown": "@iafoss thanks a lot for the code sniped will save me lost of headaches for this and other competitions.",
          "votes": 1
        },
        {
          "id": 839689,
          "postDate": "2020-05-09T14:33:49.053Z",
          "content": "<p>Thanks <a href=\"/drhabib\">@drhabib</a> for the post.</p>",
          "rawMarkdown": "Thanks @drhabib for the post.",
          "votes": 1
        },
        {
          "id": 839777,
          "postDate": "2020-05-09T15:30:03.540Z",
          "content": "<p>Calling model.float() brings the model back to full precision (fp32). I assume the model was cast to fp16 somewhere previously, either with learn.tofp16() or model.half(). </p>",
          "rawMarkdown": "Calling model.float() brings the model back to full precision (fp32). I assume the model was cast to fp16 somewhere previously, either with learn.tofp16() or model.half(). ",
          "votes": 1
        },
        {
          "id": 840049,
          "postDate": "2020-05-09T18:31:35.950Z",
          "content": "<p><a href=\"/iafoss\">@iafoss</a> <a href=\"/interneuron\">@interneuron</a> but what about it I train a model on 300x300 on my local computer, with kaggle limitations would we be able to run inference with 300x300 without running out of GPU memory?</p>",
          "rawMarkdown": "@iafoss @interneuron but what about it I train a model on 300x300 on my local computer, with kaggle limitations would we be able to run inference with 300x300 without running out of GPU memory?"
        },
        {
          "id": 840067,
          "postDate": "2020-05-09T18:48:32.237Z",
          "content": "<p>u can just use different bs at kaggle</p>",
          "rawMarkdown": "u can just use different bs at kaggle"
        },
        {
          "id": 840111,
          "postDate": "2020-05-09T19:18:09.023Z",
          "content": "<p><a href=\"/iafoss\">@iafoss</a> but when I use bz smaller than 16 I don't get good generalization, I understand because the gradient descent cant propagates enough.</p>\n\n<p>Thanks!</p>",
          "rawMarkdown": "@iafoss but when I use bz smaller than 16 I don't get good generalization, I understand because the gradient descent cant propagates enough.\n\nThanks!"
        },
        {
          "id": 840122,
          "postDate": "2020-05-09T19:31:10.200Z",
          "content": "<p>I thought u are talking about inference only, so u could use whatever fits kaggle GPUs. It is not related to bs during training.</p>",
          "rawMarkdown": "I thought u are talking about inference only, so u could use whatever fits kaggle GPUs. It is not related to bs during training."
        },
        {
          "id": 840189,
          "postDate": "2020-05-09T20:21:15.167Z",
          "content": "<p><a href=\"/iafoss\">@iafoss</a> oops... you are right, I see what you mean, thanks.</p>",
          "rawMarkdown": "@iafoss oops... you are right, I see what you mean, thanks."
        }
      ]
    },
    {
      "id": 837672,
      "postDate": "2020-05-08T00:15:39.953Z",
      "content": "<p><a href=\"/iafoss\">@iafoss</a> I imagine you load images during training, so how do you do that in a fast way? I found that switching to 12x256x256 takes 1:30 to load the data from individually pickled files per one training cycle, which I feel is way too long</p>",
      "rawMarkdown": "@iafoss I imagine you load images during training, so how do you do that in a fast way? I found that switching to 12x256x256 takes 1:30 to load the data from individually pickled files per one training cycle, which I feel is way too long",
      "replies": [
        {
          "id": 837694,
          "postDate": "2020-05-08T01:13:48.783Z",
          "content": "<p>For me loading from individual pngs works good enough even on intermediate images. Not sure if its because M2 drive, but in deepfake competition with ~2M imgs in a folder reading from a regular hardrive during training was really nearly impossible. Though, I'm not sure how to deal with full res images yet: loading takes too long(</p>",
          "rawMarkdown": "For me loading from individual pngs works good enough even on intermediate images. Not sure if its because M2 drive, but in deepfake competition with ~2M imgs in a folder reading from a regular hardrive during training was really nearly impossible. Though, I'm not sure how to deal with full res images yet: loading takes too long(",
          "votes": 1
        },
        {
          "id": 838953,
          "postDate": "2020-05-09T00:17:25.977Z",
          "content": "<p>Have you tried using pickle? In my experiments, using png is much slower than pickle files, although I think a lot of loading time during training is alleviated by pytorch's inherent asynchronous execution. </p>",
          "rawMarkdown": "Have you tried using pickle? In my experiments, using png is much slower than pickle files, although I think a lot of loading time during training is alleviated by pytorch's inherent asynchronous execution. "
        },
        {
          "id": 838957,
          "postDate": "2020-05-09T00:31:46.850Z",
          "content": "<p>I haven't tried. If you are referring to saving arrays with images into pickle without encoding into image format, the size of produced files may be quite large, though, loading may be much faster. Also, if loading takes too long, you can just go bigger model.</p>",
          "rawMarkdown": "I haven't tried. If you are referring to saving arrays with images into pickle without encoding into image format, the size of produced files may be quite large, though, loading may be much faster. Also, if loading takes too long, you can just go bigger model."
        },
        {
          "id": 838966,
          "postDate": "2020-05-09T00:43:17.517Z",
          "content": "<p>I save the data in uint8 so it's around 2.4 mb per 12x256x256x3, which is not bad. And I think it is much faster to load</p>",
          "rawMarkdown": "I save the data in uint8 so it's around 2.4 mb per 12x256x256x3, which is not bad. And I think it is much faster to load"
        },
        {
          "id": 854932,
          "postDate": "2020-05-20T13:03:44.380Z",
          "content": "<p>Hi <a href=\"/shujun717\">@shujun717</a> ,</p>\n\n<p>I had a similar problem when tring to resize the images during inference for EfficientNetB0.\ncv2.resize just wasn't doing the job fast enough. However, you can load the images/tiles in the original shape, and then use PyTorch nn.functional.interpolate(x, mode=\"bilinear\") to get some fast resizing of your tensors :) </p>\n\n<p>Hope this works</p>",
          "rawMarkdown": "Hi @shujun717 ,\n\nI had a similar problem when tring to resize the images during inference for EfficientNetB0.\ncv2.resize just wasn't doing the job fast enough. However, you can load the images/tiles in the original shape, and then use PyTorch nn.functional.interpolate(x, mode=\"bilinear\") to get some fast resizing of your tensors :) \n\nHope this works"
        }
      ]
    },
    {
      "id": 835505,
      "postDate": "2020-05-06T10:05:16.590Z",
      "content": "<p>well done, 95+ on horizon</p>",
      "rawMarkdown": "well done, 95+ on horizon",
      "replies": [
        {
          "id": 835940,
          "postDate": "2020-05-06T15:58:41.777Z",
          "content": "<p>Thanks. My expectation is that at the end of the competition the gold range would correspond to about 0.95-0.96.</p>",
          "rawMarkdown": "Thanks. My expectation is that at the end of the competition the gold range would correspond to about 0.95-0.96.",
          "votes": 1
        }
      ]
    },
    {
      "id": 835372,
      "postDate": "2020-05-06T08:24:01.270Z",
      "content": "<p>In my local experiments, my best single fold CV with TTA and checkpoint averaging can get up to 0.833 but somehow it does not transfer to lb (~0.75 with TTA and checkpoint averaging vs 0.78 without TTA). Maybe I have a bug in my code...</p>",
      "rawMarkdown": "In my local experiments, my best single fold CV with TTA and checkpoint averaging can get up to 0.833 but somehow it does not transfer to lb (~0.75 with TTA and checkpoint averaging vs 0.78 without TTA). Maybe I have a bug in my code...",
      "replies": [
        {
          "id": 836060,
          "postDate": "2020-05-06T17:30:52.517Z",
          "content": "<p>Also, LB score may not be stable, I don't know. I saw some strange stuff as well: 0.77 CV -&gt; 0.79 LB, 0.84 CV -&gt; 0.80 LB, 0.89 CV (single fold) -&gt; 0.90 LB.</p>",
          "rawMarkdown": "Also, LB score may not be stable, I don't know. I saw some strange stuff as well: 0.77 CV -&gt; 0.79 LB, 0.84 CV -&gt; 0.80 LB, 0.89 CV (single fold) -&gt; 0.90 LB."
        },
        {
          "id": 836180,
          "postDate": "2020-05-06T19:40:44.100Z",
          "content": "<p>Yeah so far it's quite unstable on low resolution images. I was quite surprised your public kernels got .79 lb since the val is around .77~78</p>",
          "rawMarkdown": "Yeah so far it's quite unstable on low resolution images. I was quite surprised your public kernels got .79 lb since the val is around .77~78"
        }
      ]
    },
    {
      "id": 835362,
      "postDate": "2020-05-06T08:19:05.180Z",
      "content": "<p>Densenet121 based network on 12x128x128 tiles: 0.81~0.82 cv, 4 fold lb 0.78 on lowest resolution images. Impressive work as always btw.</p>",
      "rawMarkdown": "Densenet121 based network on 12x128x128 tiles: 0.81~0.82 cv, 4 fold lb 0.78 on lowest resolution images. Impressive work as always btw.",
      "replies": [
        {
          "id": 835936,
          "postDate": "2020-05-06T15:57:34.333Z",
          "content": "<p>Thanks</p>",
          "rawMarkdown": "Thanks"
        }
      ]
    },
    {
      "id": 835356,
      "postDate": "2020-05-06T08:15:26.637Z",
      "content": "<p>Thanks for sharing.</p>\n\n<p>I have one question, did you try regression, instead of classification?\nAs I learned from APTOS2019, the top solution used regression,\nwhat do you think?</p>\n\n<p>Thanks again</p>",
      "rawMarkdown": "Thanks for sharing.\n\nI have one question, did you try regression, instead of classification?\nAs I learned from APTOS2019, the top solution used regression,\nwhat do you think?\n\nThanks again",
      "replies": [
        {
          "id": 835503,
          "postDate": "2020-05-06T10:04:32.523Z",
          "content": "<p>regression is the way to go. </p>\n\n<p>PS: In APTOS, I placed 46/2931 using softmax classification and everyone above me used regression. Use regression (or combination) it better fits the metric. </p>",
          "rawMarkdown": "regression is the way to go. \n\nPS: In APTOS, I placed 46/2931 using softmax classification and everyone above me used regression. Use regression (or combination) it better fits the metric. ",
          "votes": 6
        },
        {
          "id": 835932,
          "postDate": "2020-05-06T15:55:36.653Z",
          "content": "<p>I do use regression. Actually, it can be quite quickly checked from the confusion matrix if one posted it:\n<code>\n0.8150716393433914\n[[2446  307   38   24   41   17]\n [ 395 1790  297   61   59   14]\n [  83  367  675  130   55   31]\n [  42   80  149  656  164  135]\n [  84   74   74  145  712  156]\n [  51   28   46   90  150  850]]\n</code>\n<code>\n0.8420650490449149\n[[2246  485   98   37    5    2]\n [ 426 1555  508   99   24    4]\n [  52  391  557  250   88    3]\n [  24   63  217  435  400   87]\n [  31   66  109  195  463  381]\n [  17   44   37   93  261  763]]\n</code>\nThough, I got quite similar LB for those two experiments(do not really know why), but for classification (top one) the false predictions are more scattered, while for a different loss they are more localized near true predictions.</p>",
          "rawMarkdown": "I do use regression. Actually, it can be quite quickly checked from the confusion matrix if one posted it:\n```\n0.8150716393433914\n[[2446  307   38   24   41   17]\n [ 395 1790  297   61   59   14]\n [  83  367  675  130   55   31]\n [  42   80  149  656  164  135]\n [  84   74   74  145  712  156]\n [  51   28   46   90  150  850]]\n```\n```\n0.8420650490449149\n[[2246  485   98   37    5    2]\n [ 426 1555  508   99   24    4]\n [  52  391  557  250   88    3]\n [  24   63  217  435  400   87]\n [  31   66  109  195  463  381]\n [  17   44   37   93  261  763]]\n```\nThough, I got quite similar LB for those two experiments(do not really know why), but for classification (top one) the false predictions are more scattered, while for a different loss they are more localized near true predictions.",
          "votes": 2
        },
        {
          "id": 840559,
          "postDate": "2020-05-10T06:22:20.133Z",
          "content": "<p>I also noticed that regression gives higher local CV than classification, but the public LB does not reflect this for me.</p>",
          "rawMarkdown": "I also noticed that regression gives higher local CV than classification, but the public LB does not reflect this for me."
        },
        {
          "id": 849685,
          "postDate": "2020-05-16T01:32:31.937Z",
          "content": "<p>Do all of you train on isup or on the gleason score. I assume gleason should give better results as there is more information in it. But each time I try I get equivalent or worse results than training on isup directly. (I'm only now moving from classification to regression and will see if things are different there.)</p>",
          "rawMarkdown": "Do all of you train on isup or on the gleason score. I assume gleason should give better results as there is more information in it. But each time I try I get equivalent or worse results than training on isup directly. (I'm only now moving from classification to regression and will see if things are different there.)",
          "votes": 1
        }
      ]
    },
    {
      "id": 903613,
      "postDate": "2020-06-27T01:50:29.113Z",
      "content": "<p>Method: Based on lafoss tiling + self modified trick\nModel: Efficientnet-B4 or  Efficientnet-B1\nFold: Stratified fold 0 out of 8 folds\nLeaderboard: 0.85</p>",
      "rawMarkdown": "Method: Based on lafoss tiling + self modified trick\nModel: Efficientnet-B4 or  Efficientnet-B1\nFold: Stratified fold 0 out of 8 folds\nLeaderboard: 0.85\n\n",
      "isDeleted": true
    },
    {
      "id": 893385,
      "postDate": "2020-06-19T15:11:17.780Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 836153,
      "author_name": "Ian Pan",
      "author_url": "",
      "post_date": "2020-05-06T19:03:12.863000",
      "content": "<p>Great work. My hypothesis is that overfitting more easily occurs at lower magnifications. In my own experience, I have higher CV using level 1 vs level 0 (0.91 vs, 0.88), but LB is essentially the same (0.87). If you're using the same patch size for levels 1 and 2, then the patches for level 1 will have a much higher proportion of tissue vs. background. That might be one reason for CV-LB discrepancy. </p>",
      "votes": 10,
      "replies": [
        {
          "id": 836170,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-05-06T19:26:32.747000",
          "content": "<p>Thanks, quite interesting observation. I use different tile size, but the the fraction of background, indeed, may be different for my current setups. So the hypothesis may be that the average tissue area in train and test sets are different that creates a gap for particular setups.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 836177,
          "author_name": "Shujun",
          "author_url": "",
          "post_date": "2020-05-06T19:38:04.700000",
          "content": "<p>One thing that might be interesting to try, is to choose the tiles in a stochastic fashion, which should result in a more robust model. However, then the tile generation would have to be done during training</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 836247,
          "author_name": "Ian Pan",
          "author_url": "",
          "post_date": "2020-05-06T20:44:31.277000",
          "content": "<p>You could sample the tiles weighted by number of tissue pixels. That will add some stochasticity to the training process without interfering too much with learning by including many non-informative tiles. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 837966,
          "author_name": "Benjamin Dubreu",
          "author_url": "",
          "post_date": "2020-05-08T07:37:32.540000",
          "content": "<p><a href=\"/vaillant\">@vaillant</a> \"You could sample the tiles weighted by number of tissue pixels\" I believe that's what lafoss' approach does already. Well, more precisely, it uses a proxy for this: it tries to avoid white pixels as much as possible (by sorting the tiles according to the sum of their pixel values and taking the first N).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 838462,
          "author_name": "Ian Pan",
          "author_url": "",
          "post_date": "2020-05-08T15:37:48.287000",
          "content": "<p>I was referring to a stochastic sampling during training where each tile is sampled with probability proportional to the amount of tissue pixels for more regularization vs the deterministic approach. </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 838641,
          "author_name": "Benjamin Dubreu",
          "author_url": "",
          "post_date": "2020-05-08T17:34:15.987000",
          "content": "<p>owh ok I hadn't understood, thanks for the clarification. Quite the nice trick indeed !</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 915890,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-07-05T07:18:10.837000",
          "content": "<p><a href=\"/vaillant\">@vaillant</a> first of all congrats for being at top.\nTO your post stochastic sampling weighted by Tissues ,\n1) How would u decided the weight of tiles\n2)if  few set of tiles are  weighted less compared to other tile but carrying more cancerous region in them which should be decider for isup grade in terms of amount of abnormal cells region in overall wsi slide then  wont we miss classify that wsi to lower isup grade ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 916762,
          "author_name": "Ian Pan",
          "author_url": "",
          "post_date": "2020-07-06T02:22:38.490000",
          "content": "<p>I don't fully understand your post, but I never tried doing this. I don't really think it will make a difference. </p>\n\n<p>If you want to try it:\n1. Take all the tiles in an image and compute mean pixel value (as Iafoss does).\n2. Normalize these values so that they sum to 1. \n3. Sample each tile based on the normalized values using np.random.choice or something similar.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 842237,
      "author_name": "hirune924",
      "author_url": "",
      "post_date": "2020-05-11T09:14:51.657000",
      "content": "<p>model: se-resnet50\nfold: single fold \ncv-qwk: 0.894\nlb-qwk: 0.87</p>",
      "votes": 7,
      "replies": [
        {
          "id": 847231,
          "author_name": "wayfarer",
          "author_url": "",
          "post_date": "2020-05-14T08:42:21.723000",
          "content": "<p>hi, great results, classification or regression?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 847502,
          "author_name": "hirune924",
          "author_url": "",
          "post_date": "2020-05-14T12:52:24.627000",
          "content": "<p>This is the result of regression model. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 848149,
          "author_name": "DeepLearner",
          "author_url": "",
          "post_date": "2020-05-14T19:11:07.493000",
          "content": "<p>May I know how many epoches you used to get this kappa score? At that epoch what is the training kappa value?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 848233,
          "author_name": "TheStoneCa",
          "author_url": "",
          "post_date": "2020-05-14T20:07:15.743000",
          "content": "<p><a href=\"/hirune924\">@hirune924</a>, how you go about to do a single fold ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 848418,
          "author_name": "hirune924",
          "author_url": "",
          "post_date": "2020-05-15T00:46:08.093000",
          "content": "<p><a href=\"/bethewinner\">@bethewinner</a> I spent 60 epochs training and I'm not monitoring QWK during train. Instead, the RMSE is 0.51.\n<a href=\"/thestoneca\">@thestoneca</a> you can see the code here(<a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/145296#828631\">https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/145296#828631</a>  ).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 848423,
          "author_name": "TheStoneCa",
          "author_url": "",
          "post_date": "2020-05-15T00:50:35.307000",
          "content": "<p><a href=\"/hirune924\">@hirune924</a> thanks but I dont know if I am missing something.... </p>\n\n<p>df = pd.read_csv(os.path.join(data_dir,'train.csv'))\nskf = StratifiedKFold(n_splits=5, shuffle = True, random_state = 2020)\nfor fold, (train_index, val_index) in enumerate(skf.split(df.values, df['isup_grade'])):\n    df.loc[val_index, 'fold'] = int(fold)</p>\n\n<p>here you are creating 5 folds, and StratifiedKFold the least you can create is 2.</p>\n\n<p>Then my question here is how do you create one fold..... </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 848427,
          "author_name": "hirune924",
          "author_url": "",
          "post_date": "2020-05-15T00:56:02.053000",
          "content": "<p><a href=\"/thestoneca\">@thestoneca</a> I think it's called a single fold when you use only one fold of what you split as a 5fold. It's a single fold divided into 80:20</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 848438,
          "author_name": "TheStoneCa",
          "author_url": "",
          "post_date": "2020-05-15T01:10:56.407000",
          "content": "<p><a href=\"/hirune924\">@hirune924</a> meaning that of the 5 folds create you only you one? sorry, language barrier.....</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 848442,
          "author_name": "hirune924",
          "author_url": "",
          "post_date": "2020-05-15T01:17:40.877000",
          "content": "<p><a href=\"/thestoneca\">@thestoneca</a> I just split all data up so that train:val=80:20</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 849040,
          "author_name": "wayfarer",
          "author_url": "",
          "post_date": "2020-05-15T13:02:42.113000",
          "content": "<p><a href=\"/hirune924\">@hirune924</a>, how do you handle the images? like tiles and feed them 1 by 1? thanks.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 915887,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-07-05T07:14:04.010000",
          "content": "<p><a href=\"/hirune924\">@hirune924</a>  i tried enough of regression in beginning but was not getting better than 0.84 .\nWhat could be conributing to the score.\n1) Tiling method ?public or your own\n2) Some customization in loss used for regression ?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 840058,
      "author_name": "Shujun",
      "author_url": "",
      "post_date": "2020-05-09T18:38:53.680000",
      "content": "<p>A little update: resnet34 based architecture, 0.88 single fold val -&gt; 0.85 lb. 4 fold cv 0.87 -&gt; lb 0.87. It seems that the choice of architecture (or backbone) may not be the most important</p>",
      "votes": 5,
      "replies": [
        {
          "id": 841207,
          "author_name": "wayfarer",
          "author_url": "",
          "post_date": "2020-05-10T17:09:36.167000",
          "content": "<p>Hello, great results. Are you using tiles as Iafoss?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 841362,
          "author_name": "Shujun",
          "author_url": "",
          "post_date": "2020-05-10T18:51:08.023000",
          "content": "<p>Yes</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 841411,
          "author_name": "Benjamin Dubreu",
          "author_url": "",
          "post_date": "2020-05-10T19:27:26.400000",
          "content": "<p>Hi <a href=\"/shujun717\">@shujun717</a> ! getting CV and LB on par is quite curious ^^ Any idea how you achieved that feat ? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 841444,
          "author_name": "Shujun",
          "author_url": "",
          "post_date": "2020-05-10T19:46:14.713000",
          "content": "<p>You should ask <a href=\"/iafoss\">@iafoss</a> that. His public kernel had val score of 0.77~0.78 yet the lb score is 0.79, which was quite a surprise to me. My latest 0.88 submission had a 4 fold cv score of 0.874, but iafoss achieved 0.90 with a single model with val score of 0.88. In any case, the most immediate thing that comes to my mind is that I used a very simple backbone ~ resnet34, since I have tried efficientnets which did not work well for me</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 841484,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-05-10T20:12:43.850000",
          "content": "<p>My impression so far is that LB mostly correlates with karolinska CV. However, the training data seems to be noisy( and one needs to find a sweet spot between LB and CV(</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 842939,
          "author_name": "wayfarer",
          "author_url": "",
          "post_date": "2020-05-11T17:57:39.877000",
          "content": "<p><a href=\"/shujun717\">@shujun717</a> sorry, are u using regression or classification?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 836585,
      "author_name": "Ian Pan",
      "author_url": "",
      "post_date": "2020-05-07T04:47:37.040000",
      "content": "<p><a href=\"/iafoss\">@iafoss</a> what is the breakdown of QWK between the 2 institutions? I noticed that when my CV went from 0.88 to 0.91 with no change in LB, there was only CV improvement for Radboud data - Karolinska data performance was the same. </p>",
      "votes": 6,
      "replies": [
        {
          "id": 837125,
          "author_name": "DatNT",
          "author_url": "",
          "post_date": "2020-05-07T14:38:10.977000",
          "content": "<p>in my case, the score for karolinska subset was way worse than the one for radboud too</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 837223,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-05-07T16:08:37.267000",
          "content": "<p>It looks like this:\n<code>\nkarolinska 0.909\n[[433  41   4   2   1   0]\n [ 34 382  55   1   1   0]\n [  0  49  71  36   6   0]\n [  0   5  20  40  19   1]\n [  0   6   6  12  63  23]\n [  0   1   0   3  17  44]]\nradboud 0.8403027122576369\n[[182  34  15   5   2   0]\n [ 14 115  41  10   1   0]\n [  4  37  72  48  11   1]\n [ 10   8  21  96  72  14]\n [  2   7  11  32  93  56]\n [  2   5   9  10  61 152]]\n</code>\nMeanwhile my low res models look like this:\n<code>\nkarolinska 0.805\n[[1609  254   45    6   10    1]\n [ 300 1220  237   39   15    3]\n [  42  274  227   97   24    4]\n [   9   32   92   86   73   25]\n [  20   53   46   62  174  126]\n [  11    5   14   24   52  145]]\nradboud 0.829\n[[745 122  37  27  14   3]\n [100 405 217  65  13   2]\n [ 17 117 304 181  47   7]\n [ 16  44 129 267 346 107]\n [ 14  24  47  98 273 308]\n [  7  19  19  58 219 642]]\n</code>\nSo, it seems that <strong>public LB may account only for karolinska data</strong></p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 837244,
          "author_name": "Ian Pan",
          "author_url": "",
          "post_date": "2020-05-07T16:28:51.007000",
          "content": "<p>That is my suspicion as well. </p>\n\n<p>For LB 0.87, I have this for CV:</p>\n\n<p><code>\nKarolinska \n0.88638\nRadboud\n0.90636\nOverall\n0.90936\n</code></p>\n\n<p>The question is whether the entire test set is skewed towards Karolinska or if the public test set was randomly (or intentionally) sampled that way. </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 837267,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-05-07T16:43:28.473000",
          "content": "<p>It would be a quite bad move if organizers split test into public/private based on the provider( like public= karolinska and private=radboud</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 837316,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-05-07T17:29:07.930000",
          "content": "<p>in Bengali Private dataset was containing graphemes which were not present in Public Test .... So I will be not surprised if private is divided based on institution ...</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 837381,
          "author_name": "Ian Pan",
          "author_url": "",
          "post_date": "2020-05-07T18:18:42.593000",
          "content": "<p>It doesn't really make sense for this competition because then we can just optimize for one institution's data, and the resulting model would not be as generalizable. Note that there are no new institutions in the test set. There seems to be some evidence that at least the public test set is skewed towards Karolinska. If the entire test set has a similar distribution to the training set (roughly 50/50), that would mean the private test set is skewed towards Radboud. </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 838212,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-05-08T11:37:26.827000",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> and you reminded that shake up the competition, a nightmare 😭 \nI hope it wouldn't be the same for this comp too. 🤕 </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 838426,
          "author_name": "Wouter Bulten",
          "author_url": "",
          "post_date": "2020-05-08T15:04:52.010000",
          "content": "<p>Great discussions. Both public and private test sets contain images from both institutions. As someone already mentioned, from the patient/medical perspective the goal is to build a model that can work on datasets from multiple institutions/labs. That's also the aim of the <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/overview/miccai-2020\">paper/workshop</a>.</p>\n\n<p>Please take the data description into account regarding the labels:</p>\n\n<blockquote>\n  <p>The labels are imperfect. This is a challenging area of pathology and even experts in the field with years of experience do not always agree on how to interpret a slide. This will make training models more difficult, but increases the potential medical value of having a strong model to provide consistent ratings. All of the private test set images and most of the public test set images were graded by multiple pathologists, but this was not feasible for the training set. You can find additional details about how consistently the pathologist's labels matched <a href=\"https://zenodo.org/record/3715938#.XrVzk6gzZPa\">here</a>.</p>\n</blockquote>\n\n<p>Also, the <a href=\"https://zenodo.org/record/3715938#.XrVzk6gzZPa\">challenge document</a> contains a lot of background info on how the data was collected.</p>",
          "votes": 20,
          "replies": []
        },
        {
          "id": 838447,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-05-08T15:29:03.710000",
          "content": "<p>Thanks for the reference.</p>\n\n<p>I went quickly thru challenge document and highlighted important information just in case if you dont want to read 13 pages =) </p>\n\n<p><strong>TL DR:</strong>\n```\nTraining set: +/- 11.000 cases\nPublic test set: +/- 400 cases (with expert gradings)\nPrivate test set: +/- 400 cases (with expert gradings)</p>\n\n<p>```</p>\n\n<p><strong>d) Mention further important characteristics of the training, validation and test cases (e.g. class distribution in\nclassification tasks chosen according to real-world distribution vs. equal class distribution) and justify the choice.</strong></p>\n\n<p><code>\nCases were sampled based on the Gleason grade group.\n</code></p>\n\n<p>The test set was graded independently by three pathologists who are subspecialized in\nuropathology. A final consensus score was determined in three rounds. The non-expert students that annotated the training set were all medical students with prior experience in annotating pathology cases.</p>\n\n<p><strong>Radboudumc data</strong></p>\n\n<p>The training set contains label noise. This label noise is introduced due to several reasons, including inconclusive pathologist reports, annotation errors, errors in the original diagnosis, disagreement between pathologists. To test the level of label noise, we let students annotate the test set with the same protocol as the training set. The labels, as determined by the students, were then compared to the consensus labels set by the experts. On grade group, the accuracy was <strong>0.720 (quadratic weighted kappa 0.853)</strong>. These values indicate a high agreement, but show the presence of label errors. Given the nature of Gleason grading and the problems of rater disagreement, handling this label noise is part of the challenge.</p>\n\n<p>The test set was graded by three experts in consensus. We determined this as the best possible gold standard for this grading task. Still, due to the subjective nature of Gleason grading, some errors can still be present. </p>\n\n<p><strong>Karolinska data</strong></p>\n\n<p>All cases were retrieved from the STHLM3 study with participants from the Stockholm county, Sweden, during the years 2013-2015.</p>\n\n<p>Each file represents a single case/slide and has one grade. Each slide typically consists of two sections from the same biopsy, but there is occasionally only one. In the case of cancer in the slide, one of the sections has a pen mark adjacent to the tissue where cancer is present.</p>\n\n<p>Similarly to the cases from Radboudumc, the Karolinska cases contain label noise due to the subjective nature of the Gleason grading system</p>",
          "votes": 18,
          "replies": []
        },
        {
          "id": 838480,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-05-08T15:46:34.173000",
          "content": "<p><a href=\"/wouterbulten\">@wouterbulten</a> Thank you for clarification of the structure of the test set and additional reference material on the data preparation.\n<a href=\"/drhabib\">@drhabib</a> , really good summary, thanks.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 838645,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "2020-05-08T17:38:34.303000",
          "content": "<p><a href=\"/wouterbulten\">@wouterbulten</a> <a href=\"/drhabib\">@drhabib</a> </p>\n\n<p>\"Each file represents a single case/slide and has one grade. Each slide typically consists of two sections from the same biopsy, but there is occasionally only one. In the case of cancer in the slide, one of the sections has a pen mark adjacent to the tissue where cancer is present.\"</p>\n\n<p>Can anyone clarify what the situation is with pen marks in the test set? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 838659,
          "author_name": "Benjamin Dubreu",
          "author_url": "",
          "post_date": "2020-05-08T17:44:32.630000",
          "content": "<p><a href=\"/fergusoci\">@fergusoci</a> : I believe I read somewhere (probably in the data tab section) that there is no pen marks in the test set.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 838668,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "2020-05-08T17:54:13.543000",
          "content": "<p>Ah, thank you, <a href=\"/bdubreu\">@bdubreu</a> </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 838916,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-05-08T22:27:36.687000",
          "content": "<p>Regarding <code>\"All of the private test set images and most of the public test set images were graded by multiple pathologists, but this was not feasible for the training set\"</code> and <code>\"The labels, as determined by the students, were then compared to the consensus labels set by the experts. On grade group, the accuracy was 0.720 (quadratic weighted kappa 0.853)\"</code>, It explains why CV and LB can have an opposit  trend, even for individual components: <code>[CV 0.886 (karolinska 0.909, radboud 0.840), LB 0.90]</code> vs. <code>[CV 0.892 (karolinska 0.923, radboud 0.842), LB 0.89]</code>. \nIt seems that we will need to deal with untrustable CV and noisy LB, which can be easily overfitted(</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 838923,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-05-08T22:45:41.703000",
          "content": "<p>Time to start fine tuning <code>random_seeds</code> =) </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 915894,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-07-05T07:23:06.250000",
          "content": "<p><a href=\"/wouterbulten\">@wouterbulten</a> \nis there a possiblity that kernels public score are calculated using different portion set of  test data every time the submission happens,otherwise it is surprising that many  are having difference of .01 to 0.03 in their score every time they submit using even though same models</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 856434,
      "author_name": "Arnaud Roussel",
      "author_url": "",
      "post_date": "2020-05-21T17:50:08.520000",
      "content": "<p>I have a lot of difference between local CV and LB\nBest single fold 0.8090CV 0.85LB\nSame model with 5 fold ensemble average (0.8012 average local CV) 0.86LB</p>\n\n<p>I did the tile selection a bit different than you and used to have more \"bright\" tiles. Currently running a model with your tile selection to see what it does I'm seeing an increase in local CV but will be interesting to see how it translates to LB.</p>\n\n<p>Haven't yet output the per provider results. Will do later or for another model run.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 854879,
      "author_name": "Aksell",
      "author_url": "",
      "post_date": "2020-05-20T12:06:37.153000",
      "content": "<p>model: resnext50\nfold: single fold\ncv-qwk : 0.877\nLb-qwk : 0.89</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 904856,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-28T03:03:07.190000",
      "content": "<p>resnet34, single fold, single model 0.91 LB</p>",
      "votes": 3,
      "replies": [
        {
          "id": 904953,
          "author_name": "ilovescience",
          "author_url": "",
          "post_date": "2020-06-28T05:34:03.997000",
          "content": "<p>Quite impressive! What kind of dataset are you using? Iafoss's version, Qishen Ha's version, your own?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 915873,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-07-05T07:05:58.937000",
          "content": "<p>@zhang well done...\ni wont ask much details. Just the public kernels tiling approach + res34 is enough for you to get the score or you had to adopt different tiling and loss approach.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 857651,
      "author_name": "DrHB",
      "author_url": "",
      "post_date": "2020-05-22T19:41:28.943000",
      "content": "<p><code>\nmodel: resnet34\nfold:     1\ncv:       0.884\nkarolinska: 0.8750\nradbound:  0.857\n</code></p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 885139,
      "author_name": "Salman",
      "author_url": "",
      "post_date": "2020-06-14T00:08:02.427000",
      "content": "<p>Can you please share what is the maximum number of tiles from intermediate layer with which you managed to train network?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 885151,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-06-14T00:34:21.097000",
          "content": "<p>If u check <a href=\"https://www.kaggle.com/haqishen/panda-inference-w-36-tiles-256\">this kernel</a>, u will see some setup that is working with intermediate res layer.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 885162,
          "author_name": "Salman",
          "author_url": "",
          "post_date": "2020-06-14T01:04:14.367000",
          "content": "<p>I know there we managed like 36 tiles of 256 size. I just wanted to confirm about maximum number of patches possible through the network &gt;= 256*256. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 874494,
      "author_name": "Iafoss",
      "author_url": "",
      "post_date": "2020-06-05T03:44:57.837000",
      "content": "<p>Actually I could boost single fold single model performance to 0.91 LB. Unfortunately, based on my previous subs, ensemble doesn't give much boost, and my expectation is getting only 0.01- improvement from it at single model performance of 0.90-0.91.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 875115,
          "author_name": "Benjamin Dubreu",
          "author_url": "",
          "post_date": "2020-06-05T14:27:23.570000",
          "content": "<p>Impressive... Does that single model at 0.91 involve better preprocessing, or different modeling tricks ? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 875313,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-06-05T16:42:13.297000",
          "content": "<p>Different tiles from ones I considered before.</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 860380,
      "author_name": "fergusoci",
      "author_url": "",
      "post_date": "2020-05-25T09:00:59.837000",
      "content": "<p>I have LB 0.84 with efficientnet-b4 (CV 0.82). Efficientnet-b0 gave me LB 0.8 (CV 0.79). </p>",
      "votes": 1,
      "replies": [
        {
          "id": 860597,
          "author_name": "Claudio Fanconi",
          "author_url": "",
          "post_date": "2020-05-25T13:02:19.127000",
          "content": "<p>Do you use the orginial image sizes from the EfficientNet paper which are optimized for its compound method? Hence, 224x224 for EffNetB0 and 380x380 for EffNetB4, or do you just use 128x128 from iafoss's tiles?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 860792,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "2020-05-25T15:38:06.970000",
          "content": "<p>I used 224x224 for both b0 and b4</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 860799,
          "author_name": "Claudio Fanconi",
          "author_url": "",
          "post_date": "2020-05-25T15:41:36.783000",
          "content": "<p>Alright, thanks! I had the same result for B0 - I shall try it with B4:)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 860816,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "2020-05-25T15:55:55.820000",
          "content": "<p>Yep, there's a tradeoff between the input size and how many tiles you can fit in a batch, obviously. I kept it at 224 because otherwise I couldn't get enough tiles to fit into memory during training and it hurt the results</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 860826,
          "author_name": "Claudio Fanconi",
          "author_url": "",
          "post_date": "2020-05-25T16:02:53.120000",
          "content": "<p>Yes, I came across the same problem. May I ask if you again used 12 tiles?\nFurthermore, how many epochs did it take you for convergence with B4?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 860965,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "2020-05-25T18:25:27.797000",
          "content": "<p>I use batches of 24 (split across 3 x 1080s). I got decent convergence after about 20-30 epochs, but ran for 50 in the end.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 861005,
          "author_name": "Claudio Fanconi",
          "author_url": "",
          "post_date": "2020-05-25T19:00:19.503000",
          "content": "<p>Okay that makes sense, as the max BS I could reach in a Kaggle Notebook is 8. Thus, only using a single GPU.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 894708,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-06-20T17:20:43.720000",
          "content": "<p><a href=\"/fergusoci\">@fergusoci</a> \nhow many tiles you are able to use for a an image.\nWHat is your current CV vs LB.. \nI find hard to  get CV and LB match until i was 0.88 . </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 894736,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "2020-06-20T17:53:59.730000",
          "content": "<p>When I wrote this post I was using a different tiling approach (from <a href=\"https://developer.ibm.com/articles/an-automatic-method-to-identify-tissues-from-big-whole-slide-images-pt1/\">here</a>). Now I'm using 36 x 256 tiles, with a tiling approach more similar to iafoss's. CV and LB are pretty well aligned, not much between them.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 915883,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-07-05T07:11:00.203000",
          "content": "<p><a href=\"/fergusoci\">@fergusoci</a>  with effnet how much best cv loss are u able to get. I am not getting loss better than 0.230/0.240 </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 916313,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "2020-07-05T14:41:33.907000",
          "content": "<p>I presume you are referring to BCE loss? I'm finding that QWK is only loosely correlated to the loss. I'm getting losses in the region of 0.19-0.2 with effb0 if I include all tiff files, 0.22-0.23 if I exclude duplicates. But qwk is roughly similar (around 0.89x) in both cases.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 857585,
      "author_name": "Jeremy Berros",
      "author_url": "",
      "post_date": "2020-05-22T18:12:26.123000",
      "content": "<p>Any idea for improvement is welcome 😃 💪 </p>\n\n<p>Framework:</p>\n\n<ul>\n<li>Tensorflow, Keras </li>\n</ul>\n\n<p>Single fold:</p>\n\n<pre><code>X_train, X_val = train_test_split(train, test_size=.2, stratify=train['isup_grade'], random_state=SEED)\n</code></pre>\n\n<p>Pre processing: </p>\n\n<ul>\n<li>4X4 tiled images (384, 384, 3)</li>\n<li>Augmentations (Hor/Ver Flip, ShiftScaleRotate)</li>\n</ul>\n\n<p>Config:</p>\n\n<pre><code>LR: 1e-3 \nBS: 16\nEpoch: 40\n</code></pre>\n\n<p>Model: Seresnext50 backbone <br> \nClassification: </p>\n\n<pre><code>loss='categorical_crossentropy'\noptimizer=optimizers.Adam(lr=LR)\nmetrics=[qw_kappa_score]\n</code></pre>\n\n<p>Callback= ReduceLROnPlateau <br></p>\n\n<pre><code>CV: 0.78\nLB: 0.79\n</code></pre>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 857095,
      "author_name": "JIANJIAN",
      "author_url": "",
      "post_date": "2020-05-22T09:45:22.517000",
      "content": "<p>singlefold\nLB：0.82\nCV：0.88    Karolinska ：0.8387   Radboud：0.8806</p>\n\n<p>It's funny!  Is LB all about Karolinska？：）</p>\n\n<p>Update:LB:0.84\nCV:0.87   Karolinska ：0.8640   Radboud：0.8477</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 852885,
      "author_name": "yuvaramsingh",
      "author_url": "",
      "post_date": "2020-05-18T18:55:11.367000",
      "content": "<p>model: resnext50\nfold: single fold\ncv-qwk : 0.67\nLb-qwk : 0.60</p>\n\n<p>loss fn : CategoricalCrossentropy\noptimizer : adam\ninput : tile concatenated into 512X512X3</p>\n\n<p>it will be helpful if someone can share ideas on how i can improve my model next. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 854116,
          "author_name": "wayfarer",
          "author_url": "",
          "post_date": "2020-05-19T19:08:49.537000",
          "content": "<p>try this approach <a href=\"https://www.kaggle.com/iafoss/panda-concat-tile-pooling-starter-0-79-lb\">https://www.kaggle.com/iafoss/panda-concat-tile-pooling-starter-0-79-lb</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 849615,
      "author_name": "Arnaud Roussel",
      "author_url": "",
      "post_date": "2020-05-15T23:32:54.663000",
      "content": "<p>How long does it take you to train a 5CV 20 epoch ? I feel my biggest pain right now is that testing ideas take like 9Hours on my GTX1080 ti (for 16x128x128) so it feels very slow. I guess I can always use a cloud solution.</p>\n\n<p>Also any tip for faster iterations ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 849672,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-05-16T01:01:14.010000",
          "content": "<p>I think you should do your tests with only one fold and then when you have a good solution train 5 folds! </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 849692,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2020-05-16T01:55:06.540000",
          "content": "<p>Yeah that is kinda what I defacto ended up doing: Stopping after one fold if I didn't see any interesting improvement.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 845825,
      "author_name": "seefun",
      "author_url": "",
      "post_date": "2020-05-13T12:52:30.010000",
      "content": "<p>May I ask how many tiles you use to get 0.90 LB? \nI use 112x112x64 but it doesn't come up to higher score.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 846213,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-05-13T16:15:30.510000",
          "content": "<p>For the low res layer 12x128x128 seems to be near the most optimal setup. For intermediate resolution I have changed both: the size and the number of tiles. When you select your setup u should keep in mind that <code>N*sz*sz ~ tissue area</code>. Also, smaller tiles would eliminate white space (allow to use less tiles), but may degrade the model performance because it may be difficult to say what is the kind of tissue is shown on a given small piece of an image without seeing the surrounding. \nAnd for intermediate resolution it is not only about the tile setup, as <a href=\"/oscarrangel\">@oscarrangel</a> asked below. Just straight use of intermediate tiles with my method may not work because of small bs limited by GPU RAM, and other trick(s) should be added. My first trial on intermediate res also didn't give better results than low res, while the second trial got 0.90 LB. So just keep exploring different tricks, and u may find something even better than I use.</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 846334,
          "author_name": "TheStoneMX",
          "author_url": "",
          "post_date": "2020-05-13T17:33:25.870000",
          "content": "<p>I could not make it work, debugging it, I found the GPU error lack of memory happens on the forward pass in the train met method... so I end up buying an Nvidia RTX with 24 Ggs, so now I will have about 40 gigs of GPU memory with data-parallel.  waiting to arrive. and there is a problem with the tqdm on kaggle.... <a href=\"https://www.kaggle.com/product-feedback/150817\">https://www.kaggle.com/product-feedback/150817</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 846478,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-05-13T19:35:25.153000",
          "content": "<p>I'm not sure about batchsize, i train with batch size 6 and i can get decent results. But probably higher bs could help boost a little bit the performance!</p>\n\n<p>As this paper says, mini batch size lower than 32 might be the way to go, it usually achieves better results: <a href=\"https://arxiv.org/pdf/1804.07612.pdf\">https://arxiv.org/pdf/1804.07612.pdf</a></p>\n\n<p>But I wish i had more memory.. haha!  </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 846524,
          "author_name": "Konstantin Lopukhin",
          "author_url": "",
          "post_date": "2020-05-13T20:15:45.973000",
          "content": "<p>FWIW it's possible to get 0.86 LB on one 2080 Ti in about 9 hours of training. But yes memory is a huge problem.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 847288,
          "author_name": "A.Demyanchuk",
          "author_url": "",
          "post_date": "2020-05-14T09:22:18.307000",
          "content": "<p>9 hours of training is a lot! Is it one model only or all folds?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 847337,
          "author_name": "Konstantin Lopukhin",
          "author_url": "",
          "post_date": "2020-05-14T10:22:40.210000",
          "content": "<p>It's one model (resnet50), yes not quite as fast as I'd like to.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 847595,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-05-14T14:10:23.503000",
          "content": "<p>There might better solution for low batch training. Recently in <a href=\"https://arxiv.org/pdf/2004.02967.pdf\">this</a> paper they propose new normalization-activation layer <code>Evo-Norm-S0</code>. Not only it performs better in different task but it also extremely stable across many batch sizes (32 bs is almost same as 4096). In figure 1 you can see nice summarization of performance. </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2Fdff4da383bfcef9078554a8038553f2b%2FScreen%20Shot%202020-05-14%20at%2010.07.53%20AM.png?generation=1589465309983305&amp;alt=media\" alt=\"\"></p>\n\n<p>Pytorch code - <a href=\"https://github.com/digantamisra98/EvoNorm\">https://github.com/digantamisra98/EvoNorm</a>\nTensorflow code - <a href=\"https://github.com/sayakpaul/EvoNorms-in-TensorFlow-2\">https://github.com/sayakpaul/EvoNorms-in-TensorFlow-2</a></p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 847622,
          "author_name": "Shujun",
          "author_url": "",
          "post_date": "2020-05-14T14:26:33.130000",
          "content": "<p>Another thing that might be interesting is inplace activated batchnorm, which can help with a big issue of the competition - memory, although it has not worked for me... Have you tried anything like that?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 847638,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-05-14T14:35:14.343000",
          "content": "<p>I always used inplace... =) </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 847662,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-05-14T14:44:14.363000",
          "content": "<p>Im not quite sure what is inplace activated batchnorm? is it like an complete other batchnorm?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 868454,
          "author_name": "Shiyuan Zeng",
          "author_url": "",
          "post_date": "2020-05-31T08:48:15.430000",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> Sorry to disturb you. If I wanna replace BN with EvoNorm, how can I do it?\n<code>\ndef dfs(net):\n    for x in net.children():\n        if isinstance(x, nn.BatchNorm2d):\n            print(x, x.num_features)\n        else:\n            dfs(x)\n</code>\nI write a dfs to print all the BatchNorm2d in backbone, but I can't change the type of it by using ’x=EvoNrom‘. Is python exists a way to use quotative-vaiable in for-loop?(like the c++ code below)\n<code>\nfor ( int&amp; i = 0; i &amp;lt; N; i ++ ) { ...... }\n</code>\nOr I should copy the whole codes of my backbone and then modify it by hand?(it's so terrible ;_;)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 868980,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-05-31T15:59:06.013000",
          "content": "<p>You can consider the following code I tried for GN:\n<code>\ndef to_GN(model):\n    for child_name, child in model.named_children():\n        if isinstance(child, nn.BatchNorm2d):\n            setattr(model, child_name, nn.GroupNorm(32,child.num_features))\n        else:\n            to_GN(child)\n</code></p>\n\n<p>Though, such replacement quite destroys the pretrained weights and didn't work well for me.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 870821,
          "author_name": "Shiyuan Zeng",
          "author_url": "",
          "post_date": "2020-06-02T00:35:15.930000",
          "content": "<p>Thank you so much!💯  I considered to trun BN to EvoNorm is my batch-size becomes samller and samller when I add more tiles. But you are right, abandon the pretrained weights also a big problem☹️ </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 870880,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2020-06-02T02:24:45.877000",
          "content": "<p>While your overall batch size gets smaller, you still send a lot of tiles through the CNN if you use the tile strategy so potentially say 4 x 32 is still good enough for batchnorm in the CNN. While the tiles are correlated it probably isn't that bad and discarding pretrained weight is probably not worth it. On the other hand, for the FC part maybe I'll try something that is less dependent on batch-size.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 871001,
          "author_name": "Shiyuan Zeng",
          "author_url": "",
          "post_date": "2020-06-02T04:49:25.793000",
          "content": "<p>Your words make sense!👍  Maybe I should try iafoss's way to feed the image into model instead of seaming tiles into a large image.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 884344,
          "author_name": "Benjamin Dubreu",
          "author_url": "",
          "post_date": "2020-06-13T09:52:54.303000",
          "content": "<p>So, did any of you guys managed to make gradient accumulation work ? I managed to do grad_acc, but then, the small batch size screws the batch normalizations and results are not good in the end.\nTo that effect, I tried replacing batchNorm layers with groupNorm and EvoNorm, but both failed... Any clue on how one can effectively find a way to solve the batch-size problem ? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 885130,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2020-06-13T23:43:33.077000",
          "content": "<p>As a side note I use pytorch lightning to make grad accumulation work easily but of course it doesn't solve the batch norm issue...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 887910,
          "author_name": "Shujun",
          "author_url": "",
          "post_date": "2020-06-16T01:52:34.367000",
          "content": "<p><a href=\"/bdubreu\">@bdubreu</a> did it not work for you? That's a little surprising actually. When I used resnext50, I had to use a batch size of 4 and gradient accumulation every 12 batches, but my model turned out ok anyway (my current best 0.90 run). </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 887952,
          "author_name": "Shiyuan Zeng",
          "author_url": "",
          "post_date": "2020-06-16T03:06:24.047000",
          "content": "<p>I think what Roussel means is that accumulate gradient is just a train trick and not the same with really enlarge batch-size?\nJust a guess😜 </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 887964,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2020-06-16T03:19:36.333000",
          "content": "<p><a href=\"/cnzengshiyuan\">@cnzengshiyuan</a>  It's an intermediary solution. You get good gradients from it but it doesn't solve the batch norm issue. I still use it though.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 888069,
          "author_name": "Shiyuan Zeng",
          "author_url": "",
          "post_date": "2020-06-16T05:27:06.533000",
          "content": "<p><a href=\"/arroqc\">@arroqc</a> Thanks for your reply, I just started to use it😄 </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 888189,
          "author_name": "Benjamin Dubreu",
          "author_url": "",
          "post_date": "2020-06-16T07:13:41.463000",
          "content": "<p><a href=\"/shujun717\">@shujun717</a> thanks for telling me this. I was trying to replicate my partner's results (she has access to a 24gb GPU). So, having only 6gigs, I tried bs8 (she was doing 32) and tried everything I could to replicate her results. Never managed to. That being said, I wanted to \"fake\" a batch size of 32, so I never tried to do gradient accumulation for more than 4 batches. Maybe doing more will help, I will try that and report here !</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 890579,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2020-06-17T15:24:06.320000",
          "content": "<p>New result for me with a CV at 0.9009 and LB at 0.89 (TTA) with single fold (and not even sure if this is the best fold as this is my first try with new architecture/process).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 890808,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "2020-06-17T17:56:03.520000",
          "content": "<p><a href=\"/arroqc\">@arroqc</a> Do you know how much of the lift was down to architecture versus process? </p>\n\n<p>I've been playing around trying to get a feel for how much each of these could contribute. Thinking that how the tiles are constructed is key. I was originally selecting the top tiles based on the approach \n <a href=\"https://developer.ibm.com/articles/an-automatic-method-to-identify-tissues-from-big-whole-slide-images-pt1/\">here</a>. Switching to the (much simpler!) <a href=\"/iafoss\">@iafoss</a> method led to LB 0.83 -&gt; 0.87 with the same Efficientnet B0 architecture. Now experimenting with other approaches to tiling; it feels like using full resolution (or half resolution) would be useful, but I'm running out of space on the GPUs and gradient accumulation doesn't seem to be playing ball! Can't get the batch norm working...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 890851,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2020-06-17T18:14:18.963000",
          "content": "<p>Process is definitely most of the lift. Let's just say that my tile selection is much better than it used to be (and that this result actually uses a lower number of tiles than I used to). I am now working with it to see how far I can push this idea. Like most people I still struggle with compute time and batch norms when using something like 36 tiles.</p>\n\n<p>Also, I'm getting better results with stitched tiles than bag of tiles. The main reason is Karolinska is better with stitched (0.91CV). I am still investigating why but maybe it has to do with the higher number of benign slides with karolinska where your network can more quickly classify it by seeing all tiles at the same time. Radboud is exactly the same for both methods for me though (0.87CV).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 891209,
          "author_name": "Sugawarya",
          "author_url": "",
          "post_date": "2020-06-18T02:41:45.187000",
          "content": "<p>what is \"stitched tiles\" ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 891231,
          "author_name": "Salman",
          "author_url": "",
          "post_date": "2020-06-18T03:22:40.893000",
          "content": "<p>Did you try to maintain sequence of patches in a way? Or something similar to that?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 891461,
          "author_name": "Alex",
          "author_url": "",
          "post_date": "2020-06-18T07:36:23.890000",
          "content": "<p>I've tried gradient accumulation without success. \nThanks <a href=\"/arroqc\">@arroqc</a> for you feedbacks. I think stitched tiles means concatenated tiles ?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 891849,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2020-06-18T13:45:37.817000",
          "content": "<p>Yes by stitched I mean concatenated in a square.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 891906,
          "author_name": "Benjamin Dubreu",
          "author_url": "",
          "post_date": "2020-06-18T14:24:13.990000",
          "content": "<p><a href=\"/arroqc\">@arroqc</a> congrats on the lift ! does your better tiles involves using the masks ? All my experiments trying to use the mask for tile selection seem to fail. I haven't made a segmentation model to mask the test files and pick the tiles accordingly though, but I think most people haven't done that anyways...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 891914,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2020-06-18T14:29:35.597000",
          "content": "<p>No I don't use the masks as it seems to be very noisy and also differences between the two providers...</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 893390,
          "author_name": "SumanSudhir",
          "author_url": "",
          "post_date": "2020-06-19T15:13:19.737000",
          "content": "<p><a href=\"/arroqc\">@arroqc</a> did you get this result using regression or classification?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 893406,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2020-06-19T15:25:53.723000",
          "content": "<p>This is with the bin + binary cross entropy method similar to the current best kernel. But I haven't tried with other losses so I don't know how it compares.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 902952,
          "author_name": "Benjamin Dubreu",
          "author_url": "",
          "post_date": "2020-06-26T13:37:28.107000",
          "content": "<p><a href=\"/shujun717\">@shujun717</a> does your gradient accumulation use reduction=sum, or reduction=mean ? \nI tried numerous setups with a step every N batches, to no avail. I think my implementation is not correct (since it worked for you). When you do a step, do you divide the accumulated gradients by the number of batches you do between two steps ? </p>\n\n<p>Sorry for annoying you with the specifics, but I'm trying a whole bunch of things (even with GroupNorm, EvoNorm and the like) and nothing seems to work ^^ </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 903050,
          "author_name": "Shujun",
          "author_url": "",
          "post_date": "2020-06-26T14:43:12.823000",
          "content": "<p>I used reduction=\"none\" then torch.mean and divide by <code>gradient_accumulation_steps</code>. And \n<code>python\nif step%gradient_accumulation_steps==0:\n  optimizer.step()\n  optimizer.zero_grad()\n</code>\nThis should not be different from reduction=sum, or reduction=mean, provided that you divide your loss accordingly. I only used reduction=\"none\" to I can scale the losses individually. I don't freeze bn or anything either </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 903076,
          "author_name": "Benjamin Dubreu",
          "author_url": "",
          "post_date": "2020-06-26T14:58:41.810000",
          "content": "<p>Then I will have to take a look at everybody else's code (including yours) at the end of comp' because I don't see what I'm missing here ^^ Thanks for taking the time to reply though ! </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 903157,
          "author_name": "Shujun",
          "author_url": "",
          "post_date": "2020-06-26T16:00:43.140000",
          "content": "<p>I can take a look at your training loop if you want. It might just be some simple mistake I feel, since it worked for me the first time I tried it </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 903659,
          "author_name": "tattaka",
          "author_url": "",
          "post_date": "2020-06-27T03:12:09.453000",
          "content": "<p>I have the same problem. I have implemented the model update process as follows (refer to <a href=\"https://discuss.pytorch.org/t/why-do-we-need-to-set-the-gradients-manually-to-zero-in-pytorch/4903/20\">https://discuss.pytorch.org/t/why-do-we-need-to-set-the-gradients-manually-to-zero-in-pytorch/4903/20</a> ):\n```\ndef update_model(self, model, loss, optimizer):\n        if self.mode == \"train\":\n            need_gradient_step = (\n                self._accumulation_counter + 1) % self.accumulation_steps == 0\n            model.zero_grad()\n            if self.use_amp and amp_enable:\n                delay_unscale = not need_gradient_step\n                with amp.scale_loss(loss, optimizer, delay_unscale=delay_unscale) as scaled_loss:\n                    scaled_loss.backward()\n            else:\n                loss.backward()</p>\n\n<pre><code>        if need_gradient_step:\n            optimizer.step()\n            self._accumulation_counter = 0\n</code></pre>\n\n<p>```\nIs this a mistake?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 916978,
          "author_name": "Shujun",
          "author_url": "",
          "post_date": "2020-07-06T06:50:24.253000",
          "content": "<p>This looks fine to me, not sure why you are setting _accumulation_counter to 0 tho</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 917072,
          "author_name": "tattaka",
          "author_url": "",
          "post_date": "2020-07-06T08:05:29.830000",
          "content": "<p>Thank you reply.\n <code>accumelation_counter</code> is incremented with the training loop.\nI keep debugging...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 917115,
          "author_name": "Shujun",
          "author_url": "",
          "post_date": "2020-07-06T08:51:46.963000",
          "content": "<p>What specific bug/issue do you have with this code?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 917143,
          "author_name": "Benjamin Dubreu",
          "author_url": "",
          "post_date": "2020-07-06T09:11:15",
          "content": "<p>Hi <a href=\"/shujun717\">@shujun717</a> ! for some reason, I never saw your comment above ! Thank you so much for the offer to review my code. But as this is a competition I believe I should be able to handle this on my own. I might come back to you <em>after</em> the competition if I don't see what I did wrong by looking at everyone's code though, if that's ok for you ;)\nGood luck for the end, I hope you stay top10 ! </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 842703,
      "author_name": "Shujun",
      "author_url": "",
      "post_date": "2020-05-11T15:25:53.827000",
      "content": "<p><a href=\"/iafoss\">@iafoss</a> how many epochs do you train per fold? In my experiments, my model seems to need more than 30 epochs per fold to converge</p>",
      "votes": 1,
      "replies": [
        {
          "id": 842708,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-05-11T15:29:09.190000",
          "content": "<p>It's correct, the convergence is not very fast, and I train for more than 30 epochs.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 842712,
          "author_name": "Shujun",
          "author_url": "",
          "post_date": "2020-05-11T15:35:10.067000",
          "content": "<p>I see. Thanks</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 835690,
      "author_name": "Kaushal Shah",
      "author_url": "",
      "post_date": "2020-05-06T12:52:41.957000",
      "content": "<p><a href=\"/iafoss\">@iafoss</a> are you using masks?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 835950,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-05-06T16:03:41.897000",
          "content": "<p>In my experiment with mask aux, as I pointed out above, I got only a slight improvment. Meanwhile, the training time increased quite a bit. So, for larger resolution I didn't use it.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 836930,
          "author_name": "Kaushal Shah",
          "author_url": "",
          "post_date": "2020-05-07T11:28:16.763000",
          "content": "<p><a href=\"/iafoss\">@iafoss</a> okay thanks!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 855587,
      "author_name": "R Guo",
      "author_url": "",
      "post_date": "2020-05-21T03:43:56.537000",
      "content": "<p>model: single fold seresnext50\nCV: 0.9\nLB: 0.88</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 835514,
      "author_name": "Vlad Vaduva",
      "author_url": "",
      "post_date": "2020-05-06T10:13:58.833000",
      "content": "<p>Nice result <a href=\"/iafoss\">@iafoss</a> . If you don't mind, your tricks are related to model architecture or data augmentation/preprocessing ?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 835943,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-05-06T16:00:10.557000",
          "content": "<p>I'd say they are related to both you mentioned and and other things as well.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 842831,
      "author_name": "TheStoneMX",
      "author_url": "",
      "post_date": "2020-05-11T16:40:05.743000",
      "content": "<p>Hi there all,</p>\n\n<p>is any one know how to translate this Pytorch code into fastai ?</p>\n\n<p>~~~\nfor i, (inputs, labels) in enumerate(training_set):\n    predictions = model(inputs)                     # Forward pass\n    loss = loss_function(predictions, labels)       # Compute loss function\n    loss = loss / accumulation_steps                # Normalize our loss (if averaged)\n    loss.backward()                                 # Backward pass\n    if (i+1) % accumulation_steps == 0:             # Wait for several backward steps\n        optimizer.step()                            # Now we can do an optimizer step\n        model.zero_grad()                           # Reset gradients tensors\n        if (i+1) % evaluation_steps == 0:           # Evaluate the model when we...\n            evaluate_model() \n~~~</p>",
      "votes": 0,
      "replies": [
        {
          "id": 842898,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-05-11T17:28:51.523000",
          "content": "<p>You always can use your favorite search engine and type \"fastai gradient accumulation callback\". Specifically now fast.ai has AccumulateScheduler. Also I wrote my own one a while back in this kernel\n<a href=\"https://www.kaggle.com/iafoss/hypercolumns-pneumothorax-fastai-0-831-lb\">https://www.kaggle.com/iafoss/hypercolumns-pneumothorax-fastai-0-831-lb</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 842966,
          "author_name": "Shujun",
          "author_url": "",
          "post_date": "2020-05-11T18:19:08.093000",
          "content": "<p>Or why not just use pure pytorch?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 842988,
          "author_name": "TheStoneMX",
          "author_url": "",
          "post_date": "2020-05-11T18:33:03.227000",
          "content": "<p>@lafoss, I did search and posted the question on Fastai forum, but everyone send m me to fastai2, I have tried to upgrade your project to fastai2 but it does not works.... I guess I don't know much about fastai in order to upgrade it....</p>\n\n<p>Thanks for the post, trying to follow your advise on GPU optimization.... so I have have gradients update not when I use bz=8, but simulate bz=32</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 842991,
          "author_name": "TheStoneMX",
          "author_url": "",
          "post_date": "2020-05-11T18:35:25.507000",
          "content": "<p><a href=\"/shujun717\">@shujun717</a> , because then you will have to get into the fit_one_cycle() code which I dont w want to mess with it, I know it can be implemented with a callback call, I am learning, maybe there is a way, but I dont knnow it... \nMaybe someone else can answer your questions better.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 843005,
          "author_name": "TheStoneMX",
          "author_url": "",
          "post_date": "2020-05-11T18:45:53.627000",
          "content": "<p><a href=\"https://forums.fast.ai/t/gradient-accumulation/70968/5?u=orangelmx\">https://forums.fast.ai/t/gradient-accumulation/70968/5?u=orangelmx</a></p>\n\n<p><a href=\"https://forums.fast.ai/t/how-can-you-train-your-model-on-large-batches-when-your-gpu-can-t-hold-more-than-a-few-samples/70895/13?u=orangelmx\">https://forums.fast.ai/t/how-can-you-train-your-model-on-large-batches-when-your-gpu-can-t-hold-more-than-a-few-samples/70895/13?u=orangelmx</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 845911,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-05-13T13:41:24.973000",
          "content": "<p><a href=\"/oscarrangel\">@oscarrangel</a>  Pytorch has native <code>one cycle</code> <a href=\"https://pytorch.org/docs/stable/_modules/torch/optim/lr_scheduler.html#OneCycleLR\">https://pytorch.org/docs/stable/_modules/torch/optim/lr_scheduler.html#OneCycleLR</a> </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 849539,
          "author_name": "BachT",
          "author_url": "",
          "post_date": "2020-05-15T21:30:17.730000",
          "content": "<p>Do you know if there is a <code>find_lr</code> for Pytorch? I am also using the native <code>OneCycleLR</code>, but struggle for <code>find_lr</code>. I implemented one from <code>https://gist.github.com/NegatioN/07edc229f9d668b2b366528d94500f49</code> but the result is weird... not sure it is due to the model or due to the bug (which I do not see it... yet :( )</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F101755%2Fdfdcb92184b38ed68d392e793479d53e%2F2020-05-15_232839.png?generation=1589578214713980&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 849572,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-05-15T22:53:59.910000",
          "content": "<p>It seams that the initial value of smoothed loss may be not accounted correctly (also, if you start from untrained model, the loss should initially drop).</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 926069,
      "author_name": "TheStoneMX",
      "author_url": "",
      "post_date": "2020-07-12T13:24:25.847000",
      "content": "<p>Hi @lafoss, I was wondering if can help me how to load my model from candence pretrainedmodels, I am still learning.\n~~~\nimport sys\nfrom pathlib import Path</p>\n\n<p>sys.path.append('../input/pytorch-pretrained-models/repository/pretrained-models.pytorch-master')\nimport pretrainedmodels</p>\n\n<p>m = pretrainedmodels.se_resnext50(pretrained='imagenet')\nchildren = list(m.children())\nhead = nn.Sequential(nn.AdaptiveAvgPool2d(1), Flatten(), \n                                  nn.Linear(children[-1].in_features,200))\nmodel = nn.Sequential(nn.Sequential(*children[:-2]), head)\n~~~\nThen tries to download but as we know we dont have internet, I have the model in the directory, but I dont know how to load it,</p>\n\n<p>thanks for the help.</p>",
      "votes": -4,
      "replies": [
        {
          "id": 926193,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-07-12T15:12:31.767000",
          "content": "<p>You don't need to load pretrained weights in inference kernel, just load your model. Try to look how to build the model without pretrained weights.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 926243,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-12T15:40:09.673000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 926292,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-07-12T16:21:44.063000",
          "content": "<p>yes</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 926330,
          "author_name": "TheStoneMX",
          "author_url": "",
          "post_date": "2020-07-12T16:44:23.760000",
          "content": "<p>I am getting this error,  is looking for the model directory first... \n[Errno 2] No such file or directory: 'models/../input/pandas-models/1_july9_model.pth'</p>\n\n<p>thanks for the help!!!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 926373,
          "author_name": "TheStoneMX",
          "author_url": "",
          "post_date": "2020-07-12T17:09:15.350000",
          "content": "<p>????\n~~~\nPath('./models').mkdir(exist_ok=True, parents=True)</p>\n\n<p>!cp '../input/panda-models/july9_model.pth' './models/july9_model'</p>\n\n<p>FileNotFoundError: [Errno 2] No such file or directory: 'models/july9_model.pth'\n~~~</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 926406,
          "author_name": "TheStoneMX",
          "author_url": "",
          "post_date": "2020-07-12T17:37:52.107000",
          "content": "<p>I was able to solve it changing the name of the model to 'saved'</p>\n\n<p>~~~</p>\n\n<h1>learn.load('saved')</h1>\n\n<p>~~~</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 840802,
      "author_name": "Ankur Shukla",
      "author_url": "",
      "post_date": "2020-05-10T11:51:20.340000",
      "content": "<p>densenet121 architechture.</p>",
      "votes": -1,
      "replies": []
    },
    {
      "id": 928008,
      "author_name": "TheStoneMX",
      "author_url": "",
      "post_date": "2020-07-13T17:25:37.090000",
      "content": "<p>Hi there my frind @lafoss I was wondering if you can give me a help one more time before this ends, I havent be able to submit anymore</p>\n\n<p>the error is that  \"csv file not found.... \" ????? when I run it on draft everything works fine! as I do the commit and creates the cvs file for the submition,** Submission CSV Not Found**</p>\n\n<p>I was wondering if you see something wrong with the code?</p>\n\n<p>~~~\nTEST_PATH = '/kaggle/input/prostate-cancer-grade-assessment/test_images'</p>\n\n<p>if os.path.exists(TEST_PATH):\n    dls = dBlock.dataloaders(df_test, bs=8)</p>\n\n<pre><code>learn = Learner(dls, get_model())\nlearn.load('saved') \n\ntest_dl = dls.test_dl(df_test)\n_,_, preds = learn.get_preds(dl=test_dl, with_decoded=True)\n\ndf_test[\"isup_grade\"] = preds\nsub = df_test[[\"image_id\",\"isup_grade\"]]\nsub.to_csv('submission.csv', index=False) \n</code></pre>\n\n<p>else:\n    df_train =  df_train.loc[:5]\n    dls = dBlock.dataloaders(df_train, bs=8)</p>\n\n<pre><code>learn = Learner(dls, get_model())\nlearn.load('saved') \n\ntrain_dl = dls.test_dl(df_train)\n_,_, preds = learn.get_preds(dl=train_dl, with_decoded=True)\n\ndf_train[\"isup_grade\"] = preds\nsub = df_train[[\"image_id\",\"isup_grade\"]]\nsub.to_csv('submission.csv', index=Fals\n</code></pre>\n\n<p>~~~\nThanks a lot for all your help, I have learn a lot on this competition thanks to you.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 929580,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-07-14T18:58:53.787000",
          "content": "<p>I'd think that df_test is not opened or something similar. That is happening is u get some error if the condition in the if statement is True, and the execution of the cell is stopped before u write csv.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 882507,
      "author_name": "Salman",
      "author_url": "",
      "post_date": "2020-06-11T21:35:07.237000",
      "content": "<p><a href=\"/iafoss\">@iafoss</a> can you tell me about your best CV and LB with lowest res tiles?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 882557,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-06-11T23:13:47.337000",
          "content": "<p>As written above, on the lowest res 4 fold CV of 0.843 gives ~0.80 LB, and the max LB I got for these images is 0.82. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 867598,
      "author_name": "Adil Zouitine",
      "author_url": "",
      "post_date": "2020-05-30T12:51:08.773000",
      "content": "<p>Method: Based on lafoss tiling\nModel: ResNet18\nFold: Stratified\nLeaderboard: 0.84\nValidation: \n<code>\nQWP:0.8556\n[[487  68  16   1   3   0]\n [ 96 331  81   8   6   1]\n [  8  84 128  34  14   1]\n [  1  20  47  65  87  25]\n [  6  17  20  33  77  96]\n [  4   3   9  15  76 136]]\n</code></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 861402,
      "author_name": "Shujun",
      "author_url": "",
      "post_date": "2020-05-26T03:12:39.043000",
      "content": "<p><a href=\"/iafoss\">@iafoss</a> is your 0.90 run from a single fold or multiple folds ensembled?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 861466,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-05-26T04:15:04.253000",
          "content": "<p>As I wrote above, I got 0.90 by a single fold model. Though, 4 folds also give 0.90.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 861485,
          "author_name": "Shujun",
          "author_url": "",
          "post_date": "2020-05-26T04:31:12.717000",
          "content": "<p>Right just checking. My 0.90 run was also single fold</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 859954,
      "author_name": "GreatGameDota",
      "author_url": "",
      "post_date": "2020-05-24T22:08:57.803000",
      "content": "<p>single fold seresnext50\nCV: .810\nLB: .85</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 854529,
      "author_name": "Yovin Yahathugoda",
      "author_url": "",
      "post_date": "2020-05-20T05:01:20.027000",
      "content": "<p>model - Resnext50\nfold - single fold\nCV-QWK : 0.85\nLB-QWK: 0.81</p>",
      "votes": 0,
      "replies": [
        {
          "id": 855368,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-05-20T21:31:22.840000",
          "content": "<p>Just a few questions, did u use the lowest res layer, and did u check the score based on institution? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 855557,
          "author_name": "Yovin Yahathugoda",
          "author_url": "",
          "post_date": "2020-05-21T03:17:00.737000",
          "content": "<p><a href=\"/iafoss\">@iafoss</a> i used level 1 images and directly applied your tile method without any rescaling.\nI did not check the CV score based on the institution, but i will check them and post here.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 855559,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-05-21T03:17:51.800000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 855691,
          "author_name": "Yovin Yahathugoda",
          "author_url": "",
          "post_date": "2020-05-21T05:47:18.753000",
          "content": "<p>Update <a href=\"/iafoss\">@iafoss</a> \nmodel - Resnext50\nfold - single fold\nImage level - level 1\nLB-QWK: 0.81\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1554871%2F82d3f9e996df7869781291b6debc43eb%2Foverall.png?generation=1590039929674473&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1554871%2Fdb787fe740b7933959868278c8c3add3%2Fkarolinska.png?generation=1590039930098666&amp;alt=media\" alt=\"\">   <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1554871%2Fc9a6cb6aa1112e3245911113405716f4%2Fradboud.png?generation=1590039931743737&amp;alt=media\" alt=\"\"></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 842197,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-05-11T08:52:39.740000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 842687,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-05-11T15:01:34.557000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 842102,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-05-11T07:39:47.183000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 841695,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-05-11T00:59:05.427000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 840268,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-05-09T22:09:37.680000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 839789,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-05-09T15:39:38.067000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 839819,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-05-09T16:01:19.550000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 840360,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-05-10T00:28:48.540000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 840396,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-05-10T01:26:24.567000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 840426,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-05-10T02:45:35.740000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 840428,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-05-10T02:51:50.737000",
          "content": "",
          "votes": 3,
          "replies": []
        },
        {
          "id": 843115,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-05-11T21:03:22.943000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 843135,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-05-11T21:19:34.247000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 839187,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-05-09T06:28:06.307000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 839117,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-05-09T05:04:18.533000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 838973,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-05-09T00:56:49.847000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 838995,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-05-09T01:43:59.140000",
          "content": "",
          "votes": 5,
          "replies": []
        },
        {
          "id": 839050,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-05-09T03:05:39.987000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 839057,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-05-09T03:20:32.817000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 839062,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-05-09T03:26:35.570000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 839171,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-05-09T06:03:58.200000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 839686,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-05-09T14:32:28.313000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 839689,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-05-09T14:33:49.053000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 839777,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-05-09T15:30:03.540000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 840049,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-05-09T18:31:35.950000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 840067,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-05-09T18:48:32.237000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 840111,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-05-09T19:18:09.023000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 840122,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-05-09T19:31:10.200000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 840189,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-05-09T20:21:15.167000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 837672,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-05-08T00:15:39.953000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 837694,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-05-08T01:13:48.783000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 838953,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-05-09T00:17:25.977000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 838957,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-05-09T00:31:46.850000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 838966,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-05-09T00:43:17.517000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 854932,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-05-20T13:03:44.380000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 835505,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-05-06T10:05:16.590000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 835940,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-05-06T15:58:41.777000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 835372,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-05-06T08:24:01.270000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 836060,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-05-06T17:30:52.517000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 836180,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-05-06T19:40:44.100000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 835362,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-05-06T08:19:05.180000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 835936,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-05-06T15:57:34.333000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 835356,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-05-06T08:15:26.637000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 835503,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-05-06T10:04:32.523000",
          "content": "",
          "votes": 6,
          "replies": []
        },
        {
          "id": 835932,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-05-06T15:55:36.653000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 840559,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-05-10T06:22:20.133000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 849685,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-05-16T01:32:31.937000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 903613,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-27T01:50:29.113000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 893385,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-19T15:11:17.780000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "835094": "I'm starting a common competition thread: **What is your current best single model?**\n\nMy one:\n**0.886 single fold CV, 0.90 LB**\n```\n[[615  75  19   7   3   0]\n [ 48 497  96  11   2   0]\n [  4  86 143  84  17   1]\n [ 10  13  41 136  91  15]\n [  2  13  17  44 156  79]\n [  2   6   9  13  78 196]]\n```\n- ResNext50 based like in [my kernel](https://www.kaggle.com/iafoss/panda-concat-tile-pooling-starter-0-79-lb) but with a number of additional tricks. \n- Tiles from a tiff layer of intermediate resolution based on [this kernel](https://www.kaggle.com/iafoss/panda-16x128x128-tiles).\n- No segmentation\n\n\n**Additional observations:**\n- Low res tiff layer (12x128x128 tiles) can give up to:\n4 fold CV of 0.843, ~0.80 LB (I'd expect that LB may be not stable, and I may face with a dilemma: should I trust more LB or CV... But I didn't find any leaks so far).\n```\n[[2290  424  106   41   11    1]\n [ 463 1563  455  116   17    2]\n [  60  396  547  261   67   10]\n [  26   81  213  397  395  114]\n [  29   65  128  182  450  391]\n [  13   31   42   91  270  768]]\n```\nOn low res tiff layer the same fold I ran 0.886 model (best one above) gives:\n```\n0.853 CV\n[[580 108  18   9   4   0]\n [122 414  98  15   4   1]\n [ 14  89 154  67  11   0]\n [  5  17  71 106  86  21]\n [  6  17  35  49 124  80]\n [  2   7  13  30  74 178]]\n```\nSo **going to large images doesn't give a substantial CV boost** so far, while LB is quite different.\n\n- Classification+Segmentation aux gave me only a tiny boost when I made a direct comparison:\n0.842 -&gt; 0.843 (4 fold CV on low res tiff layer)\n\nOne also may  check the related topic on [CV vs LB match](https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/145296)",
    "836153": "Great work. My hypothesis is that overfitting more easily occurs at lower magnifications. In my own experience, I have higher CV using level 1 vs level 0 (0.91 vs, 0.88), but LB is essentially the same (0.87). If you're using the same patch size for levels 1 and 2, then the patches for level 1 will have a much higher proportion of tissue vs. background. That might be one reason for CV-LB discrepancy. ",
    "842237": "model: se-resnet50\nfold: single fold \ncv-qwk: 0.894\nlb-qwk: 0.87",
    "840058": "A little update: resnet34 based architecture, 0.88 single fold val -&gt; 0.85 lb. 4 fold cv 0.87 -&gt; lb 0.87. It seems that the choice of architecture (or backbone) may not be the most important",
    "836585": "@iafoss what is the breakdown of QWK between the 2 institutions? I noticed that when my CV went from 0.88 to 0.91 with no change in LB, there was only CV improvement for Radboud data - Karolinska data performance was the same. ",
    "856434": "I have a lot of difference between local CV and LB\nBest single fold 0.8090CV 0.85LB\nSame model with 5 fold ensemble average (0.8012 average local CV) 0.86LB\n\nI did the tile selection a bit different than you and used to have more \"bright\" tiles. Currently running a model with your tile selection to see what it does I'm seeing an increase in local CV but will be interesting to see how it translates to LB.\n\nHaven't yet output the per provider results. Will do later or for another model run.",
    "854879": "model: resnext50\nfold: single fold\ncv-qwk : 0.877\nLb-qwk : 0.89",
    "904856": "resnet34, single fold, single model 0.91 LB",
    "857651": "```\nmodel: resnet34\nfold:     1\ncv:       0.884\nkarolinska: 0.8750\nradbound:  0.857\n```\n\n",
    "885139": "Can you please share what is the maximum number of tiles from intermediate layer with which you managed to train network?",
    "874494": "Actually I could boost single fold single model performance to 0.91 LB. Unfortunately, based on my previous subs, ensemble doesn't give much boost, and my expectation is getting only 0.01- improvement from it at single model performance of 0.90-0.91.",
    "860380": "I have LB 0.84 with efficientnet-b4 (CV 0.82). Efficientnet-b0 gave me LB 0.8 (CV 0.79). \n",
    "857585": "Any idea for improvement is welcome 😃 💪 \n\nFramework:\n\n- Tensorflow, Keras \n\nSingle fold:\n\n    X_train, X_val = train_test_split(train, test_size=.2, stratify=train['isup_grade'], random_state=SEED)\n\nPre processing: \n\n- 4X4 tiled images (384, 384, 3)\n- Augmentations (Hor/Ver Flip, ShiftScaleRotate)\n\nConfig:\n\n    LR: 1e-3 \n    BS: 16\n    Epoch: 40\n\nModel: Seresnext50 backbone <br> \nClassification: \n    \n    loss='categorical_crossentropy'\n    optimizer=optimizers.Adam(lr=LR)\n    metrics=[qw_kappa_score]\n    \nCallback= ReduceLROnPlateau <br>\n\n    CV: 0.78\n    LB: 0.79",
    "857095": "singlefold\nLB：0.82\nCV：0.88    Karolinska ：0.8387   Radboud：0.8806\n\nIt's funny!  Is LB all about Karolinska？：）\n\nUpdate:LB:0.84\nCV:0.87   Karolinska ：0.8640   Radboud：0.8477\n",
    "852885": "model: resnext50\nfold: single fold\ncv-qwk : 0.67\nLb-qwk : 0.60\n\nloss fn : CategoricalCrossentropy\noptimizer : adam\ninput : tile concatenated into 512X512X3\n\nit will be helpful if someone can share ideas on how i can improve my model next. \n",
    "849615": "How long does it take you to train a 5CV 20 epoch ? I feel my biggest pain right now is that testing ideas take like 9Hours on my GTX1080 ti (for 16x128x128) so it feels very slow. I guess I can always use a cloud solution.\n\nAlso any tip for faster iterations ?",
    "845825": "May I ask how many tiles you use to get 0.90 LB? \nI use 112x112x64 but it doesn't come up to higher score.",
    "842703": "@iafoss how many epochs do you train per fold? In my experiments, my model seems to need more than 30 epochs per fold to converge",
    "835690": "@iafoss are you using masks?",
    "855587": "model: single fold seresnext50\nCV: 0.9\nLB: 0.88",
    "835514": "Nice result @iafoss . If you don't mind, your tricks are related to model architecture or data augmentation/preprocessing ?",
    "842831": "Hi there all,\n\nis any one know how to translate this Pytorch code into fastai ?\n\n~~~\nfor i, (inputs, labels) in enumerate(training_set):\n    predictions = model(inputs)                     # Forward pass\n    loss = loss_function(predictions, labels)       # Compute loss function\n    loss = loss / accumulation_steps                # Normalize our loss (if averaged)\n    loss.backward()                                 # Backward pass\n    if (i+1) % accumulation_steps == 0:             # Wait for several backward steps\n        optimizer.step()                            # Now we can do an optimizer step\n        model.zero_grad()                           # Reset gradients tensors\n        if (i+1) % evaluation_steps == 0:           # Evaluate the model when we...\n            evaluate_model() \n~~~",
    "926069": "Hi @lafoss, I was wondering if can help me how to load my model from candence pretrainedmodels, I am still learning.\n~~~\nimport sys\nfrom pathlib import Path\n\nsys.path.append('../input/pytorch-pretrained-models/repository/pretrained-models.pytorch-master')\nimport pretrainedmodels\n\nm = pretrainedmodels.se_resnext50(pretrained='imagenet')\nchildren = list(m.children())\nhead = nn.Sequential(nn.AdaptiveAvgPool2d(1), Flatten(), \n                                  nn.Linear(children[-1].in_features,200))\nmodel = nn.Sequential(nn.Sequential(*children[:-2]), head)\n~~~\nThen tries to download but as we know we dont have internet, I have the model in the directory, but I dont know how to load it,\n\nthanks for the help.",
    "840802": "densenet121 architechture.",
    "928008": "Hi there my frind @lafoss I was wondering if you can give me a help one more time before this ends, I havent be able to submit anymore\n\nthe error is that  \"csv file not found.... \" ????? when I run it on draft everything works fine! as I do the commit and creates the cvs file for the submition,** Submission CSV Not Found**\n\nI was wondering if you see something wrong with the code?\n\n~~~\nTEST_PATH = '/kaggle/input/prostate-cancer-grade-assessment/test_images'\n\nif os.path.exists(TEST_PATH):\n    dls = dBlock.dataloaders(df_test, bs=8)\n\n    learn = Learner(dls, get_model())\n    learn.load('saved') \n\n    test_dl = dls.test_dl(df_test)\n    _,_, preds = learn.get_preds(dl=test_dl, with_decoded=True)\n\n    df_test[\"isup_grade\"] = preds\n    sub = df_test[[\"image_id\",\"isup_grade\"]]\n    sub.to_csv('submission.csv', index=False) \nelse:\n    df_train =  df_train.loc[:5]\n    dls = dBlock.dataloaders(df_train, bs=8)\n\n    learn = Learner(dls, get_model())\n    learn.load('saved') \n\n    train_dl = dls.test_dl(df_train)\n    _,_, preds = learn.get_preds(dl=train_dl, with_decoded=True)\n\n    df_train[\"isup_grade\"] = preds\n    sub = df_train[[\"image_id\",\"isup_grade\"]]\n    sub.to_csv('submission.csv', index=Fals\n~~~\nThanks a lot for all your help, I have learn a lot on this competition thanks to you.",
    "882507": "@iafoss can you tell me about your best CV and LB with lowest res tiles?",
    "867598": "Method: Based on lafoss tiling\nModel: ResNet18\nFold: Stratified\nLeaderboard: 0.84\nValidation: \n```\nQWP:0.8556\n[[487  68  16   1   3   0]\n [ 96 331  81   8   6   1]\n [  8  84 128  34  14   1]\n [  1  20  47  65  87  25]\n [  6  17  20  33  77  96]\n [  4   3   9  15  76 136]]\n```",
    "861402": "@iafoss is your 0.90 run from a single fold or multiple folds ensembled?",
    "859954": "single fold seresnext50\nCV: .810\nLB: .85",
    "854529": "model - Resnext50\nfold - single fold\nCV-QWK : 0.85\nLB-QWK: 0.81",
    "842197": "Question: X-fold CV -&gt; means it was not single model, right?",
    "842102": "I train 4-5 models in the same time using built and designated platform - cnvrg.io.\nIt also enables me to get the most accurate model automatically using their conditionals feature.\nI use their community version",
    "841695": "Impressive!",
    "840268": "Good to Learn ",
    "839789": "Thank you for this. I am new to Kaggle and this is going to be my first competition. \nIf I understand correctly are you not using the 'gleason_scores' for training? \nAre you only using the 'isup_grade' for training?",
    "839187": "👍 ",
    "839117": "I am working on it",
    "838973": "@iafoss Hi there,\n\n1.- I was wondering where do you train, on your personal computer? or kaggle?\n2.- What is number of tiles per image and what dimension (sz)?\n3.- What bs are you using?\n4.- and for how many epochs ?\n\nThanks!",
    "837672": "@iafoss I imagine you load images during training, so how do you do that in a fast way? I found that switching to 12x256x256 takes 1:30 to load the data from individually pickled files per one training cycle, which I feel is way too long",
    "835505": "well done, 95+ on horizon",
    "835372": "In my local experiments, my best single fold CV with TTA and checkpoint averaging can get up to 0.833 but somehow it does not transfer to lb (~0.75 with TTA and checkpoint averaging vs 0.78 without TTA). Maybe I have a bug in my code...",
    "835362": "Densenet121 based network on 12x128x128 tiles: 0.81~0.82 cv, 4 fold lb 0.78 on lowest resolution images. Impressive work as always btw.",
    "835356": "Thanks for sharing.\n\nI have one question, did you try regression, instead of classification?\nAs I learned from APTOS2019, the top solution used regression,\nwhat do you think?\n\nThanks again",
    "903613": "Method: Based on lafoss tiling + self modified trick\nModel: Efficientnet-B4 or  Efficientnet-B1\nFold: Stratified fold 0 out of 8 folds\nLeaderboard: 0.85\n\n",
    "893385": ""
  }
}