{
  "id": 165632,
  "title": "7 things that did not work",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/165632",
  "author_name": "Vlad Vaduva",
  "post_date": "2020-07-10T12:28:29.190000",
  "votes": 34,
  "comment_count": 36,
  "views": 0,
  "content": "<p>Among other things that worked that I will share after the competition ends, here are 7 things that for me did not work, they either give the same cv/public leaderboard result as without them or lower.</p>\n\n<ol>\n<li>Images with an amount of pixels larger than 2.500.000 (sum of all patches pixels).</li>\n<li>RangerLars optimizer (Really wanted to use it in this competition but it could not give me a boost. Probably I did not use it with compatible scheduler).</li>\n<li>Very big arhitectures -&gt; overfit.</li>\n<li>Batchsizes smaller than 3 did not work for me (the updates are too noisy)</li>\n<li>Tiles selected from the biggest resolution of the .tiff files</li>\n<li>Did not manage to proper use the masks provided (this is where I thing is a large space for improvement)</li>\n<li>Any augmentation except affine transforms</li>\n<li>Training individually per data provider</li>\n</ol>",
  "messages": [
    {
      "id": 922929,
      "postDate": "2020-07-10T12:28:29.190Z",
      "content": "<p>Among other things that worked that I will share after the competition ends, here are 7 things that for me did not work, they either give the same cv/public leaderboard result as without them or lower.</p>\n\n<ol>\n<li>Images with an amount of pixels larger than 2.500.000 (sum of all patches pixels).</li>\n<li>RangerLars optimizer (Really wanted to use it in this competition but it could not give me a boost. Probably I did not use it with compatible scheduler).</li>\n<li>Very big arhitectures -&gt; overfit.</li>\n<li>Batchsizes smaller than 3 did not work for me (the updates are too noisy)</li>\n<li>Tiles selected from the biggest resolution of the .tiff files</li>\n<li>Did not manage to proper use the masks provided (this is where I thing is a large space for improvement)</li>\n<li>Any augmentation except affine transforms</li>\n<li>Training individually per data provider</li>\n</ol>",
      "rawMarkdown": "Among other things that worked that I will share after the competition ends, here are 7 things that for me did not work, they either give the same cv/public leaderboard result as without them or lower.\n\n1. Images with an amount of pixels larger than 2.500.000 (sum of all patches pixels).\n2. RangerLars optimizer (Really wanted to use it in this competition but it could not give me a boost. Probably I did not use it with compatible scheduler).\n3. Very big arhitectures -&gt; overfit.\n4. Batchsizes smaller than 3 did not work for me (the updates are too noisy)\n5. Tiles selected from the biggest resolution of the .tiff files\n5. Did not manage to proper use the masks provided (this is where I thing is a large space for improvement)\n6. Any augmentation except affine transforms\n7. Training individually per data provider",
      "votes": 34
    },
    {
      "id": 926411,
      "postDate": "2020-07-12T17:45:35.483Z",
      "content": "<p>One more: class-balanced training didn't add anything for me. I undersampled (but resampled each epoch) and it neither hurt nor helped CV or LB scores.</p>\n\n<p>I think there is quite a lot to be gained from different approaches to the inference part of the problem. The models I have aren't amazing, they only score around 0.87 but useful approaches to inference take the score to 0.89. I haven't looked at post processing yet - I'm just straight averaging - but that's next on the to-do list.</p>",
      "rawMarkdown": "One more: class-balanced training didn't add anything for me. I undersampled (but resampled each epoch) and it neither hurt nor helped CV or LB scores.\n\nI think there is quite a lot to be gained from different approaches to the inference part of the problem. The models I have aren't amazing, they only score around 0.87 but useful approaches to inference take the score to 0.89. I haven't looked at post processing yet - I'm just straight averaging - but that's next on the to-do list.",
      "votes": 5
    },
    {
      "id": 924393,
      "postDate": "2020-07-11T12:07:37.423Z",
      "content": "<p>very useful,thanks for sharing</p>",
      "rawMarkdown": "very useful,thanks for sharing",
      "votes": 1,
      "replies": [
        {
          "id": 924841,
          "postDate": "2020-07-11T16:51:23.500Z",
          "content": "<p>It is my pleasure. Hope it helps</p>",
          "rawMarkdown": "It is my pleasure. Hope it helps",
          "votes": 1
        }
      ]
    },
    {
      "id": 922935,
      "postDate": "2020-07-10T12:31:47.200Z",
      "content": "<p>which scheduler worked best for you?did you try custom scheduler?</p>",
      "rawMarkdown": "which scheduler worked best for you?did you try custom scheduler?",
      "votes": 1,
      "replies": [
        {
          "id": 922939,
          "postDate": "2020-07-10T12:36:01.617Z",
          "content": "<p>I tested CosineAnnealingWarmRestarts, CosineAnnealingLR, OneCycleLR and started with a GradualWarmupScheduler for 1 or 2 epochs</p>",
          "rawMarkdown": "I tested CosineAnnealingWarmRestarts, CosineAnnealingLR, OneCycleLR and started with a GradualWarmupScheduler for 1 or 2 epochs",
          "votes": 2
        }
      ]
    },
    {
      "id": 924070,
      "postDate": "2020-07-11T08:38:31.673Z",
      "content": "<p>did you try TTA and Post Processing? or your best submission is a single model or ensemble without TTA and Post Processing?</p>",
      "rawMarkdown": "did you try TTA and Post Processing? or your best submission is a single model or ensemble without TTA and Post Processing?",
      "votes": 2,
      "replies": [
        {
          "id": 924845,
          "postDate": "2020-07-11T16:52:06.563Z",
          "content": "<p>TTA yes. My best submission is the average of 2 folds</p>",
          "rawMarkdown": "TTA yes. My best submission is the average of 2 folds",
          "votes": 1
        }
      ]
    },
    {
      "id": 923289,
      "postDate": "2020-07-10T17:33:17.713Z",
      "content": "<p>RangerLars and other RAdam-related optimizers typically work well with a flat+cosine decay (with cosine decasy starting about 70% in)</p>",
      "rawMarkdown": "RangerLars and other RAdam-related optimizers typically work well with a flat+cosine decay (with cosine decasy starting about 70% in)",
      "votes": 2,
      "replies": [
        {
          "id": 924847,
          "postDate": "2020-07-11T16:52:30.657Z",
          "content": "<p>Will put it on my todo list</p>",
          "rawMarkdown": "Will put it on my todo list",
          "votes": 1
        }
      ]
    },
    {
      "id": 923116,
      "postDate": "2020-07-10T14:36:15.730Z",
      "content": "<p>Thank you very much for sharing again! May I ask a few questions:\n1. Have you used batch accumulation with much success?\n2. Have you gotten any boost from ensembling models?\n3. Have you attempted to train with duplicates removed?</p>",
      "rawMarkdown": "Thank you very much for sharing again! May I ask a few questions:\n1. Have you used batch accumulation with much success?\n2. Have you gotten any boost from ensembling models?\n3. Have you attempted to train with duplicates removed?",
      "votes": 2,
      "replies": [
        {
          "id": 923120,
          "postDate": "2020-07-10T14:44:26.400Z",
          "content": "<p>Hi <a href=\"/greatgamedota\">@greatgamedota</a> . Glad to see that you are doing good, we are close neighbors in the leaderboard :)\nTo answer your questions:\n1. I did not had good results at other competition with  batch accumulation so I did not tried it now.\n2. I got a little + by combining multiple folds of the same model, I am currently  working now on combining multiple model folds\n3. Yes, but did not get any improvment</p>",
          "rawMarkdown": "Hi @greatgamedota . Glad to see that you are doing good, we are close neighbors in the leaderboard :)\nTo answer your questions:\n1. I did not had good results at other competition with  batch accumulation so I did not tried it now.\n2. I got a little + by combining multiple folds of the same model, I am currently  working now on combining multiple model folds\n3. Yes, but did not get any improvment",
          "votes": 3
        }
      ]
    },
    {
      "id": 922987,
      "postDate": "2020-07-10T13:14:30.663Z",
      "content": "<p>I found the same thing with points 1 and 3, they don't give that much of a performance increase, and you'll have to effectively lower the batch size.</p>\n\n<p>Another question: Did you also exclude samples, such as samples with pen marks?</p>",
      "rawMarkdown": "I found the same thing with points 1 and 3, they don't give that much of a performance increase, and you'll have to effectively lower the batch size.\n\nAnother question: Did you also exclude samples, such as samples with pen marks?",
      "votes": 2,
      "replies": [
        {
          "id": 922989,
          "postDate": "2020-07-10T13:15:34.017Z",
          "content": "<p>No, I did not exclude the pen marks samples</p>",
          "rawMarkdown": "No, I did not exclude the pen marks samples",
          "votes": 1
        }
      ]
    },
    {
      "id": 922952,
      "postDate": "2020-07-10T12:54:25.090Z",
      "content": "<p><a href=\"/vladvdv\">@vladvdv</a> Thanks for your wonderful results (and congrats to your position)! I am currently struggling to get a stable CV-LB correlation. With stratified KF on isup I can get LB ranging from .84 to .87 on the same CV .91. Would you care to kindly disclose how CV-LB works out for you?</p>",
      "rawMarkdown": "@vladvdv Thanks for your wonderful results (and congrats to your position)! I am currently struggling to get a stable CV-LB correlation. With stratified KF on isup I can get LB ranging from .84 to .87 on the same CV .91. Would you care to kindly disclose how CV-LB works out for you?",
      "votes": 2,
      "replies": [
        {
          "id": 922979,
          "postDate": "2020-07-10T13:11:10.573Z",
          "content": "<p>Thank you, unfortunately  I had been stuck on this score for over 3 weeks, hope to find a breakthrough  to get over 0.90.\nYou have a big discrepancy between LB and CV. My advice is to use TTA for prediction and use a smaller architecture  to generalize better. \nMy CV-LB are not fully correlated in this point but I learned some patterns that help me understand how to proper evaluate the LB with me CV score</p>",
          "rawMarkdown": "Thank you, unfortunately  I had been stuck on this score for over 3 weeks, hope to find a breakthrough  to get over 0.90.\nYou have a big discrepancy between LB and CV. My advice is to use TTA for prediction and use a smaller architecture  to generalize better. \nMy CV-LB are not fully correlated in this point but I learned some patterns that help me understand how to proper evaluate the LB with me CV score",
          "votes": 3
        },
        {
          "id": 923684,
          "postDate": "2020-07-11T04:43:24.077Z",
          "content": "<p>let me tell u it needs lot of patience to reach 0.90 . very very difficult. \nand then even more to 91 . </p>",
          "rawMarkdown": "let me tell u it needs lot of patience to reach 0.90 . very very difficult. \nand then even more to 91 . ",
          "votes": 2
        },
        {
          "id": 923883,
          "postDate": "2020-07-11T06:54:54.583Z",
          "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> <a href=\"/vladvdv\">@vladvdv</a> Thanks very much! Guess I would just have to try harder with more (not so many left) submissions. :P</p>",
          "rawMarkdown": "@jaideepvalani @vladvdv Thanks very much! Guess I would just have to try harder with more (not so many left) submissions. :P",
          "votes": 1
        },
        {
          "id": 924271,
          "postDate": "2020-07-11T11:02:43.700Z",
          "content": "<p><a href=\"/roguekk007\">@roguekk007</a> In my experience, stacking tiles into single square image(Qishen Ha's method) leads to much larger local/LB gap, than stacking tiles across batch dimension. Also, using groupnorm allows to get stable training with batch size 1.</p>",
          "rawMarkdown": "@roguekk007 In my experience, stacking tiles into single square image(Qishen Ha's method) leads to much larger local/LB gap, than stacking tiles across batch dimension. Also, using groupnorm allows to get stable training with batch size 1.\n",
          "votes": 2
        },
        {
          "id": 924850,
          "postDate": "2020-07-11T16:54:03.900Z",
          "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> for reaching 0.90 is was not a hard struggle. But for going over 0.90, I am struggling for over 3 weeks without success.</p>",
          "rawMarkdown": "@jaideepvalani for reaching 0.90 is was not a hard struggle. But for going over 0.90, I am struggling for over 3 weeks without success."
        },
        {
          "id": 926718,
          "postDate": "2020-07-12T23:29:55.487Z",
          "content": "<p><a href=\"/cateek\">@cateek</a> Wonderful! Been thinking about GN myself, though discouraged by the slow training. Thanks</p>",
          "rawMarkdown": "@cateek Wonderful! Been thinking about GN myself, though discouraged by the slow training. Thanks"
        },
        {
          "id": 929612,
          "postDate": "2020-07-14T19:45:04.810Z",
          "content": "<p><a href=\"/cateek\">@cateek</a>  i tried groupnormalization in pytorch using quishen ha's kernel but i get 0.0 qwk score sometimes and first few epochs performance are worst compared to batchnorm layers,,did you try groupnorm in tensorflow? will you please share your train log of groupnorm here? thank you</p>\n\n<p>today 1 of our model got cv ~ 0.92 and lb 0.85 :(\nit is the same validation schema like  quishen ha used,we only trained on fold = 1</p>\n\n<p>yes it is a small model like @Vlad suggesting for better generalization</p>\n\n<p>selecting best 2 models going to be very difficult for us </p>",
          "rawMarkdown": "@cateek  i tried groupnormalization in pytorch using quishen ha's kernel but i get 0.0 qwk score sometimes and first few epochs performance are worst compared to batchnorm layers,,did you try groupnorm in tensorflow? will you please share your train log of groupnorm here? thank you\n\ntoday 1 of our model got cv ~ 0.92 and lb 0.85 :(\nit is the same validation schema like  quishen ha used,we only trained on fold = 1\n\nyes it is a small model like @Vlad suggesting for better generalization\n\nselecting best 2 models going to be very difficult for us "
        },
        {
          "id": 929680,
          "postDate": "2020-07-14T20:59:08.030Z",
          "content": "<p>I'm getting something like this \n<code>\nep:0 loss:0.6784 kappa_score:0.5503\nep:1 loss:0.5104 kappa_score:0.7474\nep:2 loss:0.4946 kappa_score:0.7153\nep:3 loss:0.4903 kappa_score:0.7274\nep:4 loss:0.4357 kappa_score:0.7593\n</code>\nI'm using PyTorch. This model achieves 0.9/0.89. Posted link in external data thread. \nP.S. if you're using some BN/GN in classification head, then just don't do it:) And as I said, tiling approach gives much more consistent correlation and lower LB gap</p>",
          "rawMarkdown": "I'm getting something like this \n```\nep:0 loss:0.6784 kappa_score:0.5503\nep:1 loss:0.5104 kappa_score:0.7474\nep:2 loss:0.4946 kappa_score:0.7153\nep:3 loss:0.4903 kappa_score:0.7274\nep:4 loss:0.4357 kappa_score:0.7593\n```\nI'm using PyTorch. This model achieves 0.9/0.89. Posted link in external data thread. \nP.S. if you're using some BN/GN in classification head, then just don't do it:) And as I said, tiling approach gives much more consistent correlation and lower LB gap",
          "votes": 1
        },
        {
          "id": 930247,
          "postDate": "2020-07-15T09:51:26.230Z",
          "content": "<p><a href=\"/cateek\">@cateek</a>  thank you!!\nimpressive train log\nmy groupnorm never gets above 0.5 qwk score</p>\n\n<p>i was simply replacing all the batchnorm layers to groupnorm layers from torchvision models\nfor example took resnet50 from torchvision and simply replaced all batchnorm layers to groupnorm layers and i got terrible  performance!</p>\n\n<p>i am curious- how you use groupnorm in pretrained model that is not breaking your network for this competition!</p>",
          "rawMarkdown": "@cateek  thank you!!\nimpressive train log\nmy groupnorm never gets above 0.5 qwk score\n\ni was simply replacing all the batchnorm layers to groupnorm layers from torchvision models\nfor example took resnet50 from torchvision and simply replaced all batchnorm layers to groupnorm layers and i got terrible  performance!\n\ni am curious- how you use groupnorm in pretrained model that is not breaking your network for this competition!\n\n\n"
        },
        {
          "id": 930269,
          "postDate": "2020-07-15T10:09:58.437Z",
          "content": "<p><a href=\"/cateek\">@cateek</a>  for example this is my model where if i replace all batchnorm to groupnorm then i get terrible model : </p>\n\n<p><code>enetv2(\n  (enet): ResNet(\n    (conv1): myConv2d(3, 64, kernel_size=(7, 7), stride=(2, 2), padding=(3, 3), bias=False)\n    (bn1): BatchNorm2d(64, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n    (act1): ReLU(inplace=True)\n    (maxpool): MaxPool2d(kernel_size=3, stride=2, padding=1, dilation=1, ceil_mode=False)\n    (layer1): Sequential(\n      (0): Bottleneck(\n        (conv1): myConv2d(64, 128, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn1): BatchNorm2d(128, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act1): ReLU(inplace=True)\n        (conv2): myConv2d(128, 128, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), groups=32, bias=False)\n        (bn2): BatchNorm2d(128, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act2): ReLU(inplace=True)\n        (conv3): myConv2d(128, 256, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn3): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act3): ReLU(inplace=True)\n        (downsample): Sequential(\n          (0): myConv2d(64, 256, kernel_size=(1, 1), stride=(1, 1), bias=False)\n          (1): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        )\n      )\n      (1): Bottleneck(\n        (conv1): myConv2d(256, 128, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn1): BatchNorm2d(128, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act1): ReLU(inplace=True)\n        (conv2): myConv2d(128, 128, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), groups=32, bias=False)\n        (bn2): BatchNorm2d(128, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act2): ReLU(inplace=True)\n        (conv3): myConv2d(128, 256, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn3): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act3): ReLU(inplace=True)\n      )\n      (2): Bottleneck(\n        (conv1): myConv2d(256, 128, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn1): BatchNorm2d(128, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act1): ReLU(inplace=True)\n        (conv2): myConv2d(128, 128, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), groups=32, bias=False)\n        (bn2): BatchNorm2d(128, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act2): ReLU(inplace=True)\n        (conv3): myConv2d(128, 256, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn3): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act3): ReLU(inplace=True)\n      )\n    )\n    (layer2): Sequential(\n      (0): Bottleneck(\n        (conv1): myConv2d(256, 256, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn1): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act1): ReLU(inplace=True)\n        (conv2): myConv2d(256, 256, kernel_size=(3, 3), stride=(2, 2), padding=(1, 1), groups=32, bias=False)\n        (bn2): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act2): ReLU(inplace=True)\n        (conv3): myConv2d(256, 512, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn3): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act3): ReLU(inplace=True)\n        (downsample): Sequential(\n          (0): myConv2d(256, 512, kernel_size=(1, 1), stride=(2, 2), bias=False)\n          (1): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        )\n      )\n      (1): Bottleneck(\n        (conv1): myConv2d(512, 256, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn1): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act1): ReLU(inplace=True)\n        (conv2): myConv2d(256, 256, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), groups=32, bias=False)\n        (bn2): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act2): ReLU(inplace=True)\n        (conv3): myConv2d(256, 512, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn3): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act3): ReLU(inplace=True)\n      )\n      (2): Bottleneck(\n        (conv1): myConv2d(512, 256, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn1): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act1): ReLU(inplace=True)\n        (conv2): myConv2d(256, 256, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), groups=32, bias=False)\n        (bn2): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act2): ReLU(inplace=True)\n        (conv3): myConv2d(256, 512, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn3): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act3): ReLU(inplace=True)\n      )\n      (3): Bottleneck(\n        (conv1): myConv2d(512, 256, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn1): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act1): ReLU(inplace=True)\n        (conv2): myConv2d(256, 256, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), groups=32, bias=False)\n        (bn2): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act2): ReLU(inplace=True)\n        (conv3): myConv2d(256, 512, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn3): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act3): ReLU(inplace=True)\n      )\n    )\n    (layer3): Sequential(\n      (0): Bottleneck(\n        (conv1): myConv2d(512, 512, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn1): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act1): ReLU(inplace=True)\n        (conv2): myConv2d(512, 512, kernel_size=(3, 3), stride=(2, 2), padding=(1, 1), groups=32, bias=False)\n        (bn2): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act2): ReLU(inplace=True)\n        (conv3): myConv2d(512, 1024, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn3): BatchNorm2d(1024, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act3): ReLU(inplace=True)\n        (downsample): Sequential(\n          (0): myConv2d(512, 1024, kernel_size=(1, 1), stride=(2, 2), bias=False)\n          (1): BatchNorm2d(1024, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        )\n      )\n      (1): Bottleneck(\n        (conv1): myConv2d(1024, 512, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn1): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act1): ReLU(inplace=True)\n        (conv2): myConv2d(512, 512, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), groups=32, bias=False)\n        (bn2): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act2): ReLU(inplace=True)\n        (conv3): myConv2d(512, 1024, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn3): BatchNorm2d(1024, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act3): ReLU(inplace=True)\n      )\n      (2): Bottleneck(\n        (conv1): myConv2d(1024, 512, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn1): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act1): ReLU(inplace=True)\n        (conv2): myConv2d(512, 512, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), groups=32, bias=False)\n        (bn2): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act2): ReLU(inplace=True)\n        (conv3): myConv2d(512, 1024, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn3): BatchNorm2d(1024, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act3): ReLU(inplace=True)\n      )\n      (3): Bottleneck(\n        (conv1): myConv2d(1024, 512, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn1): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act1): ReLU(inplace=True)\n        (conv2): myConv2d(512, 512, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), groups=32, bias=False)\n        (bn2): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act2): ReLU(inplace=True)\n        (conv3): myConv2d(512, 1024, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn3): BatchNorm2d(1024, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act3): ReLU(inplace=True)\n      )\n      (4): Bottleneck(\n        (conv1): myConv2d(1024, 512, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn1): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act1): ReLU(inplace=True)\n        (conv2): myConv2d(512, 512, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), groups=32, bias=False)\n        (bn2): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act2): ReLU(inplace=True)\n        (conv3): myConv2d(512, 1024, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn3): BatchNorm2d(1024, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act3): ReLU(inplace=True)\n      )\n      (5): Bottleneck(\n        (conv1): myConv2d(1024, 512, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn1): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act1): ReLU(inplace=True)\n        (conv2): myConv2d(512, 512, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), groups=32, bias=False)\n        (bn2): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act2): ReLU(inplace=True)\n        (conv3): myConv2d(512, 1024, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn3): BatchNorm2d(1024, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act3): ReLU(inplace=True)\n      )\n    )\n    (layer4): Sequential(\n      (0): Bottleneck(\n        (conv1): myConv2d(1024, 1024, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn1): BatchNorm2d(1024, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act1): ReLU(inplace=True)\n        (conv2): myConv2d(1024, 1024, kernel_size=(3, 3), stride=(2, 2), padding=(1, 1), groups=32, bias=False)\n        (bn2): BatchNorm2d(1024, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act2): ReLU(inplace=True)\n        (conv3): myConv2d(1024, 2048, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn3): BatchNorm2d(2048, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act3): ReLU(inplace=True)\n        (downsample): Sequential(\n          (0): myConv2d(1024, 2048, kernel_size=(1, 1), stride=(2, 2), bias=False)\n          (1): BatchNorm2d(2048, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        )\n      )\n      (1): Bottleneck(\n        (conv1): myConv2d(2048, 1024, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn1): BatchNorm2d(1024, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act1): ReLU(inplace=True)\n        (conv2): myConv2d(1024, 1024, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), groups=32, bias=False)\n        (bn2): BatchNorm2d(1024, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act2): ReLU(inplace=True)\n        (conv3): myConv2d(1024, 2048, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn3): BatchNorm2d(2048, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act3): ReLU(inplace=True)\n      )\n      (2): Bottleneck(\n        (conv1): myConv2d(2048, 1024, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn1): BatchNorm2d(1024, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act1): ReLU(inplace=True)\n        (conv2): myConv2d(1024, 1024, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), groups=32, bias=False)\n        (bn2): BatchNorm2d(1024, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act2): ReLU(inplace=True)\n        (conv3): myConv2d(1024, 2048, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn3): BatchNorm2d(2048, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act3): ReLU(inplace=True)\n      )\n    )\n    (global_pool): SelectAdaptivePool2d (output_size=1, pool_type=avg)\n    (fc): Identity()\n  )\n  (myfc): Linear(in_features=2048, out_features=5, bias=True)\n)</code></p>",
          "rawMarkdown": "@cateek  for example this is my model where if i replace all batchnorm to groupnorm then i get terrible model : \n\n`enetv2(\n  (enet): ResNet(\n    (conv1): myConv2d(3, 64, kernel_size=(7, 7), stride=(2, 2), padding=(3, 3), bias=False)\n    (bn1): BatchNorm2d(64, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n    (act1): ReLU(inplace=True)\n    (maxpool): MaxPool2d(kernel_size=3, stride=2, padding=1, dilation=1, ceil_mode=False)\n    (layer1): Sequential(\n      (0): Bottleneck(\n        (conv1): myConv2d(64, 128, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn1): BatchNorm2d(128, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act1): ReLU(inplace=True)\n        (conv2): myConv2d(128, 128, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), groups=32, bias=False)\n        (bn2): BatchNorm2d(128, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act2): ReLU(inplace=True)\n        (conv3): myConv2d(128, 256, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn3): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act3): ReLU(inplace=True)\n        (downsample): Sequential(\n          (0): myConv2d(64, 256, kernel_size=(1, 1), stride=(1, 1), bias=False)\n          (1): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        )\n      )\n      (1): Bottleneck(\n        (conv1): myConv2d(256, 128, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn1): BatchNorm2d(128, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act1): ReLU(inplace=True)\n        (conv2): myConv2d(128, 128, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), groups=32, bias=False)\n        (bn2): BatchNorm2d(128, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act2): ReLU(inplace=True)\n        (conv3): myConv2d(128, 256, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn3): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act3): ReLU(inplace=True)\n      )\n      (2): Bottleneck(\n        (conv1): myConv2d(256, 128, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn1): BatchNorm2d(128, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act1): ReLU(inplace=True)\n        (conv2): myConv2d(128, 128, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), groups=32, bias=False)\n        (bn2): BatchNorm2d(128, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act2): ReLU(inplace=True)\n        (conv3): myConv2d(128, 256, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn3): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act3): ReLU(inplace=True)\n      )\n    )\n    (layer2): Sequential(\n      (0): Bottleneck(\n        (conv1): myConv2d(256, 256, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn1): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act1): ReLU(inplace=True)\n        (conv2): myConv2d(256, 256, kernel_size=(3, 3), stride=(2, 2), padding=(1, 1), groups=32, bias=False)\n        (bn2): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act2): ReLU(inplace=True)\n        (conv3): myConv2d(256, 512, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn3): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act3): ReLU(inplace=True)\n        (downsample): Sequential(\n          (0): myConv2d(256, 512, kernel_size=(1, 1), stride=(2, 2), bias=False)\n          (1): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        )\n      )\n      (1): Bottleneck(\n        (conv1): myConv2d(512, 256, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn1): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act1): ReLU(inplace=True)\n        (conv2): myConv2d(256, 256, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), groups=32, bias=False)\n        (bn2): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act2): ReLU(inplace=True)\n        (conv3): myConv2d(256, 512, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn3): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act3): ReLU(inplace=True)\n      )\n      (2): Bottleneck(\n        (conv1): myConv2d(512, 256, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn1): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act1): ReLU(inplace=True)\n        (conv2): myConv2d(256, 256, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), groups=32, bias=False)\n        (bn2): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act2): ReLU(inplace=True)\n        (conv3): myConv2d(256, 512, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn3): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act3): ReLU(inplace=True)\n      )\n      (3): Bottleneck(\n        (conv1): myConv2d(512, 256, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn1): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act1): ReLU(inplace=True)\n        (conv2): myConv2d(256, 256, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), groups=32, bias=False)\n        (bn2): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act2): ReLU(inplace=True)\n        (conv3): myConv2d(256, 512, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn3): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act3): ReLU(inplace=True)\n      )\n    )\n    (layer3): Sequential(\n      (0): Bottleneck(\n        (conv1): myConv2d(512, 512, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn1): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act1): ReLU(inplace=True)\n        (conv2): myConv2d(512, 512, kernel_size=(3, 3), stride=(2, 2), padding=(1, 1), groups=32, bias=False)\n        (bn2): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act2): ReLU(inplace=True)\n        (conv3): myConv2d(512, 1024, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn3): BatchNorm2d(1024, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act3): ReLU(inplace=True)\n        (downsample): Sequential(\n          (0): myConv2d(512, 1024, kernel_size=(1, 1), stride=(2, 2), bias=False)\n          (1): BatchNorm2d(1024, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        )\n      )\n      (1): Bottleneck(\n        (conv1): myConv2d(1024, 512, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn1): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act1): ReLU(inplace=True)\n        (conv2): myConv2d(512, 512, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), groups=32, bias=False)\n        (bn2): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act2): ReLU(inplace=True)\n        (conv3): myConv2d(512, 1024, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn3): BatchNorm2d(1024, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act3): ReLU(inplace=True)\n      )\n      (2): Bottleneck(\n        (conv1): myConv2d(1024, 512, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn1): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act1): ReLU(inplace=True)\n        (conv2): myConv2d(512, 512, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), groups=32, bias=False)\n        (bn2): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act2): ReLU(inplace=True)\n        (conv3): myConv2d(512, 1024, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn3): BatchNorm2d(1024, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act3): ReLU(inplace=True)\n      )\n      (3): Bottleneck(\n        (conv1): myConv2d(1024, 512, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn1): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act1): ReLU(inplace=True)\n        (conv2): myConv2d(512, 512, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), groups=32, bias=False)\n        (bn2): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act2): ReLU(inplace=True)\n        (conv3): myConv2d(512, 1024, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn3): BatchNorm2d(1024, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act3): ReLU(inplace=True)\n      )\n      (4): Bottleneck(\n        (conv1): myConv2d(1024, 512, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn1): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act1): ReLU(inplace=True)\n        (conv2): myConv2d(512, 512, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), groups=32, bias=False)\n        (bn2): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act2): ReLU(inplace=True)\n        (conv3): myConv2d(512, 1024, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn3): BatchNorm2d(1024, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act3): ReLU(inplace=True)\n      )\n      (5): Bottleneck(\n        (conv1): myConv2d(1024, 512, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn1): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act1): ReLU(inplace=True)\n        (conv2): myConv2d(512, 512, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), groups=32, bias=False)\n        (bn2): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act2): ReLU(inplace=True)\n        (conv3): myConv2d(512, 1024, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn3): BatchNorm2d(1024, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act3): ReLU(inplace=True)\n      )\n    )\n    (layer4): Sequential(\n      (0): Bottleneck(\n        (conv1): myConv2d(1024, 1024, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn1): BatchNorm2d(1024, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act1): ReLU(inplace=True)\n        (conv2): myConv2d(1024, 1024, kernel_size=(3, 3), stride=(2, 2), padding=(1, 1), groups=32, bias=False)\n        (bn2): BatchNorm2d(1024, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act2): ReLU(inplace=True)\n        (conv3): myConv2d(1024, 2048, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn3): BatchNorm2d(2048, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act3): ReLU(inplace=True)\n        (downsample): Sequential(\n          (0): myConv2d(1024, 2048, kernel_size=(1, 1), stride=(2, 2), bias=False)\n          (1): BatchNorm2d(2048, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        )\n      )\n      (1): Bottleneck(\n        (conv1): myConv2d(2048, 1024, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn1): BatchNorm2d(1024, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act1): ReLU(inplace=True)\n        (conv2): myConv2d(1024, 1024, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), groups=32, bias=False)\n        (bn2): BatchNorm2d(1024, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act2): ReLU(inplace=True)\n        (conv3): myConv2d(1024, 2048, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn3): BatchNorm2d(2048, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act3): ReLU(inplace=True)\n      )\n      (2): Bottleneck(\n        (conv1): myConv2d(2048, 1024, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn1): BatchNorm2d(1024, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act1): ReLU(inplace=True)\n        (conv2): myConv2d(1024, 1024, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), groups=32, bias=False)\n        (bn2): BatchNorm2d(1024, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act2): ReLU(inplace=True)\n        (conv3): myConv2d(1024, 2048, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn3): BatchNorm2d(2048, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act3): ReLU(inplace=True)\n      )\n    )\n    (global_pool): SelectAdaptivePool2d (output_size=1, pool_type=avg)\n    (fc): Identity()\n  )\n  (myfc): Linear(in_features=2048, out_features=5, bias=True)\n)`\n\n"
        },
        {
          "id": 930329,
          "postDate": "2020-07-15T11:13:08.503Z",
          "content": "<p><a href=\"/mobassir\">@mobassir</a> I posted github link in external data thread. It has a couple of models with imagenet weights. If you just replace all BN layers with GN, you will definitely break pretrained weights. </p>",
          "rawMarkdown": "@mobassir I posted github link in external data thread. It has a couple of models with imagenet weights. If you just replace all BN layers with GN, you will definitely break pretrained weights. "
        },
        {
          "id": 930893,
          "postDate": "2020-07-15T19:43:43.390Z",
          "content": "<p><a href=\"/cateek\">@cateek</a> </p>\n\n<p>i tried resnet50 from here : <a href=\"https://github.com/joe-siyuan-qiao/pytorch-classification/tree/e6355f829e85ac05a71b8889f4fff77b9ab95d0b\">https://github.com/joe-siyuan-qiao/pytorch-classification/tree/e6355f829e85ac05a71b8889f4fff77b9ab95d0b</a> as you posted in external data thread,,in debug mode i can see qwk 0.00 again as like before!! just simply  changed quishen ha's efficientenet b0 with resnet50 from  that link : <a href=\"https://github.com/joe-siyuan-qiao/pytorch-classification/tree/e6355f829e85ac05a71b8889f4fff77b9ab95d0b\">https://github.com/joe-siyuan-qiao/pytorch-classification/tree/e6355f829e85ac05a71b8889f4fff77b9ab95d0b</a>\nin debug mode : \nEpoch 1, lr: 0.0000300, train loss: 0.61534, val loss: 0.60353, acc: 10.00000, qwk: 0.00000\nEpoch 2, lr: 0.0003000, train loss: 0.63444, val loss: 0.61907, acc: 10.00000, qwk: 0.00000</p>",
          "rawMarkdown": "@cateek \n\ni tried resnet50 from here : https://github.com/joe-siyuan-qiao/pytorch-classification/tree/e6355f829e85ac05a71b8889f4fff77b9ab95d0b as you posted in external data thread,,in debug mode i can see qwk 0.00 again as like before!! just simply  changed quishen ha's efficientenet b0 with resnet50 from  that link : https://github.com/joe-siyuan-qiao/pytorch-classification/tree/e6355f829e85ac05a71b8889f4fff77b9ab95d0b\nin debug mode : \nEpoch 1, lr: 0.0000300, train loss: 0.61534, val loss: 0.60353, acc: 10.00000, qwk: 0.00000\nEpoch 2, lr: 0.0003000, train loss: 0.63444, val loss: 0.61907, acc: 10.00000, qwk: 0.00000"
        },
        {
          "id": 930899,
          "postDate": "2020-07-15T19:52:45.767Z",
          "content": "<p><a href=\"/cateek\">@cateek</a>  also not  able to download weights\nthis link : <a href=\"http://cs.jhu.edu/~syqiao/WeightStandardization/R-50-GN-WS.pth.tar\">http://cs.jhu.edu/~syqiao/WeightStandardization/R-50-GN-WS.pth.tar</a> doesn't work \njust shows \"untitled\" page and never starts download\nfrom where you downloaded weights?</p>",
          "rawMarkdown": "@cateek  also not  able to download weights\nthis link : http://cs.jhu.edu/~syqiao/WeightStandardization/R-50-GN-WS.pth.tar doesn't work \njust shows \"untitled\" page and never starts download\nfrom where you downloaded weights?"
        },
        {
          "id": 930931,
          "postDate": "2020-07-15T20:21:54.500Z",
          "content": "<p><a href=\"/mobassir\">@mobassir</a> Weights link is correct, although I used ResNeXt-50. No idea why you're getting zeros in qwk, there must be some bug</p>",
          "rawMarkdown": "@mobassir Weights link is correct, although I used ResNeXt-50. No idea why you're getting zeros in qwk, there must be some bug"
        },
        {
          "id": 931504,
          "postDate": "2020-07-16T08:35:33.763Z",
          "content": "<p><a href=\"/cateek\">@cateek</a>  using resnet50 from that repo when i tried loading weight file he provided i get this error : <a href=\"https://discuss.pytorch.org/t/how-to-load-weights-of-weight-standardization-model/89418\">https://discuss.pytorch.org/t/how-to-load-weights-of-weight-standardization-model/89418</a></p>\n\n<p>i downloaded resnet50 weight  by following link he provided</p>",
          "rawMarkdown": "@cateek  using resnet50 from that repo when i tried loading weight file he provided i get this error : https://discuss.pytorch.org/t/how-to-load-weights-of-weight-standardization-model/89418\n\ni downloaded resnet50 weight  by following link he provided"
        },
        {
          "id": 931526,
          "postDate": "2020-07-16T08:48:28.613Z",
          "content": "<p>Well, haven't encountered anything like that. </p>",
          "rawMarkdown": "Well, haven't encountered anything like that. ",
          "votes": 1
        },
        {
          "id": 931527,
          "postDate": "2020-07-16T08:50:12.703Z",
          "content": "<p>sorry <a href=\"/cateek\">@cateek</a>  i disturbed you a lot in this discussion post and  thank you for answering everything,i will try resnext50,thanks</p>",
          "rawMarkdown": "sorry @cateek  i disturbed you a lot in this discussion post and  thank you for answering everything,i will try resnext50,thanks"
        }
      ]
    },
    {
      "id": 922933,
      "postDate": "2020-07-10T12:30:57.860Z",
      "content": "<p>Thanks for sharing your result.</p>",
      "rawMarkdown": "Thanks for sharing your result.",
      "votes": 2,
      "replies": [
        {
          "id": 922940,
          "postDate": "2020-07-10T12:36:50.177Z",
          "content": "<p>No problem. This is how we all learn things, by sharing and discussing about them !</p>",
          "rawMarkdown": "No problem. This is how we all learn things, by sharing and discussing about them !",
          "votes": 5
        }
      ]
    },
    {
      "id": 927378,
      "postDate": "2020-07-13T10:52:03.173Z",
      "content": "<p>well done</p>",
      "rawMarkdown": "well done"
    },
    {
      "id": 925280,
      "postDate": "2020-07-12T01:14:27.290Z",
      "content": "<p>Good to know! Totally agree with 4 :D\nMay I ask you how many epochs do you train for? For me, the valid-loss fluctuate in 20~30 epochs while train-loss keep decreasing. So I end training in 30 epoch with valid-loss almost equal to 6~7 x train-loss(local score ~0.88 and lb also 0.88, single fold). The loss I used is BCEWithLogitsLoss like Qishen Ha did.(but not the same on changing pred to label)</p>",
      "rawMarkdown": "Good to know! Totally agree with 4 :D\nMay I ask you how many epochs do you train for? For me, the valid-loss fluctuate in 20~30 epochs while train-loss keep decreasing. So I end training in 30 epoch with valid-loss almost equal to 6~7 x train-loss(local score ~0.88 and lb also 0.88, single fold). The loss I used is BCEWithLogitsLoss like Qishen Ha did.(but not the same on changing pred to label)"
    },
    {
      "id": 928335,
      "postDate": "2020-07-13T22:21:32.190Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 926411,
      "author_name": "fergusoci",
      "author_url": "",
      "post_date": "2020-07-12T17:45:35.483000",
      "content": "<p>One more: class-balanced training didn't add anything for me. I undersampled (but resampled each epoch) and it neither hurt nor helped CV or LB scores.</p>\n\n<p>I think there is quite a lot to be gained from different approaches to the inference part of the problem. The models I have aren't amazing, they only score around 0.87 but useful approaches to inference take the score to 0.89. I haven't looked at post processing yet - I'm just straight averaging - but that's next on the to-do list.</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 924393,
      "author_name": "leilaSoleymani",
      "author_url": "",
      "post_date": "2020-07-11T12:07:37.423000",
      "content": "<p>very useful,thanks for sharing</p>",
      "votes": 1,
      "replies": [
        {
          "id": 924841,
          "author_name": "Vlad Vaduva",
          "author_url": "",
          "post_date": "2020-07-11T16:51:23.500000",
          "content": "<p>It is my pleasure. Hope it helps</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 922935,
      "author_name": "Mobassir",
      "author_url": "",
      "post_date": "2020-07-10T12:31:47.200000",
      "content": "<p>which scheduler worked best for you?did you try custom scheduler?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 922939,
          "author_name": "Vlad Vaduva",
          "author_url": "",
          "post_date": "2020-07-10T12:36:01.617000",
          "content": "<p>I tested CosineAnnealingWarmRestarts, CosineAnnealingLR, OneCycleLR and started with a GradualWarmupScheduler for 1 or 2 epochs</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 924070,
      "author_name": "Mobassir",
      "author_url": "",
      "post_date": "2020-07-11T08:38:31.673000",
      "content": "<p>did you try TTA and Post Processing? or your best submission is a single model or ensemble without TTA and Post Processing?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 924845,
          "author_name": "Vlad Vaduva",
          "author_url": "",
          "post_date": "2020-07-11T16:52:06.563000",
          "content": "<p>TTA yes. My best submission is the average of 2 folds</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 923289,
      "author_name": "ilovescience",
      "author_url": "",
      "post_date": "2020-07-10T17:33:17.713000",
      "content": "<p>RangerLars and other RAdam-related optimizers typically work well with a flat+cosine decay (with cosine decasy starting about 70% in)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 924847,
          "author_name": "Vlad Vaduva",
          "author_url": "",
          "post_date": "2020-07-11T16:52:30.657000",
          "content": "<p>Will put it on my todo list</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 923116,
      "author_name": "GreatGameDota",
      "author_url": "",
      "post_date": "2020-07-10T14:36:15.730000",
      "content": "<p>Thank you very much for sharing again! May I ask a few questions:\n1. Have you used batch accumulation with much success?\n2. Have you gotten any boost from ensembling models?\n3. Have you attempted to train with duplicates removed?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 923120,
          "author_name": "Vlad Vaduva",
          "author_url": "",
          "post_date": "2020-07-10T14:44:26.400000",
          "content": "<p>Hi <a href=\"/greatgamedota\">@greatgamedota</a> . Glad to see that you are doing good, we are close neighbors in the leaderboard :)\nTo answer your questions:\n1. I did not had good results at other competition with  batch accumulation so I did not tried it now.\n2. I got a little + by combining multiple folds of the same model, I am currently  working now on combining multiple model folds\n3. Yes, but did not get any improvment</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 922987,
      "author_name": "Stephan",
      "author_url": "",
      "post_date": "2020-07-10T13:14:30.663000",
      "content": "<p>I found the same thing with points 1 and 3, they don't give that much of a performance increase, and you'll have to effectively lower the batch size.</p>\n\n<p>Another question: Did you also exclude samples, such as samples with pen marks?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 922989,
          "author_name": "Vlad Vaduva",
          "author_url": "",
          "post_date": "2020-07-10T13:15:34.017000",
          "content": "<p>No, I did not exclude the pen marks samples</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 922952,
      "author_name": "Nicholas Lyu",
      "author_url": "",
      "post_date": "2020-07-10T12:54:25.090000",
      "content": "<p><a href=\"/vladvdv\">@vladvdv</a> Thanks for your wonderful results (and congrats to your position)! I am currently struggling to get a stable CV-LB correlation. With stratified KF on isup I can get LB ranging from .84 to .87 on the same CV .91. Would you care to kindly disclose how CV-LB works out for you?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 922979,
          "author_name": "Vlad Vaduva",
          "author_url": "",
          "post_date": "2020-07-10T13:11:10.573000",
          "content": "<p>Thank you, unfortunately  I had been stuck on this score for over 3 weeks, hope to find a breakthrough  to get over 0.90.\nYou have a big discrepancy between LB and CV. My advice is to use TTA for prediction and use a smaller architecture  to generalize better. \nMy CV-LB are not fully correlated in this point but I learned some patterns that help me understand how to proper evaluate the LB with me CV score</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 923684,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-07-11T04:43:24.077000",
          "content": "<p>let me tell u it needs lot of patience to reach 0.90 . very very difficult. \nand then even more to 91 . </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 923883,
          "author_name": "Nicholas Lyu",
          "author_url": "",
          "post_date": "2020-07-11T06:54:54.583000",
          "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> <a href=\"/vladvdv\">@vladvdv</a> Thanks very much! Guess I would just have to try harder with more (not so many left) submissions. :P</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 924271,
          "author_name": "Eek The Cat",
          "author_url": "",
          "post_date": "2020-07-11T11:02:43.700000",
          "content": "<p><a href=\"/roguekk007\">@roguekk007</a> In my experience, stacking tiles into single square image(Qishen Ha's method) leads to much larger local/LB gap, than stacking tiles across batch dimension. Also, using groupnorm allows to get stable training with batch size 1.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 924850,
          "author_name": "Vlad Vaduva",
          "author_url": "",
          "post_date": "2020-07-11T16:54:03.900000",
          "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> for reaching 0.90 is was not a hard struggle. But for going over 0.90, I am struggling for over 3 weeks without success.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 926718,
          "author_name": "Nicholas Lyu",
          "author_url": "",
          "post_date": "2020-07-12T23:29:55.487000",
          "content": "<p><a href=\"/cateek\">@cateek</a> Wonderful! Been thinking about GN myself, though discouraged by the slow training. Thanks</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 929612,
          "author_name": "Mobassir",
          "author_url": "",
          "post_date": "2020-07-14T19:45:04.810000",
          "content": "<p><a href=\"/cateek\">@cateek</a>  i tried groupnormalization in pytorch using quishen ha's kernel but i get 0.0 qwk score sometimes and first few epochs performance are worst compared to batchnorm layers,,did you try groupnorm in tensorflow? will you please share your train log of groupnorm here? thank you</p>\n\n<p>today 1 of our model got cv ~ 0.92 and lb 0.85 :(\nit is the same validation schema like  quishen ha used,we only trained on fold = 1</p>\n\n<p>yes it is a small model like @Vlad suggesting for better generalization</p>\n\n<p>selecting best 2 models going to be very difficult for us </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 929680,
          "author_name": "Eek The Cat",
          "author_url": "",
          "post_date": "2020-07-14T20:59:08.030000",
          "content": "<p>I'm getting something like this \n<code>\nep:0 loss:0.6784 kappa_score:0.5503\nep:1 loss:0.5104 kappa_score:0.7474\nep:2 loss:0.4946 kappa_score:0.7153\nep:3 loss:0.4903 kappa_score:0.7274\nep:4 loss:0.4357 kappa_score:0.7593\n</code>\nI'm using PyTorch. This model achieves 0.9/0.89. Posted link in external data thread. \nP.S. if you're using some BN/GN in classification head, then just don't do it:) And as I said, tiling approach gives much more consistent correlation and lower LB gap</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 930247,
          "author_name": "Mobassir",
          "author_url": "",
          "post_date": "2020-07-15T09:51:26.230000",
          "content": "<p><a href=\"/cateek\">@cateek</a>  thank you!!\nimpressive train log\nmy groupnorm never gets above 0.5 qwk score</p>\n\n<p>i was simply replacing all the batchnorm layers to groupnorm layers from torchvision models\nfor example took resnet50 from torchvision and simply replaced all batchnorm layers to groupnorm layers and i got terrible  performance!</p>\n\n<p>i am curious- how you use groupnorm in pretrained model that is not breaking your network for this competition!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 930269,
          "author_name": "Mobassir",
          "author_url": "",
          "post_date": "2020-07-15T10:09:58.437000",
          "content": "<p><a href=\"/cateek\">@cateek</a>  for example this is my model where if i replace all batchnorm to groupnorm then i get terrible model : </p>\n\n<p><code>enetv2(\n  (enet): ResNet(\n    (conv1): myConv2d(3, 64, kernel_size=(7, 7), stride=(2, 2), padding=(3, 3), bias=False)\n    (bn1): BatchNorm2d(64, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n    (act1): ReLU(inplace=True)\n    (maxpool): MaxPool2d(kernel_size=3, stride=2, padding=1, dilation=1, ceil_mode=False)\n    (layer1): Sequential(\n      (0): Bottleneck(\n        (conv1): myConv2d(64, 128, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn1): BatchNorm2d(128, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act1): ReLU(inplace=True)\n        (conv2): myConv2d(128, 128, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), groups=32, bias=False)\n        (bn2): BatchNorm2d(128, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act2): ReLU(inplace=True)\n        (conv3): myConv2d(128, 256, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn3): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act3): ReLU(inplace=True)\n        (downsample): Sequential(\n          (0): myConv2d(64, 256, kernel_size=(1, 1), stride=(1, 1), bias=False)\n          (1): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        )\n      )\n      (1): Bottleneck(\n        (conv1): myConv2d(256, 128, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn1): BatchNorm2d(128, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act1): ReLU(inplace=True)\n        (conv2): myConv2d(128, 128, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), groups=32, bias=False)\n        (bn2): BatchNorm2d(128, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act2): ReLU(inplace=True)\n        (conv3): myConv2d(128, 256, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn3): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act3): ReLU(inplace=True)\n      )\n      (2): Bottleneck(\n        (conv1): myConv2d(256, 128, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn1): BatchNorm2d(128, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act1): ReLU(inplace=True)\n        (conv2): myConv2d(128, 128, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), groups=32, bias=False)\n        (bn2): BatchNorm2d(128, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act2): ReLU(inplace=True)\n        (conv3): myConv2d(128, 256, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn3): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act3): ReLU(inplace=True)\n      )\n    )\n    (layer2): Sequential(\n      (0): Bottleneck(\n        (conv1): myConv2d(256, 256, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn1): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act1): ReLU(inplace=True)\n        (conv2): myConv2d(256, 256, kernel_size=(3, 3), stride=(2, 2), padding=(1, 1), groups=32, bias=False)\n        (bn2): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act2): ReLU(inplace=True)\n        (conv3): myConv2d(256, 512, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn3): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act3): ReLU(inplace=True)\n        (downsample): Sequential(\n          (0): myConv2d(256, 512, kernel_size=(1, 1), stride=(2, 2), bias=False)\n          (1): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        )\n      )\n      (1): Bottleneck(\n        (conv1): myConv2d(512, 256, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn1): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act1): ReLU(inplace=True)\n        (conv2): myConv2d(256, 256, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), groups=32, bias=False)\n        (bn2): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act2): ReLU(inplace=True)\n        (conv3): myConv2d(256, 512, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn3): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act3): ReLU(inplace=True)\n      )\n      (2): Bottleneck(\n        (conv1): myConv2d(512, 256, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn1): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act1): ReLU(inplace=True)\n        (conv2): myConv2d(256, 256, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), groups=32, bias=False)\n        (bn2): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act2): ReLU(inplace=True)\n        (conv3): myConv2d(256, 512, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn3): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act3): ReLU(inplace=True)\n      )\n      (3): Bottleneck(\n        (conv1): myConv2d(512, 256, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn1): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act1): ReLU(inplace=True)\n        (conv2): myConv2d(256, 256, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), groups=32, bias=False)\n        (bn2): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act2): ReLU(inplace=True)\n        (conv3): myConv2d(256, 512, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn3): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act3): ReLU(inplace=True)\n      )\n    )\n    (layer3): Sequential(\n      (0): Bottleneck(\n        (conv1): myConv2d(512, 512, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn1): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act1): ReLU(inplace=True)\n        (conv2): myConv2d(512, 512, kernel_size=(3, 3), stride=(2, 2), padding=(1, 1), groups=32, bias=False)\n        (bn2): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act2): ReLU(inplace=True)\n        (conv3): myConv2d(512, 1024, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn3): BatchNorm2d(1024, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act3): ReLU(inplace=True)\n        (downsample): Sequential(\n          (0): myConv2d(512, 1024, kernel_size=(1, 1), stride=(2, 2), bias=False)\n          (1): BatchNorm2d(1024, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        )\n      )\n      (1): Bottleneck(\n        (conv1): myConv2d(1024, 512, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn1): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act1): ReLU(inplace=True)\n        (conv2): myConv2d(512, 512, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), groups=32, bias=False)\n        (bn2): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act2): ReLU(inplace=True)\n        (conv3): myConv2d(512, 1024, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn3): BatchNorm2d(1024, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act3): ReLU(inplace=True)\n      )\n      (2): Bottleneck(\n        (conv1): myConv2d(1024, 512, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn1): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act1): ReLU(inplace=True)\n        (conv2): myConv2d(512, 512, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), groups=32, bias=False)\n        (bn2): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act2): ReLU(inplace=True)\n        (conv3): myConv2d(512, 1024, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn3): BatchNorm2d(1024, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act3): ReLU(inplace=True)\n      )\n      (3): Bottleneck(\n        (conv1): myConv2d(1024, 512, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn1): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act1): ReLU(inplace=True)\n        (conv2): myConv2d(512, 512, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), groups=32, bias=False)\n        (bn2): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act2): ReLU(inplace=True)\n        (conv3): myConv2d(512, 1024, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn3): BatchNorm2d(1024, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act3): ReLU(inplace=True)\n      )\n      (4): Bottleneck(\n        (conv1): myConv2d(1024, 512, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn1): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act1): ReLU(inplace=True)\n        (conv2): myConv2d(512, 512, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), groups=32, bias=False)\n        (bn2): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act2): ReLU(inplace=True)\n        (conv3): myConv2d(512, 1024, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn3): BatchNorm2d(1024, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act3): ReLU(inplace=True)\n      )\n      (5): Bottleneck(\n        (conv1): myConv2d(1024, 512, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn1): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act1): ReLU(inplace=True)\n        (conv2): myConv2d(512, 512, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), groups=32, bias=False)\n        (bn2): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act2): ReLU(inplace=True)\n        (conv3): myConv2d(512, 1024, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn3): BatchNorm2d(1024, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act3): ReLU(inplace=True)\n      )\n    )\n    (layer4): Sequential(\n      (0): Bottleneck(\n        (conv1): myConv2d(1024, 1024, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn1): BatchNorm2d(1024, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act1): ReLU(inplace=True)\n        (conv2): myConv2d(1024, 1024, kernel_size=(3, 3), stride=(2, 2), padding=(1, 1), groups=32, bias=False)\n        (bn2): BatchNorm2d(1024, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act2): ReLU(inplace=True)\n        (conv3): myConv2d(1024, 2048, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn3): BatchNorm2d(2048, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act3): ReLU(inplace=True)\n        (downsample): Sequential(\n          (0): myConv2d(1024, 2048, kernel_size=(1, 1), stride=(2, 2), bias=False)\n          (1): BatchNorm2d(2048, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        )\n      )\n      (1): Bottleneck(\n        (conv1): myConv2d(2048, 1024, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn1): BatchNorm2d(1024, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act1): ReLU(inplace=True)\n        (conv2): myConv2d(1024, 1024, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), groups=32, bias=False)\n        (bn2): BatchNorm2d(1024, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act2): ReLU(inplace=True)\n        (conv3): myConv2d(1024, 2048, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn3): BatchNorm2d(2048, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act3): ReLU(inplace=True)\n      )\n      (2): Bottleneck(\n        (conv1): myConv2d(2048, 1024, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn1): BatchNorm2d(1024, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act1): ReLU(inplace=True)\n        (conv2): myConv2d(1024, 1024, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), groups=32, bias=False)\n        (bn2): BatchNorm2d(1024, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act2): ReLU(inplace=True)\n        (conv3): myConv2d(1024, 2048, kernel_size=(1, 1), stride=(1, 1), bias=False)\n        (bn3): BatchNorm2d(2048, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        (act3): ReLU(inplace=True)\n      )\n    )\n    (global_pool): SelectAdaptivePool2d (output_size=1, pool_type=avg)\n    (fc): Identity()\n  )\n  (myfc): Linear(in_features=2048, out_features=5, bias=True)\n)</code></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 930329,
          "author_name": "Eek The Cat",
          "author_url": "",
          "post_date": "2020-07-15T11:13:08.503000",
          "content": "<p><a href=\"/mobassir\">@mobassir</a> I posted github link in external data thread. It has a couple of models with imagenet weights. If you just replace all BN layers with GN, you will definitely break pretrained weights. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 930893,
          "author_name": "Mobassir",
          "author_url": "",
          "post_date": "2020-07-15T19:43:43.390000",
          "content": "<p><a href=\"/cateek\">@cateek</a> </p>\n\n<p>i tried resnet50 from here : <a href=\"https://github.com/joe-siyuan-qiao/pytorch-classification/tree/e6355f829e85ac05a71b8889f4fff77b9ab95d0b\">https://github.com/joe-siyuan-qiao/pytorch-classification/tree/e6355f829e85ac05a71b8889f4fff77b9ab95d0b</a> as you posted in external data thread,,in debug mode i can see qwk 0.00 again as like before!! just simply  changed quishen ha's efficientenet b0 with resnet50 from  that link : <a href=\"https://github.com/joe-siyuan-qiao/pytorch-classification/tree/e6355f829e85ac05a71b8889f4fff77b9ab95d0b\">https://github.com/joe-siyuan-qiao/pytorch-classification/tree/e6355f829e85ac05a71b8889f4fff77b9ab95d0b</a>\nin debug mode : \nEpoch 1, lr: 0.0000300, train loss: 0.61534, val loss: 0.60353, acc: 10.00000, qwk: 0.00000\nEpoch 2, lr: 0.0003000, train loss: 0.63444, val loss: 0.61907, acc: 10.00000, qwk: 0.00000</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 930899,
          "author_name": "Mobassir",
          "author_url": "",
          "post_date": "2020-07-15T19:52:45.767000",
          "content": "<p><a href=\"/cateek\">@cateek</a>  also not  able to download weights\nthis link : <a href=\"http://cs.jhu.edu/~syqiao/WeightStandardization/R-50-GN-WS.pth.tar\">http://cs.jhu.edu/~syqiao/WeightStandardization/R-50-GN-WS.pth.tar</a> doesn't work \njust shows \"untitled\" page and never starts download\nfrom where you downloaded weights?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 930931,
          "author_name": "Eek The Cat",
          "author_url": "",
          "post_date": "2020-07-15T20:21:54.500000",
          "content": "<p><a href=\"/mobassir\">@mobassir</a> Weights link is correct, although I used ResNeXt-50. No idea why you're getting zeros in qwk, there must be some bug</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 931504,
          "author_name": "Mobassir",
          "author_url": "",
          "post_date": "2020-07-16T08:35:33.763000",
          "content": "<p><a href=\"/cateek\">@cateek</a>  using resnet50 from that repo when i tried loading weight file he provided i get this error : <a href=\"https://discuss.pytorch.org/t/how-to-load-weights-of-weight-standardization-model/89418\">https://discuss.pytorch.org/t/how-to-load-weights-of-weight-standardization-model/89418</a></p>\n\n<p>i downloaded resnet50 weight  by following link he provided</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 931526,
          "author_name": "Eek The Cat",
          "author_url": "",
          "post_date": "2020-07-16T08:48:28.613000",
          "content": "<p>Well, haven't encountered anything like that. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 931527,
          "author_name": "Mobassir",
          "author_url": "",
          "post_date": "2020-07-16T08:50:12.703000",
          "content": "<p>sorry <a href=\"/cateek\">@cateek</a>  i disturbed you a lot in this discussion post and  thank you for answering everything,i will try resnext50,thanks</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 922933,
      "author_name": "Kurian Benoy",
      "author_url": "",
      "post_date": "2020-07-10T12:30:57.860000",
      "content": "<p>Thanks for sharing your result.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 922940,
          "author_name": "Vlad Vaduva",
          "author_url": "",
          "post_date": "2020-07-10T12:36:50.177000",
          "content": "<p>No problem. This is how we all learn things, by sharing and discussing about them !</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 927378,
      "author_name": "Anmol Jaiswal",
      "author_url": "",
      "post_date": "2020-07-13T10:52:03.173000",
      "content": "<p>well done</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 925280,
      "author_name": "Shiyuan Zeng",
      "author_url": "",
      "post_date": "2020-07-12T01:14:27.290000",
      "content": "<p>Good to know! Totally agree with 4 :D\nMay I ask you how many epochs do you train for? For me, the valid-loss fluctuate in 20~30 epochs while train-loss keep decreasing. So I end training in 30 epoch with valid-loss almost equal to 6~7 x train-loss(local score ~0.88 and lb also 0.88, single fold). The loss I used is BCEWithLogitsLoss like Qishen Ha did.(but not the same on changing pred to label)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 928335,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-13T22:21:32.190000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "922929": "Among other things that worked that I will share after the competition ends, here are 7 things that for me did not work, they either give the same cv/public leaderboard result as without them or lower.\n\n1. Images with an amount of pixels larger than 2.500.000 (sum of all patches pixels).\n2. RangerLars optimizer (Really wanted to use it in this competition but it could not give me a boost. Probably I did not use it with compatible scheduler).\n3. Very big arhitectures -&gt; overfit.\n4. Batchsizes smaller than 3 did not work for me (the updates are too noisy)\n5. Tiles selected from the biggest resolution of the .tiff files\n5. Did not manage to proper use the masks provided (this is where I thing is a large space for improvement)\n6. Any augmentation except affine transforms\n7. Training individually per data provider",
    "926411": "One more: class-balanced training didn't add anything for me. I undersampled (but resampled each epoch) and it neither hurt nor helped CV or LB scores.\n\nI think there is quite a lot to be gained from different approaches to the inference part of the problem. The models I have aren't amazing, they only score around 0.87 but useful approaches to inference take the score to 0.89. I haven't looked at post processing yet - I'm just straight averaging - but that's next on the to-do list.",
    "924393": "very useful,thanks for sharing",
    "922935": "which scheduler worked best for you?did you try custom scheduler?",
    "924070": "did you try TTA and Post Processing? or your best submission is a single model or ensemble without TTA and Post Processing?",
    "923289": "RangerLars and other RAdam-related optimizers typically work well with a flat+cosine decay (with cosine decasy starting about 70% in)",
    "923116": "Thank you very much for sharing again! May I ask a few questions:\n1. Have you used batch accumulation with much success?\n2. Have you gotten any boost from ensembling models?\n3. Have you attempted to train with duplicates removed?",
    "922987": "I found the same thing with points 1 and 3, they don't give that much of a performance increase, and you'll have to effectively lower the batch size.\n\nAnother question: Did you also exclude samples, such as samples with pen marks?",
    "922952": "@vladvdv Thanks for your wonderful results (and congrats to your position)! I am currently struggling to get a stable CV-LB correlation. With stratified KF on isup I can get LB ranging from .84 to .87 on the same CV .91. Would you care to kindly disclose how CV-LB works out for you?",
    "922933": "Thanks for sharing your result.",
    "927378": "well done",
    "925280": "Good to know! Totally agree with 4 :D\nMay I ask you how many epochs do you train for? For me, the valid-loss fluctuate in 20~30 epochs while train-loss keep decreasing. So I end training in 30 epoch with valid-loss almost equal to 6~7 x train-loss(local score ~0.88 and lb also 0.88, single fold). The loss I used is BCEWithLogitsLoss like Qishen Ha did.(but not the same on changing pred to label)",
    "928335": ""
  }
}