{
  "id": 161909,
  "title": "Reproducibility issue lb ranging from 0.89-0.91 on the same setup",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/161909",
  "author_name": "Shujun",
  "post_date": "2020-06-26T16:03:33.870000",
  "votes": 22,
  "comment_count": 19,
  "views": 0,
  "content": "<p>Has anyone else run into reproducibility issues? Maybe I should have set the seeds but I somehow can't reproduce my 0.91 run. Lb range from 0.89 to 0.91 on the same setup</p>",
  "messages": [
    {
      "id": 903165,
      "postDate": "2020-06-26T16:03:33.870Z",
      "content": "<p>Has anyone else run into reproducibility issues? Maybe I should have set the seeds but I somehow can't reproduce my 0.91 run. Lb range from 0.89 to 0.91 on the same setup</p>",
      "rawMarkdown": "Has anyone else run into reproducibility issues? Maybe I should have set the seeds but I somehow can't reproduce my 0.91 run. Lb range from 0.89 to 0.91 on the same setup",
      "votes": 22
    },
    {
      "id": 903585,
      "postDate": "2020-06-27T01:00:43.603Z",
      "content": "<p>Well Some time ago I did this experiments... Still looking for an explanation.... </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2Fca733f61dc351e048d5e405be99c5353%2FScreen%20Shot%202020-06-26%20at%208.55.36%20PM.png?generation=1593219634634111&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Well Some time ago I did this experiments... Still looking for an explanation.... \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2Fca733f61dc351e048d5e405be99c5353%2FScreen%20Shot%202020-06-26%20at%208.55.36%20PM.png?generation=1593219634634111&amp;alt=media)\n",
      "votes": 11,
      "replies": [
        {
          "id": 903611,
          "postDate": "2020-06-27T01:46:57.073Z",
          "content": "<p>this is similar to my experience</p>",
          "rawMarkdown": "this is similar to my experience\n",
          "votes": 1
        },
        {
          "id": 903617,
          "postDate": "2020-06-27T01:54:03.210Z",
          "content": "<p>btw I forgot to mention my <code>random seed</code> is <code>69420</code>... Just incase if someone wants to reproduce this results =)</p>",
          "rawMarkdown": "btw I forgot to mention my `random seed` is `69420`... Just incase if someone wants to reproduce this results =)",
          "votes": 2
        },
        {
          "id": 904144,
          "postDate": "2020-06-27T11:53:57.563Z",
          "content": "<p>I don't remember exactly who talked about fine-tuning seed two months ago but it seems to make real sense 😁 </p>",
          "rawMarkdown": "I don't remember exactly who talked about fine-tuning seed two months ago but it seems to make real sense 😁 "
        },
        {
          "id": 904999,
          "postDate": "2020-06-28T07:14:28.883Z",
          "content": "<p>Looks like one cycle policy :)</p>",
          "rawMarkdown": "Looks like one cycle policy :)"
        },
        {
          "id": 905737,
          "postDate": "2020-06-28T18:42:39.853Z",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> \ni get very analogus  CM as yours and CV as shown below still clueless why is not giving same at LB but CV less than this gets a better score . Any pointers that would help further?</p>\n\n<p><code>\n**(array([[152.,   6.,   5.,   1.,   5.,   0.],\n        [ 15.,  81.,  21.,   7.,   2.,   0.],\n        [  3.,  14.,  78.,  30.,   2.,   1.],\n        [ 11.,   4.,  21.,  89.,  39.,  13.],\n        [  2.,   1.,   7.,  23.,  87.,  30.],\n        [  2.,   0.,   2.,  12.,  28., 130.]]),\n array([[357.,  19.,   4.,   0.,   0.,   0.],\n        [ 28., 305.,  23.,   3.,   1.,   0.],\n        [  3.,  53.,  61.,  12.,   6.,   1.],\n        [  1.,   2.,   8.,  30.,  19.,   2.],\n        [  2.,   3.,   5.,  11.,  52.,   9.],\n        [  0.,   0.,   2.,   3.,  20.,  34.]]),\n array([[509.,  25.,   9.,   1.,   5.,   0.],\n        [ 43., 386.,  44.,  10.,   3.,   0.],\n        [  6.,  67., 139.,  42.,   8.,   2.],\n        [ 12.,   6.,  29., 119.,  58.,  15.],\n        [  4.,   4.,  12.,  34., 139.,  39.],\n        [  2.,   0.,   4.,  15.,  48., 164.]])**,\n 0.8673648612039228, RA\n 0.9088795478706533) KA\n</code></p>",
          "rawMarkdown": "@drhabib \ni get very analogus  CM as yours and CV as shown below still clueless why is not giving same at LB but CV less than this gets a better score . Any pointers that would help further?\n\n```\n**(array([[152.,   6.,   5.,   1.,   5.,   0.],\n        [ 15.,  81.,  21.,   7.,   2.,   0.],\n        [  3.,  14.,  78.,  30.,   2.,   1.],\n        [ 11.,   4.,  21.,  89.,  39.,  13.],\n        [  2.,   1.,   7.,  23.,  87.,  30.],\n        [  2.,   0.,   2.,  12.,  28., 130.]]),\n array([[357.,  19.,   4.,   0.,   0.,   0.],\n        [ 28., 305.,  23.,   3.,   1.,   0.],\n        [  3.,  53.,  61.,  12.,   6.,   1.],\n        [  1.,   2.,   8.,  30.,  19.,   2.],\n        [  2.,   3.,   5.,  11.,  52.,   9.],\n        [  0.,   0.,   2.,   3.,  20.,  34.]]),\n array([[509.,  25.,   9.,   1.,   5.,   0.],\n        [ 43., 386.,  44.,  10.,   3.,   0.],\n        [  6.,  67., 139.,  42.,   8.,   2.],\n        [ 12.,   6.,  29., 119.,  58.,  15.],\n        [  4.,   4.,  12.,  34., 139.,  39.],\n        [  2.,   0.,   4.,  15.,  48., 164.]])**,\n 0.8673648612039228, RA\n 0.9088795478706533) KA\n```"
        }
      ]
    },
    {
      "id": 907490,
      "postDate": "2020-06-30T02:30:03.113Z",
      "content": "<p>Nah I have the same problem in the 0.87-0.89 range where the exact same strategy gives me widely different results. While I can reproduce my pytorch results if absolutely everything is the same, there's still something wrong going on. I wonder how bad the shake up is going to be...</p>\n\n<p>Like most in this thread higher than me I also get good 0.9050 ish CV but haven't yet had LB above 0.89. Maybe I should just change the seed and keep submitting until something sticks lol.</p>\n\n<p>Time to also save top_3 checkpoints I guess.</p>",
      "rawMarkdown": "Nah I have the same problem in the 0.87-0.89 range where the exact same strategy gives me widely different results. While I can reproduce my pytorch results if absolutely everything is the same, there's still something wrong going on. I wonder how bad the shake up is going to be...\n\nLike most in this thread higher than me I also get good 0.9050 ish CV but haven't yet had LB above 0.89. Maybe I should just change the seed and keep submitting until something sticks lol.\n\nTime to also save top_3 checkpoints I guess.",
      "votes": 1
    },
    {
      "id": 904927,
      "postDate": "2020-06-28T05:09:40.913Z",
      "content": "<p>Actually it seems like at least in pytorch, it is not possible to have exactly reproducible results. </p>\n\n<p><a href=\"https://pytorch.org/docs/stable/notes/randomness.html\">https://pytorch.org/docs/stable/notes/randomness.html</a></p>",
      "rawMarkdown": "Actually it seems like at least in pytorch, it is not possible to have exactly reproducible results. \n\nhttps://pytorch.org/docs/stable/notes/randomness.html",
      "votes": 1,
      "replies": [
        {
          "id": 904950,
          "postDate": "2020-06-28T05:32:38.003Z",
          "content": "<p>My understanding from reading that page is that with the correctly set seeds, it <em>is</em> deterministic and reproducible. Is that incorrect?</p>",
          "rawMarkdown": "My understanding from reading that page is that with the correctly set seeds, it _is_ deterministic and reproducible. Is that incorrect?"
        },
        {
          "id": 904971,
          "postDate": "2020-06-28T06:11:37.967Z",
          "content": "<p>\"Completely reproducible results are not guaranteed across PyTorch releases, individual commits or different platforms. Furthermore, results need not be reproducible between CPU and GPU executions, even when using identical seeds.\"</p>\n\n<p>Also from my local experiments, this seems to be the case even when I set all the seeds across pytorch and numpy</p>",
          "rawMarkdown": "\"Completely reproducible results are not guaranteed across PyTorch releases, individual commits or different platforms. Furthermore, results need not be reproducible between CPU and GPU executions, even when using identical seeds.\"\n\nAlso from my local experiments, this seems to be the case even when I set all the seeds across pytorch and numpy",
          "votes": 1
        },
        {
          "id": 904972,
          "postDate": "2020-06-28T06:17:35.983Z",
          "content": "<p>\"However, in order to <strong>make computations deterministic on your specific problem</strong> on one specific platform and PyTorch release, there are a couple of steps to take.</p>\n\n<p>There are two pseudorandom number generators involved in PyTorch, which you will need to seed manually to make runs reproducible. Furthermore, you should ensure that all other libraries your code relies on and which use random numbers also use a fixed seed.\"</p>\n\n<p>Have you also set the cuDNN seeds?\n <code>\ntorch.backends.cudnn.deterministic = True\ntorch.backends.cudnn.benchmark = False\n</code></p>",
          "rawMarkdown": "\"However, in order to **make computations deterministic on your specific problem** on one specific platform and PyTorch release, there are a couple of steps to take.\n\nThere are two pseudorandom number generators involved in PyTorch, which you will need to seed manually to make runs reproducible. Furthermore, you should ensure that all other libraries your code relies on and which use random numbers also use a fixed seed.\"\n\nHave you also set the cuDNN seeds?\n ```\ntorch.backends.cudnn.deterministic = True\ntorch.backends.cudnn.benchmark = False\n```",
          "votes": 1
        }
      ]
    },
    {
      "id": 903623,
      "postDate": "2020-06-27T02:04:57.467Z",
      "content": "<p>Interesting insights. My current model, with all random seeds fixed(except for <code>torch.backends.cudnn.deterministic = True</code> because it slows down training), I can constantly score upper 0.90(judging from LB). </p>",
      "rawMarkdown": "Interesting insights. My current model, with all random seeds fixed(except for `torch.backends.cudnn.deterministic = True` because it slows down training), I can constantly score upper 0.90(judging from LB). ",
      "votes": 1
    },
    {
      "id": 903198,
      "postDate": "2020-06-26T16:26:14.720Z",
      "content": "<p>had the same issue - scored 0.9, then changed here and there and couldn't reproduce it, although I had all the seeds fixed. I noticed that different checkpoints (same model, same seeds) with almost the same kappa and loss lead to 0.85-0.9 range. I started to work on stability, but not that successful so far.</p>",
      "rawMarkdown": "had the same issue - scored 0.9, then changed here and there and couldn't reproduce it, although I had all the seeds fixed. I noticed that different checkpoints (same model, same seeds) with almost the same kappa and loss lead to 0.85-0.9 range. I started to work on stability, but not that successful so far.",
      "votes": 1,
      "replies": [
        {
          "id": 903373,
          "postDate": "2020-06-26T19:12:46.987Z",
          "content": "<p>Yeah, so for the same run if I select by different criteria, the lb score can be 0.89, 0.90, and 0.91. It seems theres a lot of lb noise </p>",
          "rawMarkdown": "Yeah, so for the same run if I select by different criteria, the lb score can be 0.89, 0.90, and 0.91. It seems theres a lot of lb noise ",
          "votes": 4
        }
      ]
    },
    {
      "id": 913860,
      "postDate": "2020-07-03T13:35:19.397Z",
      "content": "<p><a href=\"/shujun717\">@shujun717</a> \nhow u manage to make 0.91. I mean any thing different in terms of Tile extraction methodology ,loss used . \nI think all who are stuck at 0.90 may be pretty much using more or less same tile extraction method.  </p>",
      "rawMarkdown": "@shujun717 \nhow u manage to make 0.91. I mean any thing different in terms of Tile extraction methodology ,loss used . \nI think all who are stuck at 0.90 may be pretty much using more or less same tile extraction method.  ",
      "replies": [
        {
          "id": 916750,
          "postDate": "2020-07-06T02:05:19.633Z",
          "content": "<p>My 0.91 sub used standard tiles from Iafoss' kernel, but there is a lot of lb noise, the same run has checkpoints ranging from lb 0.89 to 0.91...</p>",
          "rawMarkdown": "My 0.91 sub used standard tiles from Iafoss' kernel, but there is a lot of lb noise, the same run has checkpoints ranging from lb 0.89 to 0.91..."
        }
      ]
    },
    {
      "id": 907969,
      "postDate": "2020-06-30T09:40:19.230Z",
      "content": "<p>I actually have the opposite - my local CV is around 0.88, while the  public was 0.9.  But I use the quite restrictive validation set without any duplicates, etc, so it's kinda by design</p>",
      "rawMarkdown": "I actually have the opposite - my local CV is around 0.88, while the  public was 0.9.  But I use the quite restrictive validation set without any duplicates, etc, so it's kinda by design"
    },
    {
      "id": 904893,
      "postDate": "2020-06-28T04:35:50.190Z",
      "content": "<p>I've also noticed this exactly reproducing the notebook by <a href=\"/haqishen\">@haqishen</a></p>\n\n<p>While frustrating, I'm glad to see I'm not going insane since others are experiencing it too.</p>",
      "rawMarkdown": "I've also noticed this exactly reproducing the notebook by @haqishen\n\nWhile frustrating, I'm glad to see I'm not going insane since others are experiencing it too.",
      "replies": [
        {
          "id": 917426,
          "postDate": "2020-07-06T13:51:07.483Z",
          "content": "<p><a href=\"/tylerashworth\">@tylerashworth</a> would you like to share your results in reproducing Ha's notebook? (Credits to Ha, it's an extremely helpful notebook). I have results ranging from .84 to .86, no .87 so far.</p>",
          "rawMarkdown": "@tylerashworth would you like to share your results in reproducing Ha's notebook? (Credits to Ha, it's an extremely helpful notebook). I have results ranging from .84 to .86, no .87 so far."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 903585,
      "author_name": "DrHB",
      "author_url": "",
      "post_date": "2020-06-27T01:00:43.603000",
      "content": "<p>Well Some time ago I did this experiments... Still looking for an explanation.... </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2Fca733f61dc351e048d5e405be99c5353%2FScreen%20Shot%202020-06-26%20at%208.55.36%20PM.png?generation=1593219634634111&amp;alt=media\" alt=\"\"></p>",
      "votes": 11,
      "replies": [
        {
          "id": 903611,
          "author_name": "Shujun",
          "author_url": "",
          "post_date": "2020-06-27T01:46:57.073000",
          "content": "<p>this is similar to my experience</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 903617,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-06-27T01:54:03.210000",
          "content": "<p>btw I forgot to mention my <code>random seed</code> is <code>69420</code>... Just incase if someone wants to reproduce this results =)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 904144,
          "author_name": "Alex",
          "author_url": "",
          "post_date": "2020-06-27T11:53:57.563000",
          "content": "<p>I don't remember exactly who talked about fine-tuning seed two months ago but it seems to make real sense 😁 </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 904999,
          "author_name": "Miroslav Valan",
          "author_url": "",
          "post_date": "2020-06-28T07:14:28.883000",
          "content": "<p>Looks like one cycle policy :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 905737,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-06-28T18:42:39.853000",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> \ni get very analogus  CM as yours and CV as shown below still clueless why is not giving same at LB but CV less than this gets a better score . Any pointers that would help further?</p>\n\n<p><code>\n**(array([[152.,   6.,   5.,   1.,   5.,   0.],\n        [ 15.,  81.,  21.,   7.,   2.,   0.],\n        [  3.,  14.,  78.,  30.,   2.,   1.],\n        [ 11.,   4.,  21.,  89.,  39.,  13.],\n        [  2.,   1.,   7.,  23.,  87.,  30.],\n        [  2.,   0.,   2.,  12.,  28., 130.]]),\n array([[357.,  19.,   4.,   0.,   0.,   0.],\n        [ 28., 305.,  23.,   3.,   1.,   0.],\n        [  3.,  53.,  61.,  12.,   6.,   1.],\n        [  1.,   2.,   8.,  30.,  19.,   2.],\n        [  2.,   3.,   5.,  11.,  52.,   9.],\n        [  0.,   0.,   2.,   3.,  20.,  34.]]),\n array([[509.,  25.,   9.,   1.,   5.,   0.],\n        [ 43., 386.,  44.,  10.,   3.,   0.],\n        [  6.,  67., 139.,  42.,   8.,   2.],\n        [ 12.,   6.,  29., 119.,  58.,  15.],\n        [  4.,   4.,  12.,  34., 139.,  39.],\n        [  2.,   0.,   4.,  15.,  48., 164.]])**,\n 0.8673648612039228, RA\n 0.9088795478706533) KA\n</code></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 907490,
      "author_name": "Arnaud Roussel",
      "author_url": "",
      "post_date": "2020-06-30T02:30:03.113000",
      "content": "<p>Nah I have the same problem in the 0.87-0.89 range where the exact same strategy gives me widely different results. While I can reproduce my pytorch results if absolutely everything is the same, there's still something wrong going on. I wonder how bad the shake up is going to be...</p>\n\n<p>Like most in this thread higher than me I also get good 0.9050 ish CV but haven't yet had LB above 0.89. Maybe I should just change the seed and keep submitting until something sticks lol.</p>\n\n<p>Time to also save top_3 checkpoints I guess.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 904927,
      "author_name": "Shujun",
      "author_url": "",
      "post_date": "2020-06-28T05:09:40.913000",
      "content": "<p>Actually it seems like at least in pytorch, it is not possible to have exactly reproducible results. </p>\n\n<p><a href=\"https://pytorch.org/docs/stable/notes/randomness.html\">https://pytorch.org/docs/stable/notes/randomness.html</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 904950,
          "author_name": "ilovescience",
          "author_url": "",
          "post_date": "2020-06-28T05:32:38.003000",
          "content": "<p>My understanding from reading that page is that with the correctly set seeds, it <em>is</em> deterministic and reproducible. Is that incorrect?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 904971,
          "author_name": "Shujun",
          "author_url": "",
          "post_date": "2020-06-28T06:11:37.967000",
          "content": "<p>\"Completely reproducible results are not guaranteed across PyTorch releases, individual commits or different platforms. Furthermore, results need not be reproducible between CPU and GPU executions, even when using identical seeds.\"</p>\n\n<p>Also from my local experiments, this seems to be the case even when I set all the seeds across pytorch and numpy</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 904972,
          "author_name": "ilovescience",
          "author_url": "",
          "post_date": "2020-06-28T06:17:35.983000",
          "content": "<p>\"However, in order to <strong>make computations deterministic on your specific problem</strong> on one specific platform and PyTorch release, there are a couple of steps to take.</p>\n\n<p>There are two pseudorandom number generators involved in PyTorch, which you will need to seed manually to make runs reproducible. Furthermore, you should ensure that all other libraries your code relies on and which use random numbers also use a fixed seed.\"</p>\n\n<p>Have you also set the cuDNN seeds?\n <code>\ntorch.backends.cudnn.deterministic = True\ntorch.backends.cudnn.benchmark = False\n</code></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 903623,
      "author_name": "RabotniKuma",
      "author_url": "",
      "post_date": "2020-06-27T02:04:57.467000",
      "content": "<p>Interesting insights. My current model, with all random seeds fixed(except for <code>torch.backends.cudnn.deterministic = True</code> because it slows down training), I can constantly score upper 0.90(judging from LB). </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 903198,
      "author_name": "abzaliev",
      "author_url": "",
      "post_date": "2020-06-26T16:26:14.720000",
      "content": "<p>had the same issue - scored 0.9, then changed here and there and couldn't reproduce it, although I had all the seeds fixed. I noticed that different checkpoints (same model, same seeds) with almost the same kappa and loss lead to 0.85-0.9 range. I started to work on stability, but not that successful so far.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 903373,
          "author_name": "Shujun",
          "author_url": "",
          "post_date": "2020-06-26T19:12:46.987000",
          "content": "<p>Yeah, so for the same run if I select by different criteria, the lb score can be 0.89, 0.90, and 0.91. It seems theres a lot of lb noise </p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 913860,
      "author_name": "Jaideep",
      "author_url": "",
      "post_date": "2020-07-03T13:35:19.397000",
      "content": "<p><a href=\"/shujun717\">@shujun717</a> \nhow u manage to make 0.91. I mean any thing different in terms of Tile extraction methodology ,loss used . \nI think all who are stuck at 0.90 may be pretty much using more or less same tile extraction method.  </p>",
      "votes": 0,
      "replies": [
        {
          "id": 916750,
          "author_name": "Shujun",
          "author_url": "",
          "post_date": "2020-07-06T02:05:19.633000",
          "content": "<p>My 0.91 sub used standard tiles from Iafoss' kernel, but there is a lot of lb noise, the same run has checkpoints ranging from lb 0.89 to 0.91...</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 907969,
      "author_name": "abzaliev",
      "author_url": "",
      "post_date": "2020-06-30T09:40:19.230000",
      "content": "<p>I actually have the opposite - my local CV is around 0.88, while the  public was 0.9.  But I use the quite restrictive validation set without any duplicates, etc, so it's kinda by design</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 904893,
      "author_name": "Tyler Ashworth",
      "author_url": "",
      "post_date": "2020-06-28T04:35:50.190000",
      "content": "<p>I've also noticed this exactly reproducing the notebook by <a href=\"/haqishen\">@haqishen</a></p>\n\n<p>While frustrating, I'm glad to see I'm not going insane since others are experiencing it too.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 917426,
          "author_name": "Nicholas Lyu",
          "author_url": "",
          "post_date": "2020-07-06T13:51:07.483000",
          "content": "<p><a href=\"/tylerashworth\">@tylerashworth</a> would you like to share your results in reproducing Ha's notebook? (Credits to Ha, it's an extremely helpful notebook). I have results ranging from .84 to .86, no .87 so far.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "903165": "Has anyone else run into reproducibility issues? Maybe I should have set the seeds but I somehow can't reproduce my 0.91 run. Lb range from 0.89 to 0.91 on the same setup",
    "903585": "Well Some time ago I did this experiments... Still looking for an explanation.... \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2Fca733f61dc351e048d5e405be99c5353%2FScreen%20Shot%202020-06-26%20at%208.55.36%20PM.png?generation=1593219634634111&amp;alt=media)\n",
    "907490": "Nah I have the same problem in the 0.87-0.89 range where the exact same strategy gives me widely different results. While I can reproduce my pytorch results if absolutely everything is the same, there's still something wrong going on. I wonder how bad the shake up is going to be...\n\nLike most in this thread higher than me I also get good 0.9050 ish CV but haven't yet had LB above 0.89. Maybe I should just change the seed and keep submitting until something sticks lol.\n\nTime to also save top_3 checkpoints I guess.",
    "904927": "Actually it seems like at least in pytorch, it is not possible to have exactly reproducible results. \n\nhttps://pytorch.org/docs/stable/notes/randomness.html",
    "903623": "Interesting insights. My current model, with all random seeds fixed(except for `torch.backends.cudnn.deterministic = True` because it slows down training), I can constantly score upper 0.90(judging from LB). ",
    "903198": "had the same issue - scored 0.9, then changed here and there and couldn't reproduce it, although I had all the seeds fixed. I noticed that different checkpoints (same model, same seeds) with almost the same kappa and loss lead to 0.85-0.9 range. I started to work on stability, but not that successful so far.",
    "913860": "@shujun717 \nhow u manage to make 0.91. I mean any thing different in terms of Tile extraction methodology ,loss used . \nI think all who are stuck at 0.90 may be pretty much using more or less same tile extraction method.  ",
    "907969": "I actually have the opposite - my local CV is around 0.88, while the  public was 0.9.  But I use the quite restrictive validation set without any duplicates, etc, so it's kinda by design",
    "904893": "I've also noticed this exactly reproducing the notebook by @haqishen\n\nWhile frustrating, I'm glad to see I'm not going insane since others are experiencing it too."
  }
}