{
  "id": 427698,
  "title": "Difference in results with different machine's ",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/427698",
  "author_name": "Monkey D Donut",
  "post_date": "2023-07-29T05:57:07.122000",
  "votes": 3,
  "comment_count": 6,
  "views": 0,
  "content": "<p>hi :),<br>\n   do the effects of training on various machines differ.Will a pretrained model that is fine-tuned using  CPU, one GPU, or multiple GPUs for the same number of epochs produce different results? <br>\nnew to kaggle, sorry if its a stupid question :) <br>\nthank you.</p>",
  "messages": [
    {
      "id": 2366418,
      "postDate": "2023-07-31T01:40:32.463Z",
      "content": "<p>Hello, </p>\n<p>Yes, you'll get different results if you aren't using any seed each time you train, irrespective of the processing unit. </p>\n<p>If you want to reproduce your results I suggest you set a seed first and use the same seed to repeat the results next time you train. Also, I found out that it's kinda difficult to reproduce the exact results when using a GPU or a TPU due to the randomness they introduce. So, for the exact replication of the results, go with a CPU and set a seed.</p>",
      "rawMarkdown": "Hello, \n\nYes, you'll get different results if you aren't using any seed each time you train, irrespective of the processing unit. \n\nIf you want to reproduce your results I suggest you set a seed first and use the same seed to repeat the results next time you train. Also, I found out that it's kinda difficult to reproduce the exact results when using a GPU or a TPU due to the randomness they introduce. So, for the exact replication of the results, go with a CPU and set a seed.",
      "votes": 1,
      "replies": [
        {
          "id": 2366604,
          "postDate": "2023-07-31T05:35:53.717Z",
          "content": "<p>thank you    </p>",
          "rawMarkdown": "thank you    "
        }
      ]
    },
    {
      "id": 2364113,
      "postDate": "2023-07-29T06:19:50.510Z",
      "content": "<p>I guess that one of the main impact will come from the batch size you will be able to set with these different settings. And this will have a consequent impact on the results, mostly in case you deal with imbalanced problems.</p>",
      "rawMarkdown": "I guess that one of the main impact will come from the batch size you will be able to set with these different settings. And this will have a consequent impact on the results, mostly in case you deal with imbalanced problems.",
      "votes": 1,
      "replies": [
        {
          "id": 2364157,
          "postDate": "2023-07-29T07:08:11.333Z",
          "content": "<p>thank you,<br>\n         so if all train aruguments are same the results should be nearly similar . do typecasting has major influnce on results?</p>",
          "rawMarkdown": "thank you,\n         so if all train aruguments are same the results should be nearly similar . do typecasting has major influnce on results?"
        }
      ]
    },
    {
      "id": 2366685,
      "postDate": "2023-07-31T06:37:49.117Z",
      "content": "<p>Hi! I have not tried multiple GPU vs single GPU setups but I could add some details on reproducibility in general </p>\n<p>1) In addition to what was mentioned by other commenters if you are using pytorch with GPU you should also set <a href=\"https://pytorch.org/docs/stable/notes/randomness.html#avoiding-nondeterministic-algorithms\" target=\"_blank\">deterministic flag</a>. I have used the code below to set seeds for all the used libraries &amp; deterministic behavior for torch.</p>\n<pre><code>def seed_everything(: ):    \n    random.()\n    os.environ[] = str()\n    np.random.()\n    torch.manual_seed()\n    torch.cuda.manual_seed()\n    torch.cuda.manual_seed_all()\n    torch.backends.cudnn.deterministic = True\n    torch.backends.cudnn.benchmark = False\n</code></pre>\n<p>If you use pytorch lightning you could set it with <code>deterministic</code> argument of Trainer &amp; seed_everything function of PL.</p>\n<p><em>Note</em>: some algorithms in pytorch does not have a deterministic CUDA implementation for backward (e. g. bilinear interpolation), but in such case it will raise an error if you set the deterministic flag.</p>\n<p>2) Also for batch size: if you are using multiple machines with different VRAM size and training separately on them (e. g. different CV folds), you could use gradient accumulation to get effectively equal batch size. E. g. you have larger machine with 24Gb VRAM and smaller machine with 16Gb, BS 64 fits into larger machine but not into smaller. Then, using BS 32 + gradient accumulation of 2 on small machine is effectively equal to BS 64. <em>Note</em>: as far as I know accumulation does not take into account batch normalization, so if your model has BNs, results will be different and smaller BS should be used with care.</p>",
      "rawMarkdown": "Hi! I have not tried multiple GPU vs single GPU setups but I could add some details on reproducibility in general \n\n1) In addition to what was mentioned by other commenters if you are using pytorch with GPU you should also set [deterministic flag](https://pytorch.org/docs/stable/notes/randomness.html#avoiding-nondeterministic-algorithms). I have used the code below to set seeds for all the used libraries & deterministic behavior for torch.\n\n```\ndef seed_everything(seed: int):    \n    random.seed(seed)\n    os.environ['PYTHONHASHSEED'] = str(seed)\n    np.random.seed(seed)\n    torch.manual_seed(seed)\n    torch.cuda.manual_seed(seed)\n    torch.cuda.manual_seed_all(seed)\n    torch.backends.cudnn.deterministic = True\n    torch.backends.cudnn.benchmark = False\n```\n\nIf you use pytorch lightning you could set it with `deterministic` argument of Trainer & seed_everything function of PL.\n\n*Note*: some algorithms in pytorch does not have a deterministic CUDA implementation for backward (e. g. bilinear interpolation), but in such case it will raise an error if you set the deterministic flag.\n\n2) Also for batch size: if you are using multiple machines with different VRAM size and training separately on them (e. g. different CV folds), you could use gradient accumulation to get effectively equal batch size. E. g. you have larger machine with 24Gb VRAM and smaller machine with 16Gb, BS 64 fits into larger machine but not into smaller. Then, using BS 32 + gradient accumulation of 2 on small machine is effectively equal to BS 64. *Note*: as far as I know accumulation does not take into account batch normalization, so if your model has BNs, results will be different and smaller BS should be used with care.",
      "votes": 2,
      "replies": [
        {
          "id": 2366842,
          "postDate": "2023-07-31T07:55:15.430Z",
          "content": "<p>thank you for the detailed explanation :)</p>",
          "rawMarkdown": "thank you for the detailed explanation :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 2364092,
      "postDate": "2023-07-29T05:57:07.123Z",
      "content": "<p>hi :),<br>\n   do the effects of training on various machines differ.Will a pretrained model that is fine-tuned using  CPU, one GPU, or multiple GPUs for the same number of epochs produce different results? <br>\nnew to kaggle, sorry if its a stupid question :) <br>\nthank you.</p>",
      "rawMarkdown": "hi :),\n   do the effects of training on various machines differ.Will a pretrained model that is fine-tuned using  CPU, one GPU, or multiple GPUs for the same number of epochs produce different results? \nnew to kaggle, sorry if its a stupid question :) \nthank you.",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 2366418,
      "author_name": "Swaroop Meher",
      "author_url": "",
      "post_date": "2023-07-31T01:40:32.463000",
      "content": "<p>Hello, </p>\n<p>Yes, you'll get different results if you aren't using any seed each time you train, irrespective of the processing unit. </p>\n<p>If you want to reproduce your results I suggest you set a seed first and use the same seed to repeat the results next time you train. Also, I found out that it's kinda difficult to reproduce the exact results when using a GPU or a TPU due to the randomness they introduce. So, for the exact replication of the results, go with a CPU and set a seed.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2366604,
          "author_name": "Monkey D Donut",
          "author_url": "",
          "post_date": "2023-07-31T05:35:53.717000",
          "content": "<p>thank you    </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2364113,
      "author_name": "FabienDaniel",
      "author_url": "",
      "post_date": "2023-07-29T06:19:50.510000",
      "content": "<p>I guess that one of the main impact will come from the batch size you will be able to set with these different settings. And this will have a consequent impact on the results, mostly in case you deal with imbalanced problems.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2364157,
          "author_name": "Monkey D Donut",
          "author_url": "",
          "post_date": "2023-07-29T07:08:11.333000",
          "content": "<p>thank you,<br>\n         so if all train aruguments are same the results should be nearly similar . do typecasting has major influnce on results?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2366685,
      "author_name": "Mikhail Kotyushev",
      "author_url": "",
      "post_date": "2023-07-31T06:37:49.117000",
      "content": "<p>Hi! I have not tried multiple GPU vs single GPU setups but I could add some details on reproducibility in general </p>\n<p>1) In addition to what was mentioned by other commenters if you are using pytorch with GPU you should also set <a href=\"https://pytorch.org/docs/stable/notes/randomness.html#avoiding-nondeterministic-algorithms\" target=\"_blank\">deterministic flag</a>. I have used the code below to set seeds for all the used libraries &amp; deterministic behavior for torch.</p>\n<pre><code>def seed_everything(: ):    \n    random.()\n    os.environ[] = str()\n    np.random.()\n    torch.manual_seed()\n    torch.cuda.manual_seed()\n    torch.cuda.manual_seed_all()\n    torch.backends.cudnn.deterministic = True\n    torch.backends.cudnn.benchmark = False\n</code></pre>\n<p>If you use pytorch lightning you could set it with <code>deterministic</code> argument of Trainer &amp; seed_everything function of PL.</p>\n<p><em>Note</em>: some algorithms in pytorch does not have a deterministic CUDA implementation for backward (e. g. bilinear interpolation), but in such case it will raise an error if you set the deterministic flag.</p>\n<p>2) Also for batch size: if you are using multiple machines with different VRAM size and training separately on them (e. g. different CV folds), you could use gradient accumulation to get effectively equal batch size. E. g. you have larger machine with 24Gb VRAM and smaller machine with 16Gb, BS 64 fits into larger machine but not into smaller. Then, using BS 32 + gradient accumulation of 2 on small machine is effectively equal to BS 64. <em>Note</em>: as far as I know accumulation does not take into account batch normalization, so if your model has BNs, results will be different and smaller BS should be used with care.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2366842,
          "author_name": "Monkey D Donut",
          "author_url": "",
          "post_date": "2023-07-31T07:55:15.430000",
          "content": "<p>thank you for the detailed explanation :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2366418": "Hello, \n\nYes, you'll get different results if you aren't using any seed each time you train, irrespective of the processing unit. \n\nIf you want to reproduce your results I suggest you set a seed first and use the same seed to repeat the results next time you train. Also, I found out that it's kinda difficult to reproduce the exact results when using a GPU or a TPU due to the randomness they introduce. So, for the exact replication of the results, go with a CPU and set a seed.",
    "2364113": "I guess that one of the main impact will come from the batch size you will be able to set with these different settings. And this will have a consequent impact on the results, mostly in case you deal with imbalanced problems.",
    "2366685": "Hi! I have not tried multiple GPU vs single GPU setups but I could add some details on reproducibility in general \n\n1) In addition to what was mentioned by other commenters if you are using pytorch with GPU you should also set [deterministic flag](https://pytorch.org/docs/stable/notes/randomness.html#avoiding-nondeterministic-algorithms). I have used the code below to set seeds for all the used libraries & deterministic behavior for torch.\n\n```\ndef seed_everything(seed: int):    \n    random.seed(seed)\n    os.environ['PYTHONHASHSEED'] = str(seed)\n    np.random.seed(seed)\n    torch.manual_seed(seed)\n    torch.cuda.manual_seed(seed)\n    torch.cuda.manual_seed_all(seed)\n    torch.backends.cudnn.deterministic = True\n    torch.backends.cudnn.benchmark = False\n```\n\nIf you use pytorch lightning you could set it with `deterministic` argument of Trainer & seed_everything function of PL.\n\n*Note*: some algorithms in pytorch does not have a deterministic CUDA implementation for backward (e. g. bilinear interpolation), but in such case it will raise an error if you set the deterministic flag.\n\n2) Also for batch size: if you are using multiple machines with different VRAM size and training separately on them (e. g. different CV folds), you could use gradient accumulation to get effectively equal batch size. E. g. you have larger machine with 24Gb VRAM and smaller machine with 16Gb, BS 64 fits into larger machine but not into smaller. Then, using BS 32 + gradient accumulation of 2 on small machine is effectively equal to BS 64. *Note*: as far as I know accumulation does not take into account batch normalization, so if your model has BNs, results will be different and smaller BS should be used with care.",
    "2364092": "hi :),\n   do the effects of training on various machines differ.Will a pretrained model that is fine-tuned using  CPU, one GPU, or multiple GPUs for the same number of epochs produce different results? \nnew to kaggle, sorry if its a stupid question :) \nthank you."
  }
}