{
  "id": 516187,
  "title": "Keras vs PyTorch Baselines LB score drops",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/516187",
  "author_name": "Julian",
  "post_date": "2024-07-01T17:30:24.488000",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi all,</p>\n<p>I noticed that since the test/sample submission data got updated all of the scores for the pytorch implementations are a good ~0.1-0.3 less than the Keras baselines, whereas the Keras baselines only dropped by about ~0-0.02. Is there any particular reason for this discrepancy in scores? Perhaps the <code>keras.layers.Normalization()</code> is somehow better at preserving the precision during normalisation/denormalisation than the PyTorch approaches or maybe it's something completely different. In any case I've seen both from the discussions and from code run on the old challenge server data that people have been able to score quite well (0.6-0.7) with simple approaches such as feed forward nets and on a subset of the challenge data. </p>\n<p>It would be really nice if someone who has already achieved 0.6+ using PyTorch could kindly share/update their baseline code (perhaps with a really simple model) on the Code section to allow better access to participants who aren't so familiar with Keras 😄. </p>\n<p>Thanks a lot in advance! </p>\n<p>---Small Update---</p>\n<p>It seems like normalization &amp; postprocessing parameters have an important role here…</p>\n<p>By making the following changes to <a href=\"https://www.kaggle.com/code/arionlazarus/leapsub2\" target=\"_blank\">this</a> PyTorch version, I was able to go from 0.45 to 0.55 LB:</p>\n<ol>\n<li>min_std=<code>1e-10</code> instead of <code>1e-6</code> (as set in the Keras baseline <a href=\"https://www.kaggle.com/code/enzosebiane/keras-baseline-seq2seq\" target=\"_blank\">notebook</a>)</li>\n<li>Replacing the q0002 values from <code>range(12, 30)</code> instead of <code>range(27)</code></li>\n</ol>",
  "messages": [
    {
      "id": 2899466,
      "postDate": "2024-07-01T17:30:24.490Z",
      "content": "<p>Hi all,</p>\n<p>I noticed that since the test/sample submission data got updated all of the scores for the pytorch implementations are a good ~0.1-0.3 less than the Keras baselines, whereas the Keras baselines only dropped by about ~0-0.02. Is there any particular reason for this discrepancy in scores? Perhaps the <code>keras.layers.Normalization()</code> is somehow better at preserving the precision during normalisation/denormalisation than the PyTorch approaches or maybe it's something completely different. In any case I've seen both from the discussions and from code run on the old challenge server data that people have been able to score quite well (0.6-0.7) with simple approaches such as feed forward nets and on a subset of the challenge data. </p>\n<p>It would be really nice if someone who has already achieved 0.6+ using PyTorch could kindly share/update their baseline code (perhaps with a really simple model) on the Code section to allow better access to participants who aren't so familiar with Keras 😄. </p>\n<p>Thanks a lot in advance! </p>\n<p>---Small Update---</p>\n<p>It seems like normalization &amp; postprocessing parameters have an important role here…</p>\n<p>By making the following changes to <a href=\"https://www.kaggle.com/code/arionlazarus/leapsub2\" target=\"_blank\">this</a> PyTorch version, I was able to go from 0.45 to 0.55 LB:</p>\n<ol>\n<li>min_std=<code>1e-10</code> instead of <code>1e-6</code> (as set in the Keras baseline <a href=\"https://www.kaggle.com/code/enzosebiane/keras-baseline-seq2seq\" target=\"_blank\">notebook</a>)</li>\n<li>Replacing the q0002 values from <code>range(12, 30)</code> instead of <code>range(27)</code></li>\n</ol>",
      "rawMarkdown": "Hi all,\n\nI noticed that since the test/sample submission data got updated all of the scores for the pytorch implementations are a good ~0.1-0.3 less than the Keras baselines, whereas the Keras baselines only dropped by about ~0-0.02. Is there any particular reason for this discrepancy in scores? Perhaps the `keras.layers.Normalization()` is somehow better at preserving the precision during normalisation/denormalisation than the PyTorch approaches or maybe it's something completely different. In any case I've seen both from the discussions and from code run on the old challenge server data that people have been able to score quite well (0.6-0.7) with simple approaches such as feed forward nets and on a subset of the challenge data. \n\nIt would be really nice if someone who has already achieved 0.6+ using PyTorch could kindly share/update their baseline code (perhaps with a really simple model) on the Code section to allow better access to participants who aren't so familiar with Keras 😄. \n\nThanks a lot in advance! \n\n---Small Update---\n\nIt seems like normalization & postprocessing parameters have an important role here...\n\nBy making the following changes to [this](https://www.kaggle.com/code/arionlazarus/leapsub2) PyTorch version, I was able to go from 0.45 to 0.55 LB:\n\n1. min_std=`1e-10` instead of `1e-6` (as set in the Keras baseline [notebook](https://www.kaggle.com/code/enzosebiane/keras-baseline-seq2seq))\n2. Replacing the q0002 values from `range(12, 30)` instead of `range(27)`",
      "votes": 1
    },
    {
      "id": 2899528,
      "postDate": "2024-07-01T17:55:59.060Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2899528,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-07-01T17:55:59.060000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2899466": "Hi all,\n\nI noticed that since the test/sample submission data got updated all of the scores for the pytorch implementations are a good ~0.1-0.3 less than the Keras baselines, whereas the Keras baselines only dropped by about ~0-0.02. Is there any particular reason for this discrepancy in scores? Perhaps the `keras.layers.Normalization()` is somehow better at preserving the precision during normalisation/denormalisation than the PyTorch approaches or maybe it's something completely different. In any case I've seen both from the discussions and from code run on the old challenge server data that people have been able to score quite well (0.6-0.7) with simple approaches such as feed forward nets and on a subset of the challenge data. \n\nIt would be really nice if someone who has already achieved 0.6+ using PyTorch could kindly share/update their baseline code (perhaps with a really simple model) on the Code section to allow better access to participants who aren't so familiar with Keras 😄. \n\nThanks a lot in advance! \n\n---Small Update---\n\nIt seems like normalization & postprocessing parameters have an important role here...\n\nBy making the following changes to [this](https://www.kaggle.com/code/arionlazarus/leapsub2) PyTorch version, I was able to go from 0.45 to 0.55 LB:\n\n1. min_std=`1e-10` instead of `1e-6` (as set in the Keras baseline [notebook](https://www.kaggle.com/code/enzosebiane/keras-baseline-seq2seq))\n2. Replacing the q0002 values from `range(12, 30)` instead of `range(27)`",
    "2899528": ""
  }
}