{
  "id": 516605,
  "title": "New paper relevant to this competition",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/516605",
  "author_name": "Jerry Lin",
  "post_date": "2024-07-03T03:39:24.586000",
  "votes": 25,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hello everyone,</p>\n<p>Thought I should let you know that there is a new paper out that's highly relevant to this competition:</p>\n<p><strong>Stable Machine-Learning Parameterization of Subgrid Processes with Real Geography and Full-physics Emulation</strong></p>\n<p><a href=\"https://arxiv.org/abs/2407.00124\" target=\"_blank\">https://arxiv.org/abs/2407.00124</a></p>\n<p>This paper couples a U-net convective parameterization to E3SM-MMF that incorporates a few new key innovations and achieves a stable integration with state-of-the-art zonal mean bias. Some of the advances, like using convective memory, are not possible with this competition while others, like using a microphysics constraint, are readily applicable.</p>\n<p>Hopefully you find some of the findings and analysis useful, and we wish you luck with the rest of the competition! Your work is highly valued by our community, and we look forward to learning from you.</p>\n<p>Best regards,</p>\n<p>Jerry</p>",
  "messages": [
    {
      "id": 2901978,
      "postDate": "2024-07-03T03:39:24.587Z",
      "content": "<p>Hello everyone,</p>\n<p>Thought I should let you know that there is a new paper out that's highly relevant to this competition:</p>\n<p><strong>Stable Machine-Learning Parameterization of Subgrid Processes with Real Geography and Full-physics Emulation</strong></p>\n<p><a href=\"https://arxiv.org/abs/2407.00124\" target=\"_blank\">https://arxiv.org/abs/2407.00124</a></p>\n<p>This paper couples a U-net convective parameterization to E3SM-MMF that incorporates a few new key innovations and achieves a stable integration with state-of-the-art zonal mean bias. Some of the advances, like using convective memory, are not possible with this competition while others, like using a microphysics constraint, are readily applicable.</p>\n<p>Hopefully you find some of the findings and analysis useful, and we wish you luck with the rest of the competition! Your work is highly valued by our community, and we look forward to learning from you.</p>\n<p>Best regards,</p>\n<p>Jerry</p>",
      "rawMarkdown": "Hello everyone,\n\nThought I should let you know that there is a new paper out that's highly relevant to this competition:\n\n**Stable Machine-Learning Parameterization of Subgrid Processes with Real Geography and Full-physics Emulation**\n\nhttps://arxiv.org/abs/2407.00124\n\nThis paper couples a U-net convective parameterization to E3SM-MMF that incorporates a few new key innovations and achieves a stable integration with state-of-the-art zonal mean bias. Some of the advances, like using convective memory, are not possible with this competition while others, like using a microphysics constraint, are readily applicable.\n\nHopefully you find some of the findings and analysis useful, and we wish you luck with the rest of the competition! Your work is highly valued by our community, and we look forward to learning from you.\n\nBest regards,\n\nJerry",
      "votes": 24
    },
    {
      "id": 2902316,
      "postDate": "2024-07-03T07:37:07.550Z",
      "content": "<p>So I tried to implement the first microphysics postprocessing for my best submission, but my score got down by 0.3. This is probably because my best submission is far from being good, and because the postprocessing depends on the quality of the prediction of the temperature tendency, if that prediction is already not that good, the postprocessing will also not make much sense.</p>\n<p>Nevertheless, I wanted to share the code I used, in case someone sees an error in it. df_test is the original test file in pandas format, and df_p_test is my best submission, also in pandas format.</p>\n<p>EDIT: I published <a href=\"https://www.kaggle.com/code/fpeccia/leap-all-postprocessing-submission-techniques\" target=\"_blank\">here</a> the correct implementation. My score improved a little using this, but not much.</p>\n<pre><code>ICE_THRESHOLD = \nLIQUID_THRESHOLD = \n i  tqdm(()):\n    new_t = df_test[] + df_p_test[]* \n\n    prev_q0002 = df_p_test[]\n    prev_q0002[new_t &lt; ICE_THRESHOLD] =  \n\n    prev_q0003 = df_p_test[]\n    prev_q0003[new_t &gt; LIQUID_THRESHOLD] =  \n\n    df_p_test[] = prev_q0002\n    df_p_test[] = prev_q0003\n</code></pre>",
      "rawMarkdown": "So I tried to implement the first microphysics postprocessing for my best submission, but my score got down by 0.3. This is probably because my best submission is far from being good, and because the postprocessing depends on the quality of the prediction of the temperature tendency, if that prediction is already not that good, the postprocessing will also not make much sense.\n\nNevertheless, I wanted to share the code I used, in case someone sees an error in it. df_test is the original test file in pandas format, and df_p_test is my best submission, also in pandas format.\n\nEDIT: I published [here](https://www.kaggle.com/code/fpeccia/leap-all-postprocessing-submission-techniques) the correct implementation. My score improved a little using this, but not much.\n\n```\nICE_THRESHOLD = 253.16\nLIQUID_THRESHOLD = 273.16\nfor i in tqdm(range(60)):\n    new_t = df_test[f\"state_t_{i}\"] + df_p_test[f\"ptend_t_{i}\"]*1200. # ptend_t unit is K/s, so I multiply by the seconds of one timestep\n    \n    prev_q0002 = df_p_test[f\"ptend_q0002_{i}\"]\n    prev_q0002[new_t < ICE_THRESHOLD] = 0. # when the temperature is below this threshold, liquid mixing ratio is zero\n    \n    prev_q0003 = df_p_test[f\"ptend_q0003_{i}\"]\n    prev_q0003[new_t > LIQUID_THRESHOLD] = 0. # when the temperature is above this threshold, ice mixing ratio is zero\n    \n    df_p_test[f\"ptend_q0002_{i}\"] = prev_q0002\n    df_p_test[f\"ptend_q0003_{i}\"] = prev_q0003\n```",
      "votes": 11,
      "replies": [
        {
          "id": 2902376,
          "postDate": "2024-07-03T08:20:09.690Z",
          "content": "<p>This implementation makes a zero correction for \"ptend_\", but shouldn't it be corrected for the \"state_\" at the next time step obtained from \"ptend_\"?</p>",
          "rawMarkdown": "This implementation makes a zero correction for \"ptend_\", but shouldn't it be corrected for the \"state_\" at the next time step obtained from \"ptend_\"?",
          "replies": [
            {
              "id": 2902459,
              "postDate": "2024-07-03T09:06:56.047Z",
              "content": "<p>You are right, I am correcting for the wrong thing :(</p>",
              "rawMarkdown": "You are right, I am correcting for the wrong thing :("
            }
          ]
        },
        {
          "id": 2902448,
          "postDate": "2024-07-03T08:58:52.827Z",
          "content": "<p>It's important to remember that the main aim of the physical constraints is to stabilize over long times, i.e., when your prediction at step n+1 is based on the prediction at step n for n&gt;&gt;1. This is at least my understanding. I doubt these constraints would be helpful in this competition, and it makes sense they will hurt the score. (But maybe they will help, regardless!).</p>",
          "rawMarkdown": "It's important to remember that the main aim of the physical constraints is to stabilize over long times, i.e., when your prediction at step n+1 is based on the prediction at step n for n>>1. This is at least my understanding. I doubt these constraints would be helpful in this competition, and it makes sense they will hurt the score. (But maybe they will help, regardless!).",
          "votes": 5
        },
        {
          "id": 2905818,
          "postDate": "2024-07-05T07:27:40.400Z",
          "content": "<p>Due to the limitation of domain knowledge, data preprocessing and postprocessing are particularly important. Thanks for sharing.</p>",
          "rawMarkdown": "Due to the limitation of domain knowledge, data preprocessing and postprocessing are particularly important. Thanks for sharing."
        },
        {
          "id": 2906393,
          "postDate": "2024-07-05T14:55:15.773Z",
          "content": "<p>This post-processing doesn't work for me</p>",
          "rawMarkdown": "This post-processing doesn't work for me"
        }
      ]
    },
    {
      "id": 2903490,
      "postDate": "2024-07-03T20:04:10.557Z",
      "content": "<p>Jerry: thanks for sharing the paper! at the bottom of page 9 and top of page 10, some internal author discussions are inadvertently left out (\"Could you confirm if… maybe trivial to add though\")</p>",
      "rawMarkdown": "Jerry: thanks for sharing the paper! at the bottom of page 9 and top of page 10, some internal author discussions are inadvertently left out (\"Could you confirm if... maybe trivial to add though\")",
      "votes": 2,
      "replies": [
        {
          "id": 2904103,
          "postDate": "2024-07-04T07:16:18.807Z",
          "content": "<p>Thanks for pointing this out! Looks like we were a bit hasty putting this out. </p>",
          "rawMarkdown": "Thanks for pointing this out! Looks like we were a bit hasty putting this out. "
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2902316,
      "author_name": "Federico Peccia",
      "author_url": "",
      "post_date": "2024-07-03T07:37:07.550000",
      "content": "<p>So I tried to implement the first microphysics postprocessing for my best submission, but my score got down by 0.3. This is probably because my best submission is far from being good, and because the postprocessing depends on the quality of the prediction of the temperature tendency, if that prediction is already not that good, the postprocessing will also not make much sense.</p>\n<p>Nevertheless, I wanted to share the code I used, in case someone sees an error in it. df_test is the original test file in pandas format, and df_p_test is my best submission, also in pandas format.</p>\n<p>EDIT: I published <a href=\"https://www.kaggle.com/code/fpeccia/leap-all-postprocessing-submission-techniques\" target=\"_blank\">here</a> the correct implementation. My score improved a little using this, but not much.</p>\n<pre><code>ICE_THRESHOLD = \nLIQUID_THRESHOLD = \n i  tqdm(()):\n    new_t = df_test[] + df_p_test[]* \n\n    prev_q0002 = df_p_test[]\n    prev_q0002[new_t &lt; ICE_THRESHOLD] =  \n\n    prev_q0003 = df_p_test[]\n    prev_q0003[new_t &gt; LIQUID_THRESHOLD] =  \n\n    df_p_test[] = prev_q0002\n    df_p_test[] = prev_q0003\n</code></pre>",
      "votes": 11,
      "replies": [
        {
          "id": 2902376,
          "author_name": "kuto",
          "author_url": "",
          "post_date": "2024-07-03T08:20:09.690000",
          "content": "<p>This implementation makes a zero correction for \"ptend_\", but shouldn't it be corrected for the \"state_\" at the next time step obtained from \"ptend_\"?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2902459,
              "author_name": "Federico Peccia",
              "author_url": "",
              "post_date": "2024-07-03T09:06:56.047000",
              "content": "<p>You are right, I am correcting for the wrong thing :(</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2902448,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2024-07-03T08:58:52.827000",
          "content": "<p>It's important to remember that the main aim of the physical constraints is to stabilize over long times, i.e., when your prediction at step n+1 is based on the prediction at step n for n&gt;&gt;1. This is at least my understanding. I doubt these constraints would be helpful in this competition, and it makes sense they will hurt the score. (But maybe they will help, regardless!).</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 2905818,
          "author_name": "ynhuhu",
          "author_url": "",
          "post_date": "2024-07-05T07:27:40.400000",
          "content": "<p>Due to the limitation of domain knowledge, data preprocessing and postprocessing are particularly important. Thanks for sharing.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2906393,
          "author_name": "Zhuoqun Li",
          "author_url": "",
          "post_date": "2024-07-05T14:55:15.773000",
          "content": "<p>This post-processing doesn't work for me</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2903490,
      "author_name": "Sukanta Basu",
      "author_url": "",
      "post_date": "2024-07-03T20:04:10.557000",
      "content": "<p>Jerry: thanks for sharing the paper! at the bottom of page 9 and top of page 10, some internal author discussions are inadvertently left out (\"Could you confirm if… maybe trivial to add though\")</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2904103,
          "author_name": "Jerry Lin",
          "author_url": "",
          "post_date": "2024-07-04T07:16:18.807000",
          "content": "<p>Thanks for pointing this out! Looks like we were a bit hasty putting this out. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2901978": "Hello everyone,\n\nThought I should let you know that there is a new paper out that's highly relevant to this competition:\n\n**Stable Machine-Learning Parameterization of Subgrid Processes with Real Geography and Full-physics Emulation**\n\nhttps://arxiv.org/abs/2407.00124\n\nThis paper couples a U-net convective parameterization to E3SM-MMF that incorporates a few new key innovations and achieves a stable integration with state-of-the-art zonal mean bias. Some of the advances, like using convective memory, are not possible with this competition while others, like using a microphysics constraint, are readily applicable.\n\nHopefully you find some of the findings and analysis useful, and we wish you luck with the rest of the competition! Your work is highly valued by our community, and we look forward to learning from you.\n\nBest regards,\n\nJerry",
    "2902316": "So I tried to implement the first microphysics postprocessing for my best submission, but my score got down by 0.3. This is probably because my best submission is far from being good, and because the postprocessing depends on the quality of the prediction of the temperature tendency, if that prediction is already not that good, the postprocessing will also not make much sense.\n\nNevertheless, I wanted to share the code I used, in case someone sees an error in it. df_test is the original test file in pandas format, and df_p_test is my best submission, also in pandas format.\n\nEDIT: I published [here](https://www.kaggle.com/code/fpeccia/leap-all-postprocessing-submission-techniques) the correct implementation. My score improved a little using this, but not much.\n\n```\nICE_THRESHOLD = 253.16\nLIQUID_THRESHOLD = 273.16\nfor i in tqdm(range(60)):\n    new_t = df_test[f\"state_t_{i}\"] + df_p_test[f\"ptend_t_{i}\"]*1200. # ptend_t unit is K/s, so I multiply by the seconds of one timestep\n    \n    prev_q0002 = df_p_test[f\"ptend_q0002_{i}\"]\n    prev_q0002[new_t < ICE_THRESHOLD] = 0. # when the temperature is below this threshold, liquid mixing ratio is zero\n    \n    prev_q0003 = df_p_test[f\"ptend_q0003_{i}\"]\n    prev_q0003[new_t > LIQUID_THRESHOLD] = 0. # when the temperature is above this threshold, ice mixing ratio is zero\n    \n    df_p_test[f\"ptend_q0002_{i}\"] = prev_q0002\n    df_p_test[f\"ptend_q0003_{i}\"] = prev_q0003\n```",
    "2903490": "Jerry: thanks for sharing the paper! at the bottom of page 9 and top of page 10, some internal author discussions are inadvertently left out (\"Could you confirm if... maybe trivial to add though\")"
  }
}