{
  "id": 495258,
  "title": "Why location and timestamps are not provided?",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/495258",
  "author_name": "Bilzard",
  "post_date": "2024-04-20T11:41:17.120000",
  "votes": 30,
  "comment_count": 22,
  "views": 0,
  "content": "<p>I want to ask host and Kaggle team about the task design.</p>\n<p>I briefly look through data tab, and found the latitude, longitude, and timestamp data is dropped on both train and test data. And I wonder why.</p>\n<p>Obviously, feeding these information will allow us to design more various modeling (e.g. time series modeling, prediction using input feature of nearby location etc.).</p>\n<p>I think adding these information won't contradict to the competition's objective to design cheap ML models to simulate climate dynamics.</p>\n<p>Are there any possibility of adding them?</p>\n<p><a href=\"https://www.kaggle.com/ashleychow\" target=\"_blank\">@ashleychow</a>, <a href=\"https://www.kaggle.com/mylesoneill\" target=\"_blank\">@mylesoneill</a></p>",
  "messages": [
    {
      "id": 2763209,
      "postDate": "2024-04-20T11:41:17.120Z",
      "content": "<p>I want to ask host and Kaggle team about the task design.</p>\n<p>I briefly look through data tab, and found the latitude, longitude, and timestamp data is dropped on both train and test data. And I wonder why.</p>\n<p>Obviously, feeding these information will allow us to design more various modeling (e.g. time series modeling, prediction using input feature of nearby location etc.).</p>\n<p>I think adding these information won't contradict to the competition's objective to design cheap ML models to simulate climate dynamics.</p>\n<p>Are there any possibility of adding them?</p>\n<p><a href=\"https://www.kaggle.com/ashleychow\" target=\"_blank\">@ashleychow</a>, <a href=\"https://www.kaggle.com/mylesoneill\" target=\"_blank\">@mylesoneill</a></p>",
      "rawMarkdown": "I want to ask host and Kaggle team about the task design.\n\nI briefly look through data tab, and found the latitude, longitude, and timestamp data is dropped on both train and test data. And I wonder why.\n\nObviously, feeding these information will allow us to design more various modeling (e.g. time series modeling, prediction using input feature of nearby location etc.).\n\nI think adding these information won't contradict to the competition's objective to design cheap ML models to simulate climate dynamics.\n\nAre there any possibility of adding them?\n\n@ashleychow, @mylesoneill",
      "votes": 30
    },
    {
      "id": 2768949,
      "postDate": "2024-04-23T05:13:57.213Z",
      "content": "<p>Hi there. This is a great question, and apologies for not addressing it from the get-go. The idea is to design ML emulators capable of closely emulating Cloud-Resolving Models (CRMs) which are not \"given\" information regarding lat/lon/time. ML emulators that properly learn the physics can hopefully generalize to out-of-sample climates; however, we would ideally like to ensure that they are actually learning the physics rather than geographic patterns.</p>",
      "rawMarkdown": "Hi there. This is a great question, and apologies for not addressing it from the get-go. The idea is to design ML emulators capable of closely emulating Cloud-Resolving Models (CRMs) which are not \"given\" information regarding lat/lon/time. ML emulators that properly learn the physics can hopefully generalize to out-of-sample climates; however, we would ideally like to ensure that they are actually learning the physics rather than geographic patterns.",
      "votes": 12,
      "replies": [
        {
          "id": 2769003,
          "postDate": "2024-04-23T05:44:26.057Z",
          "content": "<blockquote>\n  <p>we would ideally like to ensure that they are actually learning the physics rather than geographic patterns.</p>\n</blockquote>\n<p>Totally make sense. Thanks.</p>",
          "rawMarkdown": "> we would ideally like to ensure that they are actually learning the physics rather than geographic patterns.\n\nTotally make sense. Thanks.",
          "votes": 1
        },
        {
          "id": 2788391,
          "postDate": "2024-05-02T08:07:52.043Z",
          "content": "<blockquote>\n  <p>ML emulators that properly learn the physics can hopefully generalize to out-of-sample climates</p>\n</blockquote>\n<p>I think training data consist of all seasons since it is too large but should we expect unseen locations in test set?</p>",
          "rawMarkdown": "> ML emulators that properly learn the physics can hopefully generalize to out-of-sample climates\n\nI think training data consist of all seasons since it is too large but should we expect unseen locations in test set?",
          "votes": 4
        }
      ]
    },
    {
      "id": 2800888,
      "postDate": "2024-05-08T11:51:44.057Z",
      "content": "<p>While timestamp and location are not given as part of the official competition data, they can be trivially inferred from the original (<a href=\"https://huggingface.co/datasets/LEAP/ClimSim_low-res\" target=\"_blank\">https://huggingface.co/datasets/LEAP/ClimSim_low-res</a>) data: there are 384 locations, all in order, with coordinated given in \"ClimSim_low-res_grid-info.nc\" file, and timestamp is a part of each filename, so is also available. And test data comes in the same format, so also follows the same pattern.</p>\n<p>And location has a large impact on predictability: some locations produce R2 of &gt; 95%, while others struggle to produce any positive R2.</p>\n<p>I strongly suspect that all the top models use location and/or time information - i think that is the only way to generate material improvements in results vs \"unet_preds.csv\" benchmark.</p>",
      "rawMarkdown": "While timestamp and location are not given as part of the official competition data, they can be trivially inferred from the original (https://huggingface.co/datasets/LEAP/ClimSim_low-res) data: there are 384 locations, all in order, with coordinated given in \"ClimSim_low-res_grid-info.nc\" file, and timestamp is a part of each filename, so is also available. And test data comes in the same format, so also follows the same pattern.\n\nAnd location has a large impact on predictability: some locations produce R2 of > 95%, while others struggle to produce any positive R2.\n\nI strongly suspect that all the top models use location and/or time information - i think that is the only way to generate material improvements in results vs \"unet_preds.csv\" benchmark.",
      "votes": 3,
      "replies": [
        {
          "id": 2800927,
          "postDate": "2024-05-08T12:16:53.617Z",
          "content": "<p>I can confirm that score &gt;0.70 can be achieved with only kaggle dataset (no external data used, no reverse engineering)</p>",
          "rawMarkdown": "I can confirm that score >0.70 can be achieved with only kaggle dataset (no external data used, no reverse engineering)",
          "votes": 8,
          "replies": [
            {
              "id": 2801053,
              "postDate": "2024-05-08T13:06:48.440Z",
              "content": "<p>&gt;0.70 is nice, but how about your &gt;0.75, also no rev. engineering? 😅</p>",
              "rawMarkdown": "\\>0.70 is nice, but how about your \\>0.75, also no rev. engineering? 😅"
            },
            {
              "id": 2801064,
              "postDate": "2024-05-08T13:11:01.543Z",
              "content": "<p>Yes! As for now, my score is kaggle dataset only</p>",
              "rawMarkdown": "Yes! As for now, my score is kaggle dataset only",
              "votes": 3
            },
            {
              "id": 2801073,
              "postDate": "2024-05-08T13:16:41.797Z",
              "content": "<p>This is very helpful to know! And very good job on the LB, I anticipate you become GM after this comp. 💯</p>",
              "rawMarkdown": "This is very helpful to know! And very good job on the LB, I anticipate you become GM after this comp. 💯",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2764074,
      "postDate": "2024-04-20T21:27:42.203Z",
      "content": "<p>If I'm not mistaken, the data for each row is self-contained for the purpose of local physics calculations i.e. the physical calculations don't need the surrounding points in time and space for exact calculation of the target variables.<br>\nSo this is a good design of the competition imo.</p>",
      "rawMarkdown": "If I'm not mistaken, the data for each row is self-contained for the purpose of local physics calculations i.e. the physical calculations don't need the surrounding points in time and space for exact calculation of the target variables.\nSo this is a good design of the competition imo.",
      "votes": 4,
      "replies": [
        {
          "id": 2765167,
          "postDate": "2024-04-21T01:50:19.060Z",
          "content": "<p>I think this constraint (without space and time inputs) is optional because authors mentioned in their paper another task design with temporal locality.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fbd45e1887efb329c5ce4a57323dcf73c%2FScreenshot%202024-04-21%20at%2010.45.07.png?generation=1713663950689154&amp;alt=media\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F09a47aac73a7c406946720e1296164f3%2FScreenshot%202024-04-21%20at%2010.57.59.png?generation=1713664719789206&amp;alt=media\"></p>",
          "rawMarkdown": "I think this constraint (without space and time inputs) is optional because authors mentioned in their paper another task design with temporal locality.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fbd45e1887efb329c5ce4a57323dcf73c%2FScreenshot%202024-04-21%20at%2010.45.07.png?generation=1713663950689154&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F09a47aac73a7c406946720e1296164f3%2FScreenshot%202024-04-21%20at%2010.57.59.png?generation=1713664719789206&alt=media)",
          "votes": 2,
          "replies": [
            {
              "id": 2768969,
              "postDate": "2024-04-23T05:25:55.340Z",
              "content": "<p>You're correct there are a few artificial constraints as part of this competition introduced for the sake of simplicity. Spatially and temporally non-local information could indeed be beneficial, and it's possible that an operational emulator will make use of it. However, we also believe substantial headroom exists on the architecture front and it's still very possible that the innovations discovered from this Kaggle competition can make their way to something used operationally. </p>",
              "rawMarkdown": "You're correct there are a few artificial constraints as part of this competition introduced for the sake of simplicity. Spatially and temporally non-local information could indeed be beneficial, and it's possible that an operational emulator will make use of it. However, we also believe substantial headroom exists on the architecture front and it's still very possible that the innovations discovered from this Kaggle competition can make their way to something used operationally. ",
              "votes": 2
            },
            {
              "id": 2768993,
              "postDate": "2024-04-23T05:40:21.453Z",
              "content": "<p><a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> I got it. Thank you for clarification!</p>",
              "rawMarkdown": "@jerrylin96 I got it. Thank you for clarification!",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2831363,
      "postDate": "2024-05-23T16:32:53.617Z",
      "content": "<p>Just to confirm, are we allowed to develop a model that is aware of spatial/temporal information? (This is assuming that the grid info can be recovered from the test csv file).</p>",
      "rawMarkdown": "Just to confirm, are we allowed to develop a model that is aware of spatial/temporal information? (This is assuming that the grid info can be recovered from the test csv file).",
      "votes": 1,
      "replies": [
        {
          "id": 2831372,
          "postDate": "2024-05-23T16:35:52.287Z",
          "content": "<p>I don't see any reason it would not be allowed. BTW time stamp is easy to recover when looking at rolling mean of the features. Grid might be a bit harder.</p>",
          "rawMarkdown": "I don't see any reason it would not be allowed. BTW time stamp is easy to recover when looking at rolling mean of the features. Grid might be a bit harder."
        }
      ]
    },
    {
      "id": 2765769,
      "postDate": "2024-04-21T10:39:35.103Z",
      "content": "<p>I've also been thinking about designing a model that incorporates both spatial and temporal context :/<br>\nBut since we already have access to the final test set in this competition, it's possible that the coordinates and timestamps are excluded such that we cannot simply look up the target values using external climate data sources. Just a guess though…</p>",
      "rawMarkdown": "I've also been thinking about designing a model that incorporates both spatial and temporal context :/\nBut since we already have access to the final test set in this competition, it's possible that the coordinates and timestamps are excluded such that we cannot simply look up the target values using external climate data sources. Just a guess though...",
      "votes": 1,
      "replies": [
        {
          "id": 2766041,
          "postDate": "2024-04-21T13:11:17.940Z",
          "content": "<blockquote>\n  <p>it's possible that the coordinates and timestamps are excluded such that we cannot simply look up the target values using external climate data sources. </p>\n</blockquote>\n<p>I think we don't have to care about this leak, because according to the host's paper, the train/test data are generated with intense computations of simulation. Its possible to reproduce strictly saying, but maybe no-one want pay such amount of GPU computation for this competition.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fe72ac38895f031837de69edc031e2526%2FScreenshot%202024-04-21%20at%2022.08.31.png?generation=1713704931151402&amp;alt=media\"></p>",
          "rawMarkdown": "> it's possible that the coordinates and timestamps are excluded such that we cannot simply look up the target values using external climate data sources. \n\nI think we don't have to care about this leak, because according to the host's paper, the train/test data are generated with intense computations of simulation. Its possible to reproduce strictly saying, but maybe no-one want pay such amount of GPU computation for this competition.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fe72ac38895f031837de69edc031e2526%2FScreenshot%202024-04-21%20at%2022.08.31.png?generation=1713704931151402&alt=media)",
          "votes": 2,
          "replies": [
            {
              "id": 2768958,
              "postDate": "2024-04-23T05:17:18.187Z",
              "content": "<p>The test set uses held-out data not publicly available (on hugging face or otherwise). </p>",
              "rawMarkdown": "The test set uses held-out data not publicly available (on hugging face or otherwise). ",
              "votes": 5
            }
          ]
        }
      ]
    },
    {
      "id": 2765493,
      "postDate": "2024-04-21T06:58:10.583Z",
      "content": "<p>This should be addressed, otherwise it would be hard to match the performance of E3SM-MMF model</p>",
      "rawMarkdown": "This should be addressed, otherwise it would be hard to match the performance of E3SM-MMF model",
      "votes": 1,
      "replies": [
        {
          "id": 2768975,
          "postDate": "2024-04-23T05:27:50.767Z",
          "content": "<p>Please see this comment:</p>\n<p><a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/495258#2768969\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/495258#2768969</a></p>",
          "rawMarkdown": "Please see this comment:\n\nhttps://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/495258#2768969",
          "votes": 1
        }
      ]
    },
    {
      "id": 2857724,
      "postDate": "2024-06-06T05:27:11.687Z",
      "content": "<p>Does the lack of temporal information influence the validation way ?<br>\nIn host's paper, they split train / test by the temporal information, so it would be better to validate in the same way,<br>\nthough the datasets in the paper may be different from kaggle dataset.</p>\n<p>The following sentence is in<code>3 ClimSimDataset Construction</code> of the paper.</p>\n<blockquote>\n  <p>Dataset Split: The 10-year datasets are divided into: (a) a training and validation spanning the<br>\n   first 8 years (0001-02 to 0009-01; YYYY-MM), excluding the first simulated month for numerical<br>\n   spin-up, and (b) a test set spanning the remaining two years (i.e., 0009-03 to 0011-02). A one-month<br>\n   gap is intentionally introduced between the two sets to prevent test set contamination via temporal<br>\n   correlation. Both sets are stored separately in our data repositories.</p>\n</blockquote>",
      "rawMarkdown": "Does the lack of temporal information influence the validation way ?\nIn host's paper, they split train / test by the temporal information, so it would be better to validate in the same way,\nthough the datasets in the paper may be different from kaggle dataset.\n\nThe following sentence is in` 3 ClimSimDataset Construction` of the paper.\n>Dataset Split: The 10-year datasets are divided into: (a) a training and validation spanning the\n first 8 years (0001-02 to 0009-01; YYYY-MM), excluding the first simulated month for numerical\n spin-up, and (b) a test set spanning the remaining two years (i.e., 0009-03 to 0011-02). A one-month\n gap is intentionally introduced between the two sets to prevent test set contamination via temporal\n correlation. Both sets are stored separately in our data repositories."
    },
    {
      "id": 2814740,
      "postDate": "2024-05-15T14:20:42.957Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2871206,
      "postDate": "2024-06-14T05:46:21.693Z",
      "content": "<p>Thank you for support </p>",
      "rawMarkdown": "Thank you for support "
    }
  ],
  "comments": [
    {
      "id": 2768949,
      "author_name": "Jerry Lin",
      "author_url": "",
      "post_date": "2024-04-23T05:13:57.213000",
      "content": "<p>Hi there. This is a great question, and apologies for not addressing it from the get-go. The idea is to design ML emulators capable of closely emulating Cloud-Resolving Models (CRMs) which are not \"given\" information regarding lat/lon/time. ML emulators that properly learn the physics can hopefully generalize to out-of-sample climates; however, we would ideally like to ensure that they are actually learning the physics rather than geographic patterns.</p>",
      "votes": 12,
      "replies": [
        {
          "id": 2769003,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2024-04-23T05:44:26.057000",
          "content": "<blockquote>\n  <p>we would ideally like to ensure that they are actually learning the physics rather than geographic patterns.</p>\n</blockquote>\n<p>Totally make sense. Thanks.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2788391,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2024-05-02T08:07:52.043000",
          "content": "<blockquote>\n  <p>ML emulators that properly learn the physics can hopefully generalize to out-of-sample climates</p>\n</blockquote>\n<p>I think training data consist of all seasons since it is too large but should we expect unseen locations in test set?</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 2800888,
      "author_name": "Youri Matiounine",
      "author_url": "",
      "post_date": "2024-05-08T11:51:44.057000",
      "content": "<p>While timestamp and location are not given as part of the official competition data, they can be trivially inferred from the original (<a href=\"https://huggingface.co/datasets/LEAP/ClimSim_low-res\" target=\"_blank\">https://huggingface.co/datasets/LEAP/ClimSim_low-res</a>) data: there are 384 locations, all in order, with coordinated given in \"ClimSim_low-res_grid-info.nc\" file, and timestamp is a part of each filename, so is also available. And test data comes in the same format, so also follows the same pattern.</p>\n<p>And location has a large impact on predictability: some locations produce R2 of &gt; 95%, while others struggle to produce any positive R2.</p>\n<p>I strongly suspect that all the top models use location and/or time information - i think that is the only way to generate material improvements in results vs \"unet_preds.csv\" benchmark.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2800927,
          "author_name": "slime",
          "author_url": "",
          "post_date": "2024-05-08T12:16:53.617000",
          "content": "<p>I can confirm that score &gt;0.70 can be achieved with only kaggle dataset (no external data used, no reverse engineering)</p>",
          "votes": 8,
          "replies": [
            {
              "id": 2801053,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-05-08T13:06:48.440000",
              "content": "<p>&gt;0.70 is nice, but how about your &gt;0.75, also no rev. engineering? 😅</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2801064,
              "author_name": "slime",
              "author_url": "",
              "post_date": "2024-05-08T13:11:01.543000",
              "content": "<p>Yes! As for now, my score is kaggle dataset only</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2801073,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-05-08T13:16:41.797000",
              "content": "<p>This is very helpful to know! And very good job on the LB, I anticipate you become GM after this comp. 💯</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2764074,
      "author_name": "greySnow",
      "author_url": "",
      "post_date": "2024-04-20T21:27:42.203000",
      "content": "<p>If I'm not mistaken, the data for each row is self-contained for the purpose of local physics calculations i.e. the physical calculations don't need the surrounding points in time and space for exact calculation of the target variables.<br>\nSo this is a good design of the competition imo.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2765167,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2024-04-21T01:50:19.060000",
          "content": "<p>I think this constraint (without space and time inputs) is optional because authors mentioned in their paper another task design with temporal locality.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fbd45e1887efb329c5ce4a57323dcf73c%2FScreenshot%202024-04-21%20at%2010.45.07.png?generation=1713663950689154&amp;alt=media\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F09a47aac73a7c406946720e1296164f3%2FScreenshot%202024-04-21%20at%2010.57.59.png?generation=1713664719789206&amp;alt=media\"></p>",
          "votes": 2,
          "replies": [
            {
              "id": 2768969,
              "author_name": "Jerry Lin",
              "author_url": "",
              "post_date": "2024-04-23T05:25:55.340000",
              "content": "<p>You're correct there are a few artificial constraints as part of this competition introduced for the sake of simplicity. Spatially and temporally non-local information could indeed be beneficial, and it's possible that an operational emulator will make use of it. However, we also believe substantial headroom exists on the architecture front and it's still very possible that the innovations discovered from this Kaggle competition can make their way to something used operationally. </p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2768993,
              "author_name": "Bilzard",
              "author_url": "",
              "post_date": "2024-04-23T05:40:21.453000",
              "content": "<p><a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> I got it. Thank you for clarification!</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2831363,
      "author_name": "Renu Singh",
      "author_url": "",
      "post_date": "2024-05-23T16:32:53.617000",
      "content": "<p>Just to confirm, are we allowed to develop a model that is aware of spatial/temporal information? (This is assuming that the grid info can be recovered from the test csv file).</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2831372,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2024-05-23T16:35:52.287000",
          "content": "<p>I don't see any reason it would not be allowed. BTW time stamp is easy to recover when looking at rolling mean of the features. Grid might be a bit harder.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2765769,
      "author_name": "Felix Yang",
      "author_url": "",
      "post_date": "2024-04-21T10:39:35.103000",
      "content": "<p>I've also been thinking about designing a model that incorporates both spatial and temporal context :/<br>\nBut since we already have access to the final test set in this competition, it's possible that the coordinates and timestamps are excluded such that we cannot simply look up the target values using external climate data sources. Just a guess though…</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2766041,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2024-04-21T13:11:17.940000",
          "content": "<blockquote>\n  <p>it's possible that the coordinates and timestamps are excluded such that we cannot simply look up the target values using external climate data sources. </p>\n</blockquote>\n<p>I think we don't have to care about this leak, because according to the host's paper, the train/test data are generated with intense computations of simulation. Its possible to reproduce strictly saying, but maybe no-one want pay such amount of GPU computation for this competition.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fe72ac38895f031837de69edc031e2526%2FScreenshot%202024-04-21%20at%2022.08.31.png?generation=1713704931151402&amp;alt=media\"></p>",
          "votes": 2,
          "replies": [
            {
              "id": 2768958,
              "author_name": "Jerry Lin",
              "author_url": "",
              "post_date": "2024-04-23T05:17:18.187000",
              "content": "<p>The test set uses held-out data not publicly available (on hugging face or otherwise). </p>",
              "votes": 5,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2765493,
      "author_name": "slime",
      "author_url": "",
      "post_date": "2024-04-21T06:58:10.583000",
      "content": "<p>This should be addressed, otherwise it would be hard to match the performance of E3SM-MMF model</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2768975,
          "author_name": "Jerry Lin",
          "author_url": "",
          "post_date": "2024-04-23T05:27:50.767000",
          "content": "<p>Please see this comment:</p>\n<p><a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/495258#2768969\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/495258#2768969</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2857724,
      "author_name": "Aurora_blue",
      "author_url": "",
      "post_date": "2024-06-06T05:27:11.687000",
      "content": "<p>Does the lack of temporal information influence the validation way ?<br>\nIn host's paper, they split train / test by the temporal information, so it would be better to validate in the same way,<br>\nthough the datasets in the paper may be different from kaggle dataset.</p>\n<p>The following sentence is in<code>3 ClimSimDataset Construction</code> of the paper.</p>\n<blockquote>\n  <p>Dataset Split: The 10-year datasets are divided into: (a) a training and validation spanning the<br>\n   first 8 years (0001-02 to 0009-01; YYYY-MM), excluding the first simulated month for numerical<br>\n   spin-up, and (b) a test set spanning the remaining two years (i.e., 0009-03 to 0011-02). A one-month<br>\n   gap is intentionally introduced between the two sets to prevent test set contamination via temporal<br>\n   correlation. Both sets are stored separately in our data repositories.</p>\n</blockquote>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2814740,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-05-15T14:20:42.957000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2871206,
      "author_name": "Harshita Sapra",
      "author_url": "",
      "post_date": "2024-06-14T05:46:21.693000",
      "content": "<p>Thank you for support </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2763209": "I want to ask host and Kaggle team about the task design.\n\nI briefly look through data tab, and found the latitude, longitude, and timestamp data is dropped on both train and test data. And I wonder why.\n\nObviously, feeding these information will allow us to design more various modeling (e.g. time series modeling, prediction using input feature of nearby location etc.).\n\nI think adding these information won't contradict to the competition's objective to design cheap ML models to simulate climate dynamics.\n\nAre there any possibility of adding them?\n\n@ashleychow, @mylesoneill",
    "2768949": "Hi there. This is a great question, and apologies for not addressing it from the get-go. The idea is to design ML emulators capable of closely emulating Cloud-Resolving Models (CRMs) which are not \"given\" information regarding lat/lon/time. ML emulators that properly learn the physics can hopefully generalize to out-of-sample climates; however, we would ideally like to ensure that they are actually learning the physics rather than geographic patterns.",
    "2800888": "While timestamp and location are not given as part of the official competition data, they can be trivially inferred from the original (https://huggingface.co/datasets/LEAP/ClimSim_low-res) data: there are 384 locations, all in order, with coordinated given in \"ClimSim_low-res_grid-info.nc\" file, and timestamp is a part of each filename, so is also available. And test data comes in the same format, so also follows the same pattern.\n\nAnd location has a large impact on predictability: some locations produce R2 of > 95%, while others struggle to produce any positive R2.\n\nI strongly suspect that all the top models use location and/or time information - i think that is the only way to generate material improvements in results vs \"unet_preds.csv\" benchmark.",
    "2764074": "If I'm not mistaken, the data for each row is self-contained for the purpose of local physics calculations i.e. the physical calculations don't need the surrounding points in time and space for exact calculation of the target variables.\nSo this is a good design of the competition imo.",
    "2831363": "Just to confirm, are we allowed to develop a model that is aware of spatial/temporal information? (This is assuming that the grid info can be recovered from the test csv file).",
    "2765769": "I've also been thinking about designing a model that incorporates both spatial and temporal context :/\nBut since we already have access to the final test set in this competition, it's possible that the coordinates and timestamps are excluded such that we cannot simply look up the target values using external climate data sources. Just a guess though...",
    "2765493": "This should be addressed, otherwise it would be hard to match the performance of E3SM-MMF model",
    "2857724": "Does the lack of temporal information influence the validation way ?\nIn host's paper, they split train / test by the temporal information, so it would be better to validate in the same way,\nthough the datasets in the paper may be different from kaggle dataset.\n\nThe following sentence is in` 3 ClimSimDataset Construction` of the paper.\n>Dataset Split: The 10-year datasets are divided into: (a) a training and validation spanning the\n first 8 years (0001-02 to 0009-01; YYYY-MM), excluding the first simulated month for numerical\n spin-up, and (b) a test set spanning the remaining two years (i.e., 0009-03 to 0011-02). A one-month\n gap is intentionally introduced between the two sets to prevent test set contamination via temporal\n correlation. Both sets are stored separately in our data repositories.",
    "2814740": "",
    "2871206": "Thank you for support "
  }
}