{
  "id": 523726,
  "title": "Real vs Simulation",
  "url": "/competitions/ariel-data-challenge-2024/discussion/523726",
  "author_name": "Aleksey Trepetsky",
  "post_date": "2024-08-02T11:07:17.934000",
  "votes": 21,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hello!  <br>\nAm I correct in understanding that the training set contains only simulated data of exoplanets, while the test set contains only data from real exoplanets?  <br>\nIf so, how significant is the distribution shift between the two?  <br>\nAdditionally, is it possible to obtain a few examples of data from real exoplanets to see how they differ from the simulations?</p>",
  "messages": [
    {
      "id": 2944356,
      "postDate": "2024-08-02T11:07:17.933Z",
      "content": "<p>Hello!  <br>\nAm I correct in understanding that the training set contains only simulated data of exoplanets, while the test set contains only data from real exoplanets?  <br>\nIf so, how significant is the distribution shift between the two?  <br>\nAdditionally, is it possible to obtain a few examples of data from real exoplanets to see how they differ from the simulations?</p>",
      "rawMarkdown": "Hello!  \nAm I correct in understanding that the training set contains only simulated data of exoplanets, while the test set contains only data from real exoplanets?  \nIf so, how significant is the distribution shift between the two?  \nAdditionally, is it possible to obtain a few examples of data from real exoplanets to see how they differ from the simulations?",
      "votes": 21
    },
    {
      "id": 2944973,
      "postDate": "2024-08-02T22:58:30.060Z",
      "content": "<p>I can confirm that the test data are also simulated</p>",
      "rawMarkdown": "I can confirm that the test data are also simulated",
      "votes": 13,
      "replies": [
        {
          "id": 2944981,
          "postDate": "2024-08-02T23:26:05.127Z",
          "content": "<p><a href=\"https://www.kaggle.com/gordonyip\" target=\"_blank\">@gordonyip</a> Thanks you, for the conformation. </p>",
          "rawMarkdown": "@gordonyip Thanks you, for the conformation. ",
          "votes": 1
        },
        {
          "id": 2965463,
          "postDate": "2024-08-20T22:32:18.613Z",
          "content": "<p>Is the simulated data created across the competition generated in the same run and split, or is it generated in separate processes?</p>",
          "rawMarkdown": "Is the simulated data created across the competition generated in the same run and split, or is it generated in separate processes?",
          "votes": 1
        }
      ]
    },
    {
      "id": 2944370,
      "postDate": "2024-08-02T11:23:18.483Z",
      "content": "<blockquote>\n  <p><strong>A dozen of the exoplanet simulations in the test set were based directly real exoplanets. All of those cases are ignored for scoring purposes.</strong><br>\n  <a href=\"https://www.kaggle.com/alekseytrepetsky\" target=\"_blank\">@alekseytrepetsky</a> test set also having simulated data for scoring.</p>\n</blockquote>\n<hr>\n<blockquote>\n  <p>A dozen of the exoplanet simulations in the test set were based directly real exoplanets  - not for scoring.</p>\n</blockquote>",
      "rawMarkdown": "> **A dozen of the exoplanet simulations in the test set were based directly real exoplanets. All of those cases are ignored for scoring purposes.**\n@alekseytrepetsky test set also having simulated data for scoring.\n\n---\n\n> A dozen of the exoplanet simulations in the test set were based directly real exoplanets  - not for scoring.",
      "votes": 5
    },
    {
      "id": 2954605,
      "postDate": "2024-08-09T23:29:46.733Z",
      "content": "<p>My mental picture consists of two tracks:</p>\n<p>(1) Simulated test data + simulated test data ground truth<br>\nSimulated test data --&gt; ML model --&gt; ML predictions of the simulated test data ground truth</p>\n<p>(2) In 2029, actual observed data + existent but unknown ground truth<br>\nObserved data --&gt; ML model --&gt; ML predictions of the unknown ground truth</p>\n<p>Even if an ML model could do very well for (1), how can we be sure that that inference performance stays well for (2)?  Isn't there always this possibility that the particular exoplanet being observed just happens to \"not fit the bill\" (whatever that means) with the simulation algorithm, thereby rendering its ML predictions in (2) unreliable at best?</p>\n<p>P.S.: I'm aware of Set 4 where both planetary and atmospheric settings are out-range.  Let's assume the ML model does well in (1) even for Set 4.  But my question remains.</p>",
      "rawMarkdown": "My mental picture consists of two tracks:\n\n(1) Simulated test data + simulated test data ground truth\nSimulated test data --> ML model --> ML predictions of the simulated test data ground truth\n\n(2) In 2029, actual observed data + existent but unknown ground truth\nObserved data --> ML model --> ML predictions of the unknown ground truth\n\nEven if an ML model could do very well for (1), how can we be sure that that inference performance stays well for (2)?  Isn't there always this possibility that the particular exoplanet being observed just happens to \"not fit the bill\" (whatever that means) with the simulation algorithm, thereby rendering its ML predictions in (2) unreliable at best?\n\nP.S.: I'm aware of Set 4 where both planetary and atmospheric settings are out-range.  Let's assume the ML model does well in (1) even for Set 4.  But my question remains.",
      "votes": 2,
      "replies": [
        {
          "id": 2956929,
          "postDate": "2024-08-12T16:17:00.200Z",
          "content": "<p>Solving the unknown ground truth is the ultimate goal in science. Over the decades, our strategy has been to continually optimise data reduction techniques and reanalyse previous results whenever we gain new insights about the instrument or the target. A similar approach can be applied to ML models. For example, we can update our simulations with new knowledge and retrain the model.</p>\n<p>In science, we’ve had years to refine this process, while this is the first time we're proposing the same challenge to the ML community. At this stage, we're curious to see how ML will perform. With time, we’re confident that we’ll find the best ways to integrate the insights we gain from your work.</p>\n<p>In this challenge we made the test set to be different from the training, this is to simulate the situation when the data is different or unknown. Of course, future data will not always align with what we know of (our simulation), but there will be takeaways and lessons from the pipeline you are building here, and like we said, those are the insights we are looking for in this challenge. We are not looking for a silver bullet here.</p>",
          "rawMarkdown": "Solving the unknown ground truth is the ultimate goal in science. Over the decades, our strategy has been to continually optimise data reduction techniques and reanalyse previous results whenever we gain new insights about the instrument or the target. A similar approach can be applied to ML models. For example, we can update our simulations with new knowledge and retrain the model.\n\n In science, we’ve had years to refine this process, while this is the first time we're proposing the same challenge to the ML community. At this stage, we're curious to see how ML will perform. With time, we’re confident that we’ll find the best ways to integrate the insights we gain from your work.\n\nIn this challenge we made the test set to be different from the training, this is to simulate the situation when the data is different or unknown. Of course, future data will not always align with what we know of (our simulation), but there will be takeaways and lessons from the pipeline you are building here, and like we said, those are the insights we are looking for in this challenge. We are not looking for a silver bullet here.",
          "votes": 3
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2944973,
      "author_name": "Gordon Yip",
      "author_url": "",
      "post_date": "2024-08-02T22:58:30.060000",
      "content": "<p>I can confirm that the test data are also simulated</p>",
      "votes": 13,
      "replies": [
        {
          "id": 2944981,
          "author_name": "SeshuRaju 🧘‍♂️",
          "author_url": "",
          "post_date": "2024-08-02T23:26:05.127000",
          "content": "<p><a href=\"https://www.kaggle.com/gordonyip\" target=\"_blank\">@gordonyip</a> Thanks you, for the conformation. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2965463,
          "author_name": "Cody_Null",
          "author_url": "",
          "post_date": "2024-08-20T22:32:18.613000",
          "content": "<p>Is the simulated data created across the competition generated in the same run and split, or is it generated in separate processes?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2944370,
      "author_name": "SeshuRaju 🧘‍♂️",
      "author_url": "",
      "post_date": "2024-08-02T11:23:18.483000",
      "content": "<blockquote>\n  <p><strong>A dozen of the exoplanet simulations in the test set were based directly real exoplanets. All of those cases are ignored for scoring purposes.</strong><br>\n  <a href=\"https://www.kaggle.com/alekseytrepetsky\" target=\"_blank\">@alekseytrepetsky</a> test set also having simulated data for scoring.</p>\n</blockquote>\n<hr>\n<blockquote>\n  <p>A dozen of the exoplanet simulations in the test set were based directly real exoplanets  - not for scoring.</p>\n</blockquote>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 2954605,
      "author_name": "Truth Seeker",
      "author_url": "",
      "post_date": "2024-08-09T23:29:46.733000",
      "content": "<p>My mental picture consists of two tracks:</p>\n<p>(1) Simulated test data + simulated test data ground truth<br>\nSimulated test data --&gt; ML model --&gt; ML predictions of the simulated test data ground truth</p>\n<p>(2) In 2029, actual observed data + existent but unknown ground truth<br>\nObserved data --&gt; ML model --&gt; ML predictions of the unknown ground truth</p>\n<p>Even if an ML model could do very well for (1), how can we be sure that that inference performance stays well for (2)?  Isn't there always this possibility that the particular exoplanet being observed just happens to \"not fit the bill\" (whatever that means) with the simulation algorithm, thereby rendering its ML predictions in (2) unreliable at best?</p>\n<p>P.S.: I'm aware of Set 4 where both planetary and atmospheric settings are out-range.  Let's assume the ML model does well in (1) even for Set 4.  But my question remains.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2956929,
          "author_name": "Lorenzo Mugnai",
          "author_url": "",
          "post_date": "2024-08-12T16:17:00.200000",
          "content": "<p>Solving the unknown ground truth is the ultimate goal in science. Over the decades, our strategy has been to continually optimise data reduction techniques and reanalyse previous results whenever we gain new insights about the instrument or the target. A similar approach can be applied to ML models. For example, we can update our simulations with new knowledge and retrain the model.</p>\n<p>In science, we’ve had years to refine this process, while this is the first time we're proposing the same challenge to the ML community. At this stage, we're curious to see how ML will perform. With time, we’re confident that we’ll find the best ways to integrate the insights we gain from your work.</p>\n<p>In this challenge we made the test set to be different from the training, this is to simulate the situation when the data is different or unknown. Of course, future data will not always align with what we know of (our simulation), but there will be takeaways and lessons from the pipeline you are building here, and like we said, those are the insights we are looking for in this challenge. We are not looking for a silver bullet here.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2944356": "Hello!  \nAm I correct in understanding that the training set contains only simulated data of exoplanets, while the test set contains only data from real exoplanets?  \nIf so, how significant is the distribution shift between the two?  \nAdditionally, is it possible to obtain a few examples of data from real exoplanets to see how they differ from the simulations?",
    "2944973": "I can confirm that the test data are also simulated",
    "2944370": "> **A dozen of the exoplanet simulations in the test set were based directly real exoplanets. All of those cases are ignored for scoring purposes.**\n@alekseytrepetsky test set also having simulated data for scoring.\n\n---\n\n> A dozen of the exoplanet simulations in the test set were based directly real exoplanets  - not for scoring.",
    "2954605": "My mental picture consists of two tracks:\n\n(1) Simulated test data + simulated test data ground truth\nSimulated test data --> ML model --> ML predictions of the simulated test data ground truth\n\n(2) In 2029, actual observed data + existent but unknown ground truth\nObserved data --> ML model --> ML predictions of the unknown ground truth\n\nEven if an ML model could do very well for (1), how can we be sure that that inference performance stays well for (2)?  Isn't there always this possibility that the particular exoplanet being observed just happens to \"not fit the bill\" (whatever that means) with the simulation algorithm, thereby rendering its ML predictions in (2) unreliable at best?\n\nP.S.: I'm aware of Set 4 where both planetary and atmospheric settings are out-range.  Let's assume the ML model does well in (1) even for Set 4.  But my question remains."
  }
}