{
  "id": 609706,
  "title": "An alternative form for this competition: known data generation pipeline",
  "url": "/competitions/ariel-data-challenge-2025/discussion/609706",
  "author_name": "Jeroen Cottaar",
  "post_date": "2025-09-29T06:10:06.666000",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>(TLDR at bottom. And my writeup is coming soon…)</p>\n<p>In a recent discussion thread, I proposed sharing the full data generation pipeline in a reproducible form for the 2026 competition. This is completely different from the current form of the competition. I'd like to explain in a bit more detail why I think this should be considered.</p>\n<p>First, let's consider the scientific goals of this competition. As I see it, there are two key challenges to the ARIEL mission in terms of inference:</p>\n<ol>\n<li><p><strong>Determine the forward model</strong>: given an exoplanet transit depth spectrum, what are the sensor readings (including all disturbances)?</p></li>\n<li><p><strong>Determine the reverse model</strong>: invert the forward model to find the transit depths from sensor readings, including realistic uncertainty estimates.</p></li>\n</ol>\n<p>Goal (2) is fully covered by the current competition design, but goal (1) is not so straightforward. Currently, the challenge of determining the forward model seems to be introduced by the fact that the data creation pipeline is obfuscated. We are essentially asked to then reverse engineer this data creation pipeline, as a proxy for the reverse engineering of the true physics the ARIEL team will have to do.</p>\n<p>Unfortunately, the approach of obfuscating the data pipeline to simulate the unknown physics of the universe is flawed. The key issue is that in reality training labels will not be available, i.e. there will not be a set of planets for which the transit depths are known. This means that all our work that uses the provided transit depths is not relevant to the final mission.</p>\n<p>I can't come up with a good way to cover goal (1) above. One option could be to not release transit depths at all, not even for the training set. This is, after all, the situation that the ARIEL team will find itself in in reality. However, in that case the competition will almost certainly come down to test set probing, which is again not something that will be possible in reality.</p>\n<p>So the alternative I propose: accept that goal (1) is not covered, and focus fully on goal (2). Share the full synthetic data generation pipeline, to the degree that the community can create new data. Or alternatively: share the pipeline, but keep some details secret, such as the distributions you draw physical parameters from. We would then have to demonstrate the ability to learn this from data - ideally without given transit depths (but this may complicate matters too much).</p>\n<p>Another alternative to consider (regardless of whether the full pipeline is shared): consider sharing the test set data, both public and private. Rather than submitting code, we'd submit predictions directly. This is more representative of the final mission, where you don't need a 'blind' model.</p>\n<p>Note that we had a recent competition implementing both proposals above, i.e. sharing the prediction pipeline and providing the test set to competitors: the <a href=\"https://www.kaggle.com/competitions/waveform-inversion\" target=\"_blank\">Yale/UNC geophysical waveform inversion competition</a>. As far as I can tell this competition was very well received by the community.</p>\n<p>TLDR: Because the real mission won't have known ground truths, this competition isn't demonstrating the ability to learn physics. Rather, our contributions are on the inference side, i.e. inverting the physics. Embrace this, and share the full data generation pipeline. Separately, also consider releasing the test set to competitors during the competition. A <a href=\"https://www.kaggle.com/competitions/waveform-inversion\" target=\"_blank\">recent well-received competition</a> did both of these.</p>",
  "messages": [
    {
      "id": 3295565,
      "postDate": "2025-09-29T06:10:06.667Z",
      "content": "<p>(TLDR at bottom. And my writeup is coming soon…)</p>\n<p>In a recent discussion thread, I proposed sharing the full data generation pipeline in a reproducible form for the 2026 competition. This is completely different from the current form of the competition. I'd like to explain in a bit more detail why I think this should be considered.</p>\n<p>First, let's consider the scientific goals of this competition. As I see it, there are two key challenges to the ARIEL mission in terms of inference:</p>\n<ol>\n<li><p><strong>Determine the forward model</strong>: given an exoplanet transit depth spectrum, what are the sensor readings (including all disturbances)?</p></li>\n<li><p><strong>Determine the reverse model</strong>: invert the forward model to find the transit depths from sensor readings, including realistic uncertainty estimates.</p></li>\n</ol>\n<p>Goal (2) is fully covered by the current competition design, but goal (1) is not so straightforward. Currently, the challenge of determining the forward model seems to be introduced by the fact that the data creation pipeline is obfuscated. We are essentially asked to then reverse engineer this data creation pipeline, as a proxy for the reverse engineering of the true physics the ARIEL team will have to do.</p>\n<p>Unfortunately, the approach of obfuscating the data pipeline to simulate the unknown physics of the universe is flawed. The key issue is that in reality training labels will not be available, i.e. there will not be a set of planets for which the transit depths are known. This means that all our work that uses the provided transit depths is not relevant to the final mission.</p>\n<p>I can't come up with a good way to cover goal (1) above. One option could be to not release transit depths at all, not even for the training set. This is, after all, the situation that the ARIEL team will find itself in in reality. However, in that case the competition will almost certainly come down to test set probing, which is again not something that will be possible in reality.</p>\n<p>So the alternative I propose: accept that goal (1) is not covered, and focus fully on goal (2). Share the full synthetic data generation pipeline, to the degree that the community can create new data. Or alternatively: share the pipeline, but keep some details secret, such as the distributions you draw physical parameters from. We would then have to demonstrate the ability to learn this from data - ideally without given transit depths (but this may complicate matters too much).</p>\n<p>Another alternative to consider (regardless of whether the full pipeline is shared): consider sharing the test set data, both public and private. Rather than submitting code, we'd submit predictions directly. This is more representative of the final mission, where you don't need a 'blind' model.</p>\n<p>Note that we had a recent competition implementing both proposals above, i.e. sharing the prediction pipeline and providing the test set to competitors: the <a href=\"https://www.kaggle.com/competitions/waveform-inversion\" target=\"_blank\">Yale/UNC geophysical waveform inversion competition</a>. As far as I can tell this competition was very well received by the community.</p>\n<p>TLDR: Because the real mission won't have known ground truths, this competition isn't demonstrating the ability to learn physics. Rather, our contributions are on the inference side, i.e. inverting the physics. Embrace this, and share the full data generation pipeline. Separately, also consider releasing the test set to competitors during the competition. A <a href=\"https://www.kaggle.com/competitions/waveform-inversion\" target=\"_blank\">recent well-received competition</a> did both of these.</p>",
      "rawMarkdown": "(TLDR at bottom. And my writeup is coming soon...)\n\nIn a recent discussion thread, I proposed sharing the full data generation pipeline in a reproducible form for the 2026 competition. This is completely different from the current form of the competition. I'd like to explain in a bit more detail why I think this should be considered.\n\nFirst, let's consider the scientific goals of this competition. As I see it, there are two key challenges to the ARIEL mission in terms of inference:\n\n1. **Determine the forward model**: given an exoplanet transit depth spectrum, what are the sensor readings (including all disturbances)?\n\n2. **Determine the reverse model**: invert the forward model to find the transit depths from sensor readings, including realistic uncertainty estimates.\n\nGoal (2) is fully covered by the current competition design, but goal (1) is not so straightforward. Currently, the challenge of determining the forward model seems to be introduced by the fact that the data creation pipeline is obfuscated. We are essentially asked to then reverse engineer this data creation pipeline, as a proxy for the reverse engineering of the true physics the ARIEL team will have to do.\n\nUnfortunately, the approach of obfuscating the data pipeline to simulate the unknown physics of the universe is flawed. The key issue is that in reality training labels will not be available, i.e. there will not be a set of planets for which the transit depths are known. This means that all our work that uses the provided transit depths is not relevant to the final mission.\n\nI can't come up with a good way to cover goal (1) above. One option could be to not release transit depths at all, not even for the training set. This is, after all, the situation that the ARIEL team will find itself in in reality. However, in that case the competition will almost certainly come down to test set probing, which is again not something that will be possible in reality.\n\nSo the alternative I propose: accept that goal (1) is not covered, and focus fully on goal (2). Share the full synthetic data generation pipeline, to the degree that the community can create new data. Or alternatively: share the pipeline, but keep some details secret, such as the distributions you draw physical parameters from. We would then have to demonstrate the ability to learn this from data - ideally without given transit depths (but this may complicate matters too much).\n\nAnother alternative to consider (regardless of whether the full pipeline is shared): consider sharing the test set data, both public and private. Rather than submitting code, we'd submit predictions directly. This is more representative of the final mission, where you don't need a 'blind' model.\n\nNote that we had a recent competition implementing both proposals above, i.e. sharing the prediction pipeline and providing the test set to competitors: the [Yale/UNC geophysical waveform inversion competition](https://www.kaggle.com/competitions/waveform-inversion). As far as I can tell this competition was very well received by the community.\n\nTLDR: Because the real mission won't have known ground truths, this competition isn't demonstrating the ability to learn physics. Rather, our contributions are on the inference side, i.e. inverting the physics. Embrace this, and share the full data generation pipeline. Separately, also consider releasing the test set to competitors during the competition. A [recent well-received competition](https://www.kaggle.com/competitions/waveform-inversion) did both of these.",
      "votes": 2
    },
    {
      "id": 3296579,
      "postDate": "2025-10-01T08:01:32.963Z",
      "content": "<p>Great discussion - I will come back, I am in a conference and has been slow with my reply but i have been monitoring!!</p>",
      "rawMarkdown": "Great discussion - I will come back, I am in a conference and has been slow with my reply but i have been monitoring!!"
    },
    {
      "id": 3295861,
      "postDate": "2025-09-29T17:20:44.383Z",
      "content": "<p>I have the impression that a publicly-available generation pipeline would be helpful for understanding effects like the one you asked about <a href=\"https://www.kaggle.com/competitions/ariel-data-challenge-2025/discussion/602425\" target=\"_blank\">here</a>.</p>\n<p>I also have the sense that determining the forward model is a very open-ended problem, at least part of which might be framed as an effort to tune and validate the accuracy of the simulation.  With the caveat that I don't have a lot of experience in astronomy, I'd imagine that a nontrivial part of this would be done by making observations of star systems whose properties are well-known, including systems that are known not to have exoplanets (so that you can measure and characterize, in the definite absence of any transit signal, things like noise, cosmics, and drift, as they affect the real hardware on the actual satellite) as well as systems that do contain one or more exoplanets that have known transits which have been well-studied by other telescopes.   With or without an open generation pipeline, I could imagine future installments of this competition that include calibration/validation samples like these; contestants would then be tasked with identifying problems with the detector and/or simulation based on the provided calibration/validation samples, determining appropriate corrections, and applying them before making predictions.  </p>\n<p>I get that this probably only covers, at best, part of the question of determining the forward model.  But I would imagine that other parts of the forward model could be pinned down with other methods, e.g. demonstrate an ability to learn the limb-darkening law by generating transits using a randomly-chosen one of several different limb-darkening models and ask contestants to predict which one was used for this transit and what its parameters were.  But I would expect this sort of study of actual physics to be deferred until the detector is well-characterized/well-understood/well-calibrated.</p>",
      "rawMarkdown": "I have the impression that a publicly-available generation pipeline would be helpful for understanding effects like the one you asked about [here](https://www.kaggle.com/competitions/ariel-data-challenge-2025/discussion/602425).\n\nI also have the sense that determining the forward model is a very open-ended problem, at least part of which might be framed as an effort to tune and validate the accuracy of the simulation.  With the caveat that I don't have a lot of experience in astronomy, I'd imagine that a nontrivial part of this would be done by making observations of star systems whose properties are well-known, including systems that are known not to have exoplanets (so that you can measure and characterize, in the definite absence of any transit signal, things like noise, cosmics, and drift, as they affect the real hardware on the actual satellite) as well as systems that do contain one or more exoplanets that have known transits which have been well-studied by other telescopes.   With or without an open generation pipeline, I could imagine future installments of this competition that include calibration/validation samples like these; contestants would then be tasked with identifying problems with the detector and/or simulation based on the provided calibration/validation samples, determining appropriate corrections, and applying them before making predictions.  \n\nI get that this probably only covers, at best, part of the question of determining the forward model.  But I would imagine that other parts of the forward model could be pinned down with other methods, e.g. demonstrate an ability to learn the limb-darkening law by generating transits using a randomly-chosen one of several different limb-darkening models and ask contestants to predict which one was used for this transit and what its parameters were.  But I would expect this sort of study of actual physics to be deferred until the detector is well-characterized/well-understood/well-calibrated.\n"
    },
    {
      "id": 3295725,
      "postDate": "2025-09-29T11:56:34.613Z",
      "content": "<p>If I understand correctly, the classical way to extract physics from data measured with a well-known detector would be something like:</p>\n<ul>\n<li>build a physical model (Rp/Rs, P, sma, t0, i, e, LD, background, …),</li>\n<li>fold it with the detector response (time sampling, noise, systematics),</li>\n<li>and then fit to data to recover parameter posteriors with the uncertainties.<br>\nIn this way you recover physics from first principles. However, it's often with degeneracies between parameters.</li>\n</ul>\n<p>In this competition, probably the goal was to learn a mapping data -&gt; RpRs directly, and avoid the parameters degeneracy problem. If the full data generation pipeline was shared, we could build a better model via data -&gt; unfolded data -&gt; RpRs split.</p>",
      "rawMarkdown": "If I understand correctly, the classical way to extract physics from data measured with a well-known detector would be something like:\n- build a physical model (Rp/Rs, P, sma, t0, i, e, LD, background, …),\n- fold it with the detector response (time sampling, noise, systematics),\n- and then fit to data to recover parameter posteriors with the uncertainties.\nIn this way you recover physics from first principles. However, it's often with degeneracies between parameters.\n\nIn this competition, probably the goal was to learn a mapping data -> RpRs directly, and avoid the parameters degeneracy problem. If the full data generation pipeline was shared, we could build a better model via data -> unfolded data -> RpRs split."
    }
  ],
  "comments": [
    {
      "id": 3296579,
      "author_name": "Gordon Yip",
      "author_url": "",
      "post_date": "2025-10-01T08:01:32.963000",
      "content": "<p>Great discussion - I will come back, I am in a conference and has been slow with my reply but i have been monitoring!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3295861,
      "author_name": "particlebbq",
      "author_url": "",
      "post_date": "2025-09-29T17:20:44.383000",
      "content": "<p>I have the impression that a publicly-available generation pipeline would be helpful for understanding effects like the one you asked about <a href=\"https://www.kaggle.com/competitions/ariel-data-challenge-2025/discussion/602425\" target=\"_blank\">here</a>.</p>\n<p>I also have the sense that determining the forward model is a very open-ended problem, at least part of which might be framed as an effort to tune and validate the accuracy of the simulation.  With the caveat that I don't have a lot of experience in astronomy, I'd imagine that a nontrivial part of this would be done by making observations of star systems whose properties are well-known, including systems that are known not to have exoplanets (so that you can measure and characterize, in the definite absence of any transit signal, things like noise, cosmics, and drift, as they affect the real hardware on the actual satellite) as well as systems that do contain one or more exoplanets that have known transits which have been well-studied by other telescopes.   With or without an open generation pipeline, I could imagine future installments of this competition that include calibration/validation samples like these; contestants would then be tasked with identifying problems with the detector and/or simulation based on the provided calibration/validation samples, determining appropriate corrections, and applying them before making predictions.  </p>\n<p>I get that this probably only covers, at best, part of the question of determining the forward model.  But I would imagine that other parts of the forward model could be pinned down with other methods, e.g. demonstrate an ability to learn the limb-darkening law by generating transits using a randomly-chosen one of several different limb-darkening models and ask contestants to predict which one was used for this transit and what its parameters were.  But I would expect this sort of study of actual physics to be deferred until the detector is well-characterized/well-understood/well-calibrated.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3295725,
      "author_name": "Oleh Kivernyk",
      "author_url": "",
      "post_date": "2025-09-29T11:56:34.613000",
      "content": "<p>If I understand correctly, the classical way to extract physics from data measured with a well-known detector would be something like:</p>\n<ul>\n<li>build a physical model (Rp/Rs, P, sma, t0, i, e, LD, background, …),</li>\n<li>fold it with the detector response (time sampling, noise, systematics),</li>\n<li>and then fit to data to recover parameter posteriors with the uncertainties.<br>\nIn this way you recover physics from first principles. However, it's often with degeneracies between parameters.</li>\n</ul>\n<p>In this competition, probably the goal was to learn a mapping data -&gt; RpRs directly, and avoid the parameters degeneracy problem. If the full data generation pipeline was shared, we could build a better model via data -&gt; unfolded data -&gt; RpRs split.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3295565": "(TLDR at bottom. And my writeup is coming soon...)\n\nIn a recent discussion thread, I proposed sharing the full data generation pipeline in a reproducible form for the 2026 competition. This is completely different from the current form of the competition. I'd like to explain in a bit more detail why I think this should be considered.\n\nFirst, let's consider the scientific goals of this competition. As I see it, there are two key challenges to the ARIEL mission in terms of inference:\n\n1. **Determine the forward model**: given an exoplanet transit depth spectrum, what are the sensor readings (including all disturbances)?\n\n2. **Determine the reverse model**: invert the forward model to find the transit depths from sensor readings, including realistic uncertainty estimates.\n\nGoal (2) is fully covered by the current competition design, but goal (1) is not so straightforward. Currently, the challenge of determining the forward model seems to be introduced by the fact that the data creation pipeline is obfuscated. We are essentially asked to then reverse engineer this data creation pipeline, as a proxy for the reverse engineering of the true physics the ARIEL team will have to do.\n\nUnfortunately, the approach of obfuscating the data pipeline to simulate the unknown physics of the universe is flawed. The key issue is that in reality training labels will not be available, i.e. there will not be a set of planets for which the transit depths are known. This means that all our work that uses the provided transit depths is not relevant to the final mission.\n\nI can't come up with a good way to cover goal (1) above. One option could be to not release transit depths at all, not even for the training set. This is, after all, the situation that the ARIEL team will find itself in in reality. However, in that case the competition will almost certainly come down to test set probing, which is again not something that will be possible in reality.\n\nSo the alternative I propose: accept that goal (1) is not covered, and focus fully on goal (2). Share the full synthetic data generation pipeline, to the degree that the community can create new data. Or alternatively: share the pipeline, but keep some details secret, such as the distributions you draw physical parameters from. We would then have to demonstrate the ability to learn this from data - ideally without given transit depths (but this may complicate matters too much).\n\nAnother alternative to consider (regardless of whether the full pipeline is shared): consider sharing the test set data, both public and private. Rather than submitting code, we'd submit predictions directly. This is more representative of the final mission, where you don't need a 'blind' model.\n\nNote that we had a recent competition implementing both proposals above, i.e. sharing the prediction pipeline and providing the test set to competitors: the [Yale/UNC geophysical waveform inversion competition](https://www.kaggle.com/competitions/waveform-inversion). As far as I can tell this competition was very well received by the community.\n\nTLDR: Because the real mission won't have known ground truths, this competition isn't demonstrating the ability to learn physics. Rather, our contributions are on the inference side, i.e. inverting the physics. Embrace this, and share the full data generation pipeline. Separately, also consider releasing the test set to competitors during the competition. A [recent well-received competition](https://www.kaggle.com/competitions/waveform-inversion) did both of these.",
    "3296579": "Great discussion - I will come back, I am in a conference and has been slow with my reply but i have been monitoring!!",
    "3295861": "I have the impression that a publicly-available generation pipeline would be helpful for understanding effects like the one you asked about [here](https://www.kaggle.com/competitions/ariel-data-challenge-2025/discussion/602425).\n\nI also have the sense that determining the forward model is a very open-ended problem, at least part of which might be framed as an effort to tune and validate the accuracy of the simulation.  With the caveat that I don't have a lot of experience in astronomy, I'd imagine that a nontrivial part of this would be done by making observations of star systems whose properties are well-known, including systems that are known not to have exoplanets (so that you can measure and characterize, in the definite absence of any transit signal, things like noise, cosmics, and drift, as they affect the real hardware on the actual satellite) as well as systems that do contain one or more exoplanets that have known transits which have been well-studied by other telescopes.   With or without an open generation pipeline, I could imagine future installments of this competition that include calibration/validation samples like these; contestants would then be tasked with identifying problems with the detector and/or simulation based on the provided calibration/validation samples, determining appropriate corrections, and applying them before making predictions.  \n\nI get that this probably only covers, at best, part of the question of determining the forward model.  But I would imagine that other parts of the forward model could be pinned down with other methods, e.g. demonstrate an ability to learn the limb-darkening law by generating transits using a randomly-chosen one of several different limb-darkening models and ask contestants to predict which one was used for this transit and what its parameters were.  But I would expect this sort of study of actual physics to be deferred until the detector is well-characterized/well-understood/well-calibrated.\n",
    "3295725": "If I understand correctly, the classical way to extract physics from data measured with a well-known detector would be something like:\n- build a physical model (Rp/Rs, P, sma, t0, i, e, LD, background, …),\n- fold it with the detector response (time sampling, noise, systematics),\n- and then fit to data to recover parameter posteriors with the uncertainties.\nIn this way you recover physics from first principles. However, it's often with degeneracies between parameters.\n\nIn this competition, probably the goal was to learn a mapping data -> RpRs directly, and avoid the parameters degeneracy problem. If the full data generation pipeline was shared, we could build a better model via data -> unfolded data -> RpRs split."
  }
}