{
  "id": 609186,
  "title": "Last day of the competition - Thank you all. ",
  "url": "/competitions/ariel-data-challenge-2025/discussion/609186",
  "author_name": "Gordon Yip",
  "post_date": "2025-09-24T15:39:45.139000",
  "votes": 8,
  "comment_count": 39,
  "views": 0,
  "content": "<p>This is the last day of the competition. As hosts, we just want to say a few words before we hand over the glory to this year's winners.</p>\n<p>We know that this competition has been particularly challenging. A few of you described it as \"ADC2024 on steroids\" and some might even say, \"forget everything you know about ADC2024.\"</p>\n<p>In the next few weeks, we will find out who the winners are, and like many of you, we can't wait to find out their approaches.</p>\n<p>As hosts, we have been keeping a watchful eye on the progress of the competition, and we are grateful to have you with us. I mean, honestly, without you, this competition would not be a success!</p>\n<p>This year has been nothing short of surprises – from data re-generation and faulty data in the dataset to crazy transits, questions about the meaning of the metric, and many more issues that you might have discovered during your journey through the dataset and leaderboard.</p>\n<p>We just want to say thank you. Thank you for bearing with us. </p>\n<p>The competition is not perfect, but we intend to build on your help and feedback. If you think you might have learned something quite interesting, please feel free to share and tag us! We will be writing a report to document this experience.</p>\n<p>If this competition has made you learn a bit more about astronomy, about the Ariel mission, and the data we are currently dealing with, we consider this a success.</p>\n<p>Here is what we are wondering – what stuck with YOU about this competition? <br>\nThe frustrations, the breakthroughs, the random moments of enlightenment? Did this change how you think about astronomy or the Ariel mission? Did you stumble onto techniques you'd never tried before?</p>\n<p>Feel free to use this space to share.</p>\n<p>Some of you might be contacted by us for a short interview – we would like to find out about your journey. Or DM me if you want to have a chat.</p>\n<p>Last thing - our big thank you to <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a>, <a href=\"https://www.kaggle.com/elizabethpark\" target=\"_blank\">@elizabethpark</a> and many Kaggle staff for their help in this year's competition. </p>\n<p>Last last thing: Please enjoy a little <a href=\"https://imgur.com/a/0QVWXsX\" target=\"_blank\">barchart</a> race about this year's sprint!</p>",
  "messages": [
    {
      "id": 3293771,
      "postDate": "2025-09-24T15:39:45.140Z",
      "content": "<p>This is the last day of the competition. As hosts, we just want to say a few words before we hand over the glory to this year's winners.</p>\n<p>We know that this competition has been particularly challenging. A few of you described it as \"ADC2024 on steroids\" and some might even say, \"forget everything you know about ADC2024.\"</p>\n<p>In the next few weeks, we will find out who the winners are, and like many of you, we can't wait to find out their approaches.</p>\n<p>As hosts, we have been keeping a watchful eye on the progress of the competition, and we are grateful to have you with us. I mean, honestly, without you, this competition would not be a success!</p>\n<p>This year has been nothing short of surprises – from data re-generation and faulty data in the dataset to crazy transits, questions about the meaning of the metric, and many more issues that you might have discovered during your journey through the dataset and leaderboard.</p>\n<p>We just want to say thank you. Thank you for bearing with us. </p>\n<p>The competition is not perfect, but we intend to build on your help and feedback. If you think you might have learned something quite interesting, please feel free to share and tag us! We will be writing a report to document this experience.</p>\n<p>If this competition has made you learn a bit more about astronomy, about the Ariel mission, and the data we are currently dealing with, we consider this a success.</p>\n<p>Here is what we are wondering – what stuck with YOU about this competition? <br>\nThe frustrations, the breakthroughs, the random moments of enlightenment? Did this change how you think about astronomy or the Ariel mission? Did you stumble onto techniques you'd never tried before?</p>\n<p>Feel free to use this space to share.</p>\n<p>Some of you might be contacted by us for a short interview – we would like to find out about your journey. Or DM me if you want to have a chat.</p>\n<p>Last thing - our big thank you to <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a>, <a href=\"https://www.kaggle.com/elizabethpark\" target=\"_blank\">@elizabethpark</a> and many Kaggle staff for their help in this year's competition. </p>\n<p>Last last thing: Please enjoy a little <a href=\"https://imgur.com/a/0QVWXsX\" target=\"_blank\">barchart</a> race about this year's sprint!</p>",
      "rawMarkdown": "This is the last day of the competition. As hosts, we just want to say a few words before we hand over the glory to this year's winners.\n\nWe know that this competition has been particularly challenging. A few of you described it as \"ADC2024 on steroids\" and some might even say, \"forget everything you know about ADC2024.\"\n\nIn the next few weeks, we will find out who the winners are, and like many of you, we can't wait to find out their approaches.\n\nAs hosts, we have been keeping a watchful eye on the progress of the competition, and we are grateful to have you with us. I mean, honestly, without you, this competition would not be a success!\n\nThis year has been nothing short of surprises – from data re-generation and faulty data in the dataset to crazy transits, questions about the meaning of the metric, and many more issues that you might have discovered during your journey through the dataset and leaderboard.\n\nWe just want to say thank you. Thank you for bearing with us. \n\nThe competition is not perfect, but we intend to build on your help and feedback. If you think you might have learned something quite interesting, please feel free to share and tag us! We will be writing a report to document this experience.\n\nIf this competition has made you learn a bit more about astronomy, about the Ariel mission, and the data we are currently dealing with, we consider this a success.\n\nHere is what we are wondering – what stuck with YOU about this competition? \nThe frustrations, the breakthroughs, the random moments of enlightenment? Did this change how you think about astronomy or the Ariel mission? Did you stumble onto techniques you'd never tried before?\n\nFeel free to use this space to share.\n\nSome of you might be contacted by us for a short interview – we would like to find out about your journey. Or DM me if you want to have a chat.\n\nLast thing - our big thank you to @sohier, @elizabethpark and many Kaggle staff for their help in this year's competition. \n\nLast last thing: Please enjoy a little [barchart](https://imgur.com/a/0QVWXsX) race about this year's sprint!\n\n",
      "votes": 8
    },
    {
      "id": 3293812,
      "postDate": "2025-09-24T17:35:32.247Z",
      "content": "<p>You ask for feedback, here is some. Fasten your seat belt as I'll be frank.</p>\n<p>We did not spend a second on dealing with cases where transit is not contained in the provided data. We should have done it if our motivation was the LB score. But I personally could not resolve working on something useless. I enter kaggle competitions to learn something useful. I don't enter to get points.</p>\n<p>I wish you had agreed to remove these cases when I asked, 1.5 months before end of competition. That was less than 2 weeks after data regeneration.</p>\n<p>Instead you led many to spend numerous hours dealing with something that will not be useful to ARIAL mission. Being respectful of the time participants spend should be of utter importance to competition hosts.</p>\n<p>Other than that, this was a fantastic competition. Thanks for hosting it. Can't wait for next year edition, if you do one. Topic is great. learning is great. I wish we had found what we found last week earlier in the competition. We'd love a couple of weeks extension…</p>",
      "rawMarkdown": "You ask for feedback, here is some. Fasten your seat belt as I'll be frank.\n\nWe did not spend a second on dealing with cases where transit is not contained in the provided data. We should have done it if our motivation was the LB score. But I personally could not resolve working on something useless. I enter kaggle competitions to learn something useful. I don't enter to get points.\n\nI wish you had agreed to remove these cases when I asked, 1.5 months before end of competition. That was less than 2 weeks after data regeneration.\n\nInstead you led many to spend numerous hours dealing with something that will not be useful to ARIAL mission. Being respectful of the time participants spend should be of utter importance to competition hosts.\n\nOther than that, this was a fantastic competition. Thanks for hosting it. Can't wait for next year edition, if you do one. Topic is great. learning is great. I wish we had found what we found last week earlier in the competition. We'd love a couple of weeks extension...",
      "votes": 5,
      "replies": [
        {
          "id": 3293840,
          "postDate": "2025-09-24T18:44:47.877Z",
          "content": "<p>I wonder how much special methods for the clipped transits really matter; my solution doesn't involve any at least. I hope the hosts follow through with the alternative score calculation with these transits removed.</p>",
          "rawMarkdown": "I wonder how much special methods for the clipped transits really matter; my solution doesn't involve any at least. I hope the hosts follow through with the alternative score calculation with these transits removed.",
          "replies": [
            {
              "id": 3293894,
              "postDate": "2025-09-25T00:21:49.707Z",
              "content": "<p>Few teams dropped dramatically in private. I bet it is because of some of these degenerate cases. There is no lower bound on how bad a planet score can be. </p>\n<p>All we did was to put a large sigma for these, basically turning their scores to zero. But we lose at least 4% of score that way.</p>\n<p>Anyway, I am now looking forward to your writeup. Congrats one the amazing result.</p>",
              "rawMarkdown": "Few teams dropped dramatically in private. I bet it is because of some of these degenerate cases. There is no lower bound on how bad a planet score can be. \n\nAll we did was to put a large sigma for these, basically turning their scores to zero. But we lose at least 4% of score that way.\n\nAnyway, I am now looking forward to your writeup. Congrats one the amazing result.\n",
              "votes": 1
            },
            {
              "id": 3293902,
              "postDate": "2025-09-25T01:28:34.090Z",
              "content": "<p>I am one of the victims involved. I'm very curious to see how the top kagglers will review and explain this situation.</p>",
              "rawMarkdown": "I am one of the victims involved. I'm very curious to see how the top kagglers will review and explain this situation."
            },
            {
              "id": 3293906,
              "postDate": "2025-09-25T01:49:48.647Z",
              "content": "<p>I’m one of the victims as well. I’m curious to see how you’re solution was able to compensate for these situations without clipping. For us, our current methodology would produce planets that could have as high as 100000 ppm when we investigated it right after the competition. We didn’t anticipate that this would have such a huge impact as our public lb was quite good before the competition closed, and thus we didn’t do any special measurements against these situations. </p>",
              "rawMarkdown": "I’m one of the victims as well. I’m curious to see how you’re solution was able to compensate for these situations without clipping. For us, our current methodology would produce planets that could have as high as 100000 ppm when we investigated it right after the competition. We didn’t anticipate that this would have such a huge impact as our public lb was quite good before the competition closed, and thus we didn’t do any special measurements against these situations. "
            },
            {
              "id": 3294011,
              "postDate": "2025-09-25T07:34:57.093Z",
              "content": "<blockquote>\n  <p>All we did was to put a large sigma for these, basically turning their scores to zero. But we lose at least 4% of score that way.</p>\n</blockquote>\n<p>That is what I did as well. Apparently, one or more slipped my hard sample detection in the private set and dropped me to the bottom of the private leaderboard. <br>\nI think the metric should be updated for next year (I hope we get to see another round of ARIEL) and shouldn't be impacted by single bad samples too much. Can always happen that a fit badly diverges for single samples and in reality, this would be catched easily by supervising the system or looking at the results.</p>",
              "rawMarkdown": "> All we did was to put a large sigma for these, basically turning their scores to zero. But we lose at least 4% of score that way.\n\nThat is what I did as well. Apparently, one or more slipped my hard sample detection in the private set and dropped me to the bottom of the private leaderboard. \nI think the metric should be updated for next year (I hope we get to see another round of ARIEL) and shouldn't be impacted by single bad samples too much. Can always happen that a fit badly diverges for single samples and in reality, this would be catched easily by supervising the system or looking at the results."
            },
            {
              "id": 3294079,
              "postDate": "2025-09-25T10:15:32.073Z",
              "content": "<p>I had quite robust fallbacks in my approach as well, but it happened to be not so robust enough to prevent dropping to zero score. Feels quite unfair given the invested time…</p>",
              "rawMarkdown": "I had quite robust fallbacks in my approach as well, but it happened to be not so robust enough to prevent dropping to zero score. Feels quite unfair given the invested time...",
              "votes": 3
            },
            {
              "id": 3294505,
              "postDate": "2025-09-26T09:40:09.210Z",
              "content": "<p>yeah we hear you - i think bad examples do occur and they have their values, but the metric should be improved so that it wouldn't drag down the score drastically due to a few bad examples</p>",
              "rawMarkdown": "yeah we hear you - i think bad examples do occur and they have their values, but the metric should be improved so that it wouldn't drag down the score drastically due to a few bad examples",
              "votes": 1
            },
            {
              "id": 3296243,
              "postDate": "2025-09-30T13:52:17.023Z",
              "content": "<p>I have extensively tested the private LB data, and I'm now quite confident that my score is pulled down by specific LD configurations used to simulate the private data.</p>\n<p>In my fitting framework I have very strict quality controls. I control not just the quality of the fits but also the quality of the resulting physical model. I reparametrized many aspects of the model to reduce parameter degeneracies as much as possible. I also use a dynamic choice between uniform/linear/quadratic(Kipping) LDs based on NLL/AIC. If fit quality criteria falls below threshold and/or the covariance shows strong RpRs-b correlations, etc., I fallback to my \"physics unaware\" model from last year, which scors 0.428/0.428 this year.</p>\n<p>Just allowing my model to accept \"unphysical\" limb brightening magically fixed the zero score issue on the private LB:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6327791%2F75f301ab2842da3149a776ac71e8b536%2FADC.png?generation=1759238082134729&amp;alt=media\" alt=\"\"></p>\n<p>This doesn't necessarily mean that the private LB set contains examples with Limb Brightening ( <a href=\"https://www.kaggle.com/gordonyip\" target=\"_blank\">@gordonyip</a> -- could you confirm this?). It's possible that the data was simulated with more exotic LD profiles, and that allowing limb brightening in your model simply absorbs the bias that would otherwise collapse the score to zero.</p>\n<p><strong>To conclude, this highlights that GLL may not be the perfect evaluation metric for this competition</strong>. It tends to penalize very accurate physics-constrained models in favor of simpler \"physics-unaware\" models that are less sensitive to parameter degeneracies.</p>\n<p><a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> -- probably you also constrained aggressively your LD model? :)</p>",
              "rawMarkdown": "I have extensively tested the private LB data, and I'm now quite confident that my score is pulled down by specific LD configurations used to simulate the private data.\n\nIn my fitting framework I have very strict quality controls. I control not just the quality of the fits but also the quality of the resulting physical model. I reparametrized many aspects of the model to reduce parameter degeneracies as much as possible. I also use a dynamic choice between uniform/linear/quadratic(Kipping) LDs based on NLL/AIC. If fit quality criteria falls below threshold and/or the covariance shows strong RpRs-b correlations, etc., I fallback to my \"physics unaware\" model from last year, which scors 0.428/0.428 this year.\n\nJust allowing my model to accept \"unphysical\" limb brightening magically fixed the zero score issue on the private LB:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6327791%2F75f301ab2842da3149a776ac71e8b536%2FADC.png?generation=1759238082134729&alt=media)\n\nThis doesn't necessarily mean that the private LB set contains examples with Limb Brightening ( @gordonyip -- could you confirm this?). It's possible that the data was simulated with more exotic LD profiles, and that allowing limb brightening in your model simply absorbs the bias that would otherwise collapse the score to zero.\n\n**To conclude, this highlights that GLL may not be the perfect evaluation metric for this competition**. It tends to penalize very accurate physics-constrained models in favor of simpler \"physics-unaware\" models that are less sensitive to parameter degeneracies.\n\n@ilu000 -- probably you also constrained aggressively your LD model? :)\n",
              "votes": 1
            },
            {
              "id": 3296599,
              "postDate": "2025-10-01T08:42:54.883Z",
              "content": "<p>I also used constraints similar to this to keep to check the sanity of the modeling results. However, the approach I took is keeping the model free, and raising an error if the constraints are violated. This then allowed me to make deliberate choices on how to deal with these planets (there were 2 of them); in the end, I didn't do anything in my final submission, but I had a backup submission where I replaced them by dummy submission (mean and std of the other planets).</p>\n<p>I can't check right now, but I don't think I had such checks on the limb darkening parameters by the way.</p>",
              "rawMarkdown": "I also used constraints similar to this to keep to check the sanity of the modeling results. However, the approach I took is keeping the model free, and raising an error if the constraints are violated. This then allowed me to make deliberate choices on how to deal with these planets (there were 2 of them); in the end, I didn't do anything in my final submission, but I had a backup submission where I replaced them by dummy submission (mean and std of the other planets).\n\nI can't check right now, but I don't think I had such checks on the limb darkening parameters by the way.",
              "votes": 1
            },
            {
              "id": 3296602,
              "postDate": "2025-10-01T08:51:03.583Z",
              "content": "<p>Actually, I now realize I'm confused by your post above. Your model falls back to a simpler, more robust model if the limb darkening parameters are not plausible (i.e. similar to my approach explained above). But how did you end up with a score of 0.000 due to limb brightening then? Why didn't you model fall back to the robust model on those planets?</p>",
              "rawMarkdown": "Actually, I now realize I'm confused by your post above. Your model falls back to a simpler, more robust model if the limb darkening parameters are not plausible (i.e. similar to my approach explained above). But how did you end up with a score of 0.000 due to limb brightening then? Why didn't you model fall back to the robust model on those planets?"
            },
            {
              "id": 3296627,
              "postDate": "2025-10-01T09:43:06.380Z",
              "content": "<p>There are some model configurations used to simulate data that lead to degeneracy of parameters. So from the transit data  you cannot distinguish parameters apart without additional priors/constraints. They trade off against each other when you try to maximize likelihood. From my observation some planets had covariances leading to large correlations. For example, RpRs &lt;-&gt; impact parameter b &lt;-&gt; LD coeffs can impact each other so that data does not have constraining power alone.</p>\n<p>In my case, when one of the LD coefficients is stuck at 'physical' boundary - the resulting model still can have very high quality in description of transit data (very good NLL or chi2/ndof). This leads to RpRs bias, whereas the <code>b</code> parameter compensated the rest of data/model difference that the LD had problem to catch (even when b is constrained from stellar parameters with gaussian prior). </p>\n<p>So in this case we have:</p>\n<ul>\n<li>very good model description of transit data (confidence), so no fallback</li>\n<li>no extreme correlations from covariance matrix, no fallback</li>\n<li>biased RpRs with respect to RpRs(true), biased 'b' with respect to b(true) -&gt; huge penalty to GLL</li>\n<li>LD(quadratic/Kipping) did not approximate the host's custom LD(true) because of not enough freedom (or because few planets with Limb Brightening were in private data)</li>\n</ul>",
              "rawMarkdown": "There are some model configurations used to simulate data that lead to degeneracy of parameters. So from the transit data  you cannot distinguish parameters apart without additional priors/constraints. They trade off against each other when you try to maximize likelihood. From my observation some planets had covariances leading to large correlations. For example, RpRs <-> impact parameter b <-> LD coeffs can impact each other so that data does not have constraining power alone.\n\nIn my case, when one of the LD coefficients is stuck at 'physical' boundary - the resulting model still can have very high quality in description of transit data (very good NLL or chi2/ndof). This leads to RpRs bias, whereas the `b` parameter compensated the rest of data/model difference that the LD had problem to catch (even when b is constrained from stellar parameters with gaussian prior). \n\nSo in this case we have:\n - very good model description of transit data (confidence), so no fallback\n - no extreme correlations from covariance matrix, no fallback\n - biased RpRs with respect to RpRs(true), biased 'b' with respect to b(true) -> huge penalty to GLL\n - LD(quadratic/Kipping) did not approximate the host's custom LD(true) because of not enough freedom (or because few planets with Limb Brightening were in private data)"
            },
            {
              "id": 3296634,
              "postDate": "2025-10-01T10:23:56.760Z",
              "content": "<p>In our solution, we had to limit the accuracy of curve fit because it led to degraded LB while CV was improved significantly. If I was not on vacation then I'd rerun the way Jeroen did: relax constraints and trigger error if they are violated, with a fallback.</p>\n<p>We took a conservative approach and put stricter constraints than nececessa. As a result we did not drop, but didn't get a greta score either.</p>\n<p>I find it concerning if the ground truth contains physically impossible answers. I hope <a href=\"https://www.kaggle.com/gordonyip\" target=\"_blank\">@gordonyip</a> will address that shortly.</p>",
              "rawMarkdown": "In our solution, we had to limit the accuracy of curve fit because it led to degraded LB while CV was improved significantly. If I was not on vacation then I'd rerun the way Jeroen did: relax constraints and trigger error if they are violated, with a fallback.\n\nWe took a conservative approach and put stricter constraints than nececessa. As a result we did not drop, but didn't get a greta score either.\n\nI find it concerning if the ground truth contains physically impossible answers. I hope @gordonyip will address that shortly.",
              "votes": 1
            },
            {
              "id": 3296852,
              "postDate": "2025-10-01T18:56:31.543Z",
              "content": "<p>In my case there were 2 fallbacks in the private test set, but I'm not sure if they're physically unrealistic. They were caused by:</p>\n<ul>\n<li>An extremely high mean value for the transit depth (20%).</li>\n<li>An unexpected high residual of the AIRS signal (around 1.15x the noise level if I recall correctly), likely indicating the fit landed in a bad local minimum.</li>\n</ul>",
              "rawMarkdown": "In my case there were 2 fallbacks in the private test set, but I'm not sure if they're physically unrealistic. They were caused by:\n\n- An extremely high mean value for the transit depth (20%).\n- An unexpected high residual of the AIRS signal (around 1.15x the noise level if I recall correctly), likely indicating the fit landed in a bad local minimum."
            },
            {
              "id": 3297055,
              "postDate": "2025-10-02T08:34:23.743Z",
              "content": "<p>Just for completeness, I resubmitted my code with:</p>\n<ul>\n<li>Relaxed Kipping's q2 limits to [-1, 1]</li>\n<li>If q2&lt;0, I fall back to my base model (0.428/0.428)</li>\n</ul>\n<p><a href=\"https://arxiv.org/pdf/1308.0009\" target=\"_blank\">Kipping</a> parametrized the quadratic LB in terms of q1, q2 as <code>u1 = 2.0 * np.sqrt(q1) * q2; u2 = np.sqrt(q1) * (1.0 - 2.0 * q2)</code>. This parametrization reduces correlations between LD coefficients during regression, which 'must be' very useful in this competition if you are not fixing them from LD tables.  <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6327791%2Fb78de93fdbb1a5110c14b21001d69609%2Ffallback.png?generation=1759393222303450&amp;alt=media\" alt=\"\"><br>\nFrom this test, the fallback happens probably once or twice on the public LB data, while it's more often on the private LB. As I have already mentioned before, this does not necessary mean that there are planets with truly unphysically LDs. It may just contain more exotic LD profiles, where my setup prefers negative q2 to maximize likelihood.</p>",
              "rawMarkdown": "Just for completeness, I resubmitted my code with:\n - Relaxed Kipping's q2 limits to [-1, 1]\n - If q2<0, I fall back to my base model (0.428/0.428)\n\n[Kipping](https://arxiv.org/pdf/1308.0009) parametrized the quadratic LB in terms of q1, q2 as `u1 = 2.0 * np.sqrt(q1) * q2; u2 = np.sqrt(q1) * (1.0 - 2.0 * q2)`. This parametrization reduces correlations between LD coefficients during regression, which 'must be' very useful in this competition if you are not fixing them from LD tables.  \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6327791%2Fb78de93fdbb1a5110c14b21001d69609%2Ffallback.png?generation=1759393222303450&alt=media)\nFrom this test, the fallback happens probably once or twice on the public LB data, while it's more often on the private LB. As I have already mentioned before, this does not necessary mean that there are planets with truly unphysically LDs. It may just contain more exotic LD profiles, where my setup prefers negative q2 to maximize likelihood."
            }
          ]
        },
        {
          "id": 3293901,
          "postDate": "2025-09-25T00:40:03.177Z",
          "content": "<p>Indeed so. After using the phase detector to detect the abnormal planet, I couldn't think of any additional solutions.</p>",
          "rawMarkdown": "Indeed so. After using the phase detector to detect the abnormal planet, I couldn't think of any additional solutions.",
          "votes": 1
        },
        {
          "id": 3294503,
          "postDate": "2025-09-26T09:35:23.463Z",
          "content": "<p>Thanks for being honest and upfront. We do need this to make better competition. </p>\n<p>As for whether to remove the 'bad' transit or not. A few things were going through our heads at the time - <br>\nbad transits could occur due to human/measurement error (Ariel is no exception), and we have seen many datasets discarded because of bad planning, but some also managed to produce some scientific values (but I am sure you can imagine the level of engineering they have to go through… ). As hosts we wanted to know how participants could mitigate that, given these examples, or if they can. </p>\n<p>the other thing is on the organisation level,  Our consideration at the time was how much we can <code>rock</code> the boat before everyone starts leaving the ship. To be honest it is the first time we have to re-generate the dataset and we are worried if a second change might rock the boat too much</p>\n<p>so altogether we decided not to. </p>\n<p>In hindsight though- we didnt realise they are having a big impact on the metric, something we should have foreseen, and the care one must take for these bad examples may be too big of an investment for the competition. </p>\n<p>Here is my question to you and fellow Kagglers - how do you feel about data regeneration, are you happy to see them regenerated , how many times could you tolerate it and what would have been your preferred re-action from the host , learning this will help our decision making next time.</p>",
          "rawMarkdown": "Thanks for being honest and upfront. We do need this to make better competition. \n\nAs for whether to remove the 'bad' transit or not. A few things were going through our heads at the time - \nbad transits could occur due to human/measurement error (Ariel is no exception), and we have seen many datasets discarded because of bad planning, but some also managed to produce some scientific values (but I am sure you can imagine the level of engineering they have to go through... ). As hosts we wanted to know how participants could mitigate that, given these examples, or if they can. \n\nthe other thing is on the organisation level,  Our consideration at the time was how much we can `rock` the boat before everyone starts leaving the ship. To be honest it is the first time we have to re-generate the dataset and we are worried if a second change might rock the boat too much\n\nso altogether we decided not to. \n\nIn hindsight though- we didnt realise they are having a big impact on the metric, something we should have foreseen, and the care one must take for these bad examples may be too big of an investment for the competition. \n\nHere is my question to you and fellow Kagglers - how do you feel about data regeneration, are you happy to see them regenerated , how many times could you tolerate it and what would have been your preferred re-action from the host , learning this will help our decision making next time.\n",
          "votes": 1,
          "replies": [
            {
              "id": 3294515,
              "postDate": "2025-09-26T09:55:15.860Z",
              "content": "<p>I joined around the time the regeneration happened so it didn't affect me much.<br>\nIf there was another regeneration to rectify the broken transits I would welcome it.</p>\n<p>It comes down to whether our attempts to mitigate these bad transits provide scientific value.<br>\nFrom your previous statements it seems like our efforts would be better spent elsewhere.</p>",
              "rawMarkdown": "I joined around the time the regeneration happened so it didn't affect me much.\nIf there was another regeneration to rectify the broken transits I would welcome it.\n\nIt comes down to whether our attempts to mitigate these bad transits provide scientific value.\nFrom your previous statements it seems like our efforts would be better spent elsewhere."
            },
            {
              "id": 3294537,
              "postDate": "2025-09-26T10:09:21.093Z",
              "content": "<p>This competition as well as the last competition was very well done in my opinion. It is quite a unique task and I totally understand the difficulties around synthetic data.</p>\n<blockquote>\n  <p>Here is my question to you and fellow Kagglers - how do you feel about data regeneration, are you happy to see them regenerated , how many times could you tolerate it and what would have been your preferred re-action from the host , learning this will help our decision making next time.</p>\n</blockquote>\n<p>Early on in a competition, I think data re-generation is not a big issue at all, even multiple times. As long as it helps the competition to be more robust, more predictable, more scientifically correct, or to prevent leakage. Ideally, the solutions wouldn't really be impacted much, but that isn't always possible. Of course, open communication is a must here, and you did very well on that. </p>\n<p>I personally would have welcomed another regeneration quickly after the first one when it became clear that the transit length was too long for some samples. Maybe, even excluding samples from the metric would have already solved this and is only a minor change. </p>\n<p>I think the main issue is rather the unbound metric on a per-sample basis. If a single wrong sample can drive the metric to zero, things can get complicated quickly, if the metric wouldn't allow this, I guess we wouln't even have the discussion as these handful of samples would only impact the score on the third digit.</p>\n<p>That said, I am looking forward to a new round of ARIEL. </p>",
              "rawMarkdown": "This competition as well as the last competition was very well done in my opinion. It is quite a unique task and I totally understand the difficulties around synthetic data.\n\n> Here is my question to you and fellow Kagglers - how do you feel about data regeneration, are you happy to see them regenerated , how many times could you tolerate it and what would have been your preferred re-action from the host , learning this will help our decision making next time.\n\nEarly on in a competition, I think data re-generation is not a big issue at all, even multiple times. As long as it helps the competition to be more robust, more predictable, more scientifically correct, or to prevent leakage. Ideally, the solutions wouldn't really be impacted much, but that isn't always possible. Of course, open communication is a must here, and you did very well on that. \n\nI personally would have welcomed another regeneration quickly after the first one when it became clear that the transit length was too long for some samples. Maybe, even excluding samples from the metric would have already solved this and is only a minor change. \n\nI think the main issue is rather the unbound metric on a per-sample basis. If a single wrong sample can drive the metric to zero, things can get complicated quickly, if the metric wouldn't allow this, I guess we wouln't even have the discussion as these handful of samples would only impact the score on the third digit.\n\nThat said, I am looking forward to a new round of ARIEL. "
            },
            {
              "id": 3294859,
              "postDate": "2025-09-27T00:03:11.683Z",
              "content": "<p>Data regeneration was not a real issue, mostly because we cudl reuse the same pipeline as before.</p>\n<p>But not fixing the few glitches was a real mistake. it would not have rocked the boat at all. </p>\n<blockquote>\n  <p>we didnt realise they are having a big impact on the metric,</p>\n</blockquote>\n<p>I warned you about the consequence on the metric when i asked for removal. But I can't find my post. It seems you deleted the topic you had for questions from us. Why is that?</p>",
              "rawMarkdown": "Data regeneration was not a real issue, mostly because we cudl reuse the same pipeline as before.\n\nBut not fixing the few glitches was a real mistake. it would not have rocked the boat at all. \n\n> we didnt realise they are having a big impact on the metric,\n\nI warned you about the consequence on the metric when i asked for removal. But I can't find my post. It seems you deleted the topic you had for questions from us. Why is that?\n"
            },
            {
              "id": 3294909,
              "postDate": "2025-09-27T04:59:47.030Z",
              "content": "<p><a href=\"https://www.kaggle.com/competitions/ariel-data-challenge-2025/discussion/589905#3269906\" target=\"_blank\">https://www.kaggle.com/competitions/ariel-data-challenge-2025/discussion/589905#3269906</a></p>",
              "rawMarkdown": "https://www.kaggle.com/competitions/ariel-data-challenge-2025/discussion/589905#3269906",
              "votes": 1
            },
            {
              "id": 3294951,
              "postDate": "2025-09-27T07:44:04.737Z",
              "content": "<p>Thank <a href=\"https://www.kaggle.com/jeroencottaar\" target=\"_blank\">@jeroencottaar</a> , it was just unpinned. <a href=\"https://www.kaggle.com/gordonyip\" target=\"_blank\">@gordonyip</a> here is where I warned you of the consequences of bad data on the metric <a href=\"https://www.kaggle.com/competitions/ariel-data-challenge-2025/discussion/598324#3269605\" target=\"_blank\">https://www.kaggle.com/competitions/ariel-data-challenge-2025/discussion/598324#3269605</a></p>",
              "rawMarkdown": "Thank @jeroencottaar , it was just unpinned. @gordonyip here is where I warned you of the consequences of bad data on the metric https://www.kaggle.com/competitions/ariel-data-challenge-2025/discussion/598324#3269605\n\n"
            }
          ]
        }
      ]
    },
    {
      "id": 3294490,
      "postDate": "2025-09-26T09:07:12.730Z",
      "content": "<p>wow the discussion exploded, it will take me some time to go through all of them , but thank you for all the responses and discussion!</p>",
      "rawMarkdown": "wow the discussion exploded, it will take me some time to go through all of them , but thank you for all the responses and discussion!",
      "votes": 1
    },
    {
      "id": 3293948,
      "postDate": "2025-09-25T03:55:17.670Z",
      "content": "<p>Thanks for organizing a good competition!</p>\n<p>Referencing <a href=\"https://www.kaggle.com/competitions/ariel-data-challenge-2024/writeups/c-number-daiwakun-1st-place-solution\" target=\"_blank\">last year's first place solution</a>, the gain drift f(t) has a significant impact on the result. In order to accurately estimate f(t), I would say that ideally the transit should only take up about one third of the data. Since in reality the extra data is collected anyway, as CPMP pointed out, I think it would be better to provide that as well.</p>\n<p>On the other hand, I found the calibration and binning to be relatively uninteresting. I would rather that preprocessing had already been performed by the organizers, similar to what was done in <a href=\"https://www.kaggle.com/code/gordonyip/calibrating-and-binning-ariel-data\" target=\"_blank\">the public notebook by Gordon Yip</a>. With the same amount of data, the number of unique planet ids could then be significantly increased, and more data could be provided outside the transit window. Maybe the number of planets has been limited to match the data that will be collected in the future? I would argue that that is a mistake. If the simulations are accurate enough, I think it would be better to train the final model on 50,000+ simulated planets. If you want a Kaggle competition to reflect that, then you should do calibration and binning yourself.</p>",
      "rawMarkdown": "Thanks for organizing a good competition!\n\nReferencing [last year's first place solution](https://www.kaggle.com/competitions/ariel-data-challenge-2024/writeups/c-number-daiwakun-1st-place-solution), the gain drift f(t) has a significant impact on the result. In order to accurately estimate f(t), I would say that ideally the transit should only take up about one third of the data. Since in reality the extra data is collected anyway, as CPMP pointed out, I think it would be better to provide that as well.\n\nOn the other hand, I found the calibration and binning to be relatively uninteresting. I would rather that preprocessing had already been performed by the organizers, similar to what was done in [the public notebook by Gordon Yip](https://www.kaggle.com/code/gordonyip/calibrating-and-binning-ariel-data). With the same amount of data, the number of unique planet ids could then be significantly increased, and more data could be provided outside the transit window. Maybe the number of planets has been limited to match the data that will be collected in the future? I would argue that that is a mistake. If the simulations are accurate enough, I think it would be better to train the final model on 50,000+ simulated planets. If you want a Kaggle competition to reflect that, then you should do calibration and binning yourself.",
      "votes": 1,
      "replies": [
        {
          "id": 3294068,
          "postDate": "2025-09-25T09:40:31.840Z",
          "content": "<blockquote>\n  <p>the transit should only take up about one third of the data</p>\n</blockquote>\n<p>This is part of Ariel specs actually. I don't have the reference handy, but it is stated clearly. It is why I don't get we got data that did not comply with it.</p>",
          "rawMarkdown": "> the transit should only take up about one third of the data\n\nThis is part of Ariel specs actually. I don't have the reference handy, but it is stated clearly. It is why I don't get we got data that did not comply with it.",
          "votes": 1
        }
      ]
    },
    {
      "id": 3293841,
      "postDate": "2025-09-24T18:48:42.067Z",
      "content": "<p>It was another great competition!</p>\n<p>One aspect you might want to consider is being more open about the generation of the synthetic data. While the overall flow is published, there are many details that we have to figure out ourselves. But many of these, such as using a polynomial for the gain drift, are not actually physically realistic, so the value of our work there is limited. And any results we find will anyway be limited by the accuracy of your assumptions in generating the data. </p>\n<p>I wonder what this competition might have looked like if the full synthetic data generation pipeline (to the point that we can reproduce it) was released, along with a qualitative description of the changes made for the test set. </p>",
      "rawMarkdown": "It was another great competition!\n\nOne aspect you might want to consider is being more open about the generation of the synthetic data. While the overall flow is published, there are many details that we have to figure out ourselves. But many of these, such as using a polynomial for the gain drift, are not actually physically realistic, so the value of our work there is limited. And any results we find will anyway be limited by the accuracy of your assumptions in generating the data. \n\nI wonder what this competition might have looked like if the full synthetic data generation pipeline (to the point that we can reproduce it) was released, along with a qualitative description of the changes made for the test set. ",
      "votes": 2,
      "replies": [
        {
          "id": 3293843,
          "postDate": "2025-09-24T18:54:47.583Z",
          "content": "<p>Oh and a smaller one, also mentioned in a recent discussion: you could consider clipping the score per planet (perhaps only at 0, not at 1 - why do you clip at 1 anyway?). This means you don't get people scoring 0 just because a single planet went wildly wrong (which you'd likely catch easily when you actually see the data in the real mission). There is an art to making the model robust against that - but with only two tries on the private test set, it's not really fair.</p>",
          "rawMarkdown": "Oh and a smaller one, also mentioned in a recent discussion: you could consider clipping the score per planet (perhaps only at 0, not at 1 - why do you clip at 1 anyway?). This means you don't get people scoring 0 just because a single planet went wildly wrong (which you'd likely catch easily when you actually see the data in the real mission). There is an art to making the model robust against that - but with only two tries on the private test set, it's not really fair.",
          "replies": [
            {
              "id": 3293979,
              "postDate": "2025-09-25T05:49:58.997Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 3294543,
              "postDate": "2025-09-26T10:16:16.763Z",
              "content": "<p>yes clipping the score was one option we were considering. By not clipping we wanted to encourage more conservative solutions, so model that will avoid predicting the transits depth very badly and away from the GT. I think in the real mission - it is still relatively easy to catch by running the conventional pipeline, but as thing scale up, it might become less easy to spot those cases as one would really fit them to get the depth out in the classical approach.</p>\n<p>But driving the whole score to 0 is very bad , sorry!</p>",
              "rawMarkdown": "yes clipping the score was one option we were considering. By not clipping we wanted to encourage more conservative solutions, so model that will avoid predicting the transits depth very badly and away from the GT. I think in the real mission - it is still relatively easy to catch by running the conventional pipeline, but as thing scale up, it might become less easy to spot those cases as one would really fit them to get the depth out in the classical approach.\n\nBut driving the whole score to 0 is very bad , sorry!",
              "votes": 1
            }
          ]
        },
        {
          "id": 3294064,
          "postDate": "2025-09-25T09:38:44.797Z",
          "content": "<p>I disagree on disclosing the data generator.</p>\n<p>On the contrary, I think it should be made as complex as possible to prevent reverse engineering. That's the only way to assess model generalization IMHO.</p>",
          "rawMarkdown": "I disagree on disclosing the data generator.\n\nOn the contrary, I think it should be made as complex as possible to prevent reverse engineering. That's the only way to assess model generalization IMHO."
        },
        {
          "id": 3294527,
          "postDate": "2025-09-26T10:05:11.537Z",
          "content": "<p>Yes we are being very secretive about this. While disclosing the whole thing might make it clearer for each step, it also allows room for reverse engineering. Before the competition we had a long discussion with Kaggle to make sure the data generation pipeline is opaque. Reverse engineering the solution have happened before and does not provide much value to us and does not make it an enjoyable experience for those who tried to come up with alternative solution. </p>\n<p>However, the competition is quite a domain heavy one so we disclose some guidelines here and there, but locked up the full details in a paper. </p>",
          "rawMarkdown": "Yes we are being very secretive about this. While disclosing the whole thing might make it clearer for each step, it also allows room for reverse engineering. Before the competition we had a long discussion with Kaggle to make sure the data generation pipeline is opaque. Reverse engineering the solution have happened before and does not provide much value to us and does not make it an enjoyable experience for those who tried to come up with alternative solution. \n\nHowever, the competition is quite a domain heavy one so we disclose some guidelines here and there, but locked up the full details in a paper. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 3294447,
      "postDate": "2025-09-26T07:32:22.170Z",
      "content": "<p>Some feedback for next time:</p>\n<ol>\n<li>Pre-calibrate the data. I haven't seen any novel calibration methods that improve over the standard so far.</li>\n<li>Remove clipped/malformed transits and other artifacts that are not part of the mission.</li>\n<li>The metric should be more robust and less sensitive to outliers, though I don't have any specific suggestions off the top of my head.</li>\n<li>The data generation pipeline should be obfuscated as much as possible, reverse engineering your pipeline provides little scientific value.</li>\n</ol>\n<p>Thanks for the competition! First time working on astronomy and learned a lot.</p>",
      "rawMarkdown": "Some feedback for next time:\n1. Pre-calibrate the data. I haven't seen any novel calibration methods that improve over the standard so far.\n2. Remove clipped/malformed transits and other artifacts that are not part of the mission.\n3. The metric should be more robust and less sensitive to outliers, though I don't have any specific suggestions off the top of my head.\n4. The data generation pipeline should be obfuscated as much as possible, reverse engineering your pipeline provides little scientific value.\n\nThanks for the competition! First time working on astronomy and learned a lot."
    },
    {
      "id": 3294061,
      "postDate": "2025-09-25T09:35:19.373Z",
      "content": "<p>Another feedback. I gave it during the comp, but maybe it has been lost. Repeating it here.</p>\n<p>You ask for sigmas of our predictions. This makes lot of sense.</p>\n<p>But you did not provide the sigmas for the info in the star_info file. This makes no sense at all. And it looks unfair: you ask us for something you don't provide.</p>\n<p>You know about the noise in that data. We can see that noise by applying Kepler 2nd law for instance. We also see it (and that was discussed in the forum) in the discrepancy bewtween the measure transit time and the predicted transit time.</p>\n<p>Next time, if any, please provide std for every physical value you provide.</p>",
      "rawMarkdown": "Another feedback. I gave it during the comp, but maybe it has been lost. Repeating it here.\n\nYou ask for sigmas of our predictions. This makes lot of sense.\n\nBut you did not provide the sigmas for the info in the star_info file. This makes no sense at all. And it looks unfair: you ask us for something you don't provide.\n\nYou know about the noise in that data. We can see that noise by applying Kepler 2nd law for instance. We also see it (and that was discussed in the forum) in the discrepancy bewtween the measure transit time and the predicted transit time.\n\nNext time, if any, please provide std for every physical value you provide.\n\n"
    },
    {
      "id": 3293899,
      "postDate": "2025-09-25T00:38:10.613Z",
      "content": "<p>Thank you very much for your organization! This was an absolutely wonderful game. Especially, I was deeply impressed by your meticulousness in dealing with physics!</p>\n<p>At the same time, I am very much looking forward to seeing the rankings after the abnormal samples have been removed.</p>",
      "rawMarkdown": "Thank you very much for your organization! This was an absolutely wonderful game. Especially, I was deeply impressed by your meticulousness in dealing with physics!\n\nAt the same time, I am very much looking forward to seeing the rankings after the abnormal samples have been removed."
    },
    {
      "id": 3293816,
      "postDate": "2025-09-24T17:43:42.933Z",
      "content": "<p>Right back at you -- thanks for organizing this competition; it's been a fun one!  And for this thread, which promises an interesting discussion.  Here are some thoughts that come to my mind after reading your post:</p>\n<ul>\n<li>Don't feel bad about having regenerated the dataset after the competition.  I've seen organizers make the mistake of not fixing a problem that surfaced after the start of the competition, and in my opinion that screws things up far more than the need to re-process or re-tune an analysis.  In scientific work especially (as I'm sure you know), it's more important to adopt processes and mindsets that will guide you to the right answer in the end than it is to have everything right from the get-go, so kudos for bringing that practice to this competition.</li>\n<li>the approach I took for my solution is a chisquare fit for each wavelength, whose outputs feed into a neural net that does smoothing and corrects for things like limb darkening and internal fit bias.  The fits, at least, are meant to look a lot like the way we'd do data analysis back when I was a grad student in collider physics, which is to say that, to me, my solution feels (deliberately) very conventional.  But with that chisquare fit approach, I found the runtime constraint to be <em>brutal</em>.  It was only a few days ago that I managed to get things to a state where I was confident my submission would run to completion in the allowed time, and as I write this, I still don't know if the submission I have running will produce a valid score or not. (The dreaded \"submission scoring error\" got my first entry.)  I don't mean to sound like I'm complaining; the time constraint did force me to really do some creative thinking and challenge myself on what was really necessary, so I do consider it a net positive.  But please don't make it any tighter next time; you probably don't want have a situation where potentially useful approaches have to get tossed out simply because they're not fast enough.  😅</li>\n<li>In a fully conventional/traditional chisquare-fitting approach, one of the things I'd expect to have to do is binning over wavelengths: sum to get a light curve for that bin, then fit the light curve, and get an estimate of the depth parameter along with its fit error.  One of the nice things about that approach is that, so long as you don't choose overlapping bins, the error bars that you get for one bin are statistically independent from the next, in the sense that a random fluctuation in one wavelength only affects the one bin it's been assigned to.  That is a real benefit for whoever is doing downstream analysis on your results, because then they can take your fit depth measurements and the associated uncertainties and use them in a hypothesis test or whatever without having to worry about how to model correlations between the error bars of different bins.  I mention all that because the scoring metric for this competition forces participants to make a prediction for every wavelength, and doesn't enforce any statistical independence across wavelengths, so some solutions (like mine) will make it hard to interpret the spectrum (and in particular its error bars) in a downstream analysis.  Having an unusual optimization target like the one in this competition can really spur a lot of creativity, so again, not complaining.  But one of the things that might be an interesting element of a future competition would be how to interface with downstream analysis targets.  That could mean choosing a clever metric that somehow encourages/enforces statistical independence across wavelength or rewards smart binning choices, or it could mean pulling downstream analysis targets into the competition, e.g. ask participants to measure the concentration of a list of chemicals that might be present in the exoplanet atmospheres, or it could mean some other fun twist.</li>\n</ul>\n<p>Thanks again -- looking forward to next time!</p>",
      "rawMarkdown": "Right back at you -- thanks for organizing this competition; it's been a fun one!  And for this thread, which promises an interesting discussion.  Here are some thoughts that come to my mind after reading your post:\n * Don't feel bad about having regenerated the dataset after the competition.  I've seen organizers make the mistake of not fixing a problem that surfaced after the start of the competition, and in my opinion that screws things up far more than the need to re-process or re-tune an analysis.  In scientific work especially (as I'm sure you know), it's more important to adopt processes and mindsets that will guide you to the right answer in the end than it is to have everything right from the get-go, so kudos for bringing that practice to this competition.\n * the approach I took for my solution is a chisquare fit for each wavelength, whose outputs feed into a neural net that does smoothing and corrects for things like limb darkening and internal fit bias.  The fits, at least, are meant to look a lot like the way we'd do data analysis back when I was a grad student in collider physics, which is to say that, to me, my solution feels (deliberately) very conventional.  But with that chisquare fit approach, I found the runtime constraint to be *brutal*.  It was only a few days ago that I managed to get things to a state where I was confident my submission would run to completion in the allowed time, and as I write this, I still don't know if the submission I have running will produce a valid score or not. (The dreaded \"submission scoring error\" got my first entry.)  I don't mean to sound like I'm complaining; the time constraint did force me to really do some creative thinking and challenge myself on what was really necessary, so I do consider it a net positive.  But please don't make it any tighter next time; you probably don't want have a situation where potentially useful approaches have to get tossed out simply because they're not fast enough.  😅\n * In a fully conventional/traditional chisquare-fitting approach, one of the things I'd expect to have to do is binning over wavelengths: sum to get a light curve for that bin, then fit the light curve, and get an estimate of the depth parameter along with its fit error.  One of the nice things about that approach is that, so long as you don't choose overlapping bins, the error bars that you get for one bin are statistically independent from the next, in the sense that a random fluctuation in one wavelength only affects the one bin it's been assigned to.  That is a real benefit for whoever is doing downstream analysis on your results, because then they can take your fit depth measurements and the associated uncertainties and use them in a hypothesis test or whatever without having to worry about how to model correlations between the error bars of different bins.  I mention all that because the scoring metric for this competition forces participants to make a prediction for every wavelength, and doesn't enforce any statistical independence across wavelengths, so some solutions (like mine) will make it hard to interpret the spectrum (and in particular its error bars) in a downstream analysis.  Having an unusual optimization target like the one in this competition can really spur a lot of creativity, so again, not complaining.  But one of the things that might be an interesting element of a future competition would be how to interface with downstream analysis targets.  That could mean choosing a clever metric that somehow encourages/enforces statistical independence across wavelength or rewards smart binning choices, or it could mean pulling downstream analysis targets into the competition, e.g. ask participants to measure the concentration of a list of chemicals that might be present in the exoplanet atmospheres, or it could mean some other fun twist.\n\nThanks again -- looking forward to next time!",
      "replies": [
        {
          "id": 3294517,
          "postDate": "2025-09-26T09:57:13.770Z",
          "content": "<blockquote>\n  <p>Don't feel bad about having regenerated the dataset after the competition. I've seen organizers make the mistake of not fixing a problem that surfaced after the start of the competition, and in my opinion that screws things up far more than the need to re-process or re-tune an analysis. In scientific work especially (as I'm sure you know), it's more important to adopt processes and mindsets that will guide you to the right answer in the end than it is to have everything right from the get-go, so kudos for bringing that practice to this competition.</p>\n</blockquote>\n<p>Thanks that is a great validation - we were very worried when we have to do it, but i am glad we did. I have asked this in another thread but would like to ask it again in case you missed - how do you and fellow kaggler feel about post competition changes, how many times could you tolerate it before one said enough is enough?</p>\n<blockquote>\n  <p>I don't mean to sound like I'm complaining; the time constraint did force me to really do some creative thinking and challenge myself on what was really necessary, so I do consider it a net positive. But please don't make it any tighter next time; you probably don't want have a situation where potentially useful approaches have to get tossed out simply because they're not fast enough</p>\n</blockquote>\n<p>Yeah we cant do much with the time constraint and i think it has to be assessed on a case by case basis, we will definitely bring that up with Kaggle in the next round if we are hosting again!</p>\n<blockquote>\n  <p>In a fully conventional/traditional chisquare-fitting approach, one of the things I'd expect to have to do is binning over wavelengths: sum to get a light curve for that bin, then fit the light curve, and get an estimate of the depth parameter along with its fit error. One of the nice things about that approach is that, so long as you don't choose overlapping bins, the error bars that you get for one bin are statistically independent from the next, in the sense that a random fluctuation in one wavelength only affects the one bin it's been assigned to.</p>\n</blockquote>\n<p>I might need to get back to you on this because <a href=\"https://www.kaggle.com/lorenzomugnai\" target=\"_blank\">@lorenzomugnai</a> might have something to say about this. In our field we tend to fit each lightcurve first, extract the Transit Depth measurments , and then bin them down to whatever resolution we would like in downstream tasks, you should still be able to achieve statistically independence that way - but happy to be corrected here. The two approach (bin first then fits VS fit first then bin) should be analogous for noisy lightcurve with white noise, but might differ if they have alternative noise source. </p>\n<p>But we do need a smarter metric i think!</p>",
          "rawMarkdown": "> Don't feel bad about having regenerated the dataset after the competition. I've seen organizers make the mistake of not fixing a problem that surfaced after the start of the competition, and in my opinion that screws things up far more than the need to re-process or re-tune an analysis. In scientific work especially (as I'm sure you know), it's more important to adopt processes and mindsets that will guide you to the right answer in the end than it is to have everything right from the get-go, so kudos for bringing that practice to this competition.\n\nThanks that is a great validation - we were very worried when we have to do it, but i am glad we did. I have asked this in another thread but would like to ask it again in case you missed - how do you and fellow kaggler feel about post competition changes, how many times could you tolerate it before one said enough is enough?\n\n> I don't mean to sound like I'm complaining; the time constraint did force me to really do some creative thinking and challenge myself on what was really necessary, so I do consider it a net positive. But please don't make it any tighter next time; you probably don't want have a situation where potentially useful approaches have to get tossed out simply because they're not fast enough\n\nYeah we cant do much with the time constraint and i think it has to be assessed on a case by case basis, we will definitely bring that up with Kaggle in the next round if we are hosting again!\n\n>In a fully conventional/traditional chisquare-fitting approach, one of the things I'd expect to have to do is binning over wavelengths: sum to get a light curve for that bin, then fit the light curve, and get an estimate of the depth parameter along with its fit error. One of the nice things about that approach is that, so long as you don't choose overlapping bins, the error bars that you get for one bin are statistically independent from the next, in the sense that a random fluctuation in one wavelength only affects the one bin it's been assigned to.\n\nI might need to get back to you on this because @lorenzomugnai might have something to say about this. In our field we tend to fit each lightcurve first, extract the Transit Depth measurments , and then bin them down to whatever resolution we would like in downstream tasks, you should still be able to achieve statistically independence that way - but happy to be corrected here. The two approach (bin first then fits VS fit first then bin) should be analogous for noisy lightcurve with white noise, but might differ if they have alternative noise source. \n\nBut we do need a smarter metric i think!",
          "votes": 1,
          "replies": [
            {
              "id": 3294811,
              "postDate": "2025-09-26T19:39:04.363Z",
              "content": "<p><code>how do you and fellow kaggler feel about post competition changes, how many times could you tolerate it before one said enough is enough?</code></p>\n<p>I can only speak for myself, of course, but for me, it depends on the context and how it's handled and communicated.  In a case like this one, where (to my understanding) your main goal is to get the Kaggle community to help develop practical, useful analysis tools for the satellite you're designing, I'd say, \"reprocess as many times as it takes\", with a few caveats:</p>\n<ul>\n<li>promptly communicating the change to competitors is crucial to maintain fairness of the competition.  As I imagine most folks do, I try to watch the discussion boards for changes like this…but it's easy to imagine missing a post.  So this might be a question for the Kaggle team to consider, e.g. should competitors get an email or some other kind of notification like that when there's a change or reprocessing?  </li>\n<li>if a dataset regeneration happens too close to the end of the competition, I personally think it's 100% valid to extend the competition deadline a bit to give people a chance to reprocess/retrain, as this keeps the competitors' work focused on things that have real scientific value.  That said, it's also valid to argue that a deadline change could cause scheduling problems when a contest is connected to a fixed external deadline like NeurIPS abstract submission, so it's up to you and the other organizers to decide how to weigh the one concern against the other.</li>\n<li>watch out for scope creep, of course 😅</li>\n</ul>\n<p><code>Yeah we cant do much with the time constraint and i think it has to be assessed on a case by case basis, we will definitely bring that up with Kaggle in the next round if we are hosting again!</code></p>\n<p>That sounds perfect; all I can ask is that you keep it in mind and try to make sure that the time available per transit in the competition reflects what would be practical for the folks that will be working with the real data from the satellite.  😁  </p>\n<p><code>In our field we tend to fit each lightcurve first, extract the Transit Depth measurments , and then bin them down to whatever resolution we would like in downstream tasks, you should still be able to achieve statistically independence that way</code></p>\n<p>I agree completely with this approach; I didn't mean to sound like I was casting shade on fit-then-bin as opposed to bin-then-fit.  I was thinking more in terms of some details of my solution that didn't quite seem to me like they'd pass a strict peer review for a paper, like the fact that the 1d conv net that I use for predictions does not guarantee statistical independence of neighboring uncertainties like a traditional binning does.</p>\n<p><code>But we do need a smarter metric i think!</code></p>\n<p>I know there's been a lot of discussion about the quirks of the metric used in this competition, and that sort of critical commentary can have a lot of value.  (Clipping <em>would</em> make the analysis development less scary, for example.)  But FWIW, I do appreciate the concept of this competition's metric.  It's honestly pretty refreshing to see a competition where the metric cares about both the predicted signal values and their uncertainties (as opposed to just the central values), and the way this contest's metric achieves that seems quite clever to me.  </p>\n<p>Something worth celebrating about this competition is the great variety in the solutions people are posting, and I can't help but think that part of the reason for that is the metric.  It grants competitors the freedom to reframe the problem of light curve fitting as, for example, a signal processing problem where they do smoothing over wavelengths first and then try to extract fit depth, or to view the spectrum as a continuous function to be approximated with a polynomial whose parameters become the prediction targets.  And all the while, the metric forces competitors not to lose touch with the notion of an uncertainty.  It looks like a successful metric to me, in spite of its quirks.</p>\n<p>I only mention the idea about having future iterations of this competition pull in downstream prediction targets because I wonder if brainstorming in that direction might open the playing field to even more of that same kind of lateral thinking.</p>",
              "rawMarkdown": "`how do you and fellow kaggler feel about post competition changes, how many times could you tolerate it before one said enough is enough?`\n\nI can only speak for myself, of course, but for me, it depends on the context and how it's handled and communicated.  In a case like this one, where (to my understanding) your main goal is to get the Kaggle community to help develop practical, useful analysis tools for the satellite you're designing, I'd say, \"reprocess as many times as it takes\", with a few caveats:\n * promptly communicating the change to competitors is crucial to maintain fairness of the competition.  As I imagine most folks do, I try to watch the discussion boards for changes like this...but it's easy to imagine missing a post.  So this might be a question for the Kaggle team to consider, e.g. should competitors get an email or some other kind of notification like that when there's a change or reprocessing?  \n * if a dataset regeneration happens too close to the end of the competition, I personally think it's 100% valid to extend the competition deadline a bit to give people a chance to reprocess/retrain, as this keeps the competitors' work focused on things that have real scientific value.  That said, it's also valid to argue that a deadline change could cause scheduling problems when a contest is connected to a fixed external deadline like NeurIPS abstract submission, so it's up to you and the other organizers to decide how to weigh the one concern against the other.\n * watch out for scope creep, of course 😅\n\n`Yeah we cant do much with the time constraint and i think it has to be assessed on a case by case basis, we will definitely bring that up with Kaggle in the next round if we are hosting again!`\n\nThat sounds perfect; all I can ask is that you keep it in mind and try to make sure that the time available per transit in the competition reflects what would be practical for the folks that will be working with the real data from the satellite.  😁  \n\n` In our field we tend to fit each lightcurve first, extract the Transit Depth measurments , and then bin them down to whatever resolution we would like in downstream tasks, you should still be able to achieve statistically independence that way`\n\nI agree completely with this approach; I didn't mean to sound like I was casting shade on fit-then-bin as opposed to bin-then-fit.  I was thinking more in terms of some details of my solution that didn't quite seem to me like they'd pass a strict peer review for a paper, like the fact that the 1d conv net that I use for predictions does not guarantee statistical independence of neighboring uncertainties like a traditional binning does.\n\n`But we do need a smarter metric i think!`\n\nI know there's been a lot of discussion about the quirks of the metric used in this competition, and that sort of critical commentary can have a lot of value.  (Clipping *would* make the analysis development less scary, for example.)  But FWIW, I do appreciate the concept of this competition's metric.  It's honestly pretty refreshing to see a competition where the metric cares about both the predicted signal values and their uncertainties (as opposed to just the central values), and the way this contest's metric achieves that seems quite clever to me.  \n\nSomething worth celebrating about this competition is the great variety in the solutions people are posting, and I can't help but think that part of the reason for that is the metric.  It grants competitors the freedom to reframe the problem of light curve fitting as, for example, a signal processing problem where they do smoothing over wavelengths first and then try to extract fit depth, or to view the spectrum as a continuous function to be approximated with a polynomial whose parameters become the prediction targets.  And all the while, the metric forces competitors not to lose touch with the notion of an uncertainty.  It looks like a successful metric to me, in spite of its quirks.\n\nI only mention the idea about having future iterations of this competition pull in downstream prediction targets because I wonder if brainstorming in that direction might open the playing field to even more of that same kind of lateral thinking."
            }
          ]
        }
      ]
    },
    {
      "id": 3293837,
      "postDate": "2025-09-24T18:34:15.133Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 3293839,
          "postDate": "2025-09-24T18:43:42.127Z",
          "content": "<p>Midnight CET tonight.</p>",
          "rawMarkdown": "Midnight CET tonight."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3293812,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2025-09-24T17:35:32.247000",
      "content": "<p>You ask for feedback, here is some. Fasten your seat belt as I'll be frank.</p>\n<p>We did not spend a second on dealing with cases where transit is not contained in the provided data. We should have done it if our motivation was the LB score. But I personally could not resolve working on something useless. I enter kaggle competitions to learn something useful. I don't enter to get points.</p>\n<p>I wish you had agreed to remove these cases when I asked, 1.5 months before end of competition. That was less than 2 weeks after data regeneration.</p>\n<p>Instead you led many to spend numerous hours dealing with something that will not be useful to ARIAL mission. Being respectful of the time participants spend should be of utter importance to competition hosts.</p>\n<p>Other than that, this was a fantastic competition. Thanks for hosting it. Can't wait for next year edition, if you do one. Topic is great. learning is great. I wish we had found what we found last week earlier in the competition. We'd love a couple of weeks extension…</p>",
      "votes": 5,
      "replies": [
        {
          "id": 3293840,
          "author_name": "Jeroen Cottaar",
          "author_url": "",
          "post_date": "2025-09-24T18:44:47.877000",
          "content": "<p>I wonder how much special methods for the clipped transits really matter; my solution doesn't involve any at least. I hope the hosts follow through with the alternative score calculation with these transits removed.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3293894,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2025-09-25T00:21:49.707000",
              "content": "<p>Few teams dropped dramatically in private. I bet it is because of some of these degenerate cases. There is no lower bound on how bad a planet score can be. </p>\n<p>All we did was to put a large sigma for these, basically turning their scores to zero. But we lose at least 4% of score that way.</p>\n<p>Anyway, I am now looking forward to your writeup. Congrats one the amazing result.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3293902,
              "author_name": "Timmy Juicehouse",
              "author_url": "",
              "post_date": "2025-09-25T01:28:34.090000",
              "content": "<p>I am one of the victims involved. I'm very curious to see how the top kagglers will review and explain this situation.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3293906,
              "author_name": "mudesteven",
              "author_url": "",
              "post_date": "2025-09-25T01:49:48.647000",
              "content": "<p>I’m one of the victims as well. I’m curious to see how you’re solution was able to compensate for these situations without clipping. For us, our current methodology would produce planets that could have as high as 100000 ppm when we investigated it right after the competition. We didn’t anticipate that this would have such a huge impact as our public lb was quite good before the competition closed, and thus we didn’t do any special measurements against these situations. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3294011,
              "author_name": "Pascal Pfeiffer",
              "author_url": "",
              "post_date": "2025-09-25T07:34:57.093000",
              "content": "<blockquote>\n  <p>All we did was to put a large sigma for these, basically turning their scores to zero. But we lose at least 4% of score that way.</p>\n</blockquote>\n<p>That is what I did as well. Apparently, one or more slipped my hard sample detection in the private set and dropped me to the bottom of the private leaderboard. <br>\nI think the metric should be updated for next year (I hope we get to see another round of ARIEL) and shouldn't be impacted by single bad samples too much. Can always happen that a fit badly diverges for single samples and in reality, this would be catched easily by supervising the system or looking at the results.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3294079,
              "author_name": "Oleh Kivernyk",
              "author_url": "",
              "post_date": "2025-09-25T10:15:32.073000",
              "content": "<p>I had quite robust fallbacks in my approach as well, but it happened to be not so robust enough to prevent dropping to zero score. Feels quite unfair given the invested time…</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 3294505,
              "author_name": "Gordon Yip",
              "author_url": "",
              "post_date": "2025-09-26T09:40:09.210000",
              "content": "<p>yeah we hear you - i think bad examples do occur and they have their values, but the metric should be improved so that it wouldn't drag down the score drastically due to a few bad examples</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3296243,
              "author_name": "Oleh Kivernyk",
              "author_url": "",
              "post_date": "2025-09-30T13:52:17.023000",
              "content": "<p>I have extensively tested the private LB data, and I'm now quite confident that my score is pulled down by specific LD configurations used to simulate the private data.</p>\n<p>In my fitting framework I have very strict quality controls. I control not just the quality of the fits but also the quality of the resulting physical model. I reparametrized many aspects of the model to reduce parameter degeneracies as much as possible. I also use a dynamic choice between uniform/linear/quadratic(Kipping) LDs based on NLL/AIC. If fit quality criteria falls below threshold and/or the covariance shows strong RpRs-b correlations, etc., I fallback to my \"physics unaware\" model from last year, which scors 0.428/0.428 this year.</p>\n<p>Just allowing my model to accept \"unphysical\" limb brightening magically fixed the zero score issue on the private LB:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6327791%2F75f301ab2842da3149a776ac71e8b536%2FADC.png?generation=1759238082134729&amp;alt=media\" alt=\"\"></p>\n<p>This doesn't necessarily mean that the private LB set contains examples with Limb Brightening ( <a href=\"https://www.kaggle.com/gordonyip\" target=\"_blank\">@gordonyip</a> -- could you confirm this?). It's possible that the data was simulated with more exotic LD profiles, and that allowing limb brightening in your model simply absorbs the bias that would otherwise collapse the score to zero.</p>\n<p><strong>To conclude, this highlights that GLL may not be the perfect evaluation metric for this competition</strong>. It tends to penalize very accurate physics-constrained models in favor of simpler \"physics-unaware\" models that are less sensitive to parameter degeneracies.</p>\n<p><a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> -- probably you also constrained aggressively your LD model? :)</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3296599,
              "author_name": "Jeroen Cottaar",
              "author_url": "",
              "post_date": "2025-10-01T08:42:54.883000",
              "content": "<p>I also used constraints similar to this to keep to check the sanity of the modeling results. However, the approach I took is keeping the model free, and raising an error if the constraints are violated. This then allowed me to make deliberate choices on how to deal with these planets (there were 2 of them); in the end, I didn't do anything in my final submission, but I had a backup submission where I replaced them by dummy submission (mean and std of the other planets).</p>\n<p>I can't check right now, but I don't think I had such checks on the limb darkening parameters by the way.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3296602,
              "author_name": "Jeroen Cottaar",
              "author_url": "",
              "post_date": "2025-10-01T08:51:03.583000",
              "content": "<p>Actually, I now realize I'm confused by your post above. Your model falls back to a simpler, more robust model if the limb darkening parameters are not plausible (i.e. similar to my approach explained above). But how did you end up with a score of 0.000 due to limb brightening then? Why didn't you model fall back to the robust model on those planets?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3296627,
              "author_name": "Oleh Kivernyk",
              "author_url": "",
              "post_date": "2025-10-01T09:43:06.380000",
              "content": "<p>There are some model configurations used to simulate data that lead to degeneracy of parameters. So from the transit data  you cannot distinguish parameters apart without additional priors/constraints. They trade off against each other when you try to maximize likelihood. From my observation some planets had covariances leading to large correlations. For example, RpRs &lt;-&gt; impact parameter b &lt;-&gt; LD coeffs can impact each other so that data does not have constraining power alone.</p>\n<p>In my case, when one of the LD coefficients is stuck at 'physical' boundary - the resulting model still can have very high quality in description of transit data (very good NLL or chi2/ndof). This leads to RpRs bias, whereas the <code>b</code> parameter compensated the rest of data/model difference that the LD had problem to catch (even when b is constrained from stellar parameters with gaussian prior). </p>\n<p>So in this case we have:</p>\n<ul>\n<li>very good model description of transit data (confidence), so no fallback</li>\n<li>no extreme correlations from covariance matrix, no fallback</li>\n<li>biased RpRs with respect to RpRs(true), biased 'b' with respect to b(true) -&gt; huge penalty to GLL</li>\n<li>LD(quadratic/Kipping) did not approximate the host's custom LD(true) because of not enough freedom (or because few planets with Limb Brightening were in private data)</li>\n</ul>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3296634,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2025-10-01T10:23:56.760000",
              "content": "<p>In our solution, we had to limit the accuracy of curve fit because it led to degraded LB while CV was improved significantly. If I was not on vacation then I'd rerun the way Jeroen did: relax constraints and trigger error if they are violated, with a fallback.</p>\n<p>We took a conservative approach and put stricter constraints than nececessa. As a result we did not drop, but didn't get a greta score either.</p>\n<p>I find it concerning if the ground truth contains physically impossible answers. I hope <a href=\"https://www.kaggle.com/gordonyip\" target=\"_blank\">@gordonyip</a> will address that shortly.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3296852,
              "author_name": "Jeroen Cottaar",
              "author_url": "",
              "post_date": "2025-10-01T18:56:31.543000",
              "content": "<p>In my case there were 2 fallbacks in the private test set, but I'm not sure if they're physically unrealistic. They were caused by:</p>\n<ul>\n<li>An extremely high mean value for the transit depth (20%).</li>\n<li>An unexpected high residual of the AIRS signal (around 1.15x the noise level if I recall correctly), likely indicating the fit landed in a bad local minimum.</li>\n</ul>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3297055,
              "author_name": "Oleh Kivernyk",
              "author_url": "",
              "post_date": "2025-10-02T08:34:23.743000",
              "content": "<p>Just for completeness, I resubmitted my code with:</p>\n<ul>\n<li>Relaxed Kipping's q2 limits to [-1, 1]</li>\n<li>If q2&lt;0, I fall back to my base model (0.428/0.428)</li>\n</ul>\n<p><a href=\"https://arxiv.org/pdf/1308.0009\" target=\"_blank\">Kipping</a> parametrized the quadratic LB in terms of q1, q2 as <code>u1 = 2.0 * np.sqrt(q1) * q2; u2 = np.sqrt(q1) * (1.0 - 2.0 * q2)</code>. This parametrization reduces correlations between LD coefficients during regression, which 'must be' very useful in this competition if you are not fixing them from LD tables.  <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6327791%2Fb78de93fdbb1a5110c14b21001d69609%2Ffallback.png?generation=1759393222303450&amp;alt=media\" alt=\"\"><br>\nFrom this test, the fallback happens probably once or twice on the public LB data, while it's more often on the private LB. As I have already mentioned before, this does not necessary mean that there are planets with truly unphysically LDs. It may just contain more exotic LD profiles, where my setup prefers negative q2 to maximize likelihood.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 3293901,
          "author_name": "Horikita Saku",
          "author_url": "",
          "post_date": "2025-09-25T00:40:03.177000",
          "content": "<p>Indeed so. After using the phase detector to detect the abnormal planet, I couldn't think of any additional solutions.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 3294503,
          "author_name": "Gordon Yip",
          "author_url": "",
          "post_date": "2025-09-26T09:35:23.463000",
          "content": "<p>Thanks for being honest and upfront. We do need this to make better competition. </p>\n<p>As for whether to remove the 'bad' transit or not. A few things were going through our heads at the time - <br>\nbad transits could occur due to human/measurement error (Ariel is no exception), and we have seen many datasets discarded because of bad planning, but some also managed to produce some scientific values (but I am sure you can imagine the level of engineering they have to go through… ). As hosts we wanted to know how participants could mitigate that, given these examples, or if they can. </p>\n<p>the other thing is on the organisation level,  Our consideration at the time was how much we can <code>rock</code> the boat before everyone starts leaving the ship. To be honest it is the first time we have to re-generate the dataset and we are worried if a second change might rock the boat too much</p>\n<p>so altogether we decided not to. </p>\n<p>In hindsight though- we didnt realise they are having a big impact on the metric, something we should have foreseen, and the care one must take for these bad examples may be too big of an investment for the competition. </p>\n<p>Here is my question to you and fellow Kagglers - how do you feel about data regeneration, are you happy to see them regenerated , how many times could you tolerate it and what would have been your preferred re-action from the host , learning this will help our decision making next time.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3294515,
              "author_name": "sroger",
              "author_url": "",
              "post_date": "2025-09-26T09:55:15.860000",
              "content": "<p>I joined around the time the regeneration happened so it didn't affect me much.<br>\nIf there was another regeneration to rectify the broken transits I would welcome it.</p>\n<p>It comes down to whether our attempts to mitigate these bad transits provide scientific value.<br>\nFrom your previous statements it seems like our efforts would be better spent elsewhere.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3294537,
              "author_name": "Pascal Pfeiffer",
              "author_url": "",
              "post_date": "2025-09-26T10:09:21.093000",
              "content": "<p>This competition as well as the last competition was very well done in my opinion. It is quite a unique task and I totally understand the difficulties around synthetic data.</p>\n<blockquote>\n  <p>Here is my question to you and fellow Kagglers - how do you feel about data regeneration, are you happy to see them regenerated , how many times could you tolerate it and what would have been your preferred re-action from the host , learning this will help our decision making next time.</p>\n</blockquote>\n<p>Early on in a competition, I think data re-generation is not a big issue at all, even multiple times. As long as it helps the competition to be more robust, more predictable, more scientifically correct, or to prevent leakage. Ideally, the solutions wouldn't really be impacted much, but that isn't always possible. Of course, open communication is a must here, and you did very well on that. </p>\n<p>I personally would have welcomed another regeneration quickly after the first one when it became clear that the transit length was too long for some samples. Maybe, even excluding samples from the metric would have already solved this and is only a minor change. </p>\n<p>I think the main issue is rather the unbound metric on a per-sample basis. If a single wrong sample can drive the metric to zero, things can get complicated quickly, if the metric wouldn't allow this, I guess we wouln't even have the discussion as these handful of samples would only impact the score on the third digit.</p>\n<p>That said, I am looking forward to a new round of ARIEL. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3294859,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2025-09-27T00:03:11.683000",
              "content": "<p>Data regeneration was not a real issue, mostly because we cudl reuse the same pipeline as before.</p>\n<p>But not fixing the few glitches was a real mistake. it would not have rocked the boat at all. </p>\n<blockquote>\n  <p>we didnt realise they are having a big impact on the metric,</p>\n</blockquote>\n<p>I warned you about the consequence on the metric when i asked for removal. But I can't find my post. It seems you deleted the topic you had for questions from us. Why is that?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3294909,
              "author_name": "Jeroen Cottaar",
              "author_url": "",
              "post_date": "2025-09-27T04:59:47.030000",
              "content": "<p><a href=\"https://www.kaggle.com/competitions/ariel-data-challenge-2025/discussion/589905#3269906\" target=\"_blank\">https://www.kaggle.com/competitions/ariel-data-challenge-2025/discussion/589905#3269906</a></p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3294951,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2025-09-27T07:44:04.737000",
              "content": "<p>Thank <a href=\"https://www.kaggle.com/jeroencottaar\" target=\"_blank\">@jeroencottaar</a> , it was just unpinned. <a href=\"https://www.kaggle.com/gordonyip\" target=\"_blank\">@gordonyip</a> here is where I warned you of the consequences of bad data on the metric <a href=\"https://www.kaggle.com/competitions/ariel-data-challenge-2025/discussion/598324#3269605\" target=\"_blank\">https://www.kaggle.com/competitions/ariel-data-challenge-2025/discussion/598324#3269605</a></p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3294490,
      "author_name": "Gordon Yip",
      "author_url": "",
      "post_date": "2025-09-26T09:07:12.730000",
      "content": "<p>wow the discussion exploded, it will take me some time to go through all of them , but thank you for all the responses and discussion!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3293948,
      "author_name": "Thomas Dueholm Hansen",
      "author_url": "",
      "post_date": "2025-09-25T03:55:17.670000",
      "content": "<p>Thanks for organizing a good competition!</p>\n<p>Referencing <a href=\"https://www.kaggle.com/competitions/ariel-data-challenge-2024/writeups/c-number-daiwakun-1st-place-solution\" target=\"_blank\">last year's first place solution</a>, the gain drift f(t) has a significant impact on the result. In order to accurately estimate f(t), I would say that ideally the transit should only take up about one third of the data. Since in reality the extra data is collected anyway, as CPMP pointed out, I think it would be better to provide that as well.</p>\n<p>On the other hand, I found the calibration and binning to be relatively uninteresting. I would rather that preprocessing had already been performed by the organizers, similar to what was done in <a href=\"https://www.kaggle.com/code/gordonyip/calibrating-and-binning-ariel-data\" target=\"_blank\">the public notebook by Gordon Yip</a>. With the same amount of data, the number of unique planet ids could then be significantly increased, and more data could be provided outside the transit window. Maybe the number of planets has been limited to match the data that will be collected in the future? I would argue that that is a mistake. If the simulations are accurate enough, I think it would be better to train the final model on 50,000+ simulated planets. If you want a Kaggle competition to reflect that, then you should do calibration and binning yourself.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3294068,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2025-09-25T09:40:31.840000",
          "content": "<blockquote>\n  <p>the transit should only take up about one third of the data</p>\n</blockquote>\n<p>This is part of Ariel specs actually. I don't have the reference handy, but it is stated clearly. It is why I don't get we got data that did not comply with it.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3293841,
      "author_name": "Jeroen Cottaar",
      "author_url": "",
      "post_date": "2025-09-24T18:48:42.067000",
      "content": "<p>It was another great competition!</p>\n<p>One aspect you might want to consider is being more open about the generation of the synthetic data. While the overall flow is published, there are many details that we have to figure out ourselves. But many of these, such as using a polynomial for the gain drift, are not actually physically realistic, so the value of our work there is limited. And any results we find will anyway be limited by the accuracy of your assumptions in generating the data. </p>\n<p>I wonder what this competition might have looked like if the full synthetic data generation pipeline (to the point that we can reproduce it) was released, along with a qualitative description of the changes made for the test set. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 3293843,
          "author_name": "Jeroen Cottaar",
          "author_url": "",
          "post_date": "2025-09-24T18:54:47.583000",
          "content": "<p>Oh and a smaller one, also mentioned in a recent discussion: you could consider clipping the score per planet (perhaps only at 0, not at 1 - why do you clip at 1 anyway?). This means you don't get people scoring 0 just because a single planet went wildly wrong (which you'd likely catch easily when you actually see the data in the real mission). There is an art to making the model robust against that - but with only two tries on the private test set, it's not really fair.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3293979,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-09-25T05:49:58.997000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3294543,
              "author_name": "Gordon Yip",
              "author_url": "",
              "post_date": "2025-09-26T10:16:16.763000",
              "content": "<p>yes clipping the score was one option we were considering. By not clipping we wanted to encourage more conservative solutions, so model that will avoid predicting the transits depth very badly and away from the GT. I think in the real mission - it is still relatively easy to catch by running the conventional pipeline, but as thing scale up, it might become less easy to spot those cases as one would really fit them to get the depth out in the classical approach.</p>\n<p>But driving the whole score to 0 is very bad , sorry!</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 3294064,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2025-09-25T09:38:44.797000",
          "content": "<p>I disagree on disclosing the data generator.</p>\n<p>On the contrary, I think it should be made as complex as possible to prevent reverse engineering. That's the only way to assess model generalization IMHO.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3294527,
          "author_name": "Gordon Yip",
          "author_url": "",
          "post_date": "2025-09-26T10:05:11.537000",
          "content": "<p>Yes we are being very secretive about this. While disclosing the whole thing might make it clearer for each step, it also allows room for reverse engineering. Before the competition we had a long discussion with Kaggle to make sure the data generation pipeline is opaque. Reverse engineering the solution have happened before and does not provide much value to us and does not make it an enjoyable experience for those who tried to come up with alternative solution. </p>\n<p>However, the competition is quite a domain heavy one so we disclose some guidelines here and there, but locked up the full details in a paper. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 3294447,
      "author_name": "sroger",
      "author_url": "",
      "post_date": "2025-09-26T07:32:22.170000",
      "content": "<p>Some feedback for next time:</p>\n<ol>\n<li>Pre-calibrate the data. I haven't seen any novel calibration methods that improve over the standard so far.</li>\n<li>Remove clipped/malformed transits and other artifacts that are not part of the mission.</li>\n<li>The metric should be more robust and less sensitive to outliers, though I don't have any specific suggestions off the top of my head.</li>\n<li>The data generation pipeline should be obfuscated as much as possible, reverse engineering your pipeline provides little scientific value.</li>\n</ol>\n<p>Thanks for the competition! First time working on astronomy and learned a lot.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3294061,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2025-09-25T09:35:19.373000",
      "content": "<p>Another feedback. I gave it during the comp, but maybe it has been lost. Repeating it here.</p>\n<p>You ask for sigmas of our predictions. This makes lot of sense.</p>\n<p>But you did not provide the sigmas for the info in the star_info file. This makes no sense at all. And it looks unfair: you ask us for something you don't provide.</p>\n<p>You know about the noise in that data. We can see that noise by applying Kepler 2nd law for instance. We also see it (and that was discussed in the forum) in the discrepancy bewtween the measure transit time and the predicted transit time.</p>\n<p>Next time, if any, please provide std for every physical value you provide.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3293899,
      "author_name": "Horikita Saku",
      "author_url": "",
      "post_date": "2025-09-25T00:38:10.613000",
      "content": "<p>Thank you very much for your organization! This was an absolutely wonderful game. Especially, I was deeply impressed by your meticulousness in dealing with physics!</p>\n<p>At the same time, I am very much looking forward to seeing the rankings after the abnormal samples have been removed.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3293816,
      "author_name": "particlebbq",
      "author_url": "",
      "post_date": "2025-09-24T17:43:42.933000",
      "content": "<p>Right back at you -- thanks for organizing this competition; it's been a fun one!  And for this thread, which promises an interesting discussion.  Here are some thoughts that come to my mind after reading your post:</p>\n<ul>\n<li>Don't feel bad about having regenerated the dataset after the competition.  I've seen organizers make the mistake of not fixing a problem that surfaced after the start of the competition, and in my opinion that screws things up far more than the need to re-process or re-tune an analysis.  In scientific work especially (as I'm sure you know), it's more important to adopt processes and mindsets that will guide you to the right answer in the end than it is to have everything right from the get-go, so kudos for bringing that practice to this competition.</li>\n<li>the approach I took for my solution is a chisquare fit for each wavelength, whose outputs feed into a neural net that does smoothing and corrects for things like limb darkening and internal fit bias.  The fits, at least, are meant to look a lot like the way we'd do data analysis back when I was a grad student in collider physics, which is to say that, to me, my solution feels (deliberately) very conventional.  But with that chisquare fit approach, I found the runtime constraint to be <em>brutal</em>.  It was only a few days ago that I managed to get things to a state where I was confident my submission would run to completion in the allowed time, and as I write this, I still don't know if the submission I have running will produce a valid score or not. (The dreaded \"submission scoring error\" got my first entry.)  I don't mean to sound like I'm complaining; the time constraint did force me to really do some creative thinking and challenge myself on what was really necessary, so I do consider it a net positive.  But please don't make it any tighter next time; you probably don't want have a situation where potentially useful approaches have to get tossed out simply because they're not fast enough.  😅</li>\n<li>In a fully conventional/traditional chisquare-fitting approach, one of the things I'd expect to have to do is binning over wavelengths: sum to get a light curve for that bin, then fit the light curve, and get an estimate of the depth parameter along with its fit error.  One of the nice things about that approach is that, so long as you don't choose overlapping bins, the error bars that you get for one bin are statistically independent from the next, in the sense that a random fluctuation in one wavelength only affects the one bin it's been assigned to.  That is a real benefit for whoever is doing downstream analysis on your results, because then they can take your fit depth measurements and the associated uncertainties and use them in a hypothesis test or whatever without having to worry about how to model correlations between the error bars of different bins.  I mention all that because the scoring metric for this competition forces participants to make a prediction for every wavelength, and doesn't enforce any statistical independence across wavelengths, so some solutions (like mine) will make it hard to interpret the spectrum (and in particular its error bars) in a downstream analysis.  Having an unusual optimization target like the one in this competition can really spur a lot of creativity, so again, not complaining.  But one of the things that might be an interesting element of a future competition would be how to interface with downstream analysis targets.  That could mean choosing a clever metric that somehow encourages/enforces statistical independence across wavelength or rewards smart binning choices, or it could mean pulling downstream analysis targets into the competition, e.g. ask participants to measure the concentration of a list of chemicals that might be present in the exoplanet atmospheres, or it could mean some other fun twist.</li>\n</ul>\n<p>Thanks again -- looking forward to next time!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3294517,
          "author_name": "Gordon Yip",
          "author_url": "",
          "post_date": "2025-09-26T09:57:13.770000",
          "content": "<blockquote>\n  <p>Don't feel bad about having regenerated the dataset after the competition. I've seen organizers make the mistake of not fixing a problem that surfaced after the start of the competition, and in my opinion that screws things up far more than the need to re-process or re-tune an analysis. In scientific work especially (as I'm sure you know), it's more important to adopt processes and mindsets that will guide you to the right answer in the end than it is to have everything right from the get-go, so kudos for bringing that practice to this competition.</p>\n</blockquote>\n<p>Thanks that is a great validation - we were very worried when we have to do it, but i am glad we did. I have asked this in another thread but would like to ask it again in case you missed - how do you and fellow kaggler feel about post competition changes, how many times could you tolerate it before one said enough is enough?</p>\n<blockquote>\n  <p>I don't mean to sound like I'm complaining; the time constraint did force me to really do some creative thinking and challenge myself on what was really necessary, so I do consider it a net positive. But please don't make it any tighter next time; you probably don't want have a situation where potentially useful approaches have to get tossed out simply because they're not fast enough</p>\n</blockquote>\n<p>Yeah we cant do much with the time constraint and i think it has to be assessed on a case by case basis, we will definitely bring that up with Kaggle in the next round if we are hosting again!</p>\n<blockquote>\n  <p>In a fully conventional/traditional chisquare-fitting approach, one of the things I'd expect to have to do is binning over wavelengths: sum to get a light curve for that bin, then fit the light curve, and get an estimate of the depth parameter along with its fit error. One of the nice things about that approach is that, so long as you don't choose overlapping bins, the error bars that you get for one bin are statistically independent from the next, in the sense that a random fluctuation in one wavelength only affects the one bin it's been assigned to.</p>\n</blockquote>\n<p>I might need to get back to you on this because <a href=\"https://www.kaggle.com/lorenzomugnai\" target=\"_blank\">@lorenzomugnai</a> might have something to say about this. In our field we tend to fit each lightcurve first, extract the Transit Depth measurments , and then bin them down to whatever resolution we would like in downstream tasks, you should still be able to achieve statistically independence that way - but happy to be corrected here. The two approach (bin first then fits VS fit first then bin) should be analogous for noisy lightcurve with white noise, but might differ if they have alternative noise source. </p>\n<p>But we do need a smarter metric i think!</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3294811,
              "author_name": "particlebbq",
              "author_url": "",
              "post_date": "2025-09-26T19:39:04.363000",
              "content": "<p><code>how do you and fellow kaggler feel about post competition changes, how many times could you tolerate it before one said enough is enough?</code></p>\n<p>I can only speak for myself, of course, but for me, it depends on the context and how it's handled and communicated.  In a case like this one, where (to my understanding) your main goal is to get the Kaggle community to help develop practical, useful analysis tools for the satellite you're designing, I'd say, \"reprocess as many times as it takes\", with a few caveats:</p>\n<ul>\n<li>promptly communicating the change to competitors is crucial to maintain fairness of the competition.  As I imagine most folks do, I try to watch the discussion boards for changes like this…but it's easy to imagine missing a post.  So this might be a question for the Kaggle team to consider, e.g. should competitors get an email or some other kind of notification like that when there's a change or reprocessing?  </li>\n<li>if a dataset regeneration happens too close to the end of the competition, I personally think it's 100% valid to extend the competition deadline a bit to give people a chance to reprocess/retrain, as this keeps the competitors' work focused on things that have real scientific value.  That said, it's also valid to argue that a deadline change could cause scheduling problems when a contest is connected to a fixed external deadline like NeurIPS abstract submission, so it's up to you and the other organizers to decide how to weigh the one concern against the other.</li>\n<li>watch out for scope creep, of course 😅</li>\n</ul>\n<p><code>Yeah we cant do much with the time constraint and i think it has to be assessed on a case by case basis, we will definitely bring that up with Kaggle in the next round if we are hosting again!</code></p>\n<p>That sounds perfect; all I can ask is that you keep it in mind and try to make sure that the time available per transit in the competition reflects what would be practical for the folks that will be working with the real data from the satellite.  😁  </p>\n<p><code>In our field we tend to fit each lightcurve first, extract the Transit Depth measurments , and then bin them down to whatever resolution we would like in downstream tasks, you should still be able to achieve statistically independence that way</code></p>\n<p>I agree completely with this approach; I didn't mean to sound like I was casting shade on fit-then-bin as opposed to bin-then-fit.  I was thinking more in terms of some details of my solution that didn't quite seem to me like they'd pass a strict peer review for a paper, like the fact that the 1d conv net that I use for predictions does not guarantee statistical independence of neighboring uncertainties like a traditional binning does.</p>\n<p><code>But we do need a smarter metric i think!</code></p>\n<p>I know there's been a lot of discussion about the quirks of the metric used in this competition, and that sort of critical commentary can have a lot of value.  (Clipping <em>would</em> make the analysis development less scary, for example.)  But FWIW, I do appreciate the concept of this competition's metric.  It's honestly pretty refreshing to see a competition where the metric cares about both the predicted signal values and their uncertainties (as opposed to just the central values), and the way this contest's metric achieves that seems quite clever to me.  </p>\n<p>Something worth celebrating about this competition is the great variety in the solutions people are posting, and I can't help but think that part of the reason for that is the metric.  It grants competitors the freedom to reframe the problem of light curve fitting as, for example, a signal processing problem where they do smoothing over wavelengths first and then try to extract fit depth, or to view the spectrum as a continuous function to be approximated with a polynomial whose parameters become the prediction targets.  And all the while, the metric forces competitors not to lose touch with the notion of an uncertainty.  It looks like a successful metric to me, in spite of its quirks.</p>\n<p>I only mention the idea about having future iterations of this competition pull in downstream prediction targets because I wonder if brainstorming in that direction might open the playing field to even more of that same kind of lateral thinking.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3293837,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-09-24T18:34:15.133000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 3293839,
          "author_name": "Jeroen Cottaar",
          "author_url": "",
          "post_date": "2025-09-24T18:43:42.127000",
          "content": "<p>Midnight CET tonight.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3293771": "This is the last day of the competition. As hosts, we just want to say a few words before we hand over the glory to this year's winners.\n\nWe know that this competition has been particularly challenging. A few of you described it as \"ADC2024 on steroids\" and some might even say, \"forget everything you know about ADC2024.\"\n\nIn the next few weeks, we will find out who the winners are, and like many of you, we can't wait to find out their approaches.\n\nAs hosts, we have been keeping a watchful eye on the progress of the competition, and we are grateful to have you with us. I mean, honestly, without you, this competition would not be a success!\n\nThis year has been nothing short of surprises – from data re-generation and faulty data in the dataset to crazy transits, questions about the meaning of the metric, and many more issues that you might have discovered during your journey through the dataset and leaderboard.\n\nWe just want to say thank you. Thank you for bearing with us. \n\nThe competition is not perfect, but we intend to build on your help and feedback. If you think you might have learned something quite interesting, please feel free to share and tag us! We will be writing a report to document this experience.\n\nIf this competition has made you learn a bit more about astronomy, about the Ariel mission, and the data we are currently dealing with, we consider this a success.\n\nHere is what we are wondering – what stuck with YOU about this competition? \nThe frustrations, the breakthroughs, the random moments of enlightenment? Did this change how you think about astronomy or the Ariel mission? Did you stumble onto techniques you'd never tried before?\n\nFeel free to use this space to share.\n\nSome of you might be contacted by us for a short interview – we would like to find out about your journey. Or DM me if you want to have a chat.\n\nLast thing - our big thank you to @sohier, @elizabethpark and many Kaggle staff for their help in this year's competition. \n\nLast last thing: Please enjoy a little [barchart](https://imgur.com/a/0QVWXsX) race about this year's sprint!\n\n",
    "3293812": "You ask for feedback, here is some. Fasten your seat belt as I'll be frank.\n\nWe did not spend a second on dealing with cases where transit is not contained in the provided data. We should have done it if our motivation was the LB score. But I personally could not resolve working on something useless. I enter kaggle competitions to learn something useful. I don't enter to get points.\n\nI wish you had agreed to remove these cases when I asked, 1.5 months before end of competition. That was less than 2 weeks after data regeneration.\n\nInstead you led many to spend numerous hours dealing with something that will not be useful to ARIAL mission. Being respectful of the time participants spend should be of utter importance to competition hosts.\n\nOther than that, this was a fantastic competition. Thanks for hosting it. Can't wait for next year edition, if you do one. Topic is great. learning is great. I wish we had found what we found last week earlier in the competition. We'd love a couple of weeks extension...",
    "3294490": "wow the discussion exploded, it will take me some time to go through all of them , but thank you for all the responses and discussion!",
    "3293948": "Thanks for organizing a good competition!\n\nReferencing [last year's first place solution](https://www.kaggle.com/competitions/ariel-data-challenge-2024/writeups/c-number-daiwakun-1st-place-solution), the gain drift f(t) has a significant impact on the result. In order to accurately estimate f(t), I would say that ideally the transit should only take up about one third of the data. Since in reality the extra data is collected anyway, as CPMP pointed out, I think it would be better to provide that as well.\n\nOn the other hand, I found the calibration and binning to be relatively uninteresting. I would rather that preprocessing had already been performed by the organizers, similar to what was done in [the public notebook by Gordon Yip](https://www.kaggle.com/code/gordonyip/calibrating-and-binning-ariel-data). With the same amount of data, the number of unique planet ids could then be significantly increased, and more data could be provided outside the transit window. Maybe the number of planets has been limited to match the data that will be collected in the future? I would argue that that is a mistake. If the simulations are accurate enough, I think it would be better to train the final model on 50,000+ simulated planets. If you want a Kaggle competition to reflect that, then you should do calibration and binning yourself.",
    "3293841": "It was another great competition!\n\nOne aspect you might want to consider is being more open about the generation of the synthetic data. While the overall flow is published, there are many details that we have to figure out ourselves. But many of these, such as using a polynomial for the gain drift, are not actually physically realistic, so the value of our work there is limited. And any results we find will anyway be limited by the accuracy of your assumptions in generating the data. \n\nI wonder what this competition might have looked like if the full synthetic data generation pipeline (to the point that we can reproduce it) was released, along with a qualitative description of the changes made for the test set. ",
    "3294447": "Some feedback for next time:\n1. Pre-calibrate the data. I haven't seen any novel calibration methods that improve over the standard so far.\n2. Remove clipped/malformed transits and other artifacts that are not part of the mission.\n3. The metric should be more robust and less sensitive to outliers, though I don't have any specific suggestions off the top of my head.\n4. The data generation pipeline should be obfuscated as much as possible, reverse engineering your pipeline provides little scientific value.\n\nThanks for the competition! First time working on astronomy and learned a lot.",
    "3294061": "Another feedback. I gave it during the comp, but maybe it has been lost. Repeating it here.\n\nYou ask for sigmas of our predictions. This makes lot of sense.\n\nBut you did not provide the sigmas for the info in the star_info file. This makes no sense at all. And it looks unfair: you ask us for something you don't provide.\n\nYou know about the noise in that data. We can see that noise by applying Kepler 2nd law for instance. We also see it (and that was discussed in the forum) in the discrepancy bewtween the measure transit time and the predicted transit time.\n\nNext time, if any, please provide std for every physical value you provide.\n\n",
    "3293899": "Thank you very much for your organization! This was an absolutely wonderful game. Especially, I was deeply impressed by your meticulousness in dealing with physics!\n\nAt the same time, I am very much looking forward to seeing the rankings after the abnormal samples have been removed.",
    "3293816": "Right back at you -- thanks for organizing this competition; it's been a fun one!  And for this thread, which promises an interesting discussion.  Here are some thoughts that come to my mind after reading your post:\n * Don't feel bad about having regenerated the dataset after the competition.  I've seen organizers make the mistake of not fixing a problem that surfaced after the start of the competition, and in my opinion that screws things up far more than the need to re-process or re-tune an analysis.  In scientific work especially (as I'm sure you know), it's more important to adopt processes and mindsets that will guide you to the right answer in the end than it is to have everything right from the get-go, so kudos for bringing that practice to this competition.\n * the approach I took for my solution is a chisquare fit for each wavelength, whose outputs feed into a neural net that does smoothing and corrects for things like limb darkening and internal fit bias.  The fits, at least, are meant to look a lot like the way we'd do data analysis back when I was a grad student in collider physics, which is to say that, to me, my solution feels (deliberately) very conventional.  But with that chisquare fit approach, I found the runtime constraint to be *brutal*.  It was only a few days ago that I managed to get things to a state where I was confident my submission would run to completion in the allowed time, and as I write this, I still don't know if the submission I have running will produce a valid score or not. (The dreaded \"submission scoring error\" got my first entry.)  I don't mean to sound like I'm complaining; the time constraint did force me to really do some creative thinking and challenge myself on what was really necessary, so I do consider it a net positive.  But please don't make it any tighter next time; you probably don't want have a situation where potentially useful approaches have to get tossed out simply because they're not fast enough.  😅\n * In a fully conventional/traditional chisquare-fitting approach, one of the things I'd expect to have to do is binning over wavelengths: sum to get a light curve for that bin, then fit the light curve, and get an estimate of the depth parameter along with its fit error.  One of the nice things about that approach is that, so long as you don't choose overlapping bins, the error bars that you get for one bin are statistically independent from the next, in the sense that a random fluctuation in one wavelength only affects the one bin it's been assigned to.  That is a real benefit for whoever is doing downstream analysis on your results, because then they can take your fit depth measurements and the associated uncertainties and use them in a hypothesis test or whatever without having to worry about how to model correlations between the error bars of different bins.  I mention all that because the scoring metric for this competition forces participants to make a prediction for every wavelength, and doesn't enforce any statistical independence across wavelengths, so some solutions (like mine) will make it hard to interpret the spectrum (and in particular its error bars) in a downstream analysis.  Having an unusual optimization target like the one in this competition can really spur a lot of creativity, so again, not complaining.  But one of the things that might be an interesting element of a future competition would be how to interface with downstream analysis targets.  That could mean choosing a clever metric that somehow encourages/enforces statistical independence across wavelength or rewards smart binning choices, or it could mean pulling downstream analysis targets into the competition, e.g. ask participants to measure the concentration of a list of chemicals that might be present in the exoplanet atmospheres, or it could mean some other fun twist.\n\nThanks again -- looking forward to next time!",
    "3293837": ""
  }
}