{
  "id": 538139,
  "title": "Questions for us?",
  "url": "/competitions/ariel-data-challenge-2024/discussion/538139",
  "author_name": "Gordon Yip",
  "post_date": "2024-10-07T09:08:06.099000",
  "votes": 16,
  "comment_count": 56,
  "views": 0,
  "content": "<p>Since we are less than a month away from the end of this competition. Is there any question you might like to ask (about the competition, the mission or our work!?), but are too afraid to open a new topic in the discussion? Leave one here, and if we think we can help, we will answer them!</p>\n<p>Note: We cant guarantee we can answer everyone though! Sorry in advance. </p>\n<p>As always, good luck everyone! </p>",
  "messages": [
    {
      "id": 3008904,
      "postDate": "2024-10-07T09:08:06.100Z",
      "content": "<p>Since we are less than a month away from the end of this competition. Is there any question you might like to ask (about the competition, the mission or our work!?), but are too afraid to open a new topic in the discussion? Leave one here, and if we think we can help, we will answer them!</p>\n<p>Note: We cant guarantee we can answer everyone though! Sorry in advance. </p>\n<p>As always, good luck everyone! </p>",
      "rawMarkdown": "Since we are less than a month away from the end of this competition. Is there any question you might like to ask (about the competition, the mission or our work!?), but are too afraid to open a new topic in the discussion? Leave one here, and if we think we can help, we will answer them!\n\nNote: We cant guarantee we can answer everyone though! Sorry in advance. \n\nAs always, good luck everyone! ",
      "votes": 15
    },
    {
      "id": 3008972,
      "postDate": "2024-10-07T11:01:38.433Z",
      "content": "<p>I would like to know the current SOTA in terms of competition score, i.e., if you were given our dataset and applying your best current method, what would be your LB score? Can you reach 1? ('perfect' under the assumed noise). It would be even more interesting if you give the current SOTA both without compute limit and with this competition limit. </p>",
      "rawMarkdown": "I would like to know the current SOTA in terms of competition score, i.e., if you were given our dataset and applying your best current method, what would be your LB score? Can you reach 1? ('perfect' under the assumed noise). It would be even more interesting if you give the current SOTA both without compute limit and with this competition limit. ",
      "votes": 9,
      "replies": [
        {
          "id": 3009040,
          "postDate": "2024-10-07T12:57:46.623Z",
          "content": "<p>I came here to ask the same question. I tried to implement some of the methods from the articles published in March 2024, but got a bad score. But not too bad, it's still better than public notebooks. However, I'm not an expert at all, so I don't know where I could have gone wrong. Judging by the uncertainty estimates in the article, ours are pretty good.</p>",
          "rawMarkdown": "I came here to ask the same question. I tried to implement some of the methods from the articles published in March 2024, but got a bad score. But not too bad, it's still better than public notebooks. However, I'm not an expert at all, so I don't know where I could have gone wrong. Judging by the uncertainty estimates in the article, ours are pretty good.",
          "votes": 5,
          "replies": [
            {
              "id": 3012762,
              "postDate": "2024-10-09T10:55:22.480Z",
              "content": "<p>Would you mind disclosing the paper's title? Of course if you'd rather keep it to yourself that's fine :)</p>",
              "rawMarkdown": "Would you mind disclosing the paper's title? Of course if you'd rather keep it to yourself that's fine :)",
              "votes": 3
            },
            {
              "id": 3012764,
              "postDate": "2024-10-09T11:04:09.283Z",
              "content": "<p>It seems like this could be the second time I've accidentally raised the bar of a competition. This article <a href=\"https://arxiv.org/pdf/2402.15204\" target=\"_blank\">https://arxiv.org/pdf/2402.15204</a> sets the direction (2dGP + MCMC), details and implementation can be found in the list of references and links to repositories. </p>",
              "rawMarkdown": "It seems like this could be the second time I've accidentally raised the bar of a competition. This article https://arxiv.org/pdf/2402.15204 sets the direction (2dGP + MCMC), details and implementation can be found in the list of references and links to repositories. ",
              "votes": 13
            },
            {
              "id": 3012978,
              "postDate": "2024-10-09T14:44:39.690Z",
              "content": "<p>Thank you :)</p>",
              "rawMarkdown": "Thank you :)",
              "votes": 2
            },
            {
              "id": 3013366,
              "postDate": "2024-10-10T02:24:02.097Z",
              "content": "<p>In my opinion, the most challenging aspect of the data is that the light curves of different wavelengths vary greatly(i.e. the light curve is much noisier when the wavelength goes small). Can the 2dGP method really improve the estimation of different wavelengths or just the mean of all the wavelengths? Would you mind sharing some experiment results about it :)</p>",
              "rawMarkdown": "In my opinion, the most challenging aspect of the data is that the light curves of different wavelengths vary greatly(i.e. the light curve is much noisier when the wavelength goes small). Can the 2dGP method really improve the estimation of different wavelengths or just the mean of all the wavelengths? Would you mind sharing some experiment results about it :)"
            },
            {
              "id": 3013496,
              "postDate": "2024-10-10T07:50:34.357Z",
              "content": "<p>No, 2dGP does not solve THIS problem. I did not understand what \"the estimation\" is. What do you mean?</p>",
              "rawMarkdown": "No, 2dGP does not solve THIS problem. I did not understand what \"the estimation\" is. What do you mean?",
              "votes": 1
            },
            {
              "id": 3013799,
              "postDate": "2024-10-10T14:42:39.247Z",
              "content": "<p>Thanks for your reply! Sorry for my confusing terminology. I just wonder if you predict the transit depth of each wavelength or the mean transit depth of all the wavelengths(as your only_correlation notebook)? As the light curve with small wavelength is too noisy, I guess they don’t contribute to our prediction. Therefore, maybe 1d gaussian process without the wavelength dimension is enough?</p>",
              "rawMarkdown": "Thanks for your reply! Sorry for my confusing terminology. I just wonder if you predict the transit depth of each wavelength or the mean transit depth of all the wavelengths(as your only_correlation notebook)? As the light curve with small wavelength is too noisy, I guess they don’t contribute to our prediction. Therefore, maybe 1d gaussian process without the wavelength dimension is enough?"
            },
            {
              "id": 3013815,
              "postDate": "2024-10-10T14:56:19.210Z",
              "content": "<p>Yes, there is not much difference between the one-dimensional and two-dimensional version. The two-dimensional one is local enough to look for frequency-correlated noise only in the neighbors. Neighbors are usually close in terms of dynamic range and noise level. The 2dGP code written in jax is pretty fast. I liked that.</p>",
              "rawMarkdown": "Yes, there is not much difference between the one-dimensional and two-dimensional version. The two-dimensional one is local enough to look for frequency-correlated noise only in the neighbors. Neighbors are usually close in terms of dynamic range and noise level. The 2dGP code written in jax is pretty fast. I liked that."
            },
            {
              "id": 3014132,
              "postDate": "2024-10-11T00:56:55.010Z",
              "content": "<p>Thanks a lot!</p>",
              "rawMarkdown": "Thanks a lot!"
            },
            {
              "id": 3018544,
              "postDate": "2024-10-15T20:53:04.563Z",
              "content": "<p>Hi, another question if you don't mind. Did you make a submit of this method? Seems to me that on our dataset it takes a long long time for the method to complete. Or did you just check it locally and you estimated its performance based on that?</p>",
              "rawMarkdown": "Hi, another question if you don't mind. Did you make a submit of this method? Seems to me that on our dataset it takes a long long time for the method to complete. Or did you just check it locally and you estimated its performance based on that?"
            },
            {
              "id": 3018551,
              "postDate": "2024-10-15T20:58:56.210Z",
              "content": "<p>Yes, the second option. and yes, it takes a lot of time and requires unpleasant work with installing the necessary libraries. </p>",
              "rawMarkdown": "Yes, the second option. and yes, it takes a lot of time and requires unpleasant work with installing the necessary libraries. ",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3009129,
      "postDate": "2024-10-07T15:09:43.437Z",
      "content": "<p>Does the public test-set come from the same distribution as the private test-set? E.g. does public test-set have the set of stars as the training data while the private test-set has a different set? It may sound like revealing this would give a way too much, however I personally think this just levels the playing field as otherwise people will do extensive LB probing to get an advantage. If someone gambles that the public distribution is similar to the private distribution, they could get rewarded a lot, while someone that tries to build more conservative models from CV only and disregarding public LB, will get punished for thinking that public LB is a bad indicator. And since there is no skillful way to figure out whether the public test-set has in fact a similar distribution to the private test-set or not, I think it would fair to reveal it.</p>",
      "rawMarkdown": "Does the public test-set come from the same distribution as the private test-set? E.g. does public test-set have the set of stars as the training data while the private test-set has a different set? It may sound like revealing this would give a way too much, however I personally think this just levels the playing field as otherwise people will do extensive LB probing to get an advantage. If someone gambles that the public distribution is similar to the private distribution, they could get rewarded a lot, while someone that tries to build more conservative models from CV only and disregarding public LB, will get punished for thinking that public LB is a bad indicator. And since there is no skillful way to figure out whether the public test-set has in fact a similar distribution to the private test-set or not, I think it would fair to reveal it.",
      "votes": 5,
      "replies": [
        {
          "id": 3009176,
          "postDate": "2024-10-07T15:44:31.837Z",
          "content": "<p>I will second that, there is no problem in different distribution between the public and private test, but it's important to let us know (e.g. Ribonanza and Belka competitions both had significantly different distribution, but they were open about it and gave the details of the differences)</p>",
          "rawMarkdown": "I will second that, there is no problem in different distribution between the public and private test, but it's important to let us know (e.g. Ribonanza and Belka competitions both had significantly different distribution, but they were open about it and gave the details of the differences)",
          "votes": 2,
          "replies": [
            {
              "id": 3012774,
              "postDate": "2024-10-09T11:13:54.793Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/fritzcremer\" target=\"_blank\">@fritzcremer</a> and <a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a> , we are happy to confirm that the public and private test set are coming from the same distribution. The change in distribution only happens from training to test :)<br>\nwhat are the changes , you might ask - you can find more information in the references , especially in the proposal we made to NeurIPS. In a nutshell, we changed a few things in the simulation , including the characteristics of the telescope (not changing the whole design though!), atmospheric features, and in some cases, the host stars too. </p>",
              "rawMarkdown": "Hi @fritzcremer and @shlomoron , we are happy to confirm that the public and private test set are coming from the same distribution. The change in distribution only happens from training to test :)\nwhat are the changes , you might ask - you can find more information in the references , especially in the proposal we made to NeurIPS. In a nutshell, we changed a few things in the simulation , including the characteristics of the telescope (not changing the whole design though!), atmospheric features, and in some cases, the host stars too. ",
              "votes": 6
            },
            {
              "id": 3012838,
              "postDate": "2024-10-09T12:11:05.393Z",
              "content": "<p>oh and also the transit timing and transit duration, we changed that a bit to account for the fact that in some situations, the transit might not be very well centered (but we took care to exclude any grazing transit!), but only for some cases, not all. </p>",
              "rawMarkdown": "oh and also the transit timing and transit duration, we changed that a bit to account for the fact that in some situations, the transit might not be very well centered (but we took care to exclude any grazing transit!), but only for some cases, not all. ",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 3026687,
      "postDate": "2024-10-24T05:46:59.750Z",
      "content": "<p>The spectrum for the training data seems to include peaks of known gasses such as H2O and CH4 etc., but could you comment on whether we can assume the same for the test spectrum? Or could the test data include spectrums originating from non-existing virtual gasses with completely new spectrum?</p>",
      "rawMarkdown": "The spectrum for the training data seems to include peaks of known gasses such as H2O and CH4 etc., but could you comment on whether we can assume the same for the test spectrum? Or could the test data include spectrums originating from non-existing virtual gasses with completely new spectrum?",
      "votes": 3,
      "replies": [
        {
          "id": 3026827,
          "postDate": "2024-10-24T08:53:40.147Z",
          "content": "<p>You can assume that the test data will include new species, but they will always be coming from existing, known molecules, not non-existing virtual gases</p>",
          "rawMarkdown": "You can assume that the test data will include new species, but they will always be coming from existing, known molecules, not non-existing virtual gases",
          "votes": 2
        },
        {
          "id": 3026854,
          "postDate": "2024-10-24T09:32:59.280Z",
          "content": "<p>Thanks for the answer!</p>",
          "rawMarkdown": "Thanks for the answer!"
        },
        {
          "id": 3027207,
          "postDate": "2024-10-24T15:12:18.253Z",
          "content": "<p><a href=\"https://www.kaggle.com/gordonyip\" target=\"_blank\">@gordonyip</a> One more question, will new gas species appear for the private LB that did not appear in the public LB or the training data?</p>",
          "rawMarkdown": "@gordonyip One more question, will new gas species appear for the private LB that did not appear in the public LB or the training data?",
          "votes": 1
        }
      ]
    },
    {
      "id": 3012832,
      "postDate": "2024-10-09T12:08:12.947Z",
      "content": "<p>Can we assume that ingress and egress are symmetrical around the center?</p>",
      "rawMarkdown": "Can we assume that ingress and egress are symmetrical around the center?",
      "votes": 3,
      "replies": [
        {
          "id": 3012844,
          "postDate": "2024-10-09T12:15:08.213Z",
          "content": "<p>Yes you can. </p>",
          "rawMarkdown": "Yes you can. \n",
          "votes": 1,
          "replies": [
            {
              "id": 3013262,
              "postDate": "2024-10-09T19:49:46.120Z",
              "content": "<p>Isn't this contradicting your other statement?:<br>\n\"I can confirm that there are cases when some examples in the test is not as well-centered as others. If you look into the training set, there are also a few cases where it is not well-centered too.\"<br>\nOr do you mean that it is generally symmetrical, with some exceptions?</p>",
              "rawMarkdown": "Isn't this contradicting your other statement?:\n\"I can confirm that there are cases when some examples in the test is not as well-centered as others. If you look into the training set, there are also a few cases where it is not well-centered too.\"\nOr do you mean that it is generally symmetrical, with some exceptions?",
              "votes": 2
            },
            {
              "id": 3013851,
              "postDate": "2024-10-10T15:48:10.643Z",
              "content": "<p>I see where the confusion comes in, so by well centered i mean the planet is transitting the centre of the star right at the middle of the observation period (1/2 observing time, aka mid-transit time ), but that will not affect the ingress and egress moment, and they can still be symmetric, relative to the mid-transit time</p>",
              "rawMarkdown": "I see where the confusion comes in, so by well centered i mean the planet is transitting the centre of the star right at the middle of the observation period (1/2 observing time, aka mid-transit time ), but that will not affect the ingress and egress moment, and they can still be symmetric, relative to the mid-transit time"
            },
            {
              "id": 3013903,
              "postDate": "2024-10-10T17:06:51.747Z",
              "content": "<p><a href=\"https://www.kaggle.com/gordonyip\" target=\"_blank\">@gordonyip</a> Now I'm more confused, how is it possible for the time that 'the planet is transitting the centre of the star' not to be 'right at the middle of the observation period'? Should I imagine a case where observing the first half of the transmitting over the star takes  longer/shorter that the second half? How is that possible?</p>",
              "rawMarkdown": "@gordonyip Now I'm more confused, how is it possible for the time that 'the planet is transitting the centre of the star' not to be 'right at the middle of the observation period'? Should I imagine a case where observing the first half of the transmitting over the star takes  longer/shorter that the second half? How is that possible?"
            },
            {
              "id": 3014061,
              "postDate": "2024-10-10T20:33:44.577Z",
              "content": "<p><a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a> , let me try and explain this further- we define an observation period as the time spent to observe the star - believing that the transit will happen at some point. However, suppose the observation time is 7 hrs there is no guarantee that the mid-transit event (when the planet is transitting across the centre), aka. the deepest point of the transit depth, is going to happen at exactly 3.5 hours, it might be 3.4, or 3.3 or something else, it could be due to uncertainties related to the nature of the planet's orbit, the transit timing itself. I hope it helps!</p>",
              "rawMarkdown": "@shlomoron , let me try and explain this further- we define an observation period as the time spent to observe the star - believing that the transit will happen at some point. However, suppose the observation time is 7 hrs there is no guarantee that the mid-transit event (when the planet is transitting across the centre), aka. the deepest point of the transit depth, is going to happen at exactly 3.5 hours, it might be 3.4, or 3.3 or something else, it could be due to uncertainties related to the nature of the planet's orbit, the transit timing itself. I hope it helps!",
              "votes": 1
            },
            {
              "id": 3014127,
              "postDate": "2024-10-11T00:29:03.770Z",
              "content": "<p><a href=\"https://www.kaggle.com/gordonyip\" target=\"_blank\">@gordonyip</a> Ok I think I understand, so when you say \"the ingress and egress moment, and they can still be symmetric, relative to the mid-transit time\" you mean around the middle of the star (obviously) but together with the center of the star (mid-transit) they can still be asymmetric around the center of the observation, yes?<br>\nCan we at least assume that the ingress is left of the center and egress on the right? (i.e., that they are not shifted so far that both are in the first or second half of the observation)? </p>",
              "rawMarkdown": "@gordonyip Ok I think I understand, so when you say \"the ingress and egress moment, and they can still be symmetric, relative to the mid-transit time\" you mean around the middle of the star (obviously) but together with the center of the star (mid-transit) they can still be asymmetric around the center of the observation, yes?\nCan we at least assume that the ingress is left of the center and egress on the right? (i.e., that they are not shifted so far that both are in the first or second half of the observation)? ",
              "votes": 2
            },
            {
              "id": 3017909,
              "postDate": "2024-10-15T10:47:06.873Z",
              "content": "<p>Did you find answer to if we can assume the ingress is left of center and egress to the right?</p>",
              "rawMarkdown": "Did you find answer to if we can assume the ingress is left of center and egress to the right?"
            },
            {
              "id": 3017912,
              "postDate": "2024-10-15T10:55:38.977Z",
              "content": "<p>Hi sorry for the late reply - yes you can assume that ingress is left of center and egress is to the right</p>",
              "rawMarkdown": "Hi sorry for the late reply - yes you can assume that ingress is left of center and egress is to the right",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3009064,
      "postDate": "2024-10-07T13:25:43.623Z",
      "content": "<p>Where would I find some information or papers on what the readout noise is?  If I were to use it to generate a new observation, do I apply Gaussian noise to the signal frame with std equal to that frame, then go through all the steps outlined in the calibration notebook you all published, or do I apply it later in the process?</p>",
      "rawMarkdown": "Where would I find some information or papers on what the readout noise is?  If I were to use it to generate a new observation, do I apply Gaussian noise to the signal frame with std equal to that frame, then go through all the steps outlined in the calibration notebook you all published, or do I apply it later in the process?",
      "votes": 1,
      "replies": [
        {
          "id": 3010093,
          "postDate": "2024-10-08T16:46:18.673Z",
          "content": "<p>I found \"<em>ARIEL Performance Analysis Report</em>\" and \"<em>ArielRad: the Ariel radiometric model</em>\" to be insightful</p>",
          "rawMarkdown": "I found \"*ARIEL Performance Analysis Report*\" and \"*ArielRad: the Ariel radiometric model*\" to be insightful",
          "votes": 2,
          "replies": [
            {
              "id": 3012812,
              "postDate": "2024-10-09T11:55:14.997Z",
              "content": "<p>Yes those are nice resources if you want to learn about the noises :)</p>",
              "rawMarkdown": "Yes those are nice resources if you want to learn about the noises :)",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3027343,
      "postDate": "2024-10-24T17:34:35.363Z",
      "content": "<p>I have a question about how the wavelengths of the signal files map to the wavelengths of the labels.  I can see from the signal data that a slice from 39:321 is taken from the AIRS data and that is supposed to correspond to the first 282 wavelengths of the labels and that the FGS wavelength is the final 283 wavelength. Is my understanding correct? I question it because if I take the normalised flux as a function of wavelength and compare it to the label spectra the correlation is significantly better if the order of the wavelengths are reversed. Take the exoplanet at index 65 for example.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5760999%2F94892190cd2da465533047609b73db3d%2Fplanet65-not-flipped.png?generation=1729790797176652&amp;alt=media\" alt=\"current wavelength order\"></p>\n<p>But with the wavelengths flipped (reversed) this becomes<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5760999%2F80ccc38454cf91be5dd9ffa0b1dcbb3e%2Fexoplanet-65-flipped.png?generation=1729790838924472&amp;alt=media\" alt=\"wavelengths flipped\"></p>\n<p>I'm not doing anything special, I'm using the data processing from the Calibrating and Binning data workbook and I'm generating the light curves in a similar way to your started notebook, so I'm really struggling to make sense of this. </p>",
      "rawMarkdown": "I have a question about how the wavelengths of the signal files map to the wavelengths of the labels.  I can see from the signal data that a slice from 39:321 is taken from the AIRS data and that is supposed to correspond to the first 282 wavelengths of the labels and that the FGS wavelength is the final 283 wavelength. Is my understanding correct? I question it because if I take the normalised flux as a function of wavelength and compare it to the label spectra the correlation is significantly better if the order of the wavelengths are reversed. Take the exoplanet at index 65 for example.\n\n![current wavelength order](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5760999%2F94892190cd2da465533047609b73db3d%2Fplanet65-not-flipped.png?generation=1729790797176652&alt=media)\n\nBut with the wavelengths flipped (reversed) this becomes\n![wavelengths flipped](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5760999%2F80ccc38454cf91be5dd9ffa0b1dcbb3e%2Fexoplanet-65-flipped.png?generation=1729790838924472&alt=media)\n\nI'm not doing anything special, I'm using the data processing from the Calibrating and Binning data workbook and I'm generating the light curves in a similar way to your started notebook, so I'm really struggling to make sense of this. \n",
      "votes": 2,
      "replies": [
        {
          "id": 3027392,
          "postDate": "2024-10-24T18:34:03.343Z",
          "content": "<p>Yes, you have to reverse your predictions. I found this from here: <a href=\"https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/529412\" target=\"_blank\">https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/529412</a></p>\n<p>And FGS prediction should correspond to <code>wl_1</code> and <code>sigma_1</code> as described here: <a href=\"https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/527039\" target=\"_blank\">https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/527039</a></p>",
          "rawMarkdown": "Yes, you have to reverse your predictions. I found this from here: https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/529412\n\nAnd FGS prediction should correspond to `wl_1` and `sigma_1` as described here: https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/527039",
          "votes": 2
        }
      ]
    },
    {
      "id": 3025482,
      "postDate": "2024-10-22T20:15:46.067Z",
      "content": "<p>It looks like all train data are provided with uniform limb darkening modeling. Can we assume that this is the case also for the hidden dataset?</p>",
      "rawMarkdown": "It looks like all train data are provided with uniform limb darkening modeling. Can we assume that this is the case also for the hidden dataset?",
      "votes": 2,
      "replies": [
        {
          "id": 3025485,
          "postDate": "2024-10-22T20:26:04.637Z",
          "content": "<p>Yes you can !</p>",
          "rawMarkdown": "Yes you can !",
          "votes": 1
        }
      ]
    },
    {
      "id": 3018510,
      "postDate": "2024-10-15T20:02:17.863Z",
      "content": "<p>Can I ask why read noise data (read.parquet) was not included in the example of data process code ? They are the same for each planet, is it the reason that they are not useful? </p>",
      "rawMarkdown": "Can I ask why read noise data (read.parquet) was not included in the example of data process code ? They are the same for each planet, is it the reason that they are not useful? ",
      "votes": 2,
      "replies": [
        {
          "id": 3024124,
          "postDate": "2024-10-21T11:05:11.980Z",
          "content": "<p>read noise cannot be removed from the data for it's stocastic nature. The \"calibration\" file is used as an estimation of the expected noise from instrument on top of the photon noise. Hope it helps!</p>",
          "rawMarkdown": "read noise cannot be removed from the data for it's stocastic nature. The \"calibration\" file is used as an estimation of the expected noise from instrument on top of the photon noise. Hope it helps!",
          "votes": 1
        }
      ]
    },
    {
      "id": 3013254,
      "postDate": "2024-10-09T19:32:22.657Z",
      "content": "<p>This question is to anyone, really. I’m still not entirely sure on the use of the FGS in the context of the competition, besides being used for the prediction of one of the (Rp/Rs)^2 at exactly one wavelength. I understand it’s role for the mission (I think?) but not to the competition.</p>\n<p>Is there more to it? I know it’s essentially analysing a different wavelength range than the AIRS-CH0, but I fail to understand how exactly does it come into play in terms of the competition.</p>\n<p>My apologies if this has been asked/answered before, must’ve missed it, I promise I looked around!</p>",
      "rawMarkdown": "This question is to anyone, really. I’m still not entirely sure on the use of the FGS in the context of the competition, besides being used for the prediction of one of the (Rp/Rs)^2 at exactly one wavelength. I understand it’s role for the mission (I think?) but not to the competition.\n\nIs there more to it? I know it’s essentially analysing a different wavelength range than the AIRS-CH0, but I fail to understand how exactly does it come into play in terms of the competition.\n\nMy apologies if this has been asked/answered before, must’ve missed it, I promise I looked around!\n\n",
      "votes": 2,
      "replies": [
        {
          "id": 3013853,
          "postDate": "2024-10-10T15:51:28.510Z",
          "content": "<p>Indeed, if you are only looking at the target spectrum, the FGS channel only contributes to one out of 283 channels, which, you can say, not as important to AIRS-CH0 for scoring the metric. However, the FGS channels might provide some pointers to the jitter noise that is always present within the dataset, but we are not sure how visible that might be. </p>",
          "rawMarkdown": "Indeed, if you are only looking at the target spectrum, the FGS channel only contributes to one out of 283 channels, which, you can say, not as important to AIRS-CH0 for scoring the metric. However, the FGS channels might provide some pointers to the jitter noise that is always present within the dataset, but we are not sure how visible that might be. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 3013233,
      "postDate": "2024-10-09T19:07:30.723Z",
      "content": "<p>If you be on first day in competition what you with to you?</p>",
      "rawMarkdown": "If you be on first day in competition what you with to you?",
      "votes": 1
    },
    {
      "id": 3026274,
      "postDate": "2024-10-23T15:27:47.723Z",
      "content": "<p>Hi fellow Kagglers, </p>\n<p>I am new to Kaggle and I recently got a quota exceeded on GPU utilization. I am not even able to submit my notebook without GPU as my code uses cupy extensibly. DOes anyone knows a hack to extend Kaggle quota on GPUs? Any help would be highly appreciated</p>",
      "rawMarkdown": "Hi fellow Kagglers, \n\nI am new to Kaggle and I recently got a quota exceeded on GPU utilization. I am not even able to submit my notebook without GPU as my code uses cupy extensibly. DOes anyone knows a hack to extend Kaggle quota on GPUs? Any help would be highly appreciated",
      "votes": -2,
      "replies": [
        {
          "id": 3026928,
          "postDate": "2024-10-24T11:00:06.080Z",
          "content": "<p>There's no way to extend GPU quota on Kaggle.</p>",
          "rawMarkdown": "There's no way to extend GPU quota on Kaggle."
        }
      ]
    },
    {
      "id": 3027217,
      "postDate": "2024-10-24T15:19:54.607Z",
      "content": "<p>Hi, I have a question here regarding this function in the starter solution:</p>\n<p>def NN_uncertainity(model, x_test, targets_abs_max, T=5):<br>\n    predictions = []<br>\n    for _ in range(T):<br>\n        pred_norm = model.predict([x_test],verbose=0)<br>\n        pred = targets_norm_back(pred_norm, targets_abs_max)<br>\n        predictions += [pred]  <br>\n    mean, std = np.mean(np.array(predictions), axis=0), np.std(np.array(predictions), axis=0)<br>\n    return mean, std</p>\n<p>the model.perdict does not give MC_dropout prediction, i.e., all the predictions are the same, the std is 0. Does anyone know how to enable the mc_dropout in predict? It only happens in training.</p>\n<p>I googled this solution:<br>\nimport keras.backend as K</p>\n<p>class KerasDropoutPrediction(object):<br>\n    def <strong>init</strong>(self,model):<br>\n        self.f = K.function(<br>\n                [model.layers[0].input, <br>\n                 K.learning_phase()],<br>\n                [model.layers[-1].output])<br>\n    def predict(self,x, n_iter=10):<br>\n        result = []<br>\n        for _ in range(n_iter):<br>\n            result.append(self.f([x , 1]))<br>\n        result = np.array(result).reshape(n_iter,len(x)).T<br>\n        return result</p>\n<p>Howerver, there is no function in latest keras.backend. Hope this community can solve this issue:<br>\nAttributeError: module 'keras.backend' has no attribute 'function'</p>\n<p>thanks</p>",
      "rawMarkdown": "Hi, I have a question here regarding this function in the starter solution:\n\ndef NN_uncertainity(model, x_test, targets_abs_max, T=5):\n    predictions = []\n    for _ in range(T):\n        pred_norm = model.predict([x_test],verbose=0)\n        pred = targets_norm_back(pred_norm, targets_abs_max)\n        predictions += [pred]  \n    mean, std = np.mean(np.array(predictions), axis=0), np.std(np.array(predictions), axis=0)\n    return mean, std\n\nthe model.perdict does not give MC_dropout prediction, i.e., all the predictions are the same, the std is 0. Does anyone know how to enable the mc_dropout in predict? It only happens in training.\n\nI googled this solution:\nimport keras.backend as K\n\nclass KerasDropoutPrediction(object):\n    def __init__(self,model):\n        self.f = K.function(\n                [model.layers[0].input, \n                 K.learning_phase()],\n                [model.layers[-1].output])\n    def predict(self,x, n_iter=10):\n        result = []\n        for _ in range(n_iter):\n            result.append(self.f([x , 1]))\n        result = np.array(result).reshape(n_iter,len(x)).T\n        return result\n\nHowerver, there is no function in latest keras.backend. Hope this community can solve this issue:\nAttributeError: module 'keras.backend' has no attribute 'function'\n\nthanks"
    },
    {
      "id": 3024469,
      "postDate": "2024-10-21T17:12:46.710Z",
      "content": "<p>What can be the reasons for scoring error? From what I understand the following:</p>\n<ol>\n<li>Negative predictions</li>\n<li>non numeric predictions</li>\n<li>Incorrect number of columns</li>\n<li>No submission file created</li>\n<li>Wrong number of columns ( 566+1)</li>\n</ol>\n<p>If there is anything else, can you please add here. I have got this error and it runs fine on train but in test this gives this error.</p>",
      "rawMarkdown": "What can be the reasons for scoring error? From what I understand the following:\n1. Negative predictions\n2. non numeric predictions\n3. Incorrect number of columns\n4. No submission file created\n5. Wrong number of columns ( 566+1)\n\nIf there is anything else, can you please add here. I have got this error and it runs fine on train but in test this gives this error.",
      "replies": [
        {
          "id": 3024481,
          "postDate": "2024-10-21T17:19:41.670Z",
          "content": "<p>I have a model uploaded i am using to score. is it possible the code is not able to access the model by any chance?</p>",
          "rawMarkdown": "I have a model uploaded i am using to score. is it possible the code is not able to access the model by any chance?"
        },
        {
          "id": 3024484,
          "postDate": "2024-10-21T17:24:37.247Z",
          "content": "<p>submission file name : submission.csv</p>",
          "rawMarkdown": "submission file name : submission.csv"
        }
      ]
    },
    {
      "id": 3021518,
      "postDate": "2024-10-18T15:28:53.053Z",
      "content": "<p>Sorry if i missed it already in the discussion before but what is the size of the hidden test set ( in terms of number of planets) . Is it more or less than train set ( ideally should be lesser than train). an estimate would be great.</p>",
      "rawMarkdown": "Sorry if i missed it already in the discussion before but what is the size of the hidden test set ( in terms of number of planets) . Is it more or less than train set ( ideally should be lesser than train). an estimate would be great.",
      "replies": [
        {
          "id": 3021613,
          "postDate": "2024-10-18T17:09:38.897Z",
          "content": "<p>Approximately 800</p>\n<p><a href=\"https://www.kaggle.com/competitions/ariel-data-challenge-2024/data\" target=\"_blank\">https://www.kaggle.com/competitions/ariel-data-challenge-2024/data</a></p>",
          "rawMarkdown": "Approximately 800\n\nhttps://www.kaggle.com/competitions/ariel-data-challenge-2024/data"
        }
      ]
    },
    {
      "id": 3019932,
      "postDate": "2024-10-17T03:11:02.087Z",
      "content": "<p>Can you provide some data on absorption cross-sections of common gas moleculars? （for example those listed in <a href=\"https://iopscience.iop.org/article/10.3847/1538-3881/aaaf75）Hard\" target=\"_blank\">https://iopscience.iop.org/article/10.3847/1538-3881/aaaf75）Hard</a> to collect them for ones unfamilair with this field.</p>",
      "rawMarkdown": "Can you provide some data on absorption cross-sections of common gas moleculars? （for example those listed in https://iopscience.iop.org/article/10.3847/1538-3881/aaaf75）Hard to collect them for ones unfamilair with this field.",
      "replies": [
        {
          "id": 3024122,
          "postDate": "2024-10-21T11:04:21.713Z",
          "content": "<p>You can try ExoMol, <a href=\"https://www.exomol.com/data/\" target=\"_blank\">https://www.exomol.com/data/</a> and hitran <a href=\"https://hitran.org/\" target=\"_blank\">https://hitran.org/</a>, to download the molecular data</p>",
          "rawMarkdown": "You can try ExoMol, https://www.exomol.com/data/ and hitran https://hitran.org/, to download the molecular data",
          "votes": 1
        }
      ]
    },
    {
      "id": 3018000,
      "postDate": "2024-10-15T12:48:35.233Z",
      "content": "<p>Hi,</p>\n<p>I’m new to Kaggle and have a question about the 9-hour time limit for notebooks. I understand that the organizers provide preprocessed data that we can use directly. However, if I choose to do my own preprocessing on the data (e.g., calibrating or removing trends), does this additional preprocessing time also need to fit within the 9-hour limit along with the training time?</p>\n<p>Additionally, is it possible to have two separate notebooks—one for preprocessing and another for training—or must everything be done within a single notebook?</p>",
      "rawMarkdown": "Hi,\n\nI’m new to Kaggle and have a question about the 9-hour time limit for notebooks. I understand that the organizers provide preprocessed data that we can use directly. However, if I choose to do my own preprocessing on the data (e.g., calibrating or removing trends), does this additional preprocessing time also need to fit within the 9-hour limit along with the training time?\n\nAdditionally, is it possible to have two separate notebooks—one for preprocessing and another for training—or must everything be done within a single notebook?",
      "replies": [
        {
          "id": 3019074,
          "postDate": "2024-10-16T08:40:06.960Z",
          "content": "<p>Hi Matias, Indeed the preprocessing time has to be counted within the training time. There are several Jax based notebook which completed the cleaning within 20mins, you might want to have a look there. <br>\nAs for your second question - I dont know but other Kagglers should be able to help!</p>",
          "rawMarkdown": "Hi Matias, Indeed the preprocessing time has to be counted within the training time. There are several Jax based notebook which completed the cleaning within 20mins, you might want to have a look there. \nAs for your second question - I dont know but other Kagglers should be able to help!",
          "replies": [
            {
              "id": 3019104,
              "postDate": "2024-10-16T09:02:44.213Z",
              "content": "<p>Yes, you can have separate preprocessing and training notebooks. Many people do this thing when preprocessing takes a lot of time.</p>",
              "rawMarkdown": "Yes, you can have separate preprocessing and training notebooks. Many people do this thing when preprocessing takes a lot of time."
            }
          ]
        }
      ]
    },
    {
      "id": 3008975,
      "postDate": "2024-10-07T11:08:30.903Z",
      "content": "<p>I wonder if the present leaderboard score is calculated only using the revealed test set's labels or calculated across the entire test set's labels including hidden ones.<br>\nIf the previous one is correct, will the score be calculated across the entire test set's labels after deadline?<br>\nIf there are already answers to this same question, sorry in advance and please give a link below.</p>",
      "rawMarkdown": "I wonder if the present leaderboard score is calculated only using the revealed test set's labels or calculated across the entire test set's labels including hidden ones.\nIf the previous one is correct, will the score be calculated across the entire test set's labels after deadline?\nIf there are already answers to this same question, sorry in advance and please give a link below.",
      "replies": [
        {
          "id": 3009121,
          "postDate": "2024-10-07T14:59:11.020Z",
          "content": "<p>Right now the score is only calculated on a certain fraction of the test set (called public LB score). After the deadline, the score for the hidden test-set is revealed (private LB score), which will the only score relevant for the final standings. So we already are making predictions for the hidden test-set, however the scores for this set are only shown after the deadline.</p>",
          "rawMarkdown": "Right now the score is only calculated on a certain fraction of the test set (called public LB score). After the deadline, the score for the hidden test-set is revealed (private LB score), which will the only score relevant for the final standings. So we already are making predictions for the hidden test-set, however the scores for this set are only shown after the deadline.",
          "votes": 2
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3008972,
      "author_name": "greySnow",
      "author_url": "",
      "post_date": "2024-10-07T11:01:38.433000",
      "content": "<p>I would like to know the current SOTA in terms of competition score, i.e., if you were given our dataset and applying your best current method, what would be your LB score? Can you reach 1? ('perfect' under the assumed noise). It would be even more interesting if you give the current SOTA both without compute limit and with this competition limit. </p>",
      "votes": 9,
      "replies": [
        {
          "id": 3009040,
          "author_name": "Sergei Fironov",
          "author_url": "",
          "post_date": "2024-10-07T12:57:46.623000",
          "content": "<p>I came here to ask the same question. I tried to implement some of the methods from the articles published in March 2024, but got a bad score. But not too bad, it's still better than public notebooks. However, I'm not an expert at all, so I don't know where I could have gone wrong. Judging by the uncertainty estimates in the article, ours are pretty good.</p>",
          "votes": 5,
          "replies": [
            {
              "id": 3012762,
              "author_name": "GabeTheHuman",
              "author_url": "",
              "post_date": "2024-10-09T10:55:22.480000",
              "content": "<p>Would you mind disclosing the paper's title? Of course if you'd rather keep it to yourself that's fine :)</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 3012764,
              "author_name": "Sergei Fironov",
              "author_url": "",
              "post_date": "2024-10-09T11:04:09.283000",
              "content": "<p>It seems like this could be the second time I've accidentally raised the bar of a competition. This article <a href=\"https://arxiv.org/pdf/2402.15204\" target=\"_blank\">https://arxiv.org/pdf/2402.15204</a> sets the direction (2dGP + MCMC), details and implementation can be found in the list of references and links to repositories. </p>",
              "votes": 13,
              "replies": []
            },
            {
              "id": 3012978,
              "author_name": "GabeTheHuman",
              "author_url": "",
              "post_date": "2024-10-09T14:44:39.690000",
              "content": "<p>Thank you :)</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3013366,
              "author_name": "Zhu Siqi",
              "author_url": "",
              "post_date": "2024-10-10T02:24:02.097000",
              "content": "<p>In my opinion, the most challenging aspect of the data is that the light curves of different wavelengths vary greatly(i.e. the light curve is much noisier when the wavelength goes small). Can the 2dGP method really improve the estimation of different wavelengths or just the mean of all the wavelengths? Would you mind sharing some experiment results about it :)</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3013496,
              "author_name": "Sergei Fironov",
              "author_url": "",
              "post_date": "2024-10-10T07:50:34.357000",
              "content": "<p>No, 2dGP does not solve THIS problem. I did not understand what \"the estimation\" is. What do you mean?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3013799,
              "author_name": "Zhu Siqi",
              "author_url": "",
              "post_date": "2024-10-10T14:42:39.247000",
              "content": "<p>Thanks for your reply! Sorry for my confusing terminology. I just wonder if you predict the transit depth of each wavelength or the mean transit depth of all the wavelengths(as your only_correlation notebook)? As the light curve with small wavelength is too noisy, I guess they don’t contribute to our prediction. Therefore, maybe 1d gaussian process without the wavelength dimension is enough?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3013815,
              "author_name": "Sergei Fironov",
              "author_url": "",
              "post_date": "2024-10-10T14:56:19.210000",
              "content": "<p>Yes, there is not much difference between the one-dimensional and two-dimensional version. The two-dimensional one is local enough to look for frequency-correlated noise only in the neighbors. Neighbors are usually close in terms of dynamic range and noise level. The 2dGP code written in jax is pretty fast. I liked that.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3014132,
              "author_name": "Zhu Siqi",
              "author_url": "",
              "post_date": "2024-10-11T00:56:55.010000",
              "content": "<p>Thanks a lot!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3018544,
              "author_name": "GabeTheHuman",
              "author_url": "",
              "post_date": "2024-10-15T20:53:04.563000",
              "content": "<p>Hi, another question if you don't mind. Did you make a submit of this method? Seems to me that on our dataset it takes a long long time for the method to complete. Or did you just check it locally and you estimated its performance based on that?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3018551,
              "author_name": "Sergei Fironov",
              "author_url": "",
              "post_date": "2024-10-15T20:58:56.210000",
              "content": "<p>Yes, the second option. and yes, it takes a lot of time and requires unpleasant work with installing the necessary libraries. </p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3009129,
      "author_name": "Fritz Cremer",
      "author_url": "",
      "post_date": "2024-10-07T15:09:43.437000",
      "content": "<p>Does the public test-set come from the same distribution as the private test-set? E.g. does public test-set have the set of stars as the training data while the private test-set has a different set? It may sound like revealing this would give a way too much, however I personally think this just levels the playing field as otherwise people will do extensive LB probing to get an advantage. If someone gambles that the public distribution is similar to the private distribution, they could get rewarded a lot, while someone that tries to build more conservative models from CV only and disregarding public LB, will get punished for thinking that public LB is a bad indicator. And since there is no skillful way to figure out whether the public test-set has in fact a similar distribution to the private test-set or not, I think it would fair to reveal it.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 3009176,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2024-10-07T15:44:31.837000",
          "content": "<p>I will second that, there is no problem in different distribution between the public and private test, but it's important to let us know (e.g. Ribonanza and Belka competitions both had significantly different distribution, but they were open about it and gave the details of the differences)</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3012774,
              "author_name": "Gordon Yip",
              "author_url": "",
              "post_date": "2024-10-09T11:13:54.793000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/fritzcremer\" target=\"_blank\">@fritzcremer</a> and <a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a> , we are happy to confirm that the public and private test set are coming from the same distribution. The change in distribution only happens from training to test :)<br>\nwhat are the changes , you might ask - you can find more information in the references , especially in the proposal we made to NeurIPS. In a nutshell, we changed a few things in the simulation , including the characteristics of the telescope (not changing the whole design though!), atmospheric features, and in some cases, the host stars too. </p>",
              "votes": 6,
              "replies": []
            },
            {
              "id": 3012838,
              "author_name": "Gordon Yip",
              "author_url": "",
              "post_date": "2024-10-09T12:11:05.393000",
              "content": "<p>oh and also the transit timing and transit duration, we changed that a bit to account for the fact that in some situations, the transit might not be very well centered (but we took care to exclude any grazing transit!), but only for some cases, not all. </p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3026687,
      "author_name": "c-number",
      "author_url": "",
      "post_date": "2024-10-24T05:46:59.750000",
      "content": "<p>The spectrum for the training data seems to include peaks of known gasses such as H2O and CH4 etc., but could you comment on whether we can assume the same for the test spectrum? Or could the test data include spectrums originating from non-existing virtual gasses with completely new spectrum?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3026827,
          "author_name": "Gordon Yip",
          "author_url": "",
          "post_date": "2024-10-24T08:53:40.147000",
          "content": "<p>You can assume that the test data will include new species, but they will always be coming from existing, known molecules, not non-existing virtual gases</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 3026854,
          "author_name": "c-number",
          "author_url": "",
          "post_date": "2024-10-24T09:32:59.280000",
          "content": "<p>Thanks for the answer!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3027207,
          "author_name": "c-number",
          "author_url": "",
          "post_date": "2024-10-24T15:12:18.253000",
          "content": "<p><a href=\"https://www.kaggle.com/gordonyip\" target=\"_blank\">@gordonyip</a> One more question, will new gas species appear for the private LB that did not appear in the public LB or the training data?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3012832,
      "author_name": "DennisSakva",
      "author_url": "",
      "post_date": "2024-10-09T12:08:12.947000",
      "content": "<p>Can we assume that ingress and egress are symmetrical around the center?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3012844,
          "author_name": "Gordon Yip",
          "author_url": "",
          "post_date": "2024-10-09T12:15:08.213000",
          "content": "<p>Yes you can. </p>",
          "votes": 1,
          "replies": [
            {
              "id": 3013262,
              "author_name": "Fritz Cremer",
              "author_url": "",
              "post_date": "2024-10-09T19:49:46.120000",
              "content": "<p>Isn't this contradicting your other statement?:<br>\n\"I can confirm that there are cases when some examples in the test is not as well-centered as others. If you look into the training set, there are also a few cases where it is not well-centered too.\"<br>\nOr do you mean that it is generally symmetrical, with some exceptions?</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3013851,
              "author_name": "Gordon Yip",
              "author_url": "",
              "post_date": "2024-10-10T15:48:10.643000",
              "content": "<p>I see where the confusion comes in, so by well centered i mean the planet is transitting the centre of the star right at the middle of the observation period (1/2 observing time, aka mid-transit time ), but that will not affect the ingress and egress moment, and they can still be symmetric, relative to the mid-transit time</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3013903,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-10-10T17:06:51.747000",
              "content": "<p><a href=\"https://www.kaggle.com/gordonyip\" target=\"_blank\">@gordonyip</a> Now I'm more confused, how is it possible for the time that 'the planet is transitting the centre of the star' not to be 'right at the middle of the observation period'? Should I imagine a case where observing the first half of the transmitting over the star takes  longer/shorter that the second half? How is that possible?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3014061,
              "author_name": "Gordon Yip",
              "author_url": "",
              "post_date": "2024-10-10T20:33:44.577000",
              "content": "<p><a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a> , let me try and explain this further- we define an observation period as the time spent to observe the star - believing that the transit will happen at some point. However, suppose the observation time is 7 hrs there is no guarantee that the mid-transit event (when the planet is transitting across the centre), aka. the deepest point of the transit depth, is going to happen at exactly 3.5 hours, it might be 3.4, or 3.3 or something else, it could be due to uncertainties related to the nature of the planet's orbit, the transit timing itself. I hope it helps!</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3014127,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-10-11T00:29:03.770000",
              "content": "<p><a href=\"https://www.kaggle.com/gordonyip\" target=\"_blank\">@gordonyip</a> Ok I think I understand, so when you say \"the ingress and egress moment, and they can still be symmetric, relative to the mid-transit time\" you mean around the middle of the star (obviously) but together with the center of the star (mid-transit) they can still be asymmetric around the center of the observation, yes?<br>\nCan we at least assume that the ingress is left of the center and egress on the right? (i.e., that they are not shifted so far that both are in the first or second half of the observation)? </p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3017909,
              "author_name": "highDopamine",
              "author_url": "",
              "post_date": "2024-10-15T10:47:06.873000",
              "content": "<p>Did you find answer to if we can assume the ingress is left of center and egress to the right?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3017912,
              "author_name": "Gordon Yip",
              "author_url": "",
              "post_date": "2024-10-15T10:55:38.977000",
              "content": "<p>Hi sorry for the late reply - yes you can assume that ingress is left of center and egress is to the right</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3009064,
      "author_name": "Daniel Phalen",
      "author_url": "",
      "post_date": "2024-10-07T13:25:43.623000",
      "content": "<p>Where would I find some information or papers on what the readout noise is?  If I were to use it to generate a new observation, do I apply Gaussian noise to the signal frame with std equal to that frame, then go through all the steps outlined in the calibration notebook you all published, or do I apply it later in the process?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3010093,
          "author_name": "Tord Malmgren",
          "author_url": "",
          "post_date": "2024-10-08T16:46:18.673000",
          "content": "<p>I found \"<em>ARIEL Performance Analysis Report</em>\" and \"<em>ArielRad: the Ariel radiometric model</em>\" to be insightful</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3012812,
              "author_name": "Gordon Yip",
              "author_url": "",
              "post_date": "2024-10-09T11:55:14.997000",
              "content": "<p>Yes those are nice resources if you want to learn about the noises :)</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3027343,
      "author_name": "Gareth Davies",
      "author_url": "",
      "post_date": "2024-10-24T17:34:35.363000",
      "content": "<p>I have a question about how the wavelengths of the signal files map to the wavelengths of the labels.  I can see from the signal data that a slice from 39:321 is taken from the AIRS data and that is supposed to correspond to the first 282 wavelengths of the labels and that the FGS wavelength is the final 283 wavelength. Is my understanding correct? I question it because if I take the normalised flux as a function of wavelength and compare it to the label spectra the correlation is significantly better if the order of the wavelengths are reversed. Take the exoplanet at index 65 for example.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5760999%2F94892190cd2da465533047609b73db3d%2Fplanet65-not-flipped.png?generation=1729790797176652&amp;alt=media\" alt=\"current wavelength order\"></p>\n<p>But with the wavelengths flipped (reversed) this becomes<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5760999%2F80ccc38454cf91be5dd9ffa0b1dcbb3e%2Fexoplanet-65-flipped.png?generation=1729790838924472&amp;alt=media\" alt=\"wavelengths flipped\"></p>\n<p>I'm not doing anything special, I'm using the data processing from the Calibrating and Binning data workbook and I'm generating the light curves in a similar way to your started notebook, so I'm really struggling to make sense of this. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 3027392,
          "author_name": "Oleh Kivernyk",
          "author_url": "",
          "post_date": "2024-10-24T18:34:03.343000",
          "content": "<p>Yes, you have to reverse your predictions. I found this from here: <a href=\"https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/529412\" target=\"_blank\">https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/529412</a></p>\n<p>And FGS prediction should correspond to <code>wl_1</code> and <code>sigma_1</code> as described here: <a href=\"https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/527039\" target=\"_blank\">https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/527039</a></p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 3025482,
      "author_name": "Oleh Kivernyk",
      "author_url": "",
      "post_date": "2024-10-22T20:15:46.067000",
      "content": "<p>It looks like all train data are provided with uniform limb darkening modeling. Can we assume that this is the case also for the hidden dataset?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3025485,
          "author_name": "Gordon Yip",
          "author_url": "",
          "post_date": "2024-10-22T20:26:04.637000",
          "content": "<p>Yes you can !</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3018510,
      "author_name": "John",
      "author_url": "",
      "post_date": "2024-10-15T20:02:17.863000",
      "content": "<p>Can I ask why read noise data (read.parquet) was not included in the example of data process code ? They are the same for each planet, is it the reason that they are not useful? </p>",
      "votes": 2,
      "replies": [
        {
          "id": 3024124,
          "author_name": "Gordon Yip",
          "author_url": "",
          "post_date": "2024-10-21T11:05:11.980000",
          "content": "<p>read noise cannot be removed from the data for it's stocastic nature. The \"calibration\" file is used as an estimation of the expected noise from instrument on top of the photon noise. Hope it helps!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3013254,
      "author_name": "Andrei Zamfir",
      "author_url": "",
      "post_date": "2024-10-09T19:32:22.657000",
      "content": "<p>This question is to anyone, really. I’m still not entirely sure on the use of the FGS in the context of the competition, besides being used for the prediction of one of the (Rp/Rs)^2 at exactly one wavelength. I understand it’s role for the mission (I think?) but not to the competition.</p>\n<p>Is there more to it? I know it’s essentially analysing a different wavelength range than the AIRS-CH0, but I fail to understand how exactly does it come into play in terms of the competition.</p>\n<p>My apologies if this has been asked/answered before, must’ve missed it, I promise I looked around!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3013853,
          "author_name": "Gordon Yip",
          "author_url": "",
          "post_date": "2024-10-10T15:51:28.510000",
          "content": "<p>Indeed, if you are only looking at the target spectrum, the FGS channel only contributes to one out of 283 channels, which, you can say, not as important to AIRS-CH0 for scoring the metric. However, the FGS channels might provide some pointers to the jitter noise that is always present within the dataset, but we are not sure how visible that might be. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3013233,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-10-09T19:07:30.723000",
      "content": "<p>If you be on first day in competition what you with to you?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3026274,
      "author_name": "Dristi Chatterjee",
      "author_url": "",
      "post_date": "2024-10-23T15:27:47.723000",
      "content": "<p>Hi fellow Kagglers, </p>\n<p>I am new to Kaggle and I recently got a quota exceeded on GPU utilization. I am not even able to submit my notebook without GPU as my code uses cupy extensibly. DOes anyone knows a hack to extend Kaggle quota on GPUs? Any help would be highly appreciated</p>",
      "votes": -2,
      "replies": [
        {
          "id": 3026928,
          "author_name": "Araik Tamazian",
          "author_url": "",
          "post_date": "2024-10-24T11:00:06.080000",
          "content": "<p>There's no way to extend GPU quota on Kaggle.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3027217,
      "author_name": "John",
      "author_url": "",
      "post_date": "2024-10-24T15:19:54.607000",
      "content": "<p>Hi, I have a question here regarding this function in the starter solution:</p>\n<p>def NN_uncertainity(model, x_test, targets_abs_max, T=5):<br>\n    predictions = []<br>\n    for _ in range(T):<br>\n        pred_norm = model.predict([x_test],verbose=0)<br>\n        pred = targets_norm_back(pred_norm, targets_abs_max)<br>\n        predictions += [pred]  <br>\n    mean, std = np.mean(np.array(predictions), axis=0), np.std(np.array(predictions), axis=0)<br>\n    return mean, std</p>\n<p>the model.perdict does not give MC_dropout prediction, i.e., all the predictions are the same, the std is 0. Does anyone know how to enable the mc_dropout in predict? It only happens in training.</p>\n<p>I googled this solution:<br>\nimport keras.backend as K</p>\n<p>class KerasDropoutPrediction(object):<br>\n    def <strong>init</strong>(self,model):<br>\n        self.f = K.function(<br>\n                [model.layers[0].input, <br>\n                 K.learning_phase()],<br>\n                [model.layers[-1].output])<br>\n    def predict(self,x, n_iter=10):<br>\n        result = []<br>\n        for _ in range(n_iter):<br>\n            result.append(self.f([x , 1]))<br>\n        result = np.array(result).reshape(n_iter,len(x)).T<br>\n        return result</p>\n<p>Howerver, there is no function in latest keras.backend. Hope this community can solve this issue:<br>\nAttributeError: module 'keras.backend' has no attribute 'function'</p>\n<p>thanks</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3024469,
      "author_name": "Dristi Chatterjee",
      "author_url": "",
      "post_date": "2024-10-21T17:12:46.710000",
      "content": "<p>What can be the reasons for scoring error? From what I understand the following:</p>\n<ol>\n<li>Negative predictions</li>\n<li>non numeric predictions</li>\n<li>Incorrect number of columns</li>\n<li>No submission file created</li>\n<li>Wrong number of columns ( 566+1)</li>\n</ol>\n<p>If there is anything else, can you please add here. I have got this error and it runs fine on train but in test this gives this error.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3024481,
          "author_name": "Dristi Chatterjee",
          "author_url": "",
          "post_date": "2024-10-21T17:19:41.670000",
          "content": "<p>I have a model uploaded i am using to score. is it possible the code is not able to access the model by any chance?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3024484,
          "author_name": "Dristi Chatterjee",
          "author_url": "",
          "post_date": "2024-10-21T17:24:37.247000",
          "content": "<p>submission file name : submission.csv</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3021518,
      "author_name": "Dristi Chatterjee",
      "author_url": "",
      "post_date": "2024-10-18T15:28:53.053000",
      "content": "<p>Sorry if i missed it already in the discussion before but what is the size of the hidden test set ( in terms of number of planets) . Is it more or less than train set ( ideally should be lesser than train). an estimate would be great.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3021613,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2024-10-18T17:09:38.897000",
          "content": "<p>Approximately 800</p>\n<p><a href=\"https://www.kaggle.com/competitions/ariel-data-challenge-2024/data\" target=\"_blank\">https://www.kaggle.com/competitions/ariel-data-challenge-2024/data</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3019932,
      "author_name": "Ruby",
      "author_url": "",
      "post_date": "2024-10-17T03:11:02.087000",
      "content": "<p>Can you provide some data on absorption cross-sections of common gas moleculars? （for example those listed in <a href=\"https://iopscience.iop.org/article/10.3847/1538-3881/aaaf75）Hard\" target=\"_blank\">https://iopscience.iop.org/article/10.3847/1538-3881/aaaf75）Hard</a> to collect them for ones unfamilair with this field.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3024122,
          "author_name": "Gordon Yip",
          "author_url": "",
          "post_date": "2024-10-21T11:04:21.713000",
          "content": "<p>You can try ExoMol, <a href=\"https://www.exomol.com/data/\" target=\"_blank\">https://www.exomol.com/data/</a> and hitran <a href=\"https://hitran.org/\" target=\"_blank\">https://hitran.org/</a>, to download the molecular data</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3018000,
      "author_name": "Matias TorresR",
      "author_url": "",
      "post_date": "2024-10-15T12:48:35.233000",
      "content": "<p>Hi,</p>\n<p>I’m new to Kaggle and have a question about the 9-hour time limit for notebooks. I understand that the organizers provide preprocessed data that we can use directly. However, if I choose to do my own preprocessing on the data (e.g., calibrating or removing trends), does this additional preprocessing time also need to fit within the 9-hour limit along with the training time?</p>\n<p>Additionally, is it possible to have two separate notebooks—one for preprocessing and another for training—or must everything be done within a single notebook?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3019074,
          "author_name": "Gordon Yip",
          "author_url": "",
          "post_date": "2024-10-16T08:40:06.960000",
          "content": "<p>Hi Matias, Indeed the preprocessing time has to be counted within the training time. There are several Jax based notebook which completed the cleaning within 20mins, you might want to have a look there. <br>\nAs for your second question - I dont know but other Kagglers should be able to help!</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3019104,
              "author_name": "Araik Tamazian",
              "author_url": "",
              "post_date": "2024-10-16T09:02:44.213000",
              "content": "<p>Yes, you can have separate preprocessing and training notebooks. Many people do this thing when preprocessing takes a lot of time.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3008975,
      "author_name": "skhan kim",
      "author_url": "",
      "post_date": "2024-10-07T11:08:30.903000",
      "content": "<p>I wonder if the present leaderboard score is calculated only using the revealed test set's labels or calculated across the entire test set's labels including hidden ones.<br>\nIf the previous one is correct, will the score be calculated across the entire test set's labels after deadline?<br>\nIf there are already answers to this same question, sorry in advance and please give a link below.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3009121,
          "author_name": "Fritz Cremer",
          "author_url": "",
          "post_date": "2024-10-07T14:59:11.020000",
          "content": "<p>Right now the score is only calculated on a certain fraction of the test set (called public LB score). After the deadline, the score for the hidden test-set is revealed (private LB score), which will the only score relevant for the final standings. So we already are making predictions for the hidden test-set, however the scores for this set are only shown after the deadline.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3008904": "Since we are less than a month away from the end of this competition. Is there any question you might like to ask (about the competition, the mission or our work!?), but are too afraid to open a new topic in the discussion? Leave one here, and if we think we can help, we will answer them!\n\nNote: We cant guarantee we can answer everyone though! Sorry in advance. \n\nAs always, good luck everyone! ",
    "3008972": "I would like to know the current SOTA in terms of competition score, i.e., if you were given our dataset and applying your best current method, what would be your LB score? Can you reach 1? ('perfect' under the assumed noise). It would be even more interesting if you give the current SOTA both without compute limit and with this competition limit. ",
    "3009129": "Does the public test-set come from the same distribution as the private test-set? E.g. does public test-set have the set of stars as the training data while the private test-set has a different set? It may sound like revealing this would give a way too much, however I personally think this just levels the playing field as otherwise people will do extensive LB probing to get an advantage. If someone gambles that the public distribution is similar to the private distribution, they could get rewarded a lot, while someone that tries to build more conservative models from CV only and disregarding public LB, will get punished for thinking that public LB is a bad indicator. And since there is no skillful way to figure out whether the public test-set has in fact a similar distribution to the private test-set or not, I think it would fair to reveal it.",
    "3026687": "The spectrum for the training data seems to include peaks of known gasses such as H2O and CH4 etc., but could you comment on whether we can assume the same for the test spectrum? Or could the test data include spectrums originating from non-existing virtual gasses with completely new spectrum?",
    "3012832": "Can we assume that ingress and egress are symmetrical around the center?",
    "3009064": "Where would I find some information or papers on what the readout noise is?  If I were to use it to generate a new observation, do I apply Gaussian noise to the signal frame with std equal to that frame, then go through all the steps outlined in the calibration notebook you all published, or do I apply it later in the process?",
    "3027343": "I have a question about how the wavelengths of the signal files map to the wavelengths of the labels.  I can see from the signal data that a slice from 39:321 is taken from the AIRS data and that is supposed to correspond to the first 282 wavelengths of the labels and that the FGS wavelength is the final 283 wavelength. Is my understanding correct? I question it because if I take the normalised flux as a function of wavelength and compare it to the label spectra the correlation is significantly better if the order of the wavelengths are reversed. Take the exoplanet at index 65 for example.\n\n![current wavelength order](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5760999%2F94892190cd2da465533047609b73db3d%2Fplanet65-not-flipped.png?generation=1729790797176652&alt=media)\n\nBut with the wavelengths flipped (reversed) this becomes\n![wavelengths flipped](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5760999%2F80ccc38454cf91be5dd9ffa0b1dcbb3e%2Fexoplanet-65-flipped.png?generation=1729790838924472&alt=media)\n\nI'm not doing anything special, I'm using the data processing from the Calibrating and Binning data workbook and I'm generating the light curves in a similar way to your started notebook, so I'm really struggling to make sense of this. \n",
    "3025482": "It looks like all train data are provided with uniform limb darkening modeling. Can we assume that this is the case also for the hidden dataset?",
    "3018510": "Can I ask why read noise data (read.parquet) was not included in the example of data process code ? They are the same for each planet, is it the reason that they are not useful? ",
    "3013254": "This question is to anyone, really. I’m still not entirely sure on the use of the FGS in the context of the competition, besides being used for the prediction of one of the (Rp/Rs)^2 at exactly one wavelength. I understand it’s role for the mission (I think?) but not to the competition.\n\nIs there more to it? I know it’s essentially analysing a different wavelength range than the AIRS-CH0, but I fail to understand how exactly does it come into play in terms of the competition.\n\nMy apologies if this has been asked/answered before, must’ve missed it, I promise I looked around!\n\n",
    "3013233": "If you be on first day in competition what you with to you?",
    "3026274": "Hi fellow Kagglers, \n\nI am new to Kaggle and I recently got a quota exceeded on GPU utilization. I am not even able to submit my notebook without GPU as my code uses cupy extensibly. DOes anyone knows a hack to extend Kaggle quota on GPUs? Any help would be highly appreciated",
    "3027217": "Hi, I have a question here regarding this function in the starter solution:\n\ndef NN_uncertainity(model, x_test, targets_abs_max, T=5):\n    predictions = []\n    for _ in range(T):\n        pred_norm = model.predict([x_test],verbose=0)\n        pred = targets_norm_back(pred_norm, targets_abs_max)\n        predictions += [pred]  \n    mean, std = np.mean(np.array(predictions), axis=0), np.std(np.array(predictions), axis=0)\n    return mean, std\n\nthe model.perdict does not give MC_dropout prediction, i.e., all the predictions are the same, the std is 0. Does anyone know how to enable the mc_dropout in predict? It only happens in training.\n\nI googled this solution:\nimport keras.backend as K\n\nclass KerasDropoutPrediction(object):\n    def __init__(self,model):\n        self.f = K.function(\n                [model.layers[0].input, \n                 K.learning_phase()],\n                [model.layers[-1].output])\n    def predict(self,x, n_iter=10):\n        result = []\n        for _ in range(n_iter):\n            result.append(self.f([x , 1]))\n        result = np.array(result).reshape(n_iter,len(x)).T\n        return result\n\nHowerver, there is no function in latest keras.backend. Hope this community can solve this issue:\nAttributeError: module 'keras.backend' has no attribute 'function'\n\nthanks",
    "3024469": "What can be the reasons for scoring error? From what I understand the following:\n1. Negative predictions\n2. non numeric predictions\n3. Incorrect number of columns\n4. No submission file created\n5. Wrong number of columns ( 566+1)\n\nIf there is anything else, can you please add here. I have got this error and it runs fine on train but in test this gives this error.",
    "3021518": "Sorry if i missed it already in the discussion before but what is the size of the hidden test set ( in terms of number of planets) . Is it more or less than train set ( ideally should be lesser than train). an estimate would be great.",
    "3019932": "Can you provide some data on absorption cross-sections of common gas moleculars? （for example those listed in https://iopscience.iop.org/article/10.3847/1538-3881/aaaf75）Hard to collect them for ones unfamilair with this field.",
    "3018000": "Hi,\n\nI’m new to Kaggle and have a question about the 9-hour time limit for notebooks. I understand that the organizers provide preprocessed data that we can use directly. However, if I choose to do my own preprocessing on the data (e.g., calibrating or removing trends), does this additional preprocessing time also need to fit within the 9-hour limit along with the training time?\n\nAdditionally, is it possible to have two separate notebooks—one for preprocessing and another for training—or must everything be done within a single notebook?",
    "3008975": "I wonder if the present leaderboard score is calculated only using the revealed test set's labels or calculated across the entire test set's labels including hidden ones.\nIf the previous one is correct, will the score be calculated across the entire test set's labels after deadline?\nIf there are already answers to this same question, sorry in advance and please give a link below."
  }
}