{
  "id": 524184,
  "title": "How to choice of sigma* - L_ideal is important - (GLL) will get zero scores [Host answered]",
  "url": "/competitions/ariel-data-challenge-2024/discussion/524184",
  "author_name": "SeshuRaju 🧘‍♂️",
  "post_date": "2024-08-05T03:02:52.572000",
  "votes": 2,
  "comment_count": 1,
  "views": 0,
  "content": "<h1><strong>How to choice the sigma for each wavelength ?</strong></h1>\n<h3><strong>L_ideal</strong> as the case where the submission perfectly matches the ground truth values, with an uncertainty of 10 parts per million (ppm) otherwise ignored the prediction leads to score zero for the exoplanet's respective target.</h3>\n<h3>Any resources to understand this target better ?</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F1aad11469c4abec2a77ab5021a054f78%2FScreenshot%202024-08-05%20at%208.33.41AM.png?generation=1722827038461923&amp;alt=media\" alt=\"\"></p>\n<hr>\n<blockquote>\n  <p><strong>if the participants are able to pinpoint exactly the transit depths (intensities) in the test set, the best strategy is to submit their network with 10ppm as uncertainties. This conclusion builds on the assumption that the model can predict the transit depths with a very high degree of accuracy. Given our understanding of the noise within the dataset, there will be information loss (due to the noise), which may prevent the model from predicting the transit depths perfectly (at least in the test set)</strong>.</p>\n</blockquote>\n<p>Also, note that 10ppm is a very narrow value for uncertainty, and their model will be heavily penalized if the predictions are outside the uncertainty range (a characteristic of the GLL function), resulting in a worse score. <a href=\"https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/524003#2950765\" target=\"_blank\">Discussion</a></p>",
  "messages": [
    {
      "id": 2947635,
      "postDate": "2024-08-05T10:35:11.027Z",
      "content": "<p><a href=\"https://www.kaggle.com/gordonyip\" target=\"_blank\">@gordonyip</a>, any resource to understand about Sigma* target for each wavelength? </p>",
      "rawMarkdown": "@gordonyip, any resource to understand about Sigma* target for each wavelength? ",
      "votes": 2
    },
    {
      "id": 2947170,
      "postDate": "2024-08-05T03:02:52.573Z",
      "content": "<h1><strong>How to choice the sigma for each wavelength ?</strong></h1>\n<h3><strong>L_ideal</strong> as the case where the submission perfectly matches the ground truth values, with an uncertainty of 10 parts per million (ppm) otherwise ignored the prediction leads to score zero for the exoplanet's respective target.</h3>\n<h3>Any resources to understand this target better ?</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F1aad11469c4abec2a77ab5021a054f78%2FScreenshot%202024-08-05%20at%208.33.41AM.png?generation=1722827038461923&amp;alt=media\" alt=\"\"></p>\n<hr>\n<blockquote>\n  <p><strong>if the participants are able to pinpoint exactly the transit depths (intensities) in the test set, the best strategy is to submit their network with 10ppm as uncertainties. This conclusion builds on the assumption that the model can predict the transit depths with a very high degree of accuracy. Given our understanding of the noise within the dataset, there will be information loss (due to the noise), which may prevent the model from predicting the transit depths perfectly (at least in the test set)</strong>.</p>\n</blockquote>\n<p>Also, note that 10ppm is a very narrow value for uncertainty, and their model will be heavily penalized if the predictions are outside the uncertainty range (a characteristic of the GLL function), resulting in a worse score. <a href=\"https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/524003#2950765\" target=\"_blank\">Discussion</a></p>",
      "rawMarkdown": "# **How to choice the sigma for each wavelength ?**\n### **L_ideal** as the case where the submission perfectly matches the ground truth values, with an uncertainty of 10 parts per million (ppm) otherwise ignored the prediction leads to score zero for the exoplanet's respective target.\n### Any resources to understand this target better ?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F1aad11469c4abec2a77ab5021a054f78%2FScreenshot%202024-08-05%20at%208.33.41AM.png?generation=1722827038461923&alt=media)\n\n---\n\n> **if the participants are able to pinpoint exactly the transit depths (intensities) in the test set, the best strategy is to submit their network with 10ppm as uncertainties. This conclusion builds on the assumption that the model can predict the transit depths with a very high degree of accuracy. Given our understanding of the noise within the dataset, there will be information loss (due to the noise), which may prevent the model from predicting the transit depths perfectly (at least in the test set)**.\n\nAlso, note that 10ppm is a very narrow value for uncertainty, and their model will be heavily penalized if the predictions are outside the uncertainty range (a characteristic of the GLL function), resulting in a worse score. [Discussion](https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/524003#2950765)",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 2947635,
      "author_name": "SeshuRaju 🧘‍♂️",
      "author_url": "",
      "post_date": "2024-08-05T10:35:11.027000",
      "content": "<p><a href=\"https://www.kaggle.com/gordonyip\" target=\"_blank\">@gordonyip</a>, any resource to understand about Sigma* target for each wavelength? </p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2947635": "@gordonyip, any resource to understand about Sigma* target for each wavelength? ",
    "2947170": "# **How to choice the sigma for each wavelength ?**\n### **L_ideal** as the case where the submission perfectly matches the ground truth values, with an uncertainty of 10 parts per million (ppm) otherwise ignored the prediction leads to score zero for the exoplanet's respective target.\n### Any resources to understand this target better ?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F1aad11469c4abec2a77ab5021a054f78%2FScreenshot%202024-08-05%20at%208.33.41AM.png?generation=1722827038461923&alt=media)\n\n---\n\n> **if the participants are able to pinpoint exactly the transit depths (intensities) in the test set, the best strategy is to submit their network with 10ppm as uncertainties. This conclusion builds on the assumption that the model can predict the transit depths with a very high degree of accuracy. Given our understanding of the noise within the dataset, there will be information loss (due to the noise), which may prevent the model from predicting the transit depths perfectly (at least in the test set)**.\n\nAlso, note that 10ppm is a very narrow value for uncertainty, and their model will be heavily penalized if the predictions are outside the uncertainty range (a characteristic of the GLL function), resulting in a worse score. [Discussion](https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/524003#2950765)"
  }
}