{
  "id": 524287,
  "title": "Welcome to Ariel Data Challenge - Resources and notebooks",
  "url": "/competitions/ariel-data-challenge-2024/discussion/524287",
  "author_name": "Gordon Yip",
  "post_date": "2024-08-05T14:55:20.493000",
  "votes": 74,
  "comment_count": 34,
  "views": 0,
  "content": "<p>Our belated welcome to all fellow challengers. Thank you for your interest in the Ariel Data Challenge🔭. In this challenge, we are asking you to process raw observations into a data product ready for further analysis (transmission spectrum).  </p>\n<p>We understand that for some of you, this could be your first time dealing with astronomical problems, let alone exoplanets. </p>\n<p>To help you get started, we have compiled a list of materials. We will be updating the list if we think there are any knowledge gaps we should address:</p>\n<ol>\n<li>What is a transit? : A simple explanation <a href=\"https://exoplanets.nasa.gov/alien-worlds/ways-to-find-a-planet/?intent=021#/2\" target=\"_blank\"> here</a>: </li>\n<li>What is 'transit depth' and a spectrum? -- The Exoplanet Handbook (Perryman 2018) gives a good overview of transit and spectroscopy (Chap. 6, Sec. 6.12 and Chap. 11, 11.6), alternatively, have a look at <a href=\"https://arxiv.org/pdf/1001.2010v5\" target=\"_blank\">Winn (2010) </a></li>\n<li>What is an observation? What is a light curve? -- The <a href=\"https://www.ariel-datachallenge.space/workshop2024/documentation/about\" target=\"_blank\">Science</a> session gives a nice introduction to the topic, of course, the Ariel redbook is useful on that as well.</li>\n<li>Publications that could be useful -- some of you may have posted these already - <a href=\"https://www.ariel-datachallenge.space/ML/documentation/resources\" target=\"_blank\">link</a></li>\n<li>Differences between Train - test distribution. Long story short, yes, certain aspects of the test data distribution are deliberately made different from the training set. So don't be too surprised if your local score is different from your leaderboard score. We follow the same philosophy as outlined <a href=\"https://proceedings.mlr.press/v220/yip23a/yip23a.pdf\" target=\"_blank\">here</a>, section 2.2.1</li>\n<li><a href=\"https://www.kaggle.com/code/gordonyip/calibrating-astronomical-data\" target=\"_blank\">Notebook</a> on Calibrating Data by Batch , <a href=\"https://www.kaggle.com/code/gordonyip/calibrating-a-single-observation\" target=\"_blank\">Notebook</a> on calibrating a single observation.</li>\n<li><a href=\"https://www.kaggle.com/code/gordonyip/adc-2024-starter-solution\" target=\"_blank\">Starter solution</a> - there are also plenty notebooks available in the code section.</li>\n<li>Follow us on <a href=\"https://x.com/ArielTelescope\" target=\"_blank\">Twitter/X</a> - we will be sharing our updates there, stay tuned!</li>\n<li>Follow us on <a href=\"https://www.youtube.com/channel/UCMLTUdXBPNS_pDJcGNZoLYw\" target=\"_blank\">YouTube</a></li>\n</ol>\n<p>Of course, they are just the tip of an iceberg, and let us know if there are any specific aspect of the problem that you want to know. We are happy update our references. We are happy to answer any questions you may have (as long as they don't expose the test set, of course!), and I will be monitoring the discussion as much as I can. Please feel free to tag me.</p>\n<p><strong>While domain knowledge is important, please remember that it only represents what we know so far. There are phenomena within the data that we are unaware of or have not accounted for. Hence, this is why the challenge is set up - to explore new ways to denoise the data with fresh new minds like yours</strong></p>\n<p>Last but not least, please have fun! We are eager to find out what you can come up with!</p>",
  "messages": [
    {
      "id": 2947916,
      "postDate": "2024-08-05T14:55:20.493Z",
      "content": "<p>Our belated welcome to all fellow challengers. Thank you for your interest in the Ariel Data Challenge🔭. In this challenge, we are asking you to process raw observations into a data product ready for further analysis (transmission spectrum).  </p>\n<p>We understand that for some of you, this could be your first time dealing with astronomical problems, let alone exoplanets. </p>\n<p>To help you get started, we have compiled a list of materials. We will be updating the list if we think there are any knowledge gaps we should address:</p>\n<ol>\n<li>What is a transit? : A simple explanation <a href=\"https://exoplanets.nasa.gov/alien-worlds/ways-to-find-a-planet/?intent=021#/2\" target=\"_blank\"> here</a>: </li>\n<li>What is 'transit depth' and a spectrum? -- The Exoplanet Handbook (Perryman 2018) gives a good overview of transit and spectroscopy (Chap. 6, Sec. 6.12 and Chap. 11, 11.6), alternatively, have a look at <a href=\"https://arxiv.org/pdf/1001.2010v5\" target=\"_blank\">Winn (2010) </a></li>\n<li>What is an observation? What is a light curve? -- The <a href=\"https://www.ariel-datachallenge.space/workshop2024/documentation/about\" target=\"_blank\">Science</a> session gives a nice introduction to the topic, of course, the Ariel redbook is useful on that as well.</li>\n<li>Publications that could be useful -- some of you may have posted these already - <a href=\"https://www.ariel-datachallenge.space/ML/documentation/resources\" target=\"_blank\">link</a></li>\n<li>Differences between Train - test distribution. Long story short, yes, certain aspects of the test data distribution are deliberately made different from the training set. So don't be too surprised if your local score is different from your leaderboard score. We follow the same philosophy as outlined <a href=\"https://proceedings.mlr.press/v220/yip23a/yip23a.pdf\" target=\"_blank\">here</a>, section 2.2.1</li>\n<li><a href=\"https://www.kaggle.com/code/gordonyip/calibrating-astronomical-data\" target=\"_blank\">Notebook</a> on Calibrating Data by Batch , <a href=\"https://www.kaggle.com/code/gordonyip/calibrating-a-single-observation\" target=\"_blank\">Notebook</a> on calibrating a single observation.</li>\n<li><a href=\"https://www.kaggle.com/code/gordonyip/adc-2024-starter-solution\" target=\"_blank\">Starter solution</a> - there are also plenty notebooks available in the code section.</li>\n<li>Follow us on <a href=\"https://x.com/ArielTelescope\" target=\"_blank\">Twitter/X</a> - we will be sharing our updates there, stay tuned!</li>\n<li>Follow us on <a href=\"https://www.youtube.com/channel/UCMLTUdXBPNS_pDJcGNZoLYw\" target=\"_blank\">YouTube</a></li>\n</ol>\n<p>Of course, they are just the tip of an iceberg, and let us know if there are any specific aspect of the problem that you want to know. We are happy update our references. We are happy to answer any questions you may have (as long as they don't expose the test set, of course!), and I will be monitoring the discussion as much as I can. Please feel free to tag me.</p>\n<p><strong>While domain knowledge is important, please remember that it only represents what we know so far. There are phenomena within the data that we are unaware of or have not accounted for. Hence, this is why the challenge is set up - to explore new ways to denoise the data with fresh new minds like yours</strong></p>\n<p>Last but not least, please have fun! We are eager to find out what you can come up with!</p>",
      "rawMarkdown": "Our belated welcome to all fellow challengers. Thank you for your interest in the Ariel Data Challenge🔭. In this challenge, we are asking you to process raw observations into a data product ready for further analysis (transmission spectrum).  \n\nWe understand that for some of you, this could be your first time dealing with astronomical problems, let alone exoplanets. \n\nTo help you get started, we have compiled a list of materials. We will be updating the list if we think there are any knowledge gaps we should address:\n1. What is a transit? : A simple explanation [ here]( https://exoplanets.nasa.gov/alien-worlds/ways-to-find-a-planet/?intent=021#/2): \n2. What is 'transit depth' and a spectrum? -- The Exoplanet Handbook (Perryman 2018) gives a good overview of transit and spectroscopy (Chap. 6, Sec. 6.12 and Chap. 11, 11.6), alternatively, have a look at [Winn (2010) ](https://arxiv.org/pdf/1001.2010v5)\n3. What is an observation? What is a light curve? -- The [Science](https://www.ariel-datachallenge.space/workshop2024/documentation/about) session gives a nice introduction to the topic, of course, the Ariel redbook is useful on that as well.\n4. Publications that could be useful -- some of you may have posted these already - [link](https://www.ariel-datachallenge.space/ML/documentation/resources)\n5. Differences between Train - test distribution. Long story short, yes, certain aspects of the test data distribution are deliberately made different from the training set. So don't be too surprised if your local score is different from your leaderboard score. We follow the same philosophy as outlined [here](https://proceedings.mlr.press/v220/yip23a/yip23a.pdf), section 2.2.1\n6. [Notebook](https://www.kaggle.com/code/gordonyip/calibrating-astronomical-data) on Calibrating Data by Batch , [Notebook](https://www.kaggle.com/code/gordonyip/calibrating-a-single-observation) on calibrating a single observation.\n7. [Starter solution](https://www.kaggle.com/code/gordonyip/adc-2024-starter-solution) - there are also plenty notebooks available in the code section.\n8. Follow us on [Twitter/X](https://x.com/ArielTelescope) - we will be sharing our updates there, stay tuned!\n9. Follow us on [YouTube](https://www.youtube.com/channel/UCMLTUdXBPNS_pDJcGNZoLYw)\n\n\nOf course, they are just the tip of an iceberg, and let us know if there are any specific aspect of the problem that you want to know. We are happy update our references. We are happy to answer any questions you may have (as long as they don't expose the test set, of course!), and I will be monitoring the discussion as much as I can. Please feel free to tag me.\n\n**While domain knowledge is important, please remember that it only represents what we know so far. There are phenomena within the data that we are unaware of or have not accounted for. Hence, this is why the challenge is set up - to explore new ways to denoise the data with fresh new minds like yours**\n\nLast but not least, please have fun! We are eager to find out what you can come up with!",
      "votes": 73
    },
    {
      "id": 3009652,
      "postDate": "2024-10-08T07:13:50.850Z",
      "content": "<p>Why don't we release the precise score below 0.? It becomes very hard to debug with bad cases hidden …</p>",
      "rawMarkdown": "Why don't we release the precise score below 0.? It becomes very hard to debug with bad cases hidden ...",
      "votes": 3,
      "replies": [
        {
          "id": 3012819,
          "postDate": "2024-10-09T12:03:20.310Z",
          "content": "<p>sorry to hear that, but i am afraid there is not much we can do here. When we setup the 0th point on the system (since the GLL function does not have a bounded maximum), we tried to be as lenient as possible to include as much solution as possible, but obviously, that also means some solutions wont get any scores. </p>",
          "rawMarkdown": "sorry to hear that, but i am afraid there is not much we can do here. When we setup the 0th point on the system (since the GLL function does not have a bounded maximum), we tried to be as lenient as possible to include as much solution as possible, but obviously, that also means some solutions wont get any scores. ",
          "votes": 1,
          "replies": [
            {
              "id": 3024826,
              "postDate": "2024-10-22T04:00:37.873Z",
              "content": "<p>But it's possible, If yes how it can be. </p>",
              "rawMarkdown": "But it's possible, If yes how it can be. \n"
            }
          ]
        },
        {
          "id": 3034267,
          "postDate": "2024-11-02T00:17:01.653Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 3034268,
          "postDate": "2024-11-02T00:18:16.580Z",
          "content": "<p>It definitely makes debugging more challenging when bad cases aren't visible. Having detailed feedback could help participants identify issues in their solutions more effectively. It might be worth suggesting that the scoring system be adjusted to provide more transparency, even if it means just releasing additional insights or metrics that could guide debugging efforts. Your input could be valuable in improving the overall experience for everyone!</p>",
          "rawMarkdown": " It definitely makes debugging more challenging when bad cases aren't visible. Having detailed feedback could help participants identify issues in their solutions more effectively. It might be worth suggesting that the scoring system be adjusted to provide more transparency, even if it means just releasing additional insights or metrics that could guide debugging efforts. Your input could be valuable in improving the overall experience for everyone!"
        }
      ]
    },
    {
      "id": 2950984,
      "postDate": "2024-08-08T06:02:39.167Z",
      "content": "<p>1) From the four groups in the test data (Set 1, …, Set 4), can you confirm if they are all in the public and private LB?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F311485%2F44fbf3ea84c4bcfe07756f789d1b41b3%2FScreenshot%202024-08-08%20at%2007.55.43.png?generation=1723096587393487&amp;alt=media\"></p>\n<p>2) Are there any details available on how the simulations were run? </p>\n<p>Thanks a lot in advance!</p>",
      "rawMarkdown": "1) From the four groups in the test data (Set 1, ..., Set 4), can you confirm if they are all in the public and private LB?\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F311485%2F44fbf3ea84c4bcfe07756f789d1b41b3%2FScreenshot%202024-08-08%20at%2007.55.43.png?generation=1723096587393487&alt=media\" width=500>\n\n2) Are there any details available on how the simulations were run? \n\nThanks a lot in advance!",
      "votes": 6
    },
    {
      "id": 2968069,
      "postDate": "2024-08-23T13:51:37.510Z",
      "content": "<p><a href=\"https://www.kaggle.com/gordonyip\" target=\"_blank\">@gordonyip</a> Thank you for the additional information. I would like to ask a question that might be quite basic. Upon a review of the data, I don't see any difference between the training and test sets. I would like to ask what the target feature of our research is and how to identify it in the training set. Unfortunately, I couldn't find this information, if I'm not mistaken. Thank you in advance.</p>",
      "rawMarkdown": "@gordonyip Thank you for the additional information. I would like to ask a question that might be quite basic. Upon a review of the data, I don't see any difference between the training and test sets. I would like to ask what the target feature of our research is and how to identify it in the training set. Unfortunately, I couldn't find this information, if I'm not mistaken. Thank you in advance.",
      "votes": 1,
      "replies": [
        {
          "id": 2972948,
          "postDate": "2024-08-29T02:48:19.793Z",
          "content": "<p>train_labels.csv </p>",
          "rawMarkdown": "train_labels.csv ",
          "replies": [
            {
              "id": 2974172,
              "postDate": "2024-08-30T11:22:34.510Z",
              "content": "<p>Ok, thank you! I am starting to understend better.</p>\n<p>Another one question, which information contain file axis_info.parquet? In description an explanation if very poore. I woold like to understand meaning of each column in this file.</p>",
              "rawMarkdown": "Ok, thank you! I am starting to understend better.\n\nAnother one question, which information contain file axis_info.parquet? In description an explanation if very poore. I woold like to understand meaning of each column in this file."
            },
            {
              "id": 2978955,
              "postDate": "2024-09-04T12:30:47.740Z",
              "content": "<p>We will add more description to the file, but essentially, h means hour (for time axis), um mean micrometer (spectral axis), and other simply means spatial axis</p>",
              "rawMarkdown": "We will add more description to the file, but essentially, h means hour (for time axis), um mean micrometer (spectral axis), and other simply means spatial axis",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2951278,
      "postDate": "2024-08-08T13:31:24.030Z",
      "content": "<p>Are we allowed to use pretrained models from Kaggle models or Hugging Face, or do we have to build our own model from scratch?</p>",
      "rawMarkdown": "Are we allowed to use pretrained models from Kaggle models or Hugging Face, or do we have to build our own model from scratch?",
      "votes": 1,
      "replies": [
        {
          "id": 2953637,
          "postDate": "2024-08-08T22:07:35.137Z",
          "content": "<p>yes, you are allowed to use open source models </p>",
          "rawMarkdown": "yes, you are allowed to use open source models ",
          "votes": 3
        },
        {
          "id": 2975271,
          "postDate": "2024-08-31T17:08:37.543Z",
          "content": "<p>Hi Evan, If you do not mind, can I merge with your team?</p>",
          "rawMarkdown": "Hi Evan, If you do not mind, can I merge with your team?"
        }
      ]
    },
    {
      "id": 2962291,
      "postDate": "2024-08-17T12:05:26.217Z",
      "content": "<p>mark，翻译一下<br>\n我们对所有参赛者表示迟来的欢迎。感谢你们对Ariel数据挑战赛🔭的兴趣。在这个挑战中，我们要求你们将原始观测数据处理成一个准备好进行进一步分析的数据产品（传输光谱）。</p>\n<p>我们理解，对你们中的一些人来说，这可能是你们第一次处理天文问题，更不用说系外行星了。</p>\n<p>为了帮助你们开始，我们整理了一份材料清单。如果我们认为有任何知识差距需要解决，我们将更新这个列表：</p>\n<p>什么是凌星？：这里有一个简单的解释：<br>\n什么是“凌星深度”和光谱？——《系外行星手册》（Perryman 2018年）很好地概述了凌星和光谱学（第6章，第6.12节和第11章，11.6节），或者可以看看Winn（2010年）<br>\n什么是观测？什么是光变曲线？——科学会议对这个话题有一个很好的介绍，当然，Ariel红宝书在这方面也很有用。<br>\n可能有用的出版物——你们中的一些人可能已经发布了这些 - 链接<br>\n训练 - 测试分布的差异。长话短说，是的，测试数据集的某些方面是故意与训练集不同的。所以如果你的本地分数和排行榜分数不同，不要太惊讶。我们遵循这里概述的相同哲学，第2.2.1节<br>\n按批次校准数据的笔记本，校准单个观测的笔记本。<br>\n即将推出的起始解决方案！<br>\n在Twitter/X上关注我们 - 我们将在那里分享我们的更新，敬请关注！<br>\n在YouTube上关注我们<br>\n当然，它们只是冰山一角，如果你们有任何特定问题的方面想知道，请让我们知道。我们很高兴更新我们的参考资料。我们很乐意回答你们可能提出的任何问题（只要它们不暴露测试集，当然！），我会尽可能多地监控讨论。请随时标记我。</p>\n<p>虽然领域知识很重要，但请记住，它只代表我们目前所知道的内容。数据中有一些现象我们不知道或没有考虑到。因此，这就是为什么设置这个挑战 - 用像你们这样新鲜的头脑探索去噪数据的新方法</p>",
      "rawMarkdown": "mark，翻译一下\n我们对所有参赛者表示迟来的欢迎。感谢你们对Ariel数据挑战赛🔭的兴趣。在这个挑战中，我们要求你们将原始观测数据处理成一个准备好进行进一步分析的数据产品（传输光谱）。\n\n我们理解，对你们中的一些人来说，这可能是你们第一次处理天文问题，更不用说系外行星了。\n\n为了帮助你们开始，我们整理了一份材料清单。如果我们认为有任何知识差距需要解决，我们将更新这个列表：\n\n什么是凌星？：这里有一个简单的解释：\n什么是“凌星深度”和光谱？——《系外行星手册》（Perryman 2018年）很好地概述了凌星和光谱学（第6章，第6.12节和第11章，11.6节），或者可以看看Winn（2010年）\n什么是观测？什么是光变曲线？——科学会议对这个话题有一个很好的介绍，当然，Ariel红宝书在这方面也很有用。\n可能有用的出版物——你们中的一些人可能已经发布了这些 - 链接\n训练 - 测试分布的差异。长话短说，是的，测试数据集的某些方面是故意与训练集不同的。所以如果你的本地分数和排行榜分数不同，不要太惊讶。我们遵循这里概述的相同哲学，第2.2.1节\n按批次校准数据的笔记本，校准单个观测的笔记本。\n即将推出的起始解决方案！\n在Twitter/X上关注我们 - 我们将在那里分享我们的更新，敬请关注！\n在YouTube上关注我们\n当然，它们只是冰山一角，如果你们有任何特定问题的方面想知道，请让我们知道。我们很高兴更新我们的参考资料。我们很乐意回答你们可能提出的任何问题（只要它们不暴露测试集，当然！），我会尽可能多地监控讨论。请随时标记我。\n\n虽然领域知识很重要，但请记住，它只代表我们目前所知道的内容。数据中有一些现象我们不知道或没有考虑到。因此，这就是为什么设置这个挑战 - 用像你们这样新鲜的头脑探索去噪数据的新方法",
      "votes": -3,
      "replies": [
        {
          "id": 2985975,
          "postDate": "2024-09-11T08:05:55.563Z",
          "content": "<p>I'm not sure whether you are Chinese. If you are or a Chinese-speaking person, please do not use chatgpt or translation tools to mislead Chinese or other Chinese-speaking kagglers, most of the terms are wrong in Chinese.</p>\n<p>我不确定你是不是中文母语国家的人，不要用翻译软件误导大家，很多专业术语你都翻译错了。</p>",
          "rawMarkdown": "I'm not sure whether you are Chinese. If you are or a Chinese-speaking person, please do not use chatgpt or translation tools to mislead Chinese or other Chinese-speaking kagglers, most of the terms are wrong in Chinese.\n\n我不确定你是不是中文母语国家的人，不要用翻译软件误导大家，很多专业术语你都翻译错了。",
          "votes": 2,
          "replies": [
            {
              "id": 2988178,
              "postDate": "2024-09-13T13:44:56.467Z",
              "content": "<blockquote>\n  <p>I'm not sure whether you are Chinese. If you are or a Chinese-speaking person, please do not use chatgpt or translation tool to mislead Chinese or other Chinese-speaking countries, most of the terms are wrong in Chinese.</p>\n  <p>我不确定你是不是中文母语国家的人，不要用翻译软件误导大家，很多专业术语你都翻译错了。<br>\n  可以的</p>\n</blockquote>",
              "rawMarkdown": "> I'm not sure whether you are Chinese. If you are or a Chinese-speaking person, please do not use chatgpt or translation tool to mislead Chinese or other Chinese-speaking countries, most of the terms are wrong in Chinese.\n> \n> 我不确定你是不是中文母语国家的人，不要用翻译软件误导大家，很多专业术语你都翻译错了。\n可以的\n",
              "votes": -2
            }
          ]
        },
        {
          "id": 3014511,
          "postDate": "2024-10-11T09:54:19.853Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 3008430,
      "postDate": "2024-10-06T15:11:27.500Z",
      "content": "<p>Can someone help explain this problem about What is 'transit depth' and a spectrum?</p>",
      "rawMarkdown": "Can someone help explain this problem about What is 'transit depth' and a spectrum?",
      "replies": [
        {
          "id": 3008890,
          "postDate": "2024-10-07T08:40:58.040Z",
          "content": "<p>Hi Alice186, <br>\nSure! A transit depth represents the amount of light that are covered by the planet+its atmosphere during a transiting event (when a planet passes through the projected surface of the star), it creates a dip in the 'light curve'<br>\nIf you observe this event for more than one wavelength, you get transit depths at multiple wavelengths, creating what we called a spectrum. This encodes thermal or dynamical information about the planet's atmosphere, and the chemical species it contains. <br>\nFor more information - check out this link: <a href=\"https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/528233\" target=\"_blank\">https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/528233</a> </p>",
          "rawMarkdown": "Hi Alice186, \nSure! A transit depth represents the amount of light that are covered by the planet+its atmosphere during a transiting event (when a planet passes through the projected surface of the star), it creates a dip in the 'light curve'\nIf you observe this event for more than one wavelength, you get transit depths at multiple wavelengths, creating what we called a spectrum. This encodes thermal or dynamical information about the planet's atmosphere, and the chemical species it contains. \nFor more information - check out this link: https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/528233 ",
          "votes": 2
        }
      ]
    },
    {
      "id": 2954029,
      "postDate": "2024-08-09T10:01:06.953Z",
      "content": "<blockquote>\n  <p>-&gt; -&gt; Thank you, <a href=\"https://www.kaggle.com/gordonyip\" target=\"_blank\">@gordonyip</a>, for quickly addressing the questions over the past seven days with detailed explanations. Your insights helped clarify many aspects of the data. This dataset is truly unique and innovative, which has motivated me to delve deeper into the concepts over the past few days.</p>\n</blockquote>\n<hr>\n<blockquote>\n  <p>=&gt; Waiting for the <strong>Starter solution</strong></p>\n</blockquote>",
      "rawMarkdown": "> -> -> Thank you, @gordonyip, for quickly addressing the questions over the past seven days with detailed explanations. Your insights helped clarify many aspects of the data. This dataset is truly unique and innovative, which has motivated me to delve deeper into the concepts over the past few days.\n\n---\n\n> => Waiting for the **Starter solution**",
      "replies": [
        {
          "id": 2964264,
          "postDate": "2024-08-19T16:58:43.640Z",
          "content": "<p>just updated :) </p>",
          "rawMarkdown": "just updated :) ",
          "votes": 1
        }
      ]
    },
    {
      "id": 2950183,
      "postDate": "2024-08-07T10:31:46.050Z",
      "content": "<p><a href=\"https://www.kaggle.com/gordonyip\" target=\"_blank\">@gordonyip</a> </p>\n<h1>Few more questions</h1>\n<blockquote>\n  <ol>\n  <li>All transits are with single planet(similar size exoplanets) as per training data =&gt; is this same for test data too? i.e no multiple exoplanets in same transit.</li>\n  </ol>\n</blockquote>\n<hr>\n<blockquote>\n  <ol>\n  <li>Its having 2 solar systems =&gt; first with <strong>346 exoplanets</strong> and second with <strong>327 exoplanets</strong>, is remaining <strong>327 exoplanets</strong> (test) belongs to same solar systems or new solar system?</li>\n  </ol>\n</blockquote>\n<hr>\n<blockquote>\n  <ol>\n  <li>Will you share any resources for better understanding target variable wavelengths as wl_283 is what famous from previous articles, but for other targets, not much details to learn.</li>\n  </ol>\n</blockquote>",
      "rawMarkdown": "@gordonyip \n# Few more questions\n\n> 1. All transits are with single planet(similar size exoplanets) as per training data => is this same for test data too? i.e no multiple exoplanets in same transit.\n\n---\n\n> 2. Its having 2 solar systems => first with **346 exoplanets** and second with **327 exoplanets**, is remaining **327 exoplanets** (test) belongs to same solar systems or new solar system?\n\n---\n\n> 3. Will you share any resources for better understanding target variable wavelengths as wl_283 is what famous from previous articles, but for other targets, not much details to learn.\n",
      "replies": [
        {
          "id": 2950440,
          "postDate": "2024-08-07T15:10:39.570Z",
          "content": "<p><code>All transits are with single planet(similar size exoplanets) as per training data =&gt; is this same for test data too? i.e no multiple parents in same transit.</code><br>\nYes all observation feature a single transit</p>\n<p><code>Its having 2 solar systems =&gt; first with 346 exoplanets and second with 327 exoplanets, is remaining 327 exoplanets (test) belongs to same solar systems or new solar system?</code><br>\nThe 327 belongs to a new planetary system (or new host star)</p>\n<p><code>Will you share any resources for better understanding target variable wavelengths as wl_283 is what famous from previous articles, but for other targets, not much details to learn.</code><br>\nThey are the wavelength coverage for FGS1 and AIRS-Ch0. One is a photometer and the other is a spectrometer. Long story short: The wavelength coverage helps us to uncover different molecular species and different dynamical phenomena in the atmosphere, as they are the main driver that are producing different transit depths in different wavelengths. You can look at some examples <a href=\"https://academic.oup.com/rasti/article/2/1/45/6998590\" target=\"_blank\">here</a> , Figure A2<br>\nI will also update the reference list to include some references on data detrending. </p>",
          "rawMarkdown": "`All transits are with single planet(similar size exoplanets) as per training data => is this same for test data too? i.e no multiple parents in same transit.`\nYes all observation feature a single transit\n\n`Its having 2 solar systems => first with 346 exoplanets and second with 327 exoplanets, is remaining 327 exoplanets (test) belongs to same solar systems or new solar system?`\nThe 327 belongs to a new planetary system (or new host star)\n\n`Will you share any resources for better understanding target variable wavelengths as wl_283 is what famous from previous articles, but for other targets, not much details to learn.`\nThey are the wavelength coverage for FGS1 and AIRS-Ch0. One is a photometer and the other is a spectrometer. Long story short: The wavelength coverage helps us to uncover different molecular species and different dynamical phenomena in the atmosphere, as they are the main driver that are producing different transit depths in different wavelengths. You can look at some examples [here](https://academic.oup.com/rasti/article/2/1/45/6998590) , Figure A2\nI will also update the reference list to include some references on data detrending. \n\n",
          "votes": 5,
          "replies": [
            {
              "id": 2950473,
              "postDate": "2024-08-07T15:37:00.133Z",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/gordonyip\" target=\"_blank\">@gordonyip</a> for detailed explanation and i read this paper also. -&gt; This helps to build better CV.</p>",
              "rawMarkdown": "Thanks @gordonyip for detailed explanation and i read this paper also. -> This helps to build better CV.\n\n"
            },
            {
              "id": 2987288,
              "postDate": "2024-09-12T15:01:36.533Z",
              "content": "<p><a href=\"https://www.kaggle.com/gordonyip\" target=\"_blank\">@gordonyip</a> May I ask for some clarification?</p>\n<p>In the second question, <a href=\"https://www.kaggle.com/seshurajup\" target=\"_blank\">@seshurajup</a> stated that <code>327 exoplanets (test)</code>, but I think that 346 and 327 exoplanets are probably referring to the training dataset. Do we have information on how many new stars are in the hidden test dataset?</p>",
              "rawMarkdown": "@gordonyip May I ask for some clarification?\n\n In the second question, @seshurajup stated that `327 exoplanets (test)`, but I think that 346 and 327 exoplanets are probably referring to the training dataset. Do we have information on how many new stars are in the hidden test dataset?",
              "votes": 1
            },
            {
              "id": 3006226,
              "postDate": "2024-10-03T21:49:45.833Z",
              "content": "<p><a href=\"https://www.kaggle.com/gordonyip\" target=\"_blank\">@gordonyip</a>  thanks - please can we confirm if the test set is composed of 327 exoplanets of a new star type, and the other (800-327 = 473) of star type 0 or 1?</p>",
              "rawMarkdown": "@gordonyip  thanks - please can we confirm if the test set is composed of 327 exoplanets of a new star type, and the other (800-327 = 473) of star type 0 or 1?"
            }
          ]
        }
      ]
    },
    {
      "id": 2978991,
      "postDate": "2024-09-04T12:59:09.143Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2976117,
      "postDate": "2024-09-01T14:37:26.167Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 3012773,
      "postDate": "2024-10-09T11:12:09.197Z",
      "content": "<p>thanks! it works!!!</p>",
      "rawMarkdown": "thanks! it works!!!",
      "votes": 1
    },
    {
      "id": 2947971,
      "postDate": "2024-08-05T15:28:32.953Z",
      "content": "<p>Awesome, thanks! </p>",
      "rawMarkdown": "Awesome, thanks! ",
      "votes": 1
    },
    {
      "id": 3015568,
      "postDate": "2024-10-12T15:36:06.900Z",
      "content": "<p>Very challenging,thank you</p>",
      "rawMarkdown": "Very challenging,thank you"
    },
    {
      "id": 2998082,
      "postDate": "2024-09-25T08:11:10.970Z",
      "content": "<p>Thank you for information</p>",
      "rawMarkdown": "Thank you for information"
    },
    {
      "id": 2982617,
      "postDate": "2024-09-08T03:14:49.293Z",
      "content": "<p>very helpful</p>",
      "rawMarkdown": "very helpful"
    },
    {
      "id": 2984846,
      "postDate": "2024-09-10T03:48:25.507Z",
      "content": "<p>Very helpful, Thanks:)</p>",
      "rawMarkdown": "Very helpful, Thanks:)",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3009652,
      "author_name": "yuanzhe zhou",
      "author_url": "",
      "post_date": "2024-10-08T07:13:50.850000",
      "content": "<p>Why don't we release the precise score below 0.? It becomes very hard to debug with bad cases hidden …</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3012819,
          "author_name": "Gordon Yip",
          "author_url": "",
          "post_date": "2024-10-09T12:03:20.310000",
          "content": "<p>sorry to hear that, but i am afraid there is not much we can do here. When we setup the 0th point on the system (since the GLL function does not have a bounded maximum), we tried to be as lenient as possible to include as much solution as possible, but obviously, that also means some solutions wont get any scores. </p>",
          "votes": 1,
          "replies": [
            {
              "id": 3024826,
              "author_name": "Sumit Satyanarayan Sharma",
              "author_url": "",
              "post_date": "2024-10-22T04:00:37.873000",
              "content": "<p>But it's possible, If yes how it can be. </p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 3034267,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-11-02T00:17:01.653000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3034268,
          "author_name": "yiyao6442",
          "author_url": "",
          "post_date": "2024-11-02T00:18:16.580000",
          "content": "<p>It definitely makes debugging more challenging when bad cases aren't visible. Having detailed feedback could help participants identify issues in their solutions more effectively. It might be worth suggesting that the scoring system be adjusted to provide more transparency, even if it means just releasing additional insights or metrics that could guide debugging efforts. Your input could be valuable in improving the overall experience for everyone!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2950984,
      "author_name": "AGG",
      "author_url": "",
      "post_date": "2024-08-08T06:02:39.167000",
      "content": "<p>1) From the four groups in the test data (Set 1, …, Set 4), can you confirm if they are all in the public and private LB?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F311485%2F44fbf3ea84c4bcfe07756f789d1b41b3%2FScreenshot%202024-08-08%20at%2007.55.43.png?generation=1723096587393487&amp;alt=media\"></p>\n<p>2) Are there any details available on how the simulations were run? </p>\n<p>Thanks a lot in advance!</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 2968069,
      "author_name": "Anton.Tit",
      "author_url": "",
      "post_date": "2024-08-23T13:51:37.510000",
      "content": "<p><a href=\"https://www.kaggle.com/gordonyip\" target=\"_blank\">@gordonyip</a> Thank you for the additional information. I would like to ask a question that might be quite basic. Upon a review of the data, I don't see any difference between the training and test sets. I would like to ask what the target feature of our research is and how to identify it in the training set. Unfortunately, I couldn't find this information, if I'm not mistaken. Thank you in advance.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2972948,
          "author_name": "E-Max AI",
          "author_url": "",
          "post_date": "2024-08-29T02:48:19.793000",
          "content": "<p>train_labels.csv </p>",
          "votes": 0,
          "replies": [
            {
              "id": 2974172,
              "author_name": "Anton.Tit",
              "author_url": "",
              "post_date": "2024-08-30T11:22:34.510000",
              "content": "<p>Ok, thank you! I am starting to understend better.</p>\n<p>Another one question, which information contain file axis_info.parquet? In description an explanation if very poore. I woold like to understand meaning of each column in this file.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2978955,
              "author_name": "Gordon Yip",
              "author_url": "",
              "post_date": "2024-09-04T12:30:47.740000",
              "content": "<p>We will add more description to the file, but essentially, h means hour (for time axis), um mean micrometer (spectral axis), and other simply means spatial axis</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2951278,
      "author_name": "Evan Tung",
      "author_url": "",
      "post_date": "2024-08-08T13:31:24.030000",
      "content": "<p>Are we allowed to use pretrained models from Kaggle models or Hugging Face, or do we have to build our own model from scratch?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2953637,
          "author_name": "Gordon Yip",
          "author_url": "",
          "post_date": "2024-08-08T22:07:35.137000",
          "content": "<p>yes, you are allowed to use open source models </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 2975271,
          "author_name": "Jeremiah Akisanya",
          "author_url": "",
          "post_date": "2024-08-31T17:08:37.543000",
          "content": "<p>Hi Evan, If you do not mind, can I merge with your team?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2962291,
      "author_name": "E-Max AI",
      "author_url": "",
      "post_date": "2024-08-17T12:05:26.217000",
      "content": "<p>mark，翻译一下<br>\n我们对所有参赛者表示迟来的欢迎。感谢你们对Ariel数据挑战赛🔭的兴趣。在这个挑战中，我们要求你们将原始观测数据处理成一个准备好进行进一步分析的数据产品（传输光谱）。</p>\n<p>我们理解，对你们中的一些人来说，这可能是你们第一次处理天文问题，更不用说系外行星了。</p>\n<p>为了帮助你们开始，我们整理了一份材料清单。如果我们认为有任何知识差距需要解决，我们将更新这个列表：</p>\n<p>什么是凌星？：这里有一个简单的解释：<br>\n什么是“凌星深度”和光谱？——《系外行星手册》（Perryman 2018年）很好地概述了凌星和光谱学（第6章，第6.12节和第11章，11.6节），或者可以看看Winn（2010年）<br>\n什么是观测？什么是光变曲线？——科学会议对这个话题有一个很好的介绍，当然，Ariel红宝书在这方面也很有用。<br>\n可能有用的出版物——你们中的一些人可能已经发布了这些 - 链接<br>\n训练 - 测试分布的差异。长话短说，是的，测试数据集的某些方面是故意与训练集不同的。所以如果你的本地分数和排行榜分数不同，不要太惊讶。我们遵循这里概述的相同哲学，第2.2.1节<br>\n按批次校准数据的笔记本，校准单个观测的笔记本。<br>\n即将推出的起始解决方案！<br>\n在Twitter/X上关注我们 - 我们将在那里分享我们的更新，敬请关注！<br>\n在YouTube上关注我们<br>\n当然，它们只是冰山一角，如果你们有任何特定问题的方面想知道，请让我们知道。我们很高兴更新我们的参考资料。我们很乐意回答你们可能提出的任何问题（只要它们不暴露测试集，当然！），我会尽可能多地监控讨论。请随时标记我。</p>\n<p>虽然领域知识很重要，但请记住，它只代表我们目前所知道的内容。数据中有一些现象我们不知道或没有考虑到。因此，这就是为什么设置这个挑战 - 用像你们这样新鲜的头脑探索去噪数据的新方法</p>",
      "votes": -3,
      "replies": [
        {
          "id": 2985975,
          "author_name": "Timmy Juicehouse",
          "author_url": "",
          "post_date": "2024-09-11T08:05:55.563000",
          "content": "<p>I'm not sure whether you are Chinese. If you are or a Chinese-speaking person, please do not use chatgpt or translation tools to mislead Chinese or other Chinese-speaking kagglers, most of the terms are wrong in Chinese.</p>\n<p>我不确定你是不是中文母语国家的人，不要用翻译软件误导大家，很多专业术语你都翻译错了。</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2988178,
              "author_name": "EddieYan7",
              "author_url": "",
              "post_date": "2024-09-13T13:44:56.467000",
              "content": "<blockquote>\n  <p>I'm not sure whether you are Chinese. If you are or a Chinese-speaking person, please do not use chatgpt or translation tool to mislead Chinese or other Chinese-speaking countries, most of the terms are wrong in Chinese.</p>\n  <p>我不确定你是不是中文母语国家的人，不要用翻译软件误导大家，很多专业术语你都翻译错了。<br>\n  可以的</p>\n</blockquote>",
              "votes": -2,
              "replies": []
            }
          ]
        },
        {
          "id": 3014511,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-10-11T09:54:19.853000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3008430,
      "author_name": "Alice",
      "author_url": "",
      "post_date": "2024-10-06T15:11:27.500000",
      "content": "<p>Can someone help explain this problem about What is 'transit depth' and a spectrum?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3008890,
          "author_name": "Gordon Yip",
          "author_url": "",
          "post_date": "2024-10-07T08:40:58.040000",
          "content": "<p>Hi Alice186, <br>\nSure! A transit depth represents the amount of light that are covered by the planet+its atmosphere during a transiting event (when a planet passes through the projected surface of the star), it creates a dip in the 'light curve'<br>\nIf you observe this event for more than one wavelength, you get transit depths at multiple wavelengths, creating what we called a spectrum. This encodes thermal or dynamical information about the planet's atmosphere, and the chemical species it contains. <br>\nFor more information - check out this link: <a href=\"https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/528233\" target=\"_blank\">https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/528233</a> </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2954029,
      "author_name": "SeshuRaju 🧘‍♂️",
      "author_url": "",
      "post_date": "2024-08-09T10:01:06.953000",
      "content": "<blockquote>\n  <p>-&gt; -&gt; Thank you, <a href=\"https://www.kaggle.com/gordonyip\" target=\"_blank\">@gordonyip</a>, for quickly addressing the questions over the past seven days with detailed explanations. Your insights helped clarify many aspects of the data. This dataset is truly unique and innovative, which has motivated me to delve deeper into the concepts over the past few days.</p>\n</blockquote>\n<hr>\n<blockquote>\n  <p>=&gt; Waiting for the <strong>Starter solution</strong></p>\n</blockquote>",
      "votes": 0,
      "replies": [
        {
          "id": 2964264,
          "author_name": "Gordon Yip",
          "author_url": "",
          "post_date": "2024-08-19T16:58:43.640000",
          "content": "<p>just updated :) </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2950183,
      "author_name": "SeshuRaju 🧘‍♂️",
      "author_url": "",
      "post_date": "2024-08-07T10:31:46.050000",
      "content": "<p><a href=\"https://www.kaggle.com/gordonyip\" target=\"_blank\">@gordonyip</a> </p>\n<h1>Few more questions</h1>\n<blockquote>\n  <ol>\n  <li>All transits are with single planet(similar size exoplanets) as per training data =&gt; is this same for test data too? i.e no multiple exoplanets in same transit.</li>\n  </ol>\n</blockquote>\n<hr>\n<blockquote>\n  <ol>\n  <li>Its having 2 solar systems =&gt; first with <strong>346 exoplanets</strong> and second with <strong>327 exoplanets</strong>, is remaining <strong>327 exoplanets</strong> (test) belongs to same solar systems or new solar system?</li>\n  </ol>\n</blockquote>\n<hr>\n<blockquote>\n  <ol>\n  <li>Will you share any resources for better understanding target variable wavelengths as wl_283 is what famous from previous articles, but for other targets, not much details to learn.</li>\n  </ol>\n</blockquote>",
      "votes": 0,
      "replies": [
        {
          "id": 2950440,
          "author_name": "Gordon Yip",
          "author_url": "",
          "post_date": "2024-08-07T15:10:39.570000",
          "content": "<p><code>All transits are with single planet(similar size exoplanets) as per training data =&gt; is this same for test data too? i.e no multiple parents in same transit.</code><br>\nYes all observation feature a single transit</p>\n<p><code>Its having 2 solar systems =&gt; first with 346 exoplanets and second with 327 exoplanets, is remaining 327 exoplanets (test) belongs to same solar systems or new solar system?</code><br>\nThe 327 belongs to a new planetary system (or new host star)</p>\n<p><code>Will you share any resources for better understanding target variable wavelengths as wl_283 is what famous from previous articles, but for other targets, not much details to learn.</code><br>\nThey are the wavelength coverage for FGS1 and AIRS-Ch0. One is a photometer and the other is a spectrometer. Long story short: The wavelength coverage helps us to uncover different molecular species and different dynamical phenomena in the atmosphere, as they are the main driver that are producing different transit depths in different wavelengths. You can look at some examples <a href=\"https://academic.oup.com/rasti/article/2/1/45/6998590\" target=\"_blank\">here</a> , Figure A2<br>\nI will also update the reference list to include some references on data detrending. </p>",
          "votes": 5,
          "replies": [
            {
              "id": 2950473,
              "author_name": "SeshuRaju 🧘‍♂️",
              "author_url": "",
              "post_date": "2024-08-07T15:37:00.133000",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/gordonyip\" target=\"_blank\">@gordonyip</a> for detailed explanation and i read this paper also. -&gt; This helps to build better CV.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2987288,
              "author_name": "ChingYinNg",
              "author_url": "",
              "post_date": "2024-09-12T15:01:36.533000",
              "content": "<p><a href=\"https://www.kaggle.com/gordonyip\" target=\"_blank\">@gordonyip</a> May I ask for some clarification?</p>\n<p>In the second question, <a href=\"https://www.kaggle.com/seshurajup\" target=\"_blank\">@seshurajup</a> stated that <code>327 exoplanets (test)</code>, but I think that 346 and 327 exoplanets are probably referring to the training dataset. Do we have information on how many new stars are in the hidden test dataset?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3006226,
              "author_name": "Heisenger",
              "author_url": "",
              "post_date": "2024-10-03T21:49:45.833000",
              "content": "<p><a href=\"https://www.kaggle.com/gordonyip\" target=\"_blank\">@gordonyip</a>  thanks - please can we confirm if the test set is composed of 327 exoplanets of a new star type, and the other (800-327 = 473) of star type 0 or 1?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2978991,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-09-04T12:59:09.143000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2976117,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-09-01T14:37:26.167000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3012773,
      "author_name": "Alice",
      "author_url": "",
      "post_date": "2024-10-09T11:12:09.197000",
      "content": "<p>thanks! it works!!!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2947971,
      "author_name": "DennisSakva",
      "author_url": "",
      "post_date": "2024-08-05T15:28:32.953000",
      "content": "<p>Awesome, thanks! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3015568,
      "author_name": "JNL",
      "author_url": "",
      "post_date": "2024-10-12T15:36:06.900000",
      "content": "<p>Very challenging,thank you</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2998082,
      "author_name": "Danish Ghaffar",
      "author_url": "",
      "post_date": "2024-09-25T08:11:10.970000",
      "content": "<p>Thank you for information</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2982617,
      "author_name": "Rowan",
      "author_url": "",
      "post_date": "2024-09-08T03:14:49.293000",
      "content": "<p>very helpful</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2984846,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-09-10T03:48:25.507000",
      "content": "<p>Very helpful, Thanks:)</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2947916": "Our belated welcome to all fellow challengers. Thank you for your interest in the Ariel Data Challenge🔭. In this challenge, we are asking you to process raw observations into a data product ready for further analysis (transmission spectrum).  \n\nWe understand that for some of you, this could be your first time dealing with astronomical problems, let alone exoplanets. \n\nTo help you get started, we have compiled a list of materials. We will be updating the list if we think there are any knowledge gaps we should address:\n1. What is a transit? : A simple explanation [ here]( https://exoplanets.nasa.gov/alien-worlds/ways-to-find-a-planet/?intent=021#/2): \n2. What is 'transit depth' and a spectrum? -- The Exoplanet Handbook (Perryman 2018) gives a good overview of transit and spectroscopy (Chap. 6, Sec. 6.12 and Chap. 11, 11.6), alternatively, have a look at [Winn (2010) ](https://arxiv.org/pdf/1001.2010v5)\n3. What is an observation? What is a light curve? -- The [Science](https://www.ariel-datachallenge.space/workshop2024/documentation/about) session gives a nice introduction to the topic, of course, the Ariel redbook is useful on that as well.\n4. Publications that could be useful -- some of you may have posted these already - [link](https://www.ariel-datachallenge.space/ML/documentation/resources)\n5. Differences between Train - test distribution. Long story short, yes, certain aspects of the test data distribution are deliberately made different from the training set. So don't be too surprised if your local score is different from your leaderboard score. We follow the same philosophy as outlined [here](https://proceedings.mlr.press/v220/yip23a/yip23a.pdf), section 2.2.1\n6. [Notebook](https://www.kaggle.com/code/gordonyip/calibrating-astronomical-data) on Calibrating Data by Batch , [Notebook](https://www.kaggle.com/code/gordonyip/calibrating-a-single-observation) on calibrating a single observation.\n7. [Starter solution](https://www.kaggle.com/code/gordonyip/adc-2024-starter-solution) - there are also plenty notebooks available in the code section.\n8. Follow us on [Twitter/X](https://x.com/ArielTelescope) - we will be sharing our updates there, stay tuned!\n9. Follow us on [YouTube](https://www.youtube.com/channel/UCMLTUdXBPNS_pDJcGNZoLYw)\n\n\nOf course, they are just the tip of an iceberg, and let us know if there are any specific aspect of the problem that you want to know. We are happy update our references. We are happy to answer any questions you may have (as long as they don't expose the test set, of course!), and I will be monitoring the discussion as much as I can. Please feel free to tag me.\n\n**While domain knowledge is important, please remember that it only represents what we know so far. There are phenomena within the data that we are unaware of or have not accounted for. Hence, this is why the challenge is set up - to explore new ways to denoise the data with fresh new minds like yours**\n\nLast but not least, please have fun! We are eager to find out what you can come up with!",
    "3009652": "Why don't we release the precise score below 0.? It becomes very hard to debug with bad cases hidden ...",
    "2950984": "1) From the four groups in the test data (Set 1, ..., Set 4), can you confirm if they are all in the public and private LB?\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F311485%2F44fbf3ea84c4bcfe07756f789d1b41b3%2FScreenshot%202024-08-08%20at%2007.55.43.png?generation=1723096587393487&alt=media\" width=500>\n\n2) Are there any details available on how the simulations were run? \n\nThanks a lot in advance!",
    "2968069": "@gordonyip Thank you for the additional information. I would like to ask a question that might be quite basic. Upon a review of the data, I don't see any difference between the training and test sets. I would like to ask what the target feature of our research is and how to identify it in the training set. Unfortunately, I couldn't find this information, if I'm not mistaken. Thank you in advance.",
    "2951278": "Are we allowed to use pretrained models from Kaggle models or Hugging Face, or do we have to build our own model from scratch?",
    "2962291": "mark，翻译一下\n我们对所有参赛者表示迟来的欢迎。感谢你们对Ariel数据挑战赛🔭的兴趣。在这个挑战中，我们要求你们将原始观测数据处理成一个准备好进行进一步分析的数据产品（传输光谱）。\n\n我们理解，对你们中的一些人来说，这可能是你们第一次处理天文问题，更不用说系外行星了。\n\n为了帮助你们开始，我们整理了一份材料清单。如果我们认为有任何知识差距需要解决，我们将更新这个列表：\n\n什么是凌星？：这里有一个简单的解释：\n什么是“凌星深度”和光谱？——《系外行星手册》（Perryman 2018年）很好地概述了凌星和光谱学（第6章，第6.12节和第11章，11.6节），或者可以看看Winn（2010年）\n什么是观测？什么是光变曲线？——科学会议对这个话题有一个很好的介绍，当然，Ariel红宝书在这方面也很有用。\n可能有用的出版物——你们中的一些人可能已经发布了这些 - 链接\n训练 - 测试分布的差异。长话短说，是的，测试数据集的某些方面是故意与训练集不同的。所以如果你的本地分数和排行榜分数不同，不要太惊讶。我们遵循这里概述的相同哲学，第2.2.1节\n按批次校准数据的笔记本，校准单个观测的笔记本。\n即将推出的起始解决方案！\n在Twitter/X上关注我们 - 我们将在那里分享我们的更新，敬请关注！\n在YouTube上关注我们\n当然，它们只是冰山一角，如果你们有任何特定问题的方面想知道，请让我们知道。我们很高兴更新我们的参考资料。我们很乐意回答你们可能提出的任何问题（只要它们不暴露测试集，当然！），我会尽可能多地监控讨论。请随时标记我。\n\n虽然领域知识很重要，但请记住，它只代表我们目前所知道的内容。数据中有一些现象我们不知道或没有考虑到。因此，这就是为什么设置这个挑战 - 用像你们这样新鲜的头脑探索去噪数据的新方法",
    "3008430": "Can someone help explain this problem about What is 'transit depth' and a spectrum?",
    "2954029": "> -> -> Thank you, @gordonyip, for quickly addressing the questions over the past seven days with detailed explanations. Your insights helped clarify many aspects of the data. This dataset is truly unique and innovative, which has motivated me to delve deeper into the concepts over the past few days.\n\n---\n\n> => Waiting for the **Starter solution**",
    "2950183": "@gordonyip \n# Few more questions\n\n> 1. All transits are with single planet(similar size exoplanets) as per training data => is this same for test data too? i.e no multiple exoplanets in same transit.\n\n---\n\n> 2. Its having 2 solar systems => first with **346 exoplanets** and second with **327 exoplanets**, is remaining **327 exoplanets** (test) belongs to same solar systems or new solar system?\n\n---\n\n> 3. Will you share any resources for better understanding target variable wavelengths as wl_283 is what famous from previous articles, but for other targets, not much details to learn.\n",
    "2978991": "",
    "2976117": "",
    "3012773": "thanks! it works!!!",
    "2947971": "Awesome, thanks! ",
    "3015568": "Very challenging,thank you",
    "2998082": "Thank you for information",
    "2982617": "very helpful",
    "2984846": "Very helpful, Thanks:)"
  }
}