{
  "id": 598324,
  "title": "An odd transit (Planet 1843015807)",
  "url": "/competitions/ariel-data-challenge-2025/discussion/598324",
  "author_name": "Frederick Lee",
  "post_date": "2025-08-10T04:07:52.112000",
  "votes": 3,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hello! I'm looking at a rather odd transit in the train set which is so wide that the top of the transit is not observed within the frame of the prediction. It's rather unusual, so I just wanted to ask: Is this transit intentional and of the type that the Ariel mission expects to sometimes record, or is it an artifact in the data?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10771865%2Feeb0999b95c8d46815d43e0023b6aedb%2Ff81795e8-2724-499c-be45-8f445bc7f094.png?generation=1754783609820958&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 3267437,
      "postDate": "2025-08-11T07:45:57.423Z",
      "content": "<p>Yeah in reality we should have longer observation time for those long transits, but constraints in this challenge mean that we cant increase the observation time (that will mean breaking the uniformity of the data), hence the odd cases listed in this thread. </p>",
      "rawMarkdown": "Yeah in reality we should have longer observation time for those long transits, but constraints in this challenge mean that we cant increase the observation time (that will mean breaking the uniformity of the data), hence the odd cases listed in this thread. ",
      "votes": 3,
      "replies": [
        {
          "id": 3269605,
          "postDate": "2025-08-14T18:02:48.997Z",
          "content": "<p><a href=\"https://www.kaggle.com/gordonyip\" target=\"_blank\">@gordonyip</a> Issue is that these impact the score pretty heavily. Do you want the final ranking to depend on pathological cases not representative of the real Ariel use case?</p>\n<p>Wouldn't it make sense to discard these samples where ingress and/or egress are not fully contained in the data we have? I am speaking of the test set here of course.</p>\n<p>I spend time fixing these but I doubt this is useful, nor do I think I did a perfect job at it. I am pretty sure many others spent time on these samples too. If we knew for sure that nothing similar occurs in the test set then we can all save time an focus on what is useful to you.</p>",
          "rawMarkdown": "@gordonyip Issue is that these impact the score pretty heavily. Do you want the final ranking to depend on pathological cases not representative of the real Ariel use case?\n\nWouldn't it make sense to discard these samples where ingress and/or egress are not fully contained in the data we have? I am speaking of the test set here of course.\n\nI spend time fixing these but I doubt this is useful, nor do I think I did a perfect job at it. I am pretty sure many others spent time on these samples too. If we knew for sure that nothing similar occurs in the test set then we can all save time an focus on what is useful to you.\n\n",
          "votes": 2,
          "replies": [
            {
              "id": 3270223,
              "postDate": "2025-08-16T02:50:56.430Z",
              "content": "<p>I agree fully in principle, and this would have been great if done at the start. But it doesn't seem like a fair change to make so late in the competition, since people may have already put significant effort into dealing with this. Which indeed doesn't further the scientific goals, but that's what it is now…</p>",
              "rawMarkdown": "I agree fully in principle, and this would have been great if done at the start. But it doesn't seem like a fair change to make so late in the competition, since people may have already put significant effort into dealing with this. Which indeed doesn't further the scientific goals, but that's what it is now..."
            }
          ]
        }
      ]
    },
    {
      "id": 3266960,
      "postDate": "2025-08-10T04:07:52.113Z",
      "content": "<p>Hello! I'm looking at a rather odd transit in the train set which is so wide that the top of the transit is not observed within the frame of the prediction. It's rather unusual, so I just wanted to ask: Is this transit intentional and of the type that the Ariel mission expects to sometimes record, or is it an artifact in the data?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10771865%2Feeb0999b95c8d46815d43e0023b6aedb%2Ff81795e8-2724-499c-be45-8f445bc7f094.png?generation=1754783609820958&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Hello! I'm looking at a rather odd transit in the train set which is so wide that the top of the transit is not observed within the frame of the prediction. It's rather unusual, so I just wanted to ask: Is this transit intentional and of the type that the Ariel mission expects to sometimes record, or is it an artifact in the data?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10771865%2Feeb0999b95c8d46815d43e0023b6aedb%2Ff81795e8-2724-499c-be45-8f445bc7f094.png?generation=1754783609820958&alt=media)",
      "votes": 3
    },
    {
      "id": 3267307,
      "postDate": "2025-08-10T22:03:47.953Z",
      "content": "<p>I think also 158006264.</p>",
      "rawMarkdown": "I think also 158006264.",
      "votes": 1
    },
    {
      "id": 3267018,
      "postDate": "2025-08-10T07:56:34.777Z",
      "content": "<p>There is also 1124834224. Those two are super hard to determine. It reality, a the measurement would just be made longer, I think. Here, it puts a challenge on us which may not be solvable, but we can adapt the sigma predictions for such cases to limit the impact on our scores.</p>",
      "rawMarkdown": "There is also 1124834224. Those two are super hard to determine. It reality, a the measurement would just be made longer, I think. Here, it puts a challenge on us which may not be solvable, but we can adapt the sigma predictions for such cases to limit the impact on our scores.",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 3267437,
      "author_name": "Gordon Yip",
      "author_url": "",
      "post_date": "2025-08-11T07:45:57.423000",
      "content": "<p>Yeah in reality we should have longer observation time for those long transits, but constraints in this challenge mean that we cant increase the observation time (that will mean breaking the uniformity of the data), hence the odd cases listed in this thread. </p>",
      "votes": 3,
      "replies": [
        {
          "id": 3269605,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2025-08-14T18:02:48.997000",
          "content": "<p><a href=\"https://www.kaggle.com/gordonyip\" target=\"_blank\">@gordonyip</a> Issue is that these impact the score pretty heavily. Do you want the final ranking to depend on pathological cases not representative of the real Ariel use case?</p>\n<p>Wouldn't it make sense to discard these samples where ingress and/or egress are not fully contained in the data we have? I am speaking of the test set here of course.</p>\n<p>I spend time fixing these but I doubt this is useful, nor do I think I did a perfect job at it. I am pretty sure many others spent time on these samples too. If we knew for sure that nothing similar occurs in the test set then we can all save time an focus on what is useful to you.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3270223,
              "author_name": "Jeroen Cottaar",
              "author_url": "",
              "post_date": "2025-08-16T02:50:56.430000",
              "content": "<p>I agree fully in principle, and this would have been great if done at the start. But it doesn't seem like a fair change to make so late in the competition, since people may have already put significant effort into dealing with this. Which indeed doesn't further the scientific goals, but that's what it is now…</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3267307,
      "author_name": "mudesteven",
      "author_url": "",
      "post_date": "2025-08-10T22:03:47.953000",
      "content": "<p>I think also 158006264.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3267018,
      "author_name": "Pascal Pfeiffer",
      "author_url": "",
      "post_date": "2025-08-10T07:56:34.777000",
      "content": "<p>There is also 1124834224. Those two are super hard to determine. It reality, a the measurement would just be made longer, I think. Here, it puts a challenge on us which may not be solvable, but we can adapt the sigma predictions for such cases to limit the impact on our scores.</p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3267437": "Yeah in reality we should have longer observation time for those long transits, but constraints in this challenge mean that we cant increase the observation time (that will mean breaking the uniformity of the data), hence the odd cases listed in this thread. ",
    "3266960": "Hello! I'm looking at a rather odd transit in the train set which is so wide that the top of the transit is not observed within the frame of the prediction. It's rather unusual, so I just wanted to ask: Is this transit intentional and of the type that the Ariel mission expects to sometimes record, or is it an artifact in the data?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10771865%2Feeb0999b95c8d46815d43e0023b6aedb%2Ff81795e8-2724-499c-be45-8f445bc7f094.png?generation=1754783609820958&alt=media)",
    "3267307": "I think also 158006264.",
    "3267018": "There is also 1124834224. Those two are super hard to determine. It reality, a the measurement would just be made longer, I think. Here, it puts a challenge on us which may not be solvable, but we can adapt the sigma predictions for such cases to limit the impact on our scores."
  }
}