{
  "id": 530152,
  "title": "Thoughts on shakeup potential?",
  "url": "/competitions/ariel-data-challenge-2024/discussion/530152",
  "author_name": "Cody_Null",
  "post_date": "2024-08-24T23:03:14.235000",
  "votes": 13,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I wanted to put this post out just to get a feeling for other’s opinions. I was considering joining this competition, but I have mixed feelings with regard to the shakeup potential. The data is simulated, which would usually make me worry about data drift between separate runs of the simulation (could be private vs public). But I also have seen in discussions that there is a solidified pipeline for this in our case, which I think Would make it more stable? It may even be possible that the private versus public data is a different segment of one simulation. I’m certain I have missed details on this and maybe my thoughts are misguided so I was curious what others thought.</p>",
  "messages": [
    {
      "id": 2969648,
      "postDate": "2024-08-25T08:13:47.533Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/cody11null\" target=\"_blank\">@cody11null</a>, I think it's possible to survive the shakeup if one applies the right methods. As the test dataset distribution deviates massively from the training dataset's distribution, we cannot simply fit a model to the training data and expect that it will work for the test data. Even a GroupKFold on the two stars doesn't help.</p>\n<p>The public notebooks of today (August 25) are all overfitted and misleading (this criticism includes my own public notebooks). The competition asks us to submit two numbers for every planet and wavelength: the first number is the planet's transit depth and the second one (\\(\\sigma\\)) is a measure for the uncertainty of the first one. I don't criticize that these notebooks oversimplify the task by aggregating the data — that was a concious decision which is legitimate for a baseline model. The notebooks have two other, major, flaws:</p>\n<ol>\n<li>They predict the transit depth with a regression model although no regression is needed: Transit depth is defined as a simple measurement in the timeline (\\(\\frac{oot-it}{oot}\\)), and a regression model only overfits to the noise of the training data.</li>\n<li>Although some notebooks apply a correct method to compute \\(\\sigma\\) for the known stars, their \\(\\sigma\\) prediction for the unknown stars is derived from leaderboard probing. You need to find a method for computing the uncertainty \\(\\sigma\\) based on the given features rather than leaderboard probing. By the way, one notebook predicts \\(\\sigma=0.00016\\) for known stars and \\(\\sigma=0.00085\\) for unknown stars (five times higher uncertainty). These two numbers show how far away the test data are from the training data.</li>\n</ol>\n<p>This competition is not so much about machine learning. It is much more about analyzing the data, managing noise and building up some domain knowledge.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2F2c0d649ed3a382aa9582b44e398fde22%2Foot-it.png?generation=1724573541810992&amp;alt=media\" alt=\"oot\"></p>",
      "rawMarkdown": "Hi @cody11null, I think it's possible to survive the shakeup if one applies the right methods. As the test dataset distribution deviates massively from the training dataset's distribution, we cannot simply fit a model to the training data and expect that it will work for the test data. Even a GroupKFold on the two stars doesn't help.\n\nThe public notebooks of today (August 25) are all overfitted and misleading (this criticism includes my own public notebooks). The competition asks us to submit two numbers for every planet and wavelength: the first number is the planet's transit depth and the second one (\\\\(\\sigma\\\\)) is a measure for the uncertainty of the first one. I don't criticize that these notebooks oversimplify the task by aggregating the data — that was a concious decision which is legitimate for a baseline model. The notebooks have two other, major, flaws:\n1. They predict the transit depth with a regression model although no regression is needed: Transit depth is defined as a simple measurement in the timeline (\\\\(\\frac{oot-it}{oot}\\\\)), and a regression model only overfits to the noise of the training data.\n2. Although some notebooks apply a correct method to compute \\\\(\\sigma\\\\) for the known stars, their \\\\(\\sigma\\\\) prediction for the unknown stars is derived from leaderboard probing. You need to find a method for computing the uncertainty \\\\(\\sigma\\\\) based on the given features rather than leaderboard probing. By the way, one notebook predicts \\\\(\\sigma=0.00016\\\\) for known stars and \\\\(\\sigma=0.00085\\\\) for unknown stars (five times higher uncertainty). These two numbers show how far away the test data are from the training data.\n\nThis competition is not so much about machine learning. It is much more about analyzing the data, managing noise and building up some domain knowledge.\n\n![oot](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2F2c0d649ed3a382aa9582b44e398fde22%2Foot-it.png?generation=1724573541810992&alt=media)",
      "votes": 27,
      "replies": [
        {
          "id": 2969843,
          "postDate": "2024-08-25T12:53:42.760Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> , what is your opinion about the host starting notebook and its «&nbsp;methodology&nbsp;» Monte Carlo dropout … and predictions ( small variation prediction + white curve prediction) ? </p>",
          "rawMarkdown": "Hi @ambrosm , what is your opinion about the host starting notebook and its « methodology » Monte Carlo dropout … and predictions ( small variation prediction + white curve prediction) ? ",
          "votes": 2,
          "replies": [
            {
              "id": 2970130,
              "postDate": "2024-08-25T18:35:38.073Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/pourchot\" target=\"_blank\">@pourchot</a> It's the first time I see this method. I'll probably give it a try.</p>",
              "rawMarkdown": "Hi @pourchot It's the first time I see this method. I'll probably give it a try."
            },
            {
              "id": 2999245,
              "postDate": "2024-09-26T14:18:04.320Z",
              "content": "<p>hi <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> what about genetic algorithms? I am trying that and getting similar results to the public ensemble at currently 0.545</p>",
              "rawMarkdown": "hi @ambrosm what about genetic algorithms? I am trying that and getting similar results to the public ensemble at currently 0.545"
            },
            {
              "id": 3009925,
              "postDate": "2024-10-08T14:23:24.987Z",
              "content": "<p>what are genetic algorithms, can you share some resources please.</p>",
              "rawMarkdown": "what are genetic algorithms, can you share some resources please."
            }
          ]
        },
        {
          "id": 2970286,
          "postDate": "2024-08-26T00:13:39.437Z",
          "content": "<p>I agree with this take! I think the varying σ is going to make things interesting </p>",
          "rawMarkdown": "I agree with this take! I think the varying σ is going to make things interesting ",
          "votes": 2
        }
      ]
    },
    {
      "id": 2969397,
      "postDate": "2024-08-24T23:03:14.237Z",
      "content": "<p>I wanted to put this post out just to get a feeling for other’s opinions. I was considering joining this competition, but I have mixed feelings with regard to the shakeup potential. The data is simulated, which would usually make me worry about data drift between separate runs of the simulation (could be private vs public). But I also have seen in discussions that there is a solidified pipeline for this in our case, which I think Would make it more stable? It may even be possible that the private versus public data is a different segment of one simulation. I’m certain I have missed details on this and maybe my thoughts are misguided so I was curious what others thought.</p>",
      "rawMarkdown": "I wanted to put this post out just to get a feeling for other’s opinions. I was considering joining this competition, but I have mixed feelings with regard to the shakeup potential. The data is simulated, which would usually make me worry about data drift between separate runs of the simulation (could be private vs public). But I also have seen in discussions that there is a solidified pipeline for this in our case, which I think Would make it more stable? It may even be possible that the private versus public data is a different segment of one simulation. I’m certain I have missed details on this and maybe my thoughts are misguided so I was curious what others thought.",
      "votes": 13
    },
    {
      "id": 2969531,
      "postDate": "2024-08-25T05:02:46.820Z",
      "content": "<p><a href=\"https://www.kaggle.com/cody11\" target=\"_blank\">@cody11</a> in my opinion non shakeup</p>\n<ol>\n<li>As per previous Ariel competitions, no shakeup happen. </li>\n<li>Need to train models to handle new test star similar to =&gt; star 0 vs star 1 in our train dataset.</li>\n</ol>",
      "rawMarkdown": "@cody11 in my opinion non shakeup\n1. As per previous Ariel competitions, no shakeup happen. \n2. Need to train models to handle new test star similar to => star 0 vs star 1 in our train dataset."
    }
  ],
  "comments": [
    {
      "id": 2969648,
      "author_name": "AmbrosM",
      "author_url": "",
      "post_date": "2024-08-25T08:13:47.533000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/cody11null\" target=\"_blank\">@cody11null</a>, I think it's possible to survive the shakeup if one applies the right methods. As the test dataset distribution deviates massively from the training dataset's distribution, we cannot simply fit a model to the training data and expect that it will work for the test data. Even a GroupKFold on the two stars doesn't help.</p>\n<p>The public notebooks of today (August 25) are all overfitted and misleading (this criticism includes my own public notebooks). The competition asks us to submit two numbers for every planet and wavelength: the first number is the planet's transit depth and the second one (\\(\\sigma\\)) is a measure for the uncertainty of the first one. I don't criticize that these notebooks oversimplify the task by aggregating the data — that was a concious decision which is legitimate for a baseline model. The notebooks have two other, major, flaws:</p>\n<ol>\n<li>They predict the transit depth with a regression model although no regression is needed: Transit depth is defined as a simple measurement in the timeline (\\(\\frac{oot-it}{oot}\\)), and a regression model only overfits to the noise of the training data.</li>\n<li>Although some notebooks apply a correct method to compute \\(\\sigma\\) for the known stars, their \\(\\sigma\\) prediction for the unknown stars is derived from leaderboard probing. You need to find a method for computing the uncertainty \\(\\sigma\\) based on the given features rather than leaderboard probing. By the way, one notebook predicts \\(\\sigma=0.00016\\) for known stars and \\(\\sigma=0.00085\\) for unknown stars (five times higher uncertainty). These two numbers show how far away the test data are from the training data.</li>\n</ol>\n<p>This competition is not so much about machine learning. It is much more about analyzing the data, managing noise and building up some domain knowledge.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2F2c0d649ed3a382aa9582b44e398fde22%2Foot-it.png?generation=1724573541810992&amp;alt=media\" alt=\"oot\"></p>",
      "votes": 27,
      "replies": [
        {
          "id": 2969843,
          "author_name": "Laurent Pourchot",
          "author_url": "",
          "post_date": "2024-08-25T12:53:42.760000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> , what is your opinion about the host starting notebook and its «&nbsp;methodology&nbsp;» Monte Carlo dropout … and predictions ( small variation prediction + white curve prediction) ? </p>",
          "votes": 2,
          "replies": [
            {
              "id": 2970130,
              "author_name": "AmbrosM",
              "author_url": "",
              "post_date": "2024-08-25T18:35:38.073000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/pourchot\" target=\"_blank\">@pourchot</a> It's the first time I see this method. I'll probably give it a try.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2999245,
              "author_name": "Octavi Grau",
              "author_url": "",
              "post_date": "2024-09-26T14:18:04.320000",
              "content": "<p>hi <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> what about genetic algorithms? I am trying that and getting similar results to the public ensemble at currently 0.545</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3009925,
              "author_name": "highDopamine",
              "author_url": "",
              "post_date": "2024-10-08T14:23:24.987000",
              "content": "<p>what are genetic algorithms, can you share some resources please.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2970286,
          "author_name": "Cody_Null",
          "author_url": "",
          "post_date": "2024-08-26T00:13:39.437000",
          "content": "<p>I agree with this take! I think the varying σ is going to make things interesting </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2969531,
      "author_name": "SeshuRaju 🧘‍♂️",
      "author_url": "",
      "post_date": "2024-08-25T05:02:46.820000",
      "content": "<p><a href=\"https://www.kaggle.com/cody11\" target=\"_blank\">@cody11</a> in my opinion non shakeup</p>\n<ol>\n<li>As per previous Ariel competitions, no shakeup happen. </li>\n<li>Need to train models to handle new test star similar to =&gt; star 0 vs star 1 in our train dataset.</li>\n</ol>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2969648": "Hi @cody11null, I think it's possible to survive the shakeup if one applies the right methods. As the test dataset distribution deviates massively from the training dataset's distribution, we cannot simply fit a model to the training data and expect that it will work for the test data. Even a GroupKFold on the two stars doesn't help.\n\nThe public notebooks of today (August 25) are all overfitted and misleading (this criticism includes my own public notebooks). The competition asks us to submit two numbers for every planet and wavelength: the first number is the planet's transit depth and the second one (\\\\(\\sigma\\\\)) is a measure for the uncertainty of the first one. I don't criticize that these notebooks oversimplify the task by aggregating the data — that was a concious decision which is legitimate for a baseline model. The notebooks have two other, major, flaws:\n1. They predict the transit depth with a regression model although no regression is needed: Transit depth is defined as a simple measurement in the timeline (\\\\(\\frac{oot-it}{oot}\\\\)), and a regression model only overfits to the noise of the training data.\n2. Although some notebooks apply a correct method to compute \\\\(\\sigma\\\\) for the known stars, their \\\\(\\sigma\\\\) prediction for the unknown stars is derived from leaderboard probing. You need to find a method for computing the uncertainty \\\\(\\sigma\\\\) based on the given features rather than leaderboard probing. By the way, one notebook predicts \\\\(\\sigma=0.00016\\\\) for known stars and \\\\(\\sigma=0.00085\\\\) for unknown stars (five times higher uncertainty). These two numbers show how far away the test data are from the training data.\n\nThis competition is not so much about machine learning. It is much more about analyzing the data, managing noise and building up some domain knowledge.\n\n![oot](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2F2c0d649ed3a382aa9582b44e398fde22%2Foot-it.png?generation=1724573541810992&alt=media)",
    "2969397": "I wanted to put this post out just to get a feeling for other’s opinions. I was considering joining this competition, but I have mixed feelings with regard to the shakeup potential. The data is simulated, which would usually make me worry about data drift between separate runs of the simulation (could be private vs public). But I also have seen in discussions that there is a solidified pipeline for this in our case, which I think Would make it more stable? It may even be possible that the private versus public data is a different segment of one simulation. I’m certain I have missed details on this and maybe my thoughts are misguided so I was curious what others thought.",
    "2969531": "@cody11 in my opinion non shakeup\n1. As per previous Ariel competitions, no shakeup happen. \n2. Need to train models to handle new test star similar to => star 0 vs star 1 in our train dataset."
  }
}