{
  "id": 528114,
  "title": "Understanding the competition metric",
  "url": "/competitions/ariel-data-challenge-2024/discussion/528114",
  "author_name": "AmbrosM",
  "post_date": "2024-08-14T19:52:52.721000",
  "votes": 62,
  "comment_count": 7,
  "views": 0,
  "content": "<p>In this competition, we submit a set of predicted targets \\(\\mu_{user}\\) and predicted uncertainties \\(\\sigma_{user}\\). The competition metric evaluates these predictions against the ground truth pixel level spectrum (\\(y\\)) using the Gaussian Log-likelihood (GLL) function.</p>\n<p>Astronomers (and businesspeople, too) like uncertainty quantification: They not only want to see predictions, they want to know as well how confident we are with these predictions. </p>\n<p>If we submit the same \\(\\sigma_{user}\\) for all samples, the formula for the public leaderboard score can be simplified to</p>\n<p>\\[\\large{score = \\frac {\\log \\sigma_{user}^2 + \\frac{MSE(y, \\mu_{user})}{\\sigma_{user}^2} + 9.16}{-23.03 + 9.16}}\\]</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2Fd88037edd3819571c6cc55811a847e23%2Fscore.png?generation=1723665074565505&amp;alt=media\" alt=\"score\"></p>\n<p>This means:</p>\n<ol>\n<li>The score is a (linearly transformed) sum of two terms.</li>\n<li>Whatever the value of \\(\\sigma_{user}\\) is, we want to optimize the mean squared error (MSE) of the \\(\\mu_{user}\\) predictions as usual. This is a standard regression task.</li>\n<li>After our regression model has predicted \\(\\mu_{user}\\) for every sample, we may choose a \\(\\sigma_{user}\\). If we choose a very high \\(\\sigma_{user}\\), the first term of the formula (\\(\\log \\sigma_{user}^2\\)) dominates. If we choose a low (positive) \\(\\sigma_{user}\\), the second term dominates. The optimum is somewhere in between (see diagram above).</li>\n<li>The diagram furthermore shows that a low \\(\\sigma_{user}\\) is punished more than one which is too high.</li>\n<li>The optimum  \\(\\sigma_{user}\\), which maximizes the score for a given MSE, is  \\(\\sigma_{user} = \\sqrt{MSE}\\). (To see this, compute the derivative with respect to  \\(\\sigma_{user}^2\\) and set it to zero.)</li>\n</ol>\n<p>A standard recipe to optimize this scoring function is: Predict \\(\\mu_{user}\\) with any regression model which optimizes MSE and then set \\(\\sigma_{user}\\) to the root mean squared error of the out-of-fold predictions. This is what I did in version 1 of my <a href=\"https://www.kaggle.com/code/ambrosm/adc24-intro-training\" target=\"_blank\">ADC24 Intro training</a> and <a href=\"https://www.kaggle.com/code/ambrosm/adc24-intro-inference\" target=\"_blank\">ADC24 Intro inference</a> notebooks: The oof RMSE was 0.000293, I set <code>sigma_pred=0.000293</code>, the cv score was 0.259 and the public leaderboard score was 0.114.</p>\n<p>Where does this huge cv–lb difference come from? The standard recipe doesn't take into account that train and test distributions are different. In the training dataset all planets belong to two host stars numbered 0 and 1. The test dataset is more diverse: In addition to stars 0 and 1, there is at least one other (previously unseen) star. Adapting the standard recipe to this situation, I'd propose the following:</p>\n<blockquote>\n  <p>For known stars set \\(\\sigma_{user}\\) to the out-of-fold rmse, and for unknown stars set it to a higher value. Perhaps use the out-of-fold rmse of a GroupKFold to estimate \\(\\sigma_{user}\\) for the unknown stars.</p>\n</blockquote>\n<p>EDIT: The recipe is only valid as a baseline. A more sophisticated model will of course predict individual sigmas for every planet and wavelength.</p>",
  "messages": [
    {
      "id": 2959341,
      "postDate": "2024-08-14T19:52:52.720Z",
      "content": "<p>In this competition, we submit a set of predicted targets \\(\\mu_{user}\\) and predicted uncertainties \\(\\sigma_{user}\\). The competition metric evaluates these predictions against the ground truth pixel level spectrum (\\(y\\)) using the Gaussian Log-likelihood (GLL) function.</p>\n<p>Astronomers (and businesspeople, too) like uncertainty quantification: They not only want to see predictions, they want to know as well how confident we are with these predictions. </p>\n<p>If we submit the same \\(\\sigma_{user}\\) for all samples, the formula for the public leaderboard score can be simplified to</p>\n<p>\\[\\large{score = \\frac {\\log \\sigma_{user}^2 + \\frac{MSE(y, \\mu_{user})}{\\sigma_{user}^2} + 9.16}{-23.03 + 9.16}}\\]</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2Fd88037edd3819571c6cc55811a847e23%2Fscore.png?generation=1723665074565505&amp;alt=media\" alt=\"score\"></p>\n<p>This means:</p>\n<ol>\n<li>The score is a (linearly transformed) sum of two terms.</li>\n<li>Whatever the value of \\(\\sigma_{user}\\) is, we want to optimize the mean squared error (MSE) of the \\(\\mu_{user}\\) predictions as usual. This is a standard regression task.</li>\n<li>After our regression model has predicted \\(\\mu_{user}\\) for every sample, we may choose a \\(\\sigma_{user}\\). If we choose a very high \\(\\sigma_{user}\\), the first term of the formula (\\(\\log \\sigma_{user}^2\\)) dominates. If we choose a low (positive) \\(\\sigma_{user}\\), the second term dominates. The optimum is somewhere in between (see diagram above).</li>\n<li>The diagram furthermore shows that a low \\(\\sigma_{user}\\) is punished more than one which is too high.</li>\n<li>The optimum  \\(\\sigma_{user}\\), which maximizes the score for a given MSE, is  \\(\\sigma_{user} = \\sqrt{MSE}\\). (To see this, compute the derivative with respect to  \\(\\sigma_{user}^2\\) and set it to zero.)</li>\n</ol>\n<p>A standard recipe to optimize this scoring function is: Predict \\(\\mu_{user}\\) with any regression model which optimizes MSE and then set \\(\\sigma_{user}\\) to the root mean squared error of the out-of-fold predictions. This is what I did in version 1 of my <a href=\"https://www.kaggle.com/code/ambrosm/adc24-intro-training\" target=\"_blank\">ADC24 Intro training</a> and <a href=\"https://www.kaggle.com/code/ambrosm/adc24-intro-inference\" target=\"_blank\">ADC24 Intro inference</a> notebooks: The oof RMSE was 0.000293, I set <code>sigma_pred=0.000293</code>, the cv score was 0.259 and the public leaderboard score was 0.114.</p>\n<p>Where does this huge cv–lb difference come from? The standard recipe doesn't take into account that train and test distributions are different. In the training dataset all planets belong to two host stars numbered 0 and 1. The test dataset is more diverse: In addition to stars 0 and 1, there is at least one other (previously unseen) star. Adapting the standard recipe to this situation, I'd propose the following:</p>\n<blockquote>\n  <p>For known stars set \\(\\sigma_{user}\\) to the out-of-fold rmse, and for unknown stars set it to a higher value. Perhaps use the out-of-fold rmse of a GroupKFold to estimate \\(\\sigma_{user}\\) for the unknown stars.</p>\n</blockquote>\n<p>EDIT: The recipe is only valid as a baseline. A more sophisticated model will of course predict individual sigmas for every planet and wavelength.</p>",
      "rawMarkdown": "In this competition, we submit a set of predicted targets \\\\(\\mu_{user}\\\\) and predicted uncertainties \\\\(\\sigma_{user}\\\\). The competition metric evaluates these predictions against the ground truth pixel level spectrum (\\\\(y\\\\)) using the Gaussian Log-likelihood (GLL) function.\n\nAstronomers (and businesspeople, too) like uncertainty quantification: They not only want to see predictions, they want to know as well how confident we are with these predictions. \n\nIf we submit the same \\\\(\\sigma_{user}\\\\) for all samples, the formula for the public leaderboard score can be simplified to\n\n\\\\[\\large{score = \\frac {\\log \\sigma_{user}^2 + \\frac{MSE(y, \\mu_{user})}{\\sigma_{user}^2} + 9.16}{-23.03 + 9.16}}\\\\]\n\n![score](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2Fd88037edd3819571c6cc55811a847e23%2Fscore.png?generation=1723665074565505&alt=media)\n\nThis means:\n1. The score is a (linearly transformed) sum of two terms.\n2. Whatever the value of \\\\(\\sigma_{user}\\\\) is, we want to optimize the mean squared error (MSE) of the \\\\(\\mu_{user}\\\\) predictions as usual. This is a standard regression task.\n3. After our regression model has predicted \\\\(\\mu_{user}\\\\) for every sample, we may choose a \\\\(\\sigma_{user}\\\\). If we choose a very high \\\\(\\sigma_{user}\\\\), the first term of the formula (\\\\(\\log \\sigma_{user}^2\\\\)) dominates. If we choose a low (positive) \\\\(\\sigma_{user}\\\\), the second term dominates. The optimum is somewhere in between (see diagram above).\n4. The diagram furthermore shows that a low \\\\(\\sigma_{user}\\\\) is punished more than one which is too high.\n4. The optimum  \\\\(\\sigma_{user}\\\\), which maximizes the score for a given MSE, is  \\\\(\\sigma_{user} = \\sqrt{MSE}\\\\). (To see this, compute the derivative with respect to  \\\\(\\sigma_{user}^2\\\\) and set it to zero.)\n\nA standard recipe to optimize this scoring function is: Predict \\\\(\\mu_{user}\\\\) with any regression model which optimizes MSE and then set \\\\(\\sigma_{user}\\\\) to the root mean squared error of the out-of-fold predictions. This is what I did in version 1 of my [ADC24 Intro training](https://www.kaggle.com/code/ambrosm/adc24-intro-training) and [ADC24 Intro inference](https://www.kaggle.com/code/ambrosm/adc24-intro-inference) notebooks: The oof RMSE was 0.000293, I set `sigma_pred=0.000293`, the cv score was 0.259 and the public leaderboard score was 0.114.\n\nWhere does this huge cv–lb difference come from? The standard recipe doesn't take into account that train and test distributions are different. In the training dataset all planets belong to two host stars numbered 0 and 1. The test dataset is more diverse: In addition to stars 0 and 1, there is at least one other (previously unseen) star. Adapting the standard recipe to this situation, I'd propose the following:\n\n> For known stars set \\\\(\\sigma_{user}\\\\) to the out-of-fold rmse, and for unknown stars set it to a higher value. Perhaps use the out-of-fold rmse of a GroupKFold to estimate \\\\(\\sigma_{user}\\\\) for the unknown stars.\n\nEDIT: The recipe is only valid as a baseline. A more sophisticated model will of course predict individual sigmas for every planet and wavelength.\n",
      "votes": 61
    },
    {
      "id": 2960975,
      "postDate": "2024-08-16T10:18:17.833Z",
      "content": "<p><a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> Thanks you for publishing comprehensive guide.<br>\nI think I found mistake on your calculation. </p>\n<blockquote>\n  <p>For L_ref is defined using the mean and variance of the training dataset as its prediction for all instances.</p>\n</blockquote>\n<p>I calculated L_ref in your customized GLL, it should be -11.72 not -9.16.<br>\nWhen using this value, the CV of your score 0.259 becomes 0.091 which is closer to your LB score.<br>\nFeel free to point out if I make mistakes.</p>\n<pre><code> polars  pl\n\n ():\n    var = sigmas**\n    mse = (preds - targets) ** \n    score = np.log(var) + mse / var\n     score.mean()\n\n\n ():\n    gll_simple = score * (- + ) - \n    score = (gll_simple - gll_trivial) / (gll_opt - gll_trivial)\n     score\n\n\ntargets = (\n    pl.read_csv()\n    .drop()\n    .to_numpy()\n)\nx_mean = targets.mean(axis=, keepdims=)\nx_std = targets.std(axis=, keepdims=)\n\n\n(calc_simple_metric(x_mean, x_std, targets))  \n\n\n(calc_simple_metric(targets, , targets))  \n\ncorrect_score()  \n</code></pre>",
      "rawMarkdown": "@ambrosm Thanks you for publishing comprehensive guide.\nI think I found mistake on your calculation. \n\n> For L_ref is defined using the mean and variance of the training dataset as its prediction for all instances.\n\nI calculated L_ref in your customized GLL, it should be -11.72 not -9.16.\nWhen using this value, the CV of your score 0.259 becomes 0.091 which is closer to your LB score.\nFeel free to point out if I make mistakes.\n\n```python\nimport polars as pl\n\ndef calc_simple_metric(preds, sigmas, targets):\n    var = sigmas**2\n    mse = (preds - targets) ** 2\n    score = np.log(var) + mse / var\n    return score.mean()\n\n\ndef correct_score(score, gll_trivial=-11.72408405675161, gll_opt=-23.025850929940457):\n    gll_simple = score * (-23.03 + 9.16) - 9.16\n    score = (gll_simple - gll_trivial) / (gll_opt - gll_trivial)\n    return score\n\n\ntargets = (\n    pl.read_csv(\"./ariel-data-challenge-2024/train_labels.csv\")\n    .drop(\"planet_id\")\n    .to_numpy()\n)\nx_mean = targets.mean(axis=0, keepdims=True)\nx_std = targets.std(axis=0, keepdims=True)\n\n# L_ref\nprint(calc_simple_metric(x_mean, x_std, targets))  # -11.72408405675161\n\n# L_opt\nprint(calc_simple_metric(targets, 1e-5, targets))  # -23.025850929940457\n\ncorrect_score(0.259)  # 0.0909809903872373\n```",
      "votes": 8,
      "replies": [
        {
          "id": 2963138,
          "postDate": "2024-08-18T12:01:16.027Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a>, thank you for verifying my calculation. I neglected to mention that there are three different values for L_ref: These are L_ref_train, L_ref_publictest, L_ref_privatetest. The difference comes from evaluating mean and variance of the training dataset against the true labels of the training dataset, the public test set or the private test set, respectively.</p>\n<ul>\n<li>L_ref_train is -11.72, as in your calculation.</li>\n<li>For L_ref_publictest, my estimation is -9.16.</li>\n<li>I have no idea of L_ref_privatetest.</li>\n</ul>",
          "rawMarkdown": "Hi @tatamikenn, thank you for verifying my calculation. I neglected to mention that there are three different values for L_ref: These are L_ref_train, L_ref_publictest, L_ref_privatetest. The difference comes from evaluating mean and variance of the training dataset against the true labels of the training dataset, the public test set or the private test set, respectively.\n- L_ref_train is -11.72, as in your calculation.\n- For L_ref_publictest, my estimation is -9.16.\n- I have no idea of L_ref_privatetest.",
          "votes": 4,
          "replies": [
            {
              "id": 2963366,
              "postDate": "2024-08-18T16:23:03.360Z",
              "content": "<p><a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> How do you calculated L_re_publictest estimation <strong>-9.16</strong> ?</p>",
              "rawMarkdown": "@ambrosm How do you calculated L_re_publictest estimation **-9.16** ?",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2959362,
      "postDate": "2024-08-14T20:14:09.733Z",
      "content": "<p>Great rundown of the metric! I am quite excited about this competition and digging a bit through the data now to get started.</p>\n<blockquote>\n  <p>In addition to stars 0 and 1, there is at least one other (previously unseen) star. Adapting the standard recipe to this situation, I'd propose the following</p>\n</blockquote>\n<p>This is the first time that I read it like this. Was it confirmed that test includes star 0 and star 1 from training? It sounded like it contains \"only\" a new solar system.</p>",
      "rawMarkdown": "Great rundown of the metric! I am quite excited about this competition and digging a bit through the data now to get started.\n\n>  In addition to stars 0 and 1, there is at least one other (previously unseen) star. Adapting the standard recipe to this situation, I'd propose the following\n\nThis is the first time that I read it like this. Was it confirmed that test includes star 0 and star 1 from training? It sounded like it contains \"only\" a new solar system.\n",
      "votes": 5,
      "replies": [
        {
          "id": 2959687,
          "postDate": "2024-08-15T07:19:07.383Z",
          "content": "<p><a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> <a href=\"https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/524287#2950440\" target=\"_blank\">Host conformed</a> about new star for 327 exoplanets, so other 473 need to be belongs to star 0 / 1 / both</p>\n<blockquote>\n  <p>Its having 2 solar systems =&gt; first with 346 exoplanets and second with 327 exoplanets, is remaining 327 exoplanets (test) belongs to same solar systems or new solar system?<br>\n  The 327 belongs to a new planetary system (or new host star)<br>\n  800 - 327 =&gt; 473 exoplanets belongs to star 0 / 1 / both  </p>\n</blockquote>",
          "rawMarkdown": "@ilu000 [Host conformed](https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/524287#2950440) about new star for 327 exoplanets, so other 473 need to be belongs to star 0 / 1 / both\n\n> Its having 2 solar systems => first with 346 exoplanets and second with 327 exoplanets, is remaining 327 exoplanets (test) belongs to same solar systems or new solar system?\nThe 327 belongs to a new planetary system (or new host star)\n800 - 327 => 473 exoplanets belongs to star 0 / 1 / both  ",
          "votes": 3
        },
        {
          "id": 2959789,
          "postDate": "2024-08-15T09:53:49.737Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a>, the test set includes both previously seen and previously unseen stars.</p>",
          "rawMarkdown": "Hi @ilu000, the test set includes both previously seen and previously unseen stars.",
          "votes": 8
        }
      ]
    },
    {
      "id": 2994368,
      "postDate": "2024-09-20T20:39:00.970Z",
      "content": "<p>Thanks for the clear, simple presentation!<br>\nOne  comment: the \\(L_{ideal}\\) is based on \\(MSE/\\sigma^2=0\\), fair enough. But in practice this term will be at least \\((\\sigma_{sim}/10^{-5})^2\\), where \\(\\sigma_{sim}\\) is the error term in the simulated data. So, for \\(\\sigma_{sim}=10^{-5}\\) the best score is about 0.93</p>",
      "rawMarkdown": "Thanks for the clear, simple presentation!\nOne  comment: the \\\\(L_{ideal}\\\\) is based on \\\\(MSE/\\sigma^2=0\\\\), fair enough. But in practice this term will be at least \\\\((\\sigma_{sim}/10^{-5})^2\\\\), where \\\\(\\sigma_{sim}\\\\) is the error term in the simulated data. So, for \\\\(\\sigma_{sim}=10^{-5}\\\\) the best score is about 0.93",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 2960975,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2024-08-16T10:18:17.833000",
      "content": "<p><a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> Thanks you for publishing comprehensive guide.<br>\nI think I found mistake on your calculation. </p>\n<blockquote>\n  <p>For L_ref is defined using the mean and variance of the training dataset as its prediction for all instances.</p>\n</blockquote>\n<p>I calculated L_ref in your customized GLL, it should be -11.72 not -9.16.<br>\nWhen using this value, the CV of your score 0.259 becomes 0.091 which is closer to your LB score.<br>\nFeel free to point out if I make mistakes.</p>\n<pre><code> polars  pl\n\n ():\n    var = sigmas**\n    mse = (preds - targets) ** \n    score = np.log(var) + mse / var\n     score.mean()\n\n\n ():\n    gll_simple = score * (- + ) - \n    score = (gll_simple - gll_trivial) / (gll_opt - gll_trivial)\n     score\n\n\ntargets = (\n    pl.read_csv()\n    .drop()\n    .to_numpy()\n)\nx_mean = targets.mean(axis=, keepdims=)\nx_std = targets.std(axis=, keepdims=)\n\n\n(calc_simple_metric(x_mean, x_std, targets))  \n\n\n(calc_simple_metric(targets, , targets))  \n\ncorrect_score()  \n</code></pre>",
      "votes": 8,
      "replies": [
        {
          "id": 2963138,
          "author_name": "AmbrosM",
          "author_url": "",
          "post_date": "2024-08-18T12:01:16.027000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a>, thank you for verifying my calculation. I neglected to mention that there are three different values for L_ref: These are L_ref_train, L_ref_publictest, L_ref_privatetest. The difference comes from evaluating mean and variance of the training dataset against the true labels of the training dataset, the public test set or the private test set, respectively.</p>\n<ul>\n<li>L_ref_train is -11.72, as in your calculation.</li>\n<li>For L_ref_publictest, my estimation is -9.16.</li>\n<li>I have no idea of L_ref_privatetest.</li>\n</ul>",
          "votes": 4,
          "replies": [
            {
              "id": 2963366,
              "author_name": "SeshuRaju 🧘‍♂️",
              "author_url": "",
              "post_date": "2024-08-18T16:23:03.360000",
              "content": "<p><a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> How do you calculated L_re_publictest estimation <strong>-9.16</strong> ?</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2959362,
      "author_name": "Pascal Pfeiffer",
      "author_url": "",
      "post_date": "2024-08-14T20:14:09.733000",
      "content": "<p>Great rundown of the metric! I am quite excited about this competition and digging a bit through the data now to get started.</p>\n<blockquote>\n  <p>In addition to stars 0 and 1, there is at least one other (previously unseen) star. Adapting the standard recipe to this situation, I'd propose the following</p>\n</blockquote>\n<p>This is the first time that I read it like this. Was it confirmed that test includes star 0 and star 1 from training? It sounded like it contains \"only\" a new solar system.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 2959687,
          "author_name": "SeshuRaju 🧘‍♂️",
          "author_url": "",
          "post_date": "2024-08-15T07:19:07.383000",
          "content": "<p><a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> <a href=\"https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/524287#2950440\" target=\"_blank\">Host conformed</a> about new star for 327 exoplanets, so other 473 need to be belongs to star 0 / 1 / both</p>\n<blockquote>\n  <p>Its having 2 solar systems =&gt; first with 346 exoplanets and second with 327 exoplanets, is remaining 327 exoplanets (test) belongs to same solar systems or new solar system?<br>\n  The 327 belongs to a new planetary system (or new host star)<br>\n  800 - 327 =&gt; 473 exoplanets belongs to star 0 / 1 / both  </p>\n</blockquote>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 2959789,
          "author_name": "AmbrosM",
          "author_url": "",
          "post_date": "2024-08-15T09:53:49.737000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a>, the test set includes both previously seen and previously unseen stars.</p>",
          "votes": 8,
          "replies": []
        }
      ]
    },
    {
      "id": 2994368,
      "author_name": "Daniel Dewey",
      "author_url": "",
      "post_date": "2024-09-20T20:39:00.970000",
      "content": "<p>Thanks for the clear, simple presentation!<br>\nOne  comment: the \\(L_{ideal}\\) is based on \\(MSE/\\sigma^2=0\\), fair enough. But in practice this term will be at least \\((\\sigma_{sim}/10^{-5})^2\\), where \\(\\sigma_{sim}\\) is the error term in the simulated data. So, for \\(\\sigma_{sim}=10^{-5}\\) the best score is about 0.93</p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2959341": "In this competition, we submit a set of predicted targets \\\\(\\mu_{user}\\\\) and predicted uncertainties \\\\(\\sigma_{user}\\\\). The competition metric evaluates these predictions against the ground truth pixel level spectrum (\\\\(y\\\\)) using the Gaussian Log-likelihood (GLL) function.\n\nAstronomers (and businesspeople, too) like uncertainty quantification: They not only want to see predictions, they want to know as well how confident we are with these predictions. \n\nIf we submit the same \\\\(\\sigma_{user}\\\\) for all samples, the formula for the public leaderboard score can be simplified to\n\n\\\\[\\large{score = \\frac {\\log \\sigma_{user}^2 + \\frac{MSE(y, \\mu_{user})}{\\sigma_{user}^2} + 9.16}{-23.03 + 9.16}}\\\\]\n\n![score](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2Fd88037edd3819571c6cc55811a847e23%2Fscore.png?generation=1723665074565505&alt=media)\n\nThis means:\n1. The score is a (linearly transformed) sum of two terms.\n2. Whatever the value of \\\\(\\sigma_{user}\\\\) is, we want to optimize the mean squared error (MSE) of the \\\\(\\mu_{user}\\\\) predictions as usual. This is a standard regression task.\n3. After our regression model has predicted \\\\(\\mu_{user}\\\\) for every sample, we may choose a \\\\(\\sigma_{user}\\\\). If we choose a very high \\\\(\\sigma_{user}\\\\), the first term of the formula (\\\\(\\log \\sigma_{user}^2\\\\)) dominates. If we choose a low (positive) \\\\(\\sigma_{user}\\\\), the second term dominates. The optimum is somewhere in between (see diagram above).\n4. The diagram furthermore shows that a low \\\\(\\sigma_{user}\\\\) is punished more than one which is too high.\n4. The optimum  \\\\(\\sigma_{user}\\\\), which maximizes the score for a given MSE, is  \\\\(\\sigma_{user} = \\sqrt{MSE}\\\\). (To see this, compute the derivative with respect to  \\\\(\\sigma_{user}^2\\\\) and set it to zero.)\n\nA standard recipe to optimize this scoring function is: Predict \\\\(\\mu_{user}\\\\) with any regression model which optimizes MSE and then set \\\\(\\sigma_{user}\\\\) to the root mean squared error of the out-of-fold predictions. This is what I did in version 1 of my [ADC24 Intro training](https://www.kaggle.com/code/ambrosm/adc24-intro-training) and [ADC24 Intro inference](https://www.kaggle.com/code/ambrosm/adc24-intro-inference) notebooks: The oof RMSE was 0.000293, I set `sigma_pred=0.000293`, the cv score was 0.259 and the public leaderboard score was 0.114.\n\nWhere does this huge cv–lb difference come from? The standard recipe doesn't take into account that train and test distributions are different. In the training dataset all planets belong to two host stars numbered 0 and 1. The test dataset is more diverse: In addition to stars 0 and 1, there is at least one other (previously unseen) star. Adapting the standard recipe to this situation, I'd propose the following:\n\n> For known stars set \\\\(\\sigma_{user}\\\\) to the out-of-fold rmse, and for unknown stars set it to a higher value. Perhaps use the out-of-fold rmse of a GroupKFold to estimate \\\\(\\sigma_{user}\\\\) for the unknown stars.\n\nEDIT: The recipe is only valid as a baseline. A more sophisticated model will of course predict individual sigmas for every planet and wavelength.\n",
    "2960975": "@ambrosm Thanks you for publishing comprehensive guide.\nI think I found mistake on your calculation. \n\n> For L_ref is defined using the mean and variance of the training dataset as its prediction for all instances.\n\nI calculated L_ref in your customized GLL, it should be -11.72 not -9.16.\nWhen using this value, the CV of your score 0.259 becomes 0.091 which is closer to your LB score.\nFeel free to point out if I make mistakes.\n\n```python\nimport polars as pl\n\ndef calc_simple_metric(preds, sigmas, targets):\n    var = sigmas**2\n    mse = (preds - targets) ** 2\n    score = np.log(var) + mse / var\n    return score.mean()\n\n\ndef correct_score(score, gll_trivial=-11.72408405675161, gll_opt=-23.025850929940457):\n    gll_simple = score * (-23.03 + 9.16) - 9.16\n    score = (gll_simple - gll_trivial) / (gll_opt - gll_trivial)\n    return score\n\n\ntargets = (\n    pl.read_csv(\"./ariel-data-challenge-2024/train_labels.csv\")\n    .drop(\"planet_id\")\n    .to_numpy()\n)\nx_mean = targets.mean(axis=0, keepdims=True)\nx_std = targets.std(axis=0, keepdims=True)\n\n# L_ref\nprint(calc_simple_metric(x_mean, x_std, targets))  # -11.72408405675161\n\n# L_opt\nprint(calc_simple_metric(targets, 1e-5, targets))  # -23.025850929940457\n\ncorrect_score(0.259)  # 0.0909809903872373\n```",
    "2959362": "Great rundown of the metric! I am quite excited about this competition and digging a bit through the data now to get started.\n\n>  In addition to stars 0 and 1, there is at least one other (previously unseen) star. Adapting the standard recipe to this situation, I'd propose the following\n\nThis is the first time that I read it like this. Was it confirmed that test includes star 0 and star 1 from training? It sounded like it contains \"only\" a new solar system.\n",
    "2994368": "Thanks for the clear, simple presentation!\nOne  comment: the \\\\(L_{ideal}\\\\) is based on \\\\(MSE/\\sigma^2=0\\\\), fair enough. But in practice this term will be at least \\\\((\\sigma_{sim}/10^{-5})^2\\\\), where \\\\(\\sigma_{sim}\\\\) is the error term in the simulated data. So, for \\\\(\\sigma_{sim}=10^{-5}\\\\) the best score is about 0.93"
  }
}