{
  "id": 506490,
  "title": "max=2523σ: How Extreme this Competition's Setup is",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490",
  "author_name": "Bilzard",
  "post_date": "2024-05-22T05:32:25.707000",
  "votes": 42,
  "comment_count": 25,
  "views": 0,
  "content": "<p>Hi. Hope you and your fellow Sisyphus are doing well with carrying the rock.</p>\n<p>I'll show some evidence of this competition's extreme setup.</p>\n<h2>TL; DR</h2>\n<p>some target columns are significantly risky.<br>\nIf one single outlier sample could cause significant drop on private LB (~-12 with single sample).</p>\n<h2>How these graphs are created</h2>\n<p>I plotted the risk of all 368 target columns except trivial columns with σ=0.<br>\nThe risk is calculated by</p>\n<p>$$<br>\n\\text{risk} := \\max(y) / \\sigma(y)<br>\n$$</p>\n<p>The first and 2nd picture shows some columns has significantly distant outliers (up to ~2500 σ).</p>\n<p>This could cause your model significantly drops in private LB.<br>\nThe 3rd picture depicts this outlier case.<br>\nSuppose if your model predict single outlier in private test set with ptend_q0002_27 = 0 (because 99.9999% has &lt;= 0 values), whose actual value is 2523σ, you will get <code>current score - 12.24</code> in private LB.</p>\n<p>So hope your model can handle these outliers well.<br>\n(I also hope host's already removed these miserable outliers from test set).</p>\n<p>Good luck!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fd7001b68bae626d71e9e16da10f86f74%2Frisk_graph.jpeg?generation=1716355019829983&amp;alt=media\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fe359364d9949eb7c67a8c1e42496dd6b%2FScreenshot%202024-05-22%20at%2014.16.39.png?generation=1716355036821574&amp;alt=media\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F30d088e4934ee91f6b47c4862217b3de%2FScreenshot%202024-05-22%20at%2014.16.47.png?generation=1716355049231387&amp;alt=media\"></p>",
  "messages": [
    {
      "id": 2828449,
      "postDate": "2024-05-22T05:32:25.707Z",
      "content": "<p>Hi. Hope you and your fellow Sisyphus are doing well with carrying the rock.</p>\n<p>I'll show some evidence of this competition's extreme setup.</p>\n<h2>TL; DR</h2>\n<p>some target columns are significantly risky.<br>\nIf one single outlier sample could cause significant drop on private LB (~-12 with single sample).</p>\n<h2>How these graphs are created</h2>\n<p>I plotted the risk of all 368 target columns except trivial columns with σ=0.<br>\nThe risk is calculated by</p>\n<p>$$<br>\n\\text{risk} := \\max(y) / \\sigma(y)<br>\n$$</p>\n<p>The first and 2nd picture shows some columns has significantly distant outliers (up to ~2500 σ).</p>\n<p>This could cause your model significantly drops in private LB.<br>\nThe 3rd picture depicts this outlier case.<br>\nSuppose if your model predict single outlier in private test set with ptend_q0002_27 = 0 (because 99.9999% has &lt;= 0 values), whose actual value is 2523σ, you will get <code>current score - 12.24</code> in private LB.</p>\n<p>So hope your model can handle these outliers well.<br>\n(I also hope host's already removed these miserable outliers from test set).</p>\n<p>Good luck!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fd7001b68bae626d71e9e16da10f86f74%2Frisk_graph.jpeg?generation=1716355019829983&amp;alt=media\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fe359364d9949eb7c67a8c1e42496dd6b%2FScreenshot%202024-05-22%20at%2014.16.39.png?generation=1716355036821574&amp;alt=media\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F30d088e4934ee91f6b47c4862217b3de%2FScreenshot%202024-05-22%20at%2014.16.47.png?generation=1716355049231387&amp;alt=media\"></p>",
      "rawMarkdown": "Hi. Hope you and your fellow Sisyphus are doing well with carrying the rock.\n\nI'll show some evidence of this competition's extreme setup.\n\n## TL; DR\n\nsome target columns are significantly risky.\nIf one single outlier sample could cause significant drop on private LB (~-12 with single sample).\n\n## How these graphs are created\n\nI plotted the risk of all 368 target columns except trivial columns with σ=0.\nThe risk is calculated by\n\n$$\n\\text{risk} := \\max(y) / \\sigma(y)\n$$\n\nThe first and 2nd picture shows some columns has significantly distant outliers (up to ~2500 σ).\n\nThis could cause your model significantly drops in private LB.\nThe 3rd picture depicts this outlier case.\nSuppose if your model predict single outlier in private test set with ptend_q0002_27 = 0 (because 99.9999% has <= 0 values), whose actual value is 2523σ, you will get `current score - 12.24` in private LB.\n\nSo hope your model can handle these outliers well.\n(I also hope host's already removed these miserable outliers from test set).\n\nGood luck!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fd7001b68bae626d71e9e16da10f86f74%2Frisk_graph.jpeg?generation=1716355019829983&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fe359364d9949eb7c67a8c1e42496dd6b%2FScreenshot%202024-05-22%20at%2014.16.39.png?generation=1716355036821574&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F30d088e4934ee91f6b47c4862217b3de%2FScreenshot%202024-05-22%20at%2014.16.47.png?generation=1716355049231387&alt=media)",
      "votes": 42
    },
    {
      "id": 2834423,
      "postDate": "2024-05-24T18:09:36.983Z",
      "content": "<p>Hello, thank you for providing this detailed analysis of the dataset and sharing your concerns. If it helps, the private leaderboard score is only half a percent different from the public leaderboard score for the u-net benchmark submission.</p>",
      "rawMarkdown": "Hello, thank you for providing this detailed analysis of the dataset and sharing your concerns. If it helps, the private leaderboard score is only half a percent different from the public leaderboard score for the u-net benchmark submission.",
      "votes": 14,
      "replies": [
        {
          "id": 2834437,
          "postDate": "2024-05-24T18:19:40.280Z",
          "content": "<p>That's good to know but I don't think it completely negates the risk of getting negative score.</p>",
          "rawMarkdown": "That's good to know but I don't think it completely negates the risk of getting negative score.",
          "votes": 2
        },
        {
          "id": 2834672,
          "postDate": "2024-05-24T22:41:18.327Z",
          "content": "<p><a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> Thanks. It somehow helped.</p>\n<p>However, as Gunes pointed out, there left some risk of our (i.e. Kaggle community's) best performant model would be significantly affected by tiny amount of extreme outliers.</p>\n<p>Is it possible to consider to take further actions like listed in [1]?<br>\nFor example, if we have the baseline model's prediction for all test set (i.e. <code>submission.csv</code>), we can test our model's safeness by checking if our model's predictions are not significantly apart from this prediction.</p>\n<h2>Reference</h2>\n<p>[1] <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490#2832998\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490#2832998</a></p>",
          "rawMarkdown": "@jerrylin96 Thanks. It somehow helped.\n\nHowever, as Gunes pointed out, there left some risk of our (i.e. Kaggle community's) best performant model would be significantly affected by tiny amount of extreme outliers.\n\nIs it possible to consider to take further actions like listed in [1]?\nFor example, if we have the baseline model's prediction for all test set (i.e. `submission.csv`), we can test our model's safeness by checking if our model's predictions are not significantly apart from this prediction.\n\n## Reference\n\n[1] https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490#2832998",
          "votes": 2,
          "replies": [
            {
              "id": 2835414,
              "postDate": "2024-05-25T10:03:05.023Z",
              "content": "<p>Also can we get details about the experimental setup (train/test size and etc.) of the u-net model?</p>",
              "rawMarkdown": "Also can we get details about the experimental setup (train/test size and etc.) of the u-net model?",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2832998,
      "postDate": "2024-05-24T02:58:01.050Z",
      "content": "<h2>Request for Some Action</h2>\n<p>I think generally machine learning models under current competition metrics (R2) is prone to some kind of extreme outliers (i.e. 2523σ).<br>\nIf this kind of extreme outliers are in private test set, it possibly causes catastrophic shakeup in the end of the competition.</p>\n<p>Hopefully, we believe some safeguard may be helpful for all of us. like, </p>\n<ol>\n<li>clipping extreme error in prediction error like <a href=\"https://www.kaggle.com/martynoveduard\" target=\"_blank\">@martynoveduard</a> proposed[1]</li>\n<li>Fixing to use customized R2 score with softer normalization policy like <a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a>  proposed[2]</li>\n<li>Providing baseline model's prediction (i.e. submission.csv) on test set and claim that it has decent score on private test set like I proposed[3]</li>\n</ol>\n<p>c.f.) I also posted mathematical explanation for instability in R2 metrics[4].</p>\n<p><a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> <a href=\"https://www.kaggle.com/ashleychow\" target=\"_blank\">@ashleychow</a> </p>\n<p>How hosts and Kaggle team think of this risk?</p>\n<p>Thanks in advance.</p>\n<h2>Reference</h2>\n<ul>\n<li>[1] <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490#2829249\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490#2829249</a></li>\n<li>[2] <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490#2829800\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490#2829800</a></li>\n<li>[3] <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490#2829931\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490#2829931</a></li>\n<li>[4] <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490#2829926\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490#2829926</a></li>\n</ul>",
      "rawMarkdown": "## Request for Some Action\n\nI think generally machine learning models under current competition metrics (R2) is prone to some kind of extreme outliers (i.e. 2523σ).\nIf this kind of extreme outliers are in private test set, it possibly causes catastrophic shakeup in the end of the competition.\n\nHopefully, we believe some safeguard may be helpful for all of us. like, \n\n1. clipping extreme error in prediction error like @martynoveduard proposed[1]\n2. Fixing to use customized R2 score with softer normalization policy like @shlomoron  proposed[2]\n3. Providing baseline model's prediction (i.e. submission.csv) on test set and claim that it has decent score on private test set like I proposed[3]\n\nc.f.) I also posted mathematical explanation for instability in R2 metrics[4].\n\n@jerrylin96 @ashleychow \n\nHow hosts and Kaggle team think of this risk?\n\nThanks in advance.\n\n## Reference\n\n- [1] https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490#2829249\n- [2] https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490#2829800\n- [3] https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490#2829931\n- [4] https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490#2829926",
      "votes": 9
    },
    {
      "id": 2829249,
      "postDate": "2024-05-22T13:57:24.570Z",
      "content": "<p>amazing study</p>\n<p>suggestion:<br>\nmacro-avg(r2-score) -&gt; macro-avg(clip(r2-score,0,1))</p>\n<p>this way even if the participant doesn't handle the outlier, he is losing ~0.003 per-column compared to the perfectly predicted column (it hurts, but better than to have -10 score in private)</p>",
      "rawMarkdown": "amazing study\n\nsuggestion:\nmacro-avg(r2-score) -> macro-avg(clip(r2-score,0,1))\n\nthis way even if the participant doesn't handle the outlier, he is losing ~0.003 per-column compared to the perfectly predicted column (it hurts, but better than to have -10 score in private)",
      "votes": 6,
      "replies": [
        {
          "id": 2829800,
          "postDate": "2024-05-22T20:05:26.670Z",
          "content": "<p>I will suggest a fix of my own then.<br>\nDefinition of r2 (I will use Wikipedia notation):<br>\nr2 = 1-SS(res)/SS(tot) </p>\n<p>The problem: since SS(tot) is calculated only for the test set, there can be a column with very small values for the private test set, leading to SS(tot) very small, leading to a negative explosion of r2 for this column.<br>\nFix: Use SS(tot) that is calculated for all the data. So instead of SS(tot)(test): <br>\n(n(#samples in test)/N(#all samples))*SS(tot)(all samples)<br>\nBasically, ensure that SS(tot) has reasonable values. Actually, the values should just by n since the values are normalized by their SD. So it can be reduced to:<br>\nr2 = 1-SS(res)/n(#samples in test)<br>\nI think? Well, something along those lines. The general idea should be correct.</p>",
          "rawMarkdown": "I will suggest a fix of my own then.\nDefinition of r2 (I will use Wikipedia notation):\nr2 = 1-SS(res)/SS(tot) \n\nThe problem: since SS(tot) is calculated only for the test set, there can be a column with very small values for the private test set, leading to SS(tot) very small, leading to a negative explosion of r2 for this column.\nFix: Use SS(tot) that is calculated for all the data. So instead of SS(tot)(test): \n(n(#samples in test)/N(#all samples))*SS(tot)(all samples)\nBasically, ensure that SS(tot) has reasonable values. Actually, the values should just by n since the values are normalized by their SD. So it can be reduced to:\nr2 = 1-SS(res)/n(#samples in test)\nI think? Well, something along those lines. The general idea should be correct.",
          "votes": 3,
          "replies": [
            {
              "id": 2829926,
              "postDate": "2024-05-22T22:30:34.420Z",
              "content": "<p>First of all, we should claim the relationship between $R^2$ metrics and standard deviation of given dataset. By definition,</p>\n<p>$$<br>\n\\begin{align*}<br>\nR^2 &amp;= 1 - \\frac{\\frac{1}{N} \\sum_i (y_i - \\hat{y}_i)^2}{\\frac{1}{N} \\sum_i (y_i - \\bar{y})^2} \\\\<br>\n&amp;=1 - \\frac{\\frac{1}{N} \\sum_i (y_i - \\hat{y}_i)^2}{\\sigma_t^2} \\\\<br>\n&amp;=1 - \\frac{1}{N} \\sum_i (y_i/\\sigma_t - \\hat{y}_i / \\sigma_t)^2 \\\\<br>\n\\end{align*}<br>\n$$</p>\n<p>That's why I normalized all data with standard deviation of all train data. And the above metric <strong>should be affected by outlier even if we use S(tot)_train instead of S(tot)_test</strong> i.e. <strong>sample with y_i=2523σ busts R2 score</strong>.</p>",
              "rawMarkdown": "First of all, we should claim the relationship between $R^2$ metrics and standard deviation of given dataset. By definition,\n\n$$\n\\begin{align\\*}\nR^2 &= 1 - \\frac{\\frac{1}{N} \\sum_i (y_i - \\hat{y}_i)^2}{\\frac{1}{N} \\sum_i (y_i - \\bar{y})^2} \\\\\\\\\n&=1 - \\frac{\\frac{1}{N} \\sum_i (y_i - \\hat{y}_i)^2}{\\sigma_t^2} \\\\\\\\\n&=1 - \\frac{1}{N} \\sum_i (y_i/\\sigma_t - \\hat{y}_i / \\sigma_t)^2 \\\\\\\\\n\\end{align\\*}\n$$\n\nThat's why I normalized all data with standard deviation of all train data. And the above metric **should be affected by outlier even if we use S(tot)_train instead of S(tot)_test** i.e. **sample with y_i=2523σ busts R2 score**.",
              "votes": 1
            },
            {
              "id": 2829931,
              "postDate": "2024-05-22T22:34:36.333Z",
              "content": "<p>Besides, its unclear that host &amp; Kaggle team admit changing competition metrics by its design in the middle of the competition (because it could cause biased merit for some teams and not for others etc.).<br>\nSo, my current idea is <strong>asking host that providing us some measure of safeness without changing metrics</strong>.</p>\n<blockquote>\n  <p>My other thought was to ask host to publish their unet baseline build for this comp and ask clarification for it surely have decent score on private LB. This would be a benchmark of safeness.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490#2828609\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490#2828609</a></p>",
              "rawMarkdown": "Besides, its unclear that host & Kaggle team admit changing competition metrics by its design in the middle of the competition (because it could cause biased merit for some teams and not for others etc.).\nSo, my current idea is **asking host that providing us some measure of safeness without changing metrics**.\n\n> My other thought was to ask host to publish their unet baseline build for this comp and ask clarification for it surely have decent score on private LB. This would be a benchmark of safeness.\n\nhttps://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490#2828609",
              "votes": 4
            }
          ]
        }
      ]
    },
    {
      "id": 2829050,
      "postDate": "2024-05-22T11:38:15.190Z",
      "content": "<p>Let's calculate <strong>how extreme the resolution this competition requires to our model</strong>.</p>\n<p>This can be done with maximum value without normalizing with std deviation.<br>\nThe below table is statistics of target features <strong>without normalizing with std deviation</strong>.<br>\nSo, 2523σ in <code>ptend_q0002_27</code> corresponds with 6.7934e-9 [kg/kg] in actual humidity unit.<br>\nHow this amount is difficult to predict?</p>\n<p>Suppose density of dry air is 1 [kg/m^3], this amount corresponds with 6.7934 μg of vapor within 1 m^3.<br>\nI don't know how difficult this is (I appreciate kind person with sufficient domain knowledge suggest us).</p>\n<h2>Reference</h2>\n<ul>\n<li>[1] <a href=\"https://en.wikipedia.org/wiki/Density_of_air\" target=\"_blank\">https://en.wikipedia.org/wiki/Density_of_air</a></li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Ff2f917e5d3f55127ced497aec5c6d844%2FScreenshot%202024-05-22%20at%2020.30.37.png?generation=1716377449594314&amp;alt=media\"></p>",
      "rawMarkdown": "Let's calculate **how extreme the resolution this competition requires to our model**.\n\nThis can be done with maximum value without normalizing with std deviation.\nThe below table is statistics of target features **without normalizing with std deviation**.\nSo, 2523σ in `ptend_q0002_27` corresponds with 6.7934e-9 [kg/kg] in actual humidity unit.\nHow this amount is difficult to predict?\n\nSuppose density of dry air is 1 [kg/m^3], this amount corresponds with 6.7934 μg of vapor within 1 m^3.\nI don't know how difficult this is (I appreciate kind person with sufficient domain knowledge suggest us).\n\n## Reference\n\n- [1] https://en.wikipedia.org/wiki/Density_of_air\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Ff2f917e5d3f55127ced497aec5c6d844%2FScreenshot%202024-05-22%20at%2020.30.37.png?generation=1716377449594314&alt=media)",
      "votes": 1,
      "replies": [
        {
          "id": 2829109,
          "postDate": "2024-05-22T12:24:56.910Z",
          "content": "<p><strong>Difficulty vs Risk the targets</strong></p>\n<p>The below picture is plots of R2 score of one of my best models, and the risk.<br>\nIt depicts most of peaks in high risk actually shows difficulty though some are not (e.g. <code>ptend_q0002_27</code>).<br>\nThe latter case might be the model predict according to input distribution rather than fair function approximation.<br>\nThis also shows <strong>the risky columns are actually difficult to predict</strong>.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F8297a00fe5449f6df99578eca27472d8%2Fr2_vs_risk.jpeg?generation=1716380351091679&amp;alt=media\"></p>",
          "rawMarkdown": "**Difficulty vs Risk the targets**\n\nThe below picture is plots of R2 score of one of my best models, and the risk.\nIt depicts most of peaks in high risk actually shows difficulty though some are not (e.g. `ptend_q0002_27`).\nThe latter case might be the model predict according to input distribution rather than fair function approximation.\nThis also shows **the risky columns are actually difficult to predict**.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F8297a00fe5449f6df99578eca27472d8%2Fr2_vs_risk.jpeg?generation=1716380351091679&alt=media)",
          "votes": 8,
          "replies": [
            {
              "id": 2829111,
              "postDate": "2024-05-22T12:29:35.953Z",
              "content": "<p><a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> , this is an AMAZING plot. Thank you very much for sharing this with us! I am glad to see that your R2 curves look exactly like mine, although your ptend_u and ptend_v are much higher than mines. There is probably some trick to predict that ones, that only the people in the first places noticed, right? 😉</p>",
              "rawMarkdown": "@tatamikenn , this is an AMAZING plot. Thank you very much for sharing this with us! I am glad to see that your R2 curves look exactly like mine, although your ptend_u and ptend_v are much higher than mines. There is probably some trick to predict that ones, that only the people in the first places noticed, right? 😉"
            },
            {
              "id": 2829121,
              "postDate": "2024-05-22T12:39:11.047Z",
              "content": "<p>Actually, I couldn't have fond any kind of <strong>tricks</strong> in my models.<br>\nThough digging this deep could find some important insights about this competition.</p>",
              "rawMarkdown": "Actually, I couldn't have fond any kind of **tricks** in my models.\nThough digging this deep could find some important insights about this competition.",
              "votes": 2
            },
            {
              "id": 2830455,
              "postDate": "2024-05-23T07:47:20.420Z",
              "content": "<p>So, you don't have any columns with negative (and 0 after replacement) R2? Impressive!</p>",
              "rawMarkdown": "So, you don't have any columns with negative (and 0 after replacement) R2? Impressive!"
            },
            {
              "id": 2830463,
              "postDate": "2024-05-23T07:52:26.330Z",
              "content": "<p>Yep, but I have no idea on the test set.</p>",
              "rawMarkdown": "Yep, but I have no idea on the test set.",
              "votes": 2
            },
            {
              "id": 2840664,
              "postDate": "2024-05-28T07:35:49.957Z",
              "content": "<p><br>\nEdit: I had bugs</p>\n<p>Also thanks for the great visualizations, I've integrated them into my training loop!</p>",
              "rawMarkdown": "~~Just to clarify, the beginnings (after submission mask) of q1-q3 are predictable (positive r2) by model alone? (discounting manual post-processing like the q2 state replacement trick)~~\nEdit: I had bugs\n\nAlso thanks for the great visualizations, I've integrated them into my training loop!"
            }
          ]
        }
      ]
    },
    {
      "id": 2828467,
      "postDate": "2024-05-22T06:04:49.037Z",
      "content": "<p>I created my validation scheme based on this fact. Some of the targets are not stable between different folds. I do a similar simulation on 625k sized splits and found that some of the folds can get negative r2 score on q0001 12 13 14, q0002 27 28, q0003 12 13 14 targets. I have no idea what will happen if test set has those outliers.</p>",
      "rawMarkdown": "I created my validation scheme based on this fact. Some of the targets are not stable between different folds. I do a similar simulation on 625k sized splits and found that some of the folds can get negative r2 score on q0001 12 13 14, q0002 27 28, q0003 12 13 14 targets. I have no idea what will happen if test set has those outliers.",
      "votes": 2,
      "replies": [
        {
          "id": 2828477,
          "postDate": "2024-05-22T06:19:39.457Z",
          "content": "<p>Before posting this, I debated myself whether it was better to keep this fact to myself and find a clever way to dodge the issue, or to point it out to the organizers and hosts and ask for fair treatment. In the end, I realized that the easiest path was to confront everyone with this fact and share the suffering.</p>",
          "rawMarkdown": "Before posting this, I debated myself whether it was better to keep this fact to myself and find a clever way to dodge the issue, or to point it out to the organizers and hosts and ask for fair treatment. In the end, I realized that the easiest path was to confront everyone with this fact and share the suffering.",
          "votes": 10,
          "replies": [
            {
              "id": 2828535,
              "postDate": "2024-05-22T06:47:56.717Z",
              "content": "<p>So basically… the person/team that builds the model that is more robust to outliers will win?</p>",
              "rawMarkdown": "So basically... the person/team that builds the model that is more robust to outliers will win?"
            },
            {
              "id": 2828557,
              "postDate": "2024-05-22T07:00:22.040Z",
              "content": "<p>Thanks for sharing. I think keeping information like this to yourself doesn't worth it. It may still backfire at the end since the metric is not stable. Changing metric to weighted macro average r2 score would be fine for this competition imo.</p>",
              "rawMarkdown": "Thanks for sharing. I think keeping information like this to yourself doesn't worth it. It may still backfire at the end since the metric is not stable. Changing metric to weighted macro average r2 score would be fine for this competition imo."
            },
            {
              "id": 2828584,
              "postDate": "2024-05-22T07:10:11.527Z",
              "content": "<blockquote>\n  <p>Changing metric to weighted macro average r2 score would be fine for this competition imo.</p>\n</blockquote>\n<p>Yeah. I think R2 (normalizing with unbounded sigma) as it is too extreme for this competition. At least custom metrics with normalizing clipped weight (sigma is clipped with some tiny value) should have been better?</p>",
              "rawMarkdown": "> Changing metric to weighted macro average r2 score would be fine for this competition imo.\n\nYeah. I think R2 (normalizing with unbounded sigma) as it is too extreme for this competition. At least custom metrics with normalizing clipped weight (sigma is clipped with some tiny value) should have been better?"
            },
            {
              "id": 2828590,
              "postDate": "2024-05-22T07:12:04.380Z",
              "content": "<blockquote>\n  <p>So basically… the person/team that builds the model that is more robust to outliers will win?</p>\n</blockquote>\n<p>Possibly. But it depends on the test set. If host already removed these outliers, winners in public LB highly possibly wins.</p>",
              "rawMarkdown": "> So basically… the person/team that builds the model that is more robust to outliers will win?\n\nPossibly. But it depends on the test set. If host already removed these outliers, winners in public LB highly possibly wins.",
              "votes": 2
            },
            {
              "id": 2828609,
              "postDate": "2024-05-22T07:19:45.567Z",
              "content": "<p>My other thought was to ask host to publish their unet baseline build for this comp and ask clarification for it surely have decent score on private LB. This would be a benchmark of safeness.</p>",
              "rawMarkdown": "My other thought was to ask host to publish their unet baseline build for this comp and ask clarification for it surely have decent score on private LB. This would be a benchmark of safeness.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2881583,
      "postDate": "2024-06-20T19:43:53.260Z",
      "content": "<p>Cool plot, what tool did you use to generate it? Meaning, did you just use Pandas, or something more sophisticated like Apache Beam? Also, what kind of hardware did you use? Some times I run out of memory trying to analyze the whole dataset.</p>",
      "rawMarkdown": "Cool plot, what tool did you use to generate it? Meaning, did you just use Pandas, or something more sophisticated like Apache Beam? Also, what kind of hardware did you use? Some times I run out of memory trying to analyze the whole dataset.",
      "replies": [
        {
          "id": 2882300,
          "postDate": "2024-06-21T09:27:22.100Z",
          "content": "<p>I used polars on 128 GB RAM PC.</p>",
          "rawMarkdown": "I used polars on 128 GB RAM PC."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2834423,
      "author_name": "Jerry Lin",
      "author_url": "",
      "post_date": "2024-05-24T18:09:36.983000",
      "content": "<p>Hello, thank you for providing this detailed analysis of the dataset and sharing your concerns. If it helps, the private leaderboard score is only half a percent different from the public leaderboard score for the u-net benchmark submission.</p>",
      "votes": 14,
      "replies": [
        {
          "id": 2834437,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2024-05-24T18:19:40.280000",
          "content": "<p>That's good to know but I don't think it completely negates the risk of getting negative score.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2834672,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2024-05-24T22:41:18.327000",
          "content": "<p><a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> Thanks. It somehow helped.</p>\n<p>However, as Gunes pointed out, there left some risk of our (i.e. Kaggle community's) best performant model would be significantly affected by tiny amount of extreme outliers.</p>\n<p>Is it possible to consider to take further actions like listed in [1]?<br>\nFor example, if we have the baseline model's prediction for all test set (i.e. <code>submission.csv</code>), we can test our model's safeness by checking if our model's predictions are not significantly apart from this prediction.</p>\n<h2>Reference</h2>\n<p>[1] <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490#2832998\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490#2832998</a></p>",
          "votes": 2,
          "replies": [
            {
              "id": 2835414,
              "author_name": "Gunes Evitan",
              "author_url": "",
              "post_date": "2024-05-25T10:03:05.023000",
              "content": "<p>Also can we get details about the experimental setup (train/test size and etc.) of the u-net model?</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2832998,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2024-05-24T02:58:01.050000",
      "content": "<h2>Request for Some Action</h2>\n<p>I think generally machine learning models under current competition metrics (R2) is prone to some kind of extreme outliers (i.e. 2523σ).<br>\nIf this kind of extreme outliers are in private test set, it possibly causes catastrophic shakeup in the end of the competition.</p>\n<p>Hopefully, we believe some safeguard may be helpful for all of us. like, </p>\n<ol>\n<li>clipping extreme error in prediction error like <a href=\"https://www.kaggle.com/martynoveduard\" target=\"_blank\">@martynoveduard</a> proposed[1]</li>\n<li>Fixing to use customized R2 score with softer normalization policy like <a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a>  proposed[2]</li>\n<li>Providing baseline model's prediction (i.e. submission.csv) on test set and claim that it has decent score on private test set like I proposed[3]</li>\n</ol>\n<p>c.f.) I also posted mathematical explanation for instability in R2 metrics[4].</p>\n<p><a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> <a href=\"https://www.kaggle.com/ashleychow\" target=\"_blank\">@ashleychow</a> </p>\n<p>How hosts and Kaggle team think of this risk?</p>\n<p>Thanks in advance.</p>\n<h2>Reference</h2>\n<ul>\n<li>[1] <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490#2829249\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490#2829249</a></li>\n<li>[2] <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490#2829800\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490#2829800</a></li>\n<li>[3] <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490#2829931\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490#2829931</a></li>\n<li>[4] <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490#2829926\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490#2829926</a></li>\n</ul>",
      "votes": 9,
      "replies": []
    },
    {
      "id": 2829249,
      "author_name": "slime",
      "author_url": "",
      "post_date": "2024-05-22T13:57:24.570000",
      "content": "<p>amazing study</p>\n<p>suggestion:<br>\nmacro-avg(r2-score) -&gt; macro-avg(clip(r2-score,0,1))</p>\n<p>this way even if the participant doesn't handle the outlier, he is losing ~0.003 per-column compared to the perfectly predicted column (it hurts, but better than to have -10 score in private)</p>",
      "votes": 6,
      "replies": [
        {
          "id": 2829800,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2024-05-22T20:05:26.670000",
          "content": "<p>I will suggest a fix of my own then.<br>\nDefinition of r2 (I will use Wikipedia notation):<br>\nr2 = 1-SS(res)/SS(tot) </p>\n<p>The problem: since SS(tot) is calculated only for the test set, there can be a column with very small values for the private test set, leading to SS(tot) very small, leading to a negative explosion of r2 for this column.<br>\nFix: Use SS(tot) that is calculated for all the data. So instead of SS(tot)(test): <br>\n(n(#samples in test)/N(#all samples))*SS(tot)(all samples)<br>\nBasically, ensure that SS(tot) has reasonable values. Actually, the values should just by n since the values are normalized by their SD. So it can be reduced to:<br>\nr2 = 1-SS(res)/n(#samples in test)<br>\nI think? Well, something along those lines. The general idea should be correct.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2829926,
              "author_name": "Bilzard",
              "author_url": "",
              "post_date": "2024-05-22T22:30:34.420000",
              "content": "<p>First of all, we should claim the relationship between $R^2$ metrics and standard deviation of given dataset. By definition,</p>\n<p>$$<br>\n\\begin{align*}<br>\nR^2 &amp;= 1 - \\frac{\\frac{1}{N} \\sum_i (y_i - \\hat{y}_i)^2}{\\frac{1}{N} \\sum_i (y_i - \\bar{y})^2} \\\\<br>\n&amp;=1 - \\frac{\\frac{1}{N} \\sum_i (y_i - \\hat{y}_i)^2}{\\sigma_t^2} \\\\<br>\n&amp;=1 - \\frac{1}{N} \\sum_i (y_i/\\sigma_t - \\hat{y}_i / \\sigma_t)^2 \\\\<br>\n\\end{align*}<br>\n$$</p>\n<p>That's why I normalized all data with standard deviation of all train data. And the above metric <strong>should be affected by outlier even if we use S(tot)_train instead of S(tot)_test</strong> i.e. <strong>sample with y_i=2523σ busts R2 score</strong>.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2829931,
              "author_name": "Bilzard",
              "author_url": "",
              "post_date": "2024-05-22T22:34:36.333000",
              "content": "<p>Besides, its unclear that host &amp; Kaggle team admit changing competition metrics by its design in the middle of the competition (because it could cause biased merit for some teams and not for others etc.).<br>\nSo, my current idea is <strong>asking host that providing us some measure of safeness without changing metrics</strong>.</p>\n<blockquote>\n  <p>My other thought was to ask host to publish their unet baseline build for this comp and ask clarification for it surely have decent score on private LB. This would be a benchmark of safeness.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490#2828609\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490#2828609</a></p>",
              "votes": 4,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2829050,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2024-05-22T11:38:15.190000",
      "content": "<p>Let's calculate <strong>how extreme the resolution this competition requires to our model</strong>.</p>\n<p>This can be done with maximum value without normalizing with std deviation.<br>\nThe below table is statistics of target features <strong>without normalizing with std deviation</strong>.<br>\nSo, 2523σ in <code>ptend_q0002_27</code> corresponds with 6.7934e-9 [kg/kg] in actual humidity unit.<br>\nHow this amount is difficult to predict?</p>\n<p>Suppose density of dry air is 1 [kg/m^3], this amount corresponds with 6.7934 μg of vapor within 1 m^3.<br>\nI don't know how difficult this is (I appreciate kind person with sufficient domain knowledge suggest us).</p>\n<h2>Reference</h2>\n<ul>\n<li>[1] <a href=\"https://en.wikipedia.org/wiki/Density_of_air\" target=\"_blank\">https://en.wikipedia.org/wiki/Density_of_air</a></li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Ff2f917e5d3f55127ced497aec5c6d844%2FScreenshot%202024-05-22%20at%2020.30.37.png?generation=1716377449594314&amp;alt=media\"></p>",
      "votes": 1,
      "replies": [
        {
          "id": 2829109,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2024-05-22T12:24:56.910000",
          "content": "<p><strong>Difficulty vs Risk the targets</strong></p>\n<p>The below picture is plots of R2 score of one of my best models, and the risk.<br>\nIt depicts most of peaks in high risk actually shows difficulty though some are not (e.g. <code>ptend_q0002_27</code>).<br>\nThe latter case might be the model predict according to input distribution rather than fair function approximation.<br>\nThis also shows <strong>the risky columns are actually difficult to predict</strong>.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F8297a00fe5449f6df99578eca27472d8%2Fr2_vs_risk.jpeg?generation=1716380351091679&amp;alt=media\"></p>",
          "votes": 8,
          "replies": [
            {
              "id": 2829111,
              "author_name": "Federico Peccia",
              "author_url": "",
              "post_date": "2024-05-22T12:29:35.953000",
              "content": "<p><a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> , this is an AMAZING plot. Thank you very much for sharing this with us! I am glad to see that your R2 curves look exactly like mine, although your ptend_u and ptend_v are much higher than mines. There is probably some trick to predict that ones, that only the people in the first places noticed, right? 😉</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2829121,
              "author_name": "Bilzard",
              "author_url": "",
              "post_date": "2024-05-22T12:39:11.047000",
              "content": "<p>Actually, I couldn't have fond any kind of <strong>tricks</strong> in my models.<br>\nThough digging this deep could find some important insights about this competition.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2830455,
              "author_name": "DennisSakva",
              "author_url": "",
              "post_date": "2024-05-23T07:47:20.420000",
              "content": "<p>So, you don't have any columns with negative (and 0 after replacement) R2? Impressive!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2830463,
              "author_name": "Bilzard",
              "author_url": "",
              "post_date": "2024-05-23T07:52:26.330000",
              "content": "<p>Yep, but I have no idea on the test set.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2840664,
              "author_name": "sroger",
              "author_url": "",
              "post_date": "2024-05-28T07:35:49.957000",
              "content": "<p><br>\nEdit: I had bugs</p>\n<p>Also thanks for the great visualizations, I've integrated them into my training loop!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2828467,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2024-05-22T06:04:49.037000",
      "content": "<p>I created my validation scheme based on this fact. Some of the targets are not stable between different folds. I do a similar simulation on 625k sized splits and found that some of the folds can get negative r2 score on q0001 12 13 14, q0002 27 28, q0003 12 13 14 targets. I have no idea what will happen if test set has those outliers.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2828477,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2024-05-22T06:19:39.457000",
          "content": "<p>Before posting this, I debated myself whether it was better to keep this fact to myself and find a clever way to dodge the issue, or to point it out to the organizers and hosts and ask for fair treatment. In the end, I realized that the easiest path was to confront everyone with this fact and share the suffering.</p>",
          "votes": 10,
          "replies": [
            {
              "id": 2828535,
              "author_name": "Federico Peccia",
              "author_url": "",
              "post_date": "2024-05-22T06:47:56.717000",
              "content": "<p>So basically… the person/team that builds the model that is more robust to outliers will win?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2828557,
              "author_name": "Gunes Evitan",
              "author_url": "",
              "post_date": "2024-05-22T07:00:22.040000",
              "content": "<p>Thanks for sharing. I think keeping information like this to yourself doesn't worth it. It may still backfire at the end since the metric is not stable. Changing metric to weighted macro average r2 score would be fine for this competition imo.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2828584,
              "author_name": "Bilzard",
              "author_url": "",
              "post_date": "2024-05-22T07:10:11.527000",
              "content": "<blockquote>\n  <p>Changing metric to weighted macro average r2 score would be fine for this competition imo.</p>\n</blockquote>\n<p>Yeah. I think R2 (normalizing with unbounded sigma) as it is too extreme for this competition. At least custom metrics with normalizing clipped weight (sigma is clipped with some tiny value) should have been better?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2828590,
              "author_name": "Bilzard",
              "author_url": "",
              "post_date": "2024-05-22T07:12:04.380000",
              "content": "<blockquote>\n  <p>So basically… the person/team that builds the model that is more robust to outliers will win?</p>\n</blockquote>\n<p>Possibly. But it depends on the test set. If host already removed these outliers, winners in public LB highly possibly wins.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2828609,
              "author_name": "Bilzard",
              "author_url": "",
              "post_date": "2024-05-22T07:19:45.567000",
              "content": "<p>My other thought was to ask host to publish their unet baseline build for this comp and ask clarification for it surely have decent score on private LB. This would be a benchmark of safeness.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2881583,
      "author_name": "Juan D C F",
      "author_url": "",
      "post_date": "2024-06-20T19:43:53.260000",
      "content": "<p>Cool plot, what tool did you use to generate it? Meaning, did you just use Pandas, or something more sophisticated like Apache Beam? Also, what kind of hardware did you use? Some times I run out of memory trying to analyze the whole dataset.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2882300,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2024-06-21T09:27:22.100000",
          "content": "<p>I used polars on 128 GB RAM PC.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2828449": "Hi. Hope you and your fellow Sisyphus are doing well with carrying the rock.\n\nI'll show some evidence of this competition's extreme setup.\n\n## TL; DR\n\nsome target columns are significantly risky.\nIf one single outlier sample could cause significant drop on private LB (~-12 with single sample).\n\n## How these graphs are created\n\nI plotted the risk of all 368 target columns except trivial columns with σ=0.\nThe risk is calculated by\n\n$$\n\\text{risk} := \\max(y) / \\sigma(y)\n$$\n\nThe first and 2nd picture shows some columns has significantly distant outliers (up to ~2500 σ).\n\nThis could cause your model significantly drops in private LB.\nThe 3rd picture depicts this outlier case.\nSuppose if your model predict single outlier in private test set with ptend_q0002_27 = 0 (because 99.9999% has <= 0 values), whose actual value is 2523σ, you will get `current score - 12.24` in private LB.\n\nSo hope your model can handle these outliers well.\n(I also hope host's already removed these miserable outliers from test set).\n\nGood luck!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fd7001b68bae626d71e9e16da10f86f74%2Frisk_graph.jpeg?generation=1716355019829983&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fe359364d9949eb7c67a8c1e42496dd6b%2FScreenshot%202024-05-22%20at%2014.16.39.png?generation=1716355036821574&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F30d088e4934ee91f6b47c4862217b3de%2FScreenshot%202024-05-22%20at%2014.16.47.png?generation=1716355049231387&alt=media)",
    "2834423": "Hello, thank you for providing this detailed analysis of the dataset and sharing your concerns. If it helps, the private leaderboard score is only half a percent different from the public leaderboard score for the u-net benchmark submission.",
    "2832998": "## Request for Some Action\n\nI think generally machine learning models under current competition metrics (R2) is prone to some kind of extreme outliers (i.e. 2523σ).\nIf this kind of extreme outliers are in private test set, it possibly causes catastrophic shakeup in the end of the competition.\n\nHopefully, we believe some safeguard may be helpful for all of us. like, \n\n1. clipping extreme error in prediction error like @martynoveduard proposed[1]\n2. Fixing to use customized R2 score with softer normalization policy like @shlomoron  proposed[2]\n3. Providing baseline model's prediction (i.e. submission.csv) on test set and claim that it has decent score on private test set like I proposed[3]\n\nc.f.) I also posted mathematical explanation for instability in R2 metrics[4].\n\n@jerrylin96 @ashleychow \n\nHow hosts and Kaggle team think of this risk?\n\nThanks in advance.\n\n## Reference\n\n- [1] https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490#2829249\n- [2] https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490#2829800\n- [3] https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490#2829931\n- [4] https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490#2829926",
    "2829249": "amazing study\n\nsuggestion:\nmacro-avg(r2-score) -> macro-avg(clip(r2-score,0,1))\n\nthis way even if the participant doesn't handle the outlier, he is losing ~0.003 per-column compared to the perfectly predicted column (it hurts, but better than to have -10 score in private)",
    "2829050": "Let's calculate **how extreme the resolution this competition requires to our model**.\n\nThis can be done with maximum value without normalizing with std deviation.\nThe below table is statistics of target features **without normalizing with std deviation**.\nSo, 2523σ in `ptend_q0002_27` corresponds with 6.7934e-9 [kg/kg] in actual humidity unit.\nHow this amount is difficult to predict?\n\nSuppose density of dry air is 1 [kg/m^3], this amount corresponds with 6.7934 μg of vapor within 1 m^3.\nI don't know how difficult this is (I appreciate kind person with sufficient domain knowledge suggest us).\n\n## Reference\n\n- [1] https://en.wikipedia.org/wiki/Density_of_air\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Ff2f917e5d3f55127ced497aec5c6d844%2FScreenshot%202024-05-22%20at%2020.30.37.png?generation=1716377449594314&alt=media)",
    "2828467": "I created my validation scheme based on this fact. Some of the targets are not stable between different folds. I do a similar simulation on 625k sized splits and found that some of the folds can get negative r2 score on q0001 12 13 14, q0002 27 28, q0003 12 13 14 targets. I have no idea what will happen if test set has those outliers.",
    "2881583": "Cool plot, what tool did you use to generate it? Meaning, did you just use Pandas, or something more sophisticated like Apache Beam? Also, what kind of hardware did you use? Some times I run out of memory trying to analyze the whole dataset."
  }
}