{
  "id": 502658,
  "title": "Be careful about the order of y_true and y_pred when calculating R2 score",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/502658",
  "author_name": "kuto",
  "post_date": "2024-05-14T09:40:19.721000",
  "votes": 3,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I know not many people make this mistake, but I'm sharing it here because I was mistake.</p>\n<p>When calculating R2 score, be careful about the order of y_true and y_pred. This metric is asymmetric, unlike RMSE.<br>\nFor example, <br>\nin <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.metrics.r2_score.html#sklearn.metrics.r2_score\" target=\"_blank\">sklearn</a>: <code>r2_score(y_true, y_pred)</code><br>\nin <a href=\"https://lightning.ai/docs/torchmetrics/stable/regression/r2_score.html#functional-interface\" target=\"_blank\">torch metrics</a>:  <code>r2_score(y_pred, y_true)</code></p>\n<p>I had repeated the experiment with the wrong order and suffered from the score not increasing as expected. However, after correcting this, CV and LB scores improved from 0.43 to 0.58</p>\n<p>The difference in scores is due to post-processing that fills in the predictions to 0 when R2 score is less than 0, as shared in some public notebooks. (e.g., <a href=\"https://www.kaggle.com/code/titericz/giba-baseline-xgboost\" target=\"_blank\">this notebook</a>)<br>\nThe calculated my R2 score was incorrect, leading to scores unintentionally falling below 0 in several columns. This error caused to produce lower scores.</p>",
  "messages": [
    {
      "id": 2812565,
      "postDate": "2024-05-14T09:40:19.720Z",
      "content": "<p>I know not many people make this mistake, but I'm sharing it here because I was mistake.</p>\n<p>When calculating R2 score, be careful about the order of y_true and y_pred. This metric is asymmetric, unlike RMSE.<br>\nFor example, <br>\nin <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.metrics.r2_score.html#sklearn.metrics.r2_score\" target=\"_blank\">sklearn</a>: <code>r2_score(y_true, y_pred)</code><br>\nin <a href=\"https://lightning.ai/docs/torchmetrics/stable/regression/r2_score.html#functional-interface\" target=\"_blank\">torch metrics</a>:  <code>r2_score(y_pred, y_true)</code></p>\n<p>I had repeated the experiment with the wrong order and suffered from the score not increasing as expected. However, after correcting this, CV and LB scores improved from 0.43 to 0.58</p>\n<p>The difference in scores is due to post-processing that fills in the predictions to 0 when R2 score is less than 0, as shared in some public notebooks. (e.g., <a href=\"https://www.kaggle.com/code/titericz/giba-baseline-xgboost\" target=\"_blank\">this notebook</a>)<br>\nThe calculated my R2 score was incorrect, leading to scores unintentionally falling below 0 in several columns. This error caused to produce lower scores.</p>",
      "rawMarkdown": "I know not many people make this mistake, but I'm sharing it here because I was mistake.\n\nWhen calculating R2 score, be careful about the order of y_true and y_pred. This metric is asymmetric, unlike RMSE.\nFor example, \nin [sklearn](https://scikit-learn.org/stable/modules/generated/sklearn.metrics.r2_score.html#sklearn.metrics.r2_score): `r2_score(y_true, y_pred)`\nin [torch metrics](https://lightning.ai/docs/torchmetrics/stable/regression/r2_score.html#functional-interface):  `r2_score(y_pred, y_true)`\n\nI had repeated the experiment with the wrong order and suffered from the score not increasing as expected. However, after correcting this, CV and LB scores improved from 0.43 to 0.58\n\nThe difference in scores is due to post-processing that fills in the predictions to 0 when R2 score is less than 0, as shared in some public notebooks. (e.g., [this notebook](https://www.kaggle.com/code/titericz/giba-baseline-xgboost))\nThe calculated my R2 score was incorrect, leading to scores unintentionally falling below 0 in several columns. This error caused to produce lower scores.",
      "votes": 3
    },
    {
      "id": 2812955,
      "postDate": "2024-05-14T13:47:56.500Z",
      "content": "<p>Hahaha I also made this mistake several weeks ago and there was a lot of frustration 🤣</p>",
      "rawMarkdown": "Hahaha I also made this mistake several weeks ago and there was a lot of frustration 🤣"
    },
    {
      "id": 2812953,
      "postDate": "2024-05-14T13:45:48.737Z",
      "content": "<p>Hello, I checked for sklearn-based notebook (Giba's) and it seems that the R² score is calculated properly as (true, pred). Where exactly do yo think there's something wrong ?</p>",
      "rawMarkdown": "Hello, I checked for sklearn-based notebook (Giba's) and it seems that the R² score is calculated properly as (true, pred). Where exactly do yo think there's something wrong ?",
      "replies": [
        {
          "id": 2812987,
          "postDate": "2024-05-14T14:09:22.960Z",
          "content": "<p>Yes, Giba's public notebook is correct.</p>\n<p>I mean that post-processing based on r2 score <strong>with</strong> a miscalculation of r2_score is not good.</p>",
          "rawMarkdown": "Yes, Giba's public notebook is correct.\n\n I mean that post-processing based on r2 score **with** a miscalculation of r2_score is not good.",
          "votes": 1,
          "replies": [
            {
              "id": 2814369,
              "postDate": "2024-05-15T09:08:59.800Z",
              "content": "<p>Okay thank you ! Will definitely be careful about this</p>",
              "rawMarkdown": "Okay thank you ! Will definitely be careful about this"
            }
          ]
        }
      ]
    },
    {
      "id": 2822367,
      "postDate": "2024-05-18T14:53:06.477Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2812955,
      "author_name": "greySnow",
      "author_url": "",
      "post_date": "2024-05-14T13:47:56.500000",
      "content": "<p>Hahaha I also made this mistake several weeks ago and there was a lot of frustration 🤣</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2812953,
      "author_name": "Albanito",
      "author_url": "",
      "post_date": "2024-05-14T13:45:48.737000",
      "content": "<p>Hello, I checked for sklearn-based notebook (Giba's) and it seems that the R² score is calculated properly as (true, pred). Where exactly do yo think there's something wrong ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2812987,
          "author_name": "kuto",
          "author_url": "",
          "post_date": "2024-05-14T14:09:22.960000",
          "content": "<p>Yes, Giba's public notebook is correct.</p>\n<p>I mean that post-processing based on r2 score <strong>with</strong> a miscalculation of r2_score is not good.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2814369,
              "author_name": "Albanito",
              "author_url": "",
              "post_date": "2024-05-15T09:08:59.800000",
              "content": "<p>Okay thank you ! Will definitely be careful about this</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2822367,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-05-18T14:53:06.477000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2812565": "I know not many people make this mistake, but I'm sharing it here because I was mistake.\n\nWhen calculating R2 score, be careful about the order of y_true and y_pred. This metric is asymmetric, unlike RMSE.\nFor example, \nin [sklearn](https://scikit-learn.org/stable/modules/generated/sklearn.metrics.r2_score.html#sklearn.metrics.r2_score): `r2_score(y_true, y_pred)`\nin [torch metrics](https://lightning.ai/docs/torchmetrics/stable/regression/r2_score.html#functional-interface):  `r2_score(y_pred, y_true)`\n\nI had repeated the experiment with the wrong order and suffered from the score not increasing as expected. However, after correcting this, CV and LB scores improved from 0.43 to 0.58\n\nThe difference in scores is due to post-processing that fills in the predictions to 0 when R2 score is less than 0, as shared in some public notebooks. (e.g., [this notebook](https://www.kaggle.com/code/titericz/giba-baseline-xgboost))\nThe calculated my R2 score was incorrect, leading to scores unintentionally falling below 0 in several columns. This error caused to produce lower scores.",
    "2812955": "Hahaha I also made this mistake several weeks ago and there was a lot of frustration 🤣",
    "2812953": "Hello, I checked for sklearn-based notebook (Giba's) and it seems that the R² score is calculated properly as (true, pred). Where exactly do yo think there's something wrong ?",
    "2822367": ""
  }
}