{
  "id": 498806,
  "title": "How to scale input features and targets? ",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/498806",
  "author_name": "milann",
  "post_date": "2024-04-29T16:31:12.629000",
  "votes": 6,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hello everyone,</p>\n<p>I noticed that some examples in notebooks (mostly for NN models) use the following scheme for scaling input features (<a href=\"https://www.kaggle.com/code/ymatioun/leap-simple-nn/notebook\" target=\"_blank\">reference</a>):</p>\n<pre><code>\nmx = x.mean(axis=)\nsx = np.maximum(x.std(axis=), min_std)\nx = (x - mx.reshape(,-)) / sx.reshape(,-)\nxt = (xt - mx.reshape(,-)) / sx.reshape(,-)\n</code></pre>\n<p>And for targets:</p>\n<pre><code>\nmy = y.mean(axis=)\nsy = np.maximum(np.sqrt((y*y).mean(axis=)), min_std)\ny = (y - my.reshape(,-)) / sy.reshape(,-)\n</code></pre>\n<p>I understand the X part (for input features), which is equivalent to StandardScaler in scikit-learn. However, I'm not sure for the Y part (targets), since the \"standard deviation\" <code>sy</code> is calculated differently than the <code>sx</code>. </p>\n<p>I have two questions:</p>\n<ol>\n<li>Why is there a difference in calculating sy and sx?</li>\n<li>Would it be a mistake to scale targets in the same way as input features?</li>\n</ol>\n<p>Many thanks.</p>",
  "messages": [
    {
      "id": 2783245,
      "postDate": "2024-04-29T16:31:12.630Z",
      "content": "<p>Hello everyone,</p>\n<p>I noticed that some examples in notebooks (mostly for NN models) use the following scheme for scaling input features (<a href=\"https://www.kaggle.com/code/ymatioun/leap-simple-nn/notebook\" target=\"_blank\">reference</a>):</p>\n<pre><code>\nmx = x.mean(axis=)\nsx = np.maximum(x.std(axis=), min_std)\nx = (x - mx.reshape(,-)) / sx.reshape(,-)\nxt = (xt - mx.reshape(,-)) / sx.reshape(,-)\n</code></pre>\n<p>And for targets:</p>\n<pre><code>\nmy = y.mean(axis=)\nsy = np.maximum(np.sqrt((y*y).mean(axis=)), min_std)\ny = (y - my.reshape(,-)) / sy.reshape(,-)\n</code></pre>\n<p>I understand the X part (for input features), which is equivalent to StandardScaler in scikit-learn. However, I'm not sure for the Y part (targets), since the \"standard deviation\" <code>sy</code> is calculated differently than the <code>sx</code>. </p>\n<p>I have two questions:</p>\n<ol>\n<li>Why is there a difference in calculating sy and sx?</li>\n<li>Would it be a mistake to scale targets in the same way as input features?</li>\n</ol>\n<p>Many thanks.</p>",
      "rawMarkdown": "Hello everyone,\n\nI noticed that some examples in notebooks (mostly for NN models) use the following scheme for scaling input features ([reference](https://www.kaggle.com/code/ymatioun/leap-simple-nn/notebook)):\n\n```python\n# norm X\nmx = x.mean(axis=0)\nsx = np.maximum(x.std(axis=0), min_std)\nx = (x - mx.reshape(1,-1)) / sx.reshape(1,-1)\nxt = (xt - mx.reshape(1,-1)) / sx.reshape(1,-1)\n```\n\nAnd for targets:\n```python\n# norm Y\nmy = y.mean(axis=0)\nsy = np.maximum(np.sqrt((y*y).mean(axis=0)), min_std)\ny = (y - my.reshape(1,-1)) / sy.reshape(1,-1)\n```\n\nI understand the X part (for input features), which is equivalent to StandardScaler in scikit-learn. However, I'm not sure for the Y part (targets), since the \"standard deviation\" `sy ` is calculated differently than the `sx`. \n\nI have two questions:\n1. Why is there a difference in calculating sy and sx?\n2. Would it be a mistake to scale targets in the same way as input features?\n\nMany thanks.",
      "votes": 6
    },
    {
      "id": 2783490,
      "postDate": "2024-04-29T19:07:21.720Z",
      "content": "<p>I'm not sure what the intuition is behind that scaling for y, since <code>sy</code> here is not the standard deviation, but it's being used as though it were. I used the same transformation for both, and it depends on what transformation you use whether it would be appropriate for all of the feature variables and all of the target variables. It may be worthwhile using something different for features and targets, but definitely wouldn't be wrong not to. </p>",
      "rawMarkdown": "I'm not sure what the intuition is behind that scaling for y, since `sy` here is not the standard deviation, but it's being used as though it were. I used the same transformation for both, and it depends on what transformation you use whether it would be appropriate for all of the feature variables and all of the target variables. It may be worthwhile using something different for features and targets, but definitely wouldn't be wrong not to. ",
      "votes": 1,
      "replies": [
        {
          "id": 2783528,
          "postDate": "2024-04-29T19:30:10.417Z",
          "content": "<p>Thank you. For now, I will use the well-known standardization for targets as for inputs.</p>",
          "rawMarkdown": "Thank you. For now, I will use the well-known standardization for targets as for inputs.",
          "replies": [
            {
              "id": 2783555,
              "postDate": "2024-04-29T19:52:40.620Z",
              "content": "<p>objective is R2 = 1 - (y-p)^2/y^2. So each variable is scaled down by y^2 [and not by std(y)!]. So if we divide y by mean(y^2), then all the columns of y get the correct weight and can be optimized together.</p>",
              "rawMarkdown": "objective is R2 = 1 - (y-p)^2/y^2. So each variable is scaled down by y^2 [and not by std(y)!]. So if we divide y by mean(y^2), then all the columns of y get the correct weight and can be optimized together.",
              "votes": 8
            },
            {
              "id": 2783640,
              "postDate": "2024-04-29T21:02:58.410Z",
              "content": "<p>So if I understand correctly, y is scaled in that way because a model is trained to predict all targets jointly?</p>",
              "rawMarkdown": "So if I understand correctly, y is scaled in that way because a model is trained to predict all targets jointly?",
              "votes": 1
            },
            {
              "id": 2784639,
              "postDate": "2024-04-30T12:12:33.577Z",
              "content": "<p>yes. It is scaled so that model error (MSE) exactly corresponds to competition objective (R2).</p>",
              "rawMarkdown": "yes. It is scaled so that model error (MSE) exactly corresponds to competition objective (R2).",
              "votes": 2
            },
            {
              "id": 2869393,
              "postDate": "2024-06-13T04:37:48.293Z",
              "content": "<p>In Wiki I found that the definition of R2 is $$R^2=1-\\frac{SS_{res}}{S_{tot}}$$, where $$SS_{res}=\\sum_{i}(y_i-p_i)^2, SS_{tot}=\\sum_{i}(y_i-\\bar{y})^2 \\propto(std(y))^2$$, which means y should be rescaled by std(y). I don't understand why here scaling y with mean(y^2) is correct.</p>",
              "rawMarkdown": "In Wiki I found that the definition of R2 is $$R^2=1-\\frac{SS_{res}}{S_{tot}}$$, where $$SS_{res}=\\sum_{i}(y_i-p_i)^2, SS_{tot}=\\sum_{i}(y_i-\\bar{y})^2 \\propto(std(y))^2$$, which means y should be rescaled by std(y). I don't understand why here scaling y with mean(y^2) is correct."
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2783490,
      "author_name": "Jekasm19",
      "author_url": "",
      "post_date": "2024-04-29T19:07:21.720000",
      "content": "<p>I'm not sure what the intuition is behind that scaling for y, since <code>sy</code> here is not the standard deviation, but it's being used as though it were. I used the same transformation for both, and it depends on what transformation you use whether it would be appropriate for all of the feature variables and all of the target variables. It may be worthwhile using something different for features and targets, but definitely wouldn't be wrong not to. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2783528,
          "author_name": "milann",
          "author_url": "",
          "post_date": "2024-04-29T19:30:10.417000",
          "content": "<p>Thank you. For now, I will use the well-known standardization for targets as for inputs.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2783555,
              "author_name": "Youri Matiounine",
              "author_url": "",
              "post_date": "2024-04-29T19:52:40.620000",
              "content": "<p>objective is R2 = 1 - (y-p)^2/y^2. So each variable is scaled down by y^2 [and not by std(y)!]. So if we divide y by mean(y^2), then all the columns of y get the correct weight and can be optimized together.</p>",
              "votes": 8,
              "replies": []
            },
            {
              "id": 2783640,
              "author_name": "milann",
              "author_url": "",
              "post_date": "2024-04-29T21:02:58.410000",
              "content": "<p>So if I understand correctly, y is scaled in that way because a model is trained to predict all targets jointly?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2784639,
              "author_name": "Youri Matiounine",
              "author_url": "",
              "post_date": "2024-04-30T12:12:33.577000",
              "content": "<p>yes. It is scaled so that model error (MSE) exactly corresponds to competition objective (R2).</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2869393,
              "author_name": "Sijun Xu",
              "author_url": "",
              "post_date": "2024-06-13T04:37:48.293000",
              "content": "<p>In Wiki I found that the definition of R2 is $$R^2=1-\\frac{SS_{res}}{S_{tot}}$$, where $$SS_{res}=\\sum_{i}(y_i-p_i)^2, SS_{tot}=\\sum_{i}(y_i-\\bar{y})^2 \\propto(std(y))^2$$, which means y should be rescaled by std(y). I don't understand why here scaling y with mean(y^2) is correct.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2783245": "Hello everyone,\n\nI noticed that some examples in notebooks (mostly for NN models) use the following scheme for scaling input features ([reference](https://www.kaggle.com/code/ymatioun/leap-simple-nn/notebook)):\n\n```python\n# norm X\nmx = x.mean(axis=0)\nsx = np.maximum(x.std(axis=0), min_std)\nx = (x - mx.reshape(1,-1)) / sx.reshape(1,-1)\nxt = (xt - mx.reshape(1,-1)) / sx.reshape(1,-1)\n```\n\nAnd for targets:\n```python\n# norm Y\nmy = y.mean(axis=0)\nsy = np.maximum(np.sqrt((y*y).mean(axis=0)), min_std)\ny = (y - my.reshape(1,-1)) / sy.reshape(1,-1)\n```\n\nI understand the X part (for input features), which is equivalent to StandardScaler in scikit-learn. However, I'm not sure for the Y part (targets), since the \"standard deviation\" `sy ` is calculated differently than the `sx`. \n\nI have two questions:\n1. Why is there a difference in calculating sy and sx?\n2. Would it be a mistake to scale targets in the same way as input features?\n\nMany thanks.",
    "2783490": "I'm not sure what the intuition is behind that scaling for y, since `sy` here is not the standard deviation, but it's being used as though it were. I used the same transformation for both, and it depends on what transformation you use whether it would be appropriate for all of the feature variables and all of the target variables. It may be worthwhile using something different for features and targets, but definitely wouldn't be wrong not to. "
  }
}