{
  "id": 500078,
  "title": "BIG SCORE DIFFERENCE between my calculated R2 score and actual LB",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/500078",
  "author_name": "thoth000",
  "post_date": "2024-05-04T07:30:28.533000",
  "votes": 3,
  "comment_count": 3,
  "views": 0,
  "content": "<h1>Hi</h1>\n<p>I am currently having trouble with a large difference between the R2 I calculated myself for model evaluation and the actual LB.</p>\n<h2>Here is the actual environment</h2>\n<ul>\n<li>model : Simple model created with pytorch</li>\n<li>dataset : 50000 rows of data in train.csv<ul>\n<li>train : 40000  rows (80%)</li>\n<li>valid : 5000 rows (10%)</li>\n<li>test : 5000 rows (10%, This will be used for the final model evaluation <strong>R2</strong>)</li></ul></li>\n<li>loss function : nn.MSELoss()</li>\n</ul>\n<p>Let's look at the change in the TRAIN and VALID error functions during training. It is over-trained, but not that bad.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6855922%2F94252ca74109f887af437e00fbd1c874%2Fimage.png?generation=1714806746529396&amp;alt=media\"></p>\n<p>The values of the error function at the end of the study appear to be adequate.</p>\n<pre><code>  train loss: .\n  valid loss: .\n</code></pre>\n<p>I predicted <code>test_t</code> with this model and calculated the R2 score with the following function:</p>\n<pre><code> ():\n    \n    ss_res = torch.((y_true - y_pred) ** )\n\n    mean = torch.mean(y_true)\n\n    \n    ss_tot = torch.((y_true - mean) ** )\n\n    \n    r2 =  - (ss_res / ss_tot)\n\n     r2\n</code></pre>\n<p>And, I got a <strong>0.5036</strong> score on0 <code>r_square</code>.</p>\n<p>The process that follows is as follows.</p>\n<pre><code>x = pd.read_csv(test_path)\nt = model(x)\nsub = pd.read_csv(sub_path)\nsub[:, :] *= t\n\n</code></pre>\n<p>But, my submission score(LB) is <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6855922%2F737db36b1f0e274e2d2eb16619c665bf%2Fimage.png?generation=1714807819858109&amp;alt=media\"></p>\n<hr>\n<h1>What caused it?</h1>\n<p>I think</p>\n<ol>\n<li><p>Possibility that the distribution of data <code>train.csv</code> depends on the order of the data.<br>\n  For example, the possibility that sample_id is train_0 and train_1 data are similar.</p></li>\n<li><p>Data for training is small.</p></li>\n<li><p>Is this process necessary before training model? (I have not done this process)</p></li>\n</ol>\n<pre><code>\nmx = x.mean(axis=)\nsx = np.maximum(x.std(axis=), min_std)\nx = (x - mx.reshape(,-)) / sx.reshape(,-)\n  DEBUGGING:\n    xt = (xt - mx.reshape(,-)) / sx.reshape(,-)\n\n\nmy = y.mean(axis=)\nsy = np.maximum(np.sqrt((y*y).mean(axis=)), min_std)\ny = (y - my.reshape(,-)) / sy.reshape(,-)\n</code></pre>\n<hr>\n<p>I can't think of anything else.<br>\nPlease someone help me😭</p>",
  "messages": [
    {
      "id": 2792377,
      "postDate": "2024-05-04T07:30:28.533Z",
      "content": "<h1>Hi</h1>\n<p>I am currently having trouble with a large difference between the R2 I calculated myself for model evaluation and the actual LB.</p>\n<h2>Here is the actual environment</h2>\n<ul>\n<li>model : Simple model created with pytorch</li>\n<li>dataset : 50000 rows of data in train.csv<ul>\n<li>train : 40000  rows (80%)</li>\n<li>valid : 5000 rows (10%)</li>\n<li>test : 5000 rows (10%, This will be used for the final model evaluation <strong>R2</strong>)</li></ul></li>\n<li>loss function : nn.MSELoss()</li>\n</ul>\n<p>Let's look at the change in the TRAIN and VALID error functions during training. It is over-trained, but not that bad.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6855922%2F94252ca74109f887af437e00fbd1c874%2Fimage.png?generation=1714806746529396&amp;alt=media\"></p>\n<p>The values of the error function at the end of the study appear to be adequate.</p>\n<pre><code>  train loss: .\n  valid loss: .\n</code></pre>\n<p>I predicted <code>test_t</code> with this model and calculated the R2 score with the following function:</p>\n<pre><code> ():\n    \n    ss_res = torch.((y_true - y_pred) ** )\n\n    mean = torch.mean(y_true)\n\n    \n    ss_tot = torch.((y_true - mean) ** )\n\n    \n    r2 =  - (ss_res / ss_tot)\n\n     r2\n</code></pre>\n<p>And, I got a <strong>0.5036</strong> score on0 <code>r_square</code>.</p>\n<p>The process that follows is as follows.</p>\n<pre><code>x = pd.read_csv(test_path)\nt = model(x)\nsub = pd.read_csv(sub_path)\nsub[:, :] *= t\n\n</code></pre>\n<p>But, my submission score(LB) is <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6855922%2F737db36b1f0e274e2d2eb16619c665bf%2Fimage.png?generation=1714807819858109&amp;alt=media\"></p>\n<hr>\n<h1>What caused it?</h1>\n<p>I think</p>\n<ol>\n<li><p>Possibility that the distribution of data <code>train.csv</code> depends on the order of the data.<br>\n  For example, the possibility that sample_id is train_0 and train_1 data are similar.</p></li>\n<li><p>Data for training is small.</p></li>\n<li><p>Is this process necessary before training model? (I have not done this process)</p></li>\n</ol>\n<pre><code>\nmx = x.mean(axis=)\nsx = np.maximum(x.std(axis=), min_std)\nx = (x - mx.reshape(,-)) / sx.reshape(,-)\n  DEBUGGING:\n    xt = (xt - mx.reshape(,-)) / sx.reshape(,-)\n\n\nmy = y.mean(axis=)\nsy = np.maximum(np.sqrt((y*y).mean(axis=)), min_std)\ny = (y - my.reshape(,-)) / sy.reshape(,-)\n</code></pre>\n<hr>\n<p>I can't think of anything else.<br>\nPlease someone help me😭</p>",
      "rawMarkdown": "# Hi\nI am currently having trouble with a large difference between the R2 I calculated myself for model evaluation and the actual LB.\n\n## Here is the actual environment\n- model : Simple model created with pytorch\n- dataset : 50000 rows of data in train.csv\n  - train : 40000  rows (80%)\n  - valid : 5000 rows (10%)\n  - test : 5000 rows (10%, This will be used for the final model evaluation **R2**)\n- loss function : nn.MSELoss()\n\nLet's look at the change in the TRAIN and VALID error functions during training. It is over-trained, but not that bad.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6855922%2F94252ca74109f887af437e00fbd1c874%2Fimage.png?generation=1714806746529396&alt=media)\n\nThe values of the error function at the end of the study appear to be adequate.\n```\nepoch 40 train loss: 243.44\nepoch 40 valid loss: 427.88\n```\n\nI predicted `test_t` with this model and calculated the R2 score with the following function:\n```python\ndef r_squared(y_true, y_pred):\n    # SSres\n    ss_res = torch.sum((y_true - y_pred) ** 2)\n    \n    mean = torch.mean(y_true)\n    \n    # SStot\n    ss_tot = torch.sum((y_true - mean) ** 2)\n    \n    # R2\n    r2 = 1 - (ss_res / ss_tot)\n    \n    return r2\n```\nAnd, I got a **0.5036** score on0 `r_square`.\n\nThe process that follows is as follows.\n```python\nx = pd.read_csv(test_path)\nt = model(x)\nsub = pd.read_csv(sub_path)\nsub[:, 1:] *= t\n# submission\n```\n\nBut, my submission score(LB) is \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6855922%2F737db36b1f0e274e2d2eb16619c665bf%2Fimage.png?generation=1714807819858109&alt=media)\n\n---\n# What caused it?\nI think\n\n1. Possibility that the distribution of data `train.csv` depends on the order of the data.\n      For example, the possibility that sample_id is train_0 and train_1 data are similar.\n2. Data for training is small.\n\n3. Is this process necessary before training model? (I have not done this process)\n```python\n# norm X\nmx = x.mean(axis=0)\nsx = np.maximum(x.std(axis=0), min_std)\nx = (x - mx.reshape(1,-1)) / sx.reshape(1,-1)\nif not DEBUGGING:\n    xt = (xt - mx.reshape(1,-1)) / sx.reshape(1,-1)\n\n# norm Y\nmy = y.mean(axis=0)\nsy = np.maximum(np.sqrt((y*y).mean(axis=0)), min_std)\ny = (y - my.reshape(1,-1)) / sy.reshape(1,-1)\n```\n---\nI can't think of anything else.\nPlease someone help me😭",
      "votes": 3
    },
    {
      "id": 2793106,
      "postDate": "2024-05-04T14:47:07.210Z",
      "content": "<p>The code you posted at the end (for norm X and Y) is crucial for the submission, as well as multiplication with sample_submission data.<br>\nPlease check the discussion which I started <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/498806\" target=\"_blank\">here</a>.</p>\n<p>Here are my X and Y scaler classes, reference <a href=\"https://www.kaggle.com/code/ymatioun/leap-simple-nn/notebook\" target=\"_blank\">here</a>. </p>\n<pre><code> ():\n     ():\n        self.mean = \n        self.std = \n        self.min_std = min_std\n\n     ():\n        self.mean = X.mean(axis=)\n        self.std = np.maximum(X.std(axis=), self.min_std)\n\n     ():\n        X = (X - self.mean.reshape(, -)) / self.std.reshape(,-)\n         X\n</code></pre>\n<pre><code> ():\n     ():\n        self.mean = \n        self.s = \n        self.min_std=min_std\n\n     ():\n        self.mean = Y.mean(axis=)\n        self.s = np.maximum(np.sqrt((Y*Y).mean(axis=)), self.min_std)\n\n     ():\n        Y = (Y - self.mean.reshape(,-)) / self.s.reshape(,-)\n         Y\n\n     ():\n        \n         i  (self.s.shape[]):\n             self.s[i] &lt; self.min_std * :\n                Y_pred[:,i] = \n        \n        Y_pred = Y_pred * self.s.reshape(,-) + self.mean.reshape(,-)\n         Y_pred\n</code></pre>",
      "rawMarkdown": "The code you posted at the end (for norm X and Y) is crucial for the submission, as well as multiplication with sample_submission data.\nPlease check the discussion which I started [here](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/498806).\n\nHere are my X and Y scaler classes, reference [here](https://www.kaggle.com/code/ymatioun/leap-simple-nn/notebook). \n```python\nclass XScaler(object):\n    def __init__(self, min_std=1e-8):\n        self.mean = None\n        self.std = None\n        self.min_std = min_std\n\n    def fit(self, X):\n        self.mean = X.mean(axis=0)\n        self.std = np.maximum(X.std(axis=0), self.min_std)\n\n    def transform(self, X):\n        X = (X - self.mean.reshape(1, -1)) / self.std.reshape(1,-1)\n        return X\n```\n\n```python\nclass YScaler(object):\n    def __init__(self, min_std=1e-8):\n        self.mean = None\n        self.s = None\n        self.min_std=min_std\n\n    def fit(self, Y):\n        self.mean = Y.mean(axis=0)\n        self.s = np.maximum(np.sqrt((Y*Y).mean(axis=0)), self.min_std)\n\n    def transform(self, Y):\n        Y = (Y - self.mean.reshape(1,-1)) / self.s.reshape(1,-1)\n        return Y\n\n    def inverse_transform(self, Y_pred):\n        # override constant columns\n        for i in range(self.s.shape[0]):\n            if self.s[i] < self.min_std * 1.1:\n                Y_pred[:,i] = 0\n        # undo y scaling\n        Y_pred = Y_pred * self.s.reshape(1,-1) + self.mean.reshape(1,-1)\n        return Y_pred\n```",
      "votes": 1,
      "replies": [
        {
          "id": 2794228,
          "postDate": "2024-05-05T06:53:59.077Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 2793030,
      "postDate": "2024-05-04T13:56:14.213Z",
      "content": "<p>check submission file before submit</p>",
      "rawMarkdown": " check submission file before submit"
    }
  ],
  "comments": [
    {
      "id": 2793106,
      "author_name": "milann",
      "author_url": "",
      "post_date": "2024-05-04T14:47:07.210000",
      "content": "<p>The code you posted at the end (for norm X and Y) is crucial for the submission, as well as multiplication with sample_submission data.<br>\nPlease check the discussion which I started <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/498806\" target=\"_blank\">here</a>.</p>\n<p>Here are my X and Y scaler classes, reference <a href=\"https://www.kaggle.com/code/ymatioun/leap-simple-nn/notebook\" target=\"_blank\">here</a>. </p>\n<pre><code> ():\n     ():\n        self.mean = \n        self.std = \n        self.min_std = min_std\n\n     ():\n        self.mean = X.mean(axis=)\n        self.std = np.maximum(X.std(axis=), self.min_std)\n\n     ():\n        X = (X - self.mean.reshape(, -)) / self.std.reshape(,-)\n         X\n</code></pre>\n<pre><code> ():\n     ():\n        self.mean = \n        self.s = \n        self.min_std=min_std\n\n     ():\n        self.mean = Y.mean(axis=)\n        self.s = np.maximum(np.sqrt((Y*Y).mean(axis=)), self.min_std)\n\n     ():\n        Y = (Y - self.mean.reshape(,-)) / self.s.reshape(,-)\n         Y\n\n     ():\n        \n         i  (self.s.shape[]):\n             self.s[i] &lt; self.min_std * :\n                Y_pred[:,i] = \n        \n        Y_pred = Y_pred * self.s.reshape(,-) + self.mean.reshape(,-)\n         Y_pred\n</code></pre>",
      "votes": 1,
      "replies": [
        {
          "id": 2794228,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-05-05T06:53:59.077000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2793030,
      "author_name": "benkerrouche abdelbasset",
      "author_url": "",
      "post_date": "2024-05-04T13:56:14.213000",
      "content": "<p>check submission file before submit</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2792377": "# Hi\nI am currently having trouble with a large difference between the R2 I calculated myself for model evaluation and the actual LB.\n\n## Here is the actual environment\n- model : Simple model created with pytorch\n- dataset : 50000 rows of data in train.csv\n  - train : 40000  rows (80%)\n  - valid : 5000 rows (10%)\n  - test : 5000 rows (10%, This will be used for the final model evaluation **R2**)\n- loss function : nn.MSELoss()\n\nLet's look at the change in the TRAIN and VALID error functions during training. It is over-trained, but not that bad.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6855922%2F94252ca74109f887af437e00fbd1c874%2Fimage.png?generation=1714806746529396&alt=media)\n\nThe values of the error function at the end of the study appear to be adequate.\n```\nepoch 40 train loss: 243.44\nepoch 40 valid loss: 427.88\n```\n\nI predicted `test_t` with this model and calculated the R2 score with the following function:\n```python\ndef r_squared(y_true, y_pred):\n    # SSres\n    ss_res = torch.sum((y_true - y_pred) ** 2)\n    \n    mean = torch.mean(y_true)\n    \n    # SStot\n    ss_tot = torch.sum((y_true - mean) ** 2)\n    \n    # R2\n    r2 = 1 - (ss_res / ss_tot)\n    \n    return r2\n```\nAnd, I got a **0.5036** score on0 `r_square`.\n\nThe process that follows is as follows.\n```python\nx = pd.read_csv(test_path)\nt = model(x)\nsub = pd.read_csv(sub_path)\nsub[:, 1:] *= t\n# submission\n```\n\nBut, my submission score(LB) is \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6855922%2F737db36b1f0e274e2d2eb16619c665bf%2Fimage.png?generation=1714807819858109&alt=media)\n\n---\n# What caused it?\nI think\n\n1. Possibility that the distribution of data `train.csv` depends on the order of the data.\n      For example, the possibility that sample_id is train_0 and train_1 data are similar.\n2. Data for training is small.\n\n3. Is this process necessary before training model? (I have not done this process)\n```python\n# norm X\nmx = x.mean(axis=0)\nsx = np.maximum(x.std(axis=0), min_std)\nx = (x - mx.reshape(1,-1)) / sx.reshape(1,-1)\nif not DEBUGGING:\n    xt = (xt - mx.reshape(1,-1)) / sx.reshape(1,-1)\n\n# norm Y\nmy = y.mean(axis=0)\nsy = np.maximum(np.sqrt((y*y).mean(axis=0)), min_std)\ny = (y - my.reshape(1,-1)) / sy.reshape(1,-1)\n```\n---\nI can't think of anything else.\nPlease someone help me😭",
    "2793106": "The code you posted at the end (for norm X and Y) is crucial for the submission, as well as multiplication with sample_submission data.\nPlease check the discussion which I started [here](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/498806).\n\nHere are my X and Y scaler classes, reference [here](https://www.kaggle.com/code/ymatioun/leap-simple-nn/notebook). \n```python\nclass XScaler(object):\n    def __init__(self, min_std=1e-8):\n        self.mean = None\n        self.std = None\n        self.min_std = min_std\n\n    def fit(self, X):\n        self.mean = X.mean(axis=0)\n        self.std = np.maximum(X.std(axis=0), self.min_std)\n\n    def transform(self, X):\n        X = (X - self.mean.reshape(1, -1)) / self.std.reshape(1,-1)\n        return X\n```\n\n```python\nclass YScaler(object):\n    def __init__(self, min_std=1e-8):\n        self.mean = None\n        self.s = None\n        self.min_std=min_std\n\n    def fit(self, Y):\n        self.mean = Y.mean(axis=0)\n        self.s = np.maximum(np.sqrt((Y*Y).mean(axis=0)), self.min_std)\n\n    def transform(self, Y):\n        Y = (Y - self.mean.reshape(1,-1)) / self.s.reshape(1,-1)\n        return Y\n\n    def inverse_transform(self, Y_pred):\n        # override constant columns\n        for i in range(self.s.shape[0]):\n            if self.s[i] < self.min_std * 1.1:\n                Y_pred[:,i] = 0\n        # undo y scaling\n        Y_pred = Y_pred * self.s.reshape(1,-1) + self.mean.reshape(1,-1)\n        return Y_pred\n```",
    "2793030": " check submission file before submit"
  }
}