{
  "id": 513675,
  "title": "Can model with 1M parameters achieve R2 score greater than 0.76?",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/513675",
  "author_name": "Sijun Xu",
  "post_date": "2024-06-21T06:38:38.395000",
  "votes": 15,
  "comment_count": 21,
  "views": 0,
  "content": "<p>After some attempts I use a single model with ~1M parameters to get R2 socre ~0.74 on lb. The model architecture is relative simple basically just CNN1d-like layers, and the training parameters are not optimized.  I have tried transformer encoder with positional encoding but seems cannot achieve higher R2-score. The following are some thoughts to further improve but I did not try:</p>\n<ol>\n<li>Further improve model architecture. Use models with larger parameters like transformer encoders.</li>\n<li>Search training hyperparameters(lr, lr_scheduler, batch_size etc.).</li>\n</ol>\n<p>Any advice is welcom. You can share your R2 score and number of model parameters here :)</p>\n<p>[update1]<br>\nThanks to <a href=\"https://www.kaggle.com/rob1080ti\" target=\"_blank\">@rob1080ti</a> for the notification of R2 score plot. Below is the R2 score in my cv for all targets.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2469728%2F2568e4718efff2cc706e84ead0662555%2Fr2_out.png?generation=1719037269160917&amp;alt=media\"><br>\nThe definition of risk is from the discussion <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490</a> . From the plot, it seems that some targets with higher risk indeed gives lower R2 score. But there are still some low risk targets hard to predict.</p>\n<p>[update2]<br>\nInspired from the notebook <a href=\"https://www.kaggle.com/code/ucas0v0zhuoqunli/plot-on-a-map\" target=\"_blank\">https://www.kaggle.com/code/ucas0v0zhuoqunli/plot-on-a-map</a> by <a href=\"https://www.kaggle.com/ucas0v0zhuoqunli\" target=\"_blank\">@ucas0v0zhuoqunli</a> and previous discussions, I make some global R2 maps for these targets. Here is an example for heating tendency.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2469728%2F122bd64274436a2296d1dca8cf02a3d8%2Fmap_ptend_t.png?generation=1719037724620052&amp;alt=media\"><br>\nThe global structure is similar to the original paper <a href=\"https://arxiv.org/pdf/2306.08754\" target=\"_blank\">https://arxiv.org/pdf/2306.08754</a>. Maybe the location info is not fully learned by my model.</p>\n<p>[update3]<br>\n2.4M params of model, cv/lb: 0.753/0.754. By improving model architecture the performance improves!</p>",
  "messages": [
    {
      "id": 2882063,
      "postDate": "2024-06-21T06:38:38.397Z",
      "content": "<p>After some attempts I use a single model with ~1M parameters to get R2 socre ~0.74 on lb. The model architecture is relative simple basically just CNN1d-like layers, and the training parameters are not optimized.  I have tried transformer encoder with positional encoding but seems cannot achieve higher R2-score. The following are some thoughts to further improve but I did not try:</p>\n<ol>\n<li>Further improve model architecture. Use models with larger parameters like transformer encoders.</li>\n<li>Search training hyperparameters(lr, lr_scheduler, batch_size etc.).</li>\n</ol>\n<p>Any advice is welcom. You can share your R2 score and number of model parameters here :)</p>\n<p>[update1]<br>\nThanks to <a href=\"https://www.kaggle.com/rob1080ti\" target=\"_blank\">@rob1080ti</a> for the notification of R2 score plot. Below is the R2 score in my cv for all targets.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2469728%2F2568e4718efff2cc706e84ead0662555%2Fr2_out.png?generation=1719037269160917&amp;alt=media\"><br>\nThe definition of risk is from the discussion <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490</a> . From the plot, it seems that some targets with higher risk indeed gives lower R2 score. But there are still some low risk targets hard to predict.</p>\n<p>[update2]<br>\nInspired from the notebook <a href=\"https://www.kaggle.com/code/ucas0v0zhuoqunli/plot-on-a-map\" target=\"_blank\">https://www.kaggle.com/code/ucas0v0zhuoqunli/plot-on-a-map</a> by <a href=\"https://www.kaggle.com/ucas0v0zhuoqunli\" target=\"_blank\">@ucas0v0zhuoqunli</a> and previous discussions, I make some global R2 maps for these targets. Here is an example for heating tendency.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2469728%2F122bd64274436a2296d1dca8cf02a3d8%2Fmap_ptend_t.png?generation=1719037724620052&amp;alt=media\"><br>\nThe global structure is similar to the original paper <a href=\"https://arxiv.org/pdf/2306.08754\" target=\"_blank\">https://arxiv.org/pdf/2306.08754</a>. Maybe the location info is not fully learned by my model.</p>\n<p>[update3]<br>\n2.4M params of model, cv/lb: 0.753/0.754. By improving model architecture the performance improves!</p>",
      "rawMarkdown": "After some attempts I use a single model with ~1M parameters to get R2 socre ~0.74 on lb. The model architecture is relative simple basically just CNN1d-like layers, and the training parameters are not optimized.  I have tried transformer encoder with positional encoding but seems cannot achieve higher R2-score. The following are some thoughts to further improve but I did not try:\n\n1. Further improve model architecture. Use models with larger parameters like transformer encoders.\n2. Search training hyperparameters(lr, lr_scheduler, batch_size etc.).\n\nAny advice is welcom. You can share your R2 score and number of model parameters here :)\n\n[update1]\nThanks to @rob1080ti for the notification of R2 score plot. Below is the R2 score in my cv for all targets.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2469728%2F2568e4718efff2cc706e84ead0662555%2Fr2_out.png?generation=1719037269160917&alt=media)\nThe definition of risk is from the discussion https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490 . From the plot, it seems that some targets with higher risk indeed gives lower R2 score. But there are still some low risk targets hard to predict.\n\n[update2]\nInspired from the notebook https://www.kaggle.com/code/ucas0v0zhuoqunli/plot-on-a-map by @ucas0v0zhuoqunli and previous discussions, I make some global R2 maps for these targets. Here is an example for heating tendency.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2469728%2F122bd64274436a2296d1dca8cf02a3d8%2Fmap_ptend_t.png?generation=1719037724620052&alt=media)\nThe global structure is similar to the original paper https://arxiv.org/pdf/2306.08754. Maybe the location info is not fully learned by my model.\n\n[update3]\n2.4M params of model, cv/lb: 0.753/0.754. By improving model architecture the performance improves!",
      "votes": 15
    },
    {
      "id": 2883223,
      "postDate": "2024-06-21T19:10:21.123Z",
      "content": "<p>I already tried models with 1M, 5M, 10M, and 20M parameters, but I can't achieve more than 0.71. <br>\nI feel I am missing a type of connection or layer in my models or something like that. <br>\nAnother suspicion I have is that I am using float32 data instead of float64. <br>\nI tried to reproduce Amadeo's TFRecords dataset for float64 but without success.</p>",
      "rawMarkdown": "I already tried models with 1M, 5M, 10M, and 20M parameters, but I can't achieve more than 0.71. \nI feel I am missing a type of connection or layer in my models or something like that. \nAnother suspicion I have is that I am using float32 data instead of float64. \nI tried to reproduce Amadeo's TFRecords dataset for float64 but without success.",
      "votes": 1,
      "replies": [
        {
          "id": 2883458,
          "postDate": "2024-06-22T02:14:46.517Z",
          "content": "<blockquote>\n  <p>I tried to reproduce Amadeo's TFRecords dataset for float64 but without success.</p>\n</blockquote>\n<p>What does this mean? I am also trying to reproduce the dataset for float64. </p>",
          "rawMarkdown": ">I tried to reproduce Amadeo's TFRecords dataset for float64 but without success.\n\nWhat does this mean? I am also trying to reproduce the dataset for float64. ",
          "votes": 1,
          "replies": [
            {
              "id": 2883556,
              "postDate": "2024-06-22T03:46:18.480Z",
              "content": "<p>Why not use his? There are normalized and pre normalized datasets he made available. Also, I think there is code for making tfrecords in the official GitHub.</p>",
              "rawMarkdown": "Why not use his? There are normalized and pre normalized datasets he made available. Also, I think there is code for making tfrecords in the official GitHub.",
              "votes": 2
            },
            {
              "id": 2883585,
              "postDate": "2024-06-22T04:07:01.203Z",
              "content": "<p>It is just for my learning process. In fact, I used his tfrecords dataset. I did not know the code is in the  githup. I will look for it. Thanks for the tip.</p>",
              "rawMarkdown": "It is just for my learning process. In fact, I used his tfrecords dataset. I did not know the code is in the  githup. I will look for it. Thanks for the tip.",
              "votes": 2
            }
          ]
        },
        {
          "id": 2883860,
          "postDate": "2024-06-22T07:02:41Z",
          "content": "<p>In my case, I only use float32 to train, but preprocessing data should be done with float64.</p>",
          "rawMarkdown": "In my case, I only use float32 to train, but preprocessing data should be done with float64.",
          "votes": 1
        },
        {
          "id": 2884549,
          "postDate": "2024-06-22T15:34:33.123Z",
          "content": "<p><a href=\"https://www.kaggle.com/nandodmelo\" target=\"_blank\">@nandodmelo</a> Can your architecture achieve positive R2 scores on targets ptend q0003 12-15? I think my architecture would be able to achieve 0.72+ if not get negative R2 scores on these targets.</p>",
          "rawMarkdown": "@nandodmelo Can your architecture achieve positive R2 scores on targets ptend q0003 12-15? I think my architecture would be able to achieve 0.72+ if not get negative R2 scores on these targets.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2882932,
      "postDate": "2024-06-21T16:36:13.447Z",
      "content": "<p>You could compare your R2 graph with ones that others have shared to see if there are areas that look similar/different which may be attributable to the model architecture. I am not sure how much model architecture matters in comparison to preprocessing the data except for the highest scores which may be ensembles anyhow.</p>",
      "rawMarkdown": "You could compare your R2 graph with ones that others have shared to see if there are areas that look similar/different which may be attributable to the model architecture. I am not sure how much model architecture matters in comparison to preprocessing the data except for the highest scores which may be ensembles anyhow.",
      "votes": 1,
      "replies": [
        {
          "id": 2883829,
          "postDate": "2024-06-22T06:50:50.440Z",
          "content": "<p>That's correct, preprocessing the data is crucial. I am just curious that the potential of one single model in this competition. The R2 graph from my model shows there are some targets can be improved. Thanks!</p>",
          "rawMarkdown": "That's correct, preprocessing the data is crucial. I am just curious that the potential of one single model in this competition. The R2 graph from my model shows there are some targets can be improved. Thanks!",
          "votes": 3
        }
      ]
    },
    {
      "id": 2882763,
      "postDate": "2024-06-21T14:27:11.353Z",
      "content": "<p>I'm curious, too. How much data did you use for training?</p>",
      "rawMarkdown": "I'm curious, too. How much data did you use for training?",
      "votes": 1,
      "replies": [
        {
          "id": 2883813,
          "postDate": "2024-06-22T06:38:05.580Z",
          "content": "<p>I only use data provided by kaggle.</p>",
          "rawMarkdown": "I only use data provided by kaggle.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2882682,
      "postDate": "2024-06-21T13:28:54.283Z",
      "content": "<p>i am curious how many epochs did you train to achieve 0.74 with this model? i train a ~10 million transformer but i have never achieved more than 0.7. I wonder if this is because i never trained for more than 10 epochs.</p>",
      "rawMarkdown": "i am curious how many epochs did you train to achieve 0.74 with this model? i train a ~10 million transformer but i have never achieved more than 0.7. I wonder if this is because i never trained for more than 10 epochs.",
      "votes": 1,
      "replies": [
        {
          "id": 2883811,
          "postDate": "2024-06-22T06:37:36.890Z",
          "content": "<p>I trained for 50 epochs, I think there are a lot of room to improve.</p>",
          "rawMarkdown": "I trained for 50 epochs, I think there are a lot of room to improve.",
          "votes": 2,
          "replies": [
            {
              "id": 2886909,
              "postDate": "2024-06-23T22:20:15.667Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 2899608,
      "postDate": "2024-07-01T18:36:05.953Z",
      "content": "<p>I am also stuck at 0.74, I don't think its because of model size, I think it is because of the architecture or data preprocessing, does mean/variance vs RMS scaling make a big difference?</p>",
      "rawMarkdown": "I am also stuck at 0.74, I don't think its because of model size, I think it is because of the architecture or data preprocessing, does mean/variance vs RMS scaling make a big difference?",
      "votes": 2,
      "replies": [
        {
          "id": 2899835,
          "postDate": "2024-07-01T21:57:39.360Z",
          "content": "<p>I guess the question is whether you are really stuck. I thought I was stuck too and subsequently I’ve been nudging my model up with more epochs.</p>",
          "rawMarkdown": "I guess the question is whether you are really stuck. I thought I was stuck too and subsequently I’ve been nudging my model up with more epochs.",
          "replies": [
            {
              "id": 2899843,
              "postDate": "2024-07-01T22:09:16.637Z",
              "content": "<p>If it has too many parameters it gets 0.74 and starts overfitting after more epochs, and if it smaller it gets stuck at 0.74, maybe I have to use dropout.</p>",
              "rawMarkdown": "If it has too many parameters it gets 0.74 and starts overfitting after more epochs, and if it smaller it gets stuck at 0.74, maybe I have to use dropout."
            },
            {
              "id": 2899865,
              "postDate": "2024-07-01T22:21:44.870Z",
              "content": "<p>That could be wrt Dropout. Have you tried reloading your saved model with a different objective, optimizer, and or parameters?</p>",
              "rawMarkdown": "That could be wrt Dropout. Have you tried reloading your saved model with a different objective, optimizer, and or parameters?"
            }
          ]
        }
      ]
    },
    {
      "id": 2906932,
      "postDate": "2024-07-05T21:11:16.603Z",
      "content": "<p>Did you use any pre-trained model for feature extraction or only created it from scratch? </p>",
      "rawMarkdown": "Did you use any pre-trained model for feature extraction or only created it from scratch? "
    },
    {
      "id": 2894777,
      "postDate": "2024-06-28T17:21:03.247Z",
      "content": "<p>How did you draw this picture?</p>",
      "rawMarkdown": "How did you draw this picture?",
      "replies": [
        {
          "id": 2898133,
          "postDate": "2024-07-01T01:18:18.807Z",
          "content": "<p>The first R2 graph is the valididated R2 for each target. The second R2 map is the R2 score per location per target. You can find how to  draw this map in the given notebook.</p>",
          "rawMarkdown": "The first R2 graph is the valididated R2 for each target. The second R2 map is the R2 score per location per target. You can find how to  draw this map in the given notebook.",
          "replies": [
            {
              "id": 2898736,
              "postDate": "2024-07-01T10:01:20.667Z",
              "content": "<p>thanks for your reply</p>",
              "rawMarkdown": "thanks for your reply"
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2883223,
      "author_name": "Fernando Melo",
      "author_url": "",
      "post_date": "2024-06-21T19:10:21.123000",
      "content": "<p>I already tried models with 1M, 5M, 10M, and 20M parameters, but I can't achieve more than 0.71. <br>\nI feel I am missing a type of connection or layer in my models or something like that. <br>\nAnother suspicion I have is that I am using float32 data instead of float64. <br>\nI tried to reproduce Amadeo's TFRecords dataset for float64 but without success.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2883458,
          "author_name": "ibi",
          "author_url": "",
          "post_date": "2024-06-22T02:14:46.517000",
          "content": "<blockquote>\n  <p>I tried to reproduce Amadeo's TFRecords dataset for float64 but without success.</p>\n</blockquote>\n<p>What does this mean? I am also trying to reproduce the dataset for float64. </p>",
          "votes": 1,
          "replies": [
            {
              "id": 2883556,
              "author_name": "Rob Freeman",
              "author_url": "",
              "post_date": "2024-06-22T03:46:18.480000",
              "content": "<p>Why not use his? There are normalized and pre normalized datasets he made available. Also, I think there is code for making tfrecords in the official GitHub.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2883585,
              "author_name": "ibi",
              "author_url": "",
              "post_date": "2024-06-22T04:07:01.203000",
              "content": "<p>It is just for my learning process. In fact, I used his tfrecords dataset. I did not know the code is in the  githup. I will look for it. Thanks for the tip.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        },
        {
          "id": 2883860,
          "author_name": "Sijun Xu",
          "author_url": "",
          "post_date": "2024-06-22T07:02:41",
          "content": "<p>In my case, I only use float32 to train, but preprocessing data should be done with float64.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2884549,
          "author_name": "永遠的三號波",
          "author_url": "",
          "post_date": "2024-06-22T15:34:33.123000",
          "content": "<p><a href=\"https://www.kaggle.com/nandodmelo\" target=\"_blank\">@nandodmelo</a> Can your architecture achieve positive R2 scores on targets ptend q0003 12-15? I think my architecture would be able to achieve 0.72+ if not get negative R2 scores on these targets.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2882932,
      "author_name": "Rob Freeman",
      "author_url": "",
      "post_date": "2024-06-21T16:36:13.447000",
      "content": "<p>You could compare your R2 graph with ones that others have shared to see if there are areas that look similar/different which may be attributable to the model architecture. I am not sure how much model architecture matters in comparison to preprocessing the data except for the highest scores which may be ensembles anyhow.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2883829,
          "author_name": "Sijun Xu",
          "author_url": "",
          "post_date": "2024-06-22T06:50:50.440000",
          "content": "<p>That's correct, preprocessing the data is crucial. I am just curious that the potential of one single model in this competition. The R2 graph from my model shows there are some targets can be improved. Thanks!</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2882763,
      "author_name": "Zhuoqun Li",
      "author_url": "",
      "post_date": "2024-06-21T14:27:11.353000",
      "content": "<p>I'm curious, too. How much data did you use for training?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2883813,
          "author_name": "Sijun Xu",
          "author_url": "",
          "post_date": "2024-06-22T06:38:05.580000",
          "content": "<p>I only use data provided by kaggle.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2882682,
      "author_name": "Vasilis",
      "author_url": "",
      "post_date": "2024-06-21T13:28:54.283000",
      "content": "<p>i am curious how many epochs did you train to achieve 0.74 with this model? i train a ~10 million transformer but i have never achieved more than 0.7. I wonder if this is because i never trained for more than 10 epochs.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2883811,
          "author_name": "Sijun Xu",
          "author_url": "",
          "post_date": "2024-06-22T06:37:36.890000",
          "content": "<p>I trained for 50 epochs, I think there are a lot of room to improve.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2886909,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-06-23T22:20:15.667000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2899608,
      "author_name": "Emanuel Ruzak",
      "author_url": "",
      "post_date": "2024-07-01T18:36:05.953000",
      "content": "<p>I am also stuck at 0.74, I don't think its because of model size, I think it is because of the architecture or data preprocessing, does mean/variance vs RMS scaling make a big difference?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2899835,
          "author_name": "Rob Freeman",
          "author_url": "",
          "post_date": "2024-07-01T21:57:39.360000",
          "content": "<p>I guess the question is whether you are really stuck. I thought I was stuck too and subsequently I’ve been nudging my model up with more epochs.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2899843,
              "author_name": "Emanuel Ruzak",
              "author_url": "",
              "post_date": "2024-07-01T22:09:16.637000",
              "content": "<p>If it has too many parameters it gets 0.74 and starts overfitting after more epochs, and if it smaller it gets stuck at 0.74, maybe I have to use dropout.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2899865,
              "author_name": "Rob Freeman",
              "author_url": "",
              "post_date": "2024-07-01T22:21:44.870000",
              "content": "<p>That could be wrt Dropout. Have you tried reloading your saved model with a different objective, optimizer, and or parameters?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2906932,
      "author_name": "Maren Sajdaras",
      "author_url": "",
      "post_date": "2024-07-05T21:11:16.603000",
      "content": "<p>Did you use any pre-trained model for feature extraction or only created it from scratch? </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2894777,
      "author_name": "yanqiangmiffy",
      "author_url": "",
      "post_date": "2024-06-28T17:21:03.247000",
      "content": "<p>How did you draw this picture?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2898133,
          "author_name": "Sijun Xu",
          "author_url": "",
          "post_date": "2024-07-01T01:18:18.807000",
          "content": "<p>The first R2 graph is the valididated R2 for each target. The second R2 map is the R2 score per location per target. You can find how to  draw this map in the given notebook.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2898736,
              "author_name": "yanqiangmiffy",
              "author_url": "",
              "post_date": "2024-07-01T10:01:20.667000",
              "content": "<p>thanks for your reply</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2882063": "After some attempts I use a single model with ~1M parameters to get R2 socre ~0.74 on lb. The model architecture is relative simple basically just CNN1d-like layers, and the training parameters are not optimized.  I have tried transformer encoder with positional encoding but seems cannot achieve higher R2-score. The following are some thoughts to further improve but I did not try:\n\n1. Further improve model architecture. Use models with larger parameters like transformer encoders.\n2. Search training hyperparameters(lr, lr_scheduler, batch_size etc.).\n\nAny advice is welcom. You can share your R2 score and number of model parameters here :)\n\n[update1]\nThanks to @rob1080ti for the notification of R2 score plot. Below is the R2 score in my cv for all targets.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2469728%2F2568e4718efff2cc706e84ead0662555%2Fr2_out.png?generation=1719037269160917&alt=media)\nThe definition of risk is from the discussion https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490 . From the plot, it seems that some targets with higher risk indeed gives lower R2 score. But there are still some low risk targets hard to predict.\n\n[update2]\nInspired from the notebook https://www.kaggle.com/code/ucas0v0zhuoqunli/plot-on-a-map by @ucas0v0zhuoqunli and previous discussions, I make some global R2 maps for these targets. Here is an example for heating tendency.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2469728%2F122bd64274436a2296d1dca8cf02a3d8%2Fmap_ptend_t.png?generation=1719037724620052&alt=media)\nThe global structure is similar to the original paper https://arxiv.org/pdf/2306.08754. Maybe the location info is not fully learned by my model.\n\n[update3]\n2.4M params of model, cv/lb: 0.753/0.754. By improving model architecture the performance improves!",
    "2883223": "I already tried models with 1M, 5M, 10M, and 20M parameters, but I can't achieve more than 0.71. \nI feel I am missing a type of connection or layer in my models or something like that. \nAnother suspicion I have is that I am using float32 data instead of float64. \nI tried to reproduce Amadeo's TFRecords dataset for float64 but without success.",
    "2882932": "You could compare your R2 graph with ones that others have shared to see if there are areas that look similar/different which may be attributable to the model architecture. I am not sure how much model architecture matters in comparison to preprocessing the data except for the highest scores which may be ensembles anyhow.",
    "2882763": "I'm curious, too. How much data did you use for training?",
    "2882682": "i am curious how many epochs did you train to achieve 0.74 with this model? i train a ~10 million transformer but i have never achieved more than 0.7. I wonder if this is because i never trained for more than 10 epochs.",
    "2899608": "I am also stuck at 0.74, I don't think its because of model size, I think it is because of the architecture or data preprocessing, does mean/variance vs RMS scaling make a big difference?",
    "2906932": "Did you use any pre-trained model for feature extraction or only created it from scratch? ",
    "2894777": "How did you draw this picture?"
  }
}