{
  "id": 511146,
  "title": "Train val split loss difference",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/511146",
  "author_name": "Vasilis",
  "post_date": "2024-06-09T10:22:41.391000",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hello guys, i have various transformer seq-2-seq encoder-only models. I split to 9m train and 1m validation set and all of my models reach arround 0.63 r2 score. Then i split to another 9m train and 1m validation set and the models strugle at 0.56. With so many data i was under the impression that splits should not make much of a difference, i assumed that the sets would be more or less homogenous. Does it make sense to observe such a big differences between different splits? The validation sets share around 100k common rows.</p>",
  "messages": [
    {
      "id": 2863389,
      "postDate": "2024-06-09T11:26:23.463Z",
      "content": "<p>0.63 is not good and you probably have some negative r2 score targets. Fluctuation could be happening because of that. As your model gets stronger, it will start to get less and less negative r2 scores. </p>",
      "rawMarkdown": "0.63 is not good and you probably have some negative r2 score targets. Fluctuation could be happening because of that. As your model gets stronger, it will start to get less and less negative r2 scores. ",
      "votes": 5,
      "replies": [
        {
          "id": 2864655,
          "postDate": "2024-06-10T09:07:51.697Z",
          "content": "<p>Hey Gunes, i checked my losses per target variable, i have only 1 slightly negative target loss. What i noticed though is that the losses are very tiny in the first layers and they gets bigger in the higher layers, can this be because of the way i feed the data in the transformer? i start from 0 layer up to 59, i think that makes sense. Any idea why the transformer deteriorates as layers progress?</p>",
          "rawMarkdown": "Hey Gunes, i checked my losses per target variable, i have only 1 slightly negative target loss. What i noticed though is that the losses are very tiny in the first layers and they gets bigger in the higher layers, can this be because of the way i feed the data in the transformer? i start from 0 layer up to 59, i think that makes sense. Any idea why the transformer deteriorates as layers progress?",
          "replies": [
            {
              "id": 2864752,
              "postDate": "2024-06-10T09:59:14.007Z",
              "content": "<p>'I start from 0 layer up to 59'<br>\nYour architecture doesn't makes sense to me…sounds more like a decoder</p>",
              "rawMarkdown": "'I start from 0 layer up to 59'\nYour architecture doesn't makes sense to me...sounds more like a decoder",
              "votes": 2
            },
            {
              "id": 2864838,
              "postDate": "2024-06-10T10:59:58.357Z",
              "content": "<p>no its an encoder only, my train data are N, 60, 25, oh i see so the order from 0 to 59 in the encoder does not matter i guess. Anyway i observe higher loss in higher layers</p>",
              "rawMarkdown": "no its an encoder only, my train data are N, 60, 25, oh i see so the order from 0 to 59 in the encoder does not matter i guess. Anyway i observe higher loss in higher layers\n"
            }
          ]
        }
      ]
    },
    {
      "id": 2863329,
      "postDate": "2024-06-09T10:22:41.393Z",
      "content": "<p>Hello guys, i have various transformer seq-2-seq encoder-only models. I split to 9m train and 1m validation set and all of my models reach arround 0.63 r2 score. Then i split to another 9m train and 1m validation set and the models strugle at 0.56. With so many data i was under the impression that splits should not make much of a difference, i assumed that the sets would be more or less homogenous. Does it make sense to observe such a big differences between different splits? The validation sets share around 100k common rows.</p>",
      "rawMarkdown": "Hello guys, i have various transformer seq-2-seq encoder-only models. I split to 9m train and 1m validation set and all of my models reach arround 0.63 r2 score. Then i split to another 9m train and 1m validation set and the models strugle at 0.56. With so many data i was under the impression that splits should not make much of a difference, i assumed that the sets would be more or less homogenous. Does it make sense to observe such a big differences between different splits? The validation sets share around 100k common rows.",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 2863389,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2024-06-09T11:26:23.463000",
      "content": "<p>0.63 is not good and you probably have some negative r2 score targets. Fluctuation could be happening because of that. As your model gets stronger, it will start to get less and less negative r2 scores. </p>",
      "votes": 5,
      "replies": [
        {
          "id": 2864655,
          "author_name": "Vasilis",
          "author_url": "",
          "post_date": "2024-06-10T09:07:51.697000",
          "content": "<p>Hey Gunes, i checked my losses per target variable, i have only 1 slightly negative target loss. What i noticed though is that the losses are very tiny in the first layers and they gets bigger in the higher layers, can this be because of the way i feed the data in the transformer? i start from 0 layer up to 59, i think that makes sense. Any idea why the transformer deteriorates as layers progress?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2864752,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-06-10T09:59:14.007000",
              "content": "<p>'I start from 0 layer up to 59'<br>\nYour architecture doesn't makes sense to me…sounds more like a decoder</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2864838,
              "author_name": "Vasilis",
              "author_url": "",
              "post_date": "2024-06-10T10:59:58.357000",
              "content": "<p>no its an encoder only, my train data are N, 60, 25, oh i see so the order from 0 to 59 in the encoder does not matter i guess. Anyway i observe higher loss in higher layers</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2863389": "0.63 is not good and you probably have some negative r2 score targets. Fluctuation could be happening because of that. As your model gets stronger, it will start to get less and less negative r2 scores. ",
    "2863329": "Hello guys, i have various transformer seq-2-seq encoder-only models. I split to 9m train and 1m validation set and all of my models reach arround 0.63 r2 score. Then i split to another 9m train and 1m validation set and the models strugle at 0.56. With so many data i was under the impression that splits should not make much of a difference, i assumed that the sets would be more or less homogenous. Does it make sense to observe such a big differences between different splits? The validation sets share around 100k common rows."
  }
}