{
  "id": 536845,
  "title": "why linear model works best here?",
  "url": "/competitions/ariel-data-challenge-2024/discussion/536845",
  "author_name": "yuanzhe zhou",
  "post_date": "2024-09-30T09:19:53.225000",
  "votes": 20,
  "comment_count": 6,
  "views": 0,
  "content": "<p>It is surprising that the best score notebook  is based on linear regression</p>",
  "messages": [
    {
      "id": 3002670,
      "postDate": "2024-09-30T09:19:53.227Z",
      "content": "<p>It is surprising that the best score notebook  is based on linear regression</p>",
      "rawMarkdown": "It is surprising that the best score notebook  is based on linear regression",
      "votes": 18
    },
    {
      "id": 3002709,
      "postDate": "2024-09-30T10:06:23.313Z",
      "content": "<p>Probably because the test in this task differs from the training sample, and linear models are the most robust and, in addition, they are able to extrapolate (linear dependencies)</p>",
      "rawMarkdown": "Probably because the test in this task differs from the training sample, and linear models are the most robust and, in addition, they are able to extrapolate (linear dependencies)",
      "votes": 9,
      "replies": [
        {
          "id": 3008822,
          "postDate": "2024-10-07T06:36:23.660Z",
          "content": "<p>It seems impossible to predict the entire test dataset distribution using a DNN when the training datasets have significantly different distributions from the test data.<br>\nDo you think the linear model would be the best in this competition?</p>",
          "rawMarkdown": "It seems impossible to predict the entire test dataset distribution using a DNN when the training datasets have significantly different distributions from the test data.\nDo you think the linear model would be the best in this competition?"
        }
      ]
    },
    {
      "id": 3002681,
      "postDate": "2024-09-30T09:33:10.067Z",
      "content": "<p>Linear models work much better for certain problems with out-of-(training)distribution samples. E.g. for the dummy problem of predicting twice the input value: Your train dataset is x: [0, 1, 2, 3] -&gt; y: [0, 2, 4, 6]. No doubt, both an MLP with non-linear activations and a linear model will sucessfully learn this for the training dataset. But only the linear model may produce good results for x &gt;= 4. </p>\n<p>Particularly for an MLP with e.g. tanh or sigmoid we have to deal with saturation effects, which make the model lose sensitivity when the activation outputs approach their limit. The model may have learned keep the hidden-states close to 0 with a small variance, to effectively use the near linear part of the activation to learn the relationship of x -&gt; y. However, this breaks when feeding larger inputs.</p>\n<p>I guess in this competition we have to learn a function that is quite invariant to scaling of certain input dimensions. Thus a linear model can better handle much different distributions.</p>",
      "rawMarkdown": "Linear models work much better for certain problems with out-of-(training)distribution samples. E.g. for the dummy problem of predicting twice the input value: Your train dataset is x: [0, 1, 2, 3] -> y: [0, 2, 4, 6]. No doubt, both an MLP with non-linear activations and a linear model will sucessfully learn this for the training dataset. But only the linear model may produce good results for x >= 4. \n\nParticularly for an MLP with e.g. tanh or sigmoid we have to deal with saturation effects, which make the model lose sensitivity when the activation outputs approach their limit. The model may have learned keep the hidden-states close to 0 with a small variance, to effectively use the near linear part of the activation to learn the relationship of x -> y. However, this breaks when feeding larger inputs.\n\nI guess in this competition we have to learn a function that is quite invariant to scaling of certain input dimensions. Thus a linear model can better handle much different distributions.",
      "votes": 10
    },
    {
      "id": 3013460,
      "postDate": "2024-10-10T06:42:33.953Z",
      "content": "<p>I believe it is because there are a large number of features, using a linear model can achieve good fitting results. Secondly, it is caused by inconsistent data distribution.</p>",
      "rawMarkdown": "I believe it is because there are a large number of features, using a linear model can achieve good fitting results. Secondly, it is caused by inconsistent data distribution."
    },
    {
      "id": 3012328,
      "postDate": "2024-10-08T23:12:25.357Z",
      "content": "<p>But can we use CNN or attention layers to extract the better features? <br>\nI understand the training and test data distribution differ but why these architectures other than linear model on the training distribution are not robust ?</p>",
      "rawMarkdown": "But can we use CNN or attention layers to extract the better features? \nI understand the training and test data distribution differ but why these architectures other than linear model on the training distribution are not robust ?",
      "replies": [
        {
          "id": 3012860,
          "postDate": "2024-10-09T12:35:09.677Z",
          "content": "<p>With sufficient training data, it is possible that CNN or transformer architectures would achieve better results. 673 planets is a relatively small sample on which to train those types of models.</p>",
          "rawMarkdown": "With sufficient training data, it is possible that CNN or transformer architectures would achieve better results. 673 planets is a relatively small sample on which to train those types of models.",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3002709,
      "author_name": "Sergei Fironov",
      "author_url": "",
      "post_date": "2024-09-30T10:06:23.313000",
      "content": "<p>Probably because the test in this task differs from the training sample, and linear models are the most robust and, in addition, they are able to extrapolate (linear dependencies)</p>",
      "votes": 9,
      "replies": [
        {
          "id": 3008822,
          "author_name": "skhan kim",
          "author_url": "",
          "post_date": "2024-10-07T06:36:23.660000",
          "content": "<p>It seems impossible to predict the entire test dataset distribution using a DNN when the training datasets have significantly different distributions from the test data.<br>\nDo you think the linear model would be the best in this competition?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3002681,
      "author_name": "Fritz Cremer",
      "author_url": "",
      "post_date": "2024-09-30T09:33:10.067000",
      "content": "<p>Linear models work much better for certain problems with out-of-(training)distribution samples. E.g. for the dummy problem of predicting twice the input value: Your train dataset is x: [0, 1, 2, 3] -&gt; y: [0, 2, 4, 6]. No doubt, both an MLP with non-linear activations and a linear model will sucessfully learn this for the training dataset. But only the linear model may produce good results for x &gt;= 4. </p>\n<p>Particularly for an MLP with e.g. tanh or sigmoid we have to deal with saturation effects, which make the model lose sensitivity when the activation outputs approach their limit. The model may have learned keep the hidden-states close to 0 with a small variance, to effectively use the near linear part of the activation to learn the relationship of x -&gt; y. However, this breaks when feeding larger inputs.</p>\n<p>I guess in this competition we have to learn a function that is quite invariant to scaling of certain input dimensions. Thus a linear model can better handle much different distributions.</p>",
      "votes": 10,
      "replies": []
    },
    {
      "id": 3013460,
      "author_name": "zy",
      "author_url": "",
      "post_date": "2024-10-10T06:42:33.953000",
      "content": "<p>I believe it is because there are a large number of features, using a linear model can achieve good fitting results. Secondly, it is caused by inconsistent data distribution.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3012328,
      "author_name": "highDopamine",
      "author_url": "",
      "post_date": "2024-10-08T23:12:25.357000",
      "content": "<p>But can we use CNN or attention layers to extract the better features? <br>\nI understand the training and test data distribution differ but why these architectures other than linear model on the training distribution are not robust ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3012860,
          "author_name": "Steven Hewitt",
          "author_url": "",
          "post_date": "2024-10-09T12:35:09.677000",
          "content": "<p>With sufficient training data, it is possible that CNN or transformer architectures would achieve better results. 673 planets is a relatively small sample on which to train those types of models.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3002670": "It is surprising that the best score notebook  is based on linear regression",
    "3002709": "Probably because the test in this task differs from the training sample, and linear models are the most robust and, in addition, they are able to extrapolate (linear dependencies)",
    "3002681": "Linear models work much better for certain problems with out-of-(training)distribution samples. E.g. for the dummy problem of predicting twice the input value: Your train dataset is x: [0, 1, 2, 3] -> y: [0, 2, 4, 6]. No doubt, both an MLP with non-linear activations and a linear model will sucessfully learn this for the training dataset. But only the linear model may produce good results for x >= 4. \n\nParticularly for an MLP with e.g. tanh or sigmoid we have to deal with saturation effects, which make the model lose sensitivity when the activation outputs approach their limit. The model may have learned keep the hidden-states close to 0 with a small variance, to effectively use the near linear part of the activation to learn the relationship of x -> y. However, this breaks when feeding larger inputs.\n\nI guess in this competition we have to learn a function that is quite invariant to scaling of certain input dimensions. Thus a linear model can better handle much different distributions.",
    "3013460": "I believe it is because there are a large number of features, using a linear model can achieve good fitting results. Secondly, it is caused by inconsistent data distribution.",
    "3012328": "But can we use CNN or attention layers to extract the better features? \nI understand the training and test data distribution differ but why these architectures other than linear model on the training distribution are not robust ?"
  }
}