{
  "id": 527416,
  "title": "48th place solution, Ensemble + Oversampling outliers + Weighted loss + U connections",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/527416",
  "author_name": "Emanuel Ruzak",
  "post_date": "2024-08-12T02:08:40.223000",
  "votes": 2,
  "comment_count": 0,
  "views": 0,
  "content": "<h3>Introduction</h3>\n<p>In this competition we had to predict the derivative of variables representing the state of an atmospheric column. It's basically a sequence to sequence prediction task - the first 60 features show each atmospheric layer's state, and the last one covers general and ground features. I started with an MLP before moving to a transformer model.</p>\n<h5>U connections</h5>\n<p>Predicting derivatives is similar to what diffusion models do. So instead of a vanilla transformer, I took inspiration from the paper <a href=\"https://arxiv.org/abs/2209.12152\" target=\"_blank\">All are Worth Words: A ViT Backbone for Diffusion Models</a> and went with a U-ViT model. It's a transformer with U-Net-like connections, which makes sense because it lets information flow from the first layer to the last without getting interference from layers in between.</p>\n<h5>Oversampling and Weighted loss</h5>\n<p>This dataset has features and targets with very different magnitudes. This competition uses MSE loss, so a small amount of rows dominates the loss. <br>\nThis makes training unstable, since the high-impact rows don't appear in most mini-batches, but when one appears, it destabilizes the training.<br>\nThis causes the model to take about 150 epochs before achieving good performance.</p>\n<p>To reduce the training time to 5 epochs, we determined how much each row would contribute to the error if our prediction was just the target variables' mean.<br>\nThen we repeated rows based on how much their error exceeded the mean. So, for example if the mean error is 2.5 and a row has an error of 10, it gets repeated 4 times. Rows below the mean error don't get removed, just not repeated.</p>\n<p>Then I used a custom loss function - it's still MSE, but divided by how many times that row was repeated. This stabilizes training by making sure each row contributes equally to the error, without changing what we're actually trying to achieve.</p>\n<h5>Other tricks</h5>\n<p>I also used a closed form prediction for features 142 to 149 (q0002) based on this <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/502378\" target=\"_blank\">discussion</a>. <br>\nBefore training the network, I replaced those features and their corresponding targets with zeros.</p>\n<h5>Ensembling</h5>\n<p>For my final submission, I used 6 models with different seeds for the initial shuffling before the train-test split. This improved the score from 0.741 to 0.757.</p>\n<h5>Training code (for 1 model)</h5>\n<p>Here is the code for a single model: <a href=\"https://www.kaggle.com/code/emanuelruzak/fork-of-simple-bert-v23/notebook\" target=\"_blank\">https://www.kaggle.com/code/emanuelruzak/fork-of-simple-bert-v23/notebook</a></p>",
  "messages": [
    {
      "id": 2956355,
      "postDate": "2024-08-12T02:08:40.223Z",
      "content": "<h3>Introduction</h3>\n<p>In this competition we had to predict the derivative of variables representing the state of an atmospheric column. It's basically a sequence to sequence prediction task - the first 60 features show each atmospheric layer's state, and the last one covers general and ground features. I started with an MLP before moving to a transformer model.</p>\n<h5>U connections</h5>\n<p>Predicting derivatives is similar to what diffusion models do. So instead of a vanilla transformer, I took inspiration from the paper <a href=\"https://arxiv.org/abs/2209.12152\" target=\"_blank\">All are Worth Words: A ViT Backbone for Diffusion Models</a> and went with a U-ViT model. It's a transformer with U-Net-like connections, which makes sense because it lets information flow from the first layer to the last without getting interference from layers in between.</p>\n<h5>Oversampling and Weighted loss</h5>\n<p>This dataset has features and targets with very different magnitudes. This competition uses MSE loss, so a small amount of rows dominates the loss. <br>\nThis makes training unstable, since the high-impact rows don't appear in most mini-batches, but when one appears, it destabilizes the training.<br>\nThis causes the model to take about 150 epochs before achieving good performance.</p>\n<p>To reduce the training time to 5 epochs, we determined how much each row would contribute to the error if our prediction was just the target variables' mean.<br>\nThen we repeated rows based on how much their error exceeded the mean. So, for example if the mean error is 2.5 and a row has an error of 10, it gets repeated 4 times. Rows below the mean error don't get removed, just not repeated.</p>\n<p>Then I used a custom loss function - it's still MSE, but divided by how many times that row was repeated. This stabilizes training by making sure each row contributes equally to the error, without changing what we're actually trying to achieve.</p>\n<h5>Other tricks</h5>\n<p>I also used a closed form prediction for features 142 to 149 (q0002) based on this <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/502378\" target=\"_blank\">discussion</a>. <br>\nBefore training the network, I replaced those features and their corresponding targets with zeros.</p>\n<h5>Ensembling</h5>\n<p>For my final submission, I used 6 models with different seeds for the initial shuffling before the train-test split. This improved the score from 0.741 to 0.757.</p>\n<h5>Training code (for 1 model)</h5>\n<p>Here is the code for a single model: <a href=\"https://www.kaggle.com/code/emanuelruzak/fork-of-simple-bert-v23/notebook\" target=\"_blank\">https://www.kaggle.com/code/emanuelruzak/fork-of-simple-bert-v23/notebook</a></p>",
      "rawMarkdown": "###Introduction\nIn this competition we had to predict the derivative of variables representing the state of an atmospheric column. It's basically a sequence to sequence prediction task - the first 60 features show each atmospheric layer's state, and the last one covers general and ground features. I started with an MLP before moving to a transformer model.\n\n#####U connections\nPredicting derivatives is similar to what diffusion models do. So instead of a vanilla transformer, I took inspiration from the paper [All are Worth Words: A ViT Backbone for Diffusion Models](https://arxiv.org/abs/2209.12152) and went with a U-ViT model. It's a transformer with U-Net-like connections, which makes sense because it lets information flow from the first layer to the last without getting interference from layers in between.\n\n#####Oversampling and Weighted loss\nThis dataset has features and targets with very different magnitudes. This competition uses MSE loss, so a small amount of rows dominates the loss. \nThis makes training unstable, since the high-impact rows don't appear in most mini-batches, but when one appears, it destabilizes the training.\nThis causes the model to take about 150 epochs before achieving good performance.\n\nTo reduce the training time to 5 epochs, we determined how much each row would contribute to the error if our prediction was just the target variables' mean.\nThen we repeated rows based on how much their error exceeded the mean. So, for example if the mean error is 2.5 and a row has an error of 10, it gets repeated 4 times. Rows below the mean error don't get removed, just not repeated.\n\nThen I used a custom loss function - it's still MSE, but divided by how many times that row was repeated. This stabilizes training by making sure each row contributes equally to the error, without changing what we're actually trying to achieve.\n\n#####Other tricks\nI also used a closed form prediction for features 142 to 149 (q0002) based on this [discussion](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/502378). \nBefore training the network, I replaced those features and their corresponding targets with zeros.\n\n#####Ensembling\nFor my final submission, I used 6 models with different seeds for the initial shuffling before the train-test split. This improved the score from 0.741 to 0.757.\n\n#####Training code (for 1 model)\nHere is the code for a single model: https://www.kaggle.com/code/emanuelruzak/fork-of-simple-bert-v23/notebook",
      "votes": 2
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2956355": "###Introduction\nIn this competition we had to predict the derivative of variables representing the state of an atmospheric column. It's basically a sequence to sequence prediction task - the first 60 features show each atmospheric layer's state, and the last one covers general and ground features. I started with an MLP before moving to a transformer model.\n\n#####U connections\nPredicting derivatives is similar to what diffusion models do. So instead of a vanilla transformer, I took inspiration from the paper [All are Worth Words: A ViT Backbone for Diffusion Models](https://arxiv.org/abs/2209.12152) and went with a U-ViT model. It's a transformer with U-Net-like connections, which makes sense because it lets information flow from the first layer to the last without getting interference from layers in between.\n\n#####Oversampling and Weighted loss\nThis dataset has features and targets with very different magnitudes. This competition uses MSE loss, so a small amount of rows dominates the loss. \nThis makes training unstable, since the high-impact rows don't appear in most mini-batches, but when one appears, it destabilizes the training.\nThis causes the model to take about 150 epochs before achieving good performance.\n\nTo reduce the training time to 5 epochs, we determined how much each row would contribute to the error if our prediction was just the target variables' mean.\nThen we repeated rows based on how much their error exceeded the mean. So, for example if the mean error is 2.5 and a row has an error of 10, it gets repeated 4 times. Rows below the mean error don't get removed, just not repeated.\n\nThen I used a custom loss function - it's still MSE, but divided by how many times that row was repeated. This stabilizes training by making sure each row contributes equally to the error, without changing what we're actually trying to achieve.\n\n#####Other tricks\nI also used a closed form prediction for features 142 to 149 (q0002) based on this [discussion](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/502378). \nBefore training the network, I replaced those features and their corresponding targets with zeros.\n\n#####Ensembling\nFor my final submission, I used 6 models with different seeds for the initial shuffling before the train-test split. This improved the score from 0.741 to 0.757.\n\n#####Training code (for 1 model)\nHere is the code for a single model: https://www.kaggle.com/code/emanuelruzak/fork-of-simple-bert-v23/notebook"
  }
}