{
  "id": 529393,
  "title": "14th place solution",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/529393",
  "author_name": "ohkawa3",
  "post_date": "2024-08-20T14:48:43.491000",
  "votes": 8,
  "comment_count": 1,
  "views": 0,
  "content": "<h1>TL;DR</h1>\n<ul>\n<li>Combined CNN and MLP-Mixer models.</li>\n<li>Trained on 161,464,320 data points from ClimSim_low-res and ClimSim_low-res_aqua-planet datasets.</li>\n<li>Features were normalized by mean and standard deviation, then squashed using the tanh function.</li>\n<li>Targets were normalized to achieve an RMS of 1, using SmoothL1Loss (beta=0.5).</li>\n<li>ptend_q0002 values from 12 to 27 were calculated from state_q0002.</li>\n</ul>\n<h1>Preprocessing</h1>\n<h2>Feature</h2>\n<p>Two methods were applied for feature normalization. <br>\nThe first method normalized the mean and standard deviation individually, while the second normalized them by group. <br>\nIt was found that these normalized features had significant outliers, causing instability during training. <br>\nTo address this, a tanh squashing function was used, which keeps most of the value range linear while restricting the range to 1024.</p>\n<pre><code>x = np.tanh(x / ) * \n</code></pre>\n<p>All preprocessing was done in 64-bit. The two normalized features were then concatenated.</p>\n<h2>Target</h2>\n<p>As discussed in the link below, the targets were normalized to achieve an RMS of 1:<br>\n<a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/498806\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/498806</a></p>\n<h1>Model</h1>\n<p>A custom neural network model was created based on MLP-Mixer and CNN. The vector data used in this model has the following characteristics:</p>\n<ul>\n<li>Fixed length of 60.</li>\n<li>Different characteristics at each level (e.g., significant temperature differences between the surface and upper atmosphere).</li>\n<li>Strong correlations between adjacent levels.</li>\n<li>The effectiveness of global information aggregation was confirmed in experiments.</li>\n</ul>\n<p>Earlier experiments involved developing a model based on 1D CNN. It was observed that performance improved when the kernel size was increased, promoting global information aggregation. Layers such as LSTM or Transformer, capable of global information aggregation, were added, but this significantly increased computational complexity.</p>\n<p>MLP-Mixer, while not capable of handling variable-length data like LSTM or Transformer, can efficiently aggregate global information for fixed-length data. Additionally, MLP-Mixer has the advantage of learning unique filters for each of the 60 levels. Based on this MLP-Mixer, the following two changes were made:</p>\n<ul>\n<li>The channel-mixer was changed from a 1D-conv with a kernel size of 1 to a kernel size of 3.<ul>\n<li>This change allows for feature extraction from adjacent levels.</li></ul></li>\n<li>LayerNorm was replaced with BatchNorm.<ul>\n<li>For this dataset, BatchNorm was more efficient in terms of both processing time and performance.</li></ul></li>\n</ul>\n<h1>Training</h1>\n<p>Initially, only ClimSim_low-res was used. However, since there was about 1TB of storage available, ClimSim_low-res_aqua-planet was also downloaded. Training with ClimSim_low-res_aqua-planet alone resulted in a CV of 0.19, which is low but not completely unusable. When ClimSim_low-res and ClimSim_low-res_aqua-planet were mixed for training, a slight improvement in CV was observed.</p>\n<p>ClimSim_high-res was downloaded and sampled to about 1TB, but it was realized that the download would not be completed until September, so the attempt was abandoned.</p>\n<p>Training was performed on 161,464,320 data points with a batch size of 256 over 3 epochs, taking approximately 27 hours on an RTX 4090. The AdamW optimizer was used with a warmup for 7,358 iterations and TanhLRScheduler. Training was done with Amp (float16), while inference was done in float32.</p>\n<h1>Post-processing</h1>\n<ul>\n<li>Post-processing was done in 64-bit.</li>\n<li>Since normalization was applied, the data was reverted to the original scale.</li>\n<li>NaN values were filled with the mean.</li>\n<li>ptend_q0002 values from 12 to 27 were calculated from state_q0002.</li>\n<li>The predicted results were constrained within the range of the minimum and maximum values of the training data.</li>\n</ul>\n<h1>Ensemble</h1>\n<ul>\n<li>The average of 8 models with different hyper parametors(seed, depth) was used.</li>\n<li>The individual Public LB scores for each model range from 0.784 to 0.786, and the Private LB scores range from 0.778 to 0.781.</li>\n</ul>",
  "messages": [
    {
      "id": 2965102,
      "postDate": "2024-08-20T14:48:43.490Z",
      "content": "<h1>TL;DR</h1>\n<ul>\n<li>Combined CNN and MLP-Mixer models.</li>\n<li>Trained on 161,464,320 data points from ClimSim_low-res and ClimSim_low-res_aqua-planet datasets.</li>\n<li>Features were normalized by mean and standard deviation, then squashed using the tanh function.</li>\n<li>Targets were normalized to achieve an RMS of 1, using SmoothL1Loss (beta=0.5).</li>\n<li>ptend_q0002 values from 12 to 27 were calculated from state_q0002.</li>\n</ul>\n<h1>Preprocessing</h1>\n<h2>Feature</h2>\n<p>Two methods were applied for feature normalization. <br>\nThe first method normalized the mean and standard deviation individually, while the second normalized them by group. <br>\nIt was found that these normalized features had significant outliers, causing instability during training. <br>\nTo address this, a tanh squashing function was used, which keeps most of the value range linear while restricting the range to 1024.</p>\n<pre><code>x = np.tanh(x / ) * \n</code></pre>\n<p>All preprocessing was done in 64-bit. The two normalized features were then concatenated.</p>\n<h2>Target</h2>\n<p>As discussed in the link below, the targets were normalized to achieve an RMS of 1:<br>\n<a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/498806\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/498806</a></p>\n<h1>Model</h1>\n<p>A custom neural network model was created based on MLP-Mixer and CNN. The vector data used in this model has the following characteristics:</p>\n<ul>\n<li>Fixed length of 60.</li>\n<li>Different characteristics at each level (e.g., significant temperature differences between the surface and upper atmosphere).</li>\n<li>Strong correlations between adjacent levels.</li>\n<li>The effectiveness of global information aggregation was confirmed in experiments.</li>\n</ul>\n<p>Earlier experiments involved developing a model based on 1D CNN. It was observed that performance improved when the kernel size was increased, promoting global information aggregation. Layers such as LSTM or Transformer, capable of global information aggregation, were added, but this significantly increased computational complexity.</p>\n<p>MLP-Mixer, while not capable of handling variable-length data like LSTM or Transformer, can efficiently aggregate global information for fixed-length data. Additionally, MLP-Mixer has the advantage of learning unique filters for each of the 60 levels. Based on this MLP-Mixer, the following two changes were made:</p>\n<ul>\n<li>The channel-mixer was changed from a 1D-conv with a kernel size of 1 to a kernel size of 3.<ul>\n<li>This change allows for feature extraction from adjacent levels.</li></ul></li>\n<li>LayerNorm was replaced with BatchNorm.<ul>\n<li>For this dataset, BatchNorm was more efficient in terms of both processing time and performance.</li></ul></li>\n</ul>\n<h1>Training</h1>\n<p>Initially, only ClimSim_low-res was used. However, since there was about 1TB of storage available, ClimSim_low-res_aqua-planet was also downloaded. Training with ClimSim_low-res_aqua-planet alone resulted in a CV of 0.19, which is low but not completely unusable. When ClimSim_low-res and ClimSim_low-res_aqua-planet were mixed for training, a slight improvement in CV was observed.</p>\n<p>ClimSim_high-res was downloaded and sampled to about 1TB, but it was realized that the download would not be completed until September, so the attempt was abandoned.</p>\n<p>Training was performed on 161,464,320 data points with a batch size of 256 over 3 epochs, taking approximately 27 hours on an RTX 4090. The AdamW optimizer was used with a warmup for 7,358 iterations and TanhLRScheduler. Training was done with Amp (float16), while inference was done in float32.</p>\n<h1>Post-processing</h1>\n<ul>\n<li>Post-processing was done in 64-bit.</li>\n<li>Since normalization was applied, the data was reverted to the original scale.</li>\n<li>NaN values were filled with the mean.</li>\n<li>ptend_q0002 values from 12 to 27 were calculated from state_q0002.</li>\n<li>The predicted results were constrained within the range of the minimum and maximum values of the training data.</li>\n</ul>\n<h1>Ensemble</h1>\n<ul>\n<li>The average of 8 models with different hyper parametors(seed, depth) was used.</li>\n<li>The individual Public LB scores for each model range from 0.784 to 0.786, and the Private LB scores range from 0.778 to 0.781.</li>\n</ul>",
      "rawMarkdown": "# TL;DR\n+ Combined CNN and MLP-Mixer models.\n+ Trained on 161,464,320 data points from ClimSim_low-res and ClimSim_low-res_aqua-planet datasets.\n+ Features were normalized by mean and standard deviation, then squashed using the tanh function.\n+ Targets were normalized to achieve an RMS of 1, using SmoothL1Loss (beta=0.5).\n+ ptend_q0002 values from 12 to 27 were calculated from state_q0002.\n\n# Preprocessing\n## Feature\nTwo methods were applied for feature normalization. \nThe first method normalized the mean and standard deviation individually, while the second normalized them by group. \nIt was found that these normalized features had significant outliers, causing instability during training. \nTo address this, a tanh squashing function was used, which keeps most of the value range linear while restricting the range to 1024.\n```python\nx = np.tanh(x / 1024.0) * 1024.0\n```\nAll preprocessing was done in 64-bit. The two normalized features were then concatenated.\n\n## Target\nAs discussed in the link below, the targets were normalized to achieve an RMS of 1:\nhttps://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/498806\n\n# Model\nA custom neural network model was created based on MLP-Mixer and CNN. The vector data used in this model has the following characteristics:\n- Fixed length of 60.\n- Different characteristics at each level (e.g., significant temperature differences between the surface and upper atmosphere).\n- Strong correlations between adjacent levels.\n- The effectiveness of global information aggregation was confirmed in experiments.\n\nEarlier experiments involved developing a model based on 1D CNN. It was observed that performance improved when the kernel size was increased, promoting global information aggregation. Layers such as LSTM or Transformer, capable of global information aggregation, were added, but this significantly increased computational complexity.\n\nMLP-Mixer, while not capable of handling variable-length data like LSTM or Transformer, can efficiently aggregate global information for fixed-length data. Additionally, MLP-Mixer has the advantage of learning unique filters for each of the 60 levels. Based on this MLP-Mixer, the following two changes were made:\n+ The channel-mixer was changed from a 1D-conv with a kernel size of 1 to a kernel size of 3.\n  + This change allows for feature extraction from adjacent levels.\n+ LayerNorm was replaced with BatchNorm.\n  + For this dataset, BatchNorm was more efficient in terms of both processing time and performance.\n\n# Training\nInitially, only ClimSim_low-res was used. However, since there was about 1TB of storage available, ClimSim_low-res_aqua-planet was also downloaded. Training with ClimSim_low-res_aqua-planet alone resulted in a CV of 0.19, which is low but not completely unusable. When ClimSim_low-res and ClimSim_low-res_aqua-planet were mixed for training, a slight improvement in CV was observed.\n\nClimSim_high-res was downloaded and sampled to about 1TB, but it was realized that the download would not be completed until September, so the attempt was abandoned.\n\nTraining was performed on 161,464,320 data points with a batch size of 256 over 3 epochs, taking approximately 27 hours on an RTX 4090. The AdamW optimizer was used with a warmup for 7,358 iterations and TanhLRScheduler. Training was done with Amp (float16), while inference was done in float32.\n\n# Post-processing\n+ Post-processing was done in 64-bit.\n+ Since normalization was applied, the data was reverted to the original scale.\n+ NaN values were filled with the mean.\n+ ptend_q0002 values from 12 to 27 were calculated from state_q0002.\n+ The predicted results were constrained within the range of the minimum and maximum values of the training data.\n\n# Ensemble\n+ The average of 8 models with different hyper parametors(seed, depth) was used.\n+ The individual Public LB scores for each model range from 0.784 to 0.786, and the Private LB scores range from 0.778 to 0.781.\n\n",
      "votes": 8
    },
    {
      "id": 2965254,
      "postDate": "2024-08-20T17:21:08.493Z",
      "content": "<blockquote>\n  <p>but it was realized that the download would not be completed until September, so the attempt was abandoned.</p>\n</blockquote>\n<p>lol 😅</p>",
      "rawMarkdown": "> but it was realized that the download would not be completed until September, so the attempt was abandoned.\n\nlol 😅",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 2965254,
      "author_name": "slime",
      "author_url": "",
      "post_date": "2024-08-20T17:21:08.493000",
      "content": "<blockquote>\n  <p>but it was realized that the download would not be completed until September, so the attempt was abandoned.</p>\n</blockquote>\n<p>lol 😅</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2965102": "# TL;DR\n+ Combined CNN and MLP-Mixer models.\n+ Trained on 161,464,320 data points from ClimSim_low-res and ClimSim_low-res_aqua-planet datasets.\n+ Features were normalized by mean and standard deviation, then squashed using the tanh function.\n+ Targets were normalized to achieve an RMS of 1, using SmoothL1Loss (beta=0.5).\n+ ptend_q0002 values from 12 to 27 were calculated from state_q0002.\n\n# Preprocessing\n## Feature\nTwo methods were applied for feature normalization. \nThe first method normalized the mean and standard deviation individually, while the second normalized them by group. \nIt was found that these normalized features had significant outliers, causing instability during training. \nTo address this, a tanh squashing function was used, which keeps most of the value range linear while restricting the range to 1024.\n```python\nx = np.tanh(x / 1024.0) * 1024.0\n```\nAll preprocessing was done in 64-bit. The two normalized features were then concatenated.\n\n## Target\nAs discussed in the link below, the targets were normalized to achieve an RMS of 1:\nhttps://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/498806\n\n# Model\nA custom neural network model was created based on MLP-Mixer and CNN. The vector data used in this model has the following characteristics:\n- Fixed length of 60.\n- Different characteristics at each level (e.g., significant temperature differences between the surface and upper atmosphere).\n- Strong correlations between adjacent levels.\n- The effectiveness of global information aggregation was confirmed in experiments.\n\nEarlier experiments involved developing a model based on 1D CNN. It was observed that performance improved when the kernel size was increased, promoting global information aggregation. Layers such as LSTM or Transformer, capable of global information aggregation, were added, but this significantly increased computational complexity.\n\nMLP-Mixer, while not capable of handling variable-length data like LSTM or Transformer, can efficiently aggregate global information for fixed-length data. Additionally, MLP-Mixer has the advantage of learning unique filters for each of the 60 levels. Based on this MLP-Mixer, the following two changes were made:\n+ The channel-mixer was changed from a 1D-conv with a kernel size of 1 to a kernel size of 3.\n  + This change allows for feature extraction from adjacent levels.\n+ LayerNorm was replaced with BatchNorm.\n  + For this dataset, BatchNorm was more efficient in terms of both processing time and performance.\n\n# Training\nInitially, only ClimSim_low-res was used. However, since there was about 1TB of storage available, ClimSim_low-res_aqua-planet was also downloaded. Training with ClimSim_low-res_aqua-planet alone resulted in a CV of 0.19, which is low but not completely unusable. When ClimSim_low-res and ClimSim_low-res_aqua-planet were mixed for training, a slight improvement in CV was observed.\n\nClimSim_high-res was downloaded and sampled to about 1TB, but it was realized that the download would not be completed until September, so the attempt was abandoned.\n\nTraining was performed on 161,464,320 data points with a batch size of 256 over 3 epochs, taking approximately 27 hours on an RTX 4090. The AdamW optimizer was used with a warmup for 7,358 iterations and TanhLRScheduler. Training was done with Amp (float16), while inference was done in float32.\n\n# Post-processing\n+ Post-processing was done in 64-bit.\n+ Since normalization was applied, the data was reverted to the original scale.\n+ NaN values were filled with the mean.\n+ ptend_q0002 values from 12 to 27 were calculated from state_q0002.\n+ The predicted results were constrained within the range of the minimum and maximum values of the training data.\n\n# Ensemble\n+ The average of 8 models with different hyper parametors(seed, depth) was used.\n+ The individual Public LB scores for each model range from 0.784 to 0.786, and the Private LB scores range from 0.778 to 0.781.\n\n",
    "2965254": "> but it was realized that the download would not be completed until September, so the attempt was abandoned.\n\nlol 😅"
  }
}