{
  "id": 523042,
  "title": "4th place solution",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/523042",
  "author_name": "kurupical",
  "post_date": "2024-07-29T22:44:16.795000",
  "votes": 34,
  "comment_count": 1,
  "views": 0,
  "content": "<p>First of all, we want to thank Kaggle and <a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> for hosting a competiton. And I would like to thank our teammate <a href=\"https://www.kaggle.com/takoihiraokazu\" target=\"_blank\">@takoihiraokazu</a> and <a href=\"https://www.kaggle.com/kami634\" target=\"_blank\">@kami634</a>! It was a pleasure to team up with you both.</p>\n<h1>1. Kurupical part</h1>\n<h2>1-1. Summary</h2>\n<table>\n<thead>\n<tr>\n<th>#</th>\n<th>model_name</th>\n<th>cv</th>\n<th>public</th>\n<th>private</th>\n<th>training time</th>\n<th>note</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>convnext 64x3-128x3-256x27-512x3</td>\n<td>0.7881</td>\n<td>0.78577</td>\n<td>0.78147</td>\n<td>60h with 1xRTX4090</td>\n<td></td>\n</tr>\n<tr>\n<td>2</td>\n<td>convnext 96x3-192x3-384x27-768x3</td>\n<td>0.7887</td>\n<td>0.78697</td>\n<td>0.78281</td>\n<td>120h with 1xRTX4090</td>\n<td></td>\n</tr>\n<tr>\n<td>3</td>\n<td>convnext 128x3-256x3-512x27-1024x3</td>\n<td>0.78693</td>\n<td>0.78480</td>\n<td>0.78150</td>\n<td>132h with 1xRTX4090</td>\n<td>use 81.25% data</td>\n</tr>\n<tr>\n<td>4</td>\n<td>convnext 144x3-288x3-576x27-1152x3</td>\n<td>0.78547</td>\n<td>0.78317</td>\n<td>0.77903</td>\n<td>28h with 8xRTX4090</td>\n<td></td>\n</tr>\n<tr>\n<td>5</td>\n<td>transformer 512x4</td>\n<td>0.7848</td>\n<td>0.78341</td>\n<td>0.77910</td>\n<td>54h with 1xRTX4090</td>\n<td></td>\n</tr>\n<tr>\n<td>6</td>\n<td>transformer 768x4</td>\n<td>0.7843</td>\n<td>0.78350</td>\n<td>0.77966</td>\n<td>72h with 1xRTX4090</td>\n<td></td>\n</tr>\n<tr>\n<td>7</td>\n<td>convnext 144x3-288x3-576x27-1152x3</td>\n<td>0.78769</td>\n<td>0.78622</td>\n<td>0.78137</td>\n<td>16h with 8xRTX4090</td>\n<td>training 4epochs</td>\n</tr>\n</tbody>\n</table>\n<h2>1-2. Feature Engineering / Preprocessing</h2>\n<p>I use all row-les datasets.</p>\n<h3>1-2-1. Feature Engineering</h3>\n<ul>\n<li>diff<ul>\n<li>x[i] - x[i-1]</li>\n<li>x[i] - x[i-2]</li></ul></li>\n<li>mean and diff_mean in similar features<ul>\n<li>q0002, q0003</li>\n<li>state_u, state_v</li>\n<li>pbuf_*</li></ul></li>\n</ul>\n<h3>1-2-2. Preprocessing</h3>\n<ul>\n<li>Standard Scaler for feature / label<ul>\n<li>Calculate feature mean/std with both train/test datasets.</li></ul></li>\n<li>Extremely large values can make model training unstable, so features are standardized and then clipped to the range of -100 to 100.</li>\n<li>To convert to the shape of (batch_size, n_feature, 60), scalar values (e.g., state_ps) are transformed into time series data by repeating the same value 60 times.</li>\n</ul>\n<h2>1-3. Training Methods</h2>\n<h3>1-3-1. Models</h3>\n<h4>1-3-1-1. ConvNeXt</h4>\n<p>A ConvNext with inputs of shape (batch_size, n_features, 60) and outputs of shape (batch_size, 368). The pseudocode is shown below.</p>\n<pre><code> (nn.Module):\n     ():\n        (Head1D, self).__init__()\n        self.final_layer = nn.LazyConv1d(out_channels=, kernel_size=, stride=, padding=)\n        self.fc = nn.Linear( * , )\n\n     ():\n        x = self.final_layer(x)  \n        x_out = torch.cat(\n            [\n                x[:, :, :].reshape(\n                    -,  * \n                ),  \n                self.fc(x[:, :, :].reshape(-,  * )),  \n            ],\n            dim=,\n        )  \n         x_out\n\n (nn.Module):\n     ():\n        self.head = Head1D()\n        ...\n\n     ():\n        \n        ...\n\n     ():\n        x = x[]  \n\n        \n        x_features = self.forward_feature(x)  \n\n        x_pred = self.head(x_features)  \n         x_pred\n\nmodel = ConvNeXt()\nx = torch.randn(, , )\n model(x).shape == (, )\n</code></pre>\n<ul>\n<li>Based on <a href=\"https://github.com/facebookresearch/ConvNeXt/blob/d1fa8f6fef0a165b27399986cc2bdacc92777e40/models/convnext.py\" target=\"_blank\">https://github.com/facebookresearch/ConvNeXt/blob/d1fa8f6fef0a165b27399986cc2bdacc92777e40/models/convnext.py</a>  ,I've rewritten it for 1D inputs with the following considerations:<ul>\n<li>Adjusted to ensure the shape remains unchanged between input and output by setting stride=1 and padding=\"same\".</li>\n<li>Replaced LayerNorm with BatchNorm since LayerNorm was not effective.</li></ul></li>\n</ul>\n<h4>1-3-1-2. Transformer</h4>\n<p>A Transformer with input of shape  <code>(batch_size, 60, n_features)</code> and outputs of shape <code>(batch_size, 368)</code>.  The pseudocode is shown below.</p>\n<pre><code> (nn.Module):\n     ():\n        (TransformerModel, self).__init__()\n\n        N_SEQUENTIAL_COLUMNS = \n        \n        self.fc_stem = nn.Sequential(\n            nn.LazyLinear(hidden_dims),\n            nn.LayerNorm(hidden_dims),\n            nn.GELU(),\n            nn.LazyLinear(hidden_dims),\n            nn.LayerNorm(hidden_dims),\n            nn.GELU(),\n        )\n        self.position_encoder = nn.Embedding(, hidden_dims)\n\n        \n        layer = nn.TransformerEncoderLayer(\n            d_model=hidden_dims,\n            nhead=n_heads,\n            dim_feedforward=hidden_dims * ,\n            dropout=,\n            activation=,\n            batch_first=,\n        )\n        self.transformer = nn.TransformerEncoder(layer, num_layers=n_layers)        \n\n        \n        \n        self.fc_head_list_sequential = []\n        head_cnn_params = [{: , : }] * \n         _  (N_SEQUENTIAL_COLUMNS):\n            fc_head = []\n             head_cnn_param  head_cnn_params:\n                fc_head.append(\n                    nn.LazyConv1d(\n                        stride=,\n                        padding=,\n                        **head_cnn_param,\n                    )\n                )\n                fc_head.append(nn.BatchNorm1d(head_cnn_param[]))\n                fc_head.append(nn.GELU())\n            fc_head.append(\n                nn.LazyConv1d(, kernel_size=, stride=, padding=)\n            )\n            fc_head = nn.Sequential(*fc_head)\n            self.fc_head_list_sequential.append(fc_head)\n        self.fc_head_list_sequential = nn.ModuleList(self.fc_head_list_sequential)\n\n        \n        self.fc_head_scalar = nn.Linear(hidden_dims, )\n\n     ():\n\n        \n        x = self.fc_stem(x)  \n        pe = self.position_encoder(\n            torch.arange(x.shape[], device=x.device).expand(x.shape[], -)\n        )  \n        x = x + pe\n\n        \n        x = self.transformer(x)  \n\n        \n        \n        x_out = []\n         i  ():\n            x_out_ = self.fc_head_list_sequential[i](\n                x.permute(, , )\n            )  \n            x_out_ = x_out_.squeeze()  \n            x_out.append(x_out_)        \n        \n        x_out_ = self.fc_head_scalar(x.mean(dim=))  \n        x_out.append(x_out_)\n        x_out = torch.cat(x_out, dim=)  \n         x_out        \n</code></pre>\n<h3>1-3-2. Hyper Parameter</h3>\n<table>\n<thead>\n<tr>\n<th>param_name</th>\n<th>#1</th>\n<th>#2</th>\n<th>#3</th>\n<th>#4</th>\n<th>#5</th>\n<th>#6</th>\n<th>#7</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>epochs</td>\n<td>7</td>\n<td>7</td>\n<td>7</td>\n<td>7</td>\n<td>7</td>\n<td>7</td>\n<td>4</td>\n</tr>\n<tr>\n<td>lr</td>\n<td>2.5e-3</td>\n<td>2e-3</td>\n<td>2e-3</td>\n<td>1.5e-3</td>\n<td>1e-3</td>\n<td>1e-3</td>\n<td>2e-3</td>\n</tr>\n<tr>\n<td>batch_size</td>\n<td>384</td>\n<td>384</td>\n<td>384</td>\n<td>288</td>\n<td>384</td>\n<td>384</td>\n<td>288</td>\n</tr>\n<tr>\n<td>weight_decay</td>\n<td>0.05</td>\n<td>0.05</td>\n<td>0.05</td>\n<td>0.05</td>\n<td>0.01</td>\n<td>0.01</td>\n<td>0.075</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>other<ul>\n<li>optimizer: AdamW<ul>\n<li>cnn: AdamW(weight_decay=0.05 or 0.075)</li>\n<li>transformer: AdamW(weight_decay=0.01)</li></ul></li>\n<li>loss: <ul>\n<li>cnn: SmoothL1Loss(beta=0.01)</li>\n<li>transformer: SmoothL1Loss(beta=1)</li></ul></li>\n<li>scheduler<ul>\n<li>cnn: transformers.get_polynomial_decay_schedule_with_warmup(alpha=2, warmup_ratio=0.1)</li>\n<li>transformer: transformers.get_polynomial_decay_schedule_with_warmup(alpha=1, warmup_ratio=0.1), equals to get_linear_scheduler_with_warmup</li></ul></li>\n<li>The batch size of 384 is a remnant of using lat/lon leak.</li>\n<li>Transformers tend to diverge quickly with a high learning rate, so keep it as low as possible. CNNs were not as sensitive.</li>\n<li>Regarding the beta of SmoothL1Loss, smaller models tended to perform better with something close to Huber Loss, while larger models showed better performance with something closer to L1 Loss.</li></ul></li>\n</ul>\n<h2>1-4. Postprocessing</h2>\n<ul>\n<li>Set some values to state * (1 / -1200)<ul>\n<li>ptend_q0002_12 ~ ptend_q0002_28</li></ul></li>\n</ul>\n<h2>1-5. Others</h2>\n<p>This was the competition where we needed to rely on neural networks the most among all the competitions we've participated in so far.</p>\n<ul>\n<li>Common<ul>\n<li>Proper learning rate scheduling was crucial, and we spent a significant amount of time tuning it.</li>\n<li>We tested with a small amount of data (n=1.8m, n=9m) before using the full dataset (n=70m). There were cases where what worked with a small amount of data didn't work with the full dataset.</li>\n<li>The same issue occurred between small models (ConvNeXt 32x3-64x3-128x27-256x3) and large models (ConvNeXt 96x3-192x3-384x27-768x3).</li></ul></li>\n<li>ConvNeXt<ul>\n<li>Polynomial decay scheduler and high weight decay were effective.</li>\n<li>Observing the train loss on W&amp;B, it seemed that training progressed significantly when the learning rate was low. Therefore, we adopted a polynomial decay scheduler to stay at a low learning rate for a longer period. We found that training progressed too much and caused overfitting, so we controlled it with weight decay. (cv: +0.004)</li>\n<li>It seemed that the larger the model, the higher the accuracy, given proper hyperparameter settings.<ul>\n<li>Larger models converged faster, so we reduced the number of epochs.</li>\n<li>We didn't have time to test this thoroughly towards the end.</li></ul></li></ul></li>\n<li>Transformer<ul>\n<li>We tried various architectures, and attaching a CNN-head or adding positional encoding proved effective.</li>\n<li>We tested large models of 512x4 and above with n=9m, but found no significant improvement over 512x4 or even a decrease in accuracy. It is unclear whether this was due to poor hyperparameter tuning or if 512x4 was simply sufficient for this dataset.</li></ul></li>\n</ul>\n<h1>2. Kami part</h1>\n<ul>\n<li><strong>Model</strong><ul>\n<li>1D Unet-based model x 11</li></ul></li>\n<li>It is important to reshape the input to (batch, 60, dim) to explicitly input height relationships into the model.</li>\n<li>Performance improves with a lot of data and models with large parameters.</li>\n</ul>\n<h2>2-1. Features Selection / Engineering</h2>\n<ul>\n<li><strong>Data</strong><ul>\n<li>Low-resolution data<ul>\n<li>Train: Data excluding validation from [February of Year 1, February of Year 9)</li>\n<li>Validation: Approximately 641,280 instances until February of Year 8 (skipping every 7 instances similar to Kaggle data)</li></ul></li></ul></li>\n<li><strong>Input Normalization</strong><ul>\n<li>Subtract the mean and divide by the standard deviation.</li>\n<li>State_t, q0001, q0002, q0003, u, v, ozone, ch4, n2o use common normalization across all heights.<ul>\n<li>Reason: E3SM often performs operations by height, so it is preferable to standardize these features.</li></ul></li>\n<li>Only q0001, q0002, q0003 undergo exponential change and are normalized as follows: multiply by 1e9, apply log1p, then normalize.<ul>\n<li>Reason: Possibly related to the Clausius-Clapeyron equation, which is somewhat utilized by the model.</li></ul></li></ul></li>\n<li><strong>Output Normalization</strong><ul>\n<li>Subtract the mean and divide by the standard deviation.</li></ul></li>\n<li><strong>Feature Engineering</strong><ul>\n<li>Relative humidity (expresses the relative amount of q1)<ul>\n<li>Reference: <a href=\"https://www.science.org/doi/10.1126/sciadv.adj7250\" target=\"_blank\"><strong>Climate-invariant machine learning</strong></a> (<a href=\"https://pog.mit.edu/src/beucler_climate_invariant_ml_supplement_2024.pdf\" target=\"_blank\">supplementary material</a>)</li>\n<li>The calculation method for saturation vapor pressure in the paper did not align with E3SM results, so Bolton's method used in E3SM was adopted.</li></ul></li>\n<li>Ice rate: q0002 / (q0002 + q0003)</li>\n<li>Cloud water: (q2 + q3)</li>\n<li>Add categorical features such as height (0~59) information and whether q0002/q0003 are zero using a 5-dimensional embedding.</li></ul></li>\n</ul>\n<h2>2-2. Models</h2>\n<ul>\n<li><strong>GPU:</strong> V100</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>#</th>\n<th>Model Name</th>\n<th>CV</th>\n<th>LB</th>\n<th>Training Time</th>\n<th>Note</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>204_diff_last_all_lr</td>\n<td>0.7768</td>\n<td>0.77351</td>\n<td>1d 4h</td>\n<td></td>\n</tr>\n<tr>\n<td>2</td>\n<td>201_unet_multi_all_n3_restart2</td>\n<td>0.7783</td>\n<td>-</td>\n<td>22h</td>\n<td></td>\n</tr>\n<tr>\n<td>3</td>\n<td>201_unet_multi_all_512_n3</td>\n<td>0.7794</td>\n<td>-</td>\n<td>1d 9h</td>\n<td></td>\n</tr>\n<tr>\n<td>4</td>\n<td>201_unet_multi_all_384_n2</td>\n<td>0.7801</td>\n<td>-</td>\n<td>22h</td>\n<td></td>\n</tr>\n<tr>\n<td>5</td>\n<td>201_unet_multi_all</td>\n<td>0.7815</td>\n<td>-</td>\n<td>1d 7h</td>\n<td></td>\n</tr>\n<tr>\n<td>6</td>\n<td>217_fix_transformer_leak_all_cos_head64</td>\n<td>0.7817</td>\n<td>-</td>\n<td>1d 7h</td>\n<td>With transformer head</td>\n</tr>\n<tr>\n<td>7</td>\n<td>217_fix_transformer_leak_all_cos_head64_n4</td>\n<td>0.7828</td>\n<td>-</td>\n<td>1d 20h</td>\n<td>With transformer head</td>\n</tr>\n<tr>\n<td>8</td>\n<td>222_wo_transformer_all</td>\n<td>0.7839</td>\n<td>-</td>\n<td>2d 21h</td>\n<td></td>\n</tr>\n<tr>\n<td>9</td>\n<td>222_wo_transformer_all_004</td>\n<td>0.7830</td>\n<td>-</td>\n<td>2d 21h</td>\n<td>Parameter: 354 M</td>\n</tr>\n<tr>\n<td>10</td>\n<td>225_smoothl1_loss_all_005</td>\n<td>0.7833</td>\n<td>-</td>\n<td>3d 7h</td>\n<td></td>\n</tr>\n<tr>\n<td>11</td>\n<td>225_smoothl1_loss_all_beta</td>\n<td>0.7828</td>\n<td>-</td>\n<td>2d 8h</td>\n<td></td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li><strong>Base Structure:</strong> height mlp → shallow 1D Unet x 2 → height mlp</li>\n<li><strong>Initial Height MLP</strong><ul>\n<li>Apply a common weight MLP to each height.<ul>\n<li>Reason: E3SM often performs operations by height &amp; features should be easy to use later in 1D Unet.</li></ul></li></ul></li>\n<li><strong>1D Unet</strong><ul>\n<li>A 1D Unet with many channels.</li>\n<li>Start with 256 dimensions, doubling the number of channels with each convolution, repeated 3 times.</li></ul></li>\n<li><strong>Output Height MLP</strong><ul>\n<li>Apply a common weight MLP to each height.</li>\n<li>Prepare MLPs for ptend_t, q0001, q0002, q0003, u, v &amp; directly input related features with skip connections.<ul>\n<li>For example, include state_t with ptend_t.</li></ul></li></ul></li>\n<li><strong>Other Scalar Predictions</strong><ul>\n<li>Predict using the bottleneck of the 1D Unet with an MLP.</li></ul></li>\n<li><strong>Final Output</strong><ul>\n<li>The outputs of height MLP for state_t, q0001, q0002, q0003, u, v are given as 60x2 each, then expressed as x1.exp() - x2.exp() to slightly improve the score.<ul>\n<li>Reason: The exponential relationship in the Clausius-Clapeyron equation and the model aims to predict the difference before and after the change.</li></ul></li></ul></li>\n<li><strong>Training Method</strong><ul>\n<li>Ignore some labels during training.<ul>\n<li>Labels with weight 0 in the sample submission.</li>\n<li>Labels to be ignored in post-processing.</li></ul></li>\n<li>Optimizer: Adan (not Adam)</li>\n<li>Scheduler: Cosine schedule with warmup or reduce LR on plateau.</li></ul></li>\n<li><strong>Prediction</strong><ul>\n<li>Use EMA for prediction.</li></ul></li>\n</ul>\n<h2>2-3. Post-Processing</h2>\n<ul>\n<li>Set some values to state * (1 / -1200)<ul>\n<li>ptend_q0002_12 ~ ptend_q0002_28</li></ul></li>\n<li>Target: all values in some labels &amp; low or high temperature in some q2, q3 labels</li>\n</ul>\n<h1>3. Takoi part</h1>\n<h2>3-1. Summary</h2>\n<ul>\n<li>Used data from Hugging Face's LEAP/ClimSim_low-res for training</li>\n<li>Created features from the data in a time series format</li>\n<li>Developed 12 models based on LSTM</li>\n</ul>\n<h2>3-2. Features Selection / Engineering</h2>\n<ul>\n<li><p>Data</p>\n<ul>\n<li>Validation: Used the last 639,744 rows from Kaggle's train.csv</li>\n<li>Train: Used Hugging Face's LEAP/ClimSim_low-res data (excluding the 9th year and periods overlapping with the validation period)</li></ul></li>\n<li><p>Features</p>\n<ul>\n<li>Divided into time series parts and others</li>\n<li>Time series parts (considered as a series of length 60)<ul>\n<li>Original data</li>\n<li>Differences from the subsequent data points in the series</li></ul></li>\n<li>Other parts<ul>\n<li>Original data</li>\n<li>Sum of state_q0001, state_q0002, and state_q0003</li></ul></li></ul></li>\n<li><p>Preprocessing</p>\n<ul>\n<li>Features<ul>\n<li>StandardScaler<ul>\n<li>Applied StandardScaler to each column</li></ul></li></ul></li>\n<li>Target<ul>\n<li>Used the value after applying weight to the target</li>\n<li>StandardScaler<ul>\n<li>Applied StandardScaler to each column</li></ul></li></ul></li></ul></li>\n</ul>\n<h2>3-3. Models</h2>\n<ul>\n<li>The models are based on LSTM</li>\n<li>The following is the base model<ul>\n<li>Ultimately, multiple models were created by enlarging the base model or adding conv1d</li></ul></li>\n</ul>\n<pre><code> (nn.Module):\n     ():\n        (LeapRnnModel, self).__init__()\n        self.numerical_linear = nn.Sequential(\n            nn.Linear(input_numerical_size,\n                      numerical_linear_size),\n            nn.LayerNorm(numerical_linear_size)\n        )\n        self.numerical_linear2_list = nn.ModuleList(\n            [nn.Sequential(\n                nn.Linear(input_numerical_size2,\n                          numerical_linear_size2),\n                nn.LayerNorm(numerical_linear_size2)\n            )  _  ()]\n        )\n        self.numerical_linear2 = nn.Sequential(\n            nn.Linear(input_numerical_size2,\n                      numerical_linear_size2),\n            nn.LayerNorm(numerical_linear_size2)\n        )\n        self.rnn = nn.LSTM(numerical_linear_size + numerical_linear_size2,\n                           model_size,\n                           num_layers=,\n                           batch_first=,\n                           bidirectional=)\n        self.linear_out1 = nn.Sequential(\n            nn.Linear(model_size * ,\n                      linear_out),\n            nn.LayerNorm(linear_out),\n            nn.ReLU(),\n            nn.Linear(linear_out,\n                      out_size1))\n        self.layernorm = nn.LayerNorm(model_size * )\n        self.linear_out2 = nn.Sequential(\n            nn.Linear(model_size *  + numerical_linear_size2,\n                      linear_out),\n            nn.LayerNorm(linear_out),\n            nn.ReLU(),\n            nn.Linear(linear_out,\n                      out_size2))\n        self._reinitialize()\n\n     ():\n        \n         name, p  self.named_parameters():\n               name:\n                   name:\n                    nn.init.xavier_uniform_(p.data)\n                   name:\n                    nn.init.orthogonal_(p.data)\n                   name:\n                    p.data.fill_()\n                    \n                    n = p.size()\n                    p.data[(n // ):(n // )].fill_()\n                   name:\n                    p.data.fill_()\n\n     ():\n\n        numerical_embedding = self.numerical_linear(seq_array)\n        other_embedding = self.numerical_linear2(other_array)\n        numerical_embedding2_list = [\n            linear(other_array)  linear  self.numerical_linear2_list]\n        numerical_embedding2 = torch.stack(numerical_embedding2_list, dim=)\n        numerical_embedding_concat = torch.cat(\n            [numerical_embedding, numerical_embedding2], dim=)\n        output_seq, _ = self.rnn(numerical_embedding_concat)\n        output_other = torch.mean(output_seq, dim=)\n        output_other = self.layernorm(output_other)\n        output_other = torch.cat([output_other, other_embedding], dim=)\n        output_seq = self.linear_out1(output_seq)\n        output_other = self.linear_out2(output_other)\n         output_seq, output_other\n</code></pre>\n<h3>3-4. Training Method</h3>\n<ul>\n<li>loss : SmoothL1Loss</li>\n<li>scheduler : get_cosine_schedule_with_warmup</li>\n<li>optimizer : AdamW</li>\n<li>lr : 1e-3</li>\n</ul>\n<h3>3-5. Post-Processing</h3>\n<ul>\n<li>Replace values of ptend_q0002_0 to ptend_q0002_27 with the corresponding values of state_q0002_0 to state_q0002_27 divided by (-1200).</li>\n<li>Set columns with a weight of 0 in sample_submission.csv to 0.</li>\n</ul>\n<h3>3-6. Model Results</h3>\n<ul>\n<li>All the models below are trained using A100</li>\n<li>CV is evaluated using the validation data<ul>\n<li>Results after post-processing</li></ul></li>\n<li>Some experiments do not use the entire dataset mentioned above, so the approximate size of the training data is noted</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>#</th>\n<th>exp no</th>\n<th>CV</th>\n<th>Training Time(h)</th>\n<th>Data Size</th>\n<th>Note</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>124</td>\n<td>0.7812</td>\n<td>24h</td>\n<td>about 45M</td>\n<td></td>\n</tr>\n<tr>\n<td>2</td>\n<td>130</td>\n<td>0.7815</td>\n<td>30h</td>\n<td>about 45M</td>\n<td>add 1dcnn</td>\n</tr>\n<tr>\n<td>3</td>\n<td>131</td>\n<td>0.7816</td>\n<td>34h</td>\n<td>about 60M</td>\n<td></td>\n</tr>\n<tr>\n<td>4</td>\n<td>133</td>\n<td>0.7819</td>\n<td>40h</td>\n<td>about 60M</td>\n<td>add 1dcnn</td>\n</tr>\n<tr>\n<td>5</td>\n<td>134</td>\n<td>0.7817</td>\n<td>36h</td>\n<td>about 60M</td>\n<td>add 1dcnn</td>\n</tr>\n<tr>\n<td>6</td>\n<td>135</td>\n<td>0.7819</td>\n<td>42h</td>\n<td>about 70M</td>\n<td></td>\n</tr>\n<tr>\n<td>7</td>\n<td>136</td>\n<td>0.7821</td>\n<td>45h</td>\n<td>about 70M</td>\n<td>add 1dcnn</td>\n</tr>\n<tr>\n<td>8</td>\n<td>138</td>\n<td>0.7824</td>\n<td>51h</td>\n<td>about 70M</td>\n<td>add 1dcnn</td>\n</tr>\n<tr>\n<td>9</td>\n<td>139</td>\n<td>0.7822</td>\n<td>59h</td>\n<td>about 70M</td>\n<td>add 1dcnn</td>\n</tr>\n<tr>\n<td>10</td>\n<td>141</td>\n<td>0.7825</td>\n<td>116h</td>\n<td>about 70M</td>\n<td>add 1dcnn + large model</td>\n</tr>\n<tr>\n<td>11</td>\n<td>159</td>\n<td>0.7827</td>\n<td>52h</td>\n<td>about 70M</td>\n<td>add 1dcnn</td>\n</tr>\n<tr>\n<td>1２</td>\n<td>162</td>\n<td>0.7838</td>\n<td>47h</td>\n<td>about 70M</td>\n<td>add 1dcnn + large model</td>\n</tr>\n</tbody>\n</table>\n<h1>4. Ensemble/Stacking</h1>\n<h2>4-1. Summary</h2>\n<p>Ensemble 30 models.</p>\n<table>\n<thead>\n<tr>\n<th>#</th>\n<th>method</th>\n<th>public</th>\n<th>private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>nelder-mead</td>\n<td>0.79100</td>\n<td>0.78713</td>\n</tr>\n<tr>\n<td>2</td>\n<td>1d-cnn stacking</td>\n<td>0.79193</td>\n<td>0.78774</td>\n</tr>\n</tbody>\n</table>\n<h2>4-2. Nelder-mead</h2>\n<ul>\n<li>The ensemble weights are determined using Nelder-Mead based on the predictions from team members.</li>\n<li>The weights are optimized for each target group (such as ptend_t or ptend_q0001).</li>\n<li>The ensemble results are used to replace the predictions from <code>4-3. 1D-CNN Stacking</code>.</li>\n</ul>\n<h2>4-3. 1D-CNN Stacking</h2>\n<p>Create simple 1d-cnn model with inputs of  <code>(batch_size, n_models, n_labels(=368))</code> and outputs of <code>(batch_size, n_labels)</code> and 10-folds ensemble.</p>\n<pre><code> (nn.Module):\n     ():\n        (Model1DCNN, self).__init__()\n\n        conv = []\n         hidden_dim, kernel_size  (hidden_dims, kernel_sizes):\n            conv.append(\n                nn.LazyConv1d(\n                    out_channels=hidden_dim,\n                    kernel_size=kernel_size,\n                    stride=,\n                    padding=,\n                )\n            )\n            conv.append(nn.BatchNorm1d(hidden_dim))\n            conv.append(nn.GELU())\n        conv.append(\n            nn.LazyConv1d(out_channels=, kernel_size=, stride=, padding=)\n        )\n        self.conv = nn.Sequential(*conv)\n\n     ():\n        \n        x = self.conv(\n            x\n        )  \n        x = x.mean(\n            dim=\n        )  \n         x\n</code></pre>\n<p>Hyperparameter is below: </p>\n<ul>\n<li>lr: 1e-3</li>\n<li>batch_size: 256</li>\n<li>hidden_size: <code>[256, 256, 256]</code></li>\n<li>kernel_size: <code>[3, 3, 3]</code></li>\n<li>epochs: 20</li>\n<li>optimizer: <code>AdamW(weight_decay=0)</code></li>\n<li>scheduler: <code>linear scheduler with warmup</code></li>\n<li>loss: <code>SmoothL1Loss(beta=1)</code></li>\n</ul>",
  "messages": [
    {
      "id": 2940260,
      "postDate": "2024-07-29T22:44:16.797Z",
      "content": "<p>First of all, we want to thank Kaggle and <a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> for hosting a competiton. And I would like to thank our teammate <a href=\"https://www.kaggle.com/takoihiraokazu\" target=\"_blank\">@takoihiraokazu</a> and <a href=\"https://www.kaggle.com/kami634\" target=\"_blank\">@kami634</a>! It was a pleasure to team up with you both.</p>\n<h1>1. Kurupical part</h1>\n<h2>1-1. Summary</h2>\n<table>\n<thead>\n<tr>\n<th>#</th>\n<th>model_name</th>\n<th>cv</th>\n<th>public</th>\n<th>private</th>\n<th>training time</th>\n<th>note</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>convnext 64x3-128x3-256x27-512x3</td>\n<td>0.7881</td>\n<td>0.78577</td>\n<td>0.78147</td>\n<td>60h with 1xRTX4090</td>\n<td></td>\n</tr>\n<tr>\n<td>2</td>\n<td>convnext 96x3-192x3-384x27-768x3</td>\n<td>0.7887</td>\n<td>0.78697</td>\n<td>0.78281</td>\n<td>120h with 1xRTX4090</td>\n<td></td>\n</tr>\n<tr>\n<td>3</td>\n<td>convnext 128x3-256x3-512x27-1024x3</td>\n<td>0.78693</td>\n<td>0.78480</td>\n<td>0.78150</td>\n<td>132h with 1xRTX4090</td>\n<td>use 81.25% data</td>\n</tr>\n<tr>\n<td>4</td>\n<td>convnext 144x3-288x3-576x27-1152x3</td>\n<td>0.78547</td>\n<td>0.78317</td>\n<td>0.77903</td>\n<td>28h with 8xRTX4090</td>\n<td></td>\n</tr>\n<tr>\n<td>5</td>\n<td>transformer 512x4</td>\n<td>0.7848</td>\n<td>0.78341</td>\n<td>0.77910</td>\n<td>54h with 1xRTX4090</td>\n<td></td>\n</tr>\n<tr>\n<td>6</td>\n<td>transformer 768x4</td>\n<td>0.7843</td>\n<td>0.78350</td>\n<td>0.77966</td>\n<td>72h with 1xRTX4090</td>\n<td></td>\n</tr>\n<tr>\n<td>7</td>\n<td>convnext 144x3-288x3-576x27-1152x3</td>\n<td>0.78769</td>\n<td>0.78622</td>\n<td>0.78137</td>\n<td>16h with 8xRTX4090</td>\n<td>training 4epochs</td>\n</tr>\n</tbody>\n</table>\n<h2>1-2. Feature Engineering / Preprocessing</h2>\n<p>I use all row-les datasets.</p>\n<h3>1-2-1. Feature Engineering</h3>\n<ul>\n<li>diff<ul>\n<li>x[i] - x[i-1]</li>\n<li>x[i] - x[i-2]</li></ul></li>\n<li>mean and diff_mean in similar features<ul>\n<li>q0002, q0003</li>\n<li>state_u, state_v</li>\n<li>pbuf_*</li></ul></li>\n</ul>\n<h3>1-2-2. Preprocessing</h3>\n<ul>\n<li>Standard Scaler for feature / label<ul>\n<li>Calculate feature mean/std with both train/test datasets.</li></ul></li>\n<li>Extremely large values can make model training unstable, so features are standardized and then clipped to the range of -100 to 100.</li>\n<li>To convert to the shape of (batch_size, n_feature, 60), scalar values (e.g., state_ps) are transformed into time series data by repeating the same value 60 times.</li>\n</ul>\n<h2>1-3. Training Methods</h2>\n<h3>1-3-1. Models</h3>\n<h4>1-3-1-1. ConvNeXt</h4>\n<p>A ConvNext with inputs of shape (batch_size, n_features, 60) and outputs of shape (batch_size, 368). The pseudocode is shown below.</p>\n<pre><code> (nn.Module):\n     ():\n        (Head1D, self).__init__()\n        self.final_layer = nn.LazyConv1d(out_channels=, kernel_size=, stride=, padding=)\n        self.fc = nn.Linear( * , )\n\n     ():\n        x = self.final_layer(x)  \n        x_out = torch.cat(\n            [\n                x[:, :, :].reshape(\n                    -,  * \n                ),  \n                self.fc(x[:, :, :].reshape(-,  * )),  \n            ],\n            dim=,\n        )  \n         x_out\n\n (nn.Module):\n     ():\n        self.head = Head1D()\n        ...\n\n     ():\n        \n        ...\n\n     ():\n        x = x[]  \n\n        \n        x_features = self.forward_feature(x)  \n\n        x_pred = self.head(x_features)  \n         x_pred\n\nmodel = ConvNeXt()\nx = torch.randn(, , )\n model(x).shape == (, )\n</code></pre>\n<ul>\n<li>Based on <a href=\"https://github.com/facebookresearch/ConvNeXt/blob/d1fa8f6fef0a165b27399986cc2bdacc92777e40/models/convnext.py\" target=\"_blank\">https://github.com/facebookresearch/ConvNeXt/blob/d1fa8f6fef0a165b27399986cc2bdacc92777e40/models/convnext.py</a>  ,I've rewritten it for 1D inputs with the following considerations:<ul>\n<li>Adjusted to ensure the shape remains unchanged between input and output by setting stride=1 and padding=\"same\".</li>\n<li>Replaced LayerNorm with BatchNorm since LayerNorm was not effective.</li></ul></li>\n</ul>\n<h4>1-3-1-2. Transformer</h4>\n<p>A Transformer with input of shape  <code>(batch_size, 60, n_features)</code> and outputs of shape <code>(batch_size, 368)</code>.  The pseudocode is shown below.</p>\n<pre><code> (nn.Module):\n     ():\n        (TransformerModel, self).__init__()\n\n        N_SEQUENTIAL_COLUMNS = \n        \n        self.fc_stem = nn.Sequential(\n            nn.LazyLinear(hidden_dims),\n            nn.LayerNorm(hidden_dims),\n            nn.GELU(),\n            nn.LazyLinear(hidden_dims),\n            nn.LayerNorm(hidden_dims),\n            nn.GELU(),\n        )\n        self.position_encoder = nn.Embedding(, hidden_dims)\n\n        \n        layer = nn.TransformerEncoderLayer(\n            d_model=hidden_dims,\n            nhead=n_heads,\n            dim_feedforward=hidden_dims * ,\n            dropout=,\n            activation=,\n            batch_first=,\n        )\n        self.transformer = nn.TransformerEncoder(layer, num_layers=n_layers)        \n\n        \n        \n        self.fc_head_list_sequential = []\n        head_cnn_params = [{: , : }] * \n         _  (N_SEQUENTIAL_COLUMNS):\n            fc_head = []\n             head_cnn_param  head_cnn_params:\n                fc_head.append(\n                    nn.LazyConv1d(\n                        stride=,\n                        padding=,\n                        **head_cnn_param,\n                    )\n                )\n                fc_head.append(nn.BatchNorm1d(head_cnn_param[]))\n                fc_head.append(nn.GELU())\n            fc_head.append(\n                nn.LazyConv1d(, kernel_size=, stride=, padding=)\n            )\n            fc_head = nn.Sequential(*fc_head)\n            self.fc_head_list_sequential.append(fc_head)\n        self.fc_head_list_sequential = nn.ModuleList(self.fc_head_list_sequential)\n\n        \n        self.fc_head_scalar = nn.Linear(hidden_dims, )\n\n     ():\n\n        \n        x = self.fc_stem(x)  \n        pe = self.position_encoder(\n            torch.arange(x.shape[], device=x.device).expand(x.shape[], -)\n        )  \n        x = x + pe\n\n        \n        x = self.transformer(x)  \n\n        \n        \n        x_out = []\n         i  ():\n            x_out_ = self.fc_head_list_sequential[i](\n                x.permute(, , )\n            )  \n            x_out_ = x_out_.squeeze()  \n            x_out.append(x_out_)        \n        \n        x_out_ = self.fc_head_scalar(x.mean(dim=))  \n        x_out.append(x_out_)\n        x_out = torch.cat(x_out, dim=)  \n         x_out        \n</code></pre>\n<h3>1-3-2. Hyper Parameter</h3>\n<table>\n<thead>\n<tr>\n<th>param_name</th>\n<th>#1</th>\n<th>#2</th>\n<th>#3</th>\n<th>#4</th>\n<th>#5</th>\n<th>#6</th>\n<th>#7</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>epochs</td>\n<td>7</td>\n<td>7</td>\n<td>7</td>\n<td>7</td>\n<td>7</td>\n<td>7</td>\n<td>4</td>\n</tr>\n<tr>\n<td>lr</td>\n<td>2.5e-3</td>\n<td>2e-3</td>\n<td>2e-3</td>\n<td>1.5e-3</td>\n<td>1e-3</td>\n<td>1e-3</td>\n<td>2e-3</td>\n</tr>\n<tr>\n<td>batch_size</td>\n<td>384</td>\n<td>384</td>\n<td>384</td>\n<td>288</td>\n<td>384</td>\n<td>384</td>\n<td>288</td>\n</tr>\n<tr>\n<td>weight_decay</td>\n<td>0.05</td>\n<td>0.05</td>\n<td>0.05</td>\n<td>0.05</td>\n<td>0.01</td>\n<td>0.01</td>\n<td>0.075</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>other<ul>\n<li>optimizer: AdamW<ul>\n<li>cnn: AdamW(weight_decay=0.05 or 0.075)</li>\n<li>transformer: AdamW(weight_decay=0.01)</li></ul></li>\n<li>loss: <ul>\n<li>cnn: SmoothL1Loss(beta=0.01)</li>\n<li>transformer: SmoothL1Loss(beta=1)</li></ul></li>\n<li>scheduler<ul>\n<li>cnn: transformers.get_polynomial_decay_schedule_with_warmup(alpha=2, warmup_ratio=0.1)</li>\n<li>transformer: transformers.get_polynomial_decay_schedule_with_warmup(alpha=1, warmup_ratio=0.1), equals to get_linear_scheduler_with_warmup</li></ul></li>\n<li>The batch size of 384 is a remnant of using lat/lon leak.</li>\n<li>Transformers tend to diverge quickly with a high learning rate, so keep it as low as possible. CNNs were not as sensitive.</li>\n<li>Regarding the beta of SmoothL1Loss, smaller models tended to perform better with something close to Huber Loss, while larger models showed better performance with something closer to L1 Loss.</li></ul></li>\n</ul>\n<h2>1-4. Postprocessing</h2>\n<ul>\n<li>Set some values to state * (1 / -1200)<ul>\n<li>ptend_q0002_12 ~ ptend_q0002_28</li></ul></li>\n</ul>\n<h2>1-5. Others</h2>\n<p>This was the competition where we needed to rely on neural networks the most among all the competitions we've participated in so far.</p>\n<ul>\n<li>Common<ul>\n<li>Proper learning rate scheduling was crucial, and we spent a significant amount of time tuning it.</li>\n<li>We tested with a small amount of data (n=1.8m, n=9m) before using the full dataset (n=70m). There were cases where what worked with a small amount of data didn't work with the full dataset.</li>\n<li>The same issue occurred between small models (ConvNeXt 32x3-64x3-128x27-256x3) and large models (ConvNeXt 96x3-192x3-384x27-768x3).</li></ul></li>\n<li>ConvNeXt<ul>\n<li>Polynomial decay scheduler and high weight decay were effective.</li>\n<li>Observing the train loss on W&amp;B, it seemed that training progressed significantly when the learning rate was low. Therefore, we adopted a polynomial decay scheduler to stay at a low learning rate for a longer period. We found that training progressed too much and caused overfitting, so we controlled it with weight decay. (cv: +0.004)</li>\n<li>It seemed that the larger the model, the higher the accuracy, given proper hyperparameter settings.<ul>\n<li>Larger models converged faster, so we reduced the number of epochs.</li>\n<li>We didn't have time to test this thoroughly towards the end.</li></ul></li></ul></li>\n<li>Transformer<ul>\n<li>We tried various architectures, and attaching a CNN-head or adding positional encoding proved effective.</li>\n<li>We tested large models of 512x4 and above with n=9m, but found no significant improvement over 512x4 or even a decrease in accuracy. It is unclear whether this was due to poor hyperparameter tuning or if 512x4 was simply sufficient for this dataset.</li></ul></li>\n</ul>\n<h1>2. Kami part</h1>\n<ul>\n<li><strong>Model</strong><ul>\n<li>1D Unet-based model x 11</li></ul></li>\n<li>It is important to reshape the input to (batch, 60, dim) to explicitly input height relationships into the model.</li>\n<li>Performance improves with a lot of data and models with large parameters.</li>\n</ul>\n<h2>2-1. Features Selection / Engineering</h2>\n<ul>\n<li><strong>Data</strong><ul>\n<li>Low-resolution data<ul>\n<li>Train: Data excluding validation from [February of Year 1, February of Year 9)</li>\n<li>Validation: Approximately 641,280 instances until February of Year 8 (skipping every 7 instances similar to Kaggle data)</li></ul></li></ul></li>\n<li><strong>Input Normalization</strong><ul>\n<li>Subtract the mean and divide by the standard deviation.</li>\n<li>State_t, q0001, q0002, q0003, u, v, ozone, ch4, n2o use common normalization across all heights.<ul>\n<li>Reason: E3SM often performs operations by height, so it is preferable to standardize these features.</li></ul></li>\n<li>Only q0001, q0002, q0003 undergo exponential change and are normalized as follows: multiply by 1e9, apply log1p, then normalize.<ul>\n<li>Reason: Possibly related to the Clausius-Clapeyron equation, which is somewhat utilized by the model.</li></ul></li></ul></li>\n<li><strong>Output Normalization</strong><ul>\n<li>Subtract the mean and divide by the standard deviation.</li></ul></li>\n<li><strong>Feature Engineering</strong><ul>\n<li>Relative humidity (expresses the relative amount of q1)<ul>\n<li>Reference: <a href=\"https://www.science.org/doi/10.1126/sciadv.adj7250\" target=\"_blank\"><strong>Climate-invariant machine learning</strong></a> (<a href=\"https://pog.mit.edu/src/beucler_climate_invariant_ml_supplement_2024.pdf\" target=\"_blank\">supplementary material</a>)</li>\n<li>The calculation method for saturation vapor pressure in the paper did not align with E3SM results, so Bolton's method used in E3SM was adopted.</li></ul></li>\n<li>Ice rate: q0002 / (q0002 + q0003)</li>\n<li>Cloud water: (q2 + q3)</li>\n<li>Add categorical features such as height (0~59) information and whether q0002/q0003 are zero using a 5-dimensional embedding.</li></ul></li>\n</ul>\n<h2>2-2. Models</h2>\n<ul>\n<li><strong>GPU:</strong> V100</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>#</th>\n<th>Model Name</th>\n<th>CV</th>\n<th>LB</th>\n<th>Training Time</th>\n<th>Note</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>204_diff_last_all_lr</td>\n<td>0.7768</td>\n<td>0.77351</td>\n<td>1d 4h</td>\n<td></td>\n</tr>\n<tr>\n<td>2</td>\n<td>201_unet_multi_all_n3_restart2</td>\n<td>0.7783</td>\n<td>-</td>\n<td>22h</td>\n<td></td>\n</tr>\n<tr>\n<td>3</td>\n<td>201_unet_multi_all_512_n3</td>\n<td>0.7794</td>\n<td>-</td>\n<td>1d 9h</td>\n<td></td>\n</tr>\n<tr>\n<td>4</td>\n<td>201_unet_multi_all_384_n2</td>\n<td>0.7801</td>\n<td>-</td>\n<td>22h</td>\n<td></td>\n</tr>\n<tr>\n<td>5</td>\n<td>201_unet_multi_all</td>\n<td>0.7815</td>\n<td>-</td>\n<td>1d 7h</td>\n<td></td>\n</tr>\n<tr>\n<td>6</td>\n<td>217_fix_transformer_leak_all_cos_head64</td>\n<td>0.7817</td>\n<td>-</td>\n<td>1d 7h</td>\n<td>With transformer head</td>\n</tr>\n<tr>\n<td>7</td>\n<td>217_fix_transformer_leak_all_cos_head64_n4</td>\n<td>0.7828</td>\n<td>-</td>\n<td>1d 20h</td>\n<td>With transformer head</td>\n</tr>\n<tr>\n<td>8</td>\n<td>222_wo_transformer_all</td>\n<td>0.7839</td>\n<td>-</td>\n<td>2d 21h</td>\n<td></td>\n</tr>\n<tr>\n<td>9</td>\n<td>222_wo_transformer_all_004</td>\n<td>0.7830</td>\n<td>-</td>\n<td>2d 21h</td>\n<td>Parameter: 354 M</td>\n</tr>\n<tr>\n<td>10</td>\n<td>225_smoothl1_loss_all_005</td>\n<td>0.7833</td>\n<td>-</td>\n<td>3d 7h</td>\n<td></td>\n</tr>\n<tr>\n<td>11</td>\n<td>225_smoothl1_loss_all_beta</td>\n<td>0.7828</td>\n<td>-</td>\n<td>2d 8h</td>\n<td></td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li><strong>Base Structure:</strong> height mlp → shallow 1D Unet x 2 → height mlp</li>\n<li><strong>Initial Height MLP</strong><ul>\n<li>Apply a common weight MLP to each height.<ul>\n<li>Reason: E3SM often performs operations by height &amp; features should be easy to use later in 1D Unet.</li></ul></li></ul></li>\n<li><strong>1D Unet</strong><ul>\n<li>A 1D Unet with many channels.</li>\n<li>Start with 256 dimensions, doubling the number of channels with each convolution, repeated 3 times.</li></ul></li>\n<li><strong>Output Height MLP</strong><ul>\n<li>Apply a common weight MLP to each height.</li>\n<li>Prepare MLPs for ptend_t, q0001, q0002, q0003, u, v &amp; directly input related features with skip connections.<ul>\n<li>For example, include state_t with ptend_t.</li></ul></li></ul></li>\n<li><strong>Other Scalar Predictions</strong><ul>\n<li>Predict using the bottleneck of the 1D Unet with an MLP.</li></ul></li>\n<li><strong>Final Output</strong><ul>\n<li>The outputs of height MLP for state_t, q0001, q0002, q0003, u, v are given as 60x2 each, then expressed as x1.exp() - x2.exp() to slightly improve the score.<ul>\n<li>Reason: The exponential relationship in the Clausius-Clapeyron equation and the model aims to predict the difference before and after the change.</li></ul></li></ul></li>\n<li><strong>Training Method</strong><ul>\n<li>Ignore some labels during training.<ul>\n<li>Labels with weight 0 in the sample submission.</li>\n<li>Labels to be ignored in post-processing.</li></ul></li>\n<li>Optimizer: Adan (not Adam)</li>\n<li>Scheduler: Cosine schedule with warmup or reduce LR on plateau.</li></ul></li>\n<li><strong>Prediction</strong><ul>\n<li>Use EMA for prediction.</li></ul></li>\n</ul>\n<h2>2-3. Post-Processing</h2>\n<ul>\n<li>Set some values to state * (1 / -1200)<ul>\n<li>ptend_q0002_12 ~ ptend_q0002_28</li></ul></li>\n<li>Target: all values in some labels &amp; low or high temperature in some q2, q3 labels</li>\n</ul>\n<h1>3. Takoi part</h1>\n<h2>3-1. Summary</h2>\n<ul>\n<li>Used data from Hugging Face's LEAP/ClimSim_low-res for training</li>\n<li>Created features from the data in a time series format</li>\n<li>Developed 12 models based on LSTM</li>\n</ul>\n<h2>3-2. Features Selection / Engineering</h2>\n<ul>\n<li><p>Data</p>\n<ul>\n<li>Validation: Used the last 639,744 rows from Kaggle's train.csv</li>\n<li>Train: Used Hugging Face's LEAP/ClimSim_low-res data (excluding the 9th year and periods overlapping with the validation period)</li></ul></li>\n<li><p>Features</p>\n<ul>\n<li>Divided into time series parts and others</li>\n<li>Time series parts (considered as a series of length 60)<ul>\n<li>Original data</li>\n<li>Differences from the subsequent data points in the series</li></ul></li>\n<li>Other parts<ul>\n<li>Original data</li>\n<li>Sum of state_q0001, state_q0002, and state_q0003</li></ul></li></ul></li>\n<li><p>Preprocessing</p>\n<ul>\n<li>Features<ul>\n<li>StandardScaler<ul>\n<li>Applied StandardScaler to each column</li></ul></li></ul></li>\n<li>Target<ul>\n<li>Used the value after applying weight to the target</li>\n<li>StandardScaler<ul>\n<li>Applied StandardScaler to each column</li></ul></li></ul></li></ul></li>\n</ul>\n<h2>3-3. Models</h2>\n<ul>\n<li>The models are based on LSTM</li>\n<li>The following is the base model<ul>\n<li>Ultimately, multiple models were created by enlarging the base model or adding conv1d</li></ul></li>\n</ul>\n<pre><code> (nn.Module):\n     ():\n        (LeapRnnModel, self).__init__()\n        self.numerical_linear = nn.Sequential(\n            nn.Linear(input_numerical_size,\n                      numerical_linear_size),\n            nn.LayerNorm(numerical_linear_size)\n        )\n        self.numerical_linear2_list = nn.ModuleList(\n            [nn.Sequential(\n                nn.Linear(input_numerical_size2,\n                          numerical_linear_size2),\n                nn.LayerNorm(numerical_linear_size2)\n            )  _  ()]\n        )\n        self.numerical_linear2 = nn.Sequential(\n            nn.Linear(input_numerical_size2,\n                      numerical_linear_size2),\n            nn.LayerNorm(numerical_linear_size2)\n        )\n        self.rnn = nn.LSTM(numerical_linear_size + numerical_linear_size2,\n                           model_size,\n                           num_layers=,\n                           batch_first=,\n                           bidirectional=)\n        self.linear_out1 = nn.Sequential(\n            nn.Linear(model_size * ,\n                      linear_out),\n            nn.LayerNorm(linear_out),\n            nn.ReLU(),\n            nn.Linear(linear_out,\n                      out_size1))\n        self.layernorm = nn.LayerNorm(model_size * )\n        self.linear_out2 = nn.Sequential(\n            nn.Linear(model_size *  + numerical_linear_size2,\n                      linear_out),\n            nn.LayerNorm(linear_out),\n            nn.ReLU(),\n            nn.Linear(linear_out,\n                      out_size2))\n        self._reinitialize()\n\n     ():\n        \n         name, p  self.named_parameters():\n               name:\n                   name:\n                    nn.init.xavier_uniform_(p.data)\n                   name:\n                    nn.init.orthogonal_(p.data)\n                   name:\n                    p.data.fill_()\n                    \n                    n = p.size()\n                    p.data[(n // ):(n // )].fill_()\n                   name:\n                    p.data.fill_()\n\n     ():\n\n        numerical_embedding = self.numerical_linear(seq_array)\n        other_embedding = self.numerical_linear2(other_array)\n        numerical_embedding2_list = [\n            linear(other_array)  linear  self.numerical_linear2_list]\n        numerical_embedding2 = torch.stack(numerical_embedding2_list, dim=)\n        numerical_embedding_concat = torch.cat(\n            [numerical_embedding, numerical_embedding2], dim=)\n        output_seq, _ = self.rnn(numerical_embedding_concat)\n        output_other = torch.mean(output_seq, dim=)\n        output_other = self.layernorm(output_other)\n        output_other = torch.cat([output_other, other_embedding], dim=)\n        output_seq = self.linear_out1(output_seq)\n        output_other = self.linear_out2(output_other)\n         output_seq, output_other\n</code></pre>\n<h3>3-4. Training Method</h3>\n<ul>\n<li>loss : SmoothL1Loss</li>\n<li>scheduler : get_cosine_schedule_with_warmup</li>\n<li>optimizer : AdamW</li>\n<li>lr : 1e-3</li>\n</ul>\n<h3>3-5. Post-Processing</h3>\n<ul>\n<li>Replace values of ptend_q0002_0 to ptend_q0002_27 with the corresponding values of state_q0002_0 to state_q0002_27 divided by (-1200).</li>\n<li>Set columns with a weight of 0 in sample_submission.csv to 0.</li>\n</ul>\n<h3>3-6. Model Results</h3>\n<ul>\n<li>All the models below are trained using A100</li>\n<li>CV is evaluated using the validation data<ul>\n<li>Results after post-processing</li></ul></li>\n<li>Some experiments do not use the entire dataset mentioned above, so the approximate size of the training data is noted</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>#</th>\n<th>exp no</th>\n<th>CV</th>\n<th>Training Time(h)</th>\n<th>Data Size</th>\n<th>Note</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>124</td>\n<td>0.7812</td>\n<td>24h</td>\n<td>about 45M</td>\n<td></td>\n</tr>\n<tr>\n<td>2</td>\n<td>130</td>\n<td>0.7815</td>\n<td>30h</td>\n<td>about 45M</td>\n<td>add 1dcnn</td>\n</tr>\n<tr>\n<td>3</td>\n<td>131</td>\n<td>0.7816</td>\n<td>34h</td>\n<td>about 60M</td>\n<td></td>\n</tr>\n<tr>\n<td>4</td>\n<td>133</td>\n<td>0.7819</td>\n<td>40h</td>\n<td>about 60M</td>\n<td>add 1dcnn</td>\n</tr>\n<tr>\n<td>5</td>\n<td>134</td>\n<td>0.7817</td>\n<td>36h</td>\n<td>about 60M</td>\n<td>add 1dcnn</td>\n</tr>\n<tr>\n<td>6</td>\n<td>135</td>\n<td>0.7819</td>\n<td>42h</td>\n<td>about 70M</td>\n<td></td>\n</tr>\n<tr>\n<td>7</td>\n<td>136</td>\n<td>0.7821</td>\n<td>45h</td>\n<td>about 70M</td>\n<td>add 1dcnn</td>\n</tr>\n<tr>\n<td>8</td>\n<td>138</td>\n<td>0.7824</td>\n<td>51h</td>\n<td>about 70M</td>\n<td>add 1dcnn</td>\n</tr>\n<tr>\n<td>9</td>\n<td>139</td>\n<td>0.7822</td>\n<td>59h</td>\n<td>about 70M</td>\n<td>add 1dcnn</td>\n</tr>\n<tr>\n<td>10</td>\n<td>141</td>\n<td>0.7825</td>\n<td>116h</td>\n<td>about 70M</td>\n<td>add 1dcnn + large model</td>\n</tr>\n<tr>\n<td>11</td>\n<td>159</td>\n<td>0.7827</td>\n<td>52h</td>\n<td>about 70M</td>\n<td>add 1dcnn</td>\n</tr>\n<tr>\n<td>1２</td>\n<td>162</td>\n<td>0.7838</td>\n<td>47h</td>\n<td>about 70M</td>\n<td>add 1dcnn + large model</td>\n</tr>\n</tbody>\n</table>\n<h1>4. Ensemble/Stacking</h1>\n<h2>4-1. Summary</h2>\n<p>Ensemble 30 models.</p>\n<table>\n<thead>\n<tr>\n<th>#</th>\n<th>method</th>\n<th>public</th>\n<th>private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>nelder-mead</td>\n<td>0.79100</td>\n<td>0.78713</td>\n</tr>\n<tr>\n<td>2</td>\n<td>1d-cnn stacking</td>\n<td>0.79193</td>\n<td>0.78774</td>\n</tr>\n</tbody>\n</table>\n<h2>4-2. Nelder-mead</h2>\n<ul>\n<li>The ensemble weights are determined using Nelder-Mead based on the predictions from team members.</li>\n<li>The weights are optimized for each target group (such as ptend_t or ptend_q0001).</li>\n<li>The ensemble results are used to replace the predictions from <code>4-3. 1D-CNN Stacking</code>.</li>\n</ul>\n<h2>4-3. 1D-CNN Stacking</h2>\n<p>Create simple 1d-cnn model with inputs of  <code>(batch_size, n_models, n_labels(=368))</code> and outputs of <code>(batch_size, n_labels)</code> and 10-folds ensemble.</p>\n<pre><code> (nn.Module):\n     ():\n        (Model1DCNN, self).__init__()\n\n        conv = []\n         hidden_dim, kernel_size  (hidden_dims, kernel_sizes):\n            conv.append(\n                nn.LazyConv1d(\n                    out_channels=hidden_dim,\n                    kernel_size=kernel_size,\n                    stride=,\n                    padding=,\n                )\n            )\n            conv.append(nn.BatchNorm1d(hidden_dim))\n            conv.append(nn.GELU())\n        conv.append(\n            nn.LazyConv1d(out_channels=, kernel_size=, stride=, padding=)\n        )\n        self.conv = nn.Sequential(*conv)\n\n     ():\n        \n        x = self.conv(\n            x\n        )  \n        x = x.mean(\n            dim=\n        )  \n         x\n</code></pre>\n<p>Hyperparameter is below: </p>\n<ul>\n<li>lr: 1e-3</li>\n<li>batch_size: 256</li>\n<li>hidden_size: <code>[256, 256, 256]</code></li>\n<li>kernel_size: <code>[3, 3, 3]</code></li>\n<li>epochs: 20</li>\n<li>optimizer: <code>AdamW(weight_decay=0)</code></li>\n<li>scheduler: <code>linear scheduler with warmup</code></li>\n<li>loss: <code>SmoothL1Loss(beta=1)</code></li>\n</ul>",
      "rawMarkdown": "First of all, we want to thank Kaggle and @jerrylin96 for hosting a competiton. And I would like to thank our teammate @takoihiraokazu and @kami634! It was a pleasure to team up with you both.\n\n# 1. Kurupical part\n\n## 1-1. Summary\n\n| # | model_name | cv | public | private | training time | note |\n| ---- | ---- | ---- | ---- | ---- | ---- | ---- |\n| 1 | convnext 64x3-128x3-256x27-512x3 | 0.7881 | 0.78577 | 0.78147 |60h with 1xRTX4090 | |\n| 2 | convnext 96x3-192x3-384x27-768x3 | 0.7887 | 0.78697 | 0.78281 | 120h with 1xRTX4090 | |\n| 3 | convnext 128x3-256x3-512x27-1024x3 | 0.78693 | 0.78480 | 0.78150 | 132h with 1xRTX4090 | use 81.25% data |\n| 4 | convnext 144x3-288x3-576x27-1152x3 | 0.78547 | 0.78317 | 0.77903 | 28h with 8xRTX4090 |  |\n| 5 | transformer 512x4 | 0.7848 | 0.78341 | 0.77910 | 54h with 1xRTX4090 |  |\n| 6 | transformer 768x4 | 0.7843 | 0.78350 | 0.77966 | 72h with 1xRTX4090 |  |\n| 7 | convnext 144x3-288x3-576x27-1152x3 | 0.78769 | 0.78622 | 0.78137 | 16h with 8xRTX4090 | training 4epochs |\n\n## 1-2. Feature Engineering / Preprocessing\nI use all row-les datasets.\n\n### 1-2-1. Feature Engineering\n\n- diff\n  - x[i] - x[i-1]\n  - x[i] - x[i-2]\n- mean and diff_mean in similar features\n  - q0002, q0003\n  - state_u, state_v\n  - pbuf_*\n\n### 1-2-2. Preprocessing\n- Standard Scaler for feature / label\n  - Calculate feature mean/std with both train/test datasets.\n- Extremely large values can make model training unstable, so features are standardized and then clipped to the range of -100 to 100.\n- To convert to the shape of (batch_size, n_feature, 60), scalar values (e.g., state_ps) are transformed into time series data by repeating the same value 60 times.\n\n## 1-3. Training Methods\n### 1-3-1. Models\n#### 1-3-1-1. ConvNeXt\n\nA ConvNext with inputs of shape (batch_size, n_features, 60) and outputs of shape (batch_size, 368). The pseudocode is shown below.\n\n```python\nclass Head1D(nn.Module):\n    def __init__(self):\n        super(Head1D, self).__init__()\n        self.final_layer = nn.LazyConv1d(out_channels=14, kernel_size=1, stride=1, padding=\"same\")\n        self.fc = nn.Linear(8 * 60, 8)\n\n    def forward(self, x):\n        x = self.final_layer(x)  # (hidden_size, 60) -> (14, 60)\n        x_out = torch.cat(\n            [\n                x[:, :6, :].reshape(\n                    -1, 6 * 60\n                ),  # shape = (bs, 360) ptend_t, ptend_q0001, ptend_q0002, ptend_q0003, ptend_u, ptend_v\n                self.fc(x[:, 6:, :].reshape(-1, 8 * 60)),  # shape = (bs, 8)\n            ],\n            dim=1,\n        )  # shape = (bs, 360 + 8)\n        return x_out\n\nclass ConvNeXt(nn.Module):\n    def __init__(self):\n        self.head = Head1D()\n        ...\n\n    def forward_feature(x):\n        # convnext \n        ...\n        \n    def forward(self, x):\n        x = x[\"feature\"]  # shape = (bs, 384, n_features)\n           \n        # convnext part\n        x_features = self.forward_feature(x)  # shape = (bs, n_features, 60) -> (bs, hidden_dims, 60)\n        \n        x_pred = self.head(x_features)  # shape = (bs, hidden_dims, 60) -> (bs, 368)\n        return x_pred\n    \nmodel = ConvNeXt()\nx = torch.randn(4, 25, 60)\nassert model(x).shape == (4, 368)\n\n```\n\n- Based on https://github.com/facebookresearch/ConvNeXt/blob/d1fa8f6fef0a165b27399986cc2bdacc92777e40/models/convnext.py  ,I've rewritten it for 1D inputs with the following considerations:\n  - Adjusted to ensure the shape remains unchanged between input and output by setting stride=1 and padding=\"same\".\n  - Replaced LayerNorm with BatchNorm since LayerNorm was not effective.\n\n#### 1-3-1-2. Transformer\n\nA Transformer with input of shape  ``(batch_size, 60, n_features)`` and outputs of shape ``(batch_size, 368)``.  The pseudocode is shown below.\n\n```python\nclass TransformerModel(nn.Module):\n    def __init__(\n        self,\n        hidden_dims: int,\n        n_layers: int,\n        n_heads: int,\n        head_mode: str,\n        dropout: int = 0,\n    ):\n        super(TransformerModel, self).__init__()\n        \n        N_SEQUENTIAL_COLUMNS = 6\n        # stem\n        self.fc_stem = nn.Sequential(\n            nn.LazyLinear(hidden_dims),\n            nn.LayerNorm(hidden_dims),\n            nn.GELU(),\n            nn.LazyLinear(hidden_dims),\n            nn.LayerNorm(hidden_dims),\n            nn.GELU(),\n        )\n        self.position_encoder = nn.Embedding(60, hidden_dims)\n        \n        # transformer\n        layer = nn.TransformerEncoderLayer(\n            d_model=hidden_dims,\n            nhead=n_heads,\n            dim_feedforward=hidden_dims * 4,\n            dropout=0,\n            activation=\"gelu\",\n            batch_first=True,\n        )\n        self.transformer = nn.TransformerEncoder(layer, num_layers=n_layers)        \n        \n        # head\n        ## 1. cnn (for sequential (60x6))\n        self.fc_head_list_sequential = []\n        head_cnn_params = [{\"out_channels\": 256, \"kernel_size\": 3}] * 3\n        for _ in range(N_SEQUENTIAL_COLUMNS):\n            fc_head = []\n            for head_cnn_param in head_cnn_params:\n                fc_head.append(\n                    nn.LazyConv1d(\n                        stride=1,\n                        padding=\"same\",\n                        **head_cnn_param,\n                    )\n                )\n                fc_head.append(nn.BatchNorm1d(head_cnn_param[\"out_channels\"]))\n                fc_head.append(nn.GELU())\n            fc_head.append(\n                nn.LazyConv1d(1, kernel_size=3, stride=1, padding=\"same\")\n            )\n            fc_head = nn.Sequential(*fc_head)\n            self.fc_head_list_sequential.append(fc_head)\n        self.fc_head_list_sequential = nn.ModuleList(self.fc_head_list_sequential)\n        \n        ## 2. linear (for scalar (8))\n        self.fc_head_scalar = nn.Linear(hidden_dims, 8)\n        \n    def forward(self, x):\n\n        # stem\n        x = self.fc_stem(x)  # (bs, 60, n_features) -> (bs, 60, hidden_dims)\n        pe = self.position_encoder(\n            torch.arange(x.shape[1], device=x.device).expand(x.shape[0], -1)\n        )  # (60) -> (bs, 60, hidden_dims)\n        x = x + pe\n        \n        # transformer\n        x = self.transformer(x)  # (bs, 60, hidden_dims) -> (bs, 60, hidden_dims)\n        \n        # head\n        ## 1. cnn\n        x_out = []\n        for i in range(6):\n            x_out_ = self.fc_head_list_sequential[i](\n                x.permute(0, 2, 1)\n            )  # (bs, hidden_dims, 60) -> (bs, 1, 60)\n            x_out_ = x_out_.squeeze(1)  # (bs, 1, 60) -> (bs, 60)\n            x_out.append(x_out_)        \n        ## 2. linear\n        x_out_ = self.fc_head_scalar(x.mean(dim=1))  # (bs, hidden_dims) -> (bs, 8)\n        x_out.append(x_out_)\n        x_out = torch.cat(x_out, dim=1)  # (bs, 60*6+8)\n        return x_out        \n```\n\n\n### 1-3-2. Hyper Parameter\n| param_name | #1 | #2 | #3 | #4 | #5 | #6 | #7 | \n| --- | --- | --- | --- | --- | --- | --- | --- |\n| epochs | 7 | 7 | 7 | 7 | 7 | 7 | 4 |\n| lr | 2.5e-3 | 2e-3 | 2e-3 | 1.5e-3 | 1e-3 | 1e-3 | 2e-3 |\n| batch_size | 384 | 384 | 384 | 288 | 384 | 384 | 288 |\n| weight_decay | 0.05 | 0.05 | 0.05 | 0.05 | 0.01 | 0.01 | 0.075 |\n\n- other\n    - optimizer: AdamW\n        - cnn: AdamW(weight_decay=0.05 or 0.075)\n        - transformer: AdamW(weight_decay=0.01)\n    - loss: \n        - cnn: SmoothL1Loss(beta=0.01)\n        - transformer: SmoothL1Loss(beta=1)\n    - scheduler\n        - cnn: transformers.get_polynomial_decay_schedule_with_warmup(alpha=2, warmup_ratio=0.1)\n        - transformer: transformers.get_polynomial_decay_schedule_with_warmup(alpha=1, warmup_ratio=0.1), equals to get_linear_scheduler_with_warmup\n    - The batch size of 384 is a remnant of using lat/lon leak.\n    - Transformers tend to diverge quickly with a high learning rate, so keep it as low as possible. CNNs were not as sensitive.\n    - Regarding the beta of SmoothL1Loss, smaller models tended to perform better with something close to Huber Loss, while larger models showed better performance with something closer to L1 Loss.\n\n## 1-4. Postprocessing\n- Set some values to state * (1 / -1200)\n    - ptend_q0002_12 ~ ptend_q0002_28\n\n## 1-5. Others\n\nThis was the competition where we needed to rely on neural networks the most among all the competitions we've participated in so far.\n\n- Common\n    - Proper learning rate scheduling was crucial, and we spent a significant amount of time tuning it.\n    - We tested with a small amount of data (n=1.8m, n=9m) before using the full dataset (n=70m). There were cases where what worked with a small amount of data didn't work with the full dataset.\n    - The same issue occurred between small models (ConvNeXt 32x3-64x3-128x27-256x3) and large models (ConvNeXt 96x3-192x3-384x27-768x3).\n- ConvNeXt\n    - Polynomial decay scheduler and high weight decay were effective.\n    - Observing the train loss on W&B, it seemed that training progressed significantly when the learning rate was low. Therefore, we adopted a polynomial decay scheduler to stay at a low learning rate for a longer period. We found that training progressed too much and caused overfitting, so we controlled it with weight decay. (cv: +0.004)\n    - It seemed that the larger the model, the higher the accuracy, given proper hyperparameter settings.\n        - Larger models converged faster, so we reduced the number of epochs.\n        - We didn't have time to test this thoroughly towards the end.\n- Transformer\n    - We tried various architectures, and attaching a CNN-head or adding positional encoding proved effective.\n    - We tested large models of 512x4 and above with n=9m, but found no significant improvement over 512x4 or even a decrease in accuracy. It is unclear whether this was due to poor hyperparameter tuning or if 512x4 was simply sufficient for this dataset.\n\n\n\n# 2. Kami part\n- **Model**\n    - 1D Unet-based model x 11\n- It is important to reshape the input to (batch, 60, dim) to explicitly input height relationships into the model.\n- Performance improves with a lot of data and models with large parameters.\n\n## 2-1. Features Selection / Engineering\n\n- **Data**\n    - Low-resolution data\n        - Train: Data excluding validation from [February of Year 1, February of Year 9)\n        - Validation: Approximately 641,280 instances until February of Year 8 (skipping every 7 instances similar to Kaggle data)\n- **Input Normalization**\n    - Subtract the mean and divide by the standard deviation.\n    - State_t, q0001, q0002, q0003, u, v, ozone, ch4, n2o use common normalization across all heights.\n        - Reason: E3SM often performs operations by height, so it is preferable to standardize these features.\n    - Only q0001, q0002, q0003 undergo exponential change and are normalized as follows: multiply by 1e9, apply log1p, then normalize.\n        - Reason: Possibly related to the Clausius-Clapeyron equation, which is somewhat utilized by the model.\n- **Output Normalization**\n    - Subtract the mean and divide by the standard deviation.\n- **Feature Engineering**\n    - Relative humidity (expresses the relative amount of q1)\n        - Reference: [**Climate-invariant machine learning**](https://www.science.org/doi/10.1126/sciadv.adj7250) ([supplementary material](https://pog.mit.edu/src/beucler_climate_invariant_ml_supplement_2024.pdf))\n        - The calculation method for saturation vapor pressure in the paper did not align with E3SM results, so Bolton's method used in E3SM was adopted.\n    - Ice rate: q0002 / (q0002 + q0003)\n    - Cloud water: (q2 + q3)\n    - Add categorical features such as height (0~59) information and whether q0002/q0003 are zero using a 5-dimensional embedding.\n\n## 2-2. Models\n\n- **GPU:** V100\n\n| # | Model Name | CV | LB | Training Time | Note |\n| --- | --- | --- | --- | --- | --- |\n| 1 | 204_diff_last_all_lr | 0.7768 | 0.77351 | 1d 4h | |\n| 2 | 201_unet_multi_all_n3_restart2 | 0.7783 | - | 22h | |\n| 3 | 201_unet_multi_all_512_n3 | 0.7794 | - | 1d 9h | |\n| 4 | 201_unet_multi_all_384_n2 | 0.7801 | - | 22h | |\n| 5 | 201_unet_multi_all | 0.7815 | - | 1d 7h | |\n| 6 | 217_fix_transformer_leak_all_cos_head64 | 0.7817 | - | 1d 7h | With transformer head |\n| 7 | 217_fix_transformer_leak_all_cos_head64_n4 | 0.7828 | - | 1d 20h | With transformer head |\n| 8 | 222_wo_transformer_all | 0.7839 | - | 2d 21h | |\n| 9 | 222_wo_transformer_all_004 | 0.7830 | - | 2d 21h | Parameter: 354 M |\n| 10 | 225_smoothl1_loss_all_005 | 0.7833 | - | 3d 7h | |\n| 11 | 225_smoothl1_loss_all_beta | 0.7828 | - | 2d 8h | |\n\n- **Base Structure:** height mlp → shallow 1D Unet x 2 → height mlp\n- **Initial Height MLP**\n    - Apply a common weight MLP to each height.\n        - Reason: E3SM often performs operations by height & features should be easy to use later in 1D Unet.\n- **1D Unet**\n    - A 1D Unet with many channels.\n    - Start with 256 dimensions, doubling the number of channels with each convolution, repeated 3 times.\n- **Output Height MLP**\n    - Apply a common weight MLP to each height.\n    - Prepare MLPs for ptend_t, q0001, q0002, q0003, u, v & directly input related features with skip connections.\n        - For example, include state_t with ptend_t.\n- **Other Scalar Predictions**\n    - Predict using the bottleneck of the 1D Unet with an MLP.\n- **Final Output**\n    - The outputs of height MLP for state_t, q0001, q0002, q0003, u, v are given as 60x2 each, then expressed as x1.exp() - x2.exp() to slightly improve the score.\n        - Reason: The exponential relationship in the Clausius-Clapeyron equation and the model aims to predict the difference before and after the change.\n- **Training Method**\n    - Ignore some labels during training.\n        - Labels with weight 0 in the sample submission.\n        - Labels to be ignored in post-processing.\n    - Optimizer: Adan (not Adam)\n    - Scheduler: Cosine schedule with warmup or reduce LR on plateau.\n- **Prediction**\n    - Use EMA for prediction.\n\n## 2-3. Post-Processing\n\n- Set some values to state * (1 / -1200)\n    - ptend_q0002_12 ~ ptend_q0002_28\n- Target: all values in some labels & low or high temperature in some q2, q3 labels\n\n\n\n# 3. Takoi part\n## 3-1. Summary\n- Used data from Hugging Face's LEAP/ClimSim_low-res for training\n- Created features from the data in a time series format\n- Developed 12 models based on LSTM\n\n## 3-2. Features Selection / Engineering\n- Data\n    - Validation: Used the last 639,744 rows from Kaggle's train.csv\n    - Train: Used Hugging Face's LEAP/ClimSim_low-res data (excluding the 9th year and periods overlapping with the validation period)\n\n- Features\n    - Divided into time series parts and others\n    - Time series parts (considered as a series of length 60)\n        - Original data\n        - Differences from the subsequent data points in the series\n    - Other parts\n        - Original data\n        - Sum of state_q0001, state_q0002, and state_q0003\n\n- Preprocessing\n    - Features\n        - StandardScaler\n            - Applied StandardScaler to each column\n    - Target\n        - Used the value after applying weight to the target\n        - StandardScaler\n            - Applied StandardScaler to each column\n\n## 3-3. Models\n- The models are based on LSTM\n- The following is the base model\n    - Ultimately, multiple models were created by enlarging the base model or adding conv1d\n\n\n```python\nclass LeapRnnModel(nn.Module):\n    def __init__(\n            self,\n            input_numerical_size=9 * 2,\n            numerical_linear_size=64,\n            input_numerical_size2=17,\n            numerical_linear_size2=64,\n            model_size=256 * 2,\n            linear_out=256,\n            out_size1=6,\n            out_size2=8):\n        super(LeapRnnModel, self).__init__()\n        self.numerical_linear = nn.Sequential(\n            nn.Linear(input_numerical_size,\n                      numerical_linear_size),\n            nn.LayerNorm(numerical_linear_size)\n        )\n        self.numerical_linear2_list = nn.ModuleList(\n            [nn.Sequential(\n                nn.Linear(input_numerical_size2,\n                          numerical_linear_size2),\n                nn.LayerNorm(numerical_linear_size2)\n            ) for _ in range(60)]\n        )\n        self.numerical_linear2 = nn.Sequential(\n            nn.Linear(input_numerical_size2,\n                      numerical_linear_size2),\n            nn.LayerNorm(numerical_linear_size2)\n        )\n        self.rnn = nn.LSTM(numerical_linear_size + numerical_linear_size2,\n                           model_size,\n                           num_layers=3,\n                           batch_first=True,\n                           bidirectional=True)\n        self.linear_out1 = nn.Sequential(\n            nn.Linear(model_size * 2,\n                      linear_out),\n            nn.LayerNorm(linear_out),\n            nn.ReLU(),\n            nn.Linear(linear_out,\n                      out_size1))\n        self.layernorm = nn.LayerNorm(model_size * 2)\n        self.linear_out2 = nn.Sequential(\n            nn.Linear(model_size * 2 + numerical_linear_size2,\n                      linear_out),\n            nn.LayerNorm(linear_out),\n            nn.ReLU(),\n            nn.Linear(linear_out,\n                      out_size2))\n        self._reinitialize()\n\n    def _reinitialize(self):\n        \"\"\"\n        Tensorflow/Keras-like initialization\n        \"\"\"\n        for name, p in self.named_parameters():\n            if 'rnn' in name:\n                if 'weight_ih' in name:\n                    nn.init.xavier_uniform_(p.data)\n                elif 'weight_hh' in name:\n                    nn.init.orthogonal_(p.data)\n                elif 'bias_ih' in name:\n                    p.data.fill_(0)\n                    # Set forget-gate bias to 1\n                    n = p.size(0)\n                    p.data[(n // 4):(n // 2)].fill_(1)\n                elif 'bias_hh' in name:\n                    p.data.fill_(0)\n\n    def forward(self, seq_array,\n                other_array):\n\n        numerical_embedding = self.numerical_linear(seq_array)\n        other_embedding = self.numerical_linear2(other_array)\n        numerical_embedding2_list = [\n            linear(other_array) for linear in self.numerical_linear2_list]\n        numerical_embedding2 = torch.stack(numerical_embedding2_list, dim=1)\n        numerical_embedding_concat = torch.cat(\n            [numerical_embedding, numerical_embedding2], dim=2)\n        output_seq, _ = self.rnn(numerical_embedding_concat)\n        output_other = torch.mean(output_seq, dim=1)\n        output_other = self.layernorm(output_other)\n        output_other = torch.cat([output_other, other_embedding], dim=1)\n        output_seq = self.linear_out1(output_seq)\n        output_other = self.linear_out2(output_other)\n        return output_seq, output_other\n```\n\n\n### 3-4. Training Method\n- loss : SmoothL1Loss\n- scheduler : get_cosine_schedule_with_warmup\n- optimizer : AdamW\n- lr : 1e-3\n\n### 3-5. Post-Processing\n- Replace values of ptend_q0002_0 to ptend_q0002_27 with the corresponding values of state_q0002_0 to state_q0002_27 divided by (-1200).\n- Set columns with a weight of 0 in sample_submission.csv to 0.\n\n### 3-6. Model Results\n- All the models below are trained using A100\n- CV is evaluated using the validation data\n    - Results after post-processing\n- Some experiments do not use the entire dataset mentioned above, so the approximate size of the training data is noted\n\n| # | exp no | CV | Training Time(h) | Data Size | Note |\n| --- | --- | --- | --- | --- |--- |\n| 1 | 124| 0.7812|  24h | about 45M ||\n| 2 | 130 |0.7815 |  30h | about 45M |add 1dcnn|\n| 3 | 131 |0.7816  | 34h | about 60M ||\n| 4 | 133 |0.7819  | 40h | about 60M| add 1dcnn|\n| 5 | 134 |0.7817  | 36h | about 60M  |add 1dcnn|\n| 6 | 135 |0.7819  | 42h | about 70M ||\n| 7 | 136 |0.7821  | 45h | about 70M |add 1dcnn |\n| 8 | 138 |0.7824  | 51h | about 70M |add 1dcnn|\n| 9 | 139 | 0.7822 | 59h | about 70M |add 1dcnn |\n| 10 | 141 |0.7825  | 116h | about 70M |add 1dcnn + large model|\n| 11 | 159 |0.7827  | 52h | about 70M |add 1dcnn|\n| 1２ | 162 |0.7838  | 47h |about 70M  |add 1dcnn + large model|\n\n\n\n# 4. Ensemble/Stacking\n## 4-1. Summary\nEnsemble 30 models.\n\n| # | method | public | private |\n| --- | --- | --- | --- |\n| 1 | nelder-mead | 0.79100 | 0.78713 |\n| 2 | 1d-cnn stacking | 0.79193 | 0.78774 |\n\n## 4-2. Nelder-mead\n- The ensemble weights are determined using Nelder-Mead based on the predictions from team members.\n- The weights are optimized for each target group (such as ptend_t or ptend_q0001).\n- The ensemble results are used to replace the predictions from ``4-3. 1D-CNN Stacking``.\n\n## 4-3. 1D-CNN Stacking\n\nCreate simple 1d-cnn model with inputs of  ``(batch_size, n_models, n_labels(=368))`` and outputs of ``(batch_size, n_labels)`` and 10-folds ensemble.\n\n```python\n\nclass Model1DCNN(nn.Module):\n    def __init__(self, hidden_dims, kernel_sizes):\n        super(Model1DCNN, self).__init__()\n\n        conv = []\n        for hidden_dim, kernel_size in zip(hidden_dims, kernel_sizes):\n            conv.append(\n                nn.LazyConv1d(\n                    out_channels=hidden_dim,\n                    kernel_size=kernel_size,\n                    stride=1,\n                    padding=\"same\",\n                )\n            )\n            conv.append(nn.BatchNorm1d(hidden_dim))\n            conv.append(nn.GELU())\n        conv.append(\n            nn.LazyConv1d(out_channels=1, kernel_size=2, stride=1, padding=\"same\")\n        )\n        self.conv = nn.Sequential(*conv)\n\n    def forward(self, x):\n        # x.shape = (batch_size, n_models, seq_len=368)\n        x = self.conv(\n            x\n        )  # (batch_size, n_models, seq_len=368) -> (batch_size, hidden_size, seq_len=368)\n        x = x.mean(\n            dim=1\n        )  # (batch_size, hidden_size, seq_len=368) -> (batch_size, seq_len=368)\n        return x\n```\n\nHyperparameter is below: \n- lr: 1e-3\n- batch_size: 256\n- hidden_size: ``[256, 256, 256]``\n- kernel_size: ``[3, 3, 3]``\n- epochs: 20\n- optimizer: ``AdamW(weight_decay=0)``\n- scheduler: ``linear scheduler with warmup``\n- loss: ``SmoothL1Loss(beta=1)``\n\n",
      "votes": 34
    },
    {
      "id": 2940740,
      "postDate": "2024-07-30T12:09:22.883Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/kurupical\" target=\"_blank\">@kurupical</a>,</p>\n<p>First of all, congratulations on your impressive 4th place finish in the LEAP - Atmospheric Physics using AI (ClimSim) competition! Your detailed breakdown of the models, feature engineering, and training methods is incredibly insightful. I especially appreciated your approach to using ConvNeXt and transformers for handling the atmospheric data.</p>\n<p>I had a question about your feature engineering process. Specifically, you mentioned standardizing features and then clipping to the range of -100 to 100. Could you elaborate on the impact this had on your model's performance? Did you notice any significant improvements or stability in training after implementing this step?</p>\n<p>Thank you for sharing your solution and for any insights you can provide.</p>",
      "rawMarkdown": "Hi @kurupical,\n\nFirst of all, congratulations on your impressive 4th place finish in the LEAP - Atmospheric Physics using AI (ClimSim) competition! Your detailed breakdown of the models, feature engineering, and training methods is incredibly insightful. I especially appreciated your approach to using ConvNeXt and transformers for handling the atmospheric data.\n\nI had a question about your feature engineering process. Specifically, you mentioned standardizing features and then clipping to the range of -100 to 100. Could you elaborate on the impact this had on your model's performance? Did you notice any significant improvements or stability in training after implementing this step?\n\nThank you for sharing your solution and for any insights you can provide.",
      "votes": -1
    }
  ],
  "comments": [
    {
      "id": 2940740,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-07-30T12:09:22.883000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/kurupical\" target=\"_blank\">@kurupical</a>,</p>\n<p>First of all, congratulations on your impressive 4th place finish in the LEAP - Atmospheric Physics using AI (ClimSim) competition! Your detailed breakdown of the models, feature engineering, and training methods is incredibly insightful. I especially appreciated your approach to using ConvNeXt and transformers for handling the atmospheric data.</p>\n<p>I had a question about your feature engineering process. Specifically, you mentioned standardizing features and then clipping to the range of -100 to 100. Could you elaborate on the impact this had on your model's performance? Did you notice any significant improvements or stability in training after implementing this step?</p>\n<p>Thank you for sharing your solution and for any insights you can provide.</p>",
      "votes": -1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2940260": "First of all, we want to thank Kaggle and @jerrylin96 for hosting a competiton. And I would like to thank our teammate @takoihiraokazu and @kami634! It was a pleasure to team up with you both.\n\n# 1. Kurupical part\n\n## 1-1. Summary\n\n| # | model_name | cv | public | private | training time | note |\n| ---- | ---- | ---- | ---- | ---- | ---- | ---- |\n| 1 | convnext 64x3-128x3-256x27-512x3 | 0.7881 | 0.78577 | 0.78147 |60h with 1xRTX4090 | |\n| 2 | convnext 96x3-192x3-384x27-768x3 | 0.7887 | 0.78697 | 0.78281 | 120h with 1xRTX4090 | |\n| 3 | convnext 128x3-256x3-512x27-1024x3 | 0.78693 | 0.78480 | 0.78150 | 132h with 1xRTX4090 | use 81.25% data |\n| 4 | convnext 144x3-288x3-576x27-1152x3 | 0.78547 | 0.78317 | 0.77903 | 28h with 8xRTX4090 |  |\n| 5 | transformer 512x4 | 0.7848 | 0.78341 | 0.77910 | 54h with 1xRTX4090 |  |\n| 6 | transformer 768x4 | 0.7843 | 0.78350 | 0.77966 | 72h with 1xRTX4090 |  |\n| 7 | convnext 144x3-288x3-576x27-1152x3 | 0.78769 | 0.78622 | 0.78137 | 16h with 8xRTX4090 | training 4epochs |\n\n## 1-2. Feature Engineering / Preprocessing\nI use all row-les datasets.\n\n### 1-2-1. Feature Engineering\n\n- diff\n  - x[i] - x[i-1]\n  - x[i] - x[i-2]\n- mean and diff_mean in similar features\n  - q0002, q0003\n  - state_u, state_v\n  - pbuf_*\n\n### 1-2-2. Preprocessing\n- Standard Scaler for feature / label\n  - Calculate feature mean/std with both train/test datasets.\n- Extremely large values can make model training unstable, so features are standardized and then clipped to the range of -100 to 100.\n- To convert to the shape of (batch_size, n_feature, 60), scalar values (e.g., state_ps) are transformed into time series data by repeating the same value 60 times.\n\n## 1-3. Training Methods\n### 1-3-1. Models\n#### 1-3-1-1. ConvNeXt\n\nA ConvNext with inputs of shape (batch_size, n_features, 60) and outputs of shape (batch_size, 368). The pseudocode is shown below.\n\n```python\nclass Head1D(nn.Module):\n    def __init__(self):\n        super(Head1D, self).__init__()\n        self.final_layer = nn.LazyConv1d(out_channels=14, kernel_size=1, stride=1, padding=\"same\")\n        self.fc = nn.Linear(8 * 60, 8)\n\n    def forward(self, x):\n        x = self.final_layer(x)  # (hidden_size, 60) -> (14, 60)\n        x_out = torch.cat(\n            [\n                x[:, :6, :].reshape(\n                    -1, 6 * 60\n                ),  # shape = (bs, 360) ptend_t, ptend_q0001, ptend_q0002, ptend_q0003, ptend_u, ptend_v\n                self.fc(x[:, 6:, :].reshape(-1, 8 * 60)),  # shape = (bs, 8)\n            ],\n            dim=1,\n        )  # shape = (bs, 360 + 8)\n        return x_out\n\nclass ConvNeXt(nn.Module):\n    def __init__(self):\n        self.head = Head1D()\n        ...\n\n    def forward_feature(x):\n        # convnext \n        ...\n        \n    def forward(self, x):\n        x = x[\"feature\"]  # shape = (bs, 384, n_features)\n           \n        # convnext part\n        x_features = self.forward_feature(x)  # shape = (bs, n_features, 60) -> (bs, hidden_dims, 60)\n        \n        x_pred = self.head(x_features)  # shape = (bs, hidden_dims, 60) -> (bs, 368)\n        return x_pred\n    \nmodel = ConvNeXt()\nx = torch.randn(4, 25, 60)\nassert model(x).shape == (4, 368)\n\n```\n\n- Based on https://github.com/facebookresearch/ConvNeXt/blob/d1fa8f6fef0a165b27399986cc2bdacc92777e40/models/convnext.py  ,I've rewritten it for 1D inputs with the following considerations:\n  - Adjusted to ensure the shape remains unchanged between input and output by setting stride=1 and padding=\"same\".\n  - Replaced LayerNorm with BatchNorm since LayerNorm was not effective.\n\n#### 1-3-1-2. Transformer\n\nA Transformer with input of shape  ``(batch_size, 60, n_features)`` and outputs of shape ``(batch_size, 368)``.  The pseudocode is shown below.\n\n```python\nclass TransformerModel(nn.Module):\n    def __init__(\n        self,\n        hidden_dims: int,\n        n_layers: int,\n        n_heads: int,\n        head_mode: str,\n        dropout: int = 0,\n    ):\n        super(TransformerModel, self).__init__()\n        \n        N_SEQUENTIAL_COLUMNS = 6\n        # stem\n        self.fc_stem = nn.Sequential(\n            nn.LazyLinear(hidden_dims),\n            nn.LayerNorm(hidden_dims),\n            nn.GELU(),\n            nn.LazyLinear(hidden_dims),\n            nn.LayerNorm(hidden_dims),\n            nn.GELU(),\n        )\n        self.position_encoder = nn.Embedding(60, hidden_dims)\n        \n        # transformer\n        layer = nn.TransformerEncoderLayer(\n            d_model=hidden_dims,\n            nhead=n_heads,\n            dim_feedforward=hidden_dims * 4,\n            dropout=0,\n            activation=\"gelu\",\n            batch_first=True,\n        )\n        self.transformer = nn.TransformerEncoder(layer, num_layers=n_layers)        \n        \n        # head\n        ## 1. cnn (for sequential (60x6))\n        self.fc_head_list_sequential = []\n        head_cnn_params = [{\"out_channels\": 256, \"kernel_size\": 3}] * 3\n        for _ in range(N_SEQUENTIAL_COLUMNS):\n            fc_head = []\n            for head_cnn_param in head_cnn_params:\n                fc_head.append(\n                    nn.LazyConv1d(\n                        stride=1,\n                        padding=\"same\",\n                        **head_cnn_param,\n                    )\n                )\n                fc_head.append(nn.BatchNorm1d(head_cnn_param[\"out_channels\"]))\n                fc_head.append(nn.GELU())\n            fc_head.append(\n                nn.LazyConv1d(1, kernel_size=3, stride=1, padding=\"same\")\n            )\n            fc_head = nn.Sequential(*fc_head)\n            self.fc_head_list_sequential.append(fc_head)\n        self.fc_head_list_sequential = nn.ModuleList(self.fc_head_list_sequential)\n        \n        ## 2. linear (for scalar (8))\n        self.fc_head_scalar = nn.Linear(hidden_dims, 8)\n        \n    def forward(self, x):\n\n        # stem\n        x = self.fc_stem(x)  # (bs, 60, n_features) -> (bs, 60, hidden_dims)\n        pe = self.position_encoder(\n            torch.arange(x.shape[1], device=x.device).expand(x.shape[0], -1)\n        )  # (60) -> (bs, 60, hidden_dims)\n        x = x + pe\n        \n        # transformer\n        x = self.transformer(x)  # (bs, 60, hidden_dims) -> (bs, 60, hidden_dims)\n        \n        # head\n        ## 1. cnn\n        x_out = []\n        for i in range(6):\n            x_out_ = self.fc_head_list_sequential[i](\n                x.permute(0, 2, 1)\n            )  # (bs, hidden_dims, 60) -> (bs, 1, 60)\n            x_out_ = x_out_.squeeze(1)  # (bs, 1, 60) -> (bs, 60)\n            x_out.append(x_out_)        \n        ## 2. linear\n        x_out_ = self.fc_head_scalar(x.mean(dim=1))  # (bs, hidden_dims) -> (bs, 8)\n        x_out.append(x_out_)\n        x_out = torch.cat(x_out, dim=1)  # (bs, 60*6+8)\n        return x_out        \n```\n\n\n### 1-3-2. Hyper Parameter\n| param_name | #1 | #2 | #3 | #4 | #5 | #6 | #7 | \n| --- | --- | --- | --- | --- | --- | --- | --- |\n| epochs | 7 | 7 | 7 | 7 | 7 | 7 | 4 |\n| lr | 2.5e-3 | 2e-3 | 2e-3 | 1.5e-3 | 1e-3 | 1e-3 | 2e-3 |\n| batch_size | 384 | 384 | 384 | 288 | 384 | 384 | 288 |\n| weight_decay | 0.05 | 0.05 | 0.05 | 0.05 | 0.01 | 0.01 | 0.075 |\n\n- other\n    - optimizer: AdamW\n        - cnn: AdamW(weight_decay=0.05 or 0.075)\n        - transformer: AdamW(weight_decay=0.01)\n    - loss: \n        - cnn: SmoothL1Loss(beta=0.01)\n        - transformer: SmoothL1Loss(beta=1)\n    - scheduler\n        - cnn: transformers.get_polynomial_decay_schedule_with_warmup(alpha=2, warmup_ratio=0.1)\n        - transformer: transformers.get_polynomial_decay_schedule_with_warmup(alpha=1, warmup_ratio=0.1), equals to get_linear_scheduler_with_warmup\n    - The batch size of 384 is a remnant of using lat/lon leak.\n    - Transformers tend to diverge quickly with a high learning rate, so keep it as low as possible. CNNs were not as sensitive.\n    - Regarding the beta of SmoothL1Loss, smaller models tended to perform better with something close to Huber Loss, while larger models showed better performance with something closer to L1 Loss.\n\n## 1-4. Postprocessing\n- Set some values to state * (1 / -1200)\n    - ptend_q0002_12 ~ ptend_q0002_28\n\n## 1-5. Others\n\nThis was the competition where we needed to rely on neural networks the most among all the competitions we've participated in so far.\n\n- Common\n    - Proper learning rate scheduling was crucial, and we spent a significant amount of time tuning it.\n    - We tested with a small amount of data (n=1.8m, n=9m) before using the full dataset (n=70m). There were cases where what worked with a small amount of data didn't work with the full dataset.\n    - The same issue occurred between small models (ConvNeXt 32x3-64x3-128x27-256x3) and large models (ConvNeXt 96x3-192x3-384x27-768x3).\n- ConvNeXt\n    - Polynomial decay scheduler and high weight decay were effective.\n    - Observing the train loss on W&B, it seemed that training progressed significantly when the learning rate was low. Therefore, we adopted a polynomial decay scheduler to stay at a low learning rate for a longer period. We found that training progressed too much and caused overfitting, so we controlled it with weight decay. (cv: +0.004)\n    - It seemed that the larger the model, the higher the accuracy, given proper hyperparameter settings.\n        - Larger models converged faster, so we reduced the number of epochs.\n        - We didn't have time to test this thoroughly towards the end.\n- Transformer\n    - We tried various architectures, and attaching a CNN-head or adding positional encoding proved effective.\n    - We tested large models of 512x4 and above with n=9m, but found no significant improvement over 512x4 or even a decrease in accuracy. It is unclear whether this was due to poor hyperparameter tuning or if 512x4 was simply sufficient for this dataset.\n\n\n\n# 2. Kami part\n- **Model**\n    - 1D Unet-based model x 11\n- It is important to reshape the input to (batch, 60, dim) to explicitly input height relationships into the model.\n- Performance improves with a lot of data and models with large parameters.\n\n## 2-1. Features Selection / Engineering\n\n- **Data**\n    - Low-resolution data\n        - Train: Data excluding validation from [February of Year 1, February of Year 9)\n        - Validation: Approximately 641,280 instances until February of Year 8 (skipping every 7 instances similar to Kaggle data)\n- **Input Normalization**\n    - Subtract the mean and divide by the standard deviation.\n    - State_t, q0001, q0002, q0003, u, v, ozone, ch4, n2o use common normalization across all heights.\n        - Reason: E3SM often performs operations by height, so it is preferable to standardize these features.\n    - Only q0001, q0002, q0003 undergo exponential change and are normalized as follows: multiply by 1e9, apply log1p, then normalize.\n        - Reason: Possibly related to the Clausius-Clapeyron equation, which is somewhat utilized by the model.\n- **Output Normalization**\n    - Subtract the mean and divide by the standard deviation.\n- **Feature Engineering**\n    - Relative humidity (expresses the relative amount of q1)\n        - Reference: [**Climate-invariant machine learning**](https://www.science.org/doi/10.1126/sciadv.adj7250) ([supplementary material](https://pog.mit.edu/src/beucler_climate_invariant_ml_supplement_2024.pdf))\n        - The calculation method for saturation vapor pressure in the paper did not align with E3SM results, so Bolton's method used in E3SM was adopted.\n    - Ice rate: q0002 / (q0002 + q0003)\n    - Cloud water: (q2 + q3)\n    - Add categorical features such as height (0~59) information and whether q0002/q0003 are zero using a 5-dimensional embedding.\n\n## 2-2. Models\n\n- **GPU:** V100\n\n| # | Model Name | CV | LB | Training Time | Note |\n| --- | --- | --- | --- | --- | --- |\n| 1 | 204_diff_last_all_lr | 0.7768 | 0.77351 | 1d 4h | |\n| 2 | 201_unet_multi_all_n3_restart2 | 0.7783 | - | 22h | |\n| 3 | 201_unet_multi_all_512_n3 | 0.7794 | - | 1d 9h | |\n| 4 | 201_unet_multi_all_384_n2 | 0.7801 | - | 22h | |\n| 5 | 201_unet_multi_all | 0.7815 | - | 1d 7h | |\n| 6 | 217_fix_transformer_leak_all_cos_head64 | 0.7817 | - | 1d 7h | With transformer head |\n| 7 | 217_fix_transformer_leak_all_cos_head64_n4 | 0.7828 | - | 1d 20h | With transformer head |\n| 8 | 222_wo_transformer_all | 0.7839 | - | 2d 21h | |\n| 9 | 222_wo_transformer_all_004 | 0.7830 | - | 2d 21h | Parameter: 354 M |\n| 10 | 225_smoothl1_loss_all_005 | 0.7833 | - | 3d 7h | |\n| 11 | 225_smoothl1_loss_all_beta | 0.7828 | - | 2d 8h | |\n\n- **Base Structure:** height mlp → shallow 1D Unet x 2 → height mlp\n- **Initial Height MLP**\n    - Apply a common weight MLP to each height.\n        - Reason: E3SM often performs operations by height & features should be easy to use later in 1D Unet.\n- **1D Unet**\n    - A 1D Unet with many channels.\n    - Start with 256 dimensions, doubling the number of channels with each convolution, repeated 3 times.\n- **Output Height MLP**\n    - Apply a common weight MLP to each height.\n    - Prepare MLPs for ptend_t, q0001, q0002, q0003, u, v & directly input related features with skip connections.\n        - For example, include state_t with ptend_t.\n- **Other Scalar Predictions**\n    - Predict using the bottleneck of the 1D Unet with an MLP.\n- **Final Output**\n    - The outputs of height MLP for state_t, q0001, q0002, q0003, u, v are given as 60x2 each, then expressed as x1.exp() - x2.exp() to slightly improve the score.\n        - Reason: The exponential relationship in the Clausius-Clapeyron equation and the model aims to predict the difference before and after the change.\n- **Training Method**\n    - Ignore some labels during training.\n        - Labels with weight 0 in the sample submission.\n        - Labels to be ignored in post-processing.\n    - Optimizer: Adan (not Adam)\n    - Scheduler: Cosine schedule with warmup or reduce LR on plateau.\n- **Prediction**\n    - Use EMA for prediction.\n\n## 2-3. Post-Processing\n\n- Set some values to state * (1 / -1200)\n    - ptend_q0002_12 ~ ptend_q0002_28\n- Target: all values in some labels & low or high temperature in some q2, q3 labels\n\n\n\n# 3. Takoi part\n## 3-1. Summary\n- Used data from Hugging Face's LEAP/ClimSim_low-res for training\n- Created features from the data in a time series format\n- Developed 12 models based on LSTM\n\n## 3-2. Features Selection / Engineering\n- Data\n    - Validation: Used the last 639,744 rows from Kaggle's train.csv\n    - Train: Used Hugging Face's LEAP/ClimSim_low-res data (excluding the 9th year and periods overlapping with the validation period)\n\n- Features\n    - Divided into time series parts and others\n    - Time series parts (considered as a series of length 60)\n        - Original data\n        - Differences from the subsequent data points in the series\n    - Other parts\n        - Original data\n        - Sum of state_q0001, state_q0002, and state_q0003\n\n- Preprocessing\n    - Features\n        - StandardScaler\n            - Applied StandardScaler to each column\n    - Target\n        - Used the value after applying weight to the target\n        - StandardScaler\n            - Applied StandardScaler to each column\n\n## 3-3. Models\n- The models are based on LSTM\n- The following is the base model\n    - Ultimately, multiple models were created by enlarging the base model or adding conv1d\n\n\n```python\nclass LeapRnnModel(nn.Module):\n    def __init__(\n            self,\n            input_numerical_size=9 * 2,\n            numerical_linear_size=64,\n            input_numerical_size2=17,\n            numerical_linear_size2=64,\n            model_size=256 * 2,\n            linear_out=256,\n            out_size1=6,\n            out_size2=8):\n        super(LeapRnnModel, self).__init__()\n        self.numerical_linear = nn.Sequential(\n            nn.Linear(input_numerical_size,\n                      numerical_linear_size),\n            nn.LayerNorm(numerical_linear_size)\n        )\n        self.numerical_linear2_list = nn.ModuleList(\n            [nn.Sequential(\n                nn.Linear(input_numerical_size2,\n                          numerical_linear_size2),\n                nn.LayerNorm(numerical_linear_size2)\n            ) for _ in range(60)]\n        )\n        self.numerical_linear2 = nn.Sequential(\n            nn.Linear(input_numerical_size2,\n                      numerical_linear_size2),\n            nn.LayerNorm(numerical_linear_size2)\n        )\n        self.rnn = nn.LSTM(numerical_linear_size + numerical_linear_size2,\n                           model_size,\n                           num_layers=3,\n                           batch_first=True,\n                           bidirectional=True)\n        self.linear_out1 = nn.Sequential(\n            nn.Linear(model_size * 2,\n                      linear_out),\n            nn.LayerNorm(linear_out),\n            nn.ReLU(),\n            nn.Linear(linear_out,\n                      out_size1))\n        self.layernorm = nn.LayerNorm(model_size * 2)\n        self.linear_out2 = nn.Sequential(\n            nn.Linear(model_size * 2 + numerical_linear_size2,\n                      linear_out),\n            nn.LayerNorm(linear_out),\n            nn.ReLU(),\n            nn.Linear(linear_out,\n                      out_size2))\n        self._reinitialize()\n\n    def _reinitialize(self):\n        \"\"\"\n        Tensorflow/Keras-like initialization\n        \"\"\"\n        for name, p in self.named_parameters():\n            if 'rnn' in name:\n                if 'weight_ih' in name:\n                    nn.init.xavier_uniform_(p.data)\n                elif 'weight_hh' in name:\n                    nn.init.orthogonal_(p.data)\n                elif 'bias_ih' in name:\n                    p.data.fill_(0)\n                    # Set forget-gate bias to 1\n                    n = p.size(0)\n                    p.data[(n // 4):(n // 2)].fill_(1)\n                elif 'bias_hh' in name:\n                    p.data.fill_(0)\n\n    def forward(self, seq_array,\n                other_array):\n\n        numerical_embedding = self.numerical_linear(seq_array)\n        other_embedding = self.numerical_linear2(other_array)\n        numerical_embedding2_list = [\n            linear(other_array) for linear in self.numerical_linear2_list]\n        numerical_embedding2 = torch.stack(numerical_embedding2_list, dim=1)\n        numerical_embedding_concat = torch.cat(\n            [numerical_embedding, numerical_embedding2], dim=2)\n        output_seq, _ = self.rnn(numerical_embedding_concat)\n        output_other = torch.mean(output_seq, dim=1)\n        output_other = self.layernorm(output_other)\n        output_other = torch.cat([output_other, other_embedding], dim=1)\n        output_seq = self.linear_out1(output_seq)\n        output_other = self.linear_out2(output_other)\n        return output_seq, output_other\n```\n\n\n### 3-4. Training Method\n- loss : SmoothL1Loss\n- scheduler : get_cosine_schedule_with_warmup\n- optimizer : AdamW\n- lr : 1e-3\n\n### 3-5. Post-Processing\n- Replace values of ptend_q0002_0 to ptend_q0002_27 with the corresponding values of state_q0002_0 to state_q0002_27 divided by (-1200).\n- Set columns with a weight of 0 in sample_submission.csv to 0.\n\n### 3-6. Model Results\n- All the models below are trained using A100\n- CV is evaluated using the validation data\n    - Results after post-processing\n- Some experiments do not use the entire dataset mentioned above, so the approximate size of the training data is noted\n\n| # | exp no | CV | Training Time(h) | Data Size | Note |\n| --- | --- | --- | --- | --- |--- |\n| 1 | 124| 0.7812|  24h | about 45M ||\n| 2 | 130 |0.7815 |  30h | about 45M |add 1dcnn|\n| 3 | 131 |0.7816  | 34h | about 60M ||\n| 4 | 133 |0.7819  | 40h | about 60M| add 1dcnn|\n| 5 | 134 |0.7817  | 36h | about 60M  |add 1dcnn|\n| 6 | 135 |0.7819  | 42h | about 70M ||\n| 7 | 136 |0.7821  | 45h | about 70M |add 1dcnn |\n| 8 | 138 |0.7824  | 51h | about 70M |add 1dcnn|\n| 9 | 139 | 0.7822 | 59h | about 70M |add 1dcnn |\n| 10 | 141 |0.7825  | 116h | about 70M |add 1dcnn + large model|\n| 11 | 159 |0.7827  | 52h | about 70M |add 1dcnn|\n| 1２ | 162 |0.7838  | 47h |about 70M  |add 1dcnn + large model|\n\n\n\n# 4. Ensemble/Stacking\n## 4-1. Summary\nEnsemble 30 models.\n\n| # | method | public | private |\n| --- | --- | --- | --- |\n| 1 | nelder-mead | 0.79100 | 0.78713 |\n| 2 | 1d-cnn stacking | 0.79193 | 0.78774 |\n\n## 4-2. Nelder-mead\n- The ensemble weights are determined using Nelder-Mead based on the predictions from team members.\n- The weights are optimized for each target group (such as ptend_t or ptend_q0001).\n- The ensemble results are used to replace the predictions from ``4-3. 1D-CNN Stacking``.\n\n## 4-3. 1D-CNN Stacking\n\nCreate simple 1d-cnn model with inputs of  ``(batch_size, n_models, n_labels(=368))`` and outputs of ``(batch_size, n_labels)`` and 10-folds ensemble.\n\n```python\n\nclass Model1DCNN(nn.Module):\n    def __init__(self, hidden_dims, kernel_sizes):\n        super(Model1DCNN, self).__init__()\n\n        conv = []\n        for hidden_dim, kernel_size in zip(hidden_dims, kernel_sizes):\n            conv.append(\n                nn.LazyConv1d(\n                    out_channels=hidden_dim,\n                    kernel_size=kernel_size,\n                    stride=1,\n                    padding=\"same\",\n                )\n            )\n            conv.append(nn.BatchNorm1d(hidden_dim))\n            conv.append(nn.GELU())\n        conv.append(\n            nn.LazyConv1d(out_channels=1, kernel_size=2, stride=1, padding=\"same\")\n        )\n        self.conv = nn.Sequential(*conv)\n\n    def forward(self, x):\n        # x.shape = (batch_size, n_models, seq_len=368)\n        x = self.conv(\n            x\n        )  # (batch_size, n_models, seq_len=368) -> (batch_size, hidden_size, seq_len=368)\n        x = x.mean(\n            dim=1\n        )  # (batch_size, hidden_size, seq_len=368) -> (batch_size, seq_len=368)\n        return x\n```\n\nHyperparameter is below: \n- lr: 1e-3\n- batch_size: 256\n- hidden_size: ``[256, 256, 256]``\n- kernel_size: ``[3, 3, 3]``\n- epochs: 20\n- optimizer: ``AdamW(weight_decay=0)``\n- scheduler: ``linear scheduler with warmup``\n- loss: ``SmoothL1Loss(beta=1)``\n\n",
    "2940740": "Hi @kurupical,\n\nFirst of all, congratulations on your impressive 4th place finish in the LEAP - Atmospheric Physics using AI (ClimSim) competition! Your detailed breakdown of the models, feature engineering, and training methods is incredibly insightful. I especially appreciated your approach to using ConvNeXt and transformers for handling the atmospheric data.\n\nI had a question about your feature engineering process. Specifically, you mentioned standardizing features and then clipping to the range of -100 to 100. Could you elaborate on the impact this had on your model's performance? Did you notice any significant improvements or stability in training after implementing this step?\n\nThank you for sharing your solution and for any insights you can provide."
  }
}