{
  "id": 523063,
  "title": "1st place solution for the LEAP competition",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/523063",
  "author_name": "greySnow",
  "post_date": "2024-07-30T02:38:34.904000",
  "votes": 82,
  "comment_count": 20,
  "views": 0,
  "content": "<p>This competition was an amazing experience and a wonderful sandbox to test ideas and experiments with various techniques and architectures in-depth. Thanks to Kaggle for this fantastic platform, to Colab for providing me cheap TPU to train on, to Google as the owner of both Kaggle and Colab and for providing us with Tensorflow. When I think about it, I trained on TPU (google) on tensorflow framework (google) on Colab (google) for a competition held on Kaggle (google). It's all Google from start to end. You are amazing, and I salute you.   </p>\n<p>I also want to express my heartfelt gratitude to the competition host <a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> . Your efforts in providing us with this amazing competition and your unwavering commitment to keeping it on track, even when facing unexpected LEAK problems, are truly commendable. Thank you for your dedication and hard work.  </p>\n<p>Finally, thank you to the Kaggle community, especially to all the active and helpful people on the forum and in the code section. You make the experience so much better. I love you.  </p>\n<p>This competition was a strange one. I built the best models, but there were better Kagglers than me- people who found the leak(s) and succeeded in utilizing them. I did not. I still have much to learn. I won't pretend I'm unhappy that things turned out the way they did, but the best Kagglers deserve respect. You have mine.</p>\n<p>I know that as 1st place, and considering the significant gap in scores between my solution and the next ones, a lot of eyes would be on me to ensure I did not exploit any LEAK. Hence, I made a lot of effort to provide a complete Kaggle pipeline, from downloading the data from HF through TFRecords encoding, training, inferring, and submission, with a training example that resulted in a single model 0.79081/0.78811 public/private LB. By following my pipeline step by step, you can also construct the dataset and train your own ~0.79+ public LB model from scratch, ensuring no hiding of leaky pseudo-labels in the train set or nefarious test-set reverse engineering in inference. All of this by simply copying and running a series of Kaggle notebooks. Admittedly, there are a lot of notebooks- over 100 notebooks if you intend to download and encode to TFRecords all the data yourself- but with the massive amount of data in this competition, there is no way around it. See the complete pipeline and extra details <a href=\"https://github.com/shlomoron/LEAP-solution\" target=\"_blank\">in my GitHub</a>.</p>\n<h2>Context section</h2>\n<p><a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim\" target=\"_blank\">Business context</a>.  <br>\n<a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/data\" target=\"_blank\">Data context</a>.  </p>\n<h2>1. Overview of the Approach</h2>\n<h3>1.1. model</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6832115%2Fb55d82c6ef742f31b4c560b7bc396d51%2Fsummary_graphics.png?generation=1722307603587353&amp;alt=media\" alt=\"\"><br>\nIn one word, squeezeformer. This would not surprise anyone who followed some of the more similar competitions on Kaggle in the last year, in particular <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling\" target=\"_blank\">ASLFR</a> and <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding\" target=\"_blank\">Ribonanza</a>. I used the same modified squeezeformer blocks that I saw in the <a href=\"https://github.com/hoyso48/Stanford---Ribonanza-RNA-Folding-2nd-place-solution\" target=\"_blank\">2nd solution to Ribonanza</a> by <a href=\"https://www.kaggle.com/hoyso48\" target=\"_blank\">@hoyso48</a> . The only changes I remember doing is that I deleted the first LayerNorm in the 1Dconv block and added an ECA layer. However, there may have been more minor changes that I forgot. My Tensorflow implementation was guided by hoyso48 PyTorch implementation, so once again (3rd competition in a row), I give him my thanks.  <br>\nI used 12 block models with dimensions of 256/384/512. Before the squeezeformer blocks, I have a linear dense layer followed by LayerNorm as an encoder from the input data to the model dimensions. After the squeezeformer blocks, I have a prediction head of swish dense followed by a GLUMlp block (swiGLU followed by linear dense), both with head dimensions of 1024/2048, depending on the model. For more details, check my code.  <br>\nAt first, I used dropout layers, which helped greatly when the training data was in the order of 1M samples. However, after I saw comments in the forum about dropout being unnecessary, I experimented again when I scaled up to ~40M-80M samples, and it was indeed unnecessary when ensembling (although it still allows for higher scores for single models, at the cost of twice or thrice epochs). Since removing it allows for much faster training, I also chose to drop the dropout altogether. Optimizer was AdamW with half-cosine decay scheduler, max LR 1e-3, and weight decay = 4*LR.</p>\n<h3>1.2. Loss</h3>\n<p>I once read that the most important part of a DL model is the loss function. I was a bit skeptical then- sure, it is important, but it's not exactly complicated to choose the appropriate loss, right? Say, if our metric is R-squared, the best loss will obviously be…MSE?</p>\n<h4>1.2.1. Use MAE</h4>\n<p>It performs better than MSE in every way- it converges faster, is better at convergent, and is more stable. If I need to choose one 'secret' of this competition for a high score, aside from the notorious leaks, it is this.</p>\n<h4>1.2.2. Auxiliary loss</h4>\n<p>We had explicit spacetime data for the train set but not for the test set. Auxiliary loss is a common practice in such cases. I used MAE on the normalized latitude/longitude and the sin/cos on the day/year cycles (if it is unclear, please look at my code).<br>\nA side note on the auxiliary loss- some people speculated, in the last week of the competition, that using explicit spacetime data, even that of the training set, is considered a forbidden leak. So, first, such use was <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/519772#2919233\" target=\"_blank\">confirmed by the host to be legit</a>, as long as there is no 'hacking' of the test set (which can be done by using the auxiliary predictions to pseudo-label the test set- a thing I did not do). Second, in the last week, I trained several models without auxiliary spacetime loss, and my no-spacetime ensemble of 6 models got 0.79355/0.79092. So I win anyway; auxiliary spacetime loss is unnecessary and may not even be helpful (since my winning ensemble, while with a slightly better score, is also 13 models).  </p>\n<h4>1.2.3. Confidence head</h4>\n<p>The <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/460403\" target=\"_blank\">3rd place solution in Ribonanza</a> got a suspiciously good score for a particular model. I won't explain why it was suspicious; you will need to know Ribonanza's details for the context. In any case, this suspicious model had two unique things going for it- a strange architecture and a confidence head. I still need to experiment more with said strange architecture, but for the confidense head, it was easy to try in this competition and surprisingly effective. The idea is simple- I also predict the loss of each target. I also used MAE for the confidence loss. I have one model I trained without confidence head, and it got LB 0.78945/0.78631, so I think I would also won without it. Then again, a model with the seme specifications, but with confidense head was my best model with LB 0.79159/0.78869. So, yeah. Surprisingly effective. And yes, this second model can win 1st place in this competition by itself (0.78869 private, although not 1st in public).</p>\n<h4>1.2.4 Masked loss</h4>\n<p>I masked out the targets that are zeroed in the submission or those for which we use the ptend trick (ptend_q0002_2-ptend_q0002_26).  <br>\nFor those new at LEAP- please read <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/502484\" target=\"_blank\">this post</a> for ptend trick context.  </p>\n<p>Final thoughts on the loss function: well, it is probably the most important part of a deep learning model. lol.  </p>\n<h3>1.3. Data preparation</h3>\n<h4>1.3.1 High-res data</h4>\n<p>You may have noticed already that I used high-res data. Yes, it helps. Although I would have won without it- my only low-res data ensemble of five models got LB 0.79299/0.78951. But high-res definitely helped, although it sometimes requires a special soft-clipping treatment, as you will see. Also, training with high-res is less stable. Hence, I usually used larger batch sizes (1024/2048 compared to 512 for low-res only models). I blended high-res data with a ratio of 2:1 low: high. Lower than that, the performance is weaker; higher than that, the gains are small compared to the extra training time required.    </p>\n<h4>1.3.2 Multiple data representation</h4>\n<p>Multiple data representations is a trick I learned from 1st solution <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/434485\" target=\"_blank\">at ASLFR</a>. Without going into too much detail, the situation in ASLFR was that the data could be normalized in two different ways. I remember trying both ways, finding out what is better and sticking with it. Then the competition ended, and guess what? The 1st place normalized in the two possible ways, concatenated the two representations (with a few extra steps in between; read their summary for the full details) and sent it to the model.  <br>\nFirst, Let me separate the features that are spread over the 60-height level, which I call X_col and the features that are the same for all the levels, which I call X_col_not.  <br>\nFor X_col_not, I used only one representation: the simple normalization (x-mean)/std.   <br>\nFor X_col, I used three representations. The first the same as X_col_not, where I normalize each feature in each level with its own mean/std. i.e., for state_t, then we have (state_t_1-mean(state_t_1))/std(state_t_1), (state_t_2-mean(state_t_2))/std(state_t_2) etc. In my code, I call this representation x_col_not_norm (for x_col_not) and x_col_norm (for X_col).  <br>\nThe second representation normalizes each feature by the total mean and std over all the levels. For state_t, we have (state_t_1-mean(state_t))/std(state_t), (state_t_2-mean(state_t))/std(state_t), etc. In my code, I call this representation x_total_norm.  <br>\nFinally, the third representation is:  </p>\n<pre><code>x_col_norm_log = tf((x_col_norm-x_col_norm_min+)&gt;=, tf(x_col_norm-x_col_norm_min+),\n                                    -tf(+(x_col_norm-x_col_norm_min+)))\n</code></pre>\n<p>This is the kind of thing that trying to explain with words would never be as clear as just looking at the code.  <br>\nAfter you looked at the code and understood it, you may wonder why not use:  </p>\n<pre><code> = tf.math.log(x_col_norm-x_col_norm_min+)\n</code></pre>\n<p>Why all the extra steps with the tf.cond? See, I had a problem. I calculated x_col_norm_min only with Kaggle data, and then I scaled up my code to all HF data (which have values lower than Kaggle data x_col_norm_min) but did not want to change the normalization constant because it would break the inference pipeline. Then, I would have to use different pipelines for my old and new models. Yeah, sometimes I'm a bit lazy. Proud of it. And it turned out to be an excellent choice when I included also high-res data.  </p>\n<h4>1.3.3 Wind</h4>\n<p>$$wind = \\sqrt{(state_u)^2+(state_v)^2}   $$<br>\nIt made sense, and I wanted to include at least one 'physically justified' thing in the model (silly me, yes). It did not really help but also did not hurt the model, so it stayed. I used only the first normalization for WIND, with:  <br>\nmean(wind) = mean(mean(state_u), mean(state_v))  <br>\nAnd with:  <br>\nstd(wind) = sum(std(state_u), std(state_v)  </p>\n<p>All in all, the feature dimension is:  <br>\n9[col_features]*60[levels]*3[representations]+60[wind_levels]+16[not_col features] = 1696  </p>\n<h4>1.3.4 Features soft clipping 1</h4>\n<p>After normalization, the data has some extreme values (~±3000). This problem exists only for the first representation. The model actually handled it easily, but I preferred to play on the safe side. So, for x_col_norm and WIND, I applied the following soft clipping:  </p>\n<pre><code> = \n = cut**\n = tf.where(x_col_norm&gt;cut, x_col_norm**+cut-square_cut, x_col_norm)\n = tf.where(x_col_norm&lt;-cut, -tf.math.abs(x_col_norm)**-cut+square_cut, x_col_norm)\n</code></pre>\n<h4>1.3.5 Features soft clipping 2</h4>\n<p>In addition to the first soft clipping, I applied a second soft clipping to deal with extreme values from the high-res set. I applied the clipping on all the representations, including WIND, and after I applied the first soft clipping (1.3.3) for the relevant features:  </p>\n<pre><code> = \n = tf.math.log(cut_2)\n = tf.where(x_col_norm&gt;cut_2, tf.math.log(x_col_norm)+cut_2-log_cut, x_col_norm)\n = tf.where(x_col_norm&lt;-cut_2, -tf.math.log(-x_col_norm)-cut_2+log_cut, x_col_norm)\n</code></pre>\n<p>I chose cutoff_2 so that the soft clipping would only affect high-res data.  </p>\n<h4>1.3.6 Targets soft clipping</h4>\n<p>I did this only for the high-res targets; each target was soft-clipped if it was too extreme compared to the low-res corresponding target min/max (this differs from 1.3.4/1.3.5 in that the soft-clipping range was different for each target). In code:  </p>\n<pre><code>rescale_factor = \n x == :\n    y_norm = tf(y_norm&lt;norm_y_min*rescale_factor, norm_y_min*rescale_factor-tf(-y_norm+norm_y_min*rescale_factor), y_norm)\n    y_norm = tf(y_norm&gt;norm_y_max*rescale_factor, norm_y_max*rescale_factor+tf(+y_norm-norm_y_max*rescale_factor), y_norm)\n</code></pre>\n<h3>1.4 Post-processing</h3>\n<h4>1.4.1. Downcast and Upcast</h4>\n<p>Special care should be taken when moving between FP64 and FP32. I encoded my TFRecords with the original values in FP64. After I processed the data (see 1.3. Except for 1.3.5 which happened after the casting for no particular reason), the values were downcast to FP32 before transferring them to the model. Then I upcasted the predictions to FP64, and only then did I apply de-normalization to get the values for submission:  </p>\n<pre><code> = preds + mean_y.reshape(,-)*stds\n[:, np.where(stds_new == )] = \n = preds/np.where(stds&gt;, stds, )\n</code></pre>\n<h4>1.4.2. Mean for bad targets</h4>\n<p>This is very simple:  </p>\n<pre><code>metrics = np(, preds)    ()])\n   ((metrics)):\n     metrics&lt;:\n        preds = \n</code></pre>\n<p>In reality, it was eventually unnecessary because the only bad targets were those zeroed out in the submission or in the ptend trick range (see next bullet, 1.4.3).</p>\n<h4>1.4.3. Ptend trick</h4>\n<p>Obviously. If you are new at LEAP, look <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/502484\" target=\"_blank\">here for details</a>.  </p>\n<h3>1.5 Validation</h3>\n<p>My local validation included two validation sets of randomly selected samples from the low-res dataset (the same samples for all the models), each with a size of 100K samples. In addition, I randomly selected 100K samples from the high-res dataset and 100K samples from the low-res training set (i.e., the last validation set is on samples I also train on to judge overfitting better). All in all, I had four validation sets.  <br>\nWhen I trained on ~1M samples, the local validation was highly correlated with the public leaderboard score, but when I scaled up to all the data, it was less correlated (probably because when I used the full train set, it includes samples that are very close in time to the local validation samples and with the same latitude/longitude, so it starts to overfit even on the validation set, as opposed to the public test set that includes samples from entirely different year than those that exist in the train set). For example, I could push my local validation score to ~0.8, but on the public LB, it will get ~0.785 and be lower in score than a model that achieved, say, 0.795 but trained for fewer epochs. So, for the final models trained on all the low-res data from HF (and those I trained also on the high-res data), I validated against the public LB to get the hang of good overfitting range, and ensembling validation was done directly against the public LB. Thus, my ensemble is slightly overfitted to the public LB, although I acted in ways that reduce said overfitting (equal blending weights and including 'less successful' models if I judged them to be successful enough and anticipated them to be good based on the parameters I used). If it sounds a bit like black magic- yes, it is! Deep learning is sometimes a science, sometimes an art.  </p>\n<h2>2. Details of the submission</h2>\n<h3>2.1 Ensembling</h3>\n<p>My winning ensemble included 13 models, each a bit different (see full details in my GitHub, 'The steps to reproduce my solution' bullet 5). The best model (11) was LB 0.79159/0.78869, and the worst (2) was LB 0.78795/0.78388. Both were best/worst both in public and private LB. The full ensemble was LB 0.79410/0.79123. In addition, my low-res-data-only ensemble of 5 models has LB 0.79299/0.78951, and my no-spacetime-auxilliary-loss ensemble of 6 models has 0.79355/0.79092. When you read my solution, you may have tried to find out the 'secret sauce' that got me the 1st place, but it really was the combination that made the difference. Every single technique that I used, I think I could still get to 1st place without it.  </p>\n<h3>2.2 The helpful techniques</h3>\n<p>This is a short summary of the methods I wrote about in-depth above: Squeeseformer, wide GLUMlp prediction head, no dropout, MAE, auxiliary spacetime loss, confidence head, masked loss, multiple data representation, high-res data, features and targets soft-clipping and careful downcast/upcast.</p>\n<h3>2.3 What didn't work</h3>\n<p>Various model architectures (pure transformer, other 1Dconv/transformer combinations, Unet, dropout, smaller models, larger models, other optimizers), in short, many less optimal hyper-parameters. Log-normalization (i.e., log(x), not my log(1+x) representation which deals with different issues). MSE, MSE/MAE various combinations, weighted loss function (check my Ribonanza solution for details), levels/features masking, predict the year as additional auxiliary loss, and probably other not-very-important things that I don't remember already.  </p>\n<h3>2.4 Hardware</h3>\n<p>I trained on Kaggle and Colab TPU and inferred on Kaggle P100 GPU. Compute-wise, my experiments and training cost at least 200 bucks in Colab compute units and probably less than 300 bucks in total. If it sounds like a lot compared to your 'free' personal machine, please consider the electricity cost of using an RTX4090…</p>\n<h3>2.5 Bonus: confidence is key</h3>\n<p>Since I already had confidence predictions simply because the confidence head improved the model, I took a step further out of curiosity and checked if I could get a higher score by discarding 'low-confidence' samples and by how much I could improve. Here is a graph (only for one model, not the full ensemble) that shows the R-squared as a function of the fraction of remaining samples after I discarded the most 'low-confidence' ones.  <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6832115%2F3133894fc48f275cf79b88cee8a425aa%2Fconfidence_is_key.png?generation=1722304121419960&amp;alt=media\" alt=\"\"><br>\nAs you can see, by discarding ~0.1 of the samples, I can already get to a ~0.83+ score. This is not very surprising: if you check the data, you will see that certain months and locations have much lower scores, so a higher score can simply be achieved by discarding said months/location. However, discarding by confidence may be more precise and efficient. It can be interesting to research further, especially given this specific problem, since we can replace the simulator calculation with model predictions and then send back the low-confidence samples to the simulator.  <br>\nAs a side note, I also tried ensembling by confidence (i.e., giving lower weights to low-confidence predictions), but my preliminary research did not show promising results. Although it WAS preliminary- I got to it only on sunday of the last week of the competition, and it was either researching it more in depth or watching the Euro 2024 finals…(congrats Spain!)  </p>\n<h2>3. Sources</h2>\n<p>This is my favorite part. Thank you all the resources that have helped me!  <br>\n<a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/502484\" target=\"_blank\">Ptend trick</a> and also <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/499896#2791290\" target=\"_blank\">here originally</a> by <a href=\"https://www.kaggle.com/sakvaua\" target=\"_blank\">@sakvaua</a> and <a href=\"https://www.kaggle.com/jano123\" target=\"_blank\">@jano123</a> .<br>\n<a href=\"https://github.com/hoyso48/Stanford---Ribonanza-RNA-Folding-2nd-place-solution\" target=\"_blank\">Ribonanza 2nd solution</a> for Squeezeformer architecture guidance and insights, by <a href=\"https://www.kaggle.com/hoyso48\" target=\"_blank\">@hoyso48</a> .  <br>\n<a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/460403\" target=\"_blank\">Ribonanza 3rd place solution</a> for confidence head method, by <a href=\"https://www.kaggle.com/dankrstev\" target=\"_blank\">@dankrstev</a> .  <br>\n<a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/434485\" target=\"_blank\">ASLFR 1st solution</a> for the multiple data representation method, by <a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a> and <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> .  <br>\n<a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/514020#2884414\" target=\"_blank\">Dropout is unnecessary</a> and <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/514020#2885043\" target=\"_blank\">also here</a> by <a href=\"https://www.kaggle.com/sakvaua\" target=\"_blank\">@sakvaua</a> and <a href=\"https://www.kaggle.com/ymatioun\" target=\"_blank\">@ymatioun</a> .   <br>\nIn addition, I used the low-res and high-res data <a href=\"https://huggingface.co/LEAP\" target=\"_blank\">from HF</a>.  <br>\nAnd, of course, <a href=\"https://github.com/shlomoron/LEAP-solution\" target=\"_blank\">my GitHub</a> with links and instructions for a fully reproducible solution on a Kaggle pipeline.</p>",
  "messages": [
    {
      "id": 2940353,
      "postDate": "2024-07-30T02:38:34.903Z",
      "content": "<p>This competition was an amazing experience and a wonderful sandbox to test ideas and experiments with various techniques and architectures in-depth. Thanks to Kaggle for this fantastic platform, to Colab for providing me cheap TPU to train on, to Google as the owner of both Kaggle and Colab and for providing us with Tensorflow. When I think about it, I trained on TPU (google) on tensorflow framework (google) on Colab (google) for a competition held on Kaggle (google). It's all Google from start to end. You are amazing, and I salute you.   </p>\n<p>I also want to express my heartfelt gratitude to the competition host <a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> . Your efforts in providing us with this amazing competition and your unwavering commitment to keeping it on track, even when facing unexpected LEAK problems, are truly commendable. Thank you for your dedication and hard work.  </p>\n<p>Finally, thank you to the Kaggle community, especially to all the active and helpful people on the forum and in the code section. You make the experience so much better. I love you.  </p>\n<p>This competition was a strange one. I built the best models, but there were better Kagglers than me- people who found the leak(s) and succeeded in utilizing them. I did not. I still have much to learn. I won't pretend I'm unhappy that things turned out the way they did, but the best Kagglers deserve respect. You have mine.</p>\n<p>I know that as 1st place, and considering the significant gap in scores between my solution and the next ones, a lot of eyes would be on me to ensure I did not exploit any LEAK. Hence, I made a lot of effort to provide a complete Kaggle pipeline, from downloading the data from HF through TFRecords encoding, training, inferring, and submission, with a training example that resulted in a single model 0.79081/0.78811 public/private LB. By following my pipeline step by step, you can also construct the dataset and train your own ~0.79+ public LB model from scratch, ensuring no hiding of leaky pseudo-labels in the train set or nefarious test-set reverse engineering in inference. All of this by simply copying and running a series of Kaggle notebooks. Admittedly, there are a lot of notebooks- over 100 notebooks if you intend to download and encode to TFRecords all the data yourself- but with the massive amount of data in this competition, there is no way around it. See the complete pipeline and extra details <a href=\"https://github.com/shlomoron/LEAP-solution\" target=\"_blank\">in my GitHub</a>.</p>\n<h2>Context section</h2>\n<p><a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim\" target=\"_blank\">Business context</a>.  <br>\n<a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/data\" target=\"_blank\">Data context</a>.  </p>\n<h2>1. Overview of the Approach</h2>\n<h3>1.1. model</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6832115%2Fb55d82c6ef742f31b4c560b7bc396d51%2Fsummary_graphics.png?generation=1722307603587353&amp;alt=media\" alt=\"\"><br>\nIn one word, squeezeformer. This would not surprise anyone who followed some of the more similar competitions on Kaggle in the last year, in particular <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling\" target=\"_blank\">ASLFR</a> and <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding\" target=\"_blank\">Ribonanza</a>. I used the same modified squeezeformer blocks that I saw in the <a href=\"https://github.com/hoyso48/Stanford---Ribonanza-RNA-Folding-2nd-place-solution\" target=\"_blank\">2nd solution to Ribonanza</a> by <a href=\"https://www.kaggle.com/hoyso48\" target=\"_blank\">@hoyso48</a> . The only changes I remember doing is that I deleted the first LayerNorm in the 1Dconv block and added an ECA layer. However, there may have been more minor changes that I forgot. My Tensorflow implementation was guided by hoyso48 PyTorch implementation, so once again (3rd competition in a row), I give him my thanks.  <br>\nI used 12 block models with dimensions of 256/384/512. Before the squeezeformer blocks, I have a linear dense layer followed by LayerNorm as an encoder from the input data to the model dimensions. After the squeezeformer blocks, I have a prediction head of swish dense followed by a GLUMlp block (swiGLU followed by linear dense), both with head dimensions of 1024/2048, depending on the model. For more details, check my code.  <br>\nAt first, I used dropout layers, which helped greatly when the training data was in the order of 1M samples. However, after I saw comments in the forum about dropout being unnecessary, I experimented again when I scaled up to ~40M-80M samples, and it was indeed unnecessary when ensembling (although it still allows for higher scores for single models, at the cost of twice or thrice epochs). Since removing it allows for much faster training, I also chose to drop the dropout altogether. Optimizer was AdamW with half-cosine decay scheduler, max LR 1e-3, and weight decay = 4*LR.</p>\n<h3>1.2. Loss</h3>\n<p>I once read that the most important part of a DL model is the loss function. I was a bit skeptical then- sure, it is important, but it's not exactly complicated to choose the appropriate loss, right? Say, if our metric is R-squared, the best loss will obviously be…MSE?</p>\n<h4>1.2.1. Use MAE</h4>\n<p>It performs better than MSE in every way- it converges faster, is better at convergent, and is more stable. If I need to choose one 'secret' of this competition for a high score, aside from the notorious leaks, it is this.</p>\n<h4>1.2.2. Auxiliary loss</h4>\n<p>We had explicit spacetime data for the train set but not for the test set. Auxiliary loss is a common practice in such cases. I used MAE on the normalized latitude/longitude and the sin/cos on the day/year cycles (if it is unclear, please look at my code).<br>\nA side note on the auxiliary loss- some people speculated, in the last week of the competition, that using explicit spacetime data, even that of the training set, is considered a forbidden leak. So, first, such use was <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/519772#2919233\" target=\"_blank\">confirmed by the host to be legit</a>, as long as there is no 'hacking' of the test set (which can be done by using the auxiliary predictions to pseudo-label the test set- a thing I did not do). Second, in the last week, I trained several models without auxiliary spacetime loss, and my no-spacetime ensemble of 6 models got 0.79355/0.79092. So I win anyway; auxiliary spacetime loss is unnecessary and may not even be helpful (since my winning ensemble, while with a slightly better score, is also 13 models).  </p>\n<h4>1.2.3. Confidence head</h4>\n<p>The <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/460403\" target=\"_blank\">3rd place solution in Ribonanza</a> got a suspiciously good score for a particular model. I won't explain why it was suspicious; you will need to know Ribonanza's details for the context. In any case, this suspicious model had two unique things going for it- a strange architecture and a confidence head. I still need to experiment more with said strange architecture, but for the confidense head, it was easy to try in this competition and surprisingly effective. The idea is simple- I also predict the loss of each target. I also used MAE for the confidence loss. I have one model I trained without confidence head, and it got LB 0.78945/0.78631, so I think I would also won without it. Then again, a model with the seme specifications, but with confidense head was my best model with LB 0.79159/0.78869. So, yeah. Surprisingly effective. And yes, this second model can win 1st place in this competition by itself (0.78869 private, although not 1st in public).</p>\n<h4>1.2.4 Masked loss</h4>\n<p>I masked out the targets that are zeroed in the submission or those for which we use the ptend trick (ptend_q0002_2-ptend_q0002_26).  <br>\nFor those new at LEAP- please read <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/502484\" target=\"_blank\">this post</a> for ptend trick context.  </p>\n<p>Final thoughts on the loss function: well, it is probably the most important part of a deep learning model. lol.  </p>\n<h3>1.3. Data preparation</h3>\n<h4>1.3.1 High-res data</h4>\n<p>You may have noticed already that I used high-res data. Yes, it helps. Although I would have won without it- my only low-res data ensemble of five models got LB 0.79299/0.78951. But high-res definitely helped, although it sometimes requires a special soft-clipping treatment, as you will see. Also, training with high-res is less stable. Hence, I usually used larger batch sizes (1024/2048 compared to 512 for low-res only models). I blended high-res data with a ratio of 2:1 low: high. Lower than that, the performance is weaker; higher than that, the gains are small compared to the extra training time required.    </p>\n<h4>1.3.2 Multiple data representation</h4>\n<p>Multiple data representations is a trick I learned from 1st solution <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/434485\" target=\"_blank\">at ASLFR</a>. Without going into too much detail, the situation in ASLFR was that the data could be normalized in two different ways. I remember trying both ways, finding out what is better and sticking with it. Then the competition ended, and guess what? The 1st place normalized in the two possible ways, concatenated the two representations (with a few extra steps in between; read their summary for the full details) and sent it to the model.  <br>\nFirst, Let me separate the features that are spread over the 60-height level, which I call X_col and the features that are the same for all the levels, which I call X_col_not.  <br>\nFor X_col_not, I used only one representation: the simple normalization (x-mean)/std.   <br>\nFor X_col, I used three representations. The first the same as X_col_not, where I normalize each feature in each level with its own mean/std. i.e., for state_t, then we have (state_t_1-mean(state_t_1))/std(state_t_1), (state_t_2-mean(state_t_2))/std(state_t_2) etc. In my code, I call this representation x_col_not_norm (for x_col_not) and x_col_norm (for X_col).  <br>\nThe second representation normalizes each feature by the total mean and std over all the levels. For state_t, we have (state_t_1-mean(state_t))/std(state_t), (state_t_2-mean(state_t))/std(state_t), etc. In my code, I call this representation x_total_norm.  <br>\nFinally, the third representation is:  </p>\n<pre><code>x_col_norm_log = tf((x_col_norm-x_col_norm_min+)&gt;=, tf(x_col_norm-x_col_norm_min+),\n                                    -tf(+(x_col_norm-x_col_norm_min+)))\n</code></pre>\n<p>This is the kind of thing that trying to explain with words would never be as clear as just looking at the code.  <br>\nAfter you looked at the code and understood it, you may wonder why not use:  </p>\n<pre><code> = tf.math.log(x_col_norm-x_col_norm_min+)\n</code></pre>\n<p>Why all the extra steps with the tf.cond? See, I had a problem. I calculated x_col_norm_min only with Kaggle data, and then I scaled up my code to all HF data (which have values lower than Kaggle data x_col_norm_min) but did not want to change the normalization constant because it would break the inference pipeline. Then, I would have to use different pipelines for my old and new models. Yeah, sometimes I'm a bit lazy. Proud of it. And it turned out to be an excellent choice when I included also high-res data.  </p>\n<h4>1.3.3 Wind</h4>\n<p>$$wind = \\sqrt{(state_u)^2+(state_v)^2}   $$<br>\nIt made sense, and I wanted to include at least one 'physically justified' thing in the model (silly me, yes). It did not really help but also did not hurt the model, so it stayed. I used only the first normalization for WIND, with:  <br>\nmean(wind) = mean(mean(state_u), mean(state_v))  <br>\nAnd with:  <br>\nstd(wind) = sum(std(state_u), std(state_v)  </p>\n<p>All in all, the feature dimension is:  <br>\n9[col_features]*60[levels]*3[representations]+60[wind_levels]+16[not_col features] = 1696  </p>\n<h4>1.3.4 Features soft clipping 1</h4>\n<p>After normalization, the data has some extreme values (~±3000). This problem exists only for the first representation. The model actually handled it easily, but I preferred to play on the safe side. So, for x_col_norm and WIND, I applied the following soft clipping:  </p>\n<pre><code> = \n = cut**\n = tf.where(x_col_norm&gt;cut, x_col_norm**+cut-square_cut, x_col_norm)\n = tf.where(x_col_norm&lt;-cut, -tf.math.abs(x_col_norm)**-cut+square_cut, x_col_norm)\n</code></pre>\n<h4>1.3.5 Features soft clipping 2</h4>\n<p>In addition to the first soft clipping, I applied a second soft clipping to deal with extreme values from the high-res set. I applied the clipping on all the representations, including WIND, and after I applied the first soft clipping (1.3.3) for the relevant features:  </p>\n<pre><code> = \n = tf.math.log(cut_2)\n = tf.where(x_col_norm&gt;cut_2, tf.math.log(x_col_norm)+cut_2-log_cut, x_col_norm)\n = tf.where(x_col_norm&lt;-cut_2, -tf.math.log(-x_col_norm)-cut_2+log_cut, x_col_norm)\n</code></pre>\n<p>I chose cutoff_2 so that the soft clipping would only affect high-res data.  </p>\n<h4>1.3.6 Targets soft clipping</h4>\n<p>I did this only for the high-res targets; each target was soft-clipped if it was too extreme compared to the low-res corresponding target min/max (this differs from 1.3.4/1.3.5 in that the soft-clipping range was different for each target). In code:  </p>\n<pre><code>rescale_factor = \n x == :\n    y_norm = tf(y_norm&lt;norm_y_min*rescale_factor, norm_y_min*rescale_factor-tf(-y_norm+norm_y_min*rescale_factor), y_norm)\n    y_norm = tf(y_norm&gt;norm_y_max*rescale_factor, norm_y_max*rescale_factor+tf(+y_norm-norm_y_max*rescale_factor), y_norm)\n</code></pre>\n<h3>1.4 Post-processing</h3>\n<h4>1.4.1. Downcast and Upcast</h4>\n<p>Special care should be taken when moving between FP64 and FP32. I encoded my TFRecords with the original values in FP64. After I processed the data (see 1.3. Except for 1.3.5 which happened after the casting for no particular reason), the values were downcast to FP32 before transferring them to the model. Then I upcasted the predictions to FP64, and only then did I apply de-normalization to get the values for submission:  </p>\n<pre><code> = preds + mean_y.reshape(,-)*stds\n[:, np.where(stds_new == )] = \n = preds/np.where(stds&gt;, stds, )\n</code></pre>\n<h4>1.4.2. Mean for bad targets</h4>\n<p>This is very simple:  </p>\n<pre><code>metrics = np(, preds)    ()])\n   ((metrics)):\n     metrics&lt;:\n        preds = \n</code></pre>\n<p>In reality, it was eventually unnecessary because the only bad targets were those zeroed out in the submission or in the ptend trick range (see next bullet, 1.4.3).</p>\n<h4>1.4.3. Ptend trick</h4>\n<p>Obviously. If you are new at LEAP, look <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/502484\" target=\"_blank\">here for details</a>.  </p>\n<h3>1.5 Validation</h3>\n<p>My local validation included two validation sets of randomly selected samples from the low-res dataset (the same samples for all the models), each with a size of 100K samples. In addition, I randomly selected 100K samples from the high-res dataset and 100K samples from the low-res training set (i.e., the last validation set is on samples I also train on to judge overfitting better). All in all, I had four validation sets.  <br>\nWhen I trained on ~1M samples, the local validation was highly correlated with the public leaderboard score, but when I scaled up to all the data, it was less correlated (probably because when I used the full train set, it includes samples that are very close in time to the local validation samples and with the same latitude/longitude, so it starts to overfit even on the validation set, as opposed to the public test set that includes samples from entirely different year than those that exist in the train set). For example, I could push my local validation score to ~0.8, but on the public LB, it will get ~0.785 and be lower in score than a model that achieved, say, 0.795 but trained for fewer epochs. So, for the final models trained on all the low-res data from HF (and those I trained also on the high-res data), I validated against the public LB to get the hang of good overfitting range, and ensembling validation was done directly against the public LB. Thus, my ensemble is slightly overfitted to the public LB, although I acted in ways that reduce said overfitting (equal blending weights and including 'less successful' models if I judged them to be successful enough and anticipated them to be good based on the parameters I used). If it sounds a bit like black magic- yes, it is! Deep learning is sometimes a science, sometimes an art.  </p>\n<h2>2. Details of the submission</h2>\n<h3>2.1 Ensembling</h3>\n<p>My winning ensemble included 13 models, each a bit different (see full details in my GitHub, 'The steps to reproduce my solution' bullet 5). The best model (11) was LB 0.79159/0.78869, and the worst (2) was LB 0.78795/0.78388. Both were best/worst both in public and private LB. The full ensemble was LB 0.79410/0.79123. In addition, my low-res-data-only ensemble of 5 models has LB 0.79299/0.78951, and my no-spacetime-auxilliary-loss ensemble of 6 models has 0.79355/0.79092. When you read my solution, you may have tried to find out the 'secret sauce' that got me the 1st place, but it really was the combination that made the difference. Every single technique that I used, I think I could still get to 1st place without it.  </p>\n<h3>2.2 The helpful techniques</h3>\n<p>This is a short summary of the methods I wrote about in-depth above: Squeeseformer, wide GLUMlp prediction head, no dropout, MAE, auxiliary spacetime loss, confidence head, masked loss, multiple data representation, high-res data, features and targets soft-clipping and careful downcast/upcast.</p>\n<h3>2.3 What didn't work</h3>\n<p>Various model architectures (pure transformer, other 1Dconv/transformer combinations, Unet, dropout, smaller models, larger models, other optimizers), in short, many less optimal hyper-parameters. Log-normalization (i.e., log(x), not my log(1+x) representation which deals with different issues). MSE, MSE/MAE various combinations, weighted loss function (check my Ribonanza solution for details), levels/features masking, predict the year as additional auxiliary loss, and probably other not-very-important things that I don't remember already.  </p>\n<h3>2.4 Hardware</h3>\n<p>I trained on Kaggle and Colab TPU and inferred on Kaggle P100 GPU. Compute-wise, my experiments and training cost at least 200 bucks in Colab compute units and probably less than 300 bucks in total. If it sounds like a lot compared to your 'free' personal machine, please consider the electricity cost of using an RTX4090…</p>\n<h3>2.5 Bonus: confidence is key</h3>\n<p>Since I already had confidence predictions simply because the confidence head improved the model, I took a step further out of curiosity and checked if I could get a higher score by discarding 'low-confidence' samples and by how much I could improve. Here is a graph (only for one model, not the full ensemble) that shows the R-squared as a function of the fraction of remaining samples after I discarded the most 'low-confidence' ones.  <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6832115%2F3133894fc48f275cf79b88cee8a425aa%2Fconfidence_is_key.png?generation=1722304121419960&amp;alt=media\" alt=\"\"><br>\nAs you can see, by discarding ~0.1 of the samples, I can already get to a ~0.83+ score. This is not very surprising: if you check the data, you will see that certain months and locations have much lower scores, so a higher score can simply be achieved by discarding said months/location. However, discarding by confidence may be more precise and efficient. It can be interesting to research further, especially given this specific problem, since we can replace the simulator calculation with model predictions and then send back the low-confidence samples to the simulator.  <br>\nAs a side note, I also tried ensembling by confidence (i.e., giving lower weights to low-confidence predictions), but my preliminary research did not show promising results. Although it WAS preliminary- I got to it only on sunday of the last week of the competition, and it was either researching it more in depth or watching the Euro 2024 finals…(congrats Spain!)  </p>\n<h2>3. Sources</h2>\n<p>This is my favorite part. Thank you all the resources that have helped me!  <br>\n<a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/502484\" target=\"_blank\">Ptend trick</a> and also <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/499896#2791290\" target=\"_blank\">here originally</a> by <a href=\"https://www.kaggle.com/sakvaua\" target=\"_blank\">@sakvaua</a> and <a href=\"https://www.kaggle.com/jano123\" target=\"_blank\">@jano123</a> .<br>\n<a href=\"https://github.com/hoyso48/Stanford---Ribonanza-RNA-Folding-2nd-place-solution\" target=\"_blank\">Ribonanza 2nd solution</a> for Squeezeformer architecture guidance and insights, by <a href=\"https://www.kaggle.com/hoyso48\" target=\"_blank\">@hoyso48</a> .  <br>\n<a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/460403\" target=\"_blank\">Ribonanza 3rd place solution</a> for confidence head method, by <a href=\"https://www.kaggle.com/dankrstev\" target=\"_blank\">@dankrstev</a> .  <br>\n<a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/434485\" target=\"_blank\">ASLFR 1st solution</a> for the multiple data representation method, by <a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a> and <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> .  <br>\n<a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/514020#2884414\" target=\"_blank\">Dropout is unnecessary</a> and <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/514020#2885043\" target=\"_blank\">also here</a> by <a href=\"https://www.kaggle.com/sakvaua\" target=\"_blank\">@sakvaua</a> and <a href=\"https://www.kaggle.com/ymatioun\" target=\"_blank\">@ymatioun</a> .   <br>\nIn addition, I used the low-res and high-res data <a href=\"https://huggingface.co/LEAP\" target=\"_blank\">from HF</a>.  <br>\nAnd, of course, <a href=\"https://github.com/shlomoron/LEAP-solution\" target=\"_blank\">my GitHub</a> with links and instructions for a fully reproducible solution on a Kaggle pipeline.</p>",
      "rawMarkdown": "This competition was an amazing experience and a wonderful sandbox to test ideas and experiments with various techniques and architectures in-depth. Thanks to Kaggle for this fantastic platform, to Colab for providing me cheap TPU to train on, to Google as the owner of both Kaggle and Colab and for providing us with Tensorflow. When I think about it, I trained on TPU (google) on tensorflow framework (google) on Colab (google) for a competition held on Kaggle (google). It's all Google from start to end. You are amazing, and I salute you.   \n\nI also want to express my heartfelt gratitude to the competition host @jerrylin96 . Your efforts in providing us with this amazing competition and your unwavering commitment to keeping it on track, even when facing unexpected LEAK problems, are truly commendable. Thank you for your dedication and hard work.  \n\nFinally, thank you to the Kaggle community, especially to all the active and helpful people on the forum and in the code section. You make the experience so much better. I love you.  \n\nThis competition was a strange one. I built the best models, but there were better Kagglers than me- people who found the leak(s) and succeeded in utilizing them. I did not. I still have much to learn. I won't pretend I'm unhappy that things turned out the way they did, but the best Kagglers deserve respect. You have mine.\n\nI know that as 1st place, and considering the significant gap in scores between my solution and the next ones, a lot of eyes would be on me to ensure I did not exploit any LEAK. Hence, I made a lot of effort to provide a complete Kaggle pipeline, from downloading the data from HF through TFRecords encoding, training, inferring, and submission, with a training example that resulted in a single model 0.79081/0.78811 public/private LB. By following my pipeline step by step, you can also construct the dataset and train your own ~0.79+ public LB model from scratch, ensuring no hiding of leaky pseudo-labels in the train set or nefarious test-set reverse engineering in inference. All of this by simply copying and running a series of Kaggle notebooks. Admittedly, there are a lot of notebooks- over 100 notebooks if you intend to download and encode to TFRecords all the data yourself- but with the massive amount of data in this competition, there is no way around it. See the complete pipeline and extra details [in my GitHub](https://github.com/shlomoron/LEAP-solution).\n## Context section\n[Business context](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim).  \n[Data context](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/data).  \n## 1. Overview of the Approach  \n### 1.1. model\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6832115%2Fb55d82c6ef742f31b4c560b7bc396d51%2Fsummary_graphics.png?generation=1722307603587353&alt=media)\nIn one word, squeezeformer. This would not surprise anyone who followed some of the more similar competitions on Kaggle in the last year, in particular [ASLFR](https://www.kaggle.com/competitions/asl-fingerspelling) and [Ribonanza](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding). I used the same modified squeezeformer blocks that I saw in the [2nd solution to Ribonanza](https://github.com/hoyso48/Stanford---Ribonanza-RNA-Folding-2nd-place-solution) by @hoyso48 . The only changes I remember doing is that I deleted the first LayerNorm in the 1Dconv block and added an ECA layer. However, there may have been more minor changes that I forgot. My Tensorflow implementation was guided by hoyso48 PyTorch implementation, so once again (3rd competition in a row), I give him my thanks.  \nI used 12 block models with dimensions of 256/384/512. Before the squeezeformer blocks, I have a linear dense layer followed by LayerNorm as an encoder from the input data to the model dimensions. After the squeezeformer blocks, I have a prediction head of swish dense followed by a GLUMlp block (swiGLU followed by linear dense), both with head dimensions of 1024/2048, depending on the model. For more details, check my code.  \nAt first, I used dropout layers, which helped greatly when the training data was in the order of 1M samples. However, after I saw comments in the forum about dropout being unnecessary, I experimented again when I scaled up to ~40M-80M samples, and it was indeed unnecessary when ensembling (although it still allows for higher scores for single models, at the cost of twice or thrice epochs). Since removing it allows for much faster training, I also chose to drop the dropout altogether. Optimizer was AdamW with half-cosine decay scheduler, max LR 1e-3, and weight decay = 4*LR.\n### 1.2. Loss  \nI once read that the most important part of a DL model is the loss function. I was a bit skeptical then- sure, it is important, but it's not exactly complicated to choose the appropriate loss, right? Say, if our metric is R-squared, the best loss will obviously be...MSE?\n#### 1.2.1. Use MAE  \nIt performs better than MSE in every way- it converges faster, is better at convergent, and is more stable. If I need to choose one 'secret' of this competition for a high score, aside from the notorious leaks, it is this.\n#### 1.2.2. Auxiliary loss  \nWe had explicit spacetime data for the train set but not for the test set. Auxiliary loss is a common practice in such cases. I used MAE on the normalized latitude/longitude and the sin/cos on the day/year cycles (if it is unclear, please look at my code).\nA side note on the auxiliary loss- some people speculated, in the last week of the competition, that using explicit spacetime data, even that of the training set, is considered a forbidden leak. So, first, such use was [confirmed by the host to be legit](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/519772#2919233), as long as there is no 'hacking' of the test set (which can be done by using the auxiliary predictions to pseudo-label the test set- a thing I did not do). Second, in the last week, I trained several models without auxiliary spacetime loss, and my no-spacetime ensemble of 6 models got 0.79355/0.79092. So I win anyway; auxiliary spacetime loss is unnecessary and may not even be helpful (since my winning ensemble, while with a slightly better score, is also 13 models).  \n#### 1.2.3. Confidence head  \nThe [3rd place solution in Ribonanza](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/460403) got a suspiciously good score for a particular model. I won't explain why it was suspicious; you will need to know Ribonanza's details for the context. In any case, this suspicious model had two unique things going for it- a strange architecture and a confidence head. I still need to experiment more with said strange architecture, but for the confidense head, it was easy to try in this competition and surprisingly effective. The idea is simple- I also predict the loss of each target. I also used MAE for the confidence loss. I have one model I trained without confidence head, and it got LB 0.78945/0.78631, so I think I would also won without it. Then again, a model with the seme specifications, but with confidense head was my best model with LB 0.79159/0.78869. So, yeah. Surprisingly effective. And yes, this second model can win 1st place in this competition by itself (0.78869 private, although not 1st in public).\n#### 1.2.4 Masked loss  \nI masked out the targets that are zeroed in the submission or those for which we use the ptend trick (ptend_q0002_2-ptend_q0002_26).  \nFor those new at LEAP- please read [this post](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/502484) for ptend trick context.  \n\nFinal thoughts on the loss function: well, it is probably the most important part of a deep learning model. lol.  \n### 1.3. Data preparation\n#### 1.3.1 High-res data  \nYou may have noticed already that I used high-res data. Yes, it helps. Although I would have won without it- my only low-res data ensemble of five models got LB 0.79299/0.78951. But high-res definitely helped, although it sometimes requires a special soft-clipping treatment, as you will see. Also, training with high-res is less stable. Hence, I usually used larger batch sizes (1024/2048 compared to 512 for low-res only models). I blended high-res data with a ratio of 2:1 low: high. Lower than that, the performance is weaker; higher than that, the gains are small compared to the extra training time required.    \n#### 1.3.2 Multiple data representation\nMultiple data representations is a trick I learned from 1st solution [at ASLFR](https://www.kaggle.com/competitions/asl-fingerspelling/discussion/434485). Without going into too much detail, the situation in ASLFR was that the data could be normalized in two different ways. I remember trying both ways, finding out what is better and sticking with it. Then the competition ended, and guess what? The 1st place normalized in the two possible ways, concatenated the two representations (with a few extra steps in between; read their summary for the full details) and sent it to the model.  \nFirst, Let me separate the features that are spread over the 60-height level, which I call X_col and the features that are the same for all the levels, which I call X_col_not.  \nFor X_col_not, I used only one representation: the simple normalization (x-mean)/std.   \nFor X_col, I used three representations. The first the same as X_col_not, where I normalize each feature in each level with its own mean/std. i.e., for state_t, then we have (state_t_1-mean(state_t_1))/std(state_t_1), (state_t_2-mean(state_t_2))/std(state_t_2) etc. In my code, I call this representation x_col_not_norm (for x_col_not) and x_col_norm (for X_col).  \nThe second representation normalizes each feature by the total mean and std over all the levels. For state_t, we have (state_t_1-mean(state_t))/std(state_t), (state_t_2-mean(state_t))/std(state_t), etc. In my code, I call this representation x_total_norm.  \nFinally, the third representation is:  \n\n```  \nx_col_norm_log = tf.where((x_col_norm-x_col_norm_min+1)>=1, tf.math.log(x_col_norm-x_col_norm_min+1),\n                                    -tf.math.log(1+1-(x_col_norm-x_col_norm_min+1)))\n```\n\nThis is the kind of thing that trying to explain with words would never be as clear as just looking at the code.  \nAfter you looked at the code and understood it, you may wonder why not use:  \n\n```  \nx_col_norm_log = tf.math.log(x_col_norm-x_col_norm_min+1)\n```\n\nWhy all the extra steps with the tf.cond? See, I had a problem. I calculated x_col_norm_min only with Kaggle data, and then I scaled up my code to all HF data (which have values lower than Kaggle data x_col_norm_min) but did not want to change the normalization constant because it would break the inference pipeline. Then, I would have to use different pipelines for my old and new models. Yeah, sometimes I'm a bit lazy. Proud of it. And it turned out to be an excellent choice when I included also high-res data.  \n#### 1.3.3 Wind\n$$wind = \\sqrt{(state_u)^2+(state_v)^2}   $$\nIt made sense, and I wanted to include at least one 'physically justified' thing in the model (silly me, yes). It did not really help but also did not hurt the model, so it stayed. I used only the first normalization for WIND, with:  \nmean(wind) = mean(mean(state_u), mean(state_v))  \nAnd with:  \nstd(wind) = sum(std(state_u), std(state_v)  \n\nAll in all, the feature dimension is:  \n9[col_features]\\*60[levels]\\*3[representations]+60[wind_levels]+16[not_col features] = 1696  \n#### 1.3.4 Features soft clipping 1  \nAfter normalization, the data has some extreme values (~±3000). This problem exists only for the first representation. The model actually handled it easily, but I preferred to play on the safe side. So, for x_col_norm and WIND, I applied the following soft clipping:  \n\n```  \ncutoff = 30\nsquare_cutoff = cutoff**0.5\nx_col_norm = tf.where(x_col_norm>cutoff, x_col_norm**0.5+cutoff-square_cutoff, x_col_norm)\nx_col_norm = tf.where(x_col_norm<-cutoff, -tf.math.abs(x_col_norm)**0.5-cutoff+square_cutoff, x_col_norm)\n```\n\n#### 1.3.5 Features soft clipping 2  \nIn addition to the first soft clipping, I applied a second soft clipping to deal with extreme values from the high-res set. I applied the clipping on all the representations, including WIND, and after I applied the first soft clipping (1.3.3) for the relevant features:  \n\n```\ncutoff_2 = 86.0\nlog_cutoff = tf.math.log(cutoff_2)\nx_col_norm = tf.where(x_col_norm>cutoff_2, tf.math.log(x_col_norm)+cutoff_2-log_cutoff, x_col_norm)\nx_col_norm = tf.where(x_col_norm<-cutoff_2, -tf.math.log(-x_col_norm)-cutoff_2+log_cutoff, x_col_norm)\n```\n\nI chose cutoff_2 so that the soft clipping would only affect high-res data.  \n#### 1.3.6 Targets soft clipping  \nI did this only for the high-res targets; each target was soft-clipped if it was too extreme compared to the low-res corresponding target min/max (this differs from 1.3.4/1.3.5 in that the soft-clipping range was different for each target). In code:  \n\n```  \nrescale_factor = 1.1\nif x['res'] == 0:\n    y_norm = tf.where(y_norm<norm_y_min*rescale_factor, norm_y_min*rescale_factor-tf.math.log(1-y_norm+norm_y_min*rescale_factor), y_norm)\n    y_norm = tf.where(y_norm>norm_y_max*rescale_factor, norm_y_max*rescale_factor+tf.math.log(1+y_norm-norm_y_max*rescale_factor), y_norm)\n```\n### 1.4 Post-processing  \n#### 1.4.1. Downcast and Upcast  \nSpecial care should be taken when moving between FP64 and FP32. I encoded my TFRecords with the original values in FP64. After I processed the data (see 1.3. Except for 1.3.5 which happened after the casting for no particular reason), the values were downcast to FP32 before transferring them to the model. Then I upcasted the predictions to FP64, and only then did I apply de-normalization to get the values for submission:  \n\n```\npreds = preds + mean_y.reshape(1,-1)*stds\npreds[:, np.where(stds_new == 0)] = 0\npreds = preds/np.where(stds>0, stds, 1)\n```\n\n#### 1.4.2. Mean for bad targets  \nThis is very simple:  \n\n```\nmetrics = np.asarray([sklearn.metrics.r2_score(val_labels[:, i], preds[:, i]) for i in range(368)])\nfor i in range(len(metrics)):\n    if metrics[i]<0:\n        preds[:,i] = 0\n```\n\nIn reality, it was eventually unnecessary because the only bad targets were those zeroed out in the submission or in the ptend trick range (see next bullet, 1.4.3).\n\n#### 1.4.3. Ptend trick  \nObviously. If you are new at LEAP, look [here for details](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/502484).  \n\n### 1.5 Validation  \nMy local validation included two validation sets of randomly selected samples from the low-res dataset (the same samples for all the models), each with a size of 100K samples. In addition, I randomly selected 100K samples from the high-res dataset and 100K samples from the low-res training set (i.e., the last validation set is on samples I also train on to judge overfitting better). All in all, I had four validation sets.  \nWhen I trained on ~1M samples, the local validation was highly correlated with the public leaderboard score, but when I scaled up to all the data, it was less correlated (probably because when I used the full train set, it includes samples that are very close in time to the local validation samples and with the same latitude/longitude, so it starts to overfit even on the validation set, as opposed to the public test set that includes samples from entirely different year than those that exist in the train set). For example, I could push my local validation score to ~0.8, but on the public LB, it will get ~0.785 and be lower in score than a model that achieved, say, 0.795 but trained for fewer epochs. So, for the final models trained on all the low-res data from HF (and those I trained also on the high-res data), I validated against the public LB to get the hang of good overfitting range, and ensembling validation was done directly against the public LB. Thus, my ensemble is slightly overfitted to the public LB, although I acted in ways that reduce said overfitting (equal blending weights and including 'less successful' models if I judged them to be successful enough and anticipated them to be good based on the parameters I used). If it sounds a bit like black magic- yes, it is! Deep learning is sometimes a science, sometimes an art.  \n\n## 2. Details of the submission  \n### 2.1 Ensembling  \nMy winning ensemble included 13 models, each a bit different (see full details in my GitHub, 'The steps to reproduce my solution' bullet 5). The best model (11) was LB 0.79159/0.78869, and the worst (2) was LB 0.78795/0.78388. Both were best/worst both in public and private LB. The full ensemble was LB 0.79410/0.79123. In addition, my low-res-data-only ensemble of 5 models has LB 0.79299/0.78951, and my no-spacetime-auxilliary-loss ensemble of 6 models has 0.79355/0.79092. When you read my solution, you may have tried to find out the 'secret sauce' that got me the 1st place, but it really was the combination that made the difference. Every single technique that I used, I think I could still get to 1st place without it.  \n### 2.2 The helpful techniques\nThis is a short summary of the methods I wrote about in-depth above: Squeeseformer, wide GLUMlp prediction head, no dropout, MAE, auxiliary spacetime loss, confidence head, masked loss, multiple data representation, high-res data, features and targets soft-clipping and careful downcast/upcast.\n### 2.3 What didn't work  \nVarious model architectures (pure transformer, other 1Dconv/transformer combinations, Unet, dropout, smaller models, larger models, other optimizers), in short, many less optimal hyper-parameters. Log-normalization (i.e., log(x), not my log(1+x) representation which deals with different issues). MSE, MSE/MAE various combinations, weighted loss function (check my Ribonanza solution for details), levels/features masking, predict the year as additional auxiliary loss, and probably other not-very-important things that I don't remember already.  \n### 2.4 Hardware\nI trained on Kaggle and Colab TPU and inferred on Kaggle P100 GPU. Compute-wise, my experiments and training cost at least 200 bucks in Colab compute units and probably less than 300 bucks in total. If it sounds like a lot compared to your 'free' personal machine, please consider the electricity cost of using an RTX4090...\n### 2.5 Bonus: confidence is key\nSince I already had confidence predictions simply because the confidence head improved the model, I took a step further out of curiosity and checked if I could get a higher score by discarding 'low-confidence' samples and by how much I could improve. Here is a graph (only for one model, not the full ensemble) that shows the R-squared as a function of the fraction of remaining samples after I discarded the most 'low-confidence' ones.  \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6832115%2F3133894fc48f275cf79b88cee8a425aa%2Fconfidence_is_key.png?generation=1722304121419960&alt=media)\nAs you can see, by discarding ~0.1 of the samples, I can already get to a ~0.83+ score. This is not very surprising: if you check the data, you will see that certain months and locations have much lower scores, so a higher score can simply be achieved by discarding said months/location. However, discarding by confidence may be more precise and efficient. It can be interesting to research further, especially given this specific problem, since we can replace the simulator calculation with model predictions and then send back the low-confidence samples to the simulator.  \nAs a side note, I also tried ensembling by confidence (i.e., giving lower weights to low-confidence predictions), but my preliminary research did not show promising results. Although it WAS preliminary- I got to it only on sunday of the last week of the competition, and it was either researching it more in depth or watching the Euro 2024 finals...(congrats Spain!)  \n\n## 3. Sources\nThis is my favorite part. Thank you all the resources that have helped me!  \n[Ptend trick](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/502484) and also [here originally](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/499896#2791290) by @sakvaua and @jano123 .\n[Ribonanza 2nd solution](https://github.com/hoyso48/Stanford---Ribonanza-RNA-Folding-2nd-place-solution) for Squeezeformer architecture guidance and insights, by @hoyso48 .  \n[Ribonanza 3rd place solution](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/460403) for confidence head method, by @dankrstev .  \n[ASLFR 1st solution](https://www.kaggle.com/competitions/asl-fingerspelling/discussion/434485) for the multiple data representation method, by @darraghdog and @christofhenkel .  \n[Dropout is unnecessary](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/514020#2884414) and [also here](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/514020#2885043) by @sakvaua and @ymatioun .   \nIn addition, I used the low-res and high-res data [from HF](https://huggingface.co/LEAP).  \nAnd, of course, [my GitHub](https://github.com/shlomoron/LEAP-solution) with links and instructions for a fully reproducible solution on a Kaggle pipeline.",
      "votes": 82
    },
    {
      "id": 2940524,
      "postDate": "2024-07-30T06:48:57.120Z",
      "content": "<p>Well done, well deserved.</p>",
      "rawMarkdown": "Well done, well deserved.",
      "votes": 3
    },
    {
      "id": 2940414,
      "postDate": "2024-07-30T04:26:31.003Z",
      "content": "<p>Congratulations on winning the championship, you deserve it!😀<br>\nThank you for your writing; the solution is fascinating, and I learned a lot from it.</p>",
      "rawMarkdown": "Congratulations on winning the championship, you deserve it!😀\nThank you for your writing; the solution is fascinating, and I learned a lot from it.\n",
      "votes": 3
    },
    {
      "id": 2944303,
      "postDate": "2024-08-02T09:47:18.573Z",
      "content": "<p>Well done very inspiring !</p>",
      "rawMarkdown": "Well done very inspiring !",
      "votes": 1
    },
    {
      "id": 2943762,
      "postDate": "2024-08-01T20:20:47.213Z",
      "content": "<p>Congrats and thank you for the code! I've had a chance to tinker with your training code today and I'm looking forward to swapping out various pieces with my work as a learning exercise.</p>",
      "rawMarkdown": "Congrats and thank you for the code! I've had a chance to tinker with your training code today and I'm looking forward to swapping out various pieces with my work as a learning exercise.",
      "votes": 1
    },
    {
      "id": 2943119,
      "postDate": "2024-08-01T11:03:16.773Z",
      "content": "<p>Congratulations on winning the championship, you deserve it!!!</p>",
      "rawMarkdown": "Congratulations on winning the championship, you deserve it!!!",
      "votes": 1
    },
    {
      "id": 2942587,
      "postDate": "2024-07-31T23:48:47.880Z",
      "content": "<p>after seeing all the work that went into this, I am now surprised I managed to get into the top 100 😂 got a lot to learn. Congratulations and thanks for sharing!</p>",
      "rawMarkdown": "after seeing all the work that went into this, I am now surprised I managed to get into the top 100 😂 got a lot to learn. Congratulations and thanks for sharing!",
      "votes": 1
    },
    {
      "id": 2942417,
      "postDate": "2024-07-31T19:12:05.883Z",
      "content": "<p>Amazing work , keep publishing and congrats</p>",
      "rawMarkdown": "Amazing work , keep publishing and congrats",
      "votes": 1
    },
    {
      "id": 2942032,
      "postDate": "2024-07-31T14:50:16.900Z",
      "content": "<p>Well done for your work ! Thank you for sharing this detailed answer :) I am wondering what was the training time on TPU for your models with the full data from HF ?</p>",
      "rawMarkdown": "Well done for your work ! Thank you for sharing this detailed answer :) I am wondering what was the training time on TPU for your models with the full data from HF ?",
      "votes": 1,
      "replies": [
        {
          "id": 2942036,
          "postDate": "2024-07-31T14:54:47.253Z",
          "content": "<p>Check the training notebooks linked in my Github</p>",
          "rawMarkdown": "Check the training notebooks linked in my Github",
          "votes": 1,
          "replies": [
            {
              "id": 2942055,
              "postDate": "2024-07-31T15:12:13.867Z",
              "content": "<p>Perfect thanks ! 14800 seconds, 4 times</p>",
              "rawMarkdown": "Perfect thanks ! 14800 seconds, 4 times"
            }
          ]
        }
      ]
    },
    {
      "id": 2940478,
      "postDate": "2024-07-30T05:59:41.603Z",
      "content": "<blockquote>\n  <p>I won't explain why it was suspicious;</p>\n</blockquote>\n<p>thats enough, I get it :D </p>\n<p>thank you for publishing the solution and congratz with the win!</p>",
      "rawMarkdown": "> I won't explain why it was suspicious;\n\nthats enough, I get it :D \n\nthank you for publishing the solution and congratz with the win!",
      "votes": 2
    },
    {
      "id": 2956584,
      "postDate": "2024-08-12T08:44:19.700Z",
      "content": "<p><a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a> I guess the confidence head comes from the idea of this paper <a href=\"https://arxiv.org/pdf/1802.04865\" target=\"_blank\">https://arxiv.org/pdf/1802.04865</a></p>",
      "rawMarkdown": "@shlomoron I guess the confidence head comes from the idea of this paper https://arxiv.org/pdf/1802.04865",
      "replies": [
        {
          "id": 2956596,
          "postDate": "2024-08-12T08:57:31.540Z",
          "content": "<p>IDK, maybe? I took it from <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/460403\" target=\"_blank\">Ribonanza 3rd solution</a> and he took it from Alphafold.</p>",
          "rawMarkdown": "IDK, maybe? I took it from [Ribonanza 3rd solution](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/460403) and he took it from Alphafold.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2941857,
      "postDate": "2024-07-31T11:41:28.957Z",
      "content": "<p>Congrats to wining the Leep.</p>",
      "rawMarkdown": "Congrats to wining the Leep."
    },
    {
      "id": 2941503,
      "postDate": "2024-07-31T03:33:39.610Z",
      "content": "<p>Congratulations on winning the LEAP! Thank you as well for writing such a stunning article! <br>\nAdditionally, I'd like to point out that in the section on label processing, you employed the third transformation method.</p>",
      "rawMarkdown": "Congratulations on winning the LEAP! Thank you as well for writing such a stunning article! \nAdditionally, I'd like to point out that in the section on label processing, you employed the third transformation method."
    },
    {
      "id": 2940906,
      "postDate": "2024-07-30T14:38:19.170Z",
      "content": "<p>Congrats! Amazing work, a lot of work. Its hard for me to imagine how you managed so many experiments and variations. Nice idea the confidence loss, i was not aware of it and i dont get why it works. Also the different normalisations concept was new to me! Well deserved rank 1!</p>",
      "rawMarkdown": "Congrats! Amazing work, a lot of work. Its hard for me to imagine how you managed so many experiments and variations. Nice idea the confidence loss, i was not aware of it and i dont get why it works. Also the different normalisations concept was new to me! Well deserved rank 1!"
    },
    {
      "id": 2940843,
      "postDate": "2024-07-30T13:53:57.983Z",
      "content": "<p>Congratulations on winning the championship!</p>",
      "rawMarkdown": "Congratulations on winning the championship!\n"
    },
    {
      "id": 2940560,
      "postDate": "2024-07-30T07:52:47.710Z",
      "content": "<p>How did you load the data to colab? Did you just redownload the data to the session every time you started a colab instance? Asking because i currently use Colab for RSNA and everytime I start the notebook it takes 30 min to load the 35GB of data and I cant imagine how long it takes for the LEAP data</p>",
      "rawMarkdown": "How did you load the data to colab? Did you just redownload the data to the session every time you started a colab instance? Asking because i currently use Colab for RSNA and everytime I start the notebook it takes 30 min to load the 35GB of data and I cant imagine how long it takes for the LEAP data",
      "replies": [
        {
          "id": 2940633,
          "postDate": "2024-07-30T09:09:26.673Z",
          "content": "<p>Colab TPU works the same way as Kaggle TPU. Just feed them the path to TFRecords buckets. </p>",
          "rawMarkdown": "Colab TPU works the same way as Kaggle TPU. Just feed them the path to TFRecords buckets. "
        }
      ]
    },
    {
      "id": 2941192,
      "postDate": "2024-07-30T18:16:06.913Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing"
    }
  ],
  "comments": [
    {
      "id": 2940524,
      "author_name": "Yusef A.",
      "author_url": "",
      "post_date": "2024-07-30T06:48:57.120000",
      "content": "<p>Well done, well deserved.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2940414,
      "author_name": "Max2020",
      "author_url": "",
      "post_date": "2024-07-30T04:26:31.003000",
      "content": "<p>Congratulations on winning the championship, you deserve it!😀<br>\nThank you for your writing; the solution is fascinating, and I learned a lot from it.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2944303,
      "author_name": "Sahitya Setu",
      "author_url": "",
      "post_date": "2024-08-02T09:47:18.573000",
      "content": "<p>Well done very inspiring !</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2943762,
      "author_name": "Rob Freeman",
      "author_url": "",
      "post_date": "2024-08-01T20:20:47.213000",
      "content": "<p>Congrats and thank you for the code! I've had a chance to tinker with your training code today and I'm looking forward to swapping out various pieces with my work as a learning exercise.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2943119,
      "author_name": "LeoPolv",
      "author_url": "",
      "post_date": "2024-08-01T11:03:16.773000",
      "content": "<p>Congratulations on winning the championship, you deserve it!!!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2942587,
      "author_name": "Juan D C F",
      "author_url": "",
      "post_date": "2024-07-31T23:48:47.880000",
      "content": "<p>after seeing all the work that went into this, I am now surprised I managed to get into the top 100 😂 got a lot to learn. Congratulations and thanks for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2942417,
      "author_name": "Anish Vijay",
      "author_url": "",
      "post_date": "2024-07-31T19:12:05.883000",
      "content": "<p>Amazing work , keep publishing and congrats</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2942032,
      "author_name": "Albanito",
      "author_url": "",
      "post_date": "2024-07-31T14:50:16.900000",
      "content": "<p>Well done for your work ! Thank you for sharing this detailed answer :) I am wondering what was the training time on TPU for your models with the full data from HF ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2942036,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2024-07-31T14:54:47.253000",
          "content": "<p>Check the training notebooks linked in my Github</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2942055,
              "author_name": "Albanito",
              "author_url": "",
              "post_date": "2024-07-31T15:12:13.867000",
              "content": "<p>Perfect thanks ! 14800 seconds, 4 times</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2940478,
      "author_name": "slime",
      "author_url": "",
      "post_date": "2024-07-30T05:59:41.603000",
      "content": "<blockquote>\n  <p>I won't explain why it was suspicious;</p>\n</blockquote>\n<p>thats enough, I get it :D </p>\n<p>thank you for publishing the solution and congratz with the win!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2956584,
      "author_name": "ducnh279",
      "author_url": "",
      "post_date": "2024-08-12T08:44:19.700000",
      "content": "<p><a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a> I guess the confidence head comes from the idea of this paper <a href=\"https://arxiv.org/pdf/1802.04865\" target=\"_blank\">https://arxiv.org/pdf/1802.04865</a></p>",
      "votes": 0,
      "replies": [
        {
          "id": 2956596,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2024-08-12T08:57:31.540000",
          "content": "<p>IDK, maybe? I took it from <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/460403\" target=\"_blank\">Ribonanza 3rd solution</a> and he took it from Alphafold.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2941857,
      "author_name": "Umair Hayat",
      "author_url": "",
      "post_date": "2024-07-31T11:41:28.957000",
      "content": "<p>Congrats to wining the Leep.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2941503,
      "author_name": "Mingjie Wang",
      "author_url": "",
      "post_date": "2024-07-31T03:33:39.610000",
      "content": "<p>Congratulations on winning the LEAP! Thank you as well for writing such a stunning article! <br>\nAdditionally, I'd like to point out that in the section on label processing, you employed the third transformation method.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2940906,
      "author_name": "Vasilis",
      "author_url": "",
      "post_date": "2024-07-30T14:38:19.170000",
      "content": "<p>Congrats! Amazing work, a lot of work. Its hard for me to imagine how you managed so many experiments and variations. Nice idea the confidence loss, i was not aware of it and i dont get why it works. Also the different normalisations concept was new to me! Well deserved rank 1!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2940843,
      "author_name": "Ranamalla Nithin Reddy",
      "author_url": "",
      "post_date": "2024-07-30T13:53:57.983000",
      "content": "<p>Congratulations on winning the championship!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2940560,
      "author_name": "snehal",
      "author_url": "",
      "post_date": "2024-07-30T07:52:47.710000",
      "content": "<p>How did you load the data to colab? Did you just redownload the data to the session every time you started a colab instance? Asking because i currently use Colab for RSNA and everytime I start the notebook it takes 30 min to load the 35GB of data and I cant imagine how long it takes for the LEAP data</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2940633,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2024-07-30T09:09:26.673000",
          "content": "<p>Colab TPU works the same way as Kaggle TPU. Just feed them the path to TFRecords buckets. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2941192,
      "author_name": "Nima Abedi",
      "author_url": "",
      "post_date": "2024-07-30T18:16:06.913000",
      "content": "<p>Thanks for sharing</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2940353": "This competition was an amazing experience and a wonderful sandbox to test ideas and experiments with various techniques and architectures in-depth. Thanks to Kaggle for this fantastic platform, to Colab for providing me cheap TPU to train on, to Google as the owner of both Kaggle and Colab and for providing us with Tensorflow. When I think about it, I trained on TPU (google) on tensorflow framework (google) on Colab (google) for a competition held on Kaggle (google). It's all Google from start to end. You are amazing, and I salute you.   \n\nI also want to express my heartfelt gratitude to the competition host @jerrylin96 . Your efforts in providing us with this amazing competition and your unwavering commitment to keeping it on track, even when facing unexpected LEAK problems, are truly commendable. Thank you for your dedication and hard work.  \n\nFinally, thank you to the Kaggle community, especially to all the active and helpful people on the forum and in the code section. You make the experience so much better. I love you.  \n\nThis competition was a strange one. I built the best models, but there were better Kagglers than me- people who found the leak(s) and succeeded in utilizing them. I did not. I still have much to learn. I won't pretend I'm unhappy that things turned out the way they did, but the best Kagglers deserve respect. You have mine.\n\nI know that as 1st place, and considering the significant gap in scores between my solution and the next ones, a lot of eyes would be on me to ensure I did not exploit any LEAK. Hence, I made a lot of effort to provide a complete Kaggle pipeline, from downloading the data from HF through TFRecords encoding, training, inferring, and submission, with a training example that resulted in a single model 0.79081/0.78811 public/private LB. By following my pipeline step by step, you can also construct the dataset and train your own ~0.79+ public LB model from scratch, ensuring no hiding of leaky pseudo-labels in the train set or nefarious test-set reverse engineering in inference. All of this by simply copying and running a series of Kaggle notebooks. Admittedly, there are a lot of notebooks- over 100 notebooks if you intend to download and encode to TFRecords all the data yourself- but with the massive amount of data in this competition, there is no way around it. See the complete pipeline and extra details [in my GitHub](https://github.com/shlomoron/LEAP-solution).\n## Context section\n[Business context](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim).  \n[Data context](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/data).  \n## 1. Overview of the Approach  \n### 1.1. model\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6832115%2Fb55d82c6ef742f31b4c560b7bc396d51%2Fsummary_graphics.png?generation=1722307603587353&alt=media)\nIn one word, squeezeformer. This would not surprise anyone who followed some of the more similar competitions on Kaggle in the last year, in particular [ASLFR](https://www.kaggle.com/competitions/asl-fingerspelling) and [Ribonanza](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding). I used the same modified squeezeformer blocks that I saw in the [2nd solution to Ribonanza](https://github.com/hoyso48/Stanford---Ribonanza-RNA-Folding-2nd-place-solution) by @hoyso48 . The only changes I remember doing is that I deleted the first LayerNorm in the 1Dconv block and added an ECA layer. However, there may have been more minor changes that I forgot. My Tensorflow implementation was guided by hoyso48 PyTorch implementation, so once again (3rd competition in a row), I give him my thanks.  \nI used 12 block models with dimensions of 256/384/512. Before the squeezeformer blocks, I have a linear dense layer followed by LayerNorm as an encoder from the input data to the model dimensions. After the squeezeformer blocks, I have a prediction head of swish dense followed by a GLUMlp block (swiGLU followed by linear dense), both with head dimensions of 1024/2048, depending on the model. For more details, check my code.  \nAt first, I used dropout layers, which helped greatly when the training data was in the order of 1M samples. However, after I saw comments in the forum about dropout being unnecessary, I experimented again when I scaled up to ~40M-80M samples, and it was indeed unnecessary when ensembling (although it still allows for higher scores for single models, at the cost of twice or thrice epochs). Since removing it allows for much faster training, I also chose to drop the dropout altogether. Optimizer was AdamW with half-cosine decay scheduler, max LR 1e-3, and weight decay = 4*LR.\n### 1.2. Loss  \nI once read that the most important part of a DL model is the loss function. I was a bit skeptical then- sure, it is important, but it's not exactly complicated to choose the appropriate loss, right? Say, if our metric is R-squared, the best loss will obviously be...MSE?\n#### 1.2.1. Use MAE  \nIt performs better than MSE in every way- it converges faster, is better at convergent, and is more stable. If I need to choose one 'secret' of this competition for a high score, aside from the notorious leaks, it is this.\n#### 1.2.2. Auxiliary loss  \nWe had explicit spacetime data for the train set but not for the test set. Auxiliary loss is a common practice in such cases. I used MAE on the normalized latitude/longitude and the sin/cos on the day/year cycles (if it is unclear, please look at my code).\nA side note on the auxiliary loss- some people speculated, in the last week of the competition, that using explicit spacetime data, even that of the training set, is considered a forbidden leak. So, first, such use was [confirmed by the host to be legit](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/519772#2919233), as long as there is no 'hacking' of the test set (which can be done by using the auxiliary predictions to pseudo-label the test set- a thing I did not do). Second, in the last week, I trained several models without auxiliary spacetime loss, and my no-spacetime ensemble of 6 models got 0.79355/0.79092. So I win anyway; auxiliary spacetime loss is unnecessary and may not even be helpful (since my winning ensemble, while with a slightly better score, is also 13 models).  \n#### 1.2.3. Confidence head  \nThe [3rd place solution in Ribonanza](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/460403) got a suspiciously good score for a particular model. I won't explain why it was suspicious; you will need to know Ribonanza's details for the context. In any case, this suspicious model had two unique things going for it- a strange architecture and a confidence head. I still need to experiment more with said strange architecture, but for the confidense head, it was easy to try in this competition and surprisingly effective. The idea is simple- I also predict the loss of each target. I also used MAE for the confidence loss. I have one model I trained without confidence head, and it got LB 0.78945/0.78631, so I think I would also won without it. Then again, a model with the seme specifications, but with confidense head was my best model with LB 0.79159/0.78869. So, yeah. Surprisingly effective. And yes, this second model can win 1st place in this competition by itself (0.78869 private, although not 1st in public).\n#### 1.2.4 Masked loss  \nI masked out the targets that are zeroed in the submission or those for which we use the ptend trick (ptend_q0002_2-ptend_q0002_26).  \nFor those new at LEAP- please read [this post](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/502484) for ptend trick context.  \n\nFinal thoughts on the loss function: well, it is probably the most important part of a deep learning model. lol.  \n### 1.3. Data preparation\n#### 1.3.1 High-res data  \nYou may have noticed already that I used high-res data. Yes, it helps. Although I would have won without it- my only low-res data ensemble of five models got LB 0.79299/0.78951. But high-res definitely helped, although it sometimes requires a special soft-clipping treatment, as you will see. Also, training with high-res is less stable. Hence, I usually used larger batch sizes (1024/2048 compared to 512 for low-res only models). I blended high-res data with a ratio of 2:1 low: high. Lower than that, the performance is weaker; higher than that, the gains are small compared to the extra training time required.    \n#### 1.3.2 Multiple data representation\nMultiple data representations is a trick I learned from 1st solution [at ASLFR](https://www.kaggle.com/competitions/asl-fingerspelling/discussion/434485). Without going into too much detail, the situation in ASLFR was that the data could be normalized in two different ways. I remember trying both ways, finding out what is better and sticking with it. Then the competition ended, and guess what? The 1st place normalized in the two possible ways, concatenated the two representations (with a few extra steps in between; read their summary for the full details) and sent it to the model.  \nFirst, Let me separate the features that are spread over the 60-height level, which I call X_col and the features that are the same for all the levels, which I call X_col_not.  \nFor X_col_not, I used only one representation: the simple normalization (x-mean)/std.   \nFor X_col, I used three representations. The first the same as X_col_not, where I normalize each feature in each level with its own mean/std. i.e., for state_t, then we have (state_t_1-mean(state_t_1))/std(state_t_1), (state_t_2-mean(state_t_2))/std(state_t_2) etc. In my code, I call this representation x_col_not_norm (for x_col_not) and x_col_norm (for X_col).  \nThe second representation normalizes each feature by the total mean and std over all the levels. For state_t, we have (state_t_1-mean(state_t))/std(state_t), (state_t_2-mean(state_t))/std(state_t), etc. In my code, I call this representation x_total_norm.  \nFinally, the third representation is:  \n\n```  \nx_col_norm_log = tf.where((x_col_norm-x_col_norm_min+1)>=1, tf.math.log(x_col_norm-x_col_norm_min+1),\n                                    -tf.math.log(1+1-(x_col_norm-x_col_norm_min+1)))\n```\n\nThis is the kind of thing that trying to explain with words would never be as clear as just looking at the code.  \nAfter you looked at the code and understood it, you may wonder why not use:  \n\n```  \nx_col_norm_log = tf.math.log(x_col_norm-x_col_norm_min+1)\n```\n\nWhy all the extra steps with the tf.cond? See, I had a problem. I calculated x_col_norm_min only with Kaggle data, and then I scaled up my code to all HF data (which have values lower than Kaggle data x_col_norm_min) but did not want to change the normalization constant because it would break the inference pipeline. Then, I would have to use different pipelines for my old and new models. Yeah, sometimes I'm a bit lazy. Proud of it. And it turned out to be an excellent choice when I included also high-res data.  \n#### 1.3.3 Wind\n$$wind = \\sqrt{(state_u)^2+(state_v)^2}   $$\nIt made sense, and I wanted to include at least one 'physically justified' thing in the model (silly me, yes). It did not really help but also did not hurt the model, so it stayed. I used only the first normalization for WIND, with:  \nmean(wind) = mean(mean(state_u), mean(state_v))  \nAnd with:  \nstd(wind) = sum(std(state_u), std(state_v)  \n\nAll in all, the feature dimension is:  \n9[col_features]\\*60[levels]\\*3[representations]+60[wind_levels]+16[not_col features] = 1696  \n#### 1.3.4 Features soft clipping 1  \nAfter normalization, the data has some extreme values (~±3000). This problem exists only for the first representation. The model actually handled it easily, but I preferred to play on the safe side. So, for x_col_norm and WIND, I applied the following soft clipping:  \n\n```  \ncutoff = 30\nsquare_cutoff = cutoff**0.5\nx_col_norm = tf.where(x_col_norm>cutoff, x_col_norm**0.5+cutoff-square_cutoff, x_col_norm)\nx_col_norm = tf.where(x_col_norm<-cutoff, -tf.math.abs(x_col_norm)**0.5-cutoff+square_cutoff, x_col_norm)\n```\n\n#### 1.3.5 Features soft clipping 2  \nIn addition to the first soft clipping, I applied a second soft clipping to deal with extreme values from the high-res set. I applied the clipping on all the representations, including WIND, and after I applied the first soft clipping (1.3.3) for the relevant features:  \n\n```\ncutoff_2 = 86.0\nlog_cutoff = tf.math.log(cutoff_2)\nx_col_norm = tf.where(x_col_norm>cutoff_2, tf.math.log(x_col_norm)+cutoff_2-log_cutoff, x_col_norm)\nx_col_norm = tf.where(x_col_norm<-cutoff_2, -tf.math.log(-x_col_norm)-cutoff_2+log_cutoff, x_col_norm)\n```\n\nI chose cutoff_2 so that the soft clipping would only affect high-res data.  \n#### 1.3.6 Targets soft clipping  \nI did this only for the high-res targets; each target was soft-clipped if it was too extreme compared to the low-res corresponding target min/max (this differs from 1.3.4/1.3.5 in that the soft-clipping range was different for each target). In code:  \n\n```  \nrescale_factor = 1.1\nif x['res'] == 0:\n    y_norm = tf.where(y_norm<norm_y_min*rescale_factor, norm_y_min*rescale_factor-tf.math.log(1-y_norm+norm_y_min*rescale_factor), y_norm)\n    y_norm = tf.where(y_norm>norm_y_max*rescale_factor, norm_y_max*rescale_factor+tf.math.log(1+y_norm-norm_y_max*rescale_factor), y_norm)\n```\n### 1.4 Post-processing  \n#### 1.4.1. Downcast and Upcast  \nSpecial care should be taken when moving between FP64 and FP32. I encoded my TFRecords with the original values in FP64. After I processed the data (see 1.3. Except for 1.3.5 which happened after the casting for no particular reason), the values were downcast to FP32 before transferring them to the model. Then I upcasted the predictions to FP64, and only then did I apply de-normalization to get the values for submission:  \n\n```\npreds = preds + mean_y.reshape(1,-1)*stds\npreds[:, np.where(stds_new == 0)] = 0\npreds = preds/np.where(stds>0, stds, 1)\n```\n\n#### 1.4.2. Mean for bad targets  \nThis is very simple:  \n\n```\nmetrics = np.asarray([sklearn.metrics.r2_score(val_labels[:, i], preds[:, i]) for i in range(368)])\nfor i in range(len(metrics)):\n    if metrics[i]<0:\n        preds[:,i] = 0\n```\n\nIn reality, it was eventually unnecessary because the only bad targets were those zeroed out in the submission or in the ptend trick range (see next bullet, 1.4.3).\n\n#### 1.4.3. Ptend trick  \nObviously. If you are new at LEAP, look [here for details](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/502484).  \n\n### 1.5 Validation  \nMy local validation included two validation sets of randomly selected samples from the low-res dataset (the same samples for all the models), each with a size of 100K samples. In addition, I randomly selected 100K samples from the high-res dataset and 100K samples from the low-res training set (i.e., the last validation set is on samples I also train on to judge overfitting better). All in all, I had four validation sets.  \nWhen I trained on ~1M samples, the local validation was highly correlated with the public leaderboard score, but when I scaled up to all the data, it was less correlated (probably because when I used the full train set, it includes samples that are very close in time to the local validation samples and with the same latitude/longitude, so it starts to overfit even on the validation set, as opposed to the public test set that includes samples from entirely different year than those that exist in the train set). For example, I could push my local validation score to ~0.8, but on the public LB, it will get ~0.785 and be lower in score than a model that achieved, say, 0.795 but trained for fewer epochs. So, for the final models trained on all the low-res data from HF (and those I trained also on the high-res data), I validated against the public LB to get the hang of good overfitting range, and ensembling validation was done directly against the public LB. Thus, my ensemble is slightly overfitted to the public LB, although I acted in ways that reduce said overfitting (equal blending weights and including 'less successful' models if I judged them to be successful enough and anticipated them to be good based on the parameters I used). If it sounds a bit like black magic- yes, it is! Deep learning is sometimes a science, sometimes an art.  \n\n## 2. Details of the submission  \n### 2.1 Ensembling  \nMy winning ensemble included 13 models, each a bit different (see full details in my GitHub, 'The steps to reproduce my solution' bullet 5). The best model (11) was LB 0.79159/0.78869, and the worst (2) was LB 0.78795/0.78388. Both were best/worst both in public and private LB. The full ensemble was LB 0.79410/0.79123. In addition, my low-res-data-only ensemble of 5 models has LB 0.79299/0.78951, and my no-spacetime-auxilliary-loss ensemble of 6 models has 0.79355/0.79092. When you read my solution, you may have tried to find out the 'secret sauce' that got me the 1st place, but it really was the combination that made the difference. Every single technique that I used, I think I could still get to 1st place without it.  \n### 2.2 The helpful techniques\nThis is a short summary of the methods I wrote about in-depth above: Squeeseformer, wide GLUMlp prediction head, no dropout, MAE, auxiliary spacetime loss, confidence head, masked loss, multiple data representation, high-res data, features and targets soft-clipping and careful downcast/upcast.\n### 2.3 What didn't work  \nVarious model architectures (pure transformer, other 1Dconv/transformer combinations, Unet, dropout, smaller models, larger models, other optimizers), in short, many less optimal hyper-parameters. Log-normalization (i.e., log(x), not my log(1+x) representation which deals with different issues). MSE, MSE/MAE various combinations, weighted loss function (check my Ribonanza solution for details), levels/features masking, predict the year as additional auxiliary loss, and probably other not-very-important things that I don't remember already.  \n### 2.4 Hardware\nI trained on Kaggle and Colab TPU and inferred on Kaggle P100 GPU. Compute-wise, my experiments and training cost at least 200 bucks in Colab compute units and probably less than 300 bucks in total. If it sounds like a lot compared to your 'free' personal machine, please consider the electricity cost of using an RTX4090...\n### 2.5 Bonus: confidence is key\nSince I already had confidence predictions simply because the confidence head improved the model, I took a step further out of curiosity and checked if I could get a higher score by discarding 'low-confidence' samples and by how much I could improve. Here is a graph (only for one model, not the full ensemble) that shows the R-squared as a function of the fraction of remaining samples after I discarded the most 'low-confidence' ones.  \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6832115%2F3133894fc48f275cf79b88cee8a425aa%2Fconfidence_is_key.png?generation=1722304121419960&alt=media)\nAs you can see, by discarding ~0.1 of the samples, I can already get to a ~0.83+ score. This is not very surprising: if you check the data, you will see that certain months and locations have much lower scores, so a higher score can simply be achieved by discarding said months/location. However, discarding by confidence may be more precise and efficient. It can be interesting to research further, especially given this specific problem, since we can replace the simulator calculation with model predictions and then send back the low-confidence samples to the simulator.  \nAs a side note, I also tried ensembling by confidence (i.e., giving lower weights to low-confidence predictions), but my preliminary research did not show promising results. Although it WAS preliminary- I got to it only on sunday of the last week of the competition, and it was either researching it more in depth or watching the Euro 2024 finals...(congrats Spain!)  \n\n## 3. Sources\nThis is my favorite part. Thank you all the resources that have helped me!  \n[Ptend trick](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/502484) and also [here originally](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/499896#2791290) by @sakvaua and @jano123 .\n[Ribonanza 2nd solution](https://github.com/hoyso48/Stanford---Ribonanza-RNA-Folding-2nd-place-solution) for Squeezeformer architecture guidance and insights, by @hoyso48 .  \n[Ribonanza 3rd place solution](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/460403) for confidence head method, by @dankrstev .  \n[ASLFR 1st solution](https://www.kaggle.com/competitions/asl-fingerspelling/discussion/434485) for the multiple data representation method, by @darraghdog and @christofhenkel .  \n[Dropout is unnecessary](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/514020#2884414) and [also here](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/514020#2885043) by @sakvaua and @ymatioun .   \nIn addition, I used the low-res and high-res data [from HF](https://huggingface.co/LEAP).  \nAnd, of course, [my GitHub](https://github.com/shlomoron/LEAP-solution) with links and instructions for a fully reproducible solution on a Kaggle pipeline.",
    "2940524": "Well done, well deserved.",
    "2940414": "Congratulations on winning the championship, you deserve it!😀\nThank you for your writing; the solution is fascinating, and I learned a lot from it.\n",
    "2944303": "Well done very inspiring !",
    "2943762": "Congrats and thank you for the code! I've had a chance to tinker with your training code today and I'm looking forward to swapping out various pieces with my work as a learning exercise.",
    "2943119": "Congratulations on winning the championship, you deserve it!!!",
    "2942587": "after seeing all the work that went into this, I am now surprised I managed to get into the top 100 😂 got a lot to learn. Congratulations and thanks for sharing!",
    "2942417": "Amazing work , keep publishing and congrats",
    "2942032": "Well done for your work ! Thank you for sharing this detailed answer :) I am wondering what was the training time on TPU for your models with the full data from HF ?",
    "2940478": "> I won't explain why it was suspicious;\n\nthats enough, I get it :D \n\nthank you for publishing the solution and congratz with the win!",
    "2956584": "@shlomoron I guess the confidence head comes from the idea of this paper https://arxiv.org/pdf/1802.04865",
    "2941857": "Congrats to wining the Leep.",
    "2941503": "Congratulations on winning the LEAP! Thank you as well for writing such a stunning article! \nAdditionally, I'd like to point out that in the section on label processing, you employed the third transformation method.",
    "2940906": "Congrats! Amazing work, a lot of work. Its hard for me to imagine how you managed so many experiments and variations. Nice idea the confidence loss, i was not aware of it and i dont get why it works. Also the different normalisations concept was new to me! Well deserved rank 1!",
    "2940843": "Congratulations on winning the championship!\n",
    "2940560": "How did you load the data to colab? Did you just redownload the data to the session every time you started a colab instance? Asking because i currently use Colab for RSNA and everytime I start the notebook it takes 30 min to load the 35GB of data and I cant imagine how long it takes for the LEAP data",
    "2941192": "Thanks for sharing"
  }
}