{
  "id": 437726,
  "title": "[109th Place Solution] 1DConv+Transformer+Augmentation using train+supplemental data split by types",
  "url": "/competitions/asl-fingerspelling/discussion/437726",
  "author_name": "JoAI",
  "post_date": "2023-09-07T23:12:52.394000",
  "votes": 3,
  "comment_count": 2,
  "views": 0,
  "content": "<h1>Context</h1>\n<ul>\n<li>Business context: <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/overview\" target=\"_blank\">https://www.kaggle.com/competitions/asl-fingerspelling/overview</a></li>\n<li>Data context: <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/data\" target=\"_blank\">https://www.kaggle.com/competitions/asl-fingerspelling/data</a></li>\n</ul>\n<h1>Overview of the approach</h1>\n<p>I mainly followed the solutions from <a href=\"https://www.kaggle.com/code/irohith/aslfr-transformer\" target=\"_blank\">ROHITH INGILELA</a>, <a href=\"https://www.kaggle.com/code/markwijkhuizen/aslfr-transformer-training-inference\" target=\"_blank\">Mark Wijkhuizen</a> and <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406684\" target=\"_blank\">hoyso48</a>. First, I modified Mark's solution to make it work in Colab notebook with TPU, which saved me a lot of training time. To use TPU, I secondly converted the dataset to tfrecords files. I managed to get these two steps done by learning from hoyso48's notebooks with a lot of trials and errors. When I converted data, I processed the data following Rohith's dominant hand algorithm. At this point, I got LB: 0.659. Third, I added hoyso48's 1DConv architecture to the transformer model, which gave the LB a 0.015 boost. Fourth, I added most of hoyso48's augmentation methods with some parameter-tuning, which gave the LB a solid 0.033 boost. <br>\nI also tried out some of my ideas. 1) Added supplemental data with some pre-processing. Together with train/valid dataset splitting based on sequence types and participant ids, it showed slight LB improvement. 2) Labelled the dataset with types, i.e.  phone, url, address, name, sentence. Performed multitasking training, i.e. predicting the types of the sequences and predicting the phrases as the same time, which actually dropped the LB. 3) Added a penalty in the loss function of Mark's to penalize the length difference between the true phrase and predicted phrase, which didn't improve the LB. More attempts will be discussed later. </p>\n<h1>Details of the submission</h1>\n<h2>Data processing:</h2>\n<ol>\n<li>Loaded original train_landmarks and supplemental_landmarks dataset to my google drive.</li>\n<li>Split each original parquet file (contains about 1000 sequences) to multiple parquet files by sequence_id. Removed all the frames with no hand landmarks. It turned out the LB will be increased by a lot if leaving half of the missing-hand frames in the data according to <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/434353\" target=\"_blank\">CHRIS DEOTTE's solution</a>. Calculated the frame-to-phrase ratios using the number of frames that have hand landmarks, and saved them into the train_ratio.csv and supple_ratio.csv files.</li>\n<li>Created a <a href=\"https://www.kaggle.com/datasets/joqueen/aslfr-all-landmarks-nonanhand\" target=\"_blank\">new Kaggle dataset</a> with the new parquet files (one parquet means one sequence) and new csv files.</li>\n<li>Created tfrecords files from the new dataset in step 3 using the <a href=\"https://www.kaggle.com/joqueen/yu-aslfr-train-supple-cleaned-landmarks-tfr\" target=\"_blank\">Kaggle notebook</a>. <br>\n4.1. Removed the samples in supplemental dataset that have phrase length longer than 31, as the max phrase length of train dataset is 31. <br>\n4.2 Removed the duplicates in supplemental dataset that has the same participant_id and phrase, as there are a lot duplicates in phrase and in participant_id in supplemental dataset.<br>\n4.3 Removed the samples with frame-to-phrase ratios smaller than 0.5, 1, or 2. It turned out when ratio equals to 1, it gave me the best LB.<br>\n4.4 Split the data to 5 folds based on the phrase type and participant_id. Saved the data as tfrecord files as a <a href=\"https://www.kaggle.com/datasets/joqueen/aslfr-all-5fold-grpsplit-cln-1p0-31phrlen-dedup\" target=\"_blank\">new Kaggle dataset</a>.</li>\n</ol>\n<h2>Training:</h2>\n<ol>\n<li>As described in the overview, the main body of my solution is from <a href=\"https://www.kaggle.com/code/markwijkhuizen/aslfr-transformer-training-inference\" target=\"_blank\">Mark Wijkhuizen</a> and <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406684\" target=\"_blank\">hoyso48</a>, i.e. 1DConv+Transformer+Augmentation. As I trained my model using TPU in Colab, here is my <a href=\"https://colab.research.google.com/drive/1IOkmBN558p8hB_bfF-g4KO7SrvQiFshC?usp=sharing\" target=\"_blank\">Colab notebook</a> for training. I'm not going to repeat the details again here, but I think there were a few learnings/observations worth mentioning:</li>\n<li>Mark padded the phrase length to 128 on the decoder side, the same as the frame length on the encoder side. We don't have to match the dimension. The max phrase length is only 31 and already padded to 32. Instead of padding phrase further to 128, I modified the MHA block to be able to take two different dimensions, i.e. q_len and kv_len. To be honest, it didn't improve or decrease my LB, so Mark's MHA is perfectly fine, but it did save me some training time. </li>\n</ol>\n<pre><code> (tf.keras.layers.Layer):\n     (): \n        (MultiHeadAttention,self).__init__()\n        self.d_model = d_model\n        self.n_heads = n_heads\n        self.kv_len = kv_len\n        self.q_len = q_len                                              \n        self.depth = d_model //                                                \n        self.scale =  / tf.math.sqrt(tf.cast(self.depth, tf.float32))        \n        self.wq = self.fused_mha(self.depth, q_len)                             \n        self.wk = self.fused_mha(self.depth, kv_len)                            \n        self.wv = self.fused_mha(self.depth, kv_len)                            \n        self.wo = tf.keras.layers.Dense(d_model, use_bias=)\n        self.softmax = tf.keras.layers.Softmax()\n        self.reshape = tf.keras.Sequential([\n            tf.keras.layers.Permute([, , ]),\n            tf.keras.layers.Reshape([self.q_len, self.depth]),\n        ])\n        self.do = tf.keras.layers.Dropout(dropout)\n        self.supports_masking = \n\n     ():\n         tf.keras.Sequential([\n            tf.keras.layers.Dense(dim, use_bias=),\n            tf.keras.layers.Reshape([, self.n_heads, dim // self.n_heads]),\n            tf.keras.layers.Permute([, , ]),\n        ])\n\n     ():\n        Q = self.wq(q)                                                          \n        K = self.wk(k)                                                          \n        V = self.wv(v)                                                          \n        x = tf.matmul(Q, K, transpose_b=) * self.scale                      \n        x = self.softmax(x, mask=attention_mask) @ V                            \n        x = self.reshape(x)                                                     \n        x = self.wo(x)\n        x = self.do(x, training=training)\n         x\n</code></pre>\n<ol>\n<li>I increased the mha_dropout_ratio, mlp_dropout_ratio of decoder by 1.5 times and clf_dropout_ratio of the classifier by 2 times, to heavily regularize the decoder and prevent overfitting. It improved my LB by 0.006. This idea is inspired by a discussion commented by hoyso48 .</li>\n</ol>\n<pre><code>    \n    x = Decoder(n_dec_blocks,\n                dec_dim,\n                n_mha_heads,\n                mha_dropout_ratio*, \n                mlp_ratio,\n                mlp_dropout_ratio*, \n                frame_len,\n                phrase_len,\n                )(x, phrase_inp, frames_inp)\n\n    \n    x = tf.keras.Sequential([\n        tf.keras.layers.Dropout(clf_dropout_ratio*),  \n        tf.keras.layers.Dense(N_UNIQUE_CHARACTERS,\n                              activation=tf.keras.activations.linear,\n                              kernel_initializer=INIT_HE_UNIFORM,\n                              use_bias=),\n    ], name=)(x)\n</code></pre>\n<h2>Inference</h2>\n<p>My inference followed Mark's inference method. Here is the <a href=\"https://www.kaggle.com/joqueen/m12-yu-aslfr-inference\" target=\"_blank\">inference Kaggle notebook</a>, which reported both my highest public score 0.713 and my highest private score 0.665.</p>\n<h1>What tried but didn't work (well)</h1>\n<h2>Loss penalization</h2>\n<p>In Mark's customized loss function, y_true and y_pred got truncated by the y_true length. y_true length is known, so how to truncate the y_pred is not from training process. I was thinking what if the y_pred is very long, but got truncated in the loss function, which may not really represent the loss.</p>\n<pre><code> ():\n    idxs = tf.where(y_true != PAD_IDX)\n    y_true = tf.gather_nd(y_true, idxs)\n    y_pred = tf.gather_nd(y_pred, idxs)\n    y_true = tf.cast(y_true, tf.int32)\n    y_true = tf.one_hot(y_true, N_UNIQUE_CHARACTERS, axis=)\n    loss = tf.keras.losses.categorical_crossentropy(y_true, y_pred, label_smoothing=, from_logits=) \n    loss = tf.math.reduce_mean(loss)\n     loss\n</code></pre>\n<p>So I added a penalty about the length difference between the y_true and y_pred, to force the model to generate y_pred with a similar length as y_true while keep the categorical_crossentropy optimal.</p>\n<pre><code> ():\n    \n    y_true_len = tf.cast(tf.argmax(tf.cast(tf.math.equal(y_true, PAD_IDX), tf.int32),axis=), tf.int32)\n\n    _idx = tf.argmax(y_pred, axis=)\n    _ = tf.math.equal(_idx, EOS_IDX)\n    _len1 = tf.cast(tf.argmax(tf.cast(_, tf.int32), axis=), tf.int32)\n\n    _no_eos = tf.math.logical_not(tf.reduce_any(_, axis=))\n    _len2 = tf.cast(_no_eos, tf.int32) * CFG.phrase_len\n    y_pred_len = _len1 + _len2\n    \n    penal = tf.cast(tf.(y_true_len-y_pred_len), tf.float32)\n    penal = tf.math.reduce_mean(penal)\n\n    \n    idxs = tf.where(y_true != PAD_IDX)\n    y_true = tf.gather_nd(y_true, idxs)\n    y_pred = tf.gather_nd(y_pred, idxs)\n    y_true = tf.cast(y_true, tf.int32)\n    y_true = tf.one_hot(y_true, N_UNIQUE_CHARACTERS, axis=)\n    loss = tf.keras.losses.categorical_crossentropy(y_true, y_pred, label_smoothing=, from_logits=) \n    loss = tf.math.reduce_mean(loss)\n\n     loss + CFG.loss_coeff * penal\n</code></pre>\n<p>I was quite excited about this idea, and managed to code and run it through. But it still reduced my LB by 0.001 even after some fine-tuning. (sad face)</p>\n<h2>multi-tasking training</h2>\n<p>Inspired by Mark's solution, I categorized the train+supplemental data to five types, phone_number, url, address, name_like, sentence. It is obvious that each type is so different from the other types, in terms of the content, the length, the structure. I was thinking, maybe I can do multi-tasking - let the model do classification and phrase-prediction at the same time, which may force the model to figure out the different patterns for different types of samples. But this actually reduced my LB by about 0.1, which is a lot to me. </p>\n<pre><code>        model.(\n            optimizer=opt,\n            loss = [tf.keras.losses.CategoricalCrossentropy(from_logits=,label_smoothing=), loss_w_ls],\n            metrics = [[tf.keras.metrics.CategoricalAccuracy()], [TopKAccuracy()]], \n            steps_per_execution=steps_per_epoch,\n        )\n</code></pre>\n<h2>positional encoding</h2>\n<p>Mark's method used trainable positional encoding. I switched it to the sin/cos positional encoding for encoder only, and actually it reduced my LB by 0.005. But I didn't try to switch it for decoder. </p>\n<h2>a validation strategy</h2>\n<p>I was struggling a lot on designing a good validation strategy that has the CV correlated to the LB well. Out of desperation, I even kept the all the validation data in their original form - before removing the missing-hand frames and removing the samples with low frame-to-phrase ratios. And calculated the Levenshtein distance on the raw valid data. However, it was very time consuming, my model training only took 45 min - 1.5 hour, depending on the parameters. But the Levenshtein distance calculation for validation took about 3 hours. And even with all the effort, I still didn't see strong correlation. </p>\n<h2>CTC</h2>\n<p>I tried out CTC a bit following the notebooks from <a href=\"https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place\" target=\"_blank\">ROHITH INGILELA</a> and <a href=\"https://www.kaggle.com/code/shlomoron/aslfr-ctc-on-tpu\" target=\"_blank\">GREYSNOW</a>. But due to time limitation, I didn't really figure it out in time. And it looks like the majority of the top-tier solutions used CTC.</p>\n<h1>Sources</h1>\n<p><a href=\"https://www.kaggle.com/code/irohith/aslfr-transformer\" target=\"_blank\">https://www.kaggle.com/code/irohith/aslfr-transformer</a><br>\n<a href=\"https://www.kaggle.com/code/markwijkhuizen/aslfr-transformer-training-inference\" target=\"_blank\">https://www.kaggle.com/code/markwijkhuizen/aslfr-transformer-training-inference</a>) <br>\n<a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406684\" target=\"_blank\">https://www.kaggle.com/competitions/asl-signs/discussion/406684</a><br>\n<a href=\"https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place\" target=\"_blank\">https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place</a><br>\n<a href=\"https://www.kaggle.com/code/shlomoron/aslfr-ctc-on-tpu\" target=\"_blank\">https://www.kaggle.com/code/shlomoron/aslfr-ctc-on-tpu</a><br>\n<a href=\"https://www.coursera.org/learn/nlp-sequence-models?specialization=deep-learning\" target=\"_blank\">https://www.coursera.org/learn/nlp-sequence-models?specialization=deep-learning</a> (transformer network)</p>\n<p>This is my first competition and first solution write-up. Please let me know if I missed anything or if any part of my explanation was unclear. Suggestions and discussions are welcome. Thank you.</p>",
  "messages": [
    {
      "id": 2428488,
      "postDate": "2023-09-07T23:12:52.393Z",
      "content": "<h1>Context</h1>\n<ul>\n<li>Business context: <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/overview\" target=\"_blank\">https://www.kaggle.com/competitions/asl-fingerspelling/overview</a></li>\n<li>Data context: <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/data\" target=\"_blank\">https://www.kaggle.com/competitions/asl-fingerspelling/data</a></li>\n</ul>\n<h1>Overview of the approach</h1>\n<p>I mainly followed the solutions from <a href=\"https://www.kaggle.com/code/irohith/aslfr-transformer\" target=\"_blank\">ROHITH INGILELA</a>, <a href=\"https://www.kaggle.com/code/markwijkhuizen/aslfr-transformer-training-inference\" target=\"_blank\">Mark Wijkhuizen</a> and <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406684\" target=\"_blank\">hoyso48</a>. First, I modified Mark's solution to make it work in Colab notebook with TPU, which saved me a lot of training time. To use TPU, I secondly converted the dataset to tfrecords files. I managed to get these two steps done by learning from hoyso48's notebooks with a lot of trials and errors. When I converted data, I processed the data following Rohith's dominant hand algorithm. At this point, I got LB: 0.659. Third, I added hoyso48's 1DConv architecture to the transformer model, which gave the LB a 0.015 boost. Fourth, I added most of hoyso48's augmentation methods with some parameter-tuning, which gave the LB a solid 0.033 boost. <br>\nI also tried out some of my ideas. 1) Added supplemental data with some pre-processing. Together with train/valid dataset splitting based on sequence types and participant ids, it showed slight LB improvement. 2) Labelled the dataset with types, i.e.  phone, url, address, name, sentence. Performed multitasking training, i.e. predicting the types of the sequences and predicting the phrases as the same time, which actually dropped the LB. 3) Added a penalty in the loss function of Mark's to penalize the length difference between the true phrase and predicted phrase, which didn't improve the LB. More attempts will be discussed later. </p>\n<h1>Details of the submission</h1>\n<h2>Data processing:</h2>\n<ol>\n<li>Loaded original train_landmarks and supplemental_landmarks dataset to my google drive.</li>\n<li>Split each original parquet file (contains about 1000 sequences) to multiple parquet files by sequence_id. Removed all the frames with no hand landmarks. It turned out the LB will be increased by a lot if leaving half of the missing-hand frames in the data according to <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/434353\" target=\"_blank\">CHRIS DEOTTE's solution</a>. Calculated the frame-to-phrase ratios using the number of frames that have hand landmarks, and saved them into the train_ratio.csv and supple_ratio.csv files.</li>\n<li>Created a <a href=\"https://www.kaggle.com/datasets/joqueen/aslfr-all-landmarks-nonanhand\" target=\"_blank\">new Kaggle dataset</a> with the new parquet files (one parquet means one sequence) and new csv files.</li>\n<li>Created tfrecords files from the new dataset in step 3 using the <a href=\"https://www.kaggle.com/joqueen/yu-aslfr-train-supple-cleaned-landmarks-tfr\" target=\"_blank\">Kaggle notebook</a>. <br>\n4.1. Removed the samples in supplemental dataset that have phrase length longer than 31, as the max phrase length of train dataset is 31. <br>\n4.2 Removed the duplicates in supplemental dataset that has the same participant_id and phrase, as there are a lot duplicates in phrase and in participant_id in supplemental dataset.<br>\n4.3 Removed the samples with frame-to-phrase ratios smaller than 0.5, 1, or 2. It turned out when ratio equals to 1, it gave me the best LB.<br>\n4.4 Split the data to 5 folds based on the phrase type and participant_id. Saved the data as tfrecord files as a <a href=\"https://www.kaggle.com/datasets/joqueen/aslfr-all-5fold-grpsplit-cln-1p0-31phrlen-dedup\" target=\"_blank\">new Kaggle dataset</a>.</li>\n</ol>\n<h2>Training:</h2>\n<ol>\n<li>As described in the overview, the main body of my solution is from <a href=\"https://www.kaggle.com/code/markwijkhuizen/aslfr-transformer-training-inference\" target=\"_blank\">Mark Wijkhuizen</a> and <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406684\" target=\"_blank\">hoyso48</a>, i.e. 1DConv+Transformer+Augmentation. As I trained my model using TPU in Colab, here is my <a href=\"https://colab.research.google.com/drive/1IOkmBN558p8hB_bfF-g4KO7SrvQiFshC?usp=sharing\" target=\"_blank\">Colab notebook</a> for training. I'm not going to repeat the details again here, but I think there were a few learnings/observations worth mentioning:</li>\n<li>Mark padded the phrase length to 128 on the decoder side, the same as the frame length on the encoder side. We don't have to match the dimension. The max phrase length is only 31 and already padded to 32. Instead of padding phrase further to 128, I modified the MHA block to be able to take two different dimensions, i.e. q_len and kv_len. To be honest, it didn't improve or decrease my LB, so Mark's MHA is perfectly fine, but it did save me some training time. </li>\n</ol>\n<pre><code> (tf.keras.layers.Layer):\n     (): \n        (MultiHeadAttention,self).__init__()\n        self.d_model = d_model\n        self.n_heads = n_heads\n        self.kv_len = kv_len\n        self.q_len = q_len                                              \n        self.depth = d_model //                                                \n        self.scale =  / tf.math.sqrt(tf.cast(self.depth, tf.float32))        \n        self.wq = self.fused_mha(self.depth, q_len)                             \n        self.wk = self.fused_mha(self.depth, kv_len)                            \n        self.wv = self.fused_mha(self.depth, kv_len)                            \n        self.wo = tf.keras.layers.Dense(d_model, use_bias=)\n        self.softmax = tf.keras.layers.Softmax()\n        self.reshape = tf.keras.Sequential([\n            tf.keras.layers.Permute([, , ]),\n            tf.keras.layers.Reshape([self.q_len, self.depth]),\n        ])\n        self.do = tf.keras.layers.Dropout(dropout)\n        self.supports_masking = \n\n     ():\n         tf.keras.Sequential([\n            tf.keras.layers.Dense(dim, use_bias=),\n            tf.keras.layers.Reshape([, self.n_heads, dim // self.n_heads]),\n            tf.keras.layers.Permute([, , ]),\n        ])\n\n     ():\n        Q = self.wq(q)                                                          \n        K = self.wk(k)                                                          \n        V = self.wv(v)                                                          \n        x = tf.matmul(Q, K, transpose_b=) * self.scale                      \n        x = self.softmax(x, mask=attention_mask) @ V                            \n        x = self.reshape(x)                                                     \n        x = self.wo(x)\n        x = self.do(x, training=training)\n         x\n</code></pre>\n<ol>\n<li>I increased the mha_dropout_ratio, mlp_dropout_ratio of decoder by 1.5 times and clf_dropout_ratio of the classifier by 2 times, to heavily regularize the decoder and prevent overfitting. It improved my LB by 0.006. This idea is inspired by a discussion commented by hoyso48 .</li>\n</ol>\n<pre><code>    \n    x = Decoder(n_dec_blocks,\n                dec_dim,\n                n_mha_heads,\n                mha_dropout_ratio*, \n                mlp_ratio,\n                mlp_dropout_ratio*, \n                frame_len,\n                phrase_len,\n                )(x, phrase_inp, frames_inp)\n\n    \n    x = tf.keras.Sequential([\n        tf.keras.layers.Dropout(clf_dropout_ratio*),  \n        tf.keras.layers.Dense(N_UNIQUE_CHARACTERS,\n                              activation=tf.keras.activations.linear,\n                              kernel_initializer=INIT_HE_UNIFORM,\n                              use_bias=),\n    ], name=)(x)\n</code></pre>\n<h2>Inference</h2>\n<p>My inference followed Mark's inference method. Here is the <a href=\"https://www.kaggle.com/joqueen/m12-yu-aslfr-inference\" target=\"_blank\">inference Kaggle notebook</a>, which reported both my highest public score 0.713 and my highest private score 0.665.</p>\n<h1>What tried but didn't work (well)</h1>\n<h2>Loss penalization</h2>\n<p>In Mark's customized loss function, y_true and y_pred got truncated by the y_true length. y_true length is known, so how to truncate the y_pred is not from training process. I was thinking what if the y_pred is very long, but got truncated in the loss function, which may not really represent the loss.</p>\n<pre><code> ():\n    idxs = tf.where(y_true != PAD_IDX)\n    y_true = tf.gather_nd(y_true, idxs)\n    y_pred = tf.gather_nd(y_pred, idxs)\n    y_true = tf.cast(y_true, tf.int32)\n    y_true = tf.one_hot(y_true, N_UNIQUE_CHARACTERS, axis=)\n    loss = tf.keras.losses.categorical_crossentropy(y_true, y_pred, label_smoothing=, from_logits=) \n    loss = tf.math.reduce_mean(loss)\n     loss\n</code></pre>\n<p>So I added a penalty about the length difference between the y_true and y_pred, to force the model to generate y_pred with a similar length as y_true while keep the categorical_crossentropy optimal.</p>\n<pre><code> ():\n    \n    y_true_len = tf.cast(tf.argmax(tf.cast(tf.math.equal(y_true, PAD_IDX), tf.int32),axis=), tf.int32)\n\n    _idx = tf.argmax(y_pred, axis=)\n    _ = tf.math.equal(_idx, EOS_IDX)\n    _len1 = tf.cast(tf.argmax(tf.cast(_, tf.int32), axis=), tf.int32)\n\n    _no_eos = tf.math.logical_not(tf.reduce_any(_, axis=))\n    _len2 = tf.cast(_no_eos, tf.int32) * CFG.phrase_len\n    y_pred_len = _len1 + _len2\n    \n    penal = tf.cast(tf.(y_true_len-y_pred_len), tf.float32)\n    penal = tf.math.reduce_mean(penal)\n\n    \n    idxs = tf.where(y_true != PAD_IDX)\n    y_true = tf.gather_nd(y_true, idxs)\n    y_pred = tf.gather_nd(y_pred, idxs)\n    y_true = tf.cast(y_true, tf.int32)\n    y_true = tf.one_hot(y_true, N_UNIQUE_CHARACTERS, axis=)\n    loss = tf.keras.losses.categorical_crossentropy(y_true, y_pred, label_smoothing=, from_logits=) \n    loss = tf.math.reduce_mean(loss)\n\n     loss + CFG.loss_coeff * penal\n</code></pre>\n<p>I was quite excited about this idea, and managed to code and run it through. But it still reduced my LB by 0.001 even after some fine-tuning. (sad face)</p>\n<h2>multi-tasking training</h2>\n<p>Inspired by Mark's solution, I categorized the train+supplemental data to five types, phone_number, url, address, name_like, sentence. It is obvious that each type is so different from the other types, in terms of the content, the length, the structure. I was thinking, maybe I can do multi-tasking - let the model do classification and phrase-prediction at the same time, which may force the model to figure out the different patterns for different types of samples. But this actually reduced my LB by about 0.1, which is a lot to me. </p>\n<pre><code>        model.(\n            optimizer=opt,\n            loss = [tf.keras.losses.CategoricalCrossentropy(from_logits=,label_smoothing=), loss_w_ls],\n            metrics = [[tf.keras.metrics.CategoricalAccuracy()], [TopKAccuracy()]], \n            steps_per_execution=steps_per_epoch,\n        )\n</code></pre>\n<h2>positional encoding</h2>\n<p>Mark's method used trainable positional encoding. I switched it to the sin/cos positional encoding for encoder only, and actually it reduced my LB by 0.005. But I didn't try to switch it for decoder. </p>\n<h2>a validation strategy</h2>\n<p>I was struggling a lot on designing a good validation strategy that has the CV correlated to the LB well. Out of desperation, I even kept the all the validation data in their original form - before removing the missing-hand frames and removing the samples with low frame-to-phrase ratios. And calculated the Levenshtein distance on the raw valid data. However, it was very time consuming, my model training only took 45 min - 1.5 hour, depending on the parameters. But the Levenshtein distance calculation for validation took about 3 hours. And even with all the effort, I still didn't see strong correlation. </p>\n<h2>CTC</h2>\n<p>I tried out CTC a bit following the notebooks from <a href=\"https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place\" target=\"_blank\">ROHITH INGILELA</a> and <a href=\"https://www.kaggle.com/code/shlomoron/aslfr-ctc-on-tpu\" target=\"_blank\">GREYSNOW</a>. But due to time limitation, I didn't really figure it out in time. And it looks like the majority of the top-tier solutions used CTC.</p>\n<h1>Sources</h1>\n<p><a href=\"https://www.kaggle.com/code/irohith/aslfr-transformer\" target=\"_blank\">https://www.kaggle.com/code/irohith/aslfr-transformer</a><br>\n<a href=\"https://www.kaggle.com/code/markwijkhuizen/aslfr-transformer-training-inference\" target=\"_blank\">https://www.kaggle.com/code/markwijkhuizen/aslfr-transformer-training-inference</a>) <br>\n<a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406684\" target=\"_blank\">https://www.kaggle.com/competitions/asl-signs/discussion/406684</a><br>\n<a href=\"https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place\" target=\"_blank\">https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place</a><br>\n<a href=\"https://www.kaggle.com/code/shlomoron/aslfr-ctc-on-tpu\" target=\"_blank\">https://www.kaggle.com/code/shlomoron/aslfr-ctc-on-tpu</a><br>\n<a href=\"https://www.coursera.org/learn/nlp-sequence-models?specialization=deep-learning\" target=\"_blank\">https://www.coursera.org/learn/nlp-sequence-models?specialization=deep-learning</a> (transformer network)</p>\n<p>This is my first competition and first solution write-up. Please let me know if I missed anything or if any part of my explanation was unclear. Suggestions and discussions are welcome. Thank you.</p>",
      "rawMarkdown": "# Context\n- Business context: https://www.kaggle.com/competitions/asl-fingerspelling/overview\n- Data context: https://www.kaggle.com/competitions/asl-fingerspelling/data\n\n# Overview of the approach\nI mainly followed the solutions from [ROHITH INGILELA](https://www.kaggle.com/code/irohith/aslfr-transformer), [Mark Wijkhuizen](https://www.kaggle.com/code/markwijkhuizen/aslfr-transformer-training-inference) and [hoyso48](https://www.kaggle.com/competitions/asl-signs/discussion/406684). First, I modified Mark's solution to make it work in Colab notebook with TPU, which saved me a lot of training time. To use TPU, I secondly converted the dataset to tfrecords files. I managed to get these two steps done by learning from hoyso48's notebooks with a lot of trials and errors. When I converted data, I processed the data following Rohith's dominant hand algorithm. At this point, I got LB: 0.659. Third, I added hoyso48's 1DConv architecture to the transformer model, which gave the LB a 0.015 boost. Fourth, I added most of hoyso48's augmentation methods with some parameter-tuning, which gave the LB a solid 0.033 boost. \nI also tried out some of my ideas. 1) Added supplemental data with some pre-processing. Together with train/valid dataset splitting based on sequence types and participant ids, it showed slight LB improvement. 2) Labelled the dataset with types, i.e.  phone, url, address, name, sentence. Performed multitasking training, i.e. predicting the types of the sequences and predicting the phrases as the same time, which actually dropped the LB. 3) Added a penalty in the loss function of Mark's to penalize the length difference between the true phrase and predicted phrase, which didn't improve the LB. More attempts will be discussed later. \n\n# Details of the submission\n## Data processing:\n1. Loaded original train_landmarks and supplemental_landmarks dataset to my google drive.\n2. Split each original parquet file (contains about 1000 sequences) to multiple parquet files by sequence_id. Removed all the frames with no hand landmarks. It turned out the LB will be increased by a lot if leaving half of the missing-hand frames in the data according to [CHRIS DEOTTE's solution](https://www.kaggle.com/competitions/asl-fingerspelling/discussion/434353). Calculated the frame-to-phrase ratios using the number of frames that have hand landmarks, and saved them into the train_ratio.csv and supple_ratio.csv files.\n3. Created a [new Kaggle dataset](https://www.kaggle.com/datasets/joqueen/aslfr-all-landmarks-nonanhand) with the new parquet files (one parquet means one sequence) and new csv files.\n4. Created tfrecords files from the new dataset in step 3 using the [Kaggle notebook](https://www.kaggle.com/joqueen/yu-aslfr-train-supple-cleaned-landmarks-tfr). \n4.1. Removed the samples in supplemental dataset that have phrase length longer than 31, as the max phrase length of train dataset is 31. \n4.2 Removed the duplicates in supplemental dataset that has the same participant_id and phrase, as there are a lot duplicates in phrase and in participant_id in supplemental dataset.\n4.3 Removed the samples with frame-to-phrase ratios smaller than 0.5, 1, or 2. It turned out when ratio equals to 1, it gave me the best LB.\n4.4 Split the data to 5 folds based on the phrase type and participant_id. Saved the data as tfrecord files as a [new Kaggle dataset](https://www.kaggle.com/datasets/joqueen/aslfr-all-5fold-grpsplit-cln-1p0-31phrlen-dedup).\n\n## Training:\n1. As described in the overview, the main body of my solution is from [Mark Wijkhuizen](https://www.kaggle.com/code/markwijkhuizen/aslfr-transformer-training-inference) and [hoyso48](https://www.kaggle.com/competitions/asl-signs/discussion/406684), i.e. 1DConv+Transformer+Augmentation. As I trained my model using TPU in Colab, here is my [Colab notebook](https://colab.research.google.com/drive/1IOkmBN558p8hB_bfF-g4KO7SrvQiFshC?usp=sharing) for training. I'm not going to repeat the details again here, but I think there were a few learnings/observations worth mentioning:\n2. Mark padded the phrase length to 128 on the decoder side, the same as the frame length on the encoder side. We don't have to match the dimension. The max phrase length is only 31 and already padded to 32. Instead of padding phrase further to 128, I modified the MHA block to be able to take two different dimensions, i.e. q_len and kv_len. To be honest, it didn't improve or decrease my LB, so Mark's MHA is perfectly fine, but it did save me some training time. \n```python\nclass MultiHeadAttention(tf.keras.layers.Layer):\n    def __init__(self, d_model, n_heads, q_len, kv_len, dropout): \n        super(MultiHeadAttention,self).__init__()\n        self.d_model = d_model\n        self.n_heads = n_heads\n        self.kv_len = kv_len\n        self.q_len = q_len                                              \n        self.depth = d_model // 2                                               \n        self.scale = 1.0 / tf.math.sqrt(tf.cast(self.depth, tf.float32))        \n        self.wq = self.fused_mha(self.depth, q_len)                             \n        self.wk = self.fused_mha(self.depth, kv_len)                            \n        self.wv = self.fused_mha(self.depth, kv_len)                            \n        self.wo = tf.keras.layers.Dense(d_model, use_bias=False)\n        self.softmax = tf.keras.layers.Softmax()\n        self.reshape = tf.keras.Sequential([\n            tf.keras.layers.Permute([2, 1, 3]),\n            tf.keras.layers.Reshape([self.q_len, self.depth]),\n        ])\n        self.do = tf.keras.layers.Dropout(dropout)\n        self.supports_masking = True\n\n    def fused_mha(self, dim, len):\n        return tf.keras.Sequential([\n            tf.keras.layers.Dense(dim, use_bias=False),\n            tf.keras.layers.Reshape([len, self.n_heads, dim // self.n_heads]),\n            tf.keras.layers.Permute([2, 1, 3]),\n        ])\n\n    def call(self, q, k, v, attention_mask=None, training=False):\n        Q = self.wq(q)                                                          \n        K = self.wk(k)                                                          \n        V = self.wv(v)                                                          \n        x = tf.matmul(Q, K, transpose_b=True) * self.scale                      \n        x = self.softmax(x, mask=attention_mask) @ V                            \n        x = self.reshape(x)                                                     \n        x = self.wo(x)\n        x = self.do(x, training=training)\n        return x\n```\n3. I increased the mha_dropout_ratio, mlp_dropout_ratio of decoder by 1.5 times and clf_dropout_ratio of the classifier by 2 times, to heavily regularize the decoder and prevent overfitting. It improved my LB by 0.006. This idea is inspired by a discussion commented by hoyso48 .\n```python\n    # Decoder\n    x = Decoder(n_dec_blocks,\n                dec_dim,\n                n_mha_heads,\n                mha_dropout_ratio*1.5, \n                mlp_ratio,\n                mlp_dropout_ratio*1.5, \n                frame_len,\n                phrase_len,\n                )(x, phrase_inp, frames_inp)\n\n    # Classifier\n    x = tf.keras.Sequential([\n        tf.keras.layers.Dropout(clf_dropout_ratio*2),  \n        tf.keras.layers.Dense(N_UNIQUE_CHARACTERS,\n                              activation=tf.keras.activations.linear,\n                              kernel_initializer=INIT_HE_UNIFORM,\n                              use_bias=False),\n    ], name='classifier')(x)\n```\n\n## Inference\nMy inference followed Mark's inference method. Here is the [inference Kaggle notebook](https://www.kaggle.com/joqueen/m12-yu-aslfr-inference), which reported both my highest public score 0.713 and my highest private score 0.665.\n\n# What tried but didn't work (well)\n## Loss penalization\nIn Mark's customized loss function, y_true and y_pred got truncated by the y_true length. y_true length is known, so how to truncate the y_pred is not from training process. I was thinking what if the y_pred is very long, but got truncated in the loss function, which may not really represent the loss.\n```python\ndef loss_w_ls(y_true, y_pred):\n    idxs = tf.where(y_true != PAD_IDX)\n    y_true = tf.gather_nd(y_true, idxs)\n    y_pred = tf.gather_nd(y_pred, idxs)\n    y_true = tf.cast(y_true, tf.int32)\n    y_true = tf.one_hot(y_true, N_UNIQUE_CHARACTERS, axis=1)\n    loss = tf.keras.losses.categorical_crossentropy(y_true, y_pred, label_smoothing=0.25, from_logits=True) # 0.25\n    loss = tf.math.reduce_mean(loss)\n    return loss\n```\nSo I added a penalty about the length difference between the y_true and y_pred, to force the model to generate y_pred with a similar length as y_true while keep the categorical_crossentropy optimal.\n```python\ndef loss_w_ls(y_true, y_pred):\n    # penalize different len btw y_true and y_pred\n    y_true_len = tf.cast(tf.argmax(tf.cast(tf.math.equal(y_true, PAD_IDX), tf.int32),axis=1), tf.int32)\n\n    _idx = tf.argmax(y_pred, axis=2)\n    _bool = tf.math.equal(_idx, EOS_IDX)\n    _len1 = tf.cast(tf.argmax(tf.cast(_bool, tf.int32), axis=1), tf.int32)\n\n    _no_eos = tf.math.logical_not(tf.reduce_any(_bool, axis=1))\n    _len2 = tf.cast(_no_eos, tf.int32) * CFG.phrase_len\n    y_pred_len = _len1 + _len2\n    # y_pred_len = tf.cond(y_pred_eos, lambda:y_pred_len, lambda:tf.ones_like(y_pred_len)*CFG.phrase_len)\n    penal = tf.cast(tf.abs(y_true_len-y_pred_len), tf.float32)\n    penal = tf.math.reduce_mean(penal)\n\n    # Filter Pad Tokens\n    idxs = tf.where(y_true != PAD_IDX)\n    y_true = tf.gather_nd(y_true, idxs)\n    y_pred = tf.gather_nd(y_pred, idxs)\n    y_true = tf.cast(y_true, tf.int32)\n    y_true = tf.one_hot(y_true, N_UNIQUE_CHARACTERS, axis=1)\n    loss = tf.keras.losses.categorical_crossentropy(y_true, y_pred, label_smoothing=0.25, from_logits=True) # 0.25\n    loss = tf.math.reduce_mean(loss)\n\n    return loss + CFG.loss_coeff * penal\n```\nI was quite excited about this idea, and managed to code and run it through. But it still reduced my LB by 0.001 even after some fine-tuning. (sad face)\n\n## multi-tasking training\nInspired by Mark's solution, I categorized the train+supplemental data to five types, phone_number, url, address, name_like, sentence. It is obvious that each type is so different from the other types, in terms of the content, the length, the structure. I was thinking, maybe I can do multi-tasking - let the model do classification and phrase-prediction at the same time, which may force the model to figure out the different patterns for different types of samples. But this actually reduced my LB by about 0.1, which is a lot to me. \n```python\n        model.compile(\n            optimizer=opt,\n            loss = [tf.keras.losses.CategoricalCrossentropy(from_logits=True,label_smoothing=0.1), loss_w_ls],\n            metrics = [[tf.keras.metrics.CategoricalAccuracy()], [TopKAccuracy(1)]], \n            steps_per_execution=steps_per_epoch,\n        )\n```\n\n## positional encoding\nMark's method used trainable positional encoding. I switched it to the sin/cos positional encoding for encoder only, and actually it reduced my LB by 0.005. But I didn't try to switch it for decoder. \n\n## a validation strategy\nI was struggling a lot on designing a good validation strategy that has the CV correlated to the LB well. Out of desperation, I even kept the all the validation data in their original form - before removing the missing-hand frames and removing the samples with low frame-to-phrase ratios. And calculated the Levenshtein distance on the raw valid data. However, it was very time consuming, my model training only took 45 min - 1.5 hour, depending on the parameters. But the Levenshtein distance calculation for validation took about 3 hours. And even with all the effort, I still didn't see strong correlation. \n\n## CTC\nI tried out CTC a bit following the notebooks from [ROHITH INGILELA](https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place) and [GREYSNOW](https://www.kaggle.com/code/shlomoron/aslfr-ctc-on-tpu). But due to time limitation, I didn't really figure it out in time. And it looks like the majority of the top-tier solutions used CTC.\n\n# Sources\nhttps://www.kaggle.com/code/irohith/aslfr-transformer\nhttps://www.kaggle.com/code/markwijkhuizen/aslfr-transformer-training-inference) \nhttps://www.kaggle.com/competitions/asl-signs/discussion/406684\nhttps://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place\nhttps://www.kaggle.com/code/shlomoron/aslfr-ctc-on-tpu\nhttps://www.coursera.org/learn/nlp-sequence-models?specialization=deep-learning (transformer network)\n\nThis is my first competition and first solution write-up. Please let me know if I missed anything or if any part of my explanation was unclear. Suggestions and discussions are welcome. Thank you.\n",
      "votes": 3
    },
    {
      "id": 2546624,
      "postDate": "2023-12-02T16:12:06.853Z",
      "content": "<p>Nice work!</p>",
      "rawMarkdown": "Nice work!",
      "votes": 1
    },
    {
      "id": 2677401,
      "postDate": "2024-03-02T06:42:26.337Z",
      "content": "<p>I can understand this. All other solutions are just tons of lines of code. Thanks for a great share :)</p>",
      "rawMarkdown": "I can understand this. All other solutions are just tons of lines of code. Thanks for a great share :)"
    }
  ],
  "comments": [
    {
      "id": 2546624,
      "author_name": "Ho Dinh Trieu",
      "author_url": "",
      "post_date": "2023-12-02T16:12:06.853000",
      "content": "<p>Nice work!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2677401,
      "author_name": "Mihir Sutaria",
      "author_url": "",
      "post_date": "2024-03-02T06:42:26.337000",
      "content": "<p>I can understand this. All other solutions are just tons of lines of code. Thanks for a great share :)</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2428488": "# Context\n- Business context: https://www.kaggle.com/competitions/asl-fingerspelling/overview\n- Data context: https://www.kaggle.com/competitions/asl-fingerspelling/data\n\n# Overview of the approach\nI mainly followed the solutions from [ROHITH INGILELA](https://www.kaggle.com/code/irohith/aslfr-transformer), [Mark Wijkhuizen](https://www.kaggle.com/code/markwijkhuizen/aslfr-transformer-training-inference) and [hoyso48](https://www.kaggle.com/competitions/asl-signs/discussion/406684). First, I modified Mark's solution to make it work in Colab notebook with TPU, which saved me a lot of training time. To use TPU, I secondly converted the dataset to tfrecords files. I managed to get these two steps done by learning from hoyso48's notebooks with a lot of trials and errors. When I converted data, I processed the data following Rohith's dominant hand algorithm. At this point, I got LB: 0.659. Third, I added hoyso48's 1DConv architecture to the transformer model, which gave the LB a 0.015 boost. Fourth, I added most of hoyso48's augmentation methods with some parameter-tuning, which gave the LB a solid 0.033 boost. \nI also tried out some of my ideas. 1) Added supplemental data with some pre-processing. Together with train/valid dataset splitting based on sequence types and participant ids, it showed slight LB improvement. 2) Labelled the dataset with types, i.e.  phone, url, address, name, sentence. Performed multitasking training, i.e. predicting the types of the sequences and predicting the phrases as the same time, which actually dropped the LB. 3) Added a penalty in the loss function of Mark's to penalize the length difference between the true phrase and predicted phrase, which didn't improve the LB. More attempts will be discussed later. \n\n# Details of the submission\n## Data processing:\n1. Loaded original train_landmarks and supplemental_landmarks dataset to my google drive.\n2. Split each original parquet file (contains about 1000 sequences) to multiple parquet files by sequence_id. Removed all the frames with no hand landmarks. It turned out the LB will be increased by a lot if leaving half of the missing-hand frames in the data according to [CHRIS DEOTTE's solution](https://www.kaggle.com/competitions/asl-fingerspelling/discussion/434353). Calculated the frame-to-phrase ratios using the number of frames that have hand landmarks, and saved them into the train_ratio.csv and supple_ratio.csv files.\n3. Created a [new Kaggle dataset](https://www.kaggle.com/datasets/joqueen/aslfr-all-landmarks-nonanhand) with the new parquet files (one parquet means one sequence) and new csv files.\n4. Created tfrecords files from the new dataset in step 3 using the [Kaggle notebook](https://www.kaggle.com/joqueen/yu-aslfr-train-supple-cleaned-landmarks-tfr). \n4.1. Removed the samples in supplemental dataset that have phrase length longer than 31, as the max phrase length of train dataset is 31. \n4.2 Removed the duplicates in supplemental dataset that has the same participant_id and phrase, as there are a lot duplicates in phrase and in participant_id in supplemental dataset.\n4.3 Removed the samples with frame-to-phrase ratios smaller than 0.5, 1, or 2. It turned out when ratio equals to 1, it gave me the best LB.\n4.4 Split the data to 5 folds based on the phrase type and participant_id. Saved the data as tfrecord files as a [new Kaggle dataset](https://www.kaggle.com/datasets/joqueen/aslfr-all-5fold-grpsplit-cln-1p0-31phrlen-dedup).\n\n## Training:\n1. As described in the overview, the main body of my solution is from [Mark Wijkhuizen](https://www.kaggle.com/code/markwijkhuizen/aslfr-transformer-training-inference) and [hoyso48](https://www.kaggle.com/competitions/asl-signs/discussion/406684), i.e. 1DConv+Transformer+Augmentation. As I trained my model using TPU in Colab, here is my [Colab notebook](https://colab.research.google.com/drive/1IOkmBN558p8hB_bfF-g4KO7SrvQiFshC?usp=sharing) for training. I'm not going to repeat the details again here, but I think there were a few learnings/observations worth mentioning:\n2. Mark padded the phrase length to 128 on the decoder side, the same as the frame length on the encoder side. We don't have to match the dimension. The max phrase length is only 31 and already padded to 32. Instead of padding phrase further to 128, I modified the MHA block to be able to take two different dimensions, i.e. q_len and kv_len. To be honest, it didn't improve or decrease my LB, so Mark's MHA is perfectly fine, but it did save me some training time. \n```python\nclass MultiHeadAttention(tf.keras.layers.Layer):\n    def __init__(self, d_model, n_heads, q_len, kv_len, dropout): \n        super(MultiHeadAttention,self).__init__()\n        self.d_model = d_model\n        self.n_heads = n_heads\n        self.kv_len = kv_len\n        self.q_len = q_len                                              \n        self.depth = d_model // 2                                               \n        self.scale = 1.0 / tf.math.sqrt(tf.cast(self.depth, tf.float32))        \n        self.wq = self.fused_mha(self.depth, q_len)                             \n        self.wk = self.fused_mha(self.depth, kv_len)                            \n        self.wv = self.fused_mha(self.depth, kv_len)                            \n        self.wo = tf.keras.layers.Dense(d_model, use_bias=False)\n        self.softmax = tf.keras.layers.Softmax()\n        self.reshape = tf.keras.Sequential([\n            tf.keras.layers.Permute([2, 1, 3]),\n            tf.keras.layers.Reshape([self.q_len, self.depth]),\n        ])\n        self.do = tf.keras.layers.Dropout(dropout)\n        self.supports_masking = True\n\n    def fused_mha(self, dim, len):\n        return tf.keras.Sequential([\n            tf.keras.layers.Dense(dim, use_bias=False),\n            tf.keras.layers.Reshape([len, self.n_heads, dim // self.n_heads]),\n            tf.keras.layers.Permute([2, 1, 3]),\n        ])\n\n    def call(self, q, k, v, attention_mask=None, training=False):\n        Q = self.wq(q)                                                          \n        K = self.wk(k)                                                          \n        V = self.wv(v)                                                          \n        x = tf.matmul(Q, K, transpose_b=True) * self.scale                      \n        x = self.softmax(x, mask=attention_mask) @ V                            \n        x = self.reshape(x)                                                     \n        x = self.wo(x)\n        x = self.do(x, training=training)\n        return x\n```\n3. I increased the mha_dropout_ratio, mlp_dropout_ratio of decoder by 1.5 times and clf_dropout_ratio of the classifier by 2 times, to heavily regularize the decoder and prevent overfitting. It improved my LB by 0.006. This idea is inspired by a discussion commented by hoyso48 .\n```python\n    # Decoder\n    x = Decoder(n_dec_blocks,\n                dec_dim,\n                n_mha_heads,\n                mha_dropout_ratio*1.5, \n                mlp_ratio,\n                mlp_dropout_ratio*1.5, \n                frame_len,\n                phrase_len,\n                )(x, phrase_inp, frames_inp)\n\n    # Classifier\n    x = tf.keras.Sequential([\n        tf.keras.layers.Dropout(clf_dropout_ratio*2),  \n        tf.keras.layers.Dense(N_UNIQUE_CHARACTERS,\n                              activation=tf.keras.activations.linear,\n                              kernel_initializer=INIT_HE_UNIFORM,\n                              use_bias=False),\n    ], name='classifier')(x)\n```\n\n## Inference\nMy inference followed Mark's inference method. Here is the [inference Kaggle notebook](https://www.kaggle.com/joqueen/m12-yu-aslfr-inference), which reported both my highest public score 0.713 and my highest private score 0.665.\n\n# What tried but didn't work (well)\n## Loss penalization\nIn Mark's customized loss function, y_true and y_pred got truncated by the y_true length. y_true length is known, so how to truncate the y_pred is not from training process. I was thinking what if the y_pred is very long, but got truncated in the loss function, which may not really represent the loss.\n```python\ndef loss_w_ls(y_true, y_pred):\n    idxs = tf.where(y_true != PAD_IDX)\n    y_true = tf.gather_nd(y_true, idxs)\n    y_pred = tf.gather_nd(y_pred, idxs)\n    y_true = tf.cast(y_true, tf.int32)\n    y_true = tf.one_hot(y_true, N_UNIQUE_CHARACTERS, axis=1)\n    loss = tf.keras.losses.categorical_crossentropy(y_true, y_pred, label_smoothing=0.25, from_logits=True) # 0.25\n    loss = tf.math.reduce_mean(loss)\n    return loss\n```\nSo I added a penalty about the length difference between the y_true and y_pred, to force the model to generate y_pred with a similar length as y_true while keep the categorical_crossentropy optimal.\n```python\ndef loss_w_ls(y_true, y_pred):\n    # penalize different len btw y_true and y_pred\n    y_true_len = tf.cast(tf.argmax(tf.cast(tf.math.equal(y_true, PAD_IDX), tf.int32),axis=1), tf.int32)\n\n    _idx = tf.argmax(y_pred, axis=2)\n    _bool = tf.math.equal(_idx, EOS_IDX)\n    _len1 = tf.cast(tf.argmax(tf.cast(_bool, tf.int32), axis=1), tf.int32)\n\n    _no_eos = tf.math.logical_not(tf.reduce_any(_bool, axis=1))\n    _len2 = tf.cast(_no_eos, tf.int32) * CFG.phrase_len\n    y_pred_len = _len1 + _len2\n    # y_pred_len = tf.cond(y_pred_eos, lambda:y_pred_len, lambda:tf.ones_like(y_pred_len)*CFG.phrase_len)\n    penal = tf.cast(tf.abs(y_true_len-y_pred_len), tf.float32)\n    penal = tf.math.reduce_mean(penal)\n\n    # Filter Pad Tokens\n    idxs = tf.where(y_true != PAD_IDX)\n    y_true = tf.gather_nd(y_true, idxs)\n    y_pred = tf.gather_nd(y_pred, idxs)\n    y_true = tf.cast(y_true, tf.int32)\n    y_true = tf.one_hot(y_true, N_UNIQUE_CHARACTERS, axis=1)\n    loss = tf.keras.losses.categorical_crossentropy(y_true, y_pred, label_smoothing=0.25, from_logits=True) # 0.25\n    loss = tf.math.reduce_mean(loss)\n\n    return loss + CFG.loss_coeff * penal\n```\nI was quite excited about this idea, and managed to code and run it through. But it still reduced my LB by 0.001 even after some fine-tuning. (sad face)\n\n## multi-tasking training\nInspired by Mark's solution, I categorized the train+supplemental data to five types, phone_number, url, address, name_like, sentence. It is obvious that each type is so different from the other types, in terms of the content, the length, the structure. I was thinking, maybe I can do multi-tasking - let the model do classification and phrase-prediction at the same time, which may force the model to figure out the different patterns for different types of samples. But this actually reduced my LB by about 0.1, which is a lot to me. \n```python\n        model.compile(\n            optimizer=opt,\n            loss = [tf.keras.losses.CategoricalCrossentropy(from_logits=True,label_smoothing=0.1), loss_w_ls],\n            metrics = [[tf.keras.metrics.CategoricalAccuracy()], [TopKAccuracy(1)]], \n            steps_per_execution=steps_per_epoch,\n        )\n```\n\n## positional encoding\nMark's method used trainable positional encoding. I switched it to the sin/cos positional encoding for encoder only, and actually it reduced my LB by 0.005. But I didn't try to switch it for decoder. \n\n## a validation strategy\nI was struggling a lot on designing a good validation strategy that has the CV correlated to the LB well. Out of desperation, I even kept the all the validation data in their original form - before removing the missing-hand frames and removing the samples with low frame-to-phrase ratios. And calculated the Levenshtein distance on the raw valid data. However, it was very time consuming, my model training only took 45 min - 1.5 hour, depending on the parameters. But the Levenshtein distance calculation for validation took about 3 hours. And even with all the effort, I still didn't see strong correlation. \n\n## CTC\nI tried out CTC a bit following the notebooks from [ROHITH INGILELA](https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place) and [GREYSNOW](https://www.kaggle.com/code/shlomoron/aslfr-ctc-on-tpu). But due to time limitation, I didn't really figure it out in time. And it looks like the majority of the top-tier solutions used CTC.\n\n# Sources\nhttps://www.kaggle.com/code/irohith/aslfr-transformer\nhttps://www.kaggle.com/code/markwijkhuizen/aslfr-transformer-training-inference) \nhttps://www.kaggle.com/competitions/asl-signs/discussion/406684\nhttps://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place\nhttps://www.kaggle.com/code/shlomoron/aslfr-ctc-on-tpu\nhttps://www.coursera.org/learn/nlp-sequence-models?specialization=deep-learning (transformer network)\n\nThis is my first competition and first solution write-up. Please let me know if I missed anything or if any part of my explanation was unclear. Suggestions and discussions are welcome. Thank you.\n",
    "2546624": "Nice work!",
    "2677401": "I can understand this. All other solutions are just tons of lines of code. Thanks for a great share :)"
  }
}