{
  "id": 434896,
  "title": "[28th Place Solution] Pre Training on Supplemental Data gives 0.01 LB Improvement",
  "url": "/competitions/asl-fingerspelling/discussion/434896",
  "author_name": "HW",
  "post_date": "2023-08-27T04:50:46.081000",
  "votes": 6,
  "comment_count": 3,
  "views": 0,
  "content": "<h2>Acknowledgements</h2>\n<p>Firstly, I'd like to extend my deepest gratitude to my teammates:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/chellyfan\" target=\"_blank\">@chellyfan</a></li>\n<li><a href=\"https://www.kaggle.com/chuansj\" target=\"_blank\">@chuansj</a></li>\n</ul>\n<p>Their enormous effort has been instrumental in this competition. </p>\n<p>I'd like also thanks to the authors of following notebooks. Our success wouldn't have been possible without the insights and techniques we adopted from these notebooks. A big thanks to you.</p>\n<ul>\n<li><p><strong>@irohith</strong> for his excellent preprocessing and training notebooks:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place\" target=\"_blank\">ASLFR CTC based on prev comp 1st place</a></li>\n<li><a href=\"https://www.kaggle.com/code/irohith/aslfr-preprocess-dataset-tfrecords-mean-std\" target=\"_blank\">ASLFR preprocess dataset TFRecords mean std</a></li></ul></li>\n<li><p><strong>@hoyso48</strong> for his 1st place solution from a previous competition:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/hoyso48/1st-place-solution-training\" target=\"_blank\">1st Place Solution Training</a></li></ul></li>\n<li><p><strong>@greysnow</strong> for the CTC on TPU notebook:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/shlomoron/aslfr-ctc-on-tpu\" target=\"_blank\">ASLFR CTC on TPU</a></li></ul></li>\n<li><p><strong>@royalacecat</strong> for the notebook on more blocks with quantization:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/royalacecat/the-deeper-the-better\" target=\"_blank\">The deeper, the better</a></li></ul></li>\n</ul>\n<h2>TL;DR</h2>\n<p>Our solution is based on the CTC open source notebook. Here are the key improvements we introduced (all scores are public LB):</p>\n<ul>\n<li>Only filter <strong>full NaN frames</strong> when generating tfrecord files on data preprocessing. (0.687-&gt;0.705)</li>\n<li>Adding more numbers of block in the model as described in <a href=\"https://www.kaggle.com/code/royalacecat/the-deeper-the-better\" target=\"_blank\">The deeper, the better</a>. (0.705-&gt;0.716)</li>\n<li><strong>Increased frame length</strong> from 128 to 256. (0.716-&gt;0.746)</li>\n<li><strong>Enhanced feature selection</strong> by adding more face and pose indices, increase from 96 to 128 indices. (0.746-&gt;0.763)</li>\n<li>Incorporated <strong>CTC loss with label smoothing</strong> for added regularization. (0.763 -&gt;0.766)</li>\n<li><strong>Pretrained our model</strong> on the supplemental dataset. Pre train 80 epochs on supply set then 80 epochs on main dataset. (0.766-&gt;0.777)</li>\n<li>Adopted the <strong>motion feature</strong> technique from a previous competition. (0.777-&gt;0.781)</li>\n<li>Increase training epochs on main dataset from 80 to 120 (0.781-&gt;0.782)</li>\n</ul>\n<h2>Data Preprocessing</h2>\n<p>Referring to <a href=\"https://www.kaggle.com/irohith\" target=\"_blank\">@irohith</a>'s preprocessing notebook, we made changes by adding more face and pose indices. We selected the new indices based on their variance — higher variance indicates more information.</p>\n<p>Initially, the indices were:</p>\n<pre><code>LIP = [\n    , , , , , , , , , ,\n    , , , , , , , , , ,\n    , , , , , , , , , ,\n    , , , , , , , , , ,\n]\nLPOSE = [, , , , ]\nRPOSE = [, , , , ]\nPOSE = LPOSE + RPOSE\n</code></pre>\n<p>We then updated them to:</p>\n<pre><code>LIP = [\n    , , , , , , , , , ,\n    , , , , , , , , , ,\n    , , , , , , , , , ,\n    , , , , , , , , , ,\n]\n\n\nface_id = [,,,,,,,,,,,\n           ,,,,,,,,,,,,\n           ,,,,,,,,,]\n\n\n k  face_id:\n     k   LIP:\n        LIP.append(k)\n\n\nl = (LIP)\nLIP  = LIP[:(l - l/)]\n\n\nLPOSE = [, , , , , , , , , , , , , , , ]\nRPOSE = [, , , , , , , , , , , , , , , ]\nPOSE = LPOSE + RPOSE\n</code></pre>\n<p>filter out full Nan frame</p>\n<p>origin notebook:</p>\n<pre><code>hand = tf.concat([rhand, lhand], axis=)\nhand = tf.where(tf.math.is_nan(hand), , hand)\nmask = tf.math.not_equal(tf.reduce_sum(hand, axis=[, ]), )\n</code></pre>\n<p>update to </p>\n<pre><code>hand = tf.concat([rhand, lhand,lip,rpose,lpose], axis=)\nhand = tf.where(tf.math.is_nan(hand), , hand)\nmask = tf.math.not_equal(tf.reduce_sum(hand, axis=[, ]), )\n</code></pre>\n<p>Those new indices along with 256 frame length gives us about 0.02 LB improvement!</p>\n<h2>Motion feature</h2>\n<p>see  <a href=\"https://www.kaggle.com/code/hoyso48/1st-place-solution-training\" target=\"_blank\">1st Place Solution Training</a></p>\n<h2>CTC Loss with Label Smoothing</h2>\n<p>To enhance our model's stability during training, we implemented a CTC loss with an added regularization component. By combining the CTC loss with a Kullback-Leibler (KL) divergence, we introduced label smoothing to the model. Smooth weight = 0.7 works best for our model</p>\n<pre><code> ():\n    \n    label_length = tf.reduce_sum(tf.cast(labels != pad_token_idx, tf.int32), axis=-)\n    logit_length = tf.ones(tf.shape(logits)[], dtype=tf.int32) * tf.shape(logits)[]\n    ctc_loss = tf.nn.ctc_loss(\n        labels=labels,\n        logits=logits,\n        label_length=label_length,\n        logit_length=logit_length,\n        blank_index=pad_token_idx,\n        logits_time_major=\n    )\n    ctc_loss = tf.reduce_mean(ctc_loss)\n\n    \n    kl_inp = tf.nn.softmax(logits)\n\n    \n    kl_tar = tf.fill(tf.shape(logits),  / num_classes)\n\n    \n    kldiv_loss = (tf.keras.losses.KLDivergence(tf.keras.losses.Reduction.NONE)(kl_tar, kl_inp) \n                 + tf.keras.losses.KLDivergence(tf.keras.losses.Reduction.NONE)(kl_inp, kl_tar))/\n\n    kldiv_loss = tf.reduce_mean(kldiv_loss)\n\n    \n    loss = ( - weight) * ctc_loss + weight * kldiv_loss\n     loss\n</code></pre>\n<h2>Pre-training the Model on the Supplemental Dataset</h2>\n<p>During our experimentation, we found a significant insight: training solely on the <strong>supplemental dataset</strong> yielded a LB score of 0.367. This highlighted the potential value embedded within the supplemental data.</p>\n<p>We decided to first pre-train our model on the supplemental dataset, then loaded these pre-trained weights to further train the model on the main dataset. We tested multiple epochs: 40, 60, and 80 for pre-train. 80 epochs works best and give a improvement of 0.013 on the LB score.</p>\n<h2>What We Tried, But Didn't Work</h2>\n<ol>\n<li><p><strong>Augmentation</strong>:</p>\n<ul>\n<li><strong>Random Affine Transformation</strong></li>\n<li><strong>Random Part Removal</strong>: We tried dropping certain parts (like lip, lpose, rpose, lhand, rhand) by setting them to NaN with a 5% probability. </li>\n<li><strong>Feature Switch for Phrases</strong>: switch features for phrases within the same category.</li></ul>\n<p>In 2nd place solution, <a href=\"https://www.kaggle.com/hoyso48\" target=\"_blank\">@hoyso48</a> mention that some augmentation didn't improve the model during short epoch training but in longer training would have the effect. This could applied to our case as well but we haven't verified yet.</p></li>\n<li><p><strong>Switching to GIC-CTC</strong>:</p>\n<ul>\n<li><strong>Reference Paper</strong>: <a href=\"https://ieeexplore.ieee.org/document/10094820\" target=\"_blank\">Improving CTC-Based ASR Models With Gated Interlayer Collaboration</a></li>\n<li>Increasing parameters led to our model exceeding the 5-hour inference time limit.</li></ul></li>\n</ol>\n<h2>What We could Imporve</h2>\n<ol>\n<li>long epochs training. We only train 120 epochs for final submission, which is much less compare to other top place solution (usually 300-500 epochs). We already saw a 0.001 LB improve from 80 epochs to 120 epochs on the final day, but it's too late</li>\n<li>Adjust the leaning rate and weight decay more carefully. We simply use the cosine decay lr scheduler in the public notebook and set weight decay to 0.05</li>\n<li>Utilized AWP for longer epoch training</li>\n</ol>\n<h2>Notebook:</h2>\n<p><a href=\"https://www.kaggle.com/code/wuhongrui/aslfr-gcs-path\" target=\"_blank\">https://www.kaggle.com/code/wuhongrui/aslfr-gcs-path</a><br>\n<a href=\"https://www.kaggle.com/wuhongrui/aslfr-training\" target=\"_blank\">https://www.kaggle.com/wuhongrui/aslfr-training</a><br>\n<a href=\"https://www.kaggle.com/wuhongrui/ctc-final-inference-notebook\" target=\"_blank\">https://www.kaggle.com/wuhongrui/ctc-final-inference-notebook</a></p>\n<p>Training notebook is written specific for training on Google Colab. Frist use gcs-path notebook check the path in training notebook. Set USE_SUPPLY = True first to train supply set, then set  USE_VAL = True to load pre-train weight and train main data set.</p>\n<p>To use it on Kaggle TPU, change the file address to corresponding Kaggle input address and you NEED to modify the CTC loss. Check <a href=\"https://www.kaggle.com/code/shlomoron/aslfr-ctc-on-tpu\" target=\"_blank\">ASLFR CTC on TPU</a></p>",
  "messages": [
    {
      "id": 2410560,
      "postDate": "2023-08-27T04:50:46.080Z",
      "content": "<h2>Acknowledgements</h2>\n<p>Firstly, I'd like to extend my deepest gratitude to my teammates:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/chellyfan\" target=\"_blank\">@chellyfan</a></li>\n<li><a href=\"https://www.kaggle.com/chuansj\" target=\"_blank\">@chuansj</a></li>\n</ul>\n<p>Their enormous effort has been instrumental in this competition. </p>\n<p>I'd like also thanks to the authors of following notebooks. Our success wouldn't have been possible without the insights and techniques we adopted from these notebooks. A big thanks to you.</p>\n<ul>\n<li><p><strong>@irohith</strong> for his excellent preprocessing and training notebooks:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place\" target=\"_blank\">ASLFR CTC based on prev comp 1st place</a></li>\n<li><a href=\"https://www.kaggle.com/code/irohith/aslfr-preprocess-dataset-tfrecords-mean-std\" target=\"_blank\">ASLFR preprocess dataset TFRecords mean std</a></li></ul></li>\n<li><p><strong>@hoyso48</strong> for his 1st place solution from a previous competition:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/hoyso48/1st-place-solution-training\" target=\"_blank\">1st Place Solution Training</a></li></ul></li>\n<li><p><strong>@greysnow</strong> for the CTC on TPU notebook:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/shlomoron/aslfr-ctc-on-tpu\" target=\"_blank\">ASLFR CTC on TPU</a></li></ul></li>\n<li><p><strong>@royalacecat</strong> for the notebook on more blocks with quantization:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/royalacecat/the-deeper-the-better\" target=\"_blank\">The deeper, the better</a></li></ul></li>\n</ul>\n<h2>TL;DR</h2>\n<p>Our solution is based on the CTC open source notebook. Here are the key improvements we introduced (all scores are public LB):</p>\n<ul>\n<li>Only filter <strong>full NaN frames</strong> when generating tfrecord files on data preprocessing. (0.687-&gt;0.705)</li>\n<li>Adding more numbers of block in the model as described in <a href=\"https://www.kaggle.com/code/royalacecat/the-deeper-the-better\" target=\"_blank\">The deeper, the better</a>. (0.705-&gt;0.716)</li>\n<li><strong>Increased frame length</strong> from 128 to 256. (0.716-&gt;0.746)</li>\n<li><strong>Enhanced feature selection</strong> by adding more face and pose indices, increase from 96 to 128 indices. (0.746-&gt;0.763)</li>\n<li>Incorporated <strong>CTC loss with label smoothing</strong> for added regularization. (0.763 -&gt;0.766)</li>\n<li><strong>Pretrained our model</strong> on the supplemental dataset. Pre train 80 epochs on supply set then 80 epochs on main dataset. (0.766-&gt;0.777)</li>\n<li>Adopted the <strong>motion feature</strong> technique from a previous competition. (0.777-&gt;0.781)</li>\n<li>Increase training epochs on main dataset from 80 to 120 (0.781-&gt;0.782)</li>\n</ul>\n<h2>Data Preprocessing</h2>\n<p>Referring to <a href=\"https://www.kaggle.com/irohith\" target=\"_blank\">@irohith</a>'s preprocessing notebook, we made changes by adding more face and pose indices. We selected the new indices based on their variance — higher variance indicates more information.</p>\n<p>Initially, the indices were:</p>\n<pre><code>LIP = [\n    , , , , , , , , , ,\n    , , , , , , , , , ,\n    , , , , , , , , , ,\n    , , , , , , , , , ,\n]\nLPOSE = [, , , , ]\nRPOSE = [, , , , ]\nPOSE = LPOSE + RPOSE\n</code></pre>\n<p>We then updated them to:</p>\n<pre><code>LIP = [\n    , , , , , , , , , ,\n    , , , , , , , , , ,\n    , , , , , , , , , ,\n    , , , , , , , , , ,\n]\n\n\nface_id = [,,,,,,,,,,,\n           ,,,,,,,,,,,,\n           ,,,,,,,,,]\n\n\n k  face_id:\n     k   LIP:\n        LIP.append(k)\n\n\nl = (LIP)\nLIP  = LIP[:(l - l/)]\n\n\nLPOSE = [, , , , , , , , , , , , , , , ]\nRPOSE = [, , , , , , , , , , , , , , , ]\nPOSE = LPOSE + RPOSE\n</code></pre>\n<p>filter out full Nan frame</p>\n<p>origin notebook:</p>\n<pre><code>hand = tf.concat([rhand, lhand], axis=)\nhand = tf.where(tf.math.is_nan(hand), , hand)\nmask = tf.math.not_equal(tf.reduce_sum(hand, axis=[, ]), )\n</code></pre>\n<p>update to </p>\n<pre><code>hand = tf.concat([rhand, lhand,lip,rpose,lpose], axis=)\nhand = tf.where(tf.math.is_nan(hand), , hand)\nmask = tf.math.not_equal(tf.reduce_sum(hand, axis=[, ]), )\n</code></pre>\n<p>Those new indices along with 256 frame length gives us about 0.02 LB improvement!</p>\n<h2>Motion feature</h2>\n<p>see  <a href=\"https://www.kaggle.com/code/hoyso48/1st-place-solution-training\" target=\"_blank\">1st Place Solution Training</a></p>\n<h2>CTC Loss with Label Smoothing</h2>\n<p>To enhance our model's stability during training, we implemented a CTC loss with an added regularization component. By combining the CTC loss with a Kullback-Leibler (KL) divergence, we introduced label smoothing to the model. Smooth weight = 0.7 works best for our model</p>\n<pre><code> ():\n    \n    label_length = tf.reduce_sum(tf.cast(labels != pad_token_idx, tf.int32), axis=-)\n    logit_length = tf.ones(tf.shape(logits)[], dtype=tf.int32) * tf.shape(logits)[]\n    ctc_loss = tf.nn.ctc_loss(\n        labels=labels,\n        logits=logits,\n        label_length=label_length,\n        logit_length=logit_length,\n        blank_index=pad_token_idx,\n        logits_time_major=\n    )\n    ctc_loss = tf.reduce_mean(ctc_loss)\n\n    \n    kl_inp = tf.nn.softmax(logits)\n\n    \n    kl_tar = tf.fill(tf.shape(logits),  / num_classes)\n\n    \n    kldiv_loss = (tf.keras.losses.KLDivergence(tf.keras.losses.Reduction.NONE)(kl_tar, kl_inp) \n                 + tf.keras.losses.KLDivergence(tf.keras.losses.Reduction.NONE)(kl_inp, kl_tar))/\n\n    kldiv_loss = tf.reduce_mean(kldiv_loss)\n\n    \n    loss = ( - weight) * ctc_loss + weight * kldiv_loss\n     loss\n</code></pre>\n<h2>Pre-training the Model on the Supplemental Dataset</h2>\n<p>During our experimentation, we found a significant insight: training solely on the <strong>supplemental dataset</strong> yielded a LB score of 0.367. This highlighted the potential value embedded within the supplemental data.</p>\n<p>We decided to first pre-train our model on the supplemental dataset, then loaded these pre-trained weights to further train the model on the main dataset. We tested multiple epochs: 40, 60, and 80 for pre-train. 80 epochs works best and give a improvement of 0.013 on the LB score.</p>\n<h2>What We Tried, But Didn't Work</h2>\n<ol>\n<li><p><strong>Augmentation</strong>:</p>\n<ul>\n<li><strong>Random Affine Transformation</strong></li>\n<li><strong>Random Part Removal</strong>: We tried dropping certain parts (like lip, lpose, rpose, lhand, rhand) by setting them to NaN with a 5% probability. </li>\n<li><strong>Feature Switch for Phrases</strong>: switch features for phrases within the same category.</li></ul>\n<p>In 2nd place solution, <a href=\"https://www.kaggle.com/hoyso48\" target=\"_blank\">@hoyso48</a> mention that some augmentation didn't improve the model during short epoch training but in longer training would have the effect. This could applied to our case as well but we haven't verified yet.</p></li>\n<li><p><strong>Switching to GIC-CTC</strong>:</p>\n<ul>\n<li><strong>Reference Paper</strong>: <a href=\"https://ieeexplore.ieee.org/document/10094820\" target=\"_blank\">Improving CTC-Based ASR Models With Gated Interlayer Collaboration</a></li>\n<li>Increasing parameters led to our model exceeding the 5-hour inference time limit.</li></ul></li>\n</ol>\n<h2>What We could Imporve</h2>\n<ol>\n<li>long epochs training. We only train 120 epochs for final submission, which is much less compare to other top place solution (usually 300-500 epochs). We already saw a 0.001 LB improve from 80 epochs to 120 epochs on the final day, but it's too late</li>\n<li>Adjust the leaning rate and weight decay more carefully. We simply use the cosine decay lr scheduler in the public notebook and set weight decay to 0.05</li>\n<li>Utilized AWP for longer epoch training</li>\n</ol>\n<h2>Notebook:</h2>\n<p><a href=\"https://www.kaggle.com/code/wuhongrui/aslfr-gcs-path\" target=\"_blank\">https://www.kaggle.com/code/wuhongrui/aslfr-gcs-path</a><br>\n<a href=\"https://www.kaggle.com/wuhongrui/aslfr-training\" target=\"_blank\">https://www.kaggle.com/wuhongrui/aslfr-training</a><br>\n<a href=\"https://www.kaggle.com/wuhongrui/ctc-final-inference-notebook\" target=\"_blank\">https://www.kaggle.com/wuhongrui/ctc-final-inference-notebook</a></p>\n<p>Training notebook is written specific for training on Google Colab. Frist use gcs-path notebook check the path in training notebook. Set USE_SUPPLY = True first to train supply set, then set  USE_VAL = True to load pre-train weight and train main data set.</p>\n<p>To use it on Kaggle TPU, change the file address to corresponding Kaggle input address and you NEED to modify the CTC loss. Check <a href=\"https://www.kaggle.com/code/shlomoron/aslfr-ctc-on-tpu\" target=\"_blank\">ASLFR CTC on TPU</a></p>",
      "rawMarkdown": "## Acknowledgements\n\nFirstly, I'd like to extend my deepest gratitude to my teammates:\n- @chellyfan\n- @chuansj\n\nTheir enormous effort has been instrumental in this competition. \n\nI'd like also thanks to the authors of following notebooks. Our success wouldn't have been possible without the insights and techniques we adopted from these notebooks. A big thanks to you.\n\n- **@irohith** for his excellent preprocessing and training notebooks:\n  - [ASLFR CTC based on prev comp 1st place](https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place)\n  - [ASLFR preprocess dataset TFRecords mean std](https://www.kaggle.com/code/irohith/aslfr-preprocess-dataset-tfrecords-mean-std)\n\n- **@hoyso48** for his 1st place solution from a previous competition:\n  - [1st Place Solution Training](https://www.kaggle.com/code/hoyso48/1st-place-solution-training)\n\n- **@greysnow** for the CTC on TPU notebook:\n  - [ASLFR CTC on TPU](https://www.kaggle.com/code/shlomoron/aslfr-ctc-on-tpu)\n\n- **@royalacecat** for the notebook on more blocks with quantization:\n  - [The deeper, the better](https://www.kaggle.com/code/royalacecat/the-deeper-the-better)\n\n## TL;DR\n\nOur solution is based on the CTC open source notebook. Here are the key improvements we introduced (all scores are public LB):\n\n- Only filter **full NaN frames** when generating tfrecord files on data preprocessing. (0.687->0.705)\n- Adding more numbers of block in the model as described in [The deeper, the better](https://www.kaggle.com/code/royalacecat/the-deeper-the-better). (0.705->0.716)\n- **Increased frame length** from 128 to 256. (0.716->0.746)\n- **Enhanced feature selection** by adding more face and pose indices, increase from 96 to 128 indices. (0.746->0.763)\n- Incorporated **CTC loss with label smoothing** for added regularization. (0.763 ->0.766)\n- **Pretrained our model** on the supplemental dataset. Pre train 80 epochs on supply set then 80 epochs on main dataset. (0.766->0.777)\n- Adopted the **motion feature** technique from a previous competition. (0.777->0.781)\n- Increase training epochs on main dataset from 80 to 120 (0.781->0.782)\n\n## Data Preprocessing\n\nReferring to @irohith's preprocessing notebook, we made changes by adding more face and pose indices. We selected the new indices based on their variance — higher variance indicates more information.\n\nInitially, the indices were:\n\n```python\nLIP = [\n    61, 185, 40, 39, 37, 0, 267, 269, 270, 409,\n    291, 146, 91, 181, 84, 17, 314, 405, 321, 375,\n    78, 191, 80, 81, 82, 13, 312, 311, 310, 415,\n    95, 88, 178, 87, 14, 317, 402, 318, 324, 308,\n]\nLPOSE = [13, 15, 17, 19, 21]\nRPOSE = [14, 16, 18, 20, 22]\nPOSE = LPOSE + RPOSE\n```\nWe then updated them to:\n\n```python\n\nLIP = [\n    61, 185, 40, 39, 37, 0, 267, 269, 270, 409,\n    291, 146, 91, 181, 84, 17, 314, 405, 321, 375,\n    78, 191, 80, 81, 82, 13, 312, 311, 310, 415,\n    95, 88, 178, 87, 14, 317, 402, 318, 324, 308,\n]\n\n# New face_id indices\nface_id = [454,356,323,361,389,288,251,264,447,366,368,\n           401,397,435,284,301,372,345,383,367,365,352,433,\n           376,298,265,93,234,300,132,340,353,127]\n\n# Append new face_id indices to LIP if not already present\nfor k in face_id:\n    if k not in LIP:\n        LIP.append(k)\n\n# Trim the LIP list\nl = len(LIP)\nLIP  = LIP[:int(l - l/4)]\n\n# Updated POSE indices\nLPOSE = [1, 3, 5, 7, 9, 11, 13, 15, 17, 19, 21, 23, 25, 27, 29, 31]\nRPOSE = [0, 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30]\nPOSE = LPOSE + RPOSE\n```\nfilter out full Nan frame\n\norigin notebook:\n\n```python\nhand = tf.concat([rhand, lhand], axis=1)\nhand = tf.where(tf.math.is_nan(hand), 0.0, hand)\nmask = tf.math.not_equal(tf.reduce_sum(hand, axis=[1, 2]), 0.0)\n```\nupdate to \n\n```python\nhand = tf.concat([rhand, lhand,lip,rpose,lpose], axis=1)\nhand = tf.where(tf.math.is_nan(hand), 0.0, hand)\nmask = tf.math.not_equal(tf.reduce_sum(hand, axis=[1, 2]), 0.0)\n```\nThose new indices along with 256 frame length gives us about 0.02 LB improvement!\n\n## Motion feature\nsee  [1st Place Solution Training](https://www.kaggle.com/code/hoyso48/1st-place-solution-training)\n\n\n## CTC Loss with Label Smoothing\nTo enhance our model's stability during training, we implemented a CTC loss with an added regularization component. By combining the CTC loss with a Kullback-Leibler (KL) divergence, we introduced label smoothing to the model. Smooth weight = 0.7 works best for our model\n```python\ndef smooth_ctc_loss(labels, logits, num_classes = 60 , blank=0, weight=0.7):\n    # Compute CTC Loss\n    label_length = tf.reduce_sum(tf.cast(labels != pad_token_idx, tf.int32), axis=-1)\n    logit_length = tf.ones(tf.shape(logits)[0], dtype=tf.int32) * tf.shape(logits)[1]\n    ctc_loss = tf.nn.ctc_loss(\n        labels=labels,\n        logits=logits,\n        label_length=label_length,\n        logit_length=logit_length,\n        blank_index=pad_token_idx,\n        logits_time_major=False\n    )\n    ctc_loss = tf.reduce_mean(ctc_loss)\n\n    # Compute KL Divergence Loss\n    kl_inp = tf.nn.softmax(logits)\n\n    # Create the target distribution\n    kl_tar = tf.fill(tf.shape(logits), 1. / num_classes)\n\n    # Compute the KL divergence\n    kldiv_loss = (tf.keras.losses.KLDivergence(tf.keras.losses.Reduction.NONE)(kl_tar, kl_inp) \n                 + tf.keras.losses.KLDivergence(tf.keras.losses.Reduction.NONE)(kl_inp, kl_tar))/2.0\n\n    kldiv_loss = tf.reduce_mean(kldiv_loss)\n\n    # Combined Loss\n    loss = (1. - weight) * ctc_loss + weight * kldiv_loss\n    return loss\n```\n## Pre-training the Model on the Supplemental Dataset\n\nDuring our experimentation, we found a significant insight: training solely on the **supplemental dataset** yielded a LB score of 0.367. This highlighted the potential value embedded within the supplemental data.\n\nWe decided to first pre-train our model on the supplemental dataset, then loaded these pre-trained weights to further train the model on the main dataset. We tested multiple epochs: 40, 60, and 80 for pre-train. 80 epochs works best and give a improvement of 0.013 on the LB score.\n\n## What We Tried, But Didn't Work\n\n1. **Augmentation**:\n   - **Random Affine Transformation**\n   - **Random Part Removal**: We tried dropping certain parts (like lip, lpose, rpose, lhand, rhand) by setting them to NaN with a 5% probability. \n   - **Feature Switch for Phrases**: switch features for phrases within the same category.\n\n   In 2nd place solution, @hoyso48 mention that some augmentation didn't improve the model during short epoch training but in longer training would have the effect. This could applied to our case as well but we haven't verified yet.\n\n2. **Switching to GIC-CTC**:\n   - **Reference Paper**: [Improving CTC-Based ASR Models With Gated Interlayer Collaboration](https://ieeexplore.ieee.org/document/10094820)\n   - Increasing parameters led to our model exceeding the 5-hour inference time limit.\n\n\n## What We could Imporve\n1. long epochs training. We only train 120 epochs for final submission, which is much less compare to other top place solution (usually 300-500 epochs). We already saw a 0.001 LB improve from 80 epochs to 120 epochs on the final day, but it's too late\n2. Adjust the leaning rate and weight decay more carefully. We simply use the cosine decay lr scheduler in the public notebook and set weight decay to 0.05\n3. Utilized AWP for longer epoch training\n\n## Notebook:\n \nhttps://www.kaggle.com/code/wuhongrui/aslfr-gcs-path\nhttps://www.kaggle.com/wuhongrui/aslfr-training\nhttps://www.kaggle.com/wuhongrui/ctc-final-inference-notebook\n\nTraining notebook is written specific for training on Google Colab. Frist use gcs-path notebook check the path in training notebook. Set USE_SUPPLY = True first to train supply set, then set  USE_VAL = True to load pre-train weight and train main data set.\n\nTo use it on Kaggle TPU, change the file address to corresponding Kaggle input address and you NEED to modify the CTC loss. Check [ASLFR CTC on TPU](https://www.kaggle.com/code/shlomoron/aslfr-ctc-on-tpu)\n\n\n\n\n\n\n\n\n  ",
      "votes": 6
    },
    {
      "id": 2410646,
      "postDate": "2023-08-27T06:19:15.497Z",
      "content": "<p>Thank you for sharing.</p>\n<blockquote>\n  <p>We decided to first pre-train our model on the supplemental dataset, then loaded these pre-trained weights to further train the model on the main dataset. We tested multiple epochs: 40, 60, and 80 for pre-train. 80 epochs works best and give a improvement of 0.013 on the LB score.</p>\n</blockquote>\n<p>We saw this too for early models. But as soon as models became better, pretraining on supplemental did not have effect anymore. Maybe also because our better models are trained longer (just using cutmix doubles the number of optimal epochs from 200 to 400 for us) and weights initialization is not that important then.</p>",
      "rawMarkdown": "Thank you for sharing.\n\n>We decided to first pre-train our model on the supplemental dataset, then loaded these pre-trained weights to further train the model on the main dataset. We tested multiple epochs: 40, 60, and 80 for pre-train. 80 epochs works best and give a improvement of 0.013 on the LB score.\n\nWe saw this too for early models. But as soon as models became better, pretraining on supplemental did not have effect anymore. Maybe also because our better models are trained longer (just using cutmix doubles the number of optimal epochs from 200 to 400 for us) and weights initialization is not that important then.",
      "votes": 2,
      "replies": [
        {
          "id": 2410669,
          "postDate": "2023-08-27T06:38:27.903Z",
          "content": "<p>Thanks for your great insight! It's interesting to see the impact of pre-training and the number of epochs on your model's performance. We overlooked importance of the long epoch training in this competition, and as you mentioned, the weights initialization maybe is less significant for long epoch.</p>\n<p>Congratulation again for your amazing solution of 1st place in this competition !</p>",
          "rawMarkdown": "Thanks for your great insight! It's interesting to see the impact of pre-training and the number of epochs on your model's performance. We overlooked importance of the long epoch training in this competition, and as you mentioned, the weights initialization maybe is less significant for long epoch.\n\nCongratulation again for your amazing solution of 1st place in this competition !",
          "votes": 2
        }
      ]
    },
    {
      "id": 2412159,
      "postDate": "2023-08-28T05:51:35.923Z",
      "content": "<p>Congratulations. Just improving the score with long no of epochs can backfire if there is overfitting or the model encounter local minima. </p>",
      "rawMarkdown": "Congratulations. Just improving the score with long no of epochs can backfire if there is overfitting or the model encounter local minima. "
    }
  ],
  "comments": [
    {
      "id": 2410646,
      "author_name": "Dieter",
      "author_url": "",
      "post_date": "2023-08-27T06:19:15.497000",
      "content": "<p>Thank you for sharing.</p>\n<blockquote>\n  <p>We decided to first pre-train our model on the supplemental dataset, then loaded these pre-trained weights to further train the model on the main dataset. We tested multiple epochs: 40, 60, and 80 for pre-train. 80 epochs works best and give a improvement of 0.013 on the LB score.</p>\n</blockquote>\n<p>We saw this too for early models. But as soon as models became better, pretraining on supplemental did not have effect anymore. Maybe also because our better models are trained longer (just using cutmix doubles the number of optimal epochs from 200 to 400 for us) and weights initialization is not that important then.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2410669,
          "author_name": "HW",
          "author_url": "",
          "post_date": "2023-08-27T06:38:27.903000",
          "content": "<p>Thanks for your great insight! It's interesting to see the impact of pre-training and the number of epochs on your model's performance. We overlooked importance of the long epoch training in this competition, and as you mentioned, the weights initialization maybe is less significant for long epoch.</p>\n<p>Congratulation again for your amazing solution of 1st place in this competition !</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2412159,
      "author_name": "C R Suthikshn Kumar",
      "author_url": "",
      "post_date": "2023-08-28T05:51:35.923000",
      "content": "<p>Congratulations. Just improving the score with long no of epochs can backfire if there is overfitting or the model encounter local minima. </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2410560": "## Acknowledgements\n\nFirstly, I'd like to extend my deepest gratitude to my teammates:\n- @chellyfan\n- @chuansj\n\nTheir enormous effort has been instrumental in this competition. \n\nI'd like also thanks to the authors of following notebooks. Our success wouldn't have been possible without the insights and techniques we adopted from these notebooks. A big thanks to you.\n\n- **@irohith** for his excellent preprocessing and training notebooks:\n  - [ASLFR CTC based on prev comp 1st place](https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place)\n  - [ASLFR preprocess dataset TFRecords mean std](https://www.kaggle.com/code/irohith/aslfr-preprocess-dataset-tfrecords-mean-std)\n\n- **@hoyso48** for his 1st place solution from a previous competition:\n  - [1st Place Solution Training](https://www.kaggle.com/code/hoyso48/1st-place-solution-training)\n\n- **@greysnow** for the CTC on TPU notebook:\n  - [ASLFR CTC on TPU](https://www.kaggle.com/code/shlomoron/aslfr-ctc-on-tpu)\n\n- **@royalacecat** for the notebook on more blocks with quantization:\n  - [The deeper, the better](https://www.kaggle.com/code/royalacecat/the-deeper-the-better)\n\n## TL;DR\n\nOur solution is based on the CTC open source notebook. Here are the key improvements we introduced (all scores are public LB):\n\n- Only filter **full NaN frames** when generating tfrecord files on data preprocessing. (0.687->0.705)\n- Adding more numbers of block in the model as described in [The deeper, the better](https://www.kaggle.com/code/royalacecat/the-deeper-the-better). (0.705->0.716)\n- **Increased frame length** from 128 to 256. (0.716->0.746)\n- **Enhanced feature selection** by adding more face and pose indices, increase from 96 to 128 indices. (0.746->0.763)\n- Incorporated **CTC loss with label smoothing** for added regularization. (0.763 ->0.766)\n- **Pretrained our model** on the supplemental dataset. Pre train 80 epochs on supply set then 80 epochs on main dataset. (0.766->0.777)\n- Adopted the **motion feature** technique from a previous competition. (0.777->0.781)\n- Increase training epochs on main dataset from 80 to 120 (0.781->0.782)\n\n## Data Preprocessing\n\nReferring to @irohith's preprocessing notebook, we made changes by adding more face and pose indices. We selected the new indices based on their variance — higher variance indicates more information.\n\nInitially, the indices were:\n\n```python\nLIP = [\n    61, 185, 40, 39, 37, 0, 267, 269, 270, 409,\n    291, 146, 91, 181, 84, 17, 314, 405, 321, 375,\n    78, 191, 80, 81, 82, 13, 312, 311, 310, 415,\n    95, 88, 178, 87, 14, 317, 402, 318, 324, 308,\n]\nLPOSE = [13, 15, 17, 19, 21]\nRPOSE = [14, 16, 18, 20, 22]\nPOSE = LPOSE + RPOSE\n```\nWe then updated them to:\n\n```python\n\nLIP = [\n    61, 185, 40, 39, 37, 0, 267, 269, 270, 409,\n    291, 146, 91, 181, 84, 17, 314, 405, 321, 375,\n    78, 191, 80, 81, 82, 13, 312, 311, 310, 415,\n    95, 88, 178, 87, 14, 317, 402, 318, 324, 308,\n]\n\n# New face_id indices\nface_id = [454,356,323,361,389,288,251,264,447,366,368,\n           401,397,435,284,301,372,345,383,367,365,352,433,\n           376,298,265,93,234,300,132,340,353,127]\n\n# Append new face_id indices to LIP if not already present\nfor k in face_id:\n    if k not in LIP:\n        LIP.append(k)\n\n# Trim the LIP list\nl = len(LIP)\nLIP  = LIP[:int(l - l/4)]\n\n# Updated POSE indices\nLPOSE = [1, 3, 5, 7, 9, 11, 13, 15, 17, 19, 21, 23, 25, 27, 29, 31]\nRPOSE = [0, 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30]\nPOSE = LPOSE + RPOSE\n```\nfilter out full Nan frame\n\norigin notebook:\n\n```python\nhand = tf.concat([rhand, lhand], axis=1)\nhand = tf.where(tf.math.is_nan(hand), 0.0, hand)\nmask = tf.math.not_equal(tf.reduce_sum(hand, axis=[1, 2]), 0.0)\n```\nupdate to \n\n```python\nhand = tf.concat([rhand, lhand,lip,rpose,lpose], axis=1)\nhand = tf.where(tf.math.is_nan(hand), 0.0, hand)\nmask = tf.math.not_equal(tf.reduce_sum(hand, axis=[1, 2]), 0.0)\n```\nThose new indices along with 256 frame length gives us about 0.02 LB improvement!\n\n## Motion feature\nsee  [1st Place Solution Training](https://www.kaggle.com/code/hoyso48/1st-place-solution-training)\n\n\n## CTC Loss with Label Smoothing\nTo enhance our model's stability during training, we implemented a CTC loss with an added regularization component. By combining the CTC loss with a Kullback-Leibler (KL) divergence, we introduced label smoothing to the model. Smooth weight = 0.7 works best for our model\n```python\ndef smooth_ctc_loss(labels, logits, num_classes = 60 , blank=0, weight=0.7):\n    # Compute CTC Loss\n    label_length = tf.reduce_sum(tf.cast(labels != pad_token_idx, tf.int32), axis=-1)\n    logit_length = tf.ones(tf.shape(logits)[0], dtype=tf.int32) * tf.shape(logits)[1]\n    ctc_loss = tf.nn.ctc_loss(\n        labels=labels,\n        logits=logits,\n        label_length=label_length,\n        logit_length=logit_length,\n        blank_index=pad_token_idx,\n        logits_time_major=False\n    )\n    ctc_loss = tf.reduce_mean(ctc_loss)\n\n    # Compute KL Divergence Loss\n    kl_inp = tf.nn.softmax(logits)\n\n    # Create the target distribution\n    kl_tar = tf.fill(tf.shape(logits), 1. / num_classes)\n\n    # Compute the KL divergence\n    kldiv_loss = (tf.keras.losses.KLDivergence(tf.keras.losses.Reduction.NONE)(kl_tar, kl_inp) \n                 + tf.keras.losses.KLDivergence(tf.keras.losses.Reduction.NONE)(kl_inp, kl_tar))/2.0\n\n    kldiv_loss = tf.reduce_mean(kldiv_loss)\n\n    # Combined Loss\n    loss = (1. - weight) * ctc_loss + weight * kldiv_loss\n    return loss\n```\n## Pre-training the Model on the Supplemental Dataset\n\nDuring our experimentation, we found a significant insight: training solely on the **supplemental dataset** yielded a LB score of 0.367. This highlighted the potential value embedded within the supplemental data.\n\nWe decided to first pre-train our model on the supplemental dataset, then loaded these pre-trained weights to further train the model on the main dataset. We tested multiple epochs: 40, 60, and 80 for pre-train. 80 epochs works best and give a improvement of 0.013 on the LB score.\n\n## What We Tried, But Didn't Work\n\n1. **Augmentation**:\n   - **Random Affine Transformation**\n   - **Random Part Removal**: We tried dropping certain parts (like lip, lpose, rpose, lhand, rhand) by setting them to NaN with a 5% probability. \n   - **Feature Switch for Phrases**: switch features for phrases within the same category.\n\n   In 2nd place solution, @hoyso48 mention that some augmentation didn't improve the model during short epoch training but in longer training would have the effect. This could applied to our case as well but we haven't verified yet.\n\n2. **Switching to GIC-CTC**:\n   - **Reference Paper**: [Improving CTC-Based ASR Models With Gated Interlayer Collaboration](https://ieeexplore.ieee.org/document/10094820)\n   - Increasing parameters led to our model exceeding the 5-hour inference time limit.\n\n\n## What We could Imporve\n1. long epochs training. We only train 120 epochs for final submission, which is much less compare to other top place solution (usually 300-500 epochs). We already saw a 0.001 LB improve from 80 epochs to 120 epochs on the final day, but it's too late\n2. Adjust the leaning rate and weight decay more carefully. We simply use the cosine decay lr scheduler in the public notebook and set weight decay to 0.05\n3. Utilized AWP for longer epoch training\n\n## Notebook:\n \nhttps://www.kaggle.com/code/wuhongrui/aslfr-gcs-path\nhttps://www.kaggle.com/wuhongrui/aslfr-training\nhttps://www.kaggle.com/wuhongrui/ctc-final-inference-notebook\n\nTraining notebook is written specific for training on Google Colab. Frist use gcs-path notebook check the path in training notebook. Set USE_SUPPLY = True first to train supply set, then set  USE_VAL = True to load pre-train weight and train main data set.\n\nTo use it on Kaggle TPU, change the file address to corresponding Kaggle input address and you NEED to modify the CTC loss. Check [ASLFR CTC on TPU](https://www.kaggle.com/code/shlomoron/aslfr-ctc-on-tpu)\n\n\n\n\n\n\n\n\n  ",
    "2410646": "Thank you for sharing.\n\n>We decided to first pre-train our model on the supplemental dataset, then loaded these pre-trained weights to further train the model on the main dataset. We tested multiple epochs: 40, 60, and 80 for pre-train. 80 epochs works best and give a improvement of 0.013 on the LB score.\n\nWe saw this too for early models. But as soon as models became better, pretraining on supplemental did not have effect anymore. Maybe also because our better models are trained longer (just using cutmix doubles the number of optimal epochs from 200 to 400 for us) and weights initialization is not that important then.",
    "2412159": "Congratulations. Just improving the score with long no of epochs can backfire if there is overfitting or the model encounter local minima. "
  }
}