{
  "id": 434588,
  "title": "2nd place solution - Test & Compare ASR algorithms",
  "url": "/competitions/asl-fingerspelling/discussion/434588",
  "author_name": "hoyso48",
  "post_date": "2023-08-25T18:24:52.443000",
  "votes": 77,
  "comment_count": 34,
  "views": 0,
  "content": "<p>Once again, thanks to Kaggle, Google and other organizers for hosting this exciting competition. I felt this one as well as the previous competition, is well-organized so that we can try various ideas and could learn a lot from each other. Hope this kinds of competition open in kaggle more often.. :)</p>\n<h2>TLDR</h2>\n<p>My overall solution is almost identical to the previous competition. It's mainly just the adoption of ASR(Automatic Speech Recognition) algorithms (joint CTC + Attention) on top of the previous 1st place model. The details of the model, preprocessing, and training are largely similar to the previous competition, so please refer to the <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406684\" target=\"_blank\">previous competition solution</a> for more details on model/training/augmentations/etc.</p>\n<h2>ASR Algorithms</h2>\n<p>I had a deep interest in ASR, and after recognizing its correlation with the competition, I studied and experimented with various algorithms. While I couldn't find anything particularly superior to the widely used vanilla CTC among the numerous ASR algorithms, I learned a lot through various experiments, and I will mainly share those insights.</p>\n<p>As I began studying ASR for this competition, I discovered from reading papers that current NN ASR algorithm mainly consists of CTC, Attention-based, and Transducer. I implemented all three (though there are more diverse algorithms out there). In conclusion, from a baseline performance perspective, all three algorithms were quite similar and each have pros &amp; cons. Here are the insights I gathered from implementing each algorithm:</p>\n<ul>\n<li><p>CTC:</p>\n<ul>\n<li>During greedy decoding, the prediction step = O(1). Therefore, when using a Single model, GreedyCTC is the most efficient. (this is probably why many solutions use a large single model with CTC).</li>\n<li>With beam search, prediction step = O(L) (where L = encoder output length). Depending on the efficiency of the algorithm, the computational overhead isn't very large, so the overhead by introducing beam search isn't very significant. However, the performance improvement from using beam search is minimal (+0.003), making it not very tempting(as its tflite-convertible implementation is not quite trivial) compared to Attention.</li>\n<li>As discussed, there's no guarantee of alignment between model prediction timesteps, so simple average ensembling can't be applied.</li></ul></li>\n<li><p>Attention-based:</p>\n<ul>\n<li>Uses an autoregressive approach, so even with greedy decoding, prediction step = O(N) (where N = decoder output length).</li>\n<li>If the decoder is an RNN or Transformer, stateful inference(i.e. previous key, value caching with Transformer) can be used to reduce the complexity. In actual implementation with Transformer, it reduced the inference time on the CPU by about 20~30%.</li>\n<li>Beam search is possible, but unlike CTC, it requires introducing appropriate heuristics to penalize the output length. Although several attempts were made, none worked. </li>\n<li>It's easier to apply ensembling with Attention. Simply take the average of each model at each prediction step, and it works well (ensembling three models gave +0.009).</li>\n<li>More room for score improvement than CTC(ex more decoder layers with augmentations).</li></ul></li>\n<li><p>Transducer:<br>\nPrediction steps = O(L) (where L = encoder output length). Due to the most prediction steps in a greedy manner, it's not very efficient. I didn't consider it further as optimization was harder and performance was slightly lower compared to CTC or Attention. However, it has the advantage of real-time recognition in streaming mode if the encoder is causal (which wasn't relevant for this competition).</p></li>\n</ul>\n<p>Rough single model inference time on test dataset with kaggle kernel is as follows.</p>\n<blockquote>\n  <p>CTC greedy(~40mins) &lt; AttentionGreedy=CTCBeamsearch(~1h20mins) &lt;=  CTCAttentionJointGreedy(~1h30mins)</p>\n</blockquote>\n<h2>CTC-Attention Joint Training &amp; Decoding</h2>\n<p>As both CTC and Attention showed similar performance, I tried to find a method to utilize both techniques.  I primarily referred to the following two papers:</p>\n<p><em>Joint CTC-Attention Based End-To-End Speech Recognition Using Multi-Task Learning, Kim et al. 2017.</em><br>\n<a href=\"url\" target=\"_blank\">https://arxiv.org/pdf/1609.06773.pdf</a><br>\n<em>Joint CTC/attention decoding for end-to-end speech recognition, Hori et al. 2017.</em><br>\n<a href=\"url\" target=\"_blank\">https://aclanthology.org/P17-1048.pdf</a></p>\n<p>Joint CTC-Attention Training is, as the name implies, adding both a CTC decoder (single GRU layer) and an attention decoder (single Transformer decoder layer) to one encoder for multitask learning. The loss weight is set to CTC=0.25 and attention=0.75. But the Joint training itself did not bring noticeable performance boost.</p>\n<p>By using the CTC prefix score, it is possible to calculate the probability of an arbitrary output hypothesis \"h\" without depending on the output timestep of the CTC output. In other words, <strong>implementing the CTC prefix score computation allows for ensembling outputs not only between CTC models but also between CTC and Attention models</strong>. CTC-attention joint decoding with CTC weight=0.3 showed a performance improvement of +0.007~8 without much impact on inference time. It was especially challenging for me to implement it accurately, efficiently, and without any issues to be compatible with tflite.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5003978%2Fabde23b44b90f8ba6adc7d56ca874d86%2F.drawio-3.svg?generation=1692977029060458&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5003978%2F22c798e256f66322a0cc9e42b976d570%2F.drawio-4.svg?generation=1692977118893507&amp;alt=media\" alt=\"\"></p>\n<h2>Preprocessing</h2>\n<p>Similar to the previous competition solution, but simpler. Every landmark and xyz was used and standardized. Flip left-handed signer(rather than augmentation). MAX_LEN=768 was used. Any hand-crafted feature was not significant, likely due to the absence of complex relations between movements in frames.</p>\n<h2>Model</h2>\n<p>Encoder is the same as the previous competition (Stacked Conv1DBlock + TransformerBlock) but increased in size (expand ratio 2-&gt;4 in Conv1DBlock) and depth (8 layers -&gt; 17 layers). Single model has ~6.5M parameters. I applied padding='same'(rather than 'causal') and output stride=2 (requires slightly more logic for handling masking). In my case, mixing in Transformer blocks wasn't as effective as in the previous competition. Perhaps global features were less crucial in this comp. Additionally, I added one BN to input of the Conv1DBlock for more training stability(especially with awp + more epochs).</p>\n<p>CTC decoder used a single GRU layer followed by one FC layer. Attention Decoder used a single-layer Transformer decoder. Introducing augmentation to the Decoder input and adding up to 4 Decoder layers improves the performance of the attention decoder (up to +0.004). However, considering the number of parameters and inference speed, I considered it inefficient and thus used a single-layer decoder.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5003978%2F7d1f7aee69242af5e6aa5ef19c9460f6%2Fmodeldesign.drawio-2.svg?generation=1692993585923781&amp;alt=media\" alt=\"\"></p>\n<h2>Augmentations</h2>\n<ul>\n<li>Random resample (0.5x ~ 1.5x to original length)</li>\n<li>Random Affine</li>\n<li>Random Cutout</li>\n<li>Random token replacement on decoder input(prob=0.2)</li>\n</ul>\n<p>What I overlooked this time is that augmentations which showed no performance improvement or even degraded performance in shorter epochs might actually help improve performance in longer epochs. I experienced a similar phenomenon in the previous competition, but it seemed more pronounced in this one. By selecting augmentations solely based on the results from a 60epoch experiment, I think I missed out on many potentially beneficial augmentations when testing the 400epoch training in the final week.</p>\n<h2>Training</h2>\n<p>Epoch = 400<br>\nbs = 16 * num_replicas = 128<br>\nLr = 5e-4 * num_replicas = 4e-3<br>\nAWP = 0.2 starts at 0.1 * Epoch<br>\nSchedule = CosineDecay with warmup ratio 0.1<br>\nOptimizer = AdamW (slightly better than RAdam with Lookahead)<br>\nLoss = CTC(weight=0.25) + CCE with label smoothing=0.1~0.25(weight=0.75)</p>\n<p>Training takes around 14 hours with colab TPUv2-8(as colab TPU runtime recently reduced to 3~4 hours, needed 4 consecutive sessions to complete training).<br>\nLonger Epoch always gave better CV(5fold split by id) and LB but got no time to try over 400 epochs.</p>\n<h2>LB history</h2>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>public LB</th>\n<th>private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>prevcompsinglemodel + CTC (or Attention)</td>\n<td>0.76</td>\n<td>0.74</td>\n</tr>\n<tr>\n<td>+ deeper and wider model, add pose</td>\n<td>0.79</td>\n<td>0.78</td>\n</tr>\n<tr>\n<td>+ 3 seed ensemble with Attention</td>\n<td>0.80</td>\n<td>0.79</td>\n</tr>\n<tr>\n<td>+ ctc attention joint decoding</td>\n<td>0.81</td>\n<td>0.80</td>\n</tr>\n<tr>\n<td>+ use all landmarks, longer epoch</td>\n<td>0.82</td>\n<td>0.81</td>\n</tr>\n</tbody>\n</table>\n<p>Seeing solutions from other kagglers, not only the top solutions but also the public notebooks is always inspiring. Thanks to kagglers who shared their insights and ideas as always. And big congrats to <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> and <a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a> for winning this highly competitive competition!</p>\n<p>Training/Inference code: <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/436873\" target=\"_blank\">https://www.kaggle.com/competitions/asl-fingerspelling/discussion/436873</a></p>",
  "messages": [
    {
      "id": 2408632,
      "postDate": "2023-08-25T18:24:52.443Z",
      "content": "<p>Once again, thanks to Kaggle, Google and other organizers for hosting this exciting competition. I felt this one as well as the previous competition, is well-organized so that we can try various ideas and could learn a lot from each other. Hope this kinds of competition open in kaggle more often.. :)</p>\n<h2>TLDR</h2>\n<p>My overall solution is almost identical to the previous competition. It's mainly just the adoption of ASR(Automatic Speech Recognition) algorithms (joint CTC + Attention) on top of the previous 1st place model. The details of the model, preprocessing, and training are largely similar to the previous competition, so please refer to the <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406684\" target=\"_blank\">previous competition solution</a> for more details on model/training/augmentations/etc.</p>\n<h2>ASR Algorithms</h2>\n<p>I had a deep interest in ASR, and after recognizing its correlation with the competition, I studied and experimented with various algorithms. While I couldn't find anything particularly superior to the widely used vanilla CTC among the numerous ASR algorithms, I learned a lot through various experiments, and I will mainly share those insights.</p>\n<p>As I began studying ASR for this competition, I discovered from reading papers that current NN ASR algorithm mainly consists of CTC, Attention-based, and Transducer. I implemented all three (though there are more diverse algorithms out there). In conclusion, from a baseline performance perspective, all three algorithms were quite similar and each have pros &amp; cons. Here are the insights I gathered from implementing each algorithm:</p>\n<ul>\n<li><p>CTC:</p>\n<ul>\n<li>During greedy decoding, the prediction step = O(1). Therefore, when using a Single model, GreedyCTC is the most efficient. (this is probably why many solutions use a large single model with CTC).</li>\n<li>With beam search, prediction step = O(L) (where L = encoder output length). Depending on the efficiency of the algorithm, the computational overhead isn't very large, so the overhead by introducing beam search isn't very significant. However, the performance improvement from using beam search is minimal (+0.003), making it not very tempting(as its tflite-convertible implementation is not quite trivial) compared to Attention.</li>\n<li>As discussed, there's no guarantee of alignment between model prediction timesteps, so simple average ensembling can't be applied.</li></ul></li>\n<li><p>Attention-based:</p>\n<ul>\n<li>Uses an autoregressive approach, so even with greedy decoding, prediction step = O(N) (where N = decoder output length).</li>\n<li>If the decoder is an RNN or Transformer, stateful inference(i.e. previous key, value caching with Transformer) can be used to reduce the complexity. In actual implementation with Transformer, it reduced the inference time on the CPU by about 20~30%.</li>\n<li>Beam search is possible, but unlike CTC, it requires introducing appropriate heuristics to penalize the output length. Although several attempts were made, none worked. </li>\n<li>It's easier to apply ensembling with Attention. Simply take the average of each model at each prediction step, and it works well (ensembling three models gave +0.009).</li>\n<li>More room for score improvement than CTC(ex more decoder layers with augmentations).</li></ul></li>\n<li><p>Transducer:<br>\nPrediction steps = O(L) (where L = encoder output length). Due to the most prediction steps in a greedy manner, it's not very efficient. I didn't consider it further as optimization was harder and performance was slightly lower compared to CTC or Attention. However, it has the advantage of real-time recognition in streaming mode if the encoder is causal (which wasn't relevant for this competition).</p></li>\n</ul>\n<p>Rough single model inference time on test dataset with kaggle kernel is as follows.</p>\n<blockquote>\n  <p>CTC greedy(~40mins) &lt; AttentionGreedy=CTCBeamsearch(~1h20mins) &lt;=  CTCAttentionJointGreedy(~1h30mins)</p>\n</blockquote>\n<h2>CTC-Attention Joint Training &amp; Decoding</h2>\n<p>As both CTC and Attention showed similar performance, I tried to find a method to utilize both techniques.  I primarily referred to the following two papers:</p>\n<p><em>Joint CTC-Attention Based End-To-End Speech Recognition Using Multi-Task Learning, Kim et al. 2017.</em><br>\n<a href=\"url\" target=\"_blank\">https://arxiv.org/pdf/1609.06773.pdf</a><br>\n<em>Joint CTC/attention decoding for end-to-end speech recognition, Hori et al. 2017.</em><br>\n<a href=\"url\" target=\"_blank\">https://aclanthology.org/P17-1048.pdf</a></p>\n<p>Joint CTC-Attention Training is, as the name implies, adding both a CTC decoder (single GRU layer) and an attention decoder (single Transformer decoder layer) to one encoder for multitask learning. The loss weight is set to CTC=0.25 and attention=0.75. But the Joint training itself did not bring noticeable performance boost.</p>\n<p>By using the CTC prefix score, it is possible to calculate the probability of an arbitrary output hypothesis \"h\" without depending on the output timestep of the CTC output. In other words, <strong>implementing the CTC prefix score computation allows for ensembling outputs not only between CTC models but also between CTC and Attention models</strong>. CTC-attention joint decoding with CTC weight=0.3 showed a performance improvement of +0.007~8 without much impact on inference time. It was especially challenging for me to implement it accurately, efficiently, and without any issues to be compatible with tflite.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5003978%2Fabde23b44b90f8ba6adc7d56ca874d86%2F.drawio-3.svg?generation=1692977029060458&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5003978%2F22c798e256f66322a0cc9e42b976d570%2F.drawio-4.svg?generation=1692977118893507&amp;alt=media\" alt=\"\"></p>\n<h2>Preprocessing</h2>\n<p>Similar to the previous competition solution, but simpler. Every landmark and xyz was used and standardized. Flip left-handed signer(rather than augmentation). MAX_LEN=768 was used. Any hand-crafted feature was not significant, likely due to the absence of complex relations between movements in frames.</p>\n<h2>Model</h2>\n<p>Encoder is the same as the previous competition (Stacked Conv1DBlock + TransformerBlock) but increased in size (expand ratio 2-&gt;4 in Conv1DBlock) and depth (8 layers -&gt; 17 layers). Single model has ~6.5M parameters. I applied padding='same'(rather than 'causal') and output stride=2 (requires slightly more logic for handling masking). In my case, mixing in Transformer blocks wasn't as effective as in the previous competition. Perhaps global features were less crucial in this comp. Additionally, I added one BN to input of the Conv1DBlock for more training stability(especially with awp + more epochs).</p>\n<p>CTC decoder used a single GRU layer followed by one FC layer. Attention Decoder used a single-layer Transformer decoder. Introducing augmentation to the Decoder input and adding up to 4 Decoder layers improves the performance of the attention decoder (up to +0.004). However, considering the number of parameters and inference speed, I considered it inefficient and thus used a single-layer decoder.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5003978%2F7d1f7aee69242af5e6aa5ef19c9460f6%2Fmodeldesign.drawio-2.svg?generation=1692993585923781&amp;alt=media\" alt=\"\"></p>\n<h2>Augmentations</h2>\n<ul>\n<li>Random resample (0.5x ~ 1.5x to original length)</li>\n<li>Random Affine</li>\n<li>Random Cutout</li>\n<li>Random token replacement on decoder input(prob=0.2)</li>\n</ul>\n<p>What I overlooked this time is that augmentations which showed no performance improvement or even degraded performance in shorter epochs might actually help improve performance in longer epochs. I experienced a similar phenomenon in the previous competition, but it seemed more pronounced in this one. By selecting augmentations solely based on the results from a 60epoch experiment, I think I missed out on many potentially beneficial augmentations when testing the 400epoch training in the final week.</p>\n<h2>Training</h2>\n<p>Epoch = 400<br>\nbs = 16 * num_replicas = 128<br>\nLr = 5e-4 * num_replicas = 4e-3<br>\nAWP = 0.2 starts at 0.1 * Epoch<br>\nSchedule = CosineDecay with warmup ratio 0.1<br>\nOptimizer = AdamW (slightly better than RAdam with Lookahead)<br>\nLoss = CTC(weight=0.25) + CCE with label smoothing=0.1~0.25(weight=0.75)</p>\n<p>Training takes around 14 hours with colab TPUv2-8(as colab TPU runtime recently reduced to 3~4 hours, needed 4 consecutive sessions to complete training).<br>\nLonger Epoch always gave better CV(5fold split by id) and LB but got no time to try over 400 epochs.</p>\n<h2>LB history</h2>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>public LB</th>\n<th>private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>prevcompsinglemodel + CTC (or Attention)</td>\n<td>0.76</td>\n<td>0.74</td>\n</tr>\n<tr>\n<td>+ deeper and wider model, add pose</td>\n<td>0.79</td>\n<td>0.78</td>\n</tr>\n<tr>\n<td>+ 3 seed ensemble with Attention</td>\n<td>0.80</td>\n<td>0.79</td>\n</tr>\n<tr>\n<td>+ ctc attention joint decoding</td>\n<td>0.81</td>\n<td>0.80</td>\n</tr>\n<tr>\n<td>+ use all landmarks, longer epoch</td>\n<td>0.82</td>\n<td>0.81</td>\n</tr>\n</tbody>\n</table>\n<p>Seeing solutions from other kagglers, not only the top solutions but also the public notebooks is always inspiring. Thanks to kagglers who shared their insights and ideas as always. And big congrats to <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> and <a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a> for winning this highly competitive competition!</p>\n<p>Training/Inference code: <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/436873\" target=\"_blank\">https://www.kaggle.com/competitions/asl-fingerspelling/discussion/436873</a></p>",
      "rawMarkdown": "Once again, thanks to Kaggle, Google and other organizers for hosting this exciting competition. I felt this one as well as the previous competition, is well-organized so that we can try various ideas and could learn a lot from each other. Hope this kinds of competition open in kaggle more often.. :)\n\n##TLDR\nMy overall solution is almost identical to the previous competition. It's mainly just the adoption of ASR(Automatic Speech Recognition) algorithms (joint CTC + Attention) on top of the previous 1st place model. The details of the model, preprocessing, and training are largely similar to the previous competition, so please refer to the [previous competition solution](https://www.kaggle.com/competitions/asl-signs/discussion/406684) for more details on model/training/augmentations/etc.\n\n## ASR Algorithms\nI had a deep interest in ASR, and after recognizing its correlation with the competition, I studied and experimented with various algorithms. While I couldn't find anything particularly superior to the widely used vanilla CTC among the numerous ASR algorithms, I learned a lot through various experiments, and I will mainly share those insights.\n\nAs I began studying ASR for this competition, I discovered from reading papers that current NN ASR algorithm mainly consists of CTC, Attention-based, and Transducer. I implemented all three (though there are more diverse algorithms out there). In conclusion, from a baseline performance perspective, all three algorithms were quite similar and each have pros & cons. Here are the insights I gathered from implementing each algorithm:\n\n- CTC:\n    - During greedy decoding, the prediction step = O(1). Therefore, when using a Single model, GreedyCTC is the most efficient. (this is probably why many solutions use a large single model with CTC).\n    - With beam search, prediction step = O(L) (where L = encoder output length). Depending on the efficiency of the algorithm, the computational overhead isn't very large, so the overhead by introducing beam search isn't very significant. However, the performance improvement from using beam search is minimal (+0.003), making it not very tempting(as its tflite-convertible implementation is not quite trivial) compared to Attention.\n    - As discussed, there's no guarantee of alignment between model prediction timesteps, so simple average ensembling can't be applied.\n\n- Attention-based:\n    - Uses an autoregressive approach, so even with greedy decoding, prediction step = O(N) (where N = decoder output length).\n    -  If the decoder is an RNN or Transformer, stateful inference(i.e. previous key, value caching with Transformer) can be used to reduce the complexity. In actual implementation with Transformer, it reduced the inference time on the CPU by about 20~30%.\n    - Beam search is possible, but unlike CTC, it requires introducing appropriate heuristics to penalize the output length. Although several attempts were made, none worked. \n    - It's easier to apply ensembling with Attention. Simply take the average of each model at each prediction step, and it works well (ensembling three models gave +0.009).\n    - More room for score improvement than CTC(ex more decoder layers with augmentations).\n\n- Transducer:\nPrediction steps = O(L) (where L = encoder output length). Due to the most prediction steps in a greedy manner, it's not very efficient. I didn't consider it further as optimization was harder and performance was slightly lower compared to CTC or Attention. However, it has the advantage of real-time recognition in streaming mode if the encoder is causal (which wasn't relevant for this competition).\n\nRough single model inference time on test dataset with kaggle kernel is as follows.\n\n> CTC greedy(~40mins) < AttentionGreedy=CTCBeamsearch(~1h20mins) <=  CTCAttentionJointGreedy(~1h30mins)\n\n##CTC-Attention Joint Training & Decoding\nAs both CTC and Attention showed similar performance, I tried to find a method to utilize both techniques.  I primarily referred to the following two papers:\n\n*Joint CTC-Attention Based End-To-End Speech Recognition Using Multi-Task Learning, Kim et al. 2017.*\n[https://arxiv.org/pdf/1609.06773.pdf](url)\n*Joint CTC/attention decoding for end-to-end speech recognition, Hori et al. 2017.*\n[https://aclanthology.org/P17-1048.pdf](url)\n\nJoint CTC-Attention Training is, as the name implies, adding both a CTC decoder (single GRU layer) and an attention decoder (single Transformer decoder layer) to one encoder for multitask learning. The loss weight is set to CTC=0.25 and attention=0.75. But the Joint training itself did not bring noticeable performance boost.\n\nBy using the CTC prefix score, it is possible to calculate the probability of an arbitrary output hypothesis \"h\" without depending on the output timestep of the CTC output. In other words, **implementing the CTC prefix score computation allows for ensembling outputs not only between CTC models but also between CTC and Attention models**. CTC-attention joint decoding with CTC weight=0.3 showed a performance improvement of +0.007~8 without much impact on inference time. It was especially challenging for me to implement it accurately, efficiently, and without any issues to be compatible with tflite.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5003978%2Fabde23b44b90f8ba6adc7d56ca874d86%2F.drawio-3.svg?generation=1692977029060458&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5003978%2F22c798e256f66322a0cc9e42b976d570%2F.drawio-4.svg?generation=1692977118893507&alt=media)\n\n##Preprocessing\nSimilar to the previous competition solution, but simpler. Every landmark and xyz was used and standardized. Flip left-handed signer(rather than augmentation). MAX_LEN=768 was used. Any hand-crafted feature was not significant, likely due to the absence of complex relations between movements in frames.\n\n##Model\nEncoder is the same as the previous competition (Stacked Conv1DBlock + TransformerBlock) but increased in size (expand ratio 2->4 in Conv1DBlock) and depth (8 layers -> 17 layers). Single model has ~6.5M parameters. I applied padding='same'(rather than 'causal') and output stride=2 (requires slightly more logic for handling masking). In my case, mixing in Transformer blocks wasn't as effective as in the previous competition. Perhaps global features were less crucial in this comp. Additionally, I added one BN to input of the Conv1DBlock for more training stability(especially with awp + more epochs).\n\nCTC decoder used a single GRU layer followed by one FC layer. Attention Decoder used a single-layer Transformer decoder. Introducing augmentation to the Decoder input and adding up to 4 Decoder layers improves the performance of the attention decoder (up to +0.004). However, considering the number of parameters and inference speed, I considered it inefficient and thus used a single-layer decoder.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5003978%2F7d1f7aee69242af5e6aa5ef19c9460f6%2Fmodeldesign.drawio-2.svg?generation=1692993585923781&alt=media)\n##Augmentations\n* Random resample (0.5x ~ 1.5x to original length)\n* Random Affine\n* Random Cutout\n* Random token replacement on decoder input(prob=0.2)\n\nWhat I overlooked this time is that augmentations which showed no performance improvement or even degraded performance in shorter epochs might actually help improve performance in longer epochs. I experienced a similar phenomenon in the previous competition, but it seemed more pronounced in this one. By selecting augmentations solely based on the results from a 60epoch experiment, I think I missed out on many potentially beneficial augmentations when testing the 400epoch training in the final week.\n\n\n##Training\nEpoch = 400\nbs = 16 * num_replicas = 128\nLr = 5e-4 * num_replicas = 4e-3\nAWP = 0.2 starts at 0.1 * Epoch\nSchedule = CosineDecay with warmup ratio 0.1\nOptimizer = AdamW (slightly better than RAdam with Lookahead)\nLoss = CTC(weight=0.25) + CCE with label smoothing=0.1~0.25(weight=0.75)\n\nTraining takes around 14 hours with colab TPUv2-8(as colab TPU runtime recently reduced to 3~4 hours, needed 4 consecutive sessions to complete training).\nLonger Epoch always gave better CV(5fold split by id) and LB but got no time to try over 400 epochs.\n\n##LB history\n|  | public LB | private LB | \n| --- | --- | --- |\n| prevcompsinglemodel + CTC (or Attention) | 0.76 | 0.74 |\n| + deeper and wider model, add pose | 0.79 | 0.78 |\n| + 3 seed ensemble with Attention | 0.80 | 0.79 |\n| + ctc attention joint decoding | 0.81 | 0.80 | \n| + use all landmarks, longer epoch | 0.82 | 0.81 |\n\n\nSeeing solutions from other kagglers, not only the top solutions but also the public notebooks is always inspiring. Thanks to kagglers who shared their insights and ideas as always. And big congrats to @christofhenkel and @darraghdog for winning this highly competitive competition!\n\nTraining/Inference code: https://www.kaggle.com/competitions/asl-fingerspelling/discussion/436873\n\n\n",
      "votes": 76
    },
    {
      "id": 2429971,
      "postDate": "2023-09-09T00:29:02.437Z",
      "content": "<p>Thank you very much for a detailed explanation! I also learned a lot from your previous solution.</p>",
      "rawMarkdown": "Thank you very much for a detailed explanation! I also learned a lot from your previous solution.",
      "votes": 1
    },
    {
      "id": 2426159,
      "postDate": "2023-09-06T12:48:31.130Z",
      "content": "<p>Thanks for sharing. And congrats brother. </p>",
      "rawMarkdown": "Thanks for sharing. And congrats brother. ",
      "votes": 1
    },
    {
      "id": 2417749,
      "postDate": "2023-08-31T20:07:45.157Z",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/hoyso48\" target=\"_blank\">@hoyso48</a> - Congratulations on your 2nd place finish!  Will the final tflite model or winning submission.zip file be available?  I'd like to use it with an Android app idea I'm developing.  Thanks, Matthew</p>",
      "rawMarkdown": "Hello @hoyso48 - Congratulations on your 2nd place finish!  Will the final tflite model or winning submission.zip file be available?  I'd like to use it with an Android app idea I'm developing.  Thanks, Matthew\n",
      "votes": 1
    },
    {
      "id": 2413455,
      "postDate": "2023-08-28T23:46:30.657Z",
      "content": "<p>I've been stuck with TPU training for over half a month with no progress. Looking forward to seeing your code.</p>",
      "rawMarkdown": "I've been stuck with TPU training for over half a month with no progress. Looking forward to seeing your code.",
      "votes": 1
    },
    {
      "id": 2410738,
      "postDate": "2023-08-27T07:10:02.440Z",
      "content": "<p><a href=\"https://www.kaggle.com/hoyso48\" target=\"_blank\">@hoyso48</a> Congrats.<br>\nYour solutions and implementation using tf/keras are genuinely inspiring and commendable. Thank you for sharing. :)</p>",
      "rawMarkdown": "@hoyso48 Congrats.\nYour solutions and implementation using tf/keras are genuinely inspiring and commendable. Thank you for sharing. :)",
      "votes": 1
    },
    {
      "id": 2410642,
      "postDate": "2023-08-27T06:17:42.860Z",
      "content": "<p>Congrats! Thanks for sharing your solution! <br>\nI want to ask the paper you referred to in the previous ASR competition.</p>",
      "rawMarkdown": "Congrats! Thanks for sharing your solution! \nI want to ask the paper you referred to in the previous ASR competition.",
      "votes": 1,
      "replies": [
        {
          "id": 2410874,
          "postDate": "2023-08-27T08:58:32.790Z",
          "content": "<p>I did not refer to any particular paper in the previous competition. Of course, I did search through some papers on related tasks (such as action recognition and GCN models) at the beginning, but I realized that the methods from those papers were not very effective.</p>",
          "rawMarkdown": "I did not refer to any particular paper in the previous competition. Of course, I did search through some papers on related tasks (such as action recognition and GCN models) at the beginning, but I realized that the methods from those papers were not very effective."
        }
      ]
    },
    {
      "id": 2409604,
      "postDate": "2023-08-26T10:50:54.273Z",
      "content": "<p><a href=\"https://www.kaggle.com/hoyso48\" target=\"_blank\">@hoyso48</a> Thanks for the write up; this is great as always. The way you approach the problem is inspiring. </p>",
      "rawMarkdown": "@hoyso48 Thanks for the write up; this is great as always. The way you approach the problem is inspiring. ",
      "votes": 1,
      "replies": [
        {
          "id": 2410148,
          "postDate": "2023-08-26T17:12:43.810Z",
          "content": "<p>Thanks and congratulations! Your work through these two competitions was very masterful. need more time to digest, but I've already learned a lot from your solution :)</p>",
          "rawMarkdown": "Thanks and congratulations! Your work through these two competitions was very masterful. need more time to digest, but I've already learned a lot from your solution :)"
        }
      ]
    },
    {
      "id": 2409131,
      "postDate": "2023-08-26T04:44:51.867Z",
      "content": "<p>Congrats! </p>\n<p>I have a question. Did you upload the full dataset to Google Cloud? </p>",
      "rawMarkdown": "Congrats! \n\nI have a question. Did you upload the full dataset to Google Cloud? ",
      "votes": 1,
      "replies": [
        {
          "id": 2409394,
          "postDate": "2023-08-26T08:11:20.050Z",
          "content": "<p>I used KaggleDatasets(same as previous competition, converted .parquet to .tfrecords). every kaggle dataset has its own Google Cloud Storage path, so if it is public and you have the actual gcs path then you can access it anywhere with GCS.</p>",
          "rawMarkdown": "I used KaggleDatasets(same as previous competition, converted .parquet to .tfrecords). every kaggle dataset has its own Google Cloud Storage path, so if it is public and you have the actual gcs path then you can access it anywhere with GCS.",
          "votes": 1,
          "replies": [
            {
              "id": 2409397,
              "postDate": "2023-08-26T08:15:28.003Z",
              "content": "<p>Thanks for your reply! OMG, it is so simple. Thanks again I guess it's gonna help me a lot next time.</p>",
              "rawMarkdown": "Thanks for your reply! OMG, it is so simple. Thanks again I guess it's gonna help me a lot next time."
            },
            {
              "id": 2410492,
              "postDate": "2023-08-27T02:44:26.177Z",
              "content": "<p>When I convert the data format to .tfrecords, the runtime exceeds 12 hours. How did you solve this issue?Thanks!</p>",
              "rawMarkdown": "When I convert the data format to .tfrecords, the runtime exceeds 12 hours. How did you solve this issue?Thanks!"
            }
          ]
        }
      ]
    },
    {
      "id": 2409083,
      "postDate": "2023-08-26T04:02:38.487Z",
      "content": "<p>Congratulations. Thanks for sharing the details of the notebook. <br>\nI think there is a benefit of separating training and inference notebook. </p>\n<h2>- While submitting the notebook, I observe that notebook running time should be less than 9 hours.  </h2>",
      "rawMarkdown": "Congratulations. Thanks for sharing the details of the notebook. \nI think there is a benefit of separating training and inference notebook. \n- While submitting the notebook, I observe that notebook running time should be less than 9 hours.  \n- \n",
      "votes": 1
    },
    {
      "id": 2408894,
      "postDate": "2023-08-25T21:59:09.907Z",
      "content": "<p>Congratulations! I learned so much from what you share in these two competions, I could not start this competion without your previous sharing, especially your prvious code of showing stachastic path(tf.dropout with noise shape). Thanks!<br>\nI will put some time on learning your sharing, ctc+attention method is so cool. <br>\nAlso I think you might have a big boost by simply replace your encoder to squeezeformer or conformer.</p>",
      "rawMarkdown": "Congratulations! I learned so much from what you share in these two competions, I could not start this competion without your previous sharing, especially your prvious code of showing stachastic path(tf.dropout with noise shape). Thanks!\nI will put some time on learning your sharing, ctc+attention method is so cool. \nAlso I think you might have a big boost by simply replace your encoder to squeezeformer or conformer.",
      "votes": 1,
      "replies": [
        {
          "id": 2409027,
          "postDate": "2023-08-26T02:55:37.323Z",
          "content": "<p>Glad to know that my prev solution gave you some help :)<br>\nI relied on my prev model and gave minimal changes simply because it's my own implementation from prev competition with lots of trials&amp;errors..haha<br>\nDefinitely should study&amp;try more recent conv + transformer hybrid architectures next time. Thanks!</p>",
          "rawMarkdown": "Glad to know that my prev solution gave you some help :)\nI relied on my prev model and gave minimal changes simply because it's my own implementation from prev competition with lots of trials&errors..haha\nDefinitely should study&try more recent conv + transformer hybrid architectures next time. Thanks!",
          "votes": 1
        }
      ]
    },
    {
      "id": 2412509,
      "postDate": "2023-08-28T10:12:12.083Z",
      "content": "<p>Congrats on your great finish, <a href=\"https://www.kaggle.com/hoyso48\" target=\"_blank\">@hoyso48</a>. Your solution of the last ASL competition helped  me get started on this one, thanks!</p>\n<p>I thought about the joint CTC and attention decoding quite a bit but didn't find any great way to do it. Your solution is clean!</p>\n<blockquote>\n  <p>What I overlooked this time is that augmentations which showed no performance improvement or even degraded performance in shorter epochs might actually help improve performance in longer epochs.</p>\n</blockquote>\n<p>I experienced this too. Looking at the winners' huge set of augmentations there was a lot to be gained, it seems.</p>",
      "rawMarkdown": "Congrats on your great finish, @hoyso48. Your solution of the last ASL competition helped  me get started on this one, thanks!\n\nI thought about the joint CTC and attention decoding quite a bit but didn't find any great way to do it. Your solution is clean!\n\n>What I overlooked this time is that augmentations which showed no performance improvement or even degraded performance in shorter epochs might actually help improve performance in longer epochs.\n\nI experienced this too. Looking at the winners' huge set of augmentations there was a lot to be gained, it seems.",
      "votes": 2
    },
    {
      "id": 2408876,
      "postDate": "2023-08-25T21:38:34.223Z",
      "content": "<p>Thank you for sharing. I enjoyed reading how you made the combination of CTC and attention based decoding work. We tried and did not succeed there. </p>",
      "rawMarkdown": "Thank you for sharing. I enjoyed reading how you made the combination of CTC and attention based decoding work. We tried and did not succeed there. ",
      "votes": 2,
      "replies": [
        {
          "id": 2409025,
          "postDate": "2023-08-26T02:44:44.647Z",
          "content": "<p>Thanks, and congrats on winning! I've always learned from top Kagglers like you, and now I'm happy that I could also provide you with some insights :)</p>",
          "rawMarkdown": "Thanks, and congrats on winning! I've always learned from top Kagglers like you, and now I'm happy that I could also provide you with some insights :)",
          "votes": 2
        }
      ]
    },
    {
      "id": 2430983,
      "postDate": "2023-09-09T17:40:50.947Z",
      "content": "<p>Thank for your contribution, this is amazing!</p>",
      "rawMarkdown": "Thank for your contribution, this is amazing!"
    },
    {
      "id": 2411963,
      "postDate": "2023-08-28T02:10:31.523Z",
      "content": "<p>Congratulations! Thanks for your write up!</p>",
      "rawMarkdown": "Congratulations! Thanks for your write up!"
    },
    {
      "id": 2411376,
      "postDate": "2023-08-27T15:10:06.763Z",
      "content": "<p>Wow this is a great write up. Congrats!</p>",
      "rawMarkdown": "Wow this is a great write up. Congrats!"
    },
    {
      "id": 2408699,
      "postDate": "2023-08-25T19:02:12.247Z",
      "content": "<p>Thanks for sharing your solution! Did you use google colab for everything that you've shared here?</p>",
      "rawMarkdown": "Thanks for sharing your solution! Did you use google colab for everything that you've shared here?",
      "replies": [
        {
          "id": 2408712,
          "postDate": "2023-08-25T19:05:13.620Z",
          "content": "<p>Yes. I only used Colab TPUv2-8.</p>",
          "rawMarkdown": "Yes. I only used Colab TPUv2-8.",
          "votes": 2,
          "replies": [
            {
              "id": 2408749,
              "postDate": "2023-08-25T19:18:10.613Z",
              "content": "<p>Great, thanks for the info and congrats!</p>",
              "rawMarkdown": "Great, thanks for the info and congrats!",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2426082,
      "postDate": "2023-09-06T11:52:52.667Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 2412843,
      "postDate": "2023-08-28T14:23:22.357Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2416058,
      "postDate": "2023-08-30T18:57:50.227Z",
      "content": "<p>Thanks for the share <a href=\"https://www.kaggle.com/hoyso48\" target=\"_blank\">@hoyso48</a> </p>",
      "rawMarkdown": "Thanks for the share @hoyso48 ",
      "votes": 1
    },
    {
      "id": 2413033,
      "postDate": "2023-08-28T16:33:22.727Z",
      "content": "<p>Thank You for the summary sir!</p>",
      "rawMarkdown": "Thank You for the summary sir!",
      "votes": 1
    },
    {
      "id": 2412793,
      "postDate": "2023-08-28T14:03:45.520Z",
      "content": "<p>thanks for the giving a nice solutions.</p>",
      "rawMarkdown": "thanks for the giving a nice solutions.",
      "votes": 1
    },
    {
      "id": 2411085,
      "postDate": "2023-08-27T12:09:01.403Z",
      "content": "<p>thanks for sharing your works.</p>",
      "rawMarkdown": "thanks for sharing your works.",
      "votes": 1
    },
    {
      "id": 2410532,
      "postDate": "2023-08-27T04:05:39.380Z",
      "content": "<p>Thanks for this guide.</p>",
      "rawMarkdown": "Thanks for this guide.",
      "votes": 1
    },
    {
      "id": 2410236,
      "postDate": "2023-08-26T18:24:32.697Z",
      "content": "<p>Thanks for sharing! congratulations✨</p>",
      "rawMarkdown": "Thanks for sharing! congratulations✨",
      "votes": 1
    },
    {
      "id": 2409078,
      "postDate": "2023-08-26T04:01:15.050Z",
      "content": "<p>Thank You Very Much !</p>",
      "rawMarkdown": "Thank You Very Much !",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 2429971,
      "author_name": "Tong Zou",
      "author_url": "",
      "post_date": "2023-09-09T00:29:02.437000",
      "content": "<p>Thank you very much for a detailed explanation! I also learned a lot from your previous solution.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2426159,
      "author_name": "Indra Sonowal",
      "author_url": "",
      "post_date": "2023-09-06T12:48:31.130000",
      "content": "<p>Thanks for sharing. And congrats brother. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2417749,
      "author_name": "matucker",
      "author_url": "",
      "post_date": "2023-08-31T20:07:45.157000",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/hoyso48\" target=\"_blank\">@hoyso48</a> - Congratulations on your 2nd place finish!  Will the final tflite model or winning submission.zip file be available?  I'd like to use it with an Android app idea I'm developing.  Thanks, Matthew</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2413455,
      "author_name": "Scenery SunFireInk",
      "author_url": "",
      "post_date": "2023-08-28T23:46:30.657000",
      "content": "<p>I've been stuck with TPU training for over half a month with no progress. Looking forward to seeing your code.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2410738,
      "author_name": "Innat",
      "author_url": "",
      "post_date": "2023-08-27T07:10:02.440000",
      "content": "<p><a href=\"https://www.kaggle.com/hoyso48\" target=\"_blank\">@hoyso48</a> Congrats.<br>\nYour solutions and implementation using tf/keras are genuinely inspiring and commendable. Thank you for sharing. :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2410642,
      "author_name": "mingyue81",
      "author_url": "",
      "post_date": "2023-08-27T06:17:42.860000",
      "content": "<p>Congrats! Thanks for sharing your solution! <br>\nI want to ask the paper you referred to in the previous ASR competition.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2410874,
          "author_name": "hoyso48",
          "author_url": "",
          "post_date": "2023-08-27T08:58:32.790000",
          "content": "<p>I did not refer to any particular paper in the previous competition. Of course, I did search through some papers on related tasks (such as action recognition and GCN models) at the beginning, but I realized that the methods from those papers were not very effective.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2409604,
      "author_name": "Darragh",
      "author_url": "",
      "post_date": "2023-08-26T10:50:54.273000",
      "content": "<p><a href=\"https://www.kaggle.com/hoyso48\" target=\"_blank\">@hoyso48</a> Thanks for the write up; this is great as always. The way you approach the problem is inspiring. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2410148,
          "author_name": "hoyso48",
          "author_url": "",
          "post_date": "2023-08-26T17:12:43.810000",
          "content": "<p>Thanks and congratulations! Your work through these two competitions was very masterful. need more time to digest, but I've already learned a lot from your solution :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2409131,
      "author_name": "Yu Wu",
      "author_url": "",
      "post_date": "2023-08-26T04:44:51.867000",
      "content": "<p>Congrats! </p>\n<p>I have a question. Did you upload the full dataset to Google Cloud? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2409394,
          "author_name": "hoyso48",
          "author_url": "",
          "post_date": "2023-08-26T08:11:20.050000",
          "content": "<p>I used KaggleDatasets(same as previous competition, converted .parquet to .tfrecords). every kaggle dataset has its own Google Cloud Storage path, so if it is public and you have the actual gcs path then you can access it anywhere with GCS.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2409397,
              "author_name": "Yu Wu",
              "author_url": "",
              "post_date": "2023-08-26T08:15:28.003000",
              "content": "<p>Thanks for your reply! OMG, it is so simple. Thanks again I guess it's gonna help me a lot next time.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2410492,
              "author_name": "Nowgger",
              "author_url": "",
              "post_date": "2023-08-27T02:44:26.177000",
              "content": "<p>When I convert the data format to .tfrecords, the runtime exceeds 12 hours. How did you solve this issue?Thanks!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2409083,
      "author_name": "C R Suthikshn Kumar",
      "author_url": "",
      "post_date": "2023-08-26T04:02:38.487000",
      "content": "<p>Congratulations. Thanks for sharing the details of the notebook. <br>\nI think there is a benefit of separating training and inference notebook. </p>\n<h2>- While submitting the notebook, I observe that notebook running time should be less than 9 hours.  </h2>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2408894,
      "author_name": "gezi",
      "author_url": "",
      "post_date": "2023-08-25T21:59:09.907000",
      "content": "<p>Congratulations! I learned so much from what you share in these two competions, I could not start this competion without your previous sharing, especially your prvious code of showing stachastic path(tf.dropout with noise shape). Thanks!<br>\nI will put some time on learning your sharing, ctc+attention method is so cool. <br>\nAlso I think you might have a big boost by simply replace your encoder to squeezeformer or conformer.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2409027,
          "author_name": "hoyso48",
          "author_url": "",
          "post_date": "2023-08-26T02:55:37.323000",
          "content": "<p>Glad to know that my prev solution gave you some help :)<br>\nI relied on my prev model and gave minimal changes simply because it's my own implementation from prev competition with lots of trials&amp;errors..haha<br>\nDefinitely should study&amp;try more recent conv + transformer hybrid architectures next time. Thanks!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2412509,
      "author_name": "flg",
      "author_url": "",
      "post_date": "2023-08-28T10:12:12.083000",
      "content": "<p>Congrats on your great finish, <a href=\"https://www.kaggle.com/hoyso48\" target=\"_blank\">@hoyso48</a>. Your solution of the last ASL competition helped  me get started on this one, thanks!</p>\n<p>I thought about the joint CTC and attention decoding quite a bit but didn't find any great way to do it. Your solution is clean!</p>\n<blockquote>\n  <p>What I overlooked this time is that augmentations which showed no performance improvement or even degraded performance in shorter epochs might actually help improve performance in longer epochs.</p>\n</blockquote>\n<p>I experienced this too. Looking at the winners' huge set of augmentations there was a lot to be gained, it seems.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2408876,
      "author_name": "Dieter",
      "author_url": "",
      "post_date": "2023-08-25T21:38:34.223000",
      "content": "<p>Thank you for sharing. I enjoyed reading how you made the combination of CTC and attention based decoding work. We tried and did not succeed there. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 2409025,
          "author_name": "hoyso48",
          "author_url": "",
          "post_date": "2023-08-26T02:44:44.647000",
          "content": "<p>Thanks, and congrats on winning! I've always learned from top Kagglers like you, and now I'm happy that I could also provide you with some insights :)</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2430983,
      "author_name": "Artem Ponomarenko",
      "author_url": "",
      "post_date": "2023-09-09T17:40:50.947000",
      "content": "<p>Thank for your contribution, this is amazing!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2411963,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-08-28T02:10:31.523000",
      "content": "<p>Congratulations! Thanks for your write up!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2411376,
      "author_name": "Samuel Cortinhas",
      "author_url": "",
      "post_date": "2023-08-27T15:10:06.763000",
      "content": "<p>Wow this is a great write up. Congrats!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2408699,
      "author_name": "kansk42",
      "author_url": "",
      "post_date": "2023-08-25T19:02:12.247000",
      "content": "<p>Thanks for sharing your solution! Did you use google colab for everything that you've shared here?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2408712,
          "author_name": "hoyso48",
          "author_url": "",
          "post_date": "2023-08-25T19:05:13.620000",
          "content": "<p>Yes. I only used Colab TPUv2-8.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2408749,
              "author_name": "kansk42",
              "author_url": "",
              "post_date": "2023-08-25T19:18:10.613000",
              "content": "<p>Great, thanks for the info and congrats!</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2426082,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-09-06T11:52:52.667000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2412843,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-08-28T14:23:22.357000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2416058,
      "author_name": "R.CHIRANJEEVI SRINIVAS",
      "author_url": "",
      "post_date": "2023-08-30T18:57:50.227000",
      "content": "<p>Thanks for the share <a href=\"https://www.kaggle.com/hoyso48\" target=\"_blank\">@hoyso48</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2413033,
      "author_name": "Mystic Shadow",
      "author_url": "",
      "post_date": "2023-08-28T16:33:22.727000",
      "content": "<p>Thank You for the summary sir!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2412793,
      "author_name": "Al Sani",
      "author_url": "",
      "post_date": "2023-08-28T14:03:45.520000",
      "content": "<p>thanks for the giving a nice solutions.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2411085,
      "author_name": "dragon zhang",
      "author_url": "",
      "post_date": "2023-08-27T12:09:01.403000",
      "content": "<p>thanks for sharing your works.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2410532,
      "author_name": "Muhammad Usman",
      "author_url": "",
      "post_date": "2023-08-27T04:05:39.380000",
      "content": "<p>Thanks for this guide.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2410236,
      "author_name": "Shivam Kumar",
      "author_url": "",
      "post_date": "2023-08-26T18:24:32.697000",
      "content": "<p>Thanks for sharing! congratulations✨</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2409078,
      "author_name": "Mystic Shadow",
      "author_url": "",
      "post_date": "2023-08-26T04:01:15.050000",
      "content": "<p>Thank You Very Much !</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2408632": "Once again, thanks to Kaggle, Google and other organizers for hosting this exciting competition. I felt this one as well as the previous competition, is well-organized so that we can try various ideas and could learn a lot from each other. Hope this kinds of competition open in kaggle more often.. :)\n\n##TLDR\nMy overall solution is almost identical to the previous competition. It's mainly just the adoption of ASR(Automatic Speech Recognition) algorithms (joint CTC + Attention) on top of the previous 1st place model. The details of the model, preprocessing, and training are largely similar to the previous competition, so please refer to the [previous competition solution](https://www.kaggle.com/competitions/asl-signs/discussion/406684) for more details on model/training/augmentations/etc.\n\n## ASR Algorithms\nI had a deep interest in ASR, and after recognizing its correlation with the competition, I studied and experimented with various algorithms. While I couldn't find anything particularly superior to the widely used vanilla CTC among the numerous ASR algorithms, I learned a lot through various experiments, and I will mainly share those insights.\n\nAs I began studying ASR for this competition, I discovered from reading papers that current NN ASR algorithm mainly consists of CTC, Attention-based, and Transducer. I implemented all three (though there are more diverse algorithms out there). In conclusion, from a baseline performance perspective, all three algorithms were quite similar and each have pros & cons. Here are the insights I gathered from implementing each algorithm:\n\n- CTC:\n    - During greedy decoding, the prediction step = O(1). Therefore, when using a Single model, GreedyCTC is the most efficient. (this is probably why many solutions use a large single model with CTC).\n    - With beam search, prediction step = O(L) (where L = encoder output length). Depending on the efficiency of the algorithm, the computational overhead isn't very large, so the overhead by introducing beam search isn't very significant. However, the performance improvement from using beam search is minimal (+0.003), making it not very tempting(as its tflite-convertible implementation is not quite trivial) compared to Attention.\n    - As discussed, there's no guarantee of alignment between model prediction timesteps, so simple average ensembling can't be applied.\n\n- Attention-based:\n    - Uses an autoregressive approach, so even with greedy decoding, prediction step = O(N) (where N = decoder output length).\n    -  If the decoder is an RNN or Transformer, stateful inference(i.e. previous key, value caching with Transformer) can be used to reduce the complexity. In actual implementation with Transformer, it reduced the inference time on the CPU by about 20~30%.\n    - Beam search is possible, but unlike CTC, it requires introducing appropriate heuristics to penalize the output length. Although several attempts were made, none worked. \n    - It's easier to apply ensembling with Attention. Simply take the average of each model at each prediction step, and it works well (ensembling three models gave +0.009).\n    - More room for score improvement than CTC(ex more decoder layers with augmentations).\n\n- Transducer:\nPrediction steps = O(L) (where L = encoder output length). Due to the most prediction steps in a greedy manner, it's not very efficient. I didn't consider it further as optimization was harder and performance was slightly lower compared to CTC or Attention. However, it has the advantage of real-time recognition in streaming mode if the encoder is causal (which wasn't relevant for this competition).\n\nRough single model inference time on test dataset with kaggle kernel is as follows.\n\n> CTC greedy(~40mins) < AttentionGreedy=CTCBeamsearch(~1h20mins) <=  CTCAttentionJointGreedy(~1h30mins)\n\n##CTC-Attention Joint Training & Decoding\nAs both CTC and Attention showed similar performance, I tried to find a method to utilize both techniques.  I primarily referred to the following two papers:\n\n*Joint CTC-Attention Based End-To-End Speech Recognition Using Multi-Task Learning, Kim et al. 2017.*\n[https://arxiv.org/pdf/1609.06773.pdf](url)\n*Joint CTC/attention decoding for end-to-end speech recognition, Hori et al. 2017.*\n[https://aclanthology.org/P17-1048.pdf](url)\n\nJoint CTC-Attention Training is, as the name implies, adding both a CTC decoder (single GRU layer) and an attention decoder (single Transformer decoder layer) to one encoder for multitask learning. The loss weight is set to CTC=0.25 and attention=0.75. But the Joint training itself did not bring noticeable performance boost.\n\nBy using the CTC prefix score, it is possible to calculate the probability of an arbitrary output hypothesis \"h\" without depending on the output timestep of the CTC output. In other words, **implementing the CTC prefix score computation allows for ensembling outputs not only between CTC models but also between CTC and Attention models**. CTC-attention joint decoding with CTC weight=0.3 showed a performance improvement of +0.007~8 without much impact on inference time. It was especially challenging for me to implement it accurately, efficiently, and without any issues to be compatible with tflite.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5003978%2Fabde23b44b90f8ba6adc7d56ca874d86%2F.drawio-3.svg?generation=1692977029060458&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5003978%2F22c798e256f66322a0cc9e42b976d570%2F.drawio-4.svg?generation=1692977118893507&alt=media)\n\n##Preprocessing\nSimilar to the previous competition solution, but simpler. Every landmark and xyz was used and standardized. Flip left-handed signer(rather than augmentation). MAX_LEN=768 was used. Any hand-crafted feature was not significant, likely due to the absence of complex relations between movements in frames.\n\n##Model\nEncoder is the same as the previous competition (Stacked Conv1DBlock + TransformerBlock) but increased in size (expand ratio 2->4 in Conv1DBlock) and depth (8 layers -> 17 layers). Single model has ~6.5M parameters. I applied padding='same'(rather than 'causal') and output stride=2 (requires slightly more logic for handling masking). In my case, mixing in Transformer blocks wasn't as effective as in the previous competition. Perhaps global features were less crucial in this comp. Additionally, I added one BN to input of the Conv1DBlock for more training stability(especially with awp + more epochs).\n\nCTC decoder used a single GRU layer followed by one FC layer. Attention Decoder used a single-layer Transformer decoder. Introducing augmentation to the Decoder input and adding up to 4 Decoder layers improves the performance of the attention decoder (up to +0.004). However, considering the number of parameters and inference speed, I considered it inefficient and thus used a single-layer decoder.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5003978%2F7d1f7aee69242af5e6aa5ef19c9460f6%2Fmodeldesign.drawio-2.svg?generation=1692993585923781&alt=media)\n##Augmentations\n* Random resample (0.5x ~ 1.5x to original length)\n* Random Affine\n* Random Cutout\n* Random token replacement on decoder input(prob=0.2)\n\nWhat I overlooked this time is that augmentations which showed no performance improvement or even degraded performance in shorter epochs might actually help improve performance in longer epochs. I experienced a similar phenomenon in the previous competition, but it seemed more pronounced in this one. By selecting augmentations solely based on the results from a 60epoch experiment, I think I missed out on many potentially beneficial augmentations when testing the 400epoch training in the final week.\n\n\n##Training\nEpoch = 400\nbs = 16 * num_replicas = 128\nLr = 5e-4 * num_replicas = 4e-3\nAWP = 0.2 starts at 0.1 * Epoch\nSchedule = CosineDecay with warmup ratio 0.1\nOptimizer = AdamW (slightly better than RAdam with Lookahead)\nLoss = CTC(weight=0.25) + CCE with label smoothing=0.1~0.25(weight=0.75)\n\nTraining takes around 14 hours with colab TPUv2-8(as colab TPU runtime recently reduced to 3~4 hours, needed 4 consecutive sessions to complete training).\nLonger Epoch always gave better CV(5fold split by id) and LB but got no time to try over 400 epochs.\n\n##LB history\n|  | public LB | private LB | \n| --- | --- | --- |\n| prevcompsinglemodel + CTC (or Attention) | 0.76 | 0.74 |\n| + deeper and wider model, add pose | 0.79 | 0.78 |\n| + 3 seed ensemble with Attention | 0.80 | 0.79 |\n| + ctc attention joint decoding | 0.81 | 0.80 | \n| + use all landmarks, longer epoch | 0.82 | 0.81 |\n\n\nSeeing solutions from other kagglers, not only the top solutions but also the public notebooks is always inspiring. Thanks to kagglers who shared their insights and ideas as always. And big congrats to @christofhenkel and @darraghdog for winning this highly competitive competition!\n\nTraining/Inference code: https://www.kaggle.com/competitions/asl-fingerspelling/discussion/436873\n\n\n",
    "2429971": "Thank you very much for a detailed explanation! I also learned a lot from your previous solution.",
    "2426159": "Thanks for sharing. And congrats brother. ",
    "2417749": "Hello @hoyso48 - Congratulations on your 2nd place finish!  Will the final tflite model or winning submission.zip file be available?  I'd like to use it with an Android app idea I'm developing.  Thanks, Matthew\n",
    "2413455": "I've been stuck with TPU training for over half a month with no progress. Looking forward to seeing your code.",
    "2410738": "@hoyso48 Congrats.\nYour solutions and implementation using tf/keras are genuinely inspiring and commendable. Thank you for sharing. :)",
    "2410642": "Congrats! Thanks for sharing your solution! \nI want to ask the paper you referred to in the previous ASR competition.",
    "2409604": "@hoyso48 Thanks for the write up; this is great as always. The way you approach the problem is inspiring. ",
    "2409131": "Congrats! \n\nI have a question. Did you upload the full dataset to Google Cloud? ",
    "2409083": "Congratulations. Thanks for sharing the details of the notebook. \nI think there is a benefit of separating training and inference notebook. \n- While submitting the notebook, I observe that notebook running time should be less than 9 hours.  \n- \n",
    "2408894": "Congratulations! I learned so much from what you share in these two competions, I could not start this competion without your previous sharing, especially your prvious code of showing stachastic path(tf.dropout with noise shape). Thanks!\nI will put some time on learning your sharing, ctc+attention method is so cool. \nAlso I think you might have a big boost by simply replace your encoder to squeezeformer or conformer.",
    "2412509": "Congrats on your great finish, @hoyso48. Your solution of the last ASL competition helped  me get started on this one, thanks!\n\nI thought about the joint CTC and attention decoding quite a bit but didn't find any great way to do it. Your solution is clean!\n\n>What I overlooked this time is that augmentations which showed no performance improvement or even degraded performance in shorter epochs might actually help improve performance in longer epochs.\n\nI experienced this too. Looking at the winners' huge set of augmentations there was a lot to be gained, it seems.",
    "2408876": "Thank you for sharing. I enjoyed reading how you made the combination of CTC and attention based decoding work. We tried and did not succeed there. ",
    "2430983": "Thank for your contribution, this is amazing!",
    "2411963": "Congratulations! Thanks for your write up!",
    "2411376": "Wow this is a great write up. Congrats!",
    "2408699": "Thanks for sharing your solution! Did you use google colab for everything that you've shared here?",
    "2426082": "",
    "2412843": "",
    "2416058": "Thanks for the share @hoyso48 ",
    "2413033": "Thank You for the summary sir!",
    "2412793": "thanks for the giving a nice solutions.",
    "2411085": "thanks for sharing your works.",
    "2410532": "Thanks for this guide.",
    "2410236": "Thanks for sharing! congratulations✨",
    "2409078": "Thank You Very Much !"
  }
}