{
  "id": 434983,
  "title": "[4th Place Solution] Conformer Encoder-Decoder Ensemble with beam search and edit_dist optimization",
  "url": "/competitions/asl-fingerspelling/discussion/434983",
  "author_name": "flg",
  "post_date": "2023-08-27T13:12:15.809000",
  "votes": 19,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Thanks to Kaggle and the organizer for running this competition! It was a quite unique challenge and even after three months of optimizing it feels like there are still so many things to improve, which is quite special in my opinion.</p>\n<h2>TLDR</h2>\n<p>My solution is an ensemble of two encoder-decoder models. The encoder is a 12 layer adapted conformer and the decoder is a two layer regular transformer. I added various augmentations and training techniques to align the training objective with the edit_distance competition metric. For decoding I implemented a (cached) beam search for TFLite.</p>\n<h2>Challenge and plan</h2>\n<p>The goal of the competition was to translate sign language spelling from videos that were preprocessed with human pose recognition. Submissions were made as TFLite models with a limited OPs set and evaluated using the Levenshtein Edit Distance.</p>\n<p>These choices had a couple of implications for modelling:</p>\n<ul>\n<li>\"honest\" predictions are often not edit-distance-optimal, esp. when the model recognizes no characters in a phrase,<br>\nthe honest prediction \"\" achieves a score of 0.0 while \"2 a-e -aroe\" scores 0.16,<br>\nsee <a href=\"https://www.kaggle.com/code/anokas/static-greedy-baseline-0-157-lb\" target=\"_blank\">Anoka's Static Greedy Baseline</a></li>\n<li>TFLite models in this competition only allowed a very restrictive ops-set, which meant for example that all the native<br>\nimplementations of beam_search and similar algorithms were not supported</li>\n<li>there were time and size limitations to the model (40MB) causing where to \"spend\" your parameters becoming a major<br>\ndesign decision</li>\n<li>the datasets contained quite a few samples where most or all data was missing</li>\n</ul>\n<p>With these things in mind, I assumed \"making things up\" would be a significant part of good predictions and<br>\nencoder-decoder architectures seemed naturally aligned to this. Additionally having a decoder abstracts away one of the time dimensions which made developing downstream algorithms like beam search or ensembling easier. My early tests also suggested encoder-decoders to work slightly better than CTC, so I went with that architecture.</p>\n<h2>Data</h2>\n<p>My model used 214 inputs: 21 LHand, 21 RHand, 25 Pose, 1 Nose and 40 Lips points, each using x- and y-coordinates. Data was normalized, NaNs zero-filled and the deltas to t-1 and t-2 values were used as additional Features. During training a maximum of 500 frames was used with longer sequences being resized.</p>\n<p>I applied quite a few data augmentations:</p>\n<ul>\n<li>Flip left-right</li>\n<li>Resample along the time dimension</li>\n<li>Scale / Translate / Rotate</li>\n<li>Mask up to 60% of all frames (worked better than masking sequences)</li>\n<li>Spatial cutout (similar to <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406684\" target=\"_blank\">Hoyso48's 1st place solution</a>)   <br></li>\n</ul>\n<p>All of these made a significant impact.</p>\n<p>On top of that I transformed the tokens too. I used the fact that when replacing a single token in a phrase with a random token, the ground truth is still an edit-distance-optimal target. This change gave a quite nice boost of +0.008 (using smaller models). I also added single token deletions and insertions but they had a minor impact (if at all).</p>\n<p>I split the data five-fold and most of my experiments used only one fold to train with smaller models due to compute restrictions.</p>\n<h2>Model</h2>\n<p>The base of my model was a deep <a href=\"https://arxiv.org/pdf/2005.08100.pdf\" target=\"_blank\">Conformer</a> encoder followed by a two-layer transformer decoder. The encoder used twelve layers with dimension of 144. The MHSA had four heads with dim-per-head of 64 and the Convolution used a kernel of size 65.<br>\nLike the original formulation the model used two macaron-style feed forwards with an expansion factor of four. I made some additional small changes, like changing the position of the BatchNorm and adding DropPath to each Submodule of the Conformer. The model used a drop rate of 0.1 almost everywhere, only before the final classifier it used 0.3. Instead of causal padding I used same padding and explicitly zeroed-out the padded parts.  </p>\n<p>On the decoder side I tested many different configurations but ended up using a very slim, two layer transformer<br>\ndecoder. It used four attention heads with dim-per-head of 32 and a feed forward with expansion factor of only two. This was the smallest configuration that I could train without significant performance drop off. Using a small decoder was important since the autoregressive decoding is very performance intensive.</p>\n<h2>Training</h2>\n<p>A full training run for a single twelve layer encoder model took around two days on my local 3090. To experiment with different architectures, augmentations etc., I only trained shallower models on ~20% of the data for most of the competition.<br></p>\n<p>The final training used a cross entropy loss and RAdam optimizer (but AdamW with warmup worked pretty much the same) with a peak lr of 1e-3 and cosine schedule. Weight decay of 2e-6 and label smoothing 0.2 (very minor effect) were used for regularization in combination with light gaussian weight noise (had similar effect as AWP in my test, but lower overhead). I trained for 300 epochs, the first 100 of which used the supplemental data.<br></p>\n<p>I used minimum word error rate training after the model finished training. Where character-based edit distance is used as \"word error rate\". The method starts with a converged model and uses beam search to generate say the top four predictions. It then calculates each prediction's edit distance and uses this as a weight for the model's predicted probabilities. See for example <a href=\"https://arxiv.org/abs/1712.01818\" target=\"_blank\">Minimum Word Error Rate Training for Attention-based Sequence-to-Sequence Models</a>. The method's results are unstable even after optimizing it quite a bit. However, short training runs of 1-5 epochs gave very considerable gains in early testing. Unfortunately on the final large, ensembled model it was a rather modest improvement of 0.001-0.002.</p>\n<h2>Beam search and inference-time optimizations</h2>\n<p>Using an ensemble of two models, it was easy to reach the 40MB model size limit. To max out the run time dimension too, I implemented a beam search algorithm that is compatible with the restricted TFLite ops set of this competition. Using it with cached autoregressive decoding allowed me to use beam sizes of five to six (with six sometimes failing the 5h limit). This resulted in + 0.005 on the final ensemble (and even more on earlier, weaker models). The implementation was a bit tricky as there are a few edge cases like having to reorder the decoding caches when beams are changed etc. To prevent the early termination problem when decoding with beam search I used a linear length penalty of 0.15.<br></p>\n<p>On top of this I realized my model achieved an edit distance of 0.0 on low information samples (e.g. &lt; 50 frames and &lt; 5 frames with any hand showing). But we knew that a greedy prediction of e.g. \"2 a-e -aroe\" gets a score of 0.16. Since most of these low information samples seem entirely corrupted, I simply replace the model's predictions on these with a constant prediction. I used \" a-e -are\", which slightly different from the greedy one mentioned before as I optimized it towards shorter, low information sequences.<br><br>\nIn the end, adding this one line:</p>\n<p><code>x = tf.cond(num_frames &lt; 50 and num_hand_frames &lt;= 3,\n    lambda: tf.constant([[59, 0, 32, 12, 36, 0, 12, 32, 49, 36, 60]]),\n    lambda: tf.identity(x))</code></p>\n<p>gave an improvement of +0.005 across the board (local eval, private and public LB for all models). Which is as much as the whole beam search …</p>\n<h2>What worked and didn't</h2>\n<ul>\n<li>Beam search gave a decent +0.005 improvement</li>\n<li>Replacing the model's prediction on corrupt data samples with a constant default prediction gave +0.005</li>\n<li>Deeper models worked better than wider ones</li>\n<li>MWER-training gave a small improvement (+0.001 - +0.002) - however, this was with beam search k=5, with greedy decoding gains were larger (+0.005 in local eval)</li>\n<li>Replacing a single input token with a random one was a decent augmentation</li>\n<li>CTC didn't help as an auxiliary loss</li>\n<li>masking decoder input did not help (when random token replacement was used)</li>\n<li>z-coordinates did not help</li>\n</ul>\n<p>As always: really looking forward to reading everyone's solutions. Let me know if there are any questions. Code coming <em>soon</em>.</p>",
  "messages": [
    {
      "id": 2411197,
      "postDate": "2023-08-27T13:12:15.810Z",
      "content": "<p>Thanks to Kaggle and the organizer for running this competition! It was a quite unique challenge and even after three months of optimizing it feels like there are still so many things to improve, which is quite special in my opinion.</p>\n<h2>TLDR</h2>\n<p>My solution is an ensemble of two encoder-decoder models. The encoder is a 12 layer adapted conformer and the decoder is a two layer regular transformer. I added various augmentations and training techniques to align the training objective with the edit_distance competition metric. For decoding I implemented a (cached) beam search for TFLite.</p>\n<h2>Challenge and plan</h2>\n<p>The goal of the competition was to translate sign language spelling from videos that were preprocessed with human pose recognition. Submissions were made as TFLite models with a limited OPs set and evaluated using the Levenshtein Edit Distance.</p>\n<p>These choices had a couple of implications for modelling:</p>\n<ul>\n<li>\"honest\" predictions are often not edit-distance-optimal, esp. when the model recognizes no characters in a phrase,<br>\nthe honest prediction \"\" achieves a score of 0.0 while \"2 a-e -aroe\" scores 0.16,<br>\nsee <a href=\"https://www.kaggle.com/code/anokas/static-greedy-baseline-0-157-lb\" target=\"_blank\">Anoka's Static Greedy Baseline</a></li>\n<li>TFLite models in this competition only allowed a very restrictive ops-set, which meant for example that all the native<br>\nimplementations of beam_search and similar algorithms were not supported</li>\n<li>there were time and size limitations to the model (40MB) causing where to \"spend\" your parameters becoming a major<br>\ndesign decision</li>\n<li>the datasets contained quite a few samples where most or all data was missing</li>\n</ul>\n<p>With these things in mind, I assumed \"making things up\" would be a significant part of good predictions and<br>\nencoder-decoder architectures seemed naturally aligned to this. Additionally having a decoder abstracts away one of the time dimensions which made developing downstream algorithms like beam search or ensembling easier. My early tests also suggested encoder-decoders to work slightly better than CTC, so I went with that architecture.</p>\n<h2>Data</h2>\n<p>My model used 214 inputs: 21 LHand, 21 RHand, 25 Pose, 1 Nose and 40 Lips points, each using x- and y-coordinates. Data was normalized, NaNs zero-filled and the deltas to t-1 and t-2 values were used as additional Features. During training a maximum of 500 frames was used with longer sequences being resized.</p>\n<p>I applied quite a few data augmentations:</p>\n<ul>\n<li>Flip left-right</li>\n<li>Resample along the time dimension</li>\n<li>Scale / Translate / Rotate</li>\n<li>Mask up to 60% of all frames (worked better than masking sequences)</li>\n<li>Spatial cutout (similar to <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406684\" target=\"_blank\">Hoyso48's 1st place solution</a>)   <br></li>\n</ul>\n<p>All of these made a significant impact.</p>\n<p>On top of that I transformed the tokens too. I used the fact that when replacing a single token in a phrase with a random token, the ground truth is still an edit-distance-optimal target. This change gave a quite nice boost of +0.008 (using smaller models). I also added single token deletions and insertions but they had a minor impact (if at all).</p>\n<p>I split the data five-fold and most of my experiments used only one fold to train with smaller models due to compute restrictions.</p>\n<h2>Model</h2>\n<p>The base of my model was a deep <a href=\"https://arxiv.org/pdf/2005.08100.pdf\" target=\"_blank\">Conformer</a> encoder followed by a two-layer transformer decoder. The encoder used twelve layers with dimension of 144. The MHSA had four heads with dim-per-head of 64 and the Convolution used a kernel of size 65.<br>\nLike the original formulation the model used two macaron-style feed forwards with an expansion factor of four. I made some additional small changes, like changing the position of the BatchNorm and adding DropPath to each Submodule of the Conformer. The model used a drop rate of 0.1 almost everywhere, only before the final classifier it used 0.3. Instead of causal padding I used same padding and explicitly zeroed-out the padded parts.  </p>\n<p>On the decoder side I tested many different configurations but ended up using a very slim, two layer transformer<br>\ndecoder. It used four attention heads with dim-per-head of 32 and a feed forward with expansion factor of only two. This was the smallest configuration that I could train without significant performance drop off. Using a small decoder was important since the autoregressive decoding is very performance intensive.</p>\n<h2>Training</h2>\n<p>A full training run for a single twelve layer encoder model took around two days on my local 3090. To experiment with different architectures, augmentations etc., I only trained shallower models on ~20% of the data for most of the competition.<br></p>\n<p>The final training used a cross entropy loss and RAdam optimizer (but AdamW with warmup worked pretty much the same) with a peak lr of 1e-3 and cosine schedule. Weight decay of 2e-6 and label smoothing 0.2 (very minor effect) were used for regularization in combination with light gaussian weight noise (had similar effect as AWP in my test, but lower overhead). I trained for 300 epochs, the first 100 of which used the supplemental data.<br></p>\n<p>I used minimum word error rate training after the model finished training. Where character-based edit distance is used as \"word error rate\". The method starts with a converged model and uses beam search to generate say the top four predictions. It then calculates each prediction's edit distance and uses this as a weight for the model's predicted probabilities. See for example <a href=\"https://arxiv.org/abs/1712.01818\" target=\"_blank\">Minimum Word Error Rate Training for Attention-based Sequence-to-Sequence Models</a>. The method's results are unstable even after optimizing it quite a bit. However, short training runs of 1-5 epochs gave very considerable gains in early testing. Unfortunately on the final large, ensembled model it was a rather modest improvement of 0.001-0.002.</p>\n<h2>Beam search and inference-time optimizations</h2>\n<p>Using an ensemble of two models, it was easy to reach the 40MB model size limit. To max out the run time dimension too, I implemented a beam search algorithm that is compatible with the restricted TFLite ops set of this competition. Using it with cached autoregressive decoding allowed me to use beam sizes of five to six (with six sometimes failing the 5h limit). This resulted in + 0.005 on the final ensemble (and even more on earlier, weaker models). The implementation was a bit tricky as there are a few edge cases like having to reorder the decoding caches when beams are changed etc. To prevent the early termination problem when decoding with beam search I used a linear length penalty of 0.15.<br></p>\n<p>On top of this I realized my model achieved an edit distance of 0.0 on low information samples (e.g. &lt; 50 frames and &lt; 5 frames with any hand showing). But we knew that a greedy prediction of e.g. \"2 a-e -aroe\" gets a score of 0.16. Since most of these low information samples seem entirely corrupted, I simply replace the model's predictions on these with a constant prediction. I used \" a-e -are\", which slightly different from the greedy one mentioned before as I optimized it towards shorter, low information sequences.<br><br>\nIn the end, adding this one line:</p>\n<p><code>x = tf.cond(num_frames &lt; 50 and num_hand_frames &lt;= 3,\n    lambda: tf.constant([[59, 0, 32, 12, 36, 0, 12, 32, 49, 36, 60]]),\n    lambda: tf.identity(x))</code></p>\n<p>gave an improvement of +0.005 across the board (local eval, private and public LB for all models). Which is as much as the whole beam search …</p>\n<h2>What worked and didn't</h2>\n<ul>\n<li>Beam search gave a decent +0.005 improvement</li>\n<li>Replacing the model's prediction on corrupt data samples with a constant default prediction gave +0.005</li>\n<li>Deeper models worked better than wider ones</li>\n<li>MWER-training gave a small improvement (+0.001 - +0.002) - however, this was with beam search k=5, with greedy decoding gains were larger (+0.005 in local eval)</li>\n<li>Replacing a single input token with a random one was a decent augmentation</li>\n<li>CTC didn't help as an auxiliary loss</li>\n<li>masking decoder input did not help (when random token replacement was used)</li>\n<li>z-coordinates did not help</li>\n</ul>\n<p>As always: really looking forward to reading everyone's solutions. Let me know if there are any questions. Code coming <em>soon</em>.</p>",
      "rawMarkdown": "Thanks to Kaggle and the organizer for running this competition! It was a quite unique challenge and even after three months of optimizing it feels like there are still so many things to improve, which is quite special in my opinion.\n\n## TLDR\n\nMy solution is an ensemble of two encoder-decoder models. The encoder is a 12 layer adapted conformer and the decoder is a two layer regular transformer. I added various augmentations and training techniques to align the training objective with the edit_distance competition metric. For decoding I implemented a (cached) beam search for TFLite.\n\n## Challenge and plan\n\nThe goal of the competition was to translate sign language spelling from videos that were preprocessed with human pose recognition. Submissions were made as TFLite models with a limited OPs set and evaluated using the Levenshtein Edit Distance.\n\nThese choices had a couple of implications for modelling:\n\n- \"honest\" predictions are often not edit-distance-optimal, esp. when the model recognizes no characters in a phrase,\n  the honest prediction \"\" achieves a score of 0.0 while \"2 a-e -aroe\" scores 0.16,\n  see [Anoka's Static Greedy Baseline](https://www.kaggle.com/code/anokas/static-greedy-baseline-0-157-lb)\n- TFLite models in this competition only allowed a very restrictive ops-set, which meant for example that all the native\n  implementations of beam_search and similar algorithms were not supported\n- there were time and size limitations to the model (40MB) causing where to \"spend\" your parameters becoming a major\n  design decision\n- the datasets contained quite a few samples where most or all data was missing\n\nWith these things in mind, I assumed \"making things up\" would be a significant part of good predictions and\nencoder-decoder architectures seemed naturally aligned to this. Additionally having a decoder abstracts away one of the time dimensions which made developing downstream algorithms like beam search or ensembling easier. My early tests also suggested encoder-decoders to work slightly better than CTC, so I went with that architecture.\n\n## Data\n\nMy model used 214 inputs: 21 LHand, 21 RHand, 25 Pose, 1 Nose and 40 Lips points, each using x- and y-coordinates. Data was normalized, NaNs zero-filled and the deltas to t-1 and t-2 values were used as additional Features. During training a maximum of 500 frames was used with longer sequences being resized.\n\nI applied quite a few data augmentations:\n\n- Flip left-right\n- Resample along the time dimension\n- Scale / Translate / Rotate\n- Mask up to 60% of all frames (worked better than masking sequences)\n- Spatial cutout (similar to [Hoyso48's 1st place solution](https://www.kaggle.com/competitions/asl-signs/discussion/406684))   <br>\n\n\nAll of these made a significant impact.\n\nOn top of that I transformed the tokens too. I used the fact that when replacing a single token in a phrase with a random token, the ground truth is still an edit-distance-optimal target. This change gave a quite nice boost of +0.008 (using smaller models). I also added single token deletions and insertions but they had a minor impact (if at all).\n\nI split the data five-fold and most of my experiments used only one fold to train with smaller models due to compute restrictions.\n\n\n## Model\n\nThe base of my model was a deep [Conformer](https://arxiv.org/pdf/2005.08100.pdf) encoder followed by a two-layer transformer decoder. The encoder used twelve layers with dimension of 144. The MHSA had four heads with dim-per-head of 64 and the Convolution used a kernel of size 65.\nLike the original formulation the model used two macaron-style feed forwards with an expansion factor of four. I made some additional small changes, like changing the position of the BatchNorm and adding DropPath to each Submodule of the Conformer. The model used a drop rate of 0.1 almost everywhere, only before the final classifier it used 0.3. Instead of causal padding I used same padding and explicitly zeroed-out the padded parts.  \n \nOn the decoder side I tested many different configurations but ended up using a very slim, two layer transformer\ndecoder. It used four attention heads with dim-per-head of 32 and a feed forward with expansion factor of only two. This was the smallest configuration that I could train without significant performance drop off. Using a small decoder was important since the autoregressive decoding is very performance intensive.\n\n## Training\n\nA full training run for a single twelve layer encoder model took around two days on my local 3090. To experiment with different architectures, augmentations etc., I only trained shallower models on ~20% of the data for most of the competition.<br>\n\nThe final training used a cross entropy loss and RAdam optimizer (but AdamW with warmup worked pretty much the same) with a peak lr of 1e-3 and cosine schedule. Weight decay of 2e-6 and label smoothing 0.2 (very minor effect) were used for regularization in combination with light gaussian weight noise (had similar effect as AWP in my test, but lower overhead). I trained for 300 epochs, the first 100 of which used the supplemental data.<br>\n\nI used minimum word error rate training after the model finished training. Where character-based edit distance is used as \"word error rate\". The method starts with a converged model and uses beam search to generate say the top four predictions. It then calculates each prediction's edit distance and uses this as a weight for the model's predicted probabilities. See for example [Minimum Word Error Rate Training for Attention-based Sequence-to-Sequence Models](https://arxiv.org/abs/1712.01818). The method's results are unstable even after optimizing it quite a bit. However, short training runs of 1-5 epochs gave very considerable gains in early testing. Unfortunately on the final large, ensembled model it was a rather modest improvement of 0.001-0.002.\n\n## Beam search and inference-time optimizations\n\nUsing an ensemble of two models, it was easy to reach the 40MB model size limit. To max out the run time dimension too, I implemented a beam search algorithm that is compatible with the restricted TFLite ops set of this competition. Using it with cached autoregressive decoding allowed me to use beam sizes of five to six (with six sometimes failing the 5h limit). This resulted in + 0.005 on the final ensemble (and even more on earlier, weaker models). The implementation was a bit tricky as there are a few edge cases like having to reorder the decoding caches when beams are changed etc. To prevent the early termination problem when decoding with beam search I used a linear length penalty of 0.15.<br>\n\nOn top of this I realized my model achieved an edit distance of 0.0 on low information samples (e.g. < 50 frames and < 5 frames with any hand showing). But we knew that a greedy prediction of e.g. \"2 a-e -aroe\" gets a score of 0.16. Since most of these low information samples seem entirely corrupted, I simply replace the model's predictions on these with a constant prediction. I used \" a-e -are\", which slightly different from the greedy one mentioned before as I optimized it towards shorter, low information sequences.<br>\nIn the end, adding this one line:\n\n```x = tf.cond(num_frames < 50 and num_hand_frames <= 3,\n    lambda: tf.constant([[59, 0, 32, 12, 36, 0, 12, 32, 49, 36, 60]]),\n    lambda: tf.identity(x))```\n\ngave an improvement of +0.005 across the board (local eval, private and public LB for all models). Which is as much as the whole beam search ...\n\n## What worked and didn't\n\n- Beam search gave a decent +0.005 improvement\n- Replacing the model's prediction on corrupt data samples with a constant default prediction gave +0.005\n- Deeper models worked better than wider ones\n- MWER-training gave a small improvement (+0.001 - +0.002) - however, this was with beam search k=5, with greedy decoding gains were larger (+0.005 in local eval)\n- Replacing a single input token with a random one was a decent augmentation\n- CTC didn't help as an auxiliary loss\n- masking decoder input did not help (when random token replacement was used)\n- z-coordinates did not help\n\nAs always: really looking forward to reading everyone's solutions. Let me know if there are any questions. Code coming _soon_.",
      "votes": 19
    },
    {
      "id": 2412137,
      "postDate": "2023-08-28T05:33:20.380Z",
      "content": "<p>Congratulations. I observe that sometime Data augmentation can backfire as two finger spellings may be confusing to the model. <br>\nFor example:  flipping 'C'  180 degrees may make it a 'G'. There are many such instances.</p>",
      "rawMarkdown": "Congratulations. I observe that sometime Data augmentation can backfire as two finger spellings may be confusing to the model. \nFor example:  flipping 'C'  180 degrees may make it a 'G'. There are many such instances.",
      "votes": 1,
      "replies": [
        {
          "id": 2412447,
          "postDate": "2023-08-28T09:27:05.303Z",
          "content": "<p>The flip left right augmentation that I used, mirrors the whole body and reassigns the data so that left hand becomes right hand, left body becomes right body etc. It only makes it look like the sign was made with the other hand. Since most of the data is right-handed this helped improve left-handed predictions.</p>",
          "rawMarkdown": "The flip left right augmentation that I used, mirrors the whole body and reassigns the data so that left hand becomes right hand, left body becomes right body etc. It only makes it look like the sign was made with the other hand. Since most of the data is right-handed this helped improve left-handed predictions.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2411931,
      "postDate": "2023-08-28T00:57:03.523Z",
      "content": "<blockquote>\n  <p>Replacing a single input token with a random one was a decent augmentation</p>\n</blockquote>\n<p>Replacing a single input token with a random one may appear to be a method that disrupts the training data without a clear rationale. Why is this approach effective as an augmentation technique?</p>",
      "rawMarkdown": ">Replacing a single input token with a random one was a decent augmentation\n \nReplacing a single input token with a random one may appear to be a method that disrupts the training data without a clear rationale. Why is this approach effective as an augmentation technique?",
      "votes": 1,
      "replies": [
        {
          "id": 2412432,
          "postDate": "2023-08-28T09:16:34.913Z",
          "content": "<p>It is an effective way to avoid exposure bias. For causal prediction, the model might predict a wrong character. We want the model can still make correct predictions after previous wrong predictions. Applying this augmentation is to mimic this during inference.</p>",
          "rawMarkdown": "It is an effective way to avoid exposure bias. For causal prediction, the model might predict a wrong character. We want the model can still make correct predictions after previous wrong predictions. Applying this augmentation is to mimic this during inference.",
          "votes": 3,
          "replies": [
            {
              "id": 2412461,
              "postDate": "2023-08-28T09:35:31.290Z",
              "content": "<p>You're so clever, you've really enlightened me.</p>",
              "rawMarkdown": "You're so clever, you've really enlightened me."
            }
          ]
        },
        {
          "id": 2412457,
          "postDate": "2023-08-28T09:31:38.990Z",
          "content": "<p>Similar to what <a href=\"https://www.kaggle.com/baohaoliao\" target=\"_blank\">@baohaoliao</a> said: during training with cross entropy the model always receives perfect input phrases. There are never any wrong predictions in the tokens leading up to the current one. However, during inference predictions are made autoregressively which will make mistakes (if only for missing data). This can disrupt its predictions quite a bit and the augmentation is supposed to help with that. Another similar method I and some competitors tried is masking parts of the token input, which follows a similar logic.</p>",
          "rawMarkdown": "Similar to what @baohaoliao said: during training with cross entropy the model always receives perfect input phrases. There are never any wrong predictions in the tokens leading up to the current one. However, during inference predictions are made autoregressively which will make mistakes (if only for missing data). This can disrupt its predictions quite a bit and the augmentation is supposed to help with that. Another similar method I and some competitors tried is masking parts of the token input, which follows a similar logic.",
          "votes": 2,
          "replies": [
            {
              "id": 2412481,
              "postDate": "2023-08-28T09:45:43.613Z",
              "content": "<p>Your explanation is more detailed. I finally understand the rationale behind this augmentation, but does this augmentation lead to the training set not being able to fit perfectly, resulting in the test set score not reaching 1?</p>",
              "rawMarkdown": "Your explanation is more detailed. I finally understand the rationale behind this augmentation, but does this augmentation lead to the training set not being able to fit perfectly, resulting in the test set score not reaching 1?"
            },
            {
              "id": 2412547,
              "postDate": "2023-08-28T10:59:21.070Z",
              "content": "<p>The idea is to disrupt the model's training, so it does less memorization of training data and more generalization. Ideally this should make the train score go down but the test score go up.</p>",
              "rawMarkdown": "The idea is to disrupt the model's training, so it does less memorization of training data and more generalization. Ideally this should make the train score go down but the test score go up.",
              "votes": 2
            },
            {
              "id": 2412702,
              "postDate": "2023-08-28T12:52:10.900Z",
              "content": "<p>I like your answer.</p>",
              "rawMarkdown": "I like your answer."
            }
          ]
        }
      ]
    },
    {
      "id": 2569238,
      "postDate": "2023-12-21T06:46:41.317Z",
      "content": "<p>May I know if code is available? if so could you please share the link? Thank you in advance</p>",
      "rawMarkdown": "May I know if code is available? if so could you please share the link? Thank you in advance"
    },
    {
      "id": 2420782,
      "postDate": "2023-09-02T22:23:25.543Z",
      "content": "<p>Hello! Thank you for your write-up!  Would you be able to share what stddev you used for your gaussian noise?</p>",
      "rawMarkdown": "Hello! Thank you for your write-up!  Would you be able to share what stddev you used for your gaussian noise?"
    },
    {
      "id": 2413030,
      "postDate": "2023-08-28T16:31:37.190Z",
      "content": "<p>Superb effort!</p>",
      "rawMarkdown": "Superb effort!"
    },
    {
      "id": 2426075,
      "postDate": "2023-09-06T11:41:59.767Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2412137,
      "author_name": "C R Suthikshn Kumar",
      "author_url": "",
      "post_date": "2023-08-28T05:33:20.380000",
      "content": "<p>Congratulations. I observe that sometime Data augmentation can backfire as two finger spellings may be confusing to the model. <br>\nFor example:  flipping 'C'  180 degrees may make it a 'G'. There are many such instances.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2412447,
          "author_name": "flg",
          "author_url": "",
          "post_date": "2023-08-28T09:27:05.303000",
          "content": "<p>The flip left right augmentation that I used, mirrors the whole body and reassigns the data so that left hand becomes right hand, left body becomes right body etc. It only makes it look like the sign was made with the other hand. Since most of the data is right-handed this helped improve left-handed predictions.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2411931,
      "author_name": "Scenery SunFireInk",
      "author_url": "",
      "post_date": "2023-08-28T00:57:03.523000",
      "content": "<blockquote>\n  <p>Replacing a single input token with a random one was a decent augmentation</p>\n</blockquote>\n<p>Replacing a single input token with a random one may appear to be a method that disrupts the training data without a clear rationale. Why is this approach effective as an augmentation technique?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2412432,
          "author_name": "bliao",
          "author_url": "",
          "post_date": "2023-08-28T09:16:34.913000",
          "content": "<p>It is an effective way to avoid exposure bias. For causal prediction, the model might predict a wrong character. We want the model can still make correct predictions after previous wrong predictions. Applying this augmentation is to mimic this during inference.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2412461,
              "author_name": "Scenery SunFireInk",
              "author_url": "",
              "post_date": "2023-08-28T09:35:31.290000",
              "content": "<p>You're so clever, you've really enlightened me.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2412457,
          "author_name": "flg",
          "author_url": "",
          "post_date": "2023-08-28T09:31:38.990000",
          "content": "<p>Similar to what <a href=\"https://www.kaggle.com/baohaoliao\" target=\"_blank\">@baohaoliao</a> said: during training with cross entropy the model always receives perfect input phrases. There are never any wrong predictions in the tokens leading up to the current one. However, during inference predictions are made autoregressively which will make mistakes (if only for missing data). This can disrupt its predictions quite a bit and the augmentation is supposed to help with that. Another similar method I and some competitors tried is masking parts of the token input, which follows a similar logic.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2412481,
              "author_name": "Scenery SunFireInk",
              "author_url": "",
              "post_date": "2023-08-28T09:45:43.613000",
              "content": "<p>Your explanation is more detailed. I finally understand the rationale behind this augmentation, but does this augmentation lead to the training set not being able to fit perfectly, resulting in the test set score not reaching 1?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2412547,
              "author_name": "flg",
              "author_url": "",
              "post_date": "2023-08-28T10:59:21.070000",
              "content": "<p>The idea is to disrupt the model's training, so it does less memorization of training data and more generalization. Ideally this should make the train score go down but the test score go up.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2412702,
              "author_name": "Scenery SunFireInk",
              "author_url": "",
              "post_date": "2023-08-28T12:52:10.900000",
              "content": "<p>I like your answer.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2569238,
      "author_name": "G V S King MAP",
      "author_url": "",
      "post_date": "2023-12-21T06:46:41.317000",
      "content": "<p>May I know if code is available? if so could you please share the link? Thank you in advance</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2420782,
      "author_name": "Andy Atkinson",
      "author_url": "",
      "post_date": "2023-09-02T22:23:25.543000",
      "content": "<p>Hello! Thank you for your write-up!  Would you be able to share what stddev you used for your gaussian noise?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2413030,
      "author_name": "Mystic Shadow",
      "author_url": "",
      "post_date": "2023-08-28T16:31:37.190000",
      "content": "<p>Superb effort!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2426075,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-09-06T11:41:59.767000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2411197": "Thanks to Kaggle and the organizer for running this competition! It was a quite unique challenge and even after three months of optimizing it feels like there are still so many things to improve, which is quite special in my opinion.\n\n## TLDR\n\nMy solution is an ensemble of two encoder-decoder models. The encoder is a 12 layer adapted conformer and the decoder is a two layer regular transformer. I added various augmentations and training techniques to align the training objective with the edit_distance competition metric. For decoding I implemented a (cached) beam search for TFLite.\n\n## Challenge and plan\n\nThe goal of the competition was to translate sign language spelling from videos that were preprocessed with human pose recognition. Submissions were made as TFLite models with a limited OPs set and evaluated using the Levenshtein Edit Distance.\n\nThese choices had a couple of implications for modelling:\n\n- \"honest\" predictions are often not edit-distance-optimal, esp. when the model recognizes no characters in a phrase,\n  the honest prediction \"\" achieves a score of 0.0 while \"2 a-e -aroe\" scores 0.16,\n  see [Anoka's Static Greedy Baseline](https://www.kaggle.com/code/anokas/static-greedy-baseline-0-157-lb)\n- TFLite models in this competition only allowed a very restrictive ops-set, which meant for example that all the native\n  implementations of beam_search and similar algorithms were not supported\n- there were time and size limitations to the model (40MB) causing where to \"spend\" your parameters becoming a major\n  design decision\n- the datasets contained quite a few samples where most or all data was missing\n\nWith these things in mind, I assumed \"making things up\" would be a significant part of good predictions and\nencoder-decoder architectures seemed naturally aligned to this. Additionally having a decoder abstracts away one of the time dimensions which made developing downstream algorithms like beam search or ensembling easier. My early tests also suggested encoder-decoders to work slightly better than CTC, so I went with that architecture.\n\n## Data\n\nMy model used 214 inputs: 21 LHand, 21 RHand, 25 Pose, 1 Nose and 40 Lips points, each using x- and y-coordinates. Data was normalized, NaNs zero-filled and the deltas to t-1 and t-2 values were used as additional Features. During training a maximum of 500 frames was used with longer sequences being resized.\n\nI applied quite a few data augmentations:\n\n- Flip left-right\n- Resample along the time dimension\n- Scale / Translate / Rotate\n- Mask up to 60% of all frames (worked better than masking sequences)\n- Spatial cutout (similar to [Hoyso48's 1st place solution](https://www.kaggle.com/competitions/asl-signs/discussion/406684))   <br>\n\n\nAll of these made a significant impact.\n\nOn top of that I transformed the tokens too. I used the fact that when replacing a single token in a phrase with a random token, the ground truth is still an edit-distance-optimal target. This change gave a quite nice boost of +0.008 (using smaller models). I also added single token deletions and insertions but they had a minor impact (if at all).\n\nI split the data five-fold and most of my experiments used only one fold to train with smaller models due to compute restrictions.\n\n\n## Model\n\nThe base of my model was a deep [Conformer](https://arxiv.org/pdf/2005.08100.pdf) encoder followed by a two-layer transformer decoder. The encoder used twelve layers with dimension of 144. The MHSA had four heads with dim-per-head of 64 and the Convolution used a kernel of size 65.\nLike the original formulation the model used two macaron-style feed forwards with an expansion factor of four. I made some additional small changes, like changing the position of the BatchNorm and adding DropPath to each Submodule of the Conformer. The model used a drop rate of 0.1 almost everywhere, only before the final classifier it used 0.3. Instead of causal padding I used same padding and explicitly zeroed-out the padded parts.  \n \nOn the decoder side I tested many different configurations but ended up using a very slim, two layer transformer\ndecoder. It used four attention heads with dim-per-head of 32 and a feed forward with expansion factor of only two. This was the smallest configuration that I could train without significant performance drop off. Using a small decoder was important since the autoregressive decoding is very performance intensive.\n\n## Training\n\nA full training run for a single twelve layer encoder model took around two days on my local 3090. To experiment with different architectures, augmentations etc., I only trained shallower models on ~20% of the data for most of the competition.<br>\n\nThe final training used a cross entropy loss and RAdam optimizer (but AdamW with warmup worked pretty much the same) with a peak lr of 1e-3 and cosine schedule. Weight decay of 2e-6 and label smoothing 0.2 (very minor effect) were used for regularization in combination with light gaussian weight noise (had similar effect as AWP in my test, but lower overhead). I trained for 300 epochs, the first 100 of which used the supplemental data.<br>\n\nI used minimum word error rate training after the model finished training. Where character-based edit distance is used as \"word error rate\". The method starts with a converged model and uses beam search to generate say the top four predictions. It then calculates each prediction's edit distance and uses this as a weight for the model's predicted probabilities. See for example [Minimum Word Error Rate Training for Attention-based Sequence-to-Sequence Models](https://arxiv.org/abs/1712.01818). The method's results are unstable even after optimizing it quite a bit. However, short training runs of 1-5 epochs gave very considerable gains in early testing. Unfortunately on the final large, ensembled model it was a rather modest improvement of 0.001-0.002.\n\n## Beam search and inference-time optimizations\n\nUsing an ensemble of two models, it was easy to reach the 40MB model size limit. To max out the run time dimension too, I implemented a beam search algorithm that is compatible with the restricted TFLite ops set of this competition. Using it with cached autoregressive decoding allowed me to use beam sizes of five to six (with six sometimes failing the 5h limit). This resulted in + 0.005 on the final ensemble (and even more on earlier, weaker models). The implementation was a bit tricky as there are a few edge cases like having to reorder the decoding caches when beams are changed etc. To prevent the early termination problem when decoding with beam search I used a linear length penalty of 0.15.<br>\n\nOn top of this I realized my model achieved an edit distance of 0.0 on low information samples (e.g. < 50 frames and < 5 frames with any hand showing). But we knew that a greedy prediction of e.g. \"2 a-e -aroe\" gets a score of 0.16. Since most of these low information samples seem entirely corrupted, I simply replace the model's predictions on these with a constant prediction. I used \" a-e -are\", which slightly different from the greedy one mentioned before as I optimized it towards shorter, low information sequences.<br>\nIn the end, adding this one line:\n\n```x = tf.cond(num_frames < 50 and num_hand_frames <= 3,\n    lambda: tf.constant([[59, 0, 32, 12, 36, 0, 12, 32, 49, 36, 60]]),\n    lambda: tf.identity(x))```\n\ngave an improvement of +0.005 across the board (local eval, private and public LB for all models). Which is as much as the whole beam search ...\n\n## What worked and didn't\n\n- Beam search gave a decent +0.005 improvement\n- Replacing the model's prediction on corrupt data samples with a constant default prediction gave +0.005\n- Deeper models worked better than wider ones\n- MWER-training gave a small improvement (+0.001 - +0.002) - however, this was with beam search k=5, with greedy decoding gains were larger (+0.005 in local eval)\n- Replacing a single input token with a random one was a decent augmentation\n- CTC didn't help as an auxiliary loss\n- masking decoder input did not help (when random token replacement was used)\n- z-coordinates did not help\n\nAs always: really looking forward to reading everyone's solutions. Let me know if there are any questions. Code coming _soon_.",
    "2412137": "Congratulations. I observe that sometime Data augmentation can backfire as two finger spellings may be confusing to the model. \nFor example:  flipping 'C'  180 degrees may make it a 'G'. There are many such instances.",
    "2411931": ">Replacing a single input token with a random one was a decent augmentation\n \nReplacing a single input token with a random one may appear to be a method that disrupts the training data without a clear rationale. Why is this approach effective as an augmentation technique?",
    "2569238": "May I know if code is available? if so could you please share the link? Thank you in advance",
    "2420782": "Hello! Thank you for your write-up!  Would you be able to share what stddev you used for your gaussian noise?",
    "2413030": "Superb effort!",
    "2426075": ""
  }
}