{
  "id": 434485,
  "title": "[1st place solution] Improved Squeezeformer + TransformerDecoder + Clever augmentations",
  "url": "/competitions/asl-fingerspelling/discussion/434485",
  "author_name": "Dieter",
  "post_date": "2023-08-25T11:06:20.410000",
  "votes": 242,
  "comment_count": 93,
  "views": 0,
  "content": "<p>Thanks to kaggle and everyone involved for hosting such an interesting competition. It was a great extension to the isolated sign language classification and it was very interesting to see how much of speech-to-text research could also be applied to sign language fingerspelling. As always it was a great teaming experience with <a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a> </p>\n<h2>TLDR</h2>\n<p>Our solution is based on a single encoder-decoder architecture. The encoder is a significantly improved version of Squeezeformer, where the feature extraction was adapted to handle mediapipe landmarks instead of speech signals. The decoder is a simple 2-layer transformer. We additionally predicted a confidence score to identify corrupted examples which can be useful for post-processing. We also introduced efficient and creative augmentations to regularize the model, where the most important ones were CutMix, FingerDropout and TimeStretch, DecoderInput Masking. We used pytorch for developing and training our models and then manually translated model architecture and ported weights to tensorflow from which we exported to tf-lite.</p>\n<h2>Cross validation</h2>\n<p>We split the training data into 4 folds by signer. In the beginning we had nearly perfect correlation between CV and public LB with this approach. With higher scores improvements on CV reflected a bit less on LB, mostly due to the fact that the LB score was always decently higher and hence saturated earlier. Most of the time we only trained and tracked the score of fold0 and not all folds. </p>\n<h2>Data preprocessing</h2>\n<p>In total 130 key points were used. These consisted of 21 key points from each hand, 6 pose key points from each arm, and the remaining 76 from the face (lips, nose, eyes). Locally the 130 key points were cached to .npy files for fast data loading. <br>\nPrior to data augmentations, the data was normalized with std/mean and nans were zero filled. </p>\n<h2>Augmentations</h2>\n<p>Augmentations were essential to prevent overfitting, generalize to new signers and enable deep models. We used augmentations which were popular in the first ASL competition but also came up with a lot of new creative augmentations and some have proven to be very effective. </p>\n<ul>\n<li>Resizing along time axis.</li>\n<li>Shift the sequence along the time axis. </li>\n<li>Windowed resizing along time axis (similar to warping).</li>\n<li>Left-right flip of keypoints. </li>\n<li>Cutmix of samples timewise - draw a random percentage between [0,1] and cut 2 sequences and related phrases at that percentage and mix. Mixing only within same signer was best</li>\n<li>Spatial affine - scale, shear, shift and rotate. </li>\n<li>Drop/Zero-fill between 2 to 6 different fingers over 2 to 3 time windows. </li>\n<li>Drop/Zero-fill either all face landmarks or all pose landmarks. </li>\n<li>In rare cases (~5% of samples) drop/zero-fill all hand landmarks. </li>\n<li>Temporal masking (zero-fill) in windows of different sizes or counts. </li>\n<li>Spatial masking </li>\n</ul>\n<p>Most augmentations were applied to 50% of the samples, except for resizing and spatial affine which were applied to ~80% of samples.  </p>\n<p>After augmentation, samples with more than 384 frames were resized along time axis, with channel-wise linear interpolation. Samples of less than 384 were padded to 384 for training only. tf-lite ran on variable length samples. <br>\nNo frames were dropped in preprocessing. </p>\n<h2>Model</h2>\n<p>In general, we observed that a deeper model gives significant gains (if we are able to prevent overfitting). As a consequence not only regularization techniques like augmentations are essential, but also every improvement in computational efficiency creates space to use deeper models and hence is equally important as the model architecture itself.</p>\n<p>Our model consists of 3 parts, Feature Extraction, Encoder, Decoder which are shown in the image below</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2Fdc7c0fc365f54fb4d35f03da1748ed3b%2FScreenshot%202023-08-25%20at%2012.47.53.png?generation=1692960500706760&amp;alt=media\" alt=\"\"></p>\n<p>We interpret the data like a 3 channel image, where width is defined by the number of frames, height is given by the number of the selected 130 landmarks and channels are given by raw xyz coordinates. <br>\nThe feature extraction is based on a 2D convolution followed by batchnorm and a linear layer on the flattened features to extract features per frame. We have 5 of this feature extraction modules, one for all landmarks at once and one per landmark type (left_hand, right_hand, face, pose). The all-landmark module outputs 208 dim vector/ frame. The other 4 output 52 dim vectors each which are then concatenated to have also 208 dims. We had those two 208-dim vectors per frame and get a (batch_size x 384 x 208) input for our encoder where 384 is the maximum sequence length we chose.</p>\n<p>The main component of our model is an encoder which was adapted from the Squeezeformer architecture. We did not use the actual “squeeze” idea, i.e. a temporal Unet, but used the general architecture of Squeezeformer Blocks which consist of a combination of MultiHeadSelfAttention (MHSA), Convolution and FeedForward modules. We made several improvements to this architecture:<br>\nASR conformers (and Squeezeformer) use relative positional encoding which allow the self-attention module to generalize better on different input lengths. Relative positional encoding is performance intensive, as well as using many parameters, as they are stored separately in each layer. Replacing this with Llama attention which uses rotary embeddings sped up training ~2X and tf-lite inference approx ~3X allowing larger models to be used. In addition, we cached the rotary embeddings once and fed them into each layer with the input data so they are not duplicated in each layer. This resulted in 20% less parameters in the model. We saw no benefit in using time reduction which was introduced with Squeezeformer. So in our model all layers had the same sequence length as the original input. As suggested in the Squeezeformer paper, the pre-Layer Norm from the Macaron structure is redundant and was replaced with a learnable scaling layer which scales and shifts the activations. </p>\n<p>For decoding we used a simple 2 layer transformer decoder which is similar to hugging faces <a href=\"https://github.com/huggingface/transformers/blob/main/src/transformers/models/speech_to_text/modeling_speech_to_text.py#L857\" target=\"_blank\">Speech2TextDecoder</a>, which outputs a sequence prediction. We then used cross entropy loss for training our model End2end. An extra cross entropy auxiliary loss of the reversed sequence was used. A causal decoder mainly uses encoder cross attentions for the sequence's beginning and previous characters and cross attentions for the end. To improve the model's accuracy, we use a separate causal decoder on the reversed sequence as an auxiliary loss, making the model rely more on the encoder cross attention for the label's end. For decoder inference early stopping and past key value caching was used which sped up inference significantly.  It should be noted that we found a transformer based Decoder superior to a CTC based decoding, even in a setting where computational efficiency matters a lot. </p>\n<p>Additionally to the decoder, we also added a single linear layer to take the features of the first token of the encoder output to predict a confidence score, which helps to identify garbage data and can be used in post-processing. As a target for this we used normalized levensthein distance clipped to [0,1] of OOF predictions of a decent previous model.</p>\n<p>We explored and trained all our models in pytorch, but whenever we deemed it good enough for a submission we translated each component manually to tensorflow and ported weights from our pytorch models. </p>\n<h2>Training procedure</h2>\n<p>Models were trained with a cosine learning rate schedule for 400 epochs with peak LR of 0.0045, weight decay of 0.08, 10 epochs warmup, mixed precision and an effective batch size of 512 samples. Dropout of 0.1 was used in the transformer encoder/decoder layers. It was important to train with mixed precision in order to leverage fp16 inference without performance drop. Training with fp32 and using tf-lite fp16 inference causes a drop of ~0.01 in CV vs LB. </p>\n<p>tf-lite inference used a single sample without padding, while model training was performed with time padded mini-batches. To avoid the model learning with pads, and ensure optimal inference runtime, the feature extractor and Macaron structure encoder layers were masked time wise during training. This needed to be manually implemented in pytorch on each layer. This took some effort, but paid off by significantly speeding up inference. As <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406684\" target=\"_blank\">1st place team</a> in the previous ISL competition explained, this is much easier to do in tensorflow as keras has an off the shelf masking layer. </p>\n<p>Model training and tf-lite inference ran in fp16, so tf-lite files consumed almost half the disk size of of fp32. This was important, as the 40MB size was our limitation in the end. The two final model seeds measured 39988kb. </p>\n<h2>Postprocessing</h2>\n<p>The main idea of our postprocessing is to replace poor predictions with a dummy phrase which has a small levensthein distance to the train/ test data. <a href=\"https://www.kaggle.com/anokas\" target=\"_blank\">@anokas</a> showed in his <a href=\"https://www.kaggle.com/code/anokas/static-greedy-baseline-0-157-lb\" target=\"_blank\">notebook</a> why '2 a-e -aroe' is a good candidate for that. Most of the poor prediction resulted from corrupted input data often only a few frames long. We used a confidence score predicted by our model as basis. Whenever the confidence score is below 0.15 or the sequence is shorter than 15 frames we replace the prediction with '2 a-e -aroe'.</p>\n<h2>Supplemental Data</h2>\n<p>We only marginally profited from using the supplemental data. We think the main reason is that although there are 50k samples in this supplemental data there are only 500 unique phrases, and hence the model rather learns to classify then to actually decode character-by-character. We tried a lot of approaches but only the following one gave a small boost (0.838 -&gt; 0.839): <br>\nFirst we group the supplemental data by phrase, which only leaves us with 500 groups. In each epoch of training we add one sample per group to the training dataset for our model. That means in each epoch we use 50k samples of the training data and only 500 samples of the supplemental data. </p>\n<h2>Ensembling</h2>\n<p>Our final submission is a 2-seed ensemble of our model trained on the complete training data (fullfit). We average resulting logits in each decoding step for ensembling.</p>\n<h2>What did not help</h2>\n<ul>\n<li>Fully using supplemental data</li>\n<li>Using edit distance as loss (tried different approaches)</li>\n<li>CTC loss (even as an auxiliary loss it hurt score)</li>\n<li>Label smoothing</li>\n<li>AWP - kept getting nans with FP16</li>\n<li>TTA (flip/stretch)</li>\n<li>Mixup of hidden layers &amp; Specaugment++</li>\n<li>Beam search decoding (too costly)</li>\n</ul>\n<h2>Ablation study (roughly)</h2>\n<h4>Augmentations</h4>\n<ul>\n<li>Cutmix +0.005</li>\n<li>FingerDropout +0.005</li>\n<li>Face/PoseDropout +0.005</li>\n<li>masking decoder inputs +0.003</li>\n</ul>\n<h4>Model improvements</h4>\n<ul>\n<li>CNN Feature extraction +0.005</li>\n<li>2-branch Feature extraction with indiv norm +0.003</li>\n<li>Squeezeformer over 1stplace Net of 1 round +0.005</li>\n<li>Decoder over CTC +0.003</li>\n<li>Confidence over simple rules for post-processing +0.002</li>\n</ul>\n<h4>Efficiency Improvements:</h4>\n<ul>\n<li>Deeper model due to fp16 +0.003</li>\n<li>Deeper model due to llama attention +0.003</li>\n<li>Deeper model due to masking/ variable sequence len +0.005</li>\n<li>Deeper model due to caching/ early stopping in decoder +0.005</li>\n</ul>\n<h4>Postprocessing</h4>\n<ul>\n<li>Replace bad predictions with dummy phrase +0.006</li>\n</ul>\n<h2>Used tools/ repos</h2>\n<ul>\n<li>Pytorch/Tensorflow/Tf-lite (no onnx this time)</li>\n<li>Huggingface</li>\n<li>Albumentations (adapted their framework for using things like OneOf or Compose, but wrote our own augmentation implementations)</li>\n<li>Neptune.ai was our MLOps stack to track compare and share models. Below are example training runs using different model parameters and hardware (a100 card vs kaggle kernel). The data loading in the kaggle kernel below was slow and could probably be sped up with some work.  </li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2F4693e3379ea6a1631bb3cb7b9044d42a%2FScreenshot%202023-08-25%20at%2012.58.07.png?generation=1692961107306495&amp;alt=media\" alt=\"\"></p>\n<h2>Code &amp; model weights</h2>\n<p><a href=\"https://github.com/ChristofHenkel/kaggle-asl-fingerspelling-1st-place-solution\" target=\"_blank\">https://github.com/ChristofHenkel/kaggle-asl-fingerspelling-1st-place-solution</a></p>\n<h2>Paper</h2>\n<p>tbd.  Due to the novelty of our approach we are thinking about summarizing it in a paper.</p>\n<p><strong>Thank you for reading, questions welcome</strong></p>",
  "messages": [
    {
      "id": 2407974,
      "postDate": "2023-08-25T11:06:20.410Z",
      "content": "<p>Thanks to kaggle and everyone involved for hosting such an interesting competition. It was a great extension to the isolated sign language classification and it was very interesting to see how much of speech-to-text research could also be applied to sign language fingerspelling. As always it was a great teaming experience with <a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a> </p>\n<h2>TLDR</h2>\n<p>Our solution is based on a single encoder-decoder architecture. The encoder is a significantly improved version of Squeezeformer, where the feature extraction was adapted to handle mediapipe landmarks instead of speech signals. The decoder is a simple 2-layer transformer. We additionally predicted a confidence score to identify corrupted examples which can be useful for post-processing. We also introduced efficient and creative augmentations to regularize the model, where the most important ones were CutMix, FingerDropout and TimeStretch, DecoderInput Masking. We used pytorch for developing and training our models and then manually translated model architecture and ported weights to tensorflow from which we exported to tf-lite.</p>\n<h2>Cross validation</h2>\n<p>We split the training data into 4 folds by signer. In the beginning we had nearly perfect correlation between CV and public LB with this approach. With higher scores improvements on CV reflected a bit less on LB, mostly due to the fact that the LB score was always decently higher and hence saturated earlier. Most of the time we only trained and tracked the score of fold0 and not all folds. </p>\n<h2>Data preprocessing</h2>\n<p>In total 130 key points were used. These consisted of 21 key points from each hand, 6 pose key points from each arm, and the remaining 76 from the face (lips, nose, eyes). Locally the 130 key points were cached to .npy files for fast data loading. <br>\nPrior to data augmentations, the data was normalized with std/mean and nans were zero filled. </p>\n<h2>Augmentations</h2>\n<p>Augmentations were essential to prevent overfitting, generalize to new signers and enable deep models. We used augmentations which were popular in the first ASL competition but also came up with a lot of new creative augmentations and some have proven to be very effective. </p>\n<ul>\n<li>Resizing along time axis.</li>\n<li>Shift the sequence along the time axis. </li>\n<li>Windowed resizing along time axis (similar to warping).</li>\n<li>Left-right flip of keypoints. </li>\n<li>Cutmix of samples timewise - draw a random percentage between [0,1] and cut 2 sequences and related phrases at that percentage and mix. Mixing only within same signer was best</li>\n<li>Spatial affine - scale, shear, shift and rotate. </li>\n<li>Drop/Zero-fill between 2 to 6 different fingers over 2 to 3 time windows. </li>\n<li>Drop/Zero-fill either all face landmarks or all pose landmarks. </li>\n<li>In rare cases (~5% of samples) drop/zero-fill all hand landmarks. </li>\n<li>Temporal masking (zero-fill) in windows of different sizes or counts. </li>\n<li>Spatial masking </li>\n</ul>\n<p>Most augmentations were applied to 50% of the samples, except for resizing and spatial affine which were applied to ~80% of samples.  </p>\n<p>After augmentation, samples with more than 384 frames were resized along time axis, with channel-wise linear interpolation. Samples of less than 384 were padded to 384 for training only. tf-lite ran on variable length samples. <br>\nNo frames were dropped in preprocessing. </p>\n<h2>Model</h2>\n<p>In general, we observed that a deeper model gives significant gains (if we are able to prevent overfitting). As a consequence not only regularization techniques like augmentations are essential, but also every improvement in computational efficiency creates space to use deeper models and hence is equally important as the model architecture itself.</p>\n<p>Our model consists of 3 parts, Feature Extraction, Encoder, Decoder which are shown in the image below</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2Fdc7c0fc365f54fb4d35f03da1748ed3b%2FScreenshot%202023-08-25%20at%2012.47.53.png?generation=1692960500706760&amp;alt=media\" alt=\"\"></p>\n<p>We interpret the data like a 3 channel image, where width is defined by the number of frames, height is given by the number of the selected 130 landmarks and channels are given by raw xyz coordinates. <br>\nThe feature extraction is based on a 2D convolution followed by batchnorm and a linear layer on the flattened features to extract features per frame. We have 5 of this feature extraction modules, one for all landmarks at once and one per landmark type (left_hand, right_hand, face, pose). The all-landmark module outputs 208 dim vector/ frame. The other 4 output 52 dim vectors each which are then concatenated to have also 208 dims. We had those two 208-dim vectors per frame and get a (batch_size x 384 x 208) input for our encoder where 384 is the maximum sequence length we chose.</p>\n<p>The main component of our model is an encoder which was adapted from the Squeezeformer architecture. We did not use the actual “squeeze” idea, i.e. a temporal Unet, but used the general architecture of Squeezeformer Blocks which consist of a combination of MultiHeadSelfAttention (MHSA), Convolution and FeedForward modules. We made several improvements to this architecture:<br>\nASR conformers (and Squeezeformer) use relative positional encoding which allow the self-attention module to generalize better on different input lengths. Relative positional encoding is performance intensive, as well as using many parameters, as they are stored separately in each layer. Replacing this with Llama attention which uses rotary embeddings sped up training ~2X and tf-lite inference approx ~3X allowing larger models to be used. In addition, we cached the rotary embeddings once and fed them into each layer with the input data so they are not duplicated in each layer. This resulted in 20% less parameters in the model. We saw no benefit in using time reduction which was introduced with Squeezeformer. So in our model all layers had the same sequence length as the original input. As suggested in the Squeezeformer paper, the pre-Layer Norm from the Macaron structure is redundant and was replaced with a learnable scaling layer which scales and shifts the activations. </p>\n<p>For decoding we used a simple 2 layer transformer decoder which is similar to hugging faces <a href=\"https://github.com/huggingface/transformers/blob/main/src/transformers/models/speech_to_text/modeling_speech_to_text.py#L857\" target=\"_blank\">Speech2TextDecoder</a>, which outputs a sequence prediction. We then used cross entropy loss for training our model End2end. An extra cross entropy auxiliary loss of the reversed sequence was used. A causal decoder mainly uses encoder cross attentions for the sequence's beginning and previous characters and cross attentions for the end. To improve the model's accuracy, we use a separate causal decoder on the reversed sequence as an auxiliary loss, making the model rely more on the encoder cross attention for the label's end. For decoder inference early stopping and past key value caching was used which sped up inference significantly.  It should be noted that we found a transformer based Decoder superior to a CTC based decoding, even in a setting where computational efficiency matters a lot. </p>\n<p>Additionally to the decoder, we also added a single linear layer to take the features of the first token of the encoder output to predict a confidence score, which helps to identify garbage data and can be used in post-processing. As a target for this we used normalized levensthein distance clipped to [0,1] of OOF predictions of a decent previous model.</p>\n<p>We explored and trained all our models in pytorch, but whenever we deemed it good enough for a submission we translated each component manually to tensorflow and ported weights from our pytorch models. </p>\n<h2>Training procedure</h2>\n<p>Models were trained with a cosine learning rate schedule for 400 epochs with peak LR of 0.0045, weight decay of 0.08, 10 epochs warmup, mixed precision and an effective batch size of 512 samples. Dropout of 0.1 was used in the transformer encoder/decoder layers. It was important to train with mixed precision in order to leverage fp16 inference without performance drop. Training with fp32 and using tf-lite fp16 inference causes a drop of ~0.01 in CV vs LB. </p>\n<p>tf-lite inference used a single sample without padding, while model training was performed with time padded mini-batches. To avoid the model learning with pads, and ensure optimal inference runtime, the feature extractor and Macaron structure encoder layers were masked time wise during training. This needed to be manually implemented in pytorch on each layer. This took some effort, but paid off by significantly speeding up inference. As <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406684\" target=\"_blank\">1st place team</a> in the previous ISL competition explained, this is much easier to do in tensorflow as keras has an off the shelf masking layer. </p>\n<p>Model training and tf-lite inference ran in fp16, so tf-lite files consumed almost half the disk size of of fp32. This was important, as the 40MB size was our limitation in the end. The two final model seeds measured 39988kb. </p>\n<h2>Postprocessing</h2>\n<p>The main idea of our postprocessing is to replace poor predictions with a dummy phrase which has a small levensthein distance to the train/ test data. <a href=\"https://www.kaggle.com/anokas\" target=\"_blank\">@anokas</a> showed in his <a href=\"https://www.kaggle.com/code/anokas/static-greedy-baseline-0-157-lb\" target=\"_blank\">notebook</a> why '2 a-e -aroe' is a good candidate for that. Most of the poor prediction resulted from corrupted input data often only a few frames long. We used a confidence score predicted by our model as basis. Whenever the confidence score is below 0.15 or the sequence is shorter than 15 frames we replace the prediction with '2 a-e -aroe'.</p>\n<h2>Supplemental Data</h2>\n<p>We only marginally profited from using the supplemental data. We think the main reason is that although there are 50k samples in this supplemental data there are only 500 unique phrases, and hence the model rather learns to classify then to actually decode character-by-character. We tried a lot of approaches but only the following one gave a small boost (0.838 -&gt; 0.839): <br>\nFirst we group the supplemental data by phrase, which only leaves us with 500 groups. In each epoch of training we add one sample per group to the training dataset for our model. That means in each epoch we use 50k samples of the training data and only 500 samples of the supplemental data. </p>\n<h2>Ensembling</h2>\n<p>Our final submission is a 2-seed ensemble of our model trained on the complete training data (fullfit). We average resulting logits in each decoding step for ensembling.</p>\n<h2>What did not help</h2>\n<ul>\n<li>Fully using supplemental data</li>\n<li>Using edit distance as loss (tried different approaches)</li>\n<li>CTC loss (even as an auxiliary loss it hurt score)</li>\n<li>Label smoothing</li>\n<li>AWP - kept getting nans with FP16</li>\n<li>TTA (flip/stretch)</li>\n<li>Mixup of hidden layers &amp; Specaugment++</li>\n<li>Beam search decoding (too costly)</li>\n</ul>\n<h2>Ablation study (roughly)</h2>\n<h4>Augmentations</h4>\n<ul>\n<li>Cutmix +0.005</li>\n<li>FingerDropout +0.005</li>\n<li>Face/PoseDropout +0.005</li>\n<li>masking decoder inputs +0.003</li>\n</ul>\n<h4>Model improvements</h4>\n<ul>\n<li>CNN Feature extraction +0.005</li>\n<li>2-branch Feature extraction with indiv norm +0.003</li>\n<li>Squeezeformer over 1stplace Net of 1 round +0.005</li>\n<li>Decoder over CTC +0.003</li>\n<li>Confidence over simple rules for post-processing +0.002</li>\n</ul>\n<h4>Efficiency Improvements:</h4>\n<ul>\n<li>Deeper model due to fp16 +0.003</li>\n<li>Deeper model due to llama attention +0.003</li>\n<li>Deeper model due to masking/ variable sequence len +0.005</li>\n<li>Deeper model due to caching/ early stopping in decoder +0.005</li>\n</ul>\n<h4>Postprocessing</h4>\n<ul>\n<li>Replace bad predictions with dummy phrase +0.006</li>\n</ul>\n<h2>Used tools/ repos</h2>\n<ul>\n<li>Pytorch/Tensorflow/Tf-lite (no onnx this time)</li>\n<li>Huggingface</li>\n<li>Albumentations (adapted their framework for using things like OneOf or Compose, but wrote our own augmentation implementations)</li>\n<li>Neptune.ai was our MLOps stack to track compare and share models. Below are example training runs using different model parameters and hardware (a100 card vs kaggle kernel). The data loading in the kaggle kernel below was slow and could probably be sped up with some work.  </li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2F4693e3379ea6a1631bb3cb7b9044d42a%2FScreenshot%202023-08-25%20at%2012.58.07.png?generation=1692961107306495&amp;alt=media\" alt=\"\"></p>\n<h2>Code &amp; model weights</h2>\n<p><a href=\"https://github.com/ChristofHenkel/kaggle-asl-fingerspelling-1st-place-solution\" target=\"_blank\">https://github.com/ChristofHenkel/kaggle-asl-fingerspelling-1st-place-solution</a></p>\n<h2>Paper</h2>\n<p>tbd.  Due to the novelty of our approach we are thinking about summarizing it in a paper.</p>\n<p><strong>Thank you for reading, questions welcome</strong></p>",
      "rawMarkdown": "Thanks to kaggle and everyone involved for hosting such an interesting competition. It was a great extension to the isolated sign language classification and it was very interesting to see how much of speech-to-text research could also be applied to sign language fingerspelling. As always it was a great teaming experience with @darraghdog \n\n## TLDR\nOur solution is based on a single encoder-decoder architecture. The encoder is a significantly improved version of Squeezeformer, where the feature extraction was adapted to handle mediapipe landmarks instead of speech signals. The decoder is a simple 2-layer transformer. We additionally predicted a confidence score to identify corrupted examples which can be useful for post-processing. We also introduced efficient and creative augmentations to regularize the model, where the most important ones were CutMix, FingerDropout and TimeStretch, DecoderInput Masking. We used pytorch for developing and training our models and then manually translated model architecture and ported weights to tensorflow from which we exported to tf-lite.\n\n## Cross validation\nWe split the training data into 4 folds by signer. In the beginning we had nearly perfect correlation between CV and public LB with this approach. With higher scores improvements on CV reflected a bit less on LB, mostly due to the fact that the LB score was always decently higher and hence saturated earlier. Most of the time we only trained and tracked the score of fold0 and not all folds. \n\n## Data preprocessing\nIn total 130 key points were used. These consisted of 21 key points from each hand, 6 pose key points from each arm, and the remaining 76 from the face (lips, nose, eyes). Locally the 130 key points were cached to .npy files for fast data loading. \nPrior to data augmentations, the data was normalized with std/mean and nans were zero filled. \n\n## Augmentations\nAugmentations were essential to prevent overfitting, generalize to new signers and enable deep models. We used augmentations which were popular in the first ASL competition but also came up with a lot of new creative augmentations and some have proven to be very effective. \n\n- Resizing along time axis.\n- Shift the sequence along the time axis. \n- Windowed resizing along time axis (similar to warping).\n- Left-right flip of keypoints. \n- Cutmix of samples timewise - draw a random percentage between [0,1] and cut 2 sequences and related phrases at that percentage and mix. Mixing only within same signer was best\n- Spatial affine - scale, shear, shift and rotate. \n- Drop/Zero-fill between 2 to 6 different fingers over 2 to 3 time windows. \n- Drop/Zero-fill either all face landmarks or all pose landmarks. \n- In rare cases (~5% of samples) drop/zero-fill all hand landmarks. \n- Temporal masking (zero-fill) in windows of different sizes or counts. \n- Spatial masking \n\nMost augmentations were applied to 50% of the samples, except for resizing and spatial affine which were applied to ~80% of samples.  \n\nAfter augmentation, samples with more than 384 frames were resized along time axis, with channel-wise linear interpolation. Samples of less than 384 were padded to 384 for training only. tf-lite ran on variable length samples. \nNo frames were dropped in preprocessing. \n\n## Model\nIn general, we observed that a deeper model gives significant gains (if we are able to prevent overfitting). As a consequence not only regularization techniques like augmentations are essential, but also every improvement in computational efficiency creates space to use deeper models and hence is equally important as the model architecture itself.\n\nOur model consists of 3 parts, Feature Extraction, Encoder, Decoder which are shown in the image below\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2Fdc7c0fc365f54fb4d35f03da1748ed3b%2FScreenshot%202023-08-25%20at%2012.47.53.png?generation=1692960500706760&alt=media)\n\nWe interpret the data like a 3 channel image, where width is defined by the number of frames, height is given by the number of the selected 130 landmarks and channels are given by raw xyz coordinates. \nThe feature extraction is based on a 2D convolution followed by batchnorm and a linear layer on the flattened features to extract features per frame. We have 5 of this feature extraction modules, one for all landmarks at once and one per landmark type (left_hand, right_hand, face, pose). The all-landmark module outputs 208 dim vector/ frame. The other 4 output 52 dim vectors each which are then concatenated to have also 208 dims. We had those two 208-dim vectors per frame and get a (batch_size x 384 x 208) input for our encoder where 384 is the maximum sequence length we chose.\n\nThe main component of our model is an encoder which was adapted from the Squeezeformer architecture. We did not use the actual “squeeze” idea, i.e. a temporal Unet, but used the general architecture of Squeezeformer Blocks which consist of a combination of MultiHeadSelfAttention (MHSA), Convolution and FeedForward modules. We made several improvements to this architecture:\nASR conformers (and Squeezeformer) use relative positional encoding which allow the self-attention module to generalize better on different input lengths. Relative positional encoding is performance intensive, as well as using many parameters, as they are stored separately in each layer. Replacing this with Llama attention which uses rotary embeddings sped up training ~2X and tf-lite inference approx ~3X allowing larger models to be used. In addition, we cached the rotary embeddings once and fed them into each layer with the input data so they are not duplicated in each layer. This resulted in 20% less parameters in the model. We saw no benefit in using time reduction which was introduced with Squeezeformer. So in our model all layers had the same sequence length as the original input. As suggested in the Squeezeformer paper, the pre-Layer Norm from the Macaron structure is redundant and was replaced with a learnable scaling layer which scales and shifts the activations. \n\nFor decoding we used a simple 2 layer transformer decoder which is similar to hugging faces [Speech2TextDecoder](https://github.com/huggingface/transformers/blob/main/src/transformers/models/speech_to_text/modeling_speech_to_text.py#L857), which outputs a sequence prediction. We then used cross entropy loss for training our model End2end. An extra cross entropy auxiliary loss of the reversed sequence was used. A causal decoder mainly uses encoder cross attentions for the sequence's beginning and previous characters and cross attentions for the end. To improve the model's accuracy, we use a separate causal decoder on the reversed sequence as an auxiliary loss, making the model rely more on the encoder cross attention for the label's end. For decoder inference early stopping and past key value caching was used which sped up inference significantly.  It should be noted that we found a transformer based Decoder superior to a CTC based decoding, even in a setting where computational efficiency matters a lot. \n\nAdditionally to the decoder, we also added a single linear layer to take the features of the first token of the encoder output to predict a confidence score, which helps to identify garbage data and can be used in post-processing. As a target for this we used normalized levensthein distance clipped to [0,1] of OOF predictions of a decent previous model.\n\nWe explored and trained all our models in pytorch, but whenever we deemed it good enough for a submission we translated each component manually to tensorflow and ported weights from our pytorch models. \n\n## Training procedure\nModels were trained with a cosine learning rate schedule for 400 epochs with peak LR of 0.0045, weight decay of 0.08, 10 epochs warmup, mixed precision and an effective batch size of 512 samples. Dropout of 0.1 was used in the transformer encoder/decoder layers. It was important to train with mixed precision in order to leverage fp16 inference without performance drop. Training with fp32 and using tf-lite fp16 inference causes a drop of ~0.01 in CV vs LB. \n\ntf-lite inference used a single sample without padding, while model training was performed with time padded mini-batches. To avoid the model learning with pads, and ensure optimal inference runtime, the feature extractor and Macaron structure encoder layers were masked time wise during training. This needed to be manually implemented in pytorch on each layer. This took some effort, but paid off by significantly speeding up inference. As [1st place team](https://www.kaggle.com/competitions/asl-signs/discussion/406684) in the previous ISL competition explained, this is much easier to do in tensorflow as keras has an off the shelf masking layer. \n\nModel training and tf-lite inference ran in fp16, so tf-lite files consumed almost half the disk size of of fp32. This was important, as the 40MB size was our limitation in the end. The two final model seeds measured 39988kb. \n\n## Postprocessing\nThe main idea of our postprocessing is to replace poor predictions with a dummy phrase which has a small levensthein distance to the train/ test data. @anokas showed in his [notebook](https://www.kaggle.com/code/anokas/static-greedy-baseline-0-157-lb) why '2 a-e -aroe' is a good candidate for that. Most of the poor prediction resulted from corrupted input data often only a few frames long. We used a confidence score predicted by our model as basis. Whenever the confidence score is below 0.15 or the sequence is shorter than 15 frames we replace the prediction with '2 a-e -aroe'.\n\n## Supplemental Data\nWe only marginally profited from using the supplemental data. We think the main reason is that although there are 50k samples in this supplemental data there are only 500 unique phrases, and hence the model rather learns to classify then to actually decode character-by-character. We tried a lot of approaches but only the following one gave a small boost (0.838 -> 0.839): \nFirst we group the supplemental data by phrase, which only leaves us with 500 groups. In each epoch of training we add one sample per group to the training dataset for our model. That means in each epoch we use 50k samples of the training data and only 500 samples of the supplemental data. \n\n## Ensembling\n\nOur final submission is a 2-seed ensemble of our model trained on the complete training data (fullfit). We average resulting logits in each decoding step for ensembling.\n\n## What did not help\n\n- Fully using supplemental data\n- Using edit distance as loss (tried different approaches)\n- CTC loss (even as an auxiliary loss it hurt score)\n- Label smoothing\n- AWP - kept getting nans with FP16\n- TTA (flip/stretch)\n- Mixup of hidden layers & Specaugment++\n- Beam search decoding (too costly)\n\n## Ablation study (roughly)\n\n#### Augmentations\n- Cutmix +0.005\n- FingerDropout +0.005\n- Face/PoseDropout +0.005\n- masking decoder inputs +0.003\n\n#### Model improvements\n- CNN Feature extraction +0.005\n- 2-branch Feature extraction with indiv norm +0.003\n- Squeezeformer over 1stplace Net of 1 round +0.005\n- Decoder over CTC +0.003\n- Confidence over simple rules for post-processing +0.002\n\n\n#### Efficiency Improvements:\n- Deeper model due to fp16 +0.003\n- Deeper model due to llama attention +0.003\n- Deeper model due to masking/ variable sequence len +0.005\n- Deeper model due to caching/ early stopping in decoder +0.005\n\n#### Postprocessing\n- Replace bad predictions with dummy phrase +0.006\n\n\n## Used tools/ repos\n- Pytorch/Tensorflow/Tf-lite (no onnx this time)\n- Huggingface\n- Albumentations (adapted their framework for using things like OneOf or Compose, but wrote our own augmentation implementations)\n- Neptune.ai was our MLOps stack to track compare and share models. Below are example training runs using different model parameters and hardware (a100 card vs kaggle kernel). The data loading in the kaggle kernel below was slow and could probably be sped up with some work.  \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2F4693e3379ea6a1631bb3cb7b9044d42a%2FScreenshot%202023-08-25%20at%2012.58.07.png?generation=1692961107306495&alt=media)\n\n## Code & model weights\n\nhttps://github.com/ChristofHenkel/kaggle-asl-fingerspelling-1st-place-solution\n\n## Paper\ntbd.  Due to the novelty of our approach we are thinking about summarizing it in a paper.\n\n**Thank you for reading, questions welcome**",
      "votes": 242
    },
    {
      "id": 2410650,
      "postDate": "2023-08-27T06:23:03.387Z",
      "content": "<p>Congratulations&nbsp;@christofhenkel&nbsp;and&nbsp;@darraghdog </p>\n<p>Appreciate your effort! 🙏  </p>",
      "rawMarkdown": "Congratulations @christofhenkel and @darraghdog \n\nAppreciate your effort! 🙏  \n             ",
      "votes": 3
    },
    {
      "id": 2408132,
      "postDate": "2023-08-25T13:20:46.510Z",
      "content": "<p>Super congrats <a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a> and <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> for this huge win! Thanks for sharing such detailed solution.</p>",
      "rawMarkdown": "Super congrats @darraghdog and @christofhenkel for this huge win! Thanks for sharing such detailed solution.",
      "votes": 3
    },
    {
      "id": 2407993,
      "postDate": "2023-08-25T11:17:46.003Z",
      "content": "<blockquote>\n  <p>In general, we observed that a deeper model gives significant gains (if we are able to prevent overfitting).</p>\n</blockquote>\n<p>I agree. For me, 24L 256d was better than 9L 384d, even though they have similar sizes.</p>",
      "rawMarkdown": "> In general, we observed that a deeper model gives significant gains (if we are able to prevent overfitting).\n\nI agree. For me, 24L 256d was better than 9L 384d, even though they have similar sizes.",
      "votes": 3
    },
    {
      "id": 2408367,
      "postDate": "2023-08-25T15:38:25.200Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> and <a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a> <br>\nA beautiful and clean solution. I am sure you tried a ton of different ideas how to encode these features.</p>",
      "rawMarkdown": "Congratulations @christofhenkel and @darraghdog \nA beautiful and clean solution. I am sure you tried a ton of different ideas how to encode these features.",
      "votes": 4
    },
    {
      "id": 2408273,
      "postDate": "2023-08-25T14:36:39.903Z",
      "content": "<p>I'm going to post a little comment. It might get long. Hopefully it's helpful for someone.</p>\n<p>First, Wow. I just saw Greysnow's comment saying he feels like a child… coming in 12th. Here I am 386 by copying others' work… if he's a child, I am… Anyway, really great writeup and lots of great ideas. I also look forward to seeing all the code.</p>\n<p>Well, I entered this competition as something to do and to run a couple of experiments. I've gone through a decent amount of the general learning stuff on kaggle. I don't have a background in coding (have a PhD in something else, so you know, I can probably learn things), but I love seeing what people can do in this arena and keep exploring whether I can truly get into it. I feel like I should be somewhat embarrassed to say chatGPT was my guide, and that was a big part of my experiment-- how much could this newer tool help someone as clueless as me to do anything useful. On the one hand, amazing what it could help me accomplish, but also as I think most of you know, it doesn't always take the most elegant approach, and sometimes after a few iterations of changes and errors, I'd hunker down and really try on my own and figure out how to fix the problem myself faster. My main goals were just to see how much of this I could understand, and when I had an idea, could I figure out how to implement it, all with the hope that at the end of the contest when the top solutions are posted, I could really grasp it at the least (not sure I'm there, but I've made progress). As I've posted elsewhere, I was curious whether the winner is really elegant and like totally new solutions, and/or brute force with systematic hyperparameter tuning, taking it offline to run near-infinite variations on some other dedicated hardware, etc.  I'm not sure about the offline dedicated hardware part, but I think from this writeup, there's some of both.</p>\n<p>Well I do have a few general questions that I'd like to understand better. Some of the first stuff I did with other people's code was try to simplify down to only hand landmarks. I didn't and generally still don't see what face landmarks could be helpful for in a fingerspelling dataset. I suppose it's not that I can't imagine that perhaps on a pause as someone repeats a letter, that there might be an associated facial expression the model could pick up-- but it feels quite minimal to me, and I thought there would be more benefit to not having to look at so much extra information that is mostly useless. I did come to think that some of the pose landmarks are probably helpful with regard to motion letters like \"z\" and perhaps other things-- though it still seems like a good model should be able to get all the info it needs from just the hand. I also tried ignoring the z coordinates, and I tried flipping one of the hands, thinking that it made more sense for the model to learn a single set of signs and motions, where the opposite hand ultimately takes a flipped orientation. I was not super systematic with recording how a single change affected loss or score, but more often than not whatever change I attempted gave worse performance. Trying to get rid of lhand and rhand with an xy flip of one (while I was ignoring z) definitely hurt, which surprised me. </p>\n<p>I tried adding a displacement feature by calculating the, like Euclidean norm-change from frame to frame. I was hopeful that this might help the model with double letters and spaces, which it seemed to me from output was an area of trouble, maybe unsurprisingly. However, after much effort to finally successfully implement the addition of this feature and get it all to work together, again, it hurt model performance. Maybe that makes perfect sense to people who understand better how transformers work? As I understand it, it's never entirely clear whether adding a new feature will really help a model to focus on something, or will just be redundant information. I wondered, but did not try, whether trying to identify some like biggest changes, and creating new features based on that, might help-- like maybe calculating the distance from thumb to pinky as a new feature, just as potential example. Again, is that just likely to be redundant, or does the idea not work well with the transformer architecture?</p>\n<p>I tried separating out training and validation datasets based on the participant_id, thinking it might help if the model trained on one signer and their idiosyncrasies at a time-- well and I also tried batching by signer, I guess that's really the training on one signer at a time. Didn't help.</p>\n<p>I did what others have done to make little gifs out of the signing just so I could visualize it, and I also played around with adding to the dataset by recording new landmark data from YouTube videos (and even from myself, just to do it), but it wasn't easy for me to think of how I might systematically add to the dataset without being really manual about it. That's still a question that I'm curious about and haven't answered for myself: essentially what is \"enough\" data? If you take this winning model and you train on 20% of the data, how will that compare to using 80% of it- and if we kept increasing the quantity of data here, would we continue to see improvement? </p>\n<p>I also considered augmenting the data with little rotations, but I didn't get around to implementation. I'm intrigued by the augmentation of dropping fingers and such</p>\n<p>I used greysnow's notebook and code to get things working on TPU when I ran out of GPU time, but man it was painful. It seems every one of those efforts I made above added a layer of difficulty if I tried to run it on TPU. I can only imagine the use of PyTorch and translation efforts in this notebook were quite more work.</p>\n<p>So-- a general question. While it's amazing to watch the model work and create useful predictions, the 0.84 score means, I think, that on an average 11-12-character phrase, there's nearly 2 changes necessary to get to the right phrase. From looking at the output, I think it's actually the case that you get a decent amount of perfect translation, but there are also a number that are far off. Anyway, this is the best of the bunch here, it's definitely impressive, but is it objectively \"good\"? Is it only the starting point for developing a really useful application? I don't have the context to understand what threshold must be met for people to actually want to use it. This dataset includes garbage data, which is dealt with in this notebook and others, but aside from the truly garbage data, I would think if you took a few ASL-experienced individuals and had them watch the original signing, or perhaps even the mediapipe-processed stick figures, they would be able to achieve very nearly 100% accuracy. If yes, then I'd like to think we could train models to do just as well. And in a way there's a disconnect for me in what a human can infer primarily from watching for known symbols compared with the way these models work that are not in any way informed what an \"A\" or \"B\" looks like. I also explored the idea of feeding this information in some way, but I don't know that there's a good way to do that. A different question to ask is if there were not limitations made such that this can be a TFlite model or if we could throw more GPUs and much longer training time, etc., could we then pretty easily bump this accuracy up?</p>\n<p>Now that the competition is over, do people tend to tweak further using ideas from other notebooks? Basically I wonder if someone right now could take the best ideas from different notebooks, could we quickly get to 0.9? Or if this solution has already maxed out on what's possible given the constraints. </p>\n<p>Well, thank you for posting the writeup, congratulations on the good work, and I'm interested to learn and understand more.</p>",
      "rawMarkdown": "I'm going to post a little comment. It might get long. Hopefully it's helpful for someone.\n\nFirst, Wow. I just saw Greysnow's comment saying he feels like a child... coming in 12th. Here I am 386 by copying others' work... if he's a child, I am... Anyway, really great writeup and lots of great ideas. I also look forward to seeing all the code.\n\nWell, I entered this competition as something to do and to run a couple of experiments. I've gone through a decent amount of the general learning stuff on kaggle. I don't have a background in coding (have a PhD in something else, so you know, I can probably learn things), but I love seeing what people can do in this arena and keep exploring whether I can truly get into it. I feel like I should be somewhat embarrassed to say chatGPT was my guide, and that was a big part of my experiment-- how much could this newer tool help someone as clueless as me to do anything useful. On the one hand, amazing what it could help me accomplish, but also as I think most of you know, it doesn't always take the most elegant approach, and sometimes after a few iterations of changes and errors, I'd hunker down and really try on my own and figure out how to fix the problem myself faster. My main goals were just to see how much of this I could understand, and when I had an idea, could I figure out how to implement it, all with the hope that at the end of the contest when the top solutions are posted, I could really grasp it at the least (not sure I'm there, but I've made progress). As I've posted elsewhere, I was curious whether the winner is really elegant and like totally new solutions, and/or brute force with systematic hyperparameter tuning, taking it offline to run near-infinite variations on some other dedicated hardware, etc.  I'm not sure about the offline dedicated hardware part, but I think from this writeup, there's some of both.\n\nWell I do have a few general questions that I'd like to understand better. Some of the first stuff I did with other people's code was try to simplify down to only hand landmarks. I didn't and generally still don't see what face landmarks could be helpful for in a fingerspelling dataset. I suppose it's not that I can't imagine that perhaps on a pause as someone repeats a letter, that there might be an associated facial expression the model could pick up-- but it feels quite minimal to me, and I thought there would be more benefit to not having to look at so much extra information that is mostly useless. I did come to think that some of the pose landmarks are probably helpful with regard to motion letters like \"z\" and perhaps other things-- though it still seems like a good model should be able to get all the info it needs from just the hand. I also tried ignoring the z coordinates, and I tried flipping one of the hands, thinking that it made more sense for the model to learn a single set of signs and motions, where the opposite hand ultimately takes a flipped orientation. I was not super systematic with recording how a single change affected loss or score, but more often than not whatever change I attempted gave worse performance. Trying to get rid of lhand and rhand with an xy flip of one (while I was ignoring z) definitely hurt, which surprised me. \n\nI tried adding a displacement feature by calculating the, like Euclidean norm-change from frame to frame. I was hopeful that this might help the model with double letters and spaces, which it seemed to me from output was an area of trouble, maybe unsurprisingly. However, after much effort to finally successfully implement the addition of this feature and get it all to work together, again, it hurt model performance. Maybe that makes perfect sense to people who understand better how transformers work? As I understand it, it's never entirely clear whether adding a new feature will really help a model to focus on something, or will just be redundant information. I wondered, but did not try, whether trying to identify some like biggest changes, and creating new features based on that, might help-- like maybe calculating the distance from thumb to pinky as a new feature, just as potential example. Again, is that just likely to be redundant, or does the idea not work well with the transformer architecture?\n\nI tried separating out training and validation datasets based on the participant_id, thinking it might help if the model trained on one signer and their idiosyncrasies at a time-- well and I also tried batching by signer, I guess that's really the training on one signer at a time. Didn't help.\n\nI did what others have done to make little gifs out of the signing just so I could visualize it, and I also played around with adding to the dataset by recording new landmark data from YouTube videos (and even from myself, just to do it), but it wasn't easy for me to think of how I might systematically add to the dataset without being really manual about it. That's still a question that I'm curious about and haven't answered for myself: essentially what is \"enough\" data? If you take this winning model and you train on 20% of the data, how will that compare to using 80% of it- and if we kept increasing the quantity of data here, would we continue to see improvement? \n\nI also considered augmenting the data with little rotations, but I didn't get around to implementation. I'm intrigued by the augmentation of dropping fingers and such\n\nI used greysnow's notebook and code to get things working on TPU when I ran out of GPU time, but man it was painful. It seems every one of those efforts I made above added a layer of difficulty if I tried to run it on TPU. I can only imagine the use of PyTorch and translation efforts in this notebook were quite more work.\n\nSo-- a general question. While it's amazing to watch the model work and create useful predictions, the 0.84 score means, I think, that on an average 11-12-character phrase, there's nearly 2 changes necessary to get to the right phrase. From looking at the output, I think it's actually the case that you get a decent amount of perfect translation, but there are also a number that are far off. Anyway, this is the best of the bunch here, it's definitely impressive, but is it objectively \"good\"? Is it only the starting point for developing a really useful application? I don't have the context to understand what threshold must be met for people to actually want to use it. This dataset includes garbage data, which is dealt with in this notebook and others, but aside from the truly garbage data, I would think if you took a few ASL-experienced individuals and had them watch the original signing, or perhaps even the mediapipe-processed stick figures, they would be able to achieve very nearly 100% accuracy. If yes, then I'd like to think we could train models to do just as well. And in a way there's a disconnect for me in what a human can infer primarily from watching for known symbols compared with the way these models work that are not in any way informed what an \"A\" or \"B\" looks like. I also explored the idea of feeding this information in some way, but I don't know that there's a good way to do that. A different question to ask is if there were not limitations made such that this can be a TFlite model or if we could throw more GPUs and much longer training time, etc., could we then pretty easily bump this accuracy up?\n\nNow that the competition is over, do people tend to tweak further using ideas from other notebooks? Basically I wonder if someone right now could take the best ideas from different notebooks, could we quickly get to 0.9? Or if this solution has already maxed out on what's possible given the constraints. \n\nWell, thank you for posting the writeup, congratulations on the good work, and I'm interested to learn and understand more.",
      "votes": 4,
      "replies": [
        {
          "id": 2408337,
          "postDate": "2023-08-25T15:22:22.373Z",
          "content": "<p>Let me comment on a few of your point. Wont go into details, just give some of my thoughts</p>\n<blockquote>\n  <p>Well, I entered this competition as something to do and to run a couple of experiments. I've gone through a decent amount of the general learning stuff on kaggle. I don't have a background in coding (have a PhD in something else, so you know, I can probably learn things),</p>\n</blockquote>\n<p>nobody starts with ML and directly wins competitions. Having also no python nor ML knowledge I worked my ass off and got a place of 500/2000 in my first kaggle competition, which by the way was on speech recognition 6 years ago. But with perseverance and curiosity I continued learning competition by competition and build a skill or two since then.  </p>\n<p>Regarding your point of flipping hands and using face landmarks. Some signers tend to support their signing with lip movement, which can be supportive signal. Also, not all signs are single handed. E.g. <code>+</code> which is in nearly all phone numbers uses both hands</p>\n<blockquote>\n  <p>the 0.84 score means, I think, that on an average 11-12-character phrase, there's nearly 2 changes necessary to get to the right phrase. From looking at the output, I think it's actually the case that you get a decent amount of perfect translation, but there are also a number that are far off. </p>\n</blockquote>\n<p>You need to keep in mind that 10% of the data is corrupted and without signal coming from artefacts where people just tapped the record button, and its impossible to predict the target. But those cases are not relevant to the final usability. For the sequences that contain the actual phrases, our model has a very high accuracy. </p>",
          "rawMarkdown": "Let me comment on a few of your point. Wont go into details, just give some of my thoughts\n\n>Well, I entered this competition as something to do and to run a couple of experiments. I've gone through a decent amount of the general learning stuff on kaggle. I don't have a background in coding (have a PhD in something else, so you know, I can probably learn things),\n\nnobody starts with ML and directly wins competitions. Having also no python nor ML knowledge I worked my ass off and got a place of 500/2000 in my first kaggle competition, which by the way was on speech recognition 6 years ago. But with perseverance and curiosity I continued learning competition by competition and build a skill or two since then.  \n\nRegarding your point of flipping hands and using face landmarks. Some signers tend to support their signing with lip movement, which can be supportive signal. Also, not all signs are single handed. E.g. `+` which is in nearly all phone numbers uses both hands\n\n>the 0.84 score means, I think, that on an average 11-12-character phrase, there's nearly 2 changes necessary to get to the right phrase. From looking at the output, I think it's actually the case that you get a decent amount of perfect translation, but there are also a number that are far off. \n\nYou need to keep in mind that 10% of the data is corrupted and without signal coming from artefacts where people just tapped the record button, and its impossible to predict the target. But those cases are not relevant to the final usability. For the sequences that contain the actual phrases, our model has a very high accuracy. \n\n",
          "votes": 12,
          "replies": [
            {
              "id": 2408374,
              "postDate": "2023-08-25T15:43:56.267Z",
              "content": "<p>Wow, thank you for reading and commenting. I really like hearing that you learned on the job, so to speak, gives hope. (I certainly didn't think I'd win, but the prize money gives me cover for spending a lot of time on this and neglecting other things in life just the same). </p>\n<p>I didn't realize plus symbol used both hands. Oops. I probably looked at a few phrases that didn't have + signs and just saw the entire empty data from one hand. I recognize if this was not fingerspelling but more general ASL it would make total sense to use all of the data. I just thought we could simplify here. But just the + information makes quite clear why trying to flip hands or simplify to one ruined it. Thanks for the insight there.</p>\n<p>re: accuracy, Is it not true that when you look at a validation callback, you still see a number of phrases that your model just can't get right? I essentially thought that your code effectively removed the artifactual sequences from consideration, so at least for training and validation, this would be dealt with. If the test dataset also contains artifactual data, there's not much that can be done with it, and the model doesn't have the luxury of ignoring it, even though the landmark data in no way actually corresponds to the phrase. But then I'd expect you to be able to calculate your own Levenshtein score as much higher, since it ignores artifacts, and to see a big drop once it is used on test data. I guess at the end of the day I'd like to see the performance on purely cleaned up, like independently verified, data. Thanks again.</p>",
              "rawMarkdown": "Wow, thank you for reading and commenting. I really like hearing that you learned on the job, so to speak, gives hope. (I certainly didn't think I'd win, but the prize money gives me cover for spending a lot of time on this and neglecting other things in life just the same). \n\nI didn't realize plus symbol used both hands. Oops. I probably looked at a few phrases that didn't have + signs and just saw the entire empty data from one hand. I recognize if this was not fingerspelling but more general ASL it would make total sense to use all of the data. I just thought we could simplify here. But just the + information makes quite clear why trying to flip hands or simplify to one ruined it. Thanks for the insight there.\n\nre: accuracy, Is it not true that when you look at a validation callback, you still see a number of phrases that your model just can't get right? I essentially thought that your code effectively removed the artifactual sequences from consideration, so at least for training and validation, this would be dealt with. If the test dataset also contains artifactual data, there's not much that can be done with it, and the model doesn't have the luxury of ignoring it, even though the landmark data in no way actually corresponds to the phrase. But then I'd expect you to be able to calculate your own Levenshtein score as much higher, since it ignores artifacts, and to see a big drop once it is used on test data. I guess at the end of the day I'd like to see the performance on purely cleaned up, like independently verified, data. Thanks again.\n"
            },
            {
              "id": 2408446,
              "postDate": "2023-08-25T16:23:31.597Z",
              "content": "<p>I was not talking about raw validation score but rather about the usefulness of our model in the real world, as (at least thats how I understand) it will be directly deployed in an app to help parents learn fingerspelling and support them communicate with their kids. In the real world example users, know when the did not do a proper recording, so only the quality of the model with respect to correctly recorded sequences matters. And thats where our model is quite strong. I would be very interested in human level performance for the train/ test dataset, but I would assume our model is on the same level. Dont know if hosts evaluated this.</p>",
              "rawMarkdown": "I was not talking about raw validation score but rather about the usefulness of our model in the real world, as (at least thats how I understand) it will be directly deployed in an app to help parents learn fingerspelling and support them communicate with their kids. In the real world example users, know when the did not do a proper recording, so only the quality of the model with respect to correctly recorded sequences matters. And thats where our model is quite strong. I would be very interested in human level performance for the train/ test dataset, but I would assume our model is on the same level. Dont know if hosts evaluated this.",
              "votes": 3
            }
          ]
        }
      ]
    },
    {
      "id": 2408048,
      "postDate": "2023-08-25T12:03:31.203Z",
      "content": "<p>Wow, there is so much high-level work here to learn from. I eagerly await your code. Compared to your solution, I feel like a child 😅</p>",
      "rawMarkdown": "Wow, there is so much high-level work here to learn from. I eagerly await your code. Compared to your solution, I feel like a child 😅",
      "votes": 4
    },
    {
      "id": 2447454,
      "postDate": "2023-09-20T05:13:31.733Z",
      "content": "<p>Congratulations! Your solution is really enlightening to me! Especially the part related to Augmentations.</p>",
      "rawMarkdown": "Congratulations! Your solution is really enlightening to me! Especially the part related to Augmentations.",
      "votes": 1
    },
    {
      "id": 2411226,
      "postDate": "2023-08-27T13:37:50.283Z",
      "content": "<p>great work and best wishes for your job</p>",
      "rawMarkdown": "great work and best wishes for your job",
      "votes": 1
    },
    {
      "id": 2411022,
      "postDate": "2023-08-27T11:10:17.897Z",
      "content": "<p>Congrats! Excited to see the code!</p>",
      "rawMarkdown": "Congrats! Excited to see the code!",
      "votes": 1
    },
    {
      "id": 2409932,
      "postDate": "2023-08-26T14:27:49.103Z",
      "content": "<p>This competition is about fingerspelling. However, your model uses 130 data points out of which only 27 points belong to fingers / fingerspelling. How much of the accuracy of your model is contributed by the data points related to face? </p>\n<p>In the case of 0.697 public notebook based on winning solution of the sign language competition, I found that the accuracy falls down to the extent of 0.02 if we just drop the lip related data points alone. So, how much of the increased accuracy of your model is contributed by additional 70 face related datapoints vs accuracy contributed by the improved model architecture?</p>",
      "rawMarkdown": "This competition is about fingerspelling. However, your model uses 130 data points out of which only 27 points belong to fingers / fingerspelling. How much of the accuracy of your model is contributed by the data points related to face? \n\nIn the case of 0.697 public notebook based on winning solution of the sign language competition, I found that the accuracy falls down to the extent of 0.02 if we just drop the lip related data points alone. So, how much of the increased accuracy of your model is contributed by additional 70 face related datapoints vs accuracy contributed by the improved model architecture?",
      "votes": 1,
      "replies": [
        {
          "id": 2410298,
          "postDate": "2023-08-26T19:40:41.677Z",
          "content": "<p>We started the competition with 118 key points referencing the top solution from the first competition. The only additional key points were adding the arms, from pose, which gave an additional ~0.003. We did not do experiments on removing key points. I can imagine by dropping the lips we would get a similar drop as you experienced.</p>",
          "rawMarkdown": "We started the competition with 118 key points referencing the top solution from the first competition. The only additional key points were adding the arms, from pose, which gave an additional ~0.003. We did not do experiments on removing key points. I can imagine by dropping the lips we would get a similar drop as you experienced.",
          "replies": [
            {
              "id": 2410430,
              "postDate": "2023-08-26T23:49:28.987Z",
              "content": "<p>As a research learning / conclusion from this competition, can we conclude that only way to increase the accuracy of American fingerspelling detection and translation is to include data points related to face etc., and just fingers and pose alone are not sufficient? Or is this conclusion reached in previous competition itself and is applicable for this competition also?</p>\n<p>Anyway, congratulations for winning this competition - even a week was so tiring, I could see how much of efforts your team would have put into this competition!</p>\n<p>(I came to this competition in the last week and it is my first deep learning competition. So, I was just tweaking the best public notebook parameters- those trivial experiments (which could also be misleading) made me to conclude as if no model architecture improvement can compensate for the information provided by datapoints related to lips etc. I am sure you would have done much more meaningful experiments. With your expertise in this area. what could be your insights on this?)</p>",
              "rawMarkdown": "As a research learning / conclusion from this competition, can we conclude that only way to increase the accuracy of American fingerspelling detection and translation is to include data points related to face etc., and just fingers and pose alone are not sufficient? Or is this conclusion reached in previous competition itself and is applicable for this competition also?\n\nAnyway, congratulations for winning this competition - even a week was so tiring, I could see how much of efforts your team would have put into this competition!\n\n(I came to this competition in the last week and it is my first deep learning competition. So, I was just tweaking the best public notebook parameters- those trivial experiments (which could also be misleading) made me to conclude as if no model architecture improvement can compensate for the information provided by datapoints related to lips etc. I am sure you would have done much more meaningful experiments. With your expertise in this area. what could be your insights on this?)"
            }
          ]
        }
      ]
    },
    {
      "id": 2408298,
      "postDate": "2023-08-25T14:55:40.903Z",
      "content": "<p>Congratulations for topping the LB.<br>\nThank you for sharing details of your model. Its indeed a very innovative approach.<br>\nCan you also share the running time of your notebook. Did you separate the training<br>\nand inference notebooks?</p>",
      "rawMarkdown": "Congratulations for topping the LB.\nThank you for sharing details of your model. Its indeed a very innovative approach.\nCan you also share the running time of your notebook. Did you separate the training\nand inference notebooks?",
      "votes": 1,
      "replies": [
        {
          "id": 2408301,
          "postDate": "2023-08-25T14:57:30.287Z",
          "content": "<p>training and inference is completely separate. We train locally using pytorch, then convert trained model to tf-lite and only upload submission.zip to kaggle kernel to run inference. <br>\nInference runtime is 4.5h</p>",
          "rawMarkdown": "training and inference is completely separate. We train locally using pytorch, then convert trained model to tf-lite and only upload submission.zip to kaggle kernel to run inference. \nInference runtime is 4.5h",
          "votes": 1,
          "replies": [
            {
              "id": 2408327,
              "postDate": "2023-08-25T15:15:44.737Z",
              "content": "<p>I think that addresses my question about offline. So PyTorch local. It's probably beyond scope here, but I made an effort to go offline to see if my Mac could handle any of this, tried importing the environment, but it was well above my abilities to get it to all work. I assume you're using Nvidia hardware and not a Mac.</p>",
              "rawMarkdown": "I think that addresses my question about offline. So PyTorch local. It's probably beyond scope here, but I made an effort to go offline to see if my Mac could handle any of this, tried importing the environment, but it was well above my abilities to get it to all work. I assume you're using Nvidia hardware and not a Mac."
            },
            {
              "id": 2409071,
              "postDate": "2023-08-26T03:55:11.590Z",
              "content": "<p>Thats very interesting approach. <br>\nCan you provide specifics of how to save the model after training and then reload it for inferencing in another notebook? </p>",
              "rawMarkdown": "Thats very interesting approach. \nCan you provide specifics of how to save the model after training and then reload it for inferencing in another notebook? ",
              "votes": 1
            },
            {
              "id": 2409532,
              "postDate": "2023-08-26T09:54:34.787Z",
              "content": "<ol>\n<li>train the model, and upload the model weights to a kaggle dataset</li>\n<li>create an inference kernel and link the kaggle dataset which contains the model weights</li>\n</ol>",
              "rawMarkdown": "1. train the model, and upload the model weights to a kaggle dataset\n2. create an inference kernel and link the kaggle dataset which contains the model weights",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2408032,
      "postDate": "2023-08-25T11:50:03.277Z",
      "content": "<p>\" To avoid the model learning with pads, and ensure optimal inference runtime, the feature extractor and Macaron structure encoder layers were masked time wise during training.\"</p>\n<p>i have a question. Besides speedup, is there a change in accuracy?<br>\ne.g. whisper asr model doesn't not use masking</p>\n<p>\"The Whisper feature extractor performs two operations. It first pads/truncates a batch of audio samples such that all samples have an input length of 30s. Samples shorter than 30s are padded to 30s by appending zeros to the end of the sequence (zeros in an audio signal corresponding to no signal or silence)\"</p>\n<p>it seems possible to do without masking? thanks</p>",
      "rawMarkdown": "\" To avoid the model learning with pads, and ensure optimal inference runtime, the feature extractor and Macaron structure encoder layers were masked time wise during training.\"\n\ni have a question. Besides speedup, is there a change in accuracy?\ne.g. whisper asr model doesn't not use masking\n\n\"The Whisper feature extractor performs two operations. It first pads/truncates a batch of audio samples such that all samples have an input length of 30s. Samples shorter than 30s are padded to 30s by appending zeros to the end of the sequence (zeros in an audio signal corresponding to no signal or silence)\"\n\nit seems possible to do without masking? thanks",
      "votes": 1,
      "replies": [
        {
          "id": 2408047,
          "postDate": "2023-08-25T12:03:27.900Z",
          "content": "<p>Using no pad in inference and unmasked pad in training results in lower inference score than CV.   <br>\nMasking pads in training leads to same score as no pad in inference.  <br>\nWe worked on masking until the output of a padded sample through pytorch model was the same as an unpadded sample through tfmodel - layer by layer. In the conformer structure we have layernorms, FFN, MHA and CNN. Layernorms, FFN can be masked with reshapes; the attention scores of MHA can be maskfilled, CNN was applied to whole sequence and the output was mask filled with zeros.   <br>\nMasking should be possible in whisper; although I'm not sure how it would affect pretrained weights - we trained from scratch. </p>",
          "rawMarkdown": "Using no pad in inference and unmasked pad in training results in lower inference score than CV.   \nMasking pads in training leads to same score as no pad in inference.  \nWe worked on masking until the output of a padded sample through pytorch model was the same as an unpadded sample through tfmodel - layer by layer. In the conformer structure we have layernorms, FFN, MHA and CNN. Layernorms, FFN can be masked with reshapes; the attention scores of MHA can be maskfilled, CNN was applied to whole sequence and the output was mask filled with zeros.   \nMasking should be possible in whisper; although I'm not sure how it would affect pretrained weights - we trained from scratch. ",
          "votes": 3,
          "replies": [
            {
              "id": 2408264,
              "postDate": "2023-08-25T14:30:20.937Z",
              "content": "<p>Any tips on getting the mask to flow correctly? I tried to get to this point with tf as the prior competition winner used non-padded observations for inference too.  I was able to track the existence of my pad mask through each encoder layer. Unfortunately, non-padded inference dropped my score.  I noticed softmax accepts a mask argument, tried that. Custom Conv1d with self.support_mask=True allowed the mask to continue through with 'same' style padding.  Maybe starting simple to check before adding on more complex layers would have been helpful, confirming each new layer work right would have been helpful?</p>",
              "rawMarkdown": "Any tips on getting the mask to flow correctly? I tried to get to this point with tf as the prior competition winner used non-padded observations for inference too.  I was able to track the existence of my pad mask through each encoder layer. Unfortunately, non-padded inference dropped my score.  I noticed softmax accepts a mask argument, tried that. Custom Conv1d with self.support_mask=True allowed the mask to continue through with 'same' style padding.  Maybe starting simple to check before adding on more complex layers would have been helpful, confirming each new layer work right would have been helpful?"
            },
            {
              "id": 2408432,
              "postDate": "2023-08-25T16:17:11.973Z",
              "content": "<p>We can wait for the source code as according to the competion rule, prize winner will opensource code.</p>",
              "rawMarkdown": "We can wait for the source code as according to the competion rule, prize winner will opensource code.",
              "votes": 2
            },
            {
              "id": 2409596,
              "postDate": "2023-08-26T10:42:51.597Z",
              "content": "<p>Below is the Squeezeformer block showing how we implemented the masking for the ffn and the layernorm using reshapes to ensure no pads were passed through these layers. It shows some of the input shapes for the first masking step in the <code>self.ff_mhsa</code> and <code>self.ln_ff_mhsa</code>.<br>\nThe mhsa block masked attention as normal, and within the ConvModule we just had no masking except a beginning and final step with <code>x = x.masked_fill(~mask_pad, 0.0)</code> - this was a little hacky but seemed to work. If you want I can post the <code>ConvModule</code> in reply - but we will release in the code also. </p>\n<pre><code> (nn.Module):\n     ():\n        (SqueezeformerBlock, self).__init__()\n\n        self.scale_mhsa, self.bias_mhsa = make_scale(encoder_dim)\n        self.scale_ff_mhsa, self.bias_ff_mhsa = make_scale(encoder_dim)\n        self.scale_conv, self.bias_conv = make_scale(encoder_dim)\n        self.scale_ff_conv, self.bias_ff_conv = make_scale(encoder_dim)        \n        self.mhsa_llama = LlamaAttention(LlamaConfig(hidden_size = encoder_dim, \n                                       num_attention_heads = num_attention_heads, \n                                       max_position_embeddings = ))\n        self.ln_mhsa = nn.LayerNorm(encoder_dim)\n\n        self.ff_mhsa = FeedForwardModule(\n                    encoder_dim=encoder_dim,\n                    expansion_factor=feed_forward_expansion_factor,\n                    dropout_p=feed_forward_dropout_p,\n                )\n\n        self.ln_ff_mhsa = nn.LayerNorm(encoder_dim)\n        self.conv = ConvModule(\n                    in_channels=encoder_dim,\n                    kernel_size=conv_kernel_size,\n                    expansion_factor=conv_expansion_factor,\n                    dropout_p=conv_dropout_p,\n                )\n        self.ln_conv = nn.LayerNorm(encoder_dim)\n        self.ff_conv = FeedForwardModule(\n                    encoder_dim=encoder_dim,\n                    expansion_factor=feed_forward_expansion_factor,\n                    dropout_p=feed_forward_dropout_p,\n                )\n        self.ln_ff_conv = nn.LayerNorm(encoder_dim)\n\n     ():\n        \n\n        mask_pad = ( mask).long().().unsqueeze()\n        mask_pad = ~( mask_pad.permute(, ,) * mask_pad)\n        mask_flat = mask.view(-).()\n        bs, slen, nfeats = x.shape\n\n        residual = x\n        x = x * self.scale_mhsa.to(x.dtype) + self.bias_mhsa.to(x.dtype)\n        x = residual + self.mhsa_llama(x, cos, sin, attention_mask = mask_pad.unsqueeze() )[]\n        \n        x_skip = x.view(-, x.shape[-]) \n        x = x_skip[mask_flat].unsqueeze() \n        x = self.ln_mhsa(x) \n\n        residual = x\n        x = x * self.scale_ff_mhsa.to(x.dtype) + self.bias_ff_mhsa.to(x.dtype)\n        x = residual + self.ff_mhsa(x) \n        x = self.ln_ff_mhsa(x) \n        \n        x_skip[mask_flat] = x[].to(x_skip.dtype)\n        x = x_skip.view(bs, slen, nfeats) \n\n        residual = x\n        x = x * self.scale_conv.to(x.dtype) + self.bias_conv.to(x.dtype)\n        x = residual + self.conv(x, mask_pad = mask.().unsqueeze())\n        \n        x_skip = x.view(-, x.shape[-])\n        x = x_skip[mask_flat].unsqueeze()\n        x = self.ln_conv(x)\n\n        residual = x\n        x = x * self.scale_ff_conv.to(x.dtype) + self.bias_ff_conv.to(x.dtype)\n        x = residual + self.ff_conv(x)\n        x = self.ln_ff_conv(x)\n        \n        x_skip[mask_flat] = x[].to(x_skip.dtype)\n        x = x_skip.view(bs, slen, nfeats)  \n\n         x\n</code></pre>",
              "rawMarkdown": "Below is the Squeezeformer block showing how we implemented the masking for the ffn and the layernorm using reshapes to ensure no pads were passed through these layers. It shows some of the input shapes for the first masking step in the `self.ff_mhsa` and `self.ln_ff_mhsa`.\nThe mhsa block masked attention as normal, and within the ConvModule we just had no masking except a beginning and final step with `x = x.masked_fill(~mask_pad, 0.0)` - this was a little hacky but seemed to work. If you want I can post the `ConvModule` in reply - but we will release in the code also. \n\n```\nclass SqueezeformerBlock(nn.Module):\n    def __init__(\n        self,\n        encoder_dim: int = 512,\n        num_attention_heads: int = 8,\n        feed_forward_expansion_factor: int = 4,\n        conv_expansion_factor: int = 2,\n        feed_forward_dropout_p: float = 0.1,\n        attention_dropout_p: float = 0.1,\n        conv_dropout_p: float = 0.1,\n        conv_kernel_size: int = 31,\n    ):\n        super(SqueezeformerBlock, self).__init__()\n        \n        self.scale_mhsa, self.bias_mhsa = make_scale(encoder_dim)\n        self.scale_ff_mhsa, self.bias_ff_mhsa = make_scale(encoder_dim)\n        self.scale_conv, self.bias_conv = make_scale(encoder_dim)\n        self.scale_ff_conv, self.bias_ff_conv = make_scale(encoder_dim)        \n        self.mhsa_llama = LlamaAttention(LlamaConfig(hidden_size = encoder_dim, \n                                       num_attention_heads = num_attention_heads, \n                                       max_position_embeddings = 384))\n        self.ln_mhsa = nn.LayerNorm(encoder_dim)\n        \n        self.ff_mhsa = FeedForwardModule(\n                    encoder_dim=encoder_dim,\n                    expansion_factor=feed_forward_expansion_factor,\n                    dropout_p=feed_forward_dropout_p,\n                )\n            \n        self.ln_ff_mhsa = nn.LayerNorm(encoder_dim)\n        self.conv = ConvModule(\n                    in_channels=encoder_dim,\n                    kernel_size=conv_kernel_size,\n                    expansion_factor=conv_expansion_factor,\n                    dropout_p=conv_dropout_p,\n                )\n        self.ln_conv = nn.LayerNorm(encoder_dim)\n        self.ff_conv = FeedForwardModule(\n                    encoder_dim=encoder_dim,\n                    expansion_factor=feed_forward_expansion_factor,\n                    dropout_p=feed_forward_dropout_p,\n                )\n        self.ln_ff_conv = nn.LayerNorm(encoder_dim)\n\n    def forward(self, x, cos, sin, mask):\n        # Input shapes : [64, 384, 168], [1, 1, 384, 42], [1, 1, 384, 42], [64, 384]\n\n        mask_pad = ( mask).long().bool().unsqueeze(1)\n        mask_pad = ~( mask_pad.permute(0, 2,1) * mask_pad)\n        mask_flat = mask.view(-1).bool()\n        bs, slen, nfeats = x.shape\n        \n        residual = x\n        x = x * self.scale_mhsa.to(x.dtype) + self.bias_mhsa.to(x.dtype)\n        x = residual + self.mhsa_llama(x, cos, sin, attention_mask = mask_pad.unsqueeze(1) )[0]\n        # Skip pad #1\n        x_skip = x.view(-1, x.shape[-1]) # [64, 384, 168] -> [24576, 168]\n        x = x_skip[mask_flat].unsqueeze(0) # -> [1, 11205, 168]\n        x = self.ln_mhsa(x) \n        \n        residual = x\n        x = x * self.scale_ff_mhsa.to(x.dtype) + self.bias_ff_mhsa.to(x.dtype)\n        x = residual + self.ff_mhsa(x) \n        x = self.ln_ff_mhsa(x) \n        # Unskip pad #1\n        x_skip[mask_flat] = x[0].to(x_skip.dtype)\n        x = x_skip.view(bs, slen, nfeats) # -> [64, 384, 168]\n\n        residual = x\n        x = x * self.scale_conv.to(x.dtype) + self.bias_conv.to(x.dtype)\n        x = residual + self.conv(x, mask_pad = mask.bool().unsqueeze(1))\n        # Skip pad #2\n        x_skip = x.view(-1, x.shape[-1])\n        x = x_skip[mask_flat].unsqueeze(0)\n        x = self.ln_conv(x)\n        \n        residual = x\n        x = x * self.scale_ff_conv.to(x.dtype) + self.bias_ff_conv.to(x.dtype)\n        x = residual + self.ff_conv(x)\n        x = self.ln_ff_conv(x)\n        # Unskip pad #2\n        x_skip[mask_flat] = x[0].to(x_skip.dtype)\n        x = x_skip.view(bs, slen, nfeats)  \n\n        return x\n```",
              "votes": 2
            },
            {
              "id": 2410544,
              "postDate": "2023-08-27T04:23:17.720Z",
              "content": "<p>Thank you for posting this.  I can see how you're passing the mask to the attention and conv layers.  It's ok, I can wait for your full code; will give me time to research Squeezeformer and llama attention.</p>\n<p>I used your above feedback on working to get the padded &amp; masked CV to match the non-padded CV. I identified quite a few locations in my model where my mask wasn't working right.  'same' padding in conv1d allowed some of my padding into the kernal, so I instead pre-padded the time axis manually (kernal size -1) and then used the 'valid' conv1d setting as mentioned in the 1st place solution from the last competition.  This keeps pad values out of my kernal when using self.supports_mask = True.  Or worse the mask getting deleted entirely in my positional encoding step…  Also, I was using a squeeze and excite step to weight my channels which was dropping my mask so I switched to a simple dense layer.  Now my non-padded CV matches my padded &amp; masked and i'm on to training a bigger model!  Setting up that CV test was the ticket for me.  Thanks!!!   I will likely update my encoder to this Squeezeformer sometime soon.</p>",
              "rawMarkdown": "Thank you for posting this.  I can see how you're passing the mask to the attention and conv layers.  It's ok, I can wait for your full code; will give me time to research Squeezeformer and llama attention.\n\nI used your above feedback on working to get the padded & masked CV to match the non-padded CV. I identified quite a few locations in my model where my mask wasn't working right.  'same' padding in conv1d allowed some of my padding into the kernal, so I instead pre-padded the time axis manually (kernal size -1) and then used the 'valid' conv1d setting as mentioned in the 1st place solution from the last competition.  This keeps pad values out of my kernal when using self.supports_mask = True.  Or worse the mask getting deleted entirely in my positional encoding step...  Also, I was using a squeeze and excite step to weight my channels which was dropping my mask so I switched to a simple dense layer.  Now my non-padded CV matches my padded & masked and i'm on to training a bigger model!  Setting up that CV test was the ticket for me.  Thanks!!!   I will likely update my encoder to this Squeezeformer sometime soon."
            },
            {
              "id": 2410623,
              "postDate": "2023-08-27T05:58:13.020Z",
              "content": "<p>Nice, glad it works </p>",
              "rawMarkdown": "Nice, glad it works "
            }
          ]
        }
      ]
    },
    {
      "id": 2408029,
      "postDate": "2023-08-25T11:44:48.257Z",
      "content": "<p>Congratulations to this great work. I have some questions about the implementation:</p>\n<ol>\n<li>Do you use the z coordinate? And how do you normalize the data before augmentation?</li>\n<li><code>Most of the time we only trained and tracked the score of fold0 and not all folds.</code> How do you evaluate your trained model? I.e. which fold is your validation set?</li>\n</ol>",
      "rawMarkdown": "Congratulations to this great work. I have some questions about the implementation:\n\n1. Do you use the z coordinate? And how do you normalize the data before augmentation?\n2. `Most of the time we only trained and tracked the score of fold0 and not all folds.` How do you evaluate your trained model? I.e. which fold is your validation set?\n",
      "votes": 1,
      "replies": [
        {
          "id": 2408033,
          "postDate": "2023-08-25T11:50:37.380Z",
          "content": "<p>z-coordinate was used. Normalization was performed as below.</p>\n<pre><code>        \n         = x.reshape(x.shape[],,-).permute(,,)\n         = x[~torch.isnan(x)].view(-, x.shape[-])\n         = x - nonan.mean()[None, None, :]\n         = x / nonan.std(, unbiased=False)[None, None, :]\n</code></pre>\n<p>Fold 0 was our validation set when running to check CV. Fold 0 had signers exclusively assigned to it. </p>",
          "rawMarkdown": "z-coordinate was used. Normalization was performed as below.\n```\n        # (seq_len, 3* n_landmarks) -> (seq_len, n_landmarks, 3)\n        x = x.reshape(x.shape[0],3,-1).permute(0,2,1)\n        nonan = x[~torch.isnan(x)].view(-1, x.shape[-1])\n        x = x - nonan.mean(0)[None, None, :]\n        x = x / nonan.std(0, unbiased=False)[None, None, :]\n```\n\nFold 0 was our validation set when running to check CV. Fold 0 had signers exclusively assigned to it. ",
          "votes": 1,
          "replies": [
            {
              "id": 2408072,
              "postDate": "2023-08-25T12:23:24.257Z",
              "content": "<p>BTW, how much does ensemble help? Did you try to use one bigger model that is 2x of the ensembled model? </p>",
              "rawMarkdown": "BTW, how much does ensemble help? Did you try to use one bigger model that is 2x of the ensembled model? \n"
            },
            {
              "id": 2408235,
              "postDate": "2023-08-25T14:16:04.900Z",
              "content": "<p>Ensemble of two seeds gave ~0.01+, not sure on the second point, I believe so - you can see in the neptune chart diminishing returns on model size, and more than 14 layers did not help noticeably. <br>\np.s. congrats on your gold!</p>",
              "rawMarkdown": "Ensemble of two seeds gave ~0.01+, not sure on the second point, I believe so - you can see in the neptune chart diminishing returns on model size, and more than 14 layers did not help noticeably. \np.s. congrats on your gold!",
              "votes": 4
            },
            {
              "id": 2429898,
              "postDate": "2023-09-08T22:11:40.377Z",
              "content": "<p>Thanks so much for the nice writeup. If I remember correctly, in the competition description, it suggested to discard the z axis data, as the depth info may not be accurate. If you were aware of the suggestion, may I know why you still included z axis data? Did you check the score difference between including and excluding z axis data? Thank you!</p>",
              "rawMarkdown": "Thanks so much for the nice writeup. If I remember correctly, in the competition description, it suggested to discard the z axis data, as the depth info may not be accurate. If you were aware of the suggestion, may I know why you still included z axis data? Did you check the score difference between including and excluding z axis data? Thank you!"
            }
          ]
        }
      ]
    },
    {
      "id": 2410615,
      "postDate": "2023-08-27T05:53:25.493Z",
      "content": "<p>Congratulations and great work! <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> <a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a> <br>\nStaying tuned for code and potential paper 😀</p>",
      "rawMarkdown": "Congratulations and great work! @christofhenkel @darraghdog \nStaying tuned for code and potential paper 😀",
      "votes": 2
    },
    {
      "id": 2409615,
      "postDate": "2023-08-26T10:56:32.467Z",
      "content": "<p>absolute GOAT</p>",
      "rawMarkdown": "absolute GOAT",
      "votes": 2
    },
    {
      "id": 2409337,
      "postDate": "2023-08-26T07:34:17.607Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> and <a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a>!</p>\n<p>I would like to ask a few questions about sign language knowledge because I am a deaf signer of sign language. I work as a junior researcher and developer at our university in the field of sign language recognition, focusing more on Data-Centric AI rather than model-centric AI. In sign language recognition tasks, the main challenge is the lack of sufficient dataset. Your solution based on model-centric AI is impressive, and I fully agree that if I had created a clean dataset myself, your model could have achieved close to 100% accuracy.</p>\n<p>I only started participating in Kaggle about 4 months ago, and when the \"GISLR\" competition came up and beyond, I couldn't even achieve a bronze:) , but I'm learning from your examples and others, and gradually improving my model-centric AI skills.</p>\n<p>I'm not sure if you know sign language. I would like to ask some questions:</p>\n<ul>\n<li>If you were well-versed in sign language, would you approach the solution differently? If yes, how?</li>\n<li>Have you personally studied sign language or interacted with deaf individuals while working on this project?</li>\n</ul>",
      "rawMarkdown": "Congratulations @christofhenkel and @darraghdog!\n\nI would like to ask a few questions about sign language knowledge because I am a deaf signer of sign language. I work as a junior researcher and developer at our university in the field of sign language recognition, focusing more on Data-Centric AI rather than model-centric AI. In sign language recognition tasks, the main challenge is the lack of sufficient dataset. Your solution based on model-centric AI is impressive, and I fully agree that if I had created a clean dataset myself, your model could have achieved close to 100% accuracy.\n\nI only started participating in Kaggle about 4 months ago, and when the \"GISLR\" competition came up and beyond, I couldn't even achieve a bronze:) , but I'm learning from your examples and others, and gradually improving my model-centric AI skills.\n\nI'm not sure if you know sign language. I would like to ask some questions:\n- If you were well-versed in sign language, would you approach the solution differently? If yes, how?\n- Have you personally studied sign language or interacted with deaf individuals while working on this project?",
      "votes": 2,
      "replies": [
        {
          "id": 2409530,
          "postDate": "2023-08-26T09:51:45.763Z",
          "content": "<p>Thank you for your questions.</p>\n<blockquote>\n  <p>Have you personally studied sign language or interacted with deaf individuals while working on this project?</p>\n</blockquote>\n<p>I did not know any sign language before both competitions, but during competitions I also look at the specific domain and try to understand why a model behaves as it does, with respect to mistakes. I went through the different characters on a high level, just to get an understand what is needed to spell a character (single hand? two hands? more?) and how letters, digits and special characters are different. I think understanding of the subtleties helps in model design, because ideally the model is capable at addressing those. Here, understanding fingerspelling and how mediapipe works helped coming up with meaningful augmentations that help the model generalize to new signers. For letters I even spent an hour or two training signing skills via this online game (its quite fun)<br>\n<a href=\"https://www.signlanguageforum.com/asl/fingerspelling/fingerspelling-game/\" target=\"_blank\">https://www.signlanguageforum.com/asl/fingerspelling/fingerspelling-game/</a></p>\n<blockquote>\n  <p>If you were well-versed in sign language, would you approach the solution differently? If yes, how?</p>\n</blockquote>\n<p>I dont think so. I often approach problems iteratively. Train a model, look at the predictions and understand if I can find mistakes that are domain related and can be prevented by a better model design (or by improving data quality). It would be easier to understand the mistakes if I would be well-versed in sign language but the approach would be the same</p>",
          "rawMarkdown": "Thank you for your questions.\n\n>Have you personally studied sign language or interacted with deaf individuals while working on this project?\n\nI did not know any sign language before both competitions, but during competitions I also look at the specific domain and try to understand why a model behaves as it does, with respect to mistakes. I went through the different characters on a high level, just to get an understand what is needed to spell a character (single hand? two hands? more?) and how letters, digits and special characters are different. I think understanding of the subtleties helps in model design, because ideally the model is capable at addressing those. Here, understanding fingerspelling and how mediapipe works helped coming up with meaningful augmentations that help the model generalize to new signers. For letters I even spent an hour or two training signing skills via this online game (its quite fun)\nhttps://www.signlanguageforum.com/asl/fingerspelling/fingerspelling-game/\n\n>If you were well-versed in sign language, would you approach the solution differently? If yes, how?\n\nI dont think so. I often approach problems iteratively. Train a model, look at the predictions and understand if I can find mistakes that are domain related and can be prevented by a better model design (or by improving data quality). It would be easier to understand the mistakes if I would be well-versed in sign language but the approach would be the same",
          "votes": 2,
          "replies": [
            {
              "id": 2409594,
              "postDate": "2023-08-26T10:41:43.007Z",
              "content": "<p>Thank you for your response. If in-depth knowledge of sign language is required for machine learning purposes, feel free to let me know—I'd be happy to assist you :)</p>",
              "rawMarkdown": "Thank you for your response. If in-depth knowledge of sign language is required for machine learning purposes, feel free to let me know—I'd be happy to assist you :)",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2409094,
      "postDate": "2023-08-26T04:10:24.363Z",
      "content": "<p>Congrats , A beautiful and clean solution.</p>",
      "rawMarkdown": "Congrats , A beautiful and clean solution.",
      "votes": 2
    },
    {
      "id": 2409059,
      "postDate": "2023-08-26T03:39:56.563Z",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!",
      "votes": 2
    },
    {
      "id": 2409054,
      "postDate": "2023-08-26T03:37:27.937Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> </p>\n<p>Great solution man!</p>",
      "rawMarkdown": "Congratulations @christofhenkel \n\nGreat solution man!",
      "votes": 2
    },
    {
      "id": 2408993,
      "postDate": "2023-08-26T02:07:46.250Z",
      "content": "<p>Amazing, congratulations!</p>",
      "rawMarkdown": "Amazing, congratulations!",
      "votes": 2
    },
    {
      "id": 2408969,
      "postDate": "2023-08-26T00:46:12.383Z",
      "content": "<p>Congratulations! Thanks for the write-up, great learning resource</p>",
      "rawMarkdown": "Congratulations! Thanks for the write-up, great learning resource",
      "votes": 2
    },
    {
      "id": 2408408,
      "postDate": "2023-08-25T16:09:12.233Z",
      "content": "<p>great.  Very appreciate it for sharing.</p>",
      "rawMarkdown": "great.  Very appreciate it for sharing.",
      "votes": 2
    },
    {
      "id": 2408382,
      "postDate": "2023-08-25T15:52:44.107Z",
      "content": "<p>The diagram is always on point. Congratulations!</p>",
      "rawMarkdown": "The diagram is always on point. Congratulations!",
      "votes": 2
    },
    {
      "id": 2408011,
      "postDate": "2023-08-25T11:32:35.710Z",
      "content": "<p>Excellent work! So many brilliant ideas and you noticed nearly every corner details. While most teams using ctc encoder only method, the seq2seq win 1st place😀 , and also seems using seq2seq we can better ensemble.<br>\nI have used public seq2seq notebook and also found supplement dataset not help, maybe related to seq2seq method and supplement dataset phrase distribution is very different comparing to train dataset. But for ctc encoder based method I found supplement dataset help improve a lot (more then 10 points).<br>\nFor fp16, I did not notice acc diff if train using fp32+awp then infer with converter.target_spec.supported_types = [tf.float16]. Not sure if there are other differences with your setting.</p>",
      "rawMarkdown": "Excellent work! So many brilliant ideas and you noticed nearly every corner details. While most teams using ctc encoder only method, the seq2seq win 1st place😀 , and also seems using seq2seq we can better ensemble.\nI have used public seq2seq notebook and also found supplement dataset not help, maybe related to seq2seq method and supplement dataset phrase distribution is very different comparing to train dataset. But for ctc encoder based method I found supplement dataset help improve a lot (more then 10 points).\nFor fp16, I did not notice acc diff if train using fp32+awp then infer with converter.target_spec.supported_types = [tf.float16]. Not sure if there are other differences with your setting.",
      "votes": 2
    },
    {
      "id": 3125013,
      "postDate": "2025-02-15T17:26:37.967Z",
      "content": "<p>Congrats! I learned a lot from your solution.</p>",
      "rawMarkdown": "Congrats! I learned a lot from your solution."
    },
    {
      "id": 2951253,
      "postDate": "2024-08-08T13:00:56.660Z",
      "content": "<p>I'm a little late to the game here but congrats! </p>\n<p>I'm trying to test your model with live videos and found that the model's output shape is 11, 63. I think 11 represents the sequence of characters. I'm confused about 63 since the character to prediction index only contains 58 characters. Can you help me understand, pleaes? Thanks in advance.</p>",
      "rawMarkdown": "I'm a little late to the game here but congrats! \n\nI'm trying to test your model with live videos and found that the model's output shape is 11, 63. I think 11 represents the sequence of characters. I'm confused about 63 since the character to prediction index only contains 58 characters. Can you help me understand, pleaes? Thanks in advance."
    },
    {
      "id": 2643460,
      "postDate": "2024-02-08T21:00:28.420Z",
      "content": "<p>Good job and best wishes to your job</p>",
      "rawMarkdown": "Good job and best wishes to your job"
    },
    {
      "id": 2629495,
      "postDate": "2024-01-31T20:16:43.167Z",
      "content": "<p>Congratulation <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a>! This is awesome!</p>",
      "rawMarkdown": "Congratulation @christofhenkel! This is awesome!"
    },
    {
      "id": 2598876,
      "postDate": "2024-01-12T14:40:07.730Z",
      "content": "<p>when will the research paper be out?</p>",
      "rawMarkdown": "when will the research paper be out?\n"
    },
    {
      "id": 2586254,
      "postDate": "2024-01-04T04:52:45.733Z",
      "content": "<p>Congratulations</p>",
      "rawMarkdown": "Congratulations"
    },
    {
      "id": 2573281,
      "postDate": "2023-12-24T21:03:09.747Z",
      "content": "<p>Congrats! thank you for this great work and explanation about workflow, model arch, augmentation, and all other things. This solution give a lot of insights for my graduation, so I am very happy to see like this level of work and to learn from, ty for sharing🙏🙏.</p>",
      "rawMarkdown": "Congrats! thank you for this great work and explanation about workflow, model arch, augmentation, and all other things. This solution give a lot of insights for my graduation, so I am very happy to see like this level of work and to learn from, ty for sharing🙏🙏."
    },
    {
      "id": 2553300,
      "postDate": "2023-12-08T06:17:56.640Z",
      "content": "<p>Amazing <br>\nLearned a lot from your write-up</p>",
      "rawMarkdown": "Amazing \nLearned a lot from your write-up\n"
    },
    {
      "id": 2477681,
      "postDate": "2023-10-11T13:46:13.807Z",
      "content": "<p>This is great work. Congrats! Excited to see the code!</p>",
      "rawMarkdown": "This is great work. Congrats! Excited to see the code!"
    },
    {
      "id": 2476730,
      "postDate": "2023-10-10T22:17:10.940Z",
      "content": "<p>Just came across this, what absolute gem. Thanks for sharing it!</p>",
      "rawMarkdown": "Just came across this, what absolute gem. Thanks for sharing it!"
    },
    {
      "id": 2431948,
      "postDate": "2023-09-10T14:08:32.380Z",
      "content": "<p>Congratulations. Indeed a very thorough approach to this. </p>",
      "rawMarkdown": "Congratulations. Indeed a very thorough approach to this. "
    },
    {
      "id": 2431029,
      "postDate": "2023-09-09T18:22:34.983Z",
      "content": "<p>Awesome… Explanation  </p>\n<p>congratulations for your rank🥇</p>",
      "rawMarkdown": "Awesome... Explanation  \n\ncongratulations for your rank🥇"
    },
    {
      "id": 2428188,
      "postDate": "2023-09-07T17:27:53.983Z",
      "content": "<p>Great explanation! Very informative and has good insights. Also, congratulations on your rank! 🎉 #Informative #Congratulations #Insights</p>",
      "rawMarkdown": "Great explanation! Very informative and has good insights. Also, congratulations on your rank! 🎉 #Informative #Congratulations #Insights"
    },
    {
      "id": 2419940,
      "postDate": "2023-09-02T09:27:24.797Z",
      "content": "<p>Wow, Congratulations &lt;3 </p>",
      "rawMarkdown": "Wow, Congratulations <3 "
    },
    {
      "id": 2419645,
      "postDate": "2023-09-02T06:48:44.523Z",
      "content": "<p>great ideas and work</p>",
      "rawMarkdown": "great ideas and work"
    },
    {
      "id": 2419149,
      "postDate": "2023-09-01T18:03:44.647Z",
      "content": "<p>felicidades</p>",
      "rawMarkdown": "felicidades"
    },
    {
      "id": 2419012,
      "postDate": "2023-09-01T16:33:51.303Z",
      "content": "<p>Thanks again for all your helpful comments.</p>\n<p>Any tips to avoid getting nan loss with mixed precision fp16 training?  I note you avoided AWP.  I see in tensorflow's instructions they say to convert the last softmax to float32.  I also note that you trained in pytorch and possibly that has more support in skipping nan batches or something like that.  I've been using model.fit in tensor flow which does some loss scaling to avoid underflow.  I get nan's at about the 5th or 6th epoch.  The main reason I'm interested in fp16 is because I'm running out of memory with a modest sized squeezeformer on a 3090 with a frame length of only 300 batch size 64, and I'm interested in increasing parameters.  There could be some clumsy implementation going on with my squeezeformer, but the CV is looking good so it can't be too far off. I can go to batch size of 32… but I'll be training for quite a while.</p>",
      "rawMarkdown": "Thanks again for all your helpful comments.\n\nAny tips to avoid getting nan loss with mixed precision fp16 training?  I note you avoided AWP.  I see in tensorflow's instructions they say to convert the last softmax to float32.  I also note that you trained in pytorch and possibly that has more support in skipping nan batches or something like that.  I've been using model.fit in tensor flow which does some loss scaling to avoid underflow.  I get nan's at about the 5th or 6th epoch.  The main reason I'm interested in fp16 is because I'm running out of memory with a modest sized squeezeformer on a 3090 with a frame length of only 300 batch size 64, and I'm interested in increasing parameters.  There could be some clumsy implementation going on with my squeezeformer, but the CV is looking good so it can't be too far off. I can go to batch size of 32... but I'll be training for quite a while.",
      "replies": [
        {
          "id": 2419243,
          "postDate": "2023-09-01T20:50:37.860Z",
          "content": "<p>gradient accumulation (just google it) is your friend when you want to use higher batch size. Related to NaNs while training its really hard to deal with them, hard to give some advice, since it depends on all kind of factors</p>",
          "rawMarkdown": "gradient accumulation (just google it) is your friend when you want to use higher batch size. Related to NaNs while training its really hard to deal with them, hard to give some advice, since it depends on all kind of factors",
          "votes": 1
        }
      ]
    },
    {
      "id": 2418427,
      "postDate": "2023-09-01T09:20:13.430Z",
      "content": "<p>Congratulations, <br>\nAppreciate your work.</p>",
      "rawMarkdown": "Congratulations, \nAppreciate your work."
    },
    {
      "id": 2416427,
      "postDate": "2023-08-31T03:28:48.943Z",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!"
    },
    {
      "id": 2416351,
      "postDate": "2023-08-31T01:32:14.870Z",
      "content": "<p>Wow…..<br>\nCongratulations on your achievement</p>",
      "rawMarkdown": "Wow.....\nCongratulations on your achievement\n"
    },
    {
      "id": 2416066,
      "postDate": "2023-08-30T19:04:57.600Z",
      "content": "<p>I'm curious to know more about your confidence score calculation. How did you calculate the loss and how did you combine it with the cross entropy of the decoder for training?</p>",
      "rawMarkdown": "I'm curious to know more about your confidence score calculation. How did you calculate the loss and how did you combine it with the cross entropy of the decoder for training?",
      "replies": [
        {
          "id": 2418857,
          "postDate": "2023-09-01T14:56:30.910Z",
          "content": "<p>Basically calculate levensthein distance for each data sample and clip to [0,1]. Then use that as auxiliary loss using Binary Crossentropy loss. Combine with the \"normal\" loss using weighted average with weight of 0.98 for the normal loss and 0.02 for the auxiliary loss. You can find the details here</p>\n<p><a href=\"https://github.com/ChristofHenkel/kaggle-asl-fingerspelling-1st-place-solution/blob/main/models/mdl_2_pt.py#L867-L871\" target=\"_blank\">https://github.com/ChristofHenkel/kaggle-asl-fingerspelling-1st-place-solution/blob/main/models/mdl_2_pt.py#L867-L871</a></p>",
          "rawMarkdown": "Basically calculate levensthein distance for each data sample and clip to [0,1]. Then use that as auxiliary loss using Binary Crossentropy loss. Combine with the \"normal\" loss using weighted average with weight of 0.98 for the normal loss and 0.02 for the auxiliary loss. You can find the details here\n\nhttps://github.com/ChristofHenkel/kaggle-asl-fingerspelling-1st-place-solution/blob/main/models/mdl_2_pt.py#L867-L871",
          "votes": 1,
          "replies": [
            {
              "id": 2423268,
              "postDate": "2023-09-04T14:38:01.770Z",
              "content": "<p>thank you! 🙏</p>",
              "rawMarkdown": "thank you! 🙏"
            }
          ]
        }
      ]
    },
    {
      "id": 2414448,
      "postDate": "2023-08-29T16:28:29.840Z",
      "content": "<p>good content  i like this thanku</p>",
      "rawMarkdown": "good content  i like this thanku"
    },
    {
      "id": 2413796,
      "postDate": "2023-08-29T07:02:53.900Z",
      "content": "<p>What is the meaning of OOF in this context?</p>\n<p>Thank you for the explanation, I am extracting a lot of knowledge from it.</p>",
      "rawMarkdown": "What is the meaning of OOF in this context?\n\nThank you for the explanation, I am extracting a lot of knowledge from it.",
      "replies": [
        {
          "id": 2413825,
          "postDate": "2023-08-29T07:27:32.717Z",
          "content": "<p>OOF means out-of-fold predictions. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2F91c133b7b849bc0bd9c4ce0b05dcb492%2FScreenshot%202023-08-29%20at%2009.24.15.png?generation=1693293984960629&amp;alt=media\" alt=\"\"></p>\n<p>We had you train a model for each fold and concatenate the resulting sets of validation scores. This gives you a normalized levensthein score for each data sample. We used this as an auxiliary target for training a new model. </p>",
          "rawMarkdown": "OOF means out-of-fold predictions. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2F91c133b7b849bc0bd9c4ce0b05dcb492%2FScreenshot%202023-08-29%20at%2009.24.15.png?generation=1693293984960629&alt=media)\n\nWe had you train a model for each fold and concatenate the resulting sets of validation scores. This gives you a normalized levensthein score for each data sample. We used this as an auxiliary target for training a new model. ",
          "votes": 2,
          "replies": [
            {
              "id": 2413855,
              "postDate": "2023-08-29T07:42:10.283Z",
              "content": "<p>Learned something new! Thank you so much!</p>",
              "rawMarkdown": "Learned something new! Thank you so much!"
            }
          ]
        }
      ]
    },
    {
      "id": 2413022,
      "postDate": "2023-08-28T16:29:32.157Z",
      "content": "<p>Great effort sir!</p>",
      "rawMarkdown": "Great effort sir!"
    },
    {
      "id": 2412870,
      "postDate": "2023-08-28T14:52:29.207Z",
      "content": "<p><a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> <a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a><br>\nCongratulations 🎉 and thank you for the great description!</p>\n<p>I have a question (might be for beginners).</p>\n<ul>\n<li>How and when did you decide hyper parameters of learning schedule?</li>\n</ul>\n<p>The peak LR (0.0045) and weight decay (0.08) look pretty fine grained. I'd love to know how these parameters are tuned (e.g., Peak LR is tested at 0.0005 intervals for each model) and when these are decided (e.g., tuned peak LR even when Augmentations, or fixed for each model at first, or fixed for some good model).</p>",
      "rawMarkdown": "@christofhenkel @darraghdog\nCongratulations :tada: and thank you for the great description!\n\nI have a question (might be for beginners).\n* How and when did you decide hyper parameters of learning schedule?\n\nThe peak LR (0.0045) and weight decay (0.08) look pretty fine grained. I'd love to know how these parameters are tuned (e.g., Peak LR is tested at 0.0005 intervals for each model) and when these are decided (e.g., tuned peak LR even when Augmentations, or fixed for each model at first, or fixed for some good model).",
      "replies": [
        {
          "id": 2413175,
          "postDate": "2023-08-28T17:40:42.750Z",
          "content": "<blockquote>\n  <p>The peak LR (0.0045) and weight decay (0.08) look pretty fine grained.</p>\n</blockquote>\n<p>It looks like its fine grained but it was rather the opposite. We started with some learning rate reported in the original SqueezeFormer paper, had some miscommunication within the team with respect to using DDP (which means lr / batch size should be adjusted) and ended up with a LR of 0.0045 more or less by accident, I tested a slightly different higher/ lower LR later and it did not change much, so we sticked with it. <br>\nFor weight decay we started with 0, then tried 0.01 which gave a significant improvement, then 0.05 which helped even more. Finally we tried 0.08 which was slightly better than 0.05 but in range which I would say its just noise, so we thought we keep at that. </p>",
          "rawMarkdown": ">The peak LR (0.0045) and weight decay (0.08) look pretty fine grained.\n\nIt looks like its fine grained but it was rather the opposite. We started with some learning rate reported in the original SqueezeFormer paper, had some miscommunication within the team with respect to using DDP (which means lr / batch size should be adjusted) and ended up with a LR of 0.0045 more or less by accident, I tested a slightly different higher/ lower LR later and it did not change much, so we sticked with it. \nFor weight decay we started with 0, then tried 0.01 which gave a significant improvement, then 0.05 which helped even more. Finally we tried 0.08 which was slightly better than 0.05 but in range which I would say its just noise, so we thought we keep at that. ",
          "votes": 1,
          "replies": [
            {
              "id": 2414092,
              "postDate": "2023-08-29T11:46:26.003Z",
              "content": "<p>Thank you for your response and the explanation. This is very useful as practical knowledge for me.</p>",
              "rawMarkdown": "Thank you for your response and the explanation. This is very useful as practical knowledge for me."
            }
          ]
        }
      ]
    },
    {
      "id": 2412353,
      "postDate": "2023-08-28T08:29:24.480Z",
      "content": "<p>Amazing, congratulations!</p>",
      "rawMarkdown": "Amazing, congratulations!"
    },
    {
      "id": 2412104,
      "postDate": "2023-08-28T05:04:00.737Z",
      "content": "<p>Congratulations and thank you for sharing your solution.</p>",
      "rawMarkdown": "Congratulations and thank you for sharing your solution."
    },
    {
      "id": 2412098,
      "postDate": "2023-08-28T05:00:31.753Z",
      "content": "<p>Congratulations!<br>\nThanks for sharing solution!</p>",
      "rawMarkdown": "Congratulations!\nThanks for sharing solution!"
    },
    {
      "id": 2411959,
      "postDate": "2023-08-28T02:07:21.873Z",
      "content": "<p>Congratulations! Thanks for knowledge sharing!</p>",
      "rawMarkdown": "Congratulations! Thanks for knowledge sharing!"
    },
    {
      "id": 2411248,
      "postDate": "2023-08-27T13:54:17.497Z",
      "content": "<p>Congratz on your win and thanks for the great description. Your architecture changes, especially the parameter saving and data transforms seem really, really well done!<br>\nTwo things I didn't see:</p>\n<ol>\n<li>What dimensions did you use in the encoder in the end? The screenshot shows 128, 168 and 192.</li>\n<li>What kind of hardware did you use to train?  </li>\n</ol>\n<p>Thanks!</p>",
      "rawMarkdown": "Congratz on your win and thanks for the great description. Your architecture changes, especially the parameter saving and data transforms seem really, really well done!\nTwo things I didn't see:\n1. What dimensions did you use in the encoder in the end? The screenshot shows 128, 168 and 192.\n2. What kind of hardware did you use to train?  \n\nThanks!",
      "replies": [
        {
          "id": 2411798,
          "postDate": "2023-08-27T21:35:23.137Z",
          "content": "<blockquote>\n  <p>What dimensions did you use in the encoder in the end? The screenshot shows 128, 168 and 192.</p>\n</blockquote>\n<p>Encoder was 14x SequeezeFormerBlocks with 208 dim. 14 layers seemed optimal and then we maximized the dim to fit 2 seeds within the 40MB restriction.</p>\n<blockquote>\n  <p>What kind of hardware did you use to train?</p>\n</blockquote>\n<p>I mainly used NVIDIA V100 16GB GPUs</p>",
          "rawMarkdown": ">What dimensions did you use in the encoder in the end? The screenshot shows 128, 168 and 192.\n\nEncoder was 14x SequeezeFormerBlocks with 208 dim. 14 layers seemed optimal and then we maximized the dim to fit 2 seeds within the 40MB restriction.\n\n>What kind of hardware did you use to train?\n\nI mainly used NVIDIA V100 16GB GPUs",
          "votes": 2
        }
      ]
    },
    {
      "id": 2410960,
      "postDate": "2023-08-27T09:57:44.010Z",
      "content": "<p>Congratulations </p>",
      "rawMarkdown": "Congratulations "
    },
    {
      "id": 2410861,
      "postDate": "2023-08-27T08:51:31.330Z",
      "content": "<p>By the way <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a>, did you use <a href=\"https://github.com/AlexanderLutsenko/nobuco\" target=\"_blank\">https://github.com/AlexanderLutsenko/nobuco</a> for helping with the PyTorch to TF translation or you did it solely by hand?</p>",
      "rawMarkdown": "By the way @christofhenkel, did you use https://github.com/AlexanderLutsenko/nobuco for helping with the PyTorch to TF translation or you did it solely by hand?",
      "replies": [
        {
          "id": 2410906,
          "postDate": "2023-08-27T09:20:24.790Z",
          "content": "<p>Solely by hand. What came in handy was that there is a more or less consistent implementation of squeezeformer in both here</p>\n<p>torch: <a href=\"https://github.com/upskyy/Squeezeformer/tree/main\" target=\"_blank\">https://github.com/upskyy/Squeezeformer/tree/main</a><br>\ntensorflow: <a href=\"https://github.com/kssteven418/Squeezeformer/tree/main\" target=\"_blank\">https://github.com/kssteven418/Squeezeformer/tree/main</a></p>\n<p>Also, Speech2TextDecoder and LLamaAttention is implemented in huggingface in both, torch as well as tensorflow. So its some work to translate the architecture solely by hand, but less work than you would expect. More tricky translations were the custom masking, and the caching/ early stopping of the decoder</p>",
          "rawMarkdown": "Solely by hand. What came in handy was that there is a more or less consistent implementation of squeezeformer in both here\n\ntorch: https://github.com/upskyy/Squeezeformer/tree/main\ntensorflow: https://github.com/kssteven418/Squeezeformer/tree/main\n\nAlso, Speech2TextDecoder and LLamaAttention is implemented in huggingface in both, torch as well as tensorflow. So its some work to translate the architecture solely by hand, but less work than you would expect. More tricky translations were the custom masking, and the caching/ early stopping of the decoder",
          "votes": 1
        },
        {
          "id": 2410909,
          "postDate": "2023-08-27T09:21:28.050Z",
          "content": "<p>The exercise was also useful to refresh some tensorflow skills, as I had not used it for years. </p>",
          "rawMarkdown": "The exercise was also useful to refresh some tensorflow skills, as I had not used it for years. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 2410754,
      "postDate": "2023-08-27T07:19:46.360Z",
      "content": "<p>Amazing and professional approach to find the solution. <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> </p>",
      "rawMarkdown": "Amazing and professional approach to find the solution. @christofhenkel "
    },
    {
      "id": 2425345,
      "postDate": "2023-09-05T20:12:06.553Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 2420637,
      "postDate": "2023-09-02T18:53:44.223Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2418751,
      "postDate": "2023-09-01T13:49:34.727Z",
      "rawMarkdown": "",
      "votes": -1,
      "isDeleted": true
    },
    {
      "id": 2409074,
      "postDate": "2023-08-26T03:59:02.080Z",
      "content": "<p>Thanks for new learning :)</p>",
      "rawMarkdown": "Thanks for new learning :)",
      "votes": 1
    },
    {
      "id": 2532299,
      "postDate": "2023-11-20T23:29:56.597Z",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing."
    }
  ],
  "comments": [
    {
      "id": 2410650,
      "author_name": "Dheeraj Pandey",
      "author_url": "",
      "post_date": "2023-08-27T06:23:03.387000",
      "content": "<p>Congratulations&nbsp;@christofhenkel&nbsp;and&nbsp;@darraghdog </p>\n<p>Appreciate your effort! 🙏  </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2408132,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2023-08-25T13:20:46.510000",
      "content": "<p>Super congrats <a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a> and <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> for this huge win! Thanks for sharing such detailed solution.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2407993,
      "author_name": "Jungwoo Park",
      "author_url": "",
      "post_date": "2023-08-25T11:17:46.003000",
      "content": "<blockquote>\n  <p>In general, we observed that a deeper model gives significant gains (if we are able to prevent overfitting).</p>\n</blockquote>\n<p>I agree. For me, 24L 256d was better than 9L 384d, even though they have similar sizes.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2408367,
      "author_name": "Pascal Pfeiffer",
      "author_url": "",
      "post_date": "2023-08-25T15:38:25.200000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> and <a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a> <br>\nA beautiful and clean solution. I am sure you tried a ton of different ideas how to encode these features.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 2408273,
      "author_name": "Kyle Proffitt",
      "author_url": "",
      "post_date": "2023-08-25T14:36:39.903000",
      "content": "<p>I'm going to post a little comment. It might get long. Hopefully it's helpful for someone.</p>\n<p>First, Wow. I just saw Greysnow's comment saying he feels like a child… coming in 12th. Here I am 386 by copying others' work… if he's a child, I am… Anyway, really great writeup and lots of great ideas. I also look forward to seeing all the code.</p>\n<p>Well, I entered this competition as something to do and to run a couple of experiments. I've gone through a decent amount of the general learning stuff on kaggle. I don't have a background in coding (have a PhD in something else, so you know, I can probably learn things), but I love seeing what people can do in this arena and keep exploring whether I can truly get into it. I feel like I should be somewhat embarrassed to say chatGPT was my guide, and that was a big part of my experiment-- how much could this newer tool help someone as clueless as me to do anything useful. On the one hand, amazing what it could help me accomplish, but also as I think most of you know, it doesn't always take the most elegant approach, and sometimes after a few iterations of changes and errors, I'd hunker down and really try on my own and figure out how to fix the problem myself faster. My main goals were just to see how much of this I could understand, and when I had an idea, could I figure out how to implement it, all with the hope that at the end of the contest when the top solutions are posted, I could really grasp it at the least (not sure I'm there, but I've made progress). As I've posted elsewhere, I was curious whether the winner is really elegant and like totally new solutions, and/or brute force with systematic hyperparameter tuning, taking it offline to run near-infinite variations on some other dedicated hardware, etc.  I'm not sure about the offline dedicated hardware part, but I think from this writeup, there's some of both.</p>\n<p>Well I do have a few general questions that I'd like to understand better. Some of the first stuff I did with other people's code was try to simplify down to only hand landmarks. I didn't and generally still don't see what face landmarks could be helpful for in a fingerspelling dataset. I suppose it's not that I can't imagine that perhaps on a pause as someone repeats a letter, that there might be an associated facial expression the model could pick up-- but it feels quite minimal to me, and I thought there would be more benefit to not having to look at so much extra information that is mostly useless. I did come to think that some of the pose landmarks are probably helpful with regard to motion letters like \"z\" and perhaps other things-- though it still seems like a good model should be able to get all the info it needs from just the hand. I also tried ignoring the z coordinates, and I tried flipping one of the hands, thinking that it made more sense for the model to learn a single set of signs and motions, where the opposite hand ultimately takes a flipped orientation. I was not super systematic with recording how a single change affected loss or score, but more often than not whatever change I attempted gave worse performance. Trying to get rid of lhand and rhand with an xy flip of one (while I was ignoring z) definitely hurt, which surprised me. </p>\n<p>I tried adding a displacement feature by calculating the, like Euclidean norm-change from frame to frame. I was hopeful that this might help the model with double letters and spaces, which it seemed to me from output was an area of trouble, maybe unsurprisingly. However, after much effort to finally successfully implement the addition of this feature and get it all to work together, again, it hurt model performance. Maybe that makes perfect sense to people who understand better how transformers work? As I understand it, it's never entirely clear whether adding a new feature will really help a model to focus on something, or will just be redundant information. I wondered, but did not try, whether trying to identify some like biggest changes, and creating new features based on that, might help-- like maybe calculating the distance from thumb to pinky as a new feature, just as potential example. Again, is that just likely to be redundant, or does the idea not work well with the transformer architecture?</p>\n<p>I tried separating out training and validation datasets based on the participant_id, thinking it might help if the model trained on one signer and their idiosyncrasies at a time-- well and I also tried batching by signer, I guess that's really the training on one signer at a time. Didn't help.</p>\n<p>I did what others have done to make little gifs out of the signing just so I could visualize it, and I also played around with adding to the dataset by recording new landmark data from YouTube videos (and even from myself, just to do it), but it wasn't easy for me to think of how I might systematically add to the dataset without being really manual about it. That's still a question that I'm curious about and haven't answered for myself: essentially what is \"enough\" data? If you take this winning model and you train on 20% of the data, how will that compare to using 80% of it- and if we kept increasing the quantity of data here, would we continue to see improvement? </p>\n<p>I also considered augmenting the data with little rotations, but I didn't get around to implementation. I'm intrigued by the augmentation of dropping fingers and such</p>\n<p>I used greysnow's notebook and code to get things working on TPU when I ran out of GPU time, but man it was painful. It seems every one of those efforts I made above added a layer of difficulty if I tried to run it on TPU. I can only imagine the use of PyTorch and translation efforts in this notebook were quite more work.</p>\n<p>So-- a general question. While it's amazing to watch the model work and create useful predictions, the 0.84 score means, I think, that on an average 11-12-character phrase, there's nearly 2 changes necessary to get to the right phrase. From looking at the output, I think it's actually the case that you get a decent amount of perfect translation, but there are also a number that are far off. Anyway, this is the best of the bunch here, it's definitely impressive, but is it objectively \"good\"? Is it only the starting point for developing a really useful application? I don't have the context to understand what threshold must be met for people to actually want to use it. This dataset includes garbage data, which is dealt with in this notebook and others, but aside from the truly garbage data, I would think if you took a few ASL-experienced individuals and had them watch the original signing, or perhaps even the mediapipe-processed stick figures, they would be able to achieve very nearly 100% accuracy. If yes, then I'd like to think we could train models to do just as well. And in a way there's a disconnect for me in what a human can infer primarily from watching for known symbols compared with the way these models work that are not in any way informed what an \"A\" or \"B\" looks like. I also explored the idea of feeding this information in some way, but I don't know that there's a good way to do that. A different question to ask is if there were not limitations made such that this can be a TFlite model or if we could throw more GPUs and much longer training time, etc., could we then pretty easily bump this accuracy up?</p>\n<p>Now that the competition is over, do people tend to tweak further using ideas from other notebooks? Basically I wonder if someone right now could take the best ideas from different notebooks, could we quickly get to 0.9? Or if this solution has already maxed out on what's possible given the constraints. </p>\n<p>Well, thank you for posting the writeup, congratulations on the good work, and I'm interested to learn and understand more.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2408337,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2023-08-25T15:22:22.373000",
          "content": "<p>Let me comment on a few of your point. Wont go into details, just give some of my thoughts</p>\n<blockquote>\n  <p>Well, I entered this competition as something to do and to run a couple of experiments. I've gone through a decent amount of the general learning stuff on kaggle. I don't have a background in coding (have a PhD in something else, so you know, I can probably learn things),</p>\n</blockquote>\n<p>nobody starts with ML and directly wins competitions. Having also no python nor ML knowledge I worked my ass off and got a place of 500/2000 in my first kaggle competition, which by the way was on speech recognition 6 years ago. But with perseverance and curiosity I continued learning competition by competition and build a skill or two since then.  </p>\n<p>Regarding your point of flipping hands and using face landmarks. Some signers tend to support their signing with lip movement, which can be supportive signal. Also, not all signs are single handed. E.g. <code>+</code> which is in nearly all phone numbers uses both hands</p>\n<blockquote>\n  <p>the 0.84 score means, I think, that on an average 11-12-character phrase, there's nearly 2 changes necessary to get to the right phrase. From looking at the output, I think it's actually the case that you get a decent amount of perfect translation, but there are also a number that are far off. </p>\n</blockquote>\n<p>You need to keep in mind that 10% of the data is corrupted and without signal coming from artefacts where people just tapped the record button, and its impossible to predict the target. But those cases are not relevant to the final usability. For the sequences that contain the actual phrases, our model has a very high accuracy. </p>",
          "votes": 12,
          "replies": [
            {
              "id": 2408374,
              "author_name": "Kyle Proffitt",
              "author_url": "",
              "post_date": "2023-08-25T15:43:56.267000",
              "content": "<p>Wow, thank you for reading and commenting. I really like hearing that you learned on the job, so to speak, gives hope. (I certainly didn't think I'd win, but the prize money gives me cover for spending a lot of time on this and neglecting other things in life just the same). </p>\n<p>I didn't realize plus symbol used both hands. Oops. I probably looked at a few phrases that didn't have + signs and just saw the entire empty data from one hand. I recognize if this was not fingerspelling but more general ASL it would make total sense to use all of the data. I just thought we could simplify here. But just the + information makes quite clear why trying to flip hands or simplify to one ruined it. Thanks for the insight there.</p>\n<p>re: accuracy, Is it not true that when you look at a validation callback, you still see a number of phrases that your model just can't get right? I essentially thought that your code effectively removed the artifactual sequences from consideration, so at least for training and validation, this would be dealt with. If the test dataset also contains artifactual data, there's not much that can be done with it, and the model doesn't have the luxury of ignoring it, even though the landmark data in no way actually corresponds to the phrase. But then I'd expect you to be able to calculate your own Levenshtein score as much higher, since it ignores artifacts, and to see a big drop once it is used on test data. I guess at the end of the day I'd like to see the performance on purely cleaned up, like independently verified, data. Thanks again.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2408446,
              "author_name": "Dieter",
              "author_url": "",
              "post_date": "2023-08-25T16:23:31.597000",
              "content": "<p>I was not talking about raw validation score but rather about the usefulness of our model in the real world, as (at least thats how I understand) it will be directly deployed in an app to help parents learn fingerspelling and support them communicate with their kids. In the real world example users, know when the did not do a proper recording, so only the quality of the model with respect to correctly recorded sequences matters. And thats where our model is quite strong. I would be very interested in human level performance for the train/ test dataset, but I would assume our model is on the same level. Dont know if hosts evaluated this.</p>",
              "votes": 3,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2408048,
      "author_name": "greySnow",
      "author_url": "",
      "post_date": "2023-08-25T12:03:31.203000",
      "content": "<p>Wow, there is so much high-level work here to learn from. I eagerly await your code. Compared to your solution, I feel like a child 😅</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 2447454,
      "author_name": "Zihua Meng",
      "author_url": "",
      "post_date": "2023-09-20T05:13:31.733000",
      "content": "<p>Congratulations! Your solution is really enlightening to me! Especially the part related to Augmentations.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2411226,
      "author_name": "Al Sani",
      "author_url": "",
      "post_date": "2023-08-27T13:37:50.283000",
      "content": "<p>great work and best wishes for your job</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2411022,
      "author_name": "Hendrik Zimmermann",
      "author_url": "",
      "post_date": "2023-08-27T11:10:17.897000",
      "content": "<p>Congrats! Excited to see the code!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2409932,
      "author_name": "Murugesan Narayanaswamy",
      "author_url": "",
      "post_date": "2023-08-26T14:27:49.103000",
      "content": "<p>This competition is about fingerspelling. However, your model uses 130 data points out of which only 27 points belong to fingers / fingerspelling. How much of the accuracy of your model is contributed by the data points related to face? </p>\n<p>In the case of 0.697 public notebook based on winning solution of the sign language competition, I found that the accuracy falls down to the extent of 0.02 if we just drop the lip related data points alone. So, how much of the increased accuracy of your model is contributed by additional 70 face related datapoints vs accuracy contributed by the improved model architecture?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2410298,
          "author_name": "Darragh",
          "author_url": "",
          "post_date": "2023-08-26T19:40:41.677000",
          "content": "<p>We started the competition with 118 key points referencing the top solution from the first competition. The only additional key points were adding the arms, from pose, which gave an additional ~0.003. We did not do experiments on removing key points. I can imagine by dropping the lips we would get a similar drop as you experienced.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2410430,
              "author_name": "Murugesan Narayanaswamy",
              "author_url": "",
              "post_date": "2023-08-26T23:49:28.987000",
              "content": "<p>As a research learning / conclusion from this competition, can we conclude that only way to increase the accuracy of American fingerspelling detection and translation is to include data points related to face etc., and just fingers and pose alone are not sufficient? Or is this conclusion reached in previous competition itself and is applicable for this competition also?</p>\n<p>Anyway, congratulations for winning this competition - even a week was so tiring, I could see how much of efforts your team would have put into this competition!</p>\n<p>(I came to this competition in the last week and it is my first deep learning competition. So, I was just tweaking the best public notebook parameters- those trivial experiments (which could also be misleading) made me to conclude as if no model architecture improvement can compensate for the information provided by datapoints related to lips etc. I am sure you would have done much more meaningful experiments. With your expertise in this area. what could be your insights on this?)</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2408298,
      "author_name": "C R Suthikshn Kumar",
      "author_url": "",
      "post_date": "2023-08-25T14:55:40.903000",
      "content": "<p>Congratulations for topping the LB.<br>\nThank you for sharing details of your model. Its indeed a very innovative approach.<br>\nCan you also share the running time of your notebook. Did you separate the training<br>\nand inference notebooks?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2408301,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2023-08-25T14:57:30.287000",
          "content": "<p>training and inference is completely separate. We train locally using pytorch, then convert trained model to tf-lite and only upload submission.zip to kaggle kernel to run inference. <br>\nInference runtime is 4.5h</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2408327,
              "author_name": "Kyle Proffitt",
              "author_url": "",
              "post_date": "2023-08-25T15:15:44.737000",
              "content": "<p>I think that addresses my question about offline. So PyTorch local. It's probably beyond scope here, but I made an effort to go offline to see if my Mac could handle any of this, tried importing the environment, but it was well above my abilities to get it to all work. I assume you're using Nvidia hardware and not a Mac.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2409071,
              "author_name": "C R Suthikshn Kumar",
              "author_url": "",
              "post_date": "2023-08-26T03:55:11.590000",
              "content": "<p>Thats very interesting approach. <br>\nCan you provide specifics of how to save the model after training and then reload it for inferencing in another notebook? </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2409532,
              "author_name": "Dieter",
              "author_url": "",
              "post_date": "2023-08-26T09:54:34.787000",
              "content": "<ol>\n<li>train the model, and upload the model weights to a kaggle dataset</li>\n<li>create an inference kernel and link the kaggle dataset which contains the model weights</li>\n</ol>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2408032,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-08-25T11:50:03.277000",
      "content": "<p>\" To avoid the model learning with pads, and ensure optimal inference runtime, the feature extractor and Macaron structure encoder layers were masked time wise during training.\"</p>\n<p>i have a question. Besides speedup, is there a change in accuracy?<br>\ne.g. whisper asr model doesn't not use masking</p>\n<p>\"The Whisper feature extractor performs two operations. It first pads/truncates a batch of audio samples such that all samples have an input length of 30s. Samples shorter than 30s are padded to 30s by appending zeros to the end of the sequence (zeros in an audio signal corresponding to no signal or silence)\"</p>\n<p>it seems possible to do without masking? thanks</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2408047,
          "author_name": "Darragh",
          "author_url": "",
          "post_date": "2023-08-25T12:03:27.900000",
          "content": "<p>Using no pad in inference and unmasked pad in training results in lower inference score than CV.   <br>\nMasking pads in training leads to same score as no pad in inference.  <br>\nWe worked on masking until the output of a padded sample through pytorch model was the same as an unpadded sample through tfmodel - layer by layer. In the conformer structure we have layernorms, FFN, MHA and CNN. Layernorms, FFN can be masked with reshapes; the attention scores of MHA can be maskfilled, CNN was applied to whole sequence and the output was mask filled with zeros.   <br>\nMasking should be possible in whisper; although I'm not sure how it would affect pretrained weights - we trained from scratch. </p>",
          "votes": 3,
          "replies": [
            {
              "id": 2408264,
              "author_name": "Andy Atkinson",
              "author_url": "",
              "post_date": "2023-08-25T14:30:20.937000",
              "content": "<p>Any tips on getting the mask to flow correctly? I tried to get to this point with tf as the prior competition winner used non-padded observations for inference too.  I was able to track the existence of my pad mask through each encoder layer. Unfortunately, non-padded inference dropped my score.  I noticed softmax accepts a mask argument, tried that. Custom Conv1d with self.support_mask=True allowed the mask to continue through with 'same' style padding.  Maybe starting simple to check before adding on more complex layers would have been helpful, confirming each new layer work right would have been helpful?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2408432,
              "author_name": "gezi",
              "author_url": "",
              "post_date": "2023-08-25T16:17:11.973000",
              "content": "<p>We can wait for the source code as according to the competion rule, prize winner will opensource code.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2409596,
              "author_name": "Darragh",
              "author_url": "",
              "post_date": "2023-08-26T10:42:51.597000",
              "content": "<p>Below is the Squeezeformer block showing how we implemented the masking for the ffn and the layernorm using reshapes to ensure no pads were passed through these layers. It shows some of the input shapes for the first masking step in the <code>self.ff_mhsa</code> and <code>self.ln_ff_mhsa</code>.<br>\nThe mhsa block masked attention as normal, and within the ConvModule we just had no masking except a beginning and final step with <code>x = x.masked_fill(~mask_pad, 0.0)</code> - this was a little hacky but seemed to work. If you want I can post the <code>ConvModule</code> in reply - but we will release in the code also. </p>\n<pre><code> (nn.Module):\n     ():\n        (SqueezeformerBlock, self).__init__()\n\n        self.scale_mhsa, self.bias_mhsa = make_scale(encoder_dim)\n        self.scale_ff_mhsa, self.bias_ff_mhsa = make_scale(encoder_dim)\n        self.scale_conv, self.bias_conv = make_scale(encoder_dim)\n        self.scale_ff_conv, self.bias_ff_conv = make_scale(encoder_dim)        \n        self.mhsa_llama = LlamaAttention(LlamaConfig(hidden_size = encoder_dim, \n                                       num_attention_heads = num_attention_heads, \n                                       max_position_embeddings = ))\n        self.ln_mhsa = nn.LayerNorm(encoder_dim)\n\n        self.ff_mhsa = FeedForwardModule(\n                    encoder_dim=encoder_dim,\n                    expansion_factor=feed_forward_expansion_factor,\n                    dropout_p=feed_forward_dropout_p,\n                )\n\n        self.ln_ff_mhsa = nn.LayerNorm(encoder_dim)\n        self.conv = ConvModule(\n                    in_channels=encoder_dim,\n                    kernel_size=conv_kernel_size,\n                    expansion_factor=conv_expansion_factor,\n                    dropout_p=conv_dropout_p,\n                )\n        self.ln_conv = nn.LayerNorm(encoder_dim)\n        self.ff_conv = FeedForwardModule(\n                    encoder_dim=encoder_dim,\n                    expansion_factor=feed_forward_expansion_factor,\n                    dropout_p=feed_forward_dropout_p,\n                )\n        self.ln_ff_conv = nn.LayerNorm(encoder_dim)\n\n     ():\n        \n\n        mask_pad = ( mask).long().().unsqueeze()\n        mask_pad = ~( mask_pad.permute(, ,) * mask_pad)\n        mask_flat = mask.view(-).()\n        bs, slen, nfeats = x.shape\n\n        residual = x\n        x = x * self.scale_mhsa.to(x.dtype) + self.bias_mhsa.to(x.dtype)\n        x = residual + self.mhsa_llama(x, cos, sin, attention_mask = mask_pad.unsqueeze() )[]\n        \n        x_skip = x.view(-, x.shape[-]) \n        x = x_skip[mask_flat].unsqueeze() \n        x = self.ln_mhsa(x) \n\n        residual = x\n        x = x * self.scale_ff_mhsa.to(x.dtype) + self.bias_ff_mhsa.to(x.dtype)\n        x = residual + self.ff_mhsa(x) \n        x = self.ln_ff_mhsa(x) \n        \n        x_skip[mask_flat] = x[].to(x_skip.dtype)\n        x = x_skip.view(bs, slen, nfeats) \n\n        residual = x\n        x = x * self.scale_conv.to(x.dtype) + self.bias_conv.to(x.dtype)\n        x = residual + self.conv(x, mask_pad = mask.().unsqueeze())\n        \n        x_skip = x.view(-, x.shape[-])\n        x = x_skip[mask_flat].unsqueeze()\n        x = self.ln_conv(x)\n\n        residual = x\n        x = x * self.scale_ff_conv.to(x.dtype) + self.bias_ff_conv.to(x.dtype)\n        x = residual + self.ff_conv(x)\n        x = self.ln_ff_conv(x)\n        \n        x_skip[mask_flat] = x[].to(x_skip.dtype)\n        x = x_skip.view(bs, slen, nfeats)  \n\n         x\n</code></pre>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2410544,
              "author_name": "Andy Atkinson",
              "author_url": "",
              "post_date": "2023-08-27T04:23:17.720000",
              "content": "<p>Thank you for posting this.  I can see how you're passing the mask to the attention and conv layers.  It's ok, I can wait for your full code; will give me time to research Squeezeformer and llama attention.</p>\n<p>I used your above feedback on working to get the padded &amp; masked CV to match the non-padded CV. I identified quite a few locations in my model where my mask wasn't working right.  'same' padding in conv1d allowed some of my padding into the kernal, so I instead pre-padded the time axis manually (kernal size -1) and then used the 'valid' conv1d setting as mentioned in the 1st place solution from the last competition.  This keeps pad values out of my kernal when using self.supports_mask = True.  Or worse the mask getting deleted entirely in my positional encoding step…  Also, I was using a squeeze and excite step to weight my channels which was dropping my mask so I switched to a simple dense layer.  Now my non-padded CV matches my padded &amp; masked and i'm on to training a bigger model!  Setting up that CV test was the ticket for me.  Thanks!!!   I will likely update my encoder to this Squeezeformer sometime soon.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2410623,
              "author_name": "Darragh",
              "author_url": "",
              "post_date": "2023-08-27T05:58:13.020000",
              "content": "<p>Nice, glad it works </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2408029,
      "author_name": "bliao",
      "author_url": "",
      "post_date": "2023-08-25T11:44:48.257000",
      "content": "<p>Congratulations to this great work. I have some questions about the implementation:</p>\n<ol>\n<li>Do you use the z coordinate? And how do you normalize the data before augmentation?</li>\n<li><code>Most of the time we only trained and tracked the score of fold0 and not all folds.</code> How do you evaluate your trained model? I.e. which fold is your validation set?</li>\n</ol>",
      "votes": 1,
      "replies": [
        {
          "id": 2408033,
          "author_name": "Darragh",
          "author_url": "",
          "post_date": "2023-08-25T11:50:37.380000",
          "content": "<p>z-coordinate was used. Normalization was performed as below.</p>\n<pre><code>        \n         = x.reshape(x.shape[],,-).permute(,,)\n         = x[~torch.isnan(x)].view(-, x.shape[-])\n         = x - nonan.mean()[None, None, :]\n         = x / nonan.std(, unbiased=False)[None, None, :]\n</code></pre>\n<p>Fold 0 was our validation set when running to check CV. Fold 0 had signers exclusively assigned to it. </p>",
          "votes": 1,
          "replies": [
            {
              "id": 2408072,
              "author_name": "bliao",
              "author_url": "",
              "post_date": "2023-08-25T12:23:24.257000",
              "content": "<p>BTW, how much does ensemble help? Did you try to use one bigger model that is 2x of the ensembled model? </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2408235,
              "author_name": "Darragh",
              "author_url": "",
              "post_date": "2023-08-25T14:16:04.900000",
              "content": "<p>Ensemble of two seeds gave ~0.01+, not sure on the second point, I believe so - you can see in the neptune chart diminishing returns on model size, and more than 14 layers did not help noticeably. <br>\np.s. congrats on your gold!</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2429898,
              "author_name": "JoAI",
              "author_url": "",
              "post_date": "2023-09-08T22:11:40.377000",
              "content": "<p>Thanks so much for the nice writeup. If I remember correctly, in the competition description, it suggested to discard the z axis data, as the depth info may not be accurate. If you were aware of the suggestion, may I know why you still included z axis data? Did you check the score difference between including and excluding z axis data? Thank you!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2410615,
      "author_name": "Rafi Hai",
      "author_url": "",
      "post_date": "2023-08-27T05:53:25.493000",
      "content": "<p>Congratulations and great work! <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> <a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a> <br>\nStaying tuned for code and potential paper 😀</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2409615,
      "author_name": "Kaggle God",
      "author_url": "",
      "post_date": "2023-08-26T10:56:32.467000",
      "content": "<p>absolute GOAT</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2409337,
      "author_name": "Alexey Prikhodko",
      "author_url": "",
      "post_date": "2023-08-26T07:34:17.607000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> and <a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a>!</p>\n<p>I would like to ask a few questions about sign language knowledge because I am a deaf signer of sign language. I work as a junior researcher and developer at our university in the field of sign language recognition, focusing more on Data-Centric AI rather than model-centric AI. In sign language recognition tasks, the main challenge is the lack of sufficient dataset. Your solution based on model-centric AI is impressive, and I fully agree that if I had created a clean dataset myself, your model could have achieved close to 100% accuracy.</p>\n<p>I only started participating in Kaggle about 4 months ago, and when the \"GISLR\" competition came up and beyond, I couldn't even achieve a bronze:) , but I'm learning from your examples and others, and gradually improving my model-centric AI skills.</p>\n<p>I'm not sure if you know sign language. I would like to ask some questions:</p>\n<ul>\n<li>If you were well-versed in sign language, would you approach the solution differently? If yes, how?</li>\n<li>Have you personally studied sign language or interacted with deaf individuals while working on this project?</li>\n</ul>",
      "votes": 2,
      "replies": [
        {
          "id": 2409530,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2023-08-26T09:51:45.763000",
          "content": "<p>Thank you for your questions.</p>\n<blockquote>\n  <p>Have you personally studied sign language or interacted with deaf individuals while working on this project?</p>\n</blockquote>\n<p>I did not know any sign language before both competitions, but during competitions I also look at the specific domain and try to understand why a model behaves as it does, with respect to mistakes. I went through the different characters on a high level, just to get an understand what is needed to spell a character (single hand? two hands? more?) and how letters, digits and special characters are different. I think understanding of the subtleties helps in model design, because ideally the model is capable at addressing those. Here, understanding fingerspelling and how mediapipe works helped coming up with meaningful augmentations that help the model generalize to new signers. For letters I even spent an hour or two training signing skills via this online game (its quite fun)<br>\n<a href=\"https://www.signlanguageforum.com/asl/fingerspelling/fingerspelling-game/\" target=\"_blank\">https://www.signlanguageforum.com/asl/fingerspelling/fingerspelling-game/</a></p>\n<blockquote>\n  <p>If you were well-versed in sign language, would you approach the solution differently? If yes, how?</p>\n</blockquote>\n<p>I dont think so. I often approach problems iteratively. Train a model, look at the predictions and understand if I can find mistakes that are domain related and can be prevented by a better model design (or by improving data quality). It would be easier to understand the mistakes if I would be well-versed in sign language but the approach would be the same</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2409594,
              "author_name": "Alexey Prikhodko",
              "author_url": "",
              "post_date": "2023-08-26T10:41:43.007000",
              "content": "<p>Thank you for your response. If in-depth knowledge of sign language is required for machine learning purposes, feel free to let me know—I'd be happy to assist you :)</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2409094,
      "author_name": "Hossam Nasr",
      "author_url": "",
      "post_date": "2023-08-26T04:10:24.363000",
      "content": "<p>Congrats , A beautiful and clean solution.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2409059,
      "author_name": "S.Inamura",
      "author_url": "",
      "post_date": "2023-08-26T03:39:56.563000",
      "content": "<p>Congratulations!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2409054,
      "author_name": "Muhammad Usman",
      "author_url": "",
      "post_date": "2023-08-26T03:37:27.937000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> </p>\n<p>Great solution man!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2408993,
      "author_name": "Claudio Sturaro",
      "author_url": "",
      "post_date": "2023-08-26T02:07:46.250000",
      "content": "<p>Amazing, congratulations!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2408969,
      "author_name": "Nathan Barkdull",
      "author_url": "",
      "post_date": "2023-08-26T00:46:12.383000",
      "content": "<p>Congratulations! Thanks for the write-up, great learning resource</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2408408,
      "author_name": "dragon zhang",
      "author_url": "",
      "post_date": "2023-08-25T16:09:12.233000",
      "content": "<p>great.  Very appreciate it for sharing.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2408382,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2023-08-25T15:52:44.107000",
      "content": "<p>The diagram is always on point. Congratulations!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2408011,
      "author_name": "gezi",
      "author_url": "",
      "post_date": "2023-08-25T11:32:35.710000",
      "content": "<p>Excellent work! So many brilliant ideas and you noticed nearly every corner details. While most teams using ctc encoder only method, the seq2seq win 1st place😀 , and also seems using seq2seq we can better ensemble.<br>\nI have used public seq2seq notebook and also found supplement dataset not help, maybe related to seq2seq method and supplement dataset phrase distribution is very different comparing to train dataset. But for ctc encoder based method I found supplement dataset help improve a lot (more then 10 points).<br>\nFor fp16, I did not notice acc diff if train using fp32+awp then infer with converter.target_spec.supported_types = [tf.float16]. Not sure if there are other differences with your setting.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3125013,
      "author_name": "Durgesh Tambe",
      "author_url": "",
      "post_date": "2025-02-15T17:26:37.967000",
      "content": "<p>Congrats! I learned a lot from your solution.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2951253,
      "author_name": "Kyaw Saw Htoon",
      "author_url": "",
      "post_date": "2024-08-08T13:00:56.660000",
      "content": "<p>I'm a little late to the game here but congrats! </p>\n<p>I'm trying to test your model with live videos and found that the model's output shape is 11, 63. I think 11 represents the sequence of characters. I'm confused about 63 since the character to prediction index only contains 58 characters. Can you help me understand, pleaes? Thanks in advance.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2643460,
      "author_name": "Emre Eren",
      "author_url": "",
      "post_date": "2024-02-08T21:00:28.420000",
      "content": "<p>Good job and best wishes to your job</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2629495,
      "author_name": "Joe DevonAI",
      "author_url": "",
      "post_date": "2024-01-31T20:16:43.167000",
      "content": "<p>Congratulation <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a>! This is awesome!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2598876,
      "author_name": "Anushka 298",
      "author_url": "",
      "post_date": "2024-01-12T14:40:07.730000",
      "content": "<p>when will the research paper be out?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2586254,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-04T04:52:45.733000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2573281,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-12-24T21:03:09.747000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2553300,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-12-08T06:17:56.640000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2477681,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-10-11T13:46:13.807000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2476730,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-10-10T22:17:10.940000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2431948,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-09-10T14:08:32.380000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2431029,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-09-09T18:22:34.983000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2428188,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-09-07T17:27:53.983000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2419940,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-09-02T09:27:24.797000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2419645,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-09-02T06:48:44.523000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2419149,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-09-01T18:03:44.647000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2419012,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-09-01T16:33:51.303000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2419243,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-09-01T20:50:37.860000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2418427,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-09-01T09:20:13.430000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2416427,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-08-31T03:28:48.943000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2416351,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-08-31T01:32:14.870000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2416066,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-08-30T19:04:57.600000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2418857,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-09-01T14:56:30.910000",
          "content": "",
          "votes": 1,
          "replies": [
            {
              "id": 2423268,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-09-04T14:38:01.770000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2414448,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-08-29T16:28:29.840000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2413796,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-08-29T07:02:53.900000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2413825,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-08-29T07:27:32.717000",
          "content": "",
          "votes": 2,
          "replies": [
            {
              "id": 2413855,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-08-29T07:42:10.283000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2413022,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-08-28T16:29:32.157000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2412870,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-08-28T14:52:29.207000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2413175,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-08-28T17:40:42.750000",
          "content": "",
          "votes": 1,
          "replies": [
            {
              "id": 2414092,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-08-29T11:46:26.003000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2412353,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-08-28T08:29:24.480000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2412104,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-08-28T05:04:00.737000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2412098,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-08-28T05:00:31.753000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2411959,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-08-28T02:07:21.873000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2411248,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-08-27T13:54:17.497000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2411798,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-08-27T21:35:23.137000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2410960,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-08-27T09:57:44.010000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2410861,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-08-27T08:51:31.330000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2410906,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-08-27T09:20:24.790000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2410909,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-08-27T09:21:28.050000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2410754,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-08-27T07:19:46.360000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2425345,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-09-05T20:12:06.553000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2420637,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-09-02T18:53:44.223000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2418751,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-09-01T13:49:34.727000",
      "content": "",
      "votes": -1,
      "replies": []
    },
    {
      "id": 2409074,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-08-26T03:59:02.080000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2532299,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-11-20T23:29:56.597000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2407974": "Thanks to kaggle and everyone involved for hosting such an interesting competition. It was a great extension to the isolated sign language classification and it was very interesting to see how much of speech-to-text research could also be applied to sign language fingerspelling. As always it was a great teaming experience with @darraghdog \n\n## TLDR\nOur solution is based on a single encoder-decoder architecture. The encoder is a significantly improved version of Squeezeformer, where the feature extraction was adapted to handle mediapipe landmarks instead of speech signals. The decoder is a simple 2-layer transformer. We additionally predicted a confidence score to identify corrupted examples which can be useful for post-processing. We also introduced efficient and creative augmentations to regularize the model, where the most important ones were CutMix, FingerDropout and TimeStretch, DecoderInput Masking. We used pytorch for developing and training our models and then manually translated model architecture and ported weights to tensorflow from which we exported to tf-lite.\n\n## Cross validation\nWe split the training data into 4 folds by signer. In the beginning we had nearly perfect correlation between CV and public LB with this approach. With higher scores improvements on CV reflected a bit less on LB, mostly due to the fact that the LB score was always decently higher and hence saturated earlier. Most of the time we only trained and tracked the score of fold0 and not all folds. \n\n## Data preprocessing\nIn total 130 key points were used. These consisted of 21 key points from each hand, 6 pose key points from each arm, and the remaining 76 from the face (lips, nose, eyes). Locally the 130 key points were cached to .npy files for fast data loading. \nPrior to data augmentations, the data was normalized with std/mean and nans were zero filled. \n\n## Augmentations\nAugmentations were essential to prevent overfitting, generalize to new signers and enable deep models. We used augmentations which were popular in the first ASL competition but also came up with a lot of new creative augmentations and some have proven to be very effective. \n\n- Resizing along time axis.\n- Shift the sequence along the time axis. \n- Windowed resizing along time axis (similar to warping).\n- Left-right flip of keypoints. \n- Cutmix of samples timewise - draw a random percentage between [0,1] and cut 2 sequences and related phrases at that percentage and mix. Mixing only within same signer was best\n- Spatial affine - scale, shear, shift and rotate. \n- Drop/Zero-fill between 2 to 6 different fingers over 2 to 3 time windows. \n- Drop/Zero-fill either all face landmarks or all pose landmarks. \n- In rare cases (~5% of samples) drop/zero-fill all hand landmarks. \n- Temporal masking (zero-fill) in windows of different sizes or counts. \n- Spatial masking \n\nMost augmentations were applied to 50% of the samples, except for resizing and spatial affine which were applied to ~80% of samples.  \n\nAfter augmentation, samples with more than 384 frames were resized along time axis, with channel-wise linear interpolation. Samples of less than 384 were padded to 384 for training only. tf-lite ran on variable length samples. \nNo frames were dropped in preprocessing. \n\n## Model\nIn general, we observed that a deeper model gives significant gains (if we are able to prevent overfitting). As a consequence not only regularization techniques like augmentations are essential, but also every improvement in computational efficiency creates space to use deeper models and hence is equally important as the model architecture itself.\n\nOur model consists of 3 parts, Feature Extraction, Encoder, Decoder which are shown in the image below\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2Fdc7c0fc365f54fb4d35f03da1748ed3b%2FScreenshot%202023-08-25%20at%2012.47.53.png?generation=1692960500706760&alt=media)\n\nWe interpret the data like a 3 channel image, where width is defined by the number of frames, height is given by the number of the selected 130 landmarks and channels are given by raw xyz coordinates. \nThe feature extraction is based on a 2D convolution followed by batchnorm and a linear layer on the flattened features to extract features per frame. We have 5 of this feature extraction modules, one for all landmarks at once and one per landmark type (left_hand, right_hand, face, pose). The all-landmark module outputs 208 dim vector/ frame. The other 4 output 52 dim vectors each which are then concatenated to have also 208 dims. We had those two 208-dim vectors per frame and get a (batch_size x 384 x 208) input for our encoder where 384 is the maximum sequence length we chose.\n\nThe main component of our model is an encoder which was adapted from the Squeezeformer architecture. We did not use the actual “squeeze” idea, i.e. a temporal Unet, but used the general architecture of Squeezeformer Blocks which consist of a combination of MultiHeadSelfAttention (MHSA), Convolution and FeedForward modules. We made several improvements to this architecture:\nASR conformers (and Squeezeformer) use relative positional encoding which allow the self-attention module to generalize better on different input lengths. Relative positional encoding is performance intensive, as well as using many parameters, as they are stored separately in each layer. Replacing this with Llama attention which uses rotary embeddings sped up training ~2X and tf-lite inference approx ~3X allowing larger models to be used. In addition, we cached the rotary embeddings once and fed them into each layer with the input data so they are not duplicated in each layer. This resulted in 20% less parameters in the model. We saw no benefit in using time reduction which was introduced with Squeezeformer. So in our model all layers had the same sequence length as the original input. As suggested in the Squeezeformer paper, the pre-Layer Norm from the Macaron structure is redundant and was replaced with a learnable scaling layer which scales and shifts the activations. \n\nFor decoding we used a simple 2 layer transformer decoder which is similar to hugging faces [Speech2TextDecoder](https://github.com/huggingface/transformers/blob/main/src/transformers/models/speech_to_text/modeling_speech_to_text.py#L857), which outputs a sequence prediction. We then used cross entropy loss for training our model End2end. An extra cross entropy auxiliary loss of the reversed sequence was used. A causal decoder mainly uses encoder cross attentions for the sequence's beginning and previous characters and cross attentions for the end. To improve the model's accuracy, we use a separate causal decoder on the reversed sequence as an auxiliary loss, making the model rely more on the encoder cross attention for the label's end. For decoder inference early stopping and past key value caching was used which sped up inference significantly.  It should be noted that we found a transformer based Decoder superior to a CTC based decoding, even in a setting where computational efficiency matters a lot. \n\nAdditionally to the decoder, we also added a single linear layer to take the features of the first token of the encoder output to predict a confidence score, which helps to identify garbage data and can be used in post-processing. As a target for this we used normalized levensthein distance clipped to [0,1] of OOF predictions of a decent previous model.\n\nWe explored and trained all our models in pytorch, but whenever we deemed it good enough for a submission we translated each component manually to tensorflow and ported weights from our pytorch models. \n\n## Training procedure\nModels were trained with a cosine learning rate schedule for 400 epochs with peak LR of 0.0045, weight decay of 0.08, 10 epochs warmup, mixed precision and an effective batch size of 512 samples. Dropout of 0.1 was used in the transformer encoder/decoder layers. It was important to train with mixed precision in order to leverage fp16 inference without performance drop. Training with fp32 and using tf-lite fp16 inference causes a drop of ~0.01 in CV vs LB. \n\ntf-lite inference used a single sample without padding, while model training was performed with time padded mini-batches. To avoid the model learning with pads, and ensure optimal inference runtime, the feature extractor and Macaron structure encoder layers were masked time wise during training. This needed to be manually implemented in pytorch on each layer. This took some effort, but paid off by significantly speeding up inference. As [1st place team](https://www.kaggle.com/competitions/asl-signs/discussion/406684) in the previous ISL competition explained, this is much easier to do in tensorflow as keras has an off the shelf masking layer. \n\nModel training and tf-lite inference ran in fp16, so tf-lite files consumed almost half the disk size of of fp32. This was important, as the 40MB size was our limitation in the end. The two final model seeds measured 39988kb. \n\n## Postprocessing\nThe main idea of our postprocessing is to replace poor predictions with a dummy phrase which has a small levensthein distance to the train/ test data. @anokas showed in his [notebook](https://www.kaggle.com/code/anokas/static-greedy-baseline-0-157-lb) why '2 a-e -aroe' is a good candidate for that. Most of the poor prediction resulted from corrupted input data often only a few frames long. We used a confidence score predicted by our model as basis. Whenever the confidence score is below 0.15 or the sequence is shorter than 15 frames we replace the prediction with '2 a-e -aroe'.\n\n## Supplemental Data\nWe only marginally profited from using the supplemental data. We think the main reason is that although there are 50k samples in this supplemental data there are only 500 unique phrases, and hence the model rather learns to classify then to actually decode character-by-character. We tried a lot of approaches but only the following one gave a small boost (0.838 -> 0.839): \nFirst we group the supplemental data by phrase, which only leaves us with 500 groups. In each epoch of training we add one sample per group to the training dataset for our model. That means in each epoch we use 50k samples of the training data and only 500 samples of the supplemental data. \n\n## Ensembling\n\nOur final submission is a 2-seed ensemble of our model trained on the complete training data (fullfit). We average resulting logits in each decoding step for ensembling.\n\n## What did not help\n\n- Fully using supplemental data\n- Using edit distance as loss (tried different approaches)\n- CTC loss (even as an auxiliary loss it hurt score)\n- Label smoothing\n- AWP - kept getting nans with FP16\n- TTA (flip/stretch)\n- Mixup of hidden layers & Specaugment++\n- Beam search decoding (too costly)\n\n## Ablation study (roughly)\n\n#### Augmentations\n- Cutmix +0.005\n- FingerDropout +0.005\n- Face/PoseDropout +0.005\n- masking decoder inputs +0.003\n\n#### Model improvements\n- CNN Feature extraction +0.005\n- 2-branch Feature extraction with indiv norm +0.003\n- Squeezeformer over 1stplace Net of 1 round +0.005\n- Decoder over CTC +0.003\n- Confidence over simple rules for post-processing +0.002\n\n\n#### Efficiency Improvements:\n- Deeper model due to fp16 +0.003\n- Deeper model due to llama attention +0.003\n- Deeper model due to masking/ variable sequence len +0.005\n- Deeper model due to caching/ early stopping in decoder +0.005\n\n#### Postprocessing\n- Replace bad predictions with dummy phrase +0.006\n\n\n## Used tools/ repos\n- Pytorch/Tensorflow/Tf-lite (no onnx this time)\n- Huggingface\n- Albumentations (adapted their framework for using things like OneOf or Compose, but wrote our own augmentation implementations)\n- Neptune.ai was our MLOps stack to track compare and share models. Below are example training runs using different model parameters and hardware (a100 card vs kaggle kernel). The data loading in the kaggle kernel below was slow and could probably be sped up with some work.  \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2F4693e3379ea6a1631bb3cb7b9044d42a%2FScreenshot%202023-08-25%20at%2012.58.07.png?generation=1692961107306495&alt=media)\n\n## Code & model weights\n\nhttps://github.com/ChristofHenkel/kaggle-asl-fingerspelling-1st-place-solution\n\n## Paper\ntbd.  Due to the novelty of our approach we are thinking about summarizing it in a paper.\n\n**Thank you for reading, questions welcome**",
    "2410650": "Congratulations @christofhenkel and @darraghdog \n\nAppreciate your effort! 🙏  \n             ",
    "2408132": "Super congrats @darraghdog and @christofhenkel for this huge win! Thanks for sharing such detailed solution.",
    "2407993": "> In general, we observed that a deeper model gives significant gains (if we are able to prevent overfitting).\n\nI agree. For me, 24L 256d was better than 9L 384d, even though they have similar sizes.",
    "2408367": "Congratulations @christofhenkel and @darraghdog \nA beautiful and clean solution. I am sure you tried a ton of different ideas how to encode these features.",
    "2408273": "I'm going to post a little comment. It might get long. Hopefully it's helpful for someone.\n\nFirst, Wow. I just saw Greysnow's comment saying he feels like a child... coming in 12th. Here I am 386 by copying others' work... if he's a child, I am... Anyway, really great writeup and lots of great ideas. I also look forward to seeing all the code.\n\nWell, I entered this competition as something to do and to run a couple of experiments. I've gone through a decent amount of the general learning stuff on kaggle. I don't have a background in coding (have a PhD in something else, so you know, I can probably learn things), but I love seeing what people can do in this arena and keep exploring whether I can truly get into it. I feel like I should be somewhat embarrassed to say chatGPT was my guide, and that was a big part of my experiment-- how much could this newer tool help someone as clueless as me to do anything useful. On the one hand, amazing what it could help me accomplish, but also as I think most of you know, it doesn't always take the most elegant approach, and sometimes after a few iterations of changes and errors, I'd hunker down and really try on my own and figure out how to fix the problem myself faster. My main goals were just to see how much of this I could understand, and when I had an idea, could I figure out how to implement it, all with the hope that at the end of the contest when the top solutions are posted, I could really grasp it at the least (not sure I'm there, but I've made progress). As I've posted elsewhere, I was curious whether the winner is really elegant and like totally new solutions, and/or brute force with systematic hyperparameter tuning, taking it offline to run near-infinite variations on some other dedicated hardware, etc.  I'm not sure about the offline dedicated hardware part, but I think from this writeup, there's some of both.\n\nWell I do have a few general questions that I'd like to understand better. Some of the first stuff I did with other people's code was try to simplify down to only hand landmarks. I didn't and generally still don't see what face landmarks could be helpful for in a fingerspelling dataset. I suppose it's not that I can't imagine that perhaps on a pause as someone repeats a letter, that there might be an associated facial expression the model could pick up-- but it feels quite minimal to me, and I thought there would be more benefit to not having to look at so much extra information that is mostly useless. I did come to think that some of the pose landmarks are probably helpful with regard to motion letters like \"z\" and perhaps other things-- though it still seems like a good model should be able to get all the info it needs from just the hand. I also tried ignoring the z coordinates, and I tried flipping one of the hands, thinking that it made more sense for the model to learn a single set of signs and motions, where the opposite hand ultimately takes a flipped orientation. I was not super systematic with recording how a single change affected loss or score, but more often than not whatever change I attempted gave worse performance. Trying to get rid of lhand and rhand with an xy flip of one (while I was ignoring z) definitely hurt, which surprised me. \n\nI tried adding a displacement feature by calculating the, like Euclidean norm-change from frame to frame. I was hopeful that this might help the model with double letters and spaces, which it seemed to me from output was an area of trouble, maybe unsurprisingly. However, after much effort to finally successfully implement the addition of this feature and get it all to work together, again, it hurt model performance. Maybe that makes perfect sense to people who understand better how transformers work? As I understand it, it's never entirely clear whether adding a new feature will really help a model to focus on something, or will just be redundant information. I wondered, but did not try, whether trying to identify some like biggest changes, and creating new features based on that, might help-- like maybe calculating the distance from thumb to pinky as a new feature, just as potential example. Again, is that just likely to be redundant, or does the idea not work well with the transformer architecture?\n\nI tried separating out training and validation datasets based on the participant_id, thinking it might help if the model trained on one signer and their idiosyncrasies at a time-- well and I also tried batching by signer, I guess that's really the training on one signer at a time. Didn't help.\n\nI did what others have done to make little gifs out of the signing just so I could visualize it, and I also played around with adding to the dataset by recording new landmark data from YouTube videos (and even from myself, just to do it), but it wasn't easy for me to think of how I might systematically add to the dataset without being really manual about it. That's still a question that I'm curious about and haven't answered for myself: essentially what is \"enough\" data? If you take this winning model and you train on 20% of the data, how will that compare to using 80% of it- and if we kept increasing the quantity of data here, would we continue to see improvement? \n\nI also considered augmenting the data with little rotations, but I didn't get around to implementation. I'm intrigued by the augmentation of dropping fingers and such\n\nI used greysnow's notebook and code to get things working on TPU when I ran out of GPU time, but man it was painful. It seems every one of those efforts I made above added a layer of difficulty if I tried to run it on TPU. I can only imagine the use of PyTorch and translation efforts in this notebook were quite more work.\n\nSo-- a general question. While it's amazing to watch the model work and create useful predictions, the 0.84 score means, I think, that on an average 11-12-character phrase, there's nearly 2 changes necessary to get to the right phrase. From looking at the output, I think it's actually the case that you get a decent amount of perfect translation, but there are also a number that are far off. Anyway, this is the best of the bunch here, it's definitely impressive, but is it objectively \"good\"? Is it only the starting point for developing a really useful application? I don't have the context to understand what threshold must be met for people to actually want to use it. This dataset includes garbage data, which is dealt with in this notebook and others, but aside from the truly garbage data, I would think if you took a few ASL-experienced individuals and had them watch the original signing, or perhaps even the mediapipe-processed stick figures, they would be able to achieve very nearly 100% accuracy. If yes, then I'd like to think we could train models to do just as well. And in a way there's a disconnect for me in what a human can infer primarily from watching for known symbols compared with the way these models work that are not in any way informed what an \"A\" or \"B\" looks like. I also explored the idea of feeding this information in some way, but I don't know that there's a good way to do that. A different question to ask is if there were not limitations made such that this can be a TFlite model or if we could throw more GPUs and much longer training time, etc., could we then pretty easily bump this accuracy up?\n\nNow that the competition is over, do people tend to tweak further using ideas from other notebooks? Basically I wonder if someone right now could take the best ideas from different notebooks, could we quickly get to 0.9? Or if this solution has already maxed out on what's possible given the constraints. \n\nWell, thank you for posting the writeup, congratulations on the good work, and I'm interested to learn and understand more.",
    "2408048": "Wow, there is so much high-level work here to learn from. I eagerly await your code. Compared to your solution, I feel like a child 😅",
    "2447454": "Congratulations! Your solution is really enlightening to me! Especially the part related to Augmentations.",
    "2411226": "great work and best wishes for your job",
    "2411022": "Congrats! Excited to see the code!",
    "2409932": "This competition is about fingerspelling. However, your model uses 130 data points out of which only 27 points belong to fingers / fingerspelling. How much of the accuracy of your model is contributed by the data points related to face? \n\nIn the case of 0.697 public notebook based on winning solution of the sign language competition, I found that the accuracy falls down to the extent of 0.02 if we just drop the lip related data points alone. So, how much of the increased accuracy of your model is contributed by additional 70 face related datapoints vs accuracy contributed by the improved model architecture?",
    "2408298": "Congratulations for topping the LB.\nThank you for sharing details of your model. Its indeed a very innovative approach.\nCan you also share the running time of your notebook. Did you separate the training\nand inference notebooks?",
    "2408032": "\" To avoid the model learning with pads, and ensure optimal inference runtime, the feature extractor and Macaron structure encoder layers were masked time wise during training.\"\n\ni have a question. Besides speedup, is there a change in accuracy?\ne.g. whisper asr model doesn't not use masking\n\n\"The Whisper feature extractor performs two operations. It first pads/truncates a batch of audio samples such that all samples have an input length of 30s. Samples shorter than 30s are padded to 30s by appending zeros to the end of the sequence (zeros in an audio signal corresponding to no signal or silence)\"\n\nit seems possible to do without masking? thanks",
    "2408029": "Congratulations to this great work. I have some questions about the implementation:\n\n1. Do you use the z coordinate? And how do you normalize the data before augmentation?\n2. `Most of the time we only trained and tracked the score of fold0 and not all folds.` How do you evaluate your trained model? I.e. which fold is your validation set?\n",
    "2410615": "Congratulations and great work! @christofhenkel @darraghdog \nStaying tuned for code and potential paper 😀",
    "2409615": "absolute GOAT",
    "2409337": "Congratulations @christofhenkel and @darraghdog!\n\nI would like to ask a few questions about sign language knowledge because I am a deaf signer of sign language. I work as a junior researcher and developer at our university in the field of sign language recognition, focusing more on Data-Centric AI rather than model-centric AI. In sign language recognition tasks, the main challenge is the lack of sufficient dataset. Your solution based on model-centric AI is impressive, and I fully agree that if I had created a clean dataset myself, your model could have achieved close to 100% accuracy.\n\nI only started participating in Kaggle about 4 months ago, and when the \"GISLR\" competition came up and beyond, I couldn't even achieve a bronze:) , but I'm learning from your examples and others, and gradually improving my model-centric AI skills.\n\nI'm not sure if you know sign language. I would like to ask some questions:\n- If you were well-versed in sign language, would you approach the solution differently? If yes, how?\n- Have you personally studied sign language or interacted with deaf individuals while working on this project?",
    "2409094": "Congrats , A beautiful and clean solution.",
    "2409059": "Congratulations!",
    "2409054": "Congratulations @christofhenkel \n\nGreat solution man!",
    "2408993": "Amazing, congratulations!",
    "2408969": "Congratulations! Thanks for the write-up, great learning resource",
    "2408408": "great.  Very appreciate it for sharing.",
    "2408382": "The diagram is always on point. Congratulations!",
    "2408011": "Excellent work! So many brilliant ideas and you noticed nearly every corner details. While most teams using ctc encoder only method, the seq2seq win 1st place😀 , and also seems using seq2seq we can better ensemble.\nI have used public seq2seq notebook and also found supplement dataset not help, maybe related to seq2seq method and supplement dataset phrase distribution is very different comparing to train dataset. But for ctc encoder based method I found supplement dataset help improve a lot (more then 10 points).\nFor fp16, I did not notice acc diff if train using fp32+awp then infer with converter.target_spec.supported_types = [tf.float16]. Not sure if there are other differences with your setting.",
    "3125013": "Congrats! I learned a lot from your solution.",
    "2951253": "I'm a little late to the game here but congrats! \n\nI'm trying to test your model with live videos and found that the model's output shape is 11, 63. I think 11 represents the sequence of characters. I'm confused about 63 since the character to prediction index only contains 58 characters. Can you help me understand, pleaes? Thanks in advance.",
    "2643460": "Good job and best wishes to your job",
    "2629495": "Congratulation @christofhenkel! This is awesome!",
    "2598876": "when will the research paper be out?\n",
    "2586254": "Congratulations",
    "2573281": "Congrats! thank you for this great work and explanation about workflow, model arch, augmentation, and all other things. This solution give a lot of insights for my graduation, so I am very happy to see like this level of work and to learn from, ty for sharing🙏🙏.",
    "2553300": "Amazing \nLearned a lot from your write-up\n",
    "2477681": "This is great work. Congrats! Excited to see the code!",
    "2476730": "Just came across this, what absolute gem. Thanks for sharing it!",
    "2431948": "Congratulations. Indeed a very thorough approach to this. ",
    "2431029": "Awesome... Explanation  \n\ncongratulations for your rank🥇",
    "2428188": "Great explanation! Very informative and has good insights. Also, congratulations on your rank! 🎉 #Informative #Congratulations #Insights",
    "2419940": "Wow, Congratulations <3 ",
    "2419645": "great ideas and work",
    "2419149": "felicidades",
    "2419012": "Thanks again for all your helpful comments.\n\nAny tips to avoid getting nan loss with mixed precision fp16 training?  I note you avoided AWP.  I see in tensorflow's instructions they say to convert the last softmax to float32.  I also note that you trained in pytorch and possibly that has more support in skipping nan batches or something like that.  I've been using model.fit in tensor flow which does some loss scaling to avoid underflow.  I get nan's at about the 5th or 6th epoch.  The main reason I'm interested in fp16 is because I'm running out of memory with a modest sized squeezeformer on a 3090 with a frame length of only 300 batch size 64, and I'm interested in increasing parameters.  There could be some clumsy implementation going on with my squeezeformer, but the CV is looking good so it can't be too far off. I can go to batch size of 32... but I'll be training for quite a while.",
    "2418427": "Congratulations, \nAppreciate your work.",
    "2416427": "Congratulations!",
    "2416351": "Wow.....\nCongratulations on your achievement\n",
    "2416066": "I'm curious to know more about your confidence score calculation. How did you calculate the loss and how did you combine it with the cross entropy of the decoder for training?",
    "2414448": "good content  i like this thanku",
    "2413796": "What is the meaning of OOF in this context?\n\nThank you for the explanation, I am extracting a lot of knowledge from it.",
    "2413022": "Great effort sir!",
    "2412870": "@christofhenkel @darraghdog\nCongratulations :tada: and thank you for the great description!\n\nI have a question (might be for beginners).\n* How and when did you decide hyper parameters of learning schedule?\n\nThe peak LR (0.0045) and weight decay (0.08) look pretty fine grained. I'd love to know how these parameters are tuned (e.g., Peak LR is tested at 0.0005 intervals for each model) and when these are decided (e.g., tuned peak LR even when Augmentations, or fixed for each model at first, or fixed for some good model).",
    "2412353": "Amazing, congratulations!",
    "2412104": "Congratulations and thank you for sharing your solution.",
    "2412098": "Congratulations!\nThanks for sharing solution!",
    "2411959": "Congratulations! Thanks for knowledge sharing!",
    "2411248": "Congratz on your win and thanks for the great description. Your architecture changes, especially the parameter saving and data transforms seem really, really well done!\nTwo things I didn't see:\n1. What dimensions did you use in the encoder in the end? The screenshot shows 128, 168 and 192.\n2. What kind of hardware did you use to train?  \n\nThanks!",
    "2410960": "Congratulations ",
    "2410861": "By the way @christofhenkel, did you use https://github.com/AlexanderLutsenko/nobuco for helping with the PyTorch to TF translation or you did it solely by hand?",
    "2410754": "Amazing and professional approach to find the solution. @christofhenkel ",
    "2425345": "",
    "2420637": "",
    "2418751": "",
    "2409074": "Thanks for new learning :)",
    "2532299": "Thanks for sharing."
  }
}