{
  "id": 434415,
  "title": "[5th place solution] Vanilla Transformer, Data2vec Pretraining, CutMix, and KD",
  "url": "/competitions/asl-fingerspelling/discussion/434415",
  "author_name": "Jungwoo Park",
  "post_date": "2023-08-25T06:15:16.515000",
  "votes": 51,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Thanks to Kaggle and competition hosts for this incredible and meaningful competition. Here is my solution.</p>\n<p>The code is available at <a href=\"https://github.com/affjljoo3581/Google-American-Sign-Language-Fingerspelling-Recognition\" target=\"_blank\">the github repository</a>.</p>\n<h2>TL;DR</h2>\n<ul>\n<li>Using hands, pose, and lips landmarks with 3D-based strong augmentation</li>\n<li>Vanilla transformer with conv stem and RoPE</li>\n<li>Data2vec 2.0 pretraining</li>\n<li>CTC segmentation and CutMix</li>\n<li>Knowledge distillation</li>\n</ul>\n<h2>Data Processing</h2>\n<p>I used 3d landmark points instead of 2d coordinate because it looks like applying rotation augmentation on 3d space makes models to be more robust. According to <a href=\"https://developers.google.com/mediapipe/api/solutions/java/com/google/mediapipe/tasks/components/containers/NormalizedLandmark\" target=\"_blank\">the mediapipe documentation</a>, the magnitude of z uses roughly the same scale as x. Since x and y are normalized by width and height of camera canvas and they are recorded on smartphone devices, I have to denormalize the landmarks with original aspect ratio to apply correct rotation transform. I simply estimate the aspect ratio by solving affine matrix that maps to the normalized hands at each frame to the standard hand landmarks and get scale factors from the affine matrix. The average aspect ratio is <code>0.5268970670149133</code> and I simply multiply <code>1.8979039030629028</code> to the normalized y values, i.e., <code>points *= np.array([1, 1.8979039030629028, 1])</code>.</p>\n<p>Correct Rotation Transform:<br>\n$$X' = SRX$$</p>\n<p>Wrong Rotation Transform:<br>\n$$X' = RSX$$</p>\n<p>It's important to note that applying rotation after normalizing coordinates is incorrect. Denormalize, rotate, and then normalize again. Actually, I did not normalize again because it is not necessary.</p>\n<p>For inputs, I utilized landmarks from the left hand, right hand, pose, and lips.  Initially, I focused solely on hand landmarks for the first few weeks and achieved a public LB score of 0.757. I believed fingerspelling is totally related to hand gestures only. Surprisingly, incorporating auxiliary landmarks such as pose and lips leads better performance and helped mitigate overfitting. The inclusion of additional input points contributed to better generalization.</p>\n<p>Here's a snippet of the data augmentation code I employed for both pretraining and finetuning:</p>\n<pre><code>Sequential(\n    LandmarkGroups(Normalize(), lengths=(, , )),\n    TimeFlip(p=),\n    RandomResample(limit=, p=),\n    Truncate(max_length),\n    AlignCTCLabel(),\n    LandmarkGroups(\n        transforms=(\n            FrameBlockMask(ratio=, block_size=, p=),\n            FrameBlockMask(ratio=, block_size=, p=),\n            FrameBlockMask(ratio=, block_size=, p=),\n        ),\n        lengths=(, , ),\n    ),\n    FrameNoise(ratio=, noise_stdev=, p=),\n    FeatureMask(ratio=, p=),\n    LandmarkGroups(\n        Sequential(\n            HorizontalFlip(p=),\n            RandomInterpolatedRotation(, np.pi / , p=),\n            RandomShear(limit=),\n            RandomScale(limit=),\n            RandomShift(stdev=),\n        ),\n        lengths=(, , ),\n    ),\n    Pad(max_length),\n)\n</code></pre>\n<p>The total number of input points is 75. Each component consists of 21, 21, 14, and 40 points, respectively. The hand with more <code>NaN</code> values is discarded and only the dominant hand is selected. As seen in the snippet above, spatial transformations are applied separately to the components. The landmark groups are first centered and normalized by maximum x-y values of each group. Note that z-values can exceed 1.</p>\n<h2>Model Architecture</h2>\n<p>I employed a simple Transformer encoder model similar to ViT. I used PreLN and rotary position embeddings. I replaced ViT's stem linear patch projection with a single convolutional layer. This alternation aimed to enable the model to capture relative positional differences, such as motion vectors, from the first convolutional layer. With the application of rotary embeddings, there is no length limit and also I didn't truncate input sequences at inference time. Similar to other ViT variants, I also integrated LayerDrop to mitigate overfitting.</p>\n<h2>Data2vec 2.0 Pretraining</h2>\n<p>Given that the inputs consist of 3d points and I used a normal Transformer architecture which has low inductive bias toward data attributes, I guessed it is necessary to pretrain the model to learn all about data properties. The <a href=\"https://arxiv.org/abs/2212.07525\" target=\"_blank\">Data2vec 2.0</a> method, known for its remarkable performance and efficiency across various domain, seemed promising for adaptation to landmark datasets.</p>\n<p><img src=\"https://i.ibb.co/KLM8LYh/FigureA.png\" alt=\"FigureA\"></p>\n<p>According to the paper, using multiple different masks within the same batch helps convergence and efficiency. I set $M = 8$ and $R = 0.5$, which means 50% of the input sequences are masked and there are 8 different masking patterns. After I experimented many various models, and I arrived at the following final models:</p>\n<ul>\n<li>Transformer Large (24L 1024d): 109 epochs (872 effective epochs)</li>\n<li>Transformer Small (24L 256d): 437 epochs (3496 effective epochs)</li>\n</ul>\n<p>Termination of the overall trainings was determined by training steps, not epochs, resulting in epochs that are not multiples of 10. After pretraining the model, the student parameters are used for finetuning.</p>\n<h2>CTC Segmentation and CutMix</h2>\n<p>Before explaining the finetuning part, it is essential to discuss CTC segmentation and CutMix augmentation. Check out <a href=\"https://github.com/lumaku/ctc-segmentation\" target=\"_blank\">this repository</a> and <a href=\"https://pytorch.org/audio/main/tutorials/forced_alignment_tutorial.html\" target=\"_blank\">this documentation</a> which provide information about CTC segmentation. To summarize, a CTC-trained model can detect the position of character appearances, enabling the inference of time alignment between phrases and landmark videos.</p>\n<p>Initially, I trained a Transformer Large model and created pseudo aligned labels. Using the alignments, I applied temporal CutMix augmentation which cuts random part of the original sequence and inserts a part from another random sequence at the cutting point. This technique significantly reduces overfitting and improves the performance approximately +0.02. Furthermore, I retrained the Transformer Large model with CutMix and pseudo-aligned labels like noisy student, and it achieved better performance on CV set.</p>\n<p>Moreover, I observed that while supplemental datasets without CutMix degrades the performance, they provided substantial enhancement when used with CutMix, resulting in an improvement of about +0.01 on both CV and LB.</p>\n<h2>Finetuning and Knowledge Distillation</h2>\n<p>The finetuning phase followed a standard approach. I simply utilized the CTC loss with augmentations mentioned above. Given the constraints of 40MB and 5 hours, the Transformer Large model was too extensive to be accommodated. I explored various combinations and parameter sizes, eventually setting on the Transformer Small (24L 256d) architecture. To compress the Large model into the Small model, I used knowledge distillation like DeiT to predict the hard prediction label from teacher model. I observed sharing same head for KD and CTC adversely affected the performance, so I used distinct heads. I also experimented with RNN-like heads (especially stacked BiLSTM) but there is no performance gain. It seems the model converges fast, but final performance is not changed.</p>\n<p><img src=\"https://i.ibb.co/qp1R3pZ/FigureB.png\" alt=\"FigureB\"></p>\n<p>The training epochs were as follows:</p>\n<ul>\n<li>Transformer Large (24L 1024d): 65 epochs</li>\n<li>Transformer Small (24L 256d): 830 epochs</li>\n</ul>\n<h2>CV vs LB</h2>\n<p>Contrary to the standard validation strategy (GroupKFold) that many people used, I simply split 5% of the training set to validate the model performance.</p>\n<pre><code>train_test_split(train_labels, test_size=, random_state=)\n</code></pre>\n<p>While I initially employed group kfold based on participant IDs, I encountered incongruities between cross-validation and public LB scores. The approach outlined above led to consistent improvements in both cross-validation and public LB performance.</p>\n<p><img src=\"https://i.ibb.co/DWVG0qF/FigureC.png\" alt=\"FigureC\"></p>\n<h2>What Didn't Work &amp; Future Work</h2>\n<ul>\n<li>Using external datasets did not work. Actually, I realized competition dataset is quite large and it was really hard to find large scale fingerspelling set as this competition one.</li>\n<li>Although prefix beam search showed a modest enhancement of +0.002 even with a small beam size, my implementation in tflite version is too slow, and I didn't use it finally.</li>\n<li>I experimented with various head architectures, but a single linear layer was sufficient.</li>\n<li>Conformer and transformer encoder-decoder models were not better than vanilla transformer.</li>\n<li>I also tried RandAugment and TrivialAugment but hand-crafted strong augmentations worked well.</li>\n</ul>",
  "messages": [
    {
      "id": 2407561,
      "postDate": "2023-08-25T06:15:16.517Z",
      "content": "<p>Thanks to Kaggle and competition hosts for this incredible and meaningful competition. Here is my solution.</p>\n<p>The code is available at <a href=\"https://github.com/affjljoo3581/Google-American-Sign-Language-Fingerspelling-Recognition\" target=\"_blank\">the github repository</a>.</p>\n<h2>TL;DR</h2>\n<ul>\n<li>Using hands, pose, and lips landmarks with 3D-based strong augmentation</li>\n<li>Vanilla transformer with conv stem and RoPE</li>\n<li>Data2vec 2.0 pretraining</li>\n<li>CTC segmentation and CutMix</li>\n<li>Knowledge distillation</li>\n</ul>\n<h2>Data Processing</h2>\n<p>I used 3d landmark points instead of 2d coordinate because it looks like applying rotation augmentation on 3d space makes models to be more robust. According to <a href=\"https://developers.google.com/mediapipe/api/solutions/java/com/google/mediapipe/tasks/components/containers/NormalizedLandmark\" target=\"_blank\">the mediapipe documentation</a>, the magnitude of z uses roughly the same scale as x. Since x and y are normalized by width and height of camera canvas and they are recorded on smartphone devices, I have to denormalize the landmarks with original aspect ratio to apply correct rotation transform. I simply estimate the aspect ratio by solving affine matrix that maps to the normalized hands at each frame to the standard hand landmarks and get scale factors from the affine matrix. The average aspect ratio is <code>0.5268970670149133</code> and I simply multiply <code>1.8979039030629028</code> to the normalized y values, i.e., <code>points *= np.array([1, 1.8979039030629028, 1])</code>.</p>\n<p>Correct Rotation Transform:<br>\n$$X' = SRX$$</p>\n<p>Wrong Rotation Transform:<br>\n$$X' = RSX$$</p>\n<p>It's important to note that applying rotation after normalizing coordinates is incorrect. Denormalize, rotate, and then normalize again. Actually, I did not normalize again because it is not necessary.</p>\n<p>For inputs, I utilized landmarks from the left hand, right hand, pose, and lips.  Initially, I focused solely on hand landmarks for the first few weeks and achieved a public LB score of 0.757. I believed fingerspelling is totally related to hand gestures only. Surprisingly, incorporating auxiliary landmarks such as pose and lips leads better performance and helped mitigate overfitting. The inclusion of additional input points contributed to better generalization.</p>\n<p>Here's a snippet of the data augmentation code I employed for both pretraining and finetuning:</p>\n<pre><code>Sequential(\n    LandmarkGroups(Normalize(), lengths=(, , )),\n    TimeFlip(p=),\n    RandomResample(limit=, p=),\n    Truncate(max_length),\n    AlignCTCLabel(),\n    LandmarkGroups(\n        transforms=(\n            FrameBlockMask(ratio=, block_size=, p=),\n            FrameBlockMask(ratio=, block_size=, p=),\n            FrameBlockMask(ratio=, block_size=, p=),\n        ),\n        lengths=(, , ),\n    ),\n    FrameNoise(ratio=, noise_stdev=, p=),\n    FeatureMask(ratio=, p=),\n    LandmarkGroups(\n        Sequential(\n            HorizontalFlip(p=),\n            RandomInterpolatedRotation(, np.pi / , p=),\n            RandomShear(limit=),\n            RandomScale(limit=),\n            RandomShift(stdev=),\n        ),\n        lengths=(, , ),\n    ),\n    Pad(max_length),\n)\n</code></pre>\n<p>The total number of input points is 75. Each component consists of 21, 21, 14, and 40 points, respectively. The hand with more <code>NaN</code> values is discarded and only the dominant hand is selected. As seen in the snippet above, spatial transformations are applied separately to the components. The landmark groups are first centered and normalized by maximum x-y values of each group. Note that z-values can exceed 1.</p>\n<h2>Model Architecture</h2>\n<p>I employed a simple Transformer encoder model similar to ViT. I used PreLN and rotary position embeddings. I replaced ViT's stem linear patch projection with a single convolutional layer. This alternation aimed to enable the model to capture relative positional differences, such as motion vectors, from the first convolutional layer. With the application of rotary embeddings, there is no length limit and also I didn't truncate input sequences at inference time. Similar to other ViT variants, I also integrated LayerDrop to mitigate overfitting.</p>\n<h2>Data2vec 2.0 Pretraining</h2>\n<p>Given that the inputs consist of 3d points and I used a normal Transformer architecture which has low inductive bias toward data attributes, I guessed it is necessary to pretrain the model to learn all about data properties. The <a href=\"https://arxiv.org/abs/2212.07525\" target=\"_blank\">Data2vec 2.0</a> method, known for its remarkable performance and efficiency across various domain, seemed promising for adaptation to landmark datasets.</p>\n<p><img src=\"https://i.ibb.co/KLM8LYh/FigureA.png\" alt=\"FigureA\"></p>\n<p>According to the paper, using multiple different masks within the same batch helps convergence and efficiency. I set $M = 8$ and $R = 0.5$, which means 50% of the input sequences are masked and there are 8 different masking patterns. After I experimented many various models, and I arrived at the following final models:</p>\n<ul>\n<li>Transformer Large (24L 1024d): 109 epochs (872 effective epochs)</li>\n<li>Transformer Small (24L 256d): 437 epochs (3496 effective epochs)</li>\n</ul>\n<p>Termination of the overall trainings was determined by training steps, not epochs, resulting in epochs that are not multiples of 10. After pretraining the model, the student parameters are used for finetuning.</p>\n<h2>CTC Segmentation and CutMix</h2>\n<p>Before explaining the finetuning part, it is essential to discuss CTC segmentation and CutMix augmentation. Check out <a href=\"https://github.com/lumaku/ctc-segmentation\" target=\"_blank\">this repository</a> and <a href=\"https://pytorch.org/audio/main/tutorials/forced_alignment_tutorial.html\" target=\"_blank\">this documentation</a> which provide information about CTC segmentation. To summarize, a CTC-trained model can detect the position of character appearances, enabling the inference of time alignment between phrases and landmark videos.</p>\n<p>Initially, I trained a Transformer Large model and created pseudo aligned labels. Using the alignments, I applied temporal CutMix augmentation which cuts random part of the original sequence and inserts a part from another random sequence at the cutting point. This technique significantly reduces overfitting and improves the performance approximately +0.02. Furthermore, I retrained the Transformer Large model with CutMix and pseudo-aligned labels like noisy student, and it achieved better performance on CV set.</p>\n<p>Moreover, I observed that while supplemental datasets without CutMix degrades the performance, they provided substantial enhancement when used with CutMix, resulting in an improvement of about +0.01 on both CV and LB.</p>\n<h2>Finetuning and Knowledge Distillation</h2>\n<p>The finetuning phase followed a standard approach. I simply utilized the CTC loss with augmentations mentioned above. Given the constraints of 40MB and 5 hours, the Transformer Large model was too extensive to be accommodated. I explored various combinations and parameter sizes, eventually setting on the Transformer Small (24L 256d) architecture. To compress the Large model into the Small model, I used knowledge distillation like DeiT to predict the hard prediction label from teacher model. I observed sharing same head for KD and CTC adversely affected the performance, so I used distinct heads. I also experimented with RNN-like heads (especially stacked BiLSTM) but there is no performance gain. It seems the model converges fast, but final performance is not changed.</p>\n<p><img src=\"https://i.ibb.co/qp1R3pZ/FigureB.png\" alt=\"FigureB\"></p>\n<p>The training epochs were as follows:</p>\n<ul>\n<li>Transformer Large (24L 1024d): 65 epochs</li>\n<li>Transformer Small (24L 256d): 830 epochs</li>\n</ul>\n<h2>CV vs LB</h2>\n<p>Contrary to the standard validation strategy (GroupKFold) that many people used, I simply split 5% of the training set to validate the model performance.</p>\n<pre><code>train_test_split(train_labels, test_size=, random_state=)\n</code></pre>\n<p>While I initially employed group kfold based on participant IDs, I encountered incongruities between cross-validation and public LB scores. The approach outlined above led to consistent improvements in both cross-validation and public LB performance.</p>\n<p><img src=\"https://i.ibb.co/DWVG0qF/FigureC.png\" alt=\"FigureC\"></p>\n<h2>What Didn't Work &amp; Future Work</h2>\n<ul>\n<li>Using external datasets did not work. Actually, I realized competition dataset is quite large and it was really hard to find large scale fingerspelling set as this competition one.</li>\n<li>Although prefix beam search showed a modest enhancement of +0.002 even with a small beam size, my implementation in tflite version is too slow, and I didn't use it finally.</li>\n<li>I experimented with various head architectures, but a single linear layer was sufficient.</li>\n<li>Conformer and transformer encoder-decoder models were not better than vanilla transformer.</li>\n<li>I also tried RandAugment and TrivialAugment but hand-crafted strong augmentations worked well.</li>\n</ul>",
      "rawMarkdown": "Thanks to Kaggle and competition hosts for this incredible and meaningful competition. Here is my solution.\n\nThe code is available at [the github repository](https://github.com/affjljoo3581/Google-American-Sign-Language-Fingerspelling-Recognition).\n\n## TL;DR\n* Using hands, pose, and lips landmarks with 3D-based strong augmentation\n* Vanilla transformer with conv stem and RoPE\n* Data2vec 2.0 pretraining\n* CTC segmentation and CutMix\n* Knowledge distillation\n\n## Data Processing\nI used 3d landmark points instead of 2d coordinate because it looks like applying rotation augmentation on 3d space makes models to be more robust. According to [the mediapipe documentation](https://developers.google.com/mediapipe/api/solutions/java/com/google/mediapipe/tasks/components/containers/NormalizedLandmark), the magnitude of z uses roughly the same scale as x. Since x and y are normalized by width and height of camera canvas and they are recorded on smartphone devices, I have to denormalize the landmarks with original aspect ratio to apply correct rotation transform. I simply estimate the aspect ratio by solving affine matrix that maps to the normalized hands at each frame to the standard hand landmarks and get scale factors from the affine matrix. The average aspect ratio is `0.5268970670149133` and I simply multiply `1.8979039030629028` to the normalized y values, i.e., `points *= np.array([1, 1.8979039030629028, 1])`.\n\nCorrect Rotation Transform:\n$$X' = SRX$$\n\nWrong Rotation Transform:\n$$X' = RSX$$\n\nIt's important to note that applying rotation after normalizing coordinates is incorrect. Denormalize, rotate, and then normalize again. Actually, I did not normalize again because it is not necessary.\n\nFor inputs, I utilized landmarks from the left hand, right hand, pose, and lips.  Initially, I focused solely on hand landmarks for the first few weeks and achieved a public LB score of 0.757. I believed fingerspelling is totally related to hand gestures only. Surprisingly, incorporating auxiliary landmarks such as pose and lips leads better performance and helped mitigate overfitting. The inclusion of additional input points contributed to better generalization.\n\nHere's a snippet of the data augmentation code I employed for both pretraining and finetuning:\n```python\nSequential(\n    LandmarkGroups(Normalize(), lengths=(21, 14, 40)),\n    TimeFlip(p=0.5),\n    RandomResample(limit=0.5, p=0.5),\n    Truncate(max_length),\n    AlignCTCLabel(),\n    LandmarkGroups(\n        transforms=(\n            FrameBlockMask(ratio=0.8, block_size=3, p=0.1),\n            FrameBlockMask(ratio=0.1, block_size=3, p=0.25),\n            FrameBlockMask(ratio=0.1, block_size=3, p=0.25),\n        ),\n        lengths=(21, 14, 40),\n    ),\n    FrameNoise(ratio=0.1, noise_stdev=0.3, p=0.25),\n    FeatureMask(ratio=0.1, p=0.1),\n    LandmarkGroups(\n        Sequential(\n            HorizontalFlip(p=0.5),\n            RandomInterpolatedRotation(0.2, np.pi / 4, p=0.5),\n            RandomShear(limit=0.2),\n            RandomScale(limit=0.2),\n            RandomShift(stdev=0.1),\n        ),\n        lengths=(21, 14, 40),\n    ),\n    Pad(max_length),\n)\n```\nThe total number of input points is 75. Each component consists of 21, 21, 14, and 40 points, respectively. The hand with more `NaN` values is discarded and only the dominant hand is selected. As seen in the snippet above, spatial transformations are applied separately to the components. The landmark groups are first centered and normalized by maximum x-y values of each group. Note that z-values can exceed 1.\n\n## Model Architecture\nI employed a simple Transformer encoder model similar to ViT. I used PreLN and rotary position embeddings. I replaced ViT's stem linear patch projection with a single convolutional layer. This alternation aimed to enable the model to capture relative positional differences, such as motion vectors, from the first convolutional layer. With the application of rotary embeddings, there is no length limit and also I didn't truncate input sequences at inference time. Similar to other ViT variants, I also integrated LayerDrop to mitigate overfitting.\n\n## Data2vec 2.0 Pretraining\nGiven that the inputs consist of 3d points and I used a normal Transformer architecture which has low inductive bias toward data attributes, I guessed it is necessary to pretrain the model to learn all about data properties. The [Data2vec 2.0](https://arxiv.org/abs/2212.07525) method, known for its remarkable performance and efficiency across various domain, seemed promising for adaptation to landmark datasets.\n\n![FigureA](https://i.ibb.co/KLM8LYh/FigureA.png)\n\nAccording to the paper, using multiple different masks within the same batch helps convergence and efficiency. I set $M = 8$ and $R = 0.5$, which means 50% of the input sequences are masked and there are 8 different masking patterns. After I experimented many various models, and I arrived at the following final models:\n* Transformer Large (24L 1024d): 109 epochs (872 effective epochs)\n* Transformer Small (24L 256d): 437 epochs (3496 effective epochs)\n\nTermination of the overall trainings was determined by training steps, not epochs, resulting in epochs that are not multiples of 10. After pretraining the model, the student parameters are used for finetuning.\n\n## CTC Segmentation and CutMix\nBefore explaining the finetuning part, it is essential to discuss CTC segmentation and CutMix augmentation. Check out [this repository](https://github.com/lumaku/ctc-segmentation) and [this documentation](https://pytorch.org/audio/main/tutorials/forced_alignment_tutorial.html) which provide information about CTC segmentation. To summarize, a CTC-trained model can detect the position of character appearances, enabling the inference of time alignment between phrases and landmark videos.\n\nInitially, I trained a Transformer Large model and created pseudo aligned labels. Using the alignments, I applied temporal CutMix augmentation which cuts random part of the original sequence and inserts a part from another random sequence at the cutting point. This technique significantly reduces overfitting and improves the performance approximately +0.02. Furthermore, I retrained the Transformer Large model with CutMix and pseudo-aligned labels like noisy student, and it achieved better performance on CV set.\n\nMoreover, I observed that while supplemental datasets without CutMix degrades the performance, they provided substantial enhancement when used with CutMix, resulting in an improvement of about +0.01 on both CV and LB.\n\n## Finetuning and Knowledge Distillation\nThe finetuning phase followed a standard approach. I simply utilized the CTC loss with augmentations mentioned above. Given the constraints of 40MB and 5 hours, the Transformer Large model was too extensive to be accommodated. I explored various combinations and parameter sizes, eventually setting on the Transformer Small (24L 256d) architecture. To compress the Large model into the Small model, I used knowledge distillation like DeiT to predict the hard prediction label from teacher model. I observed sharing same head for KD and CTC adversely affected the performance, so I used distinct heads. I also experimented with RNN-like heads (especially stacked BiLSTM) but there is no performance gain. It seems the model converges fast, but final performance is not changed.\n\n![FigureB](https://i.ibb.co/qp1R3pZ/FigureB.png)\n\nThe training epochs were as follows:\n* Transformer Large (24L 1024d): 65 epochs\n* Transformer Small (24L 256d): 830 epochs\n\n## CV vs LB\nContrary to the standard validation strategy (GroupKFold) that many people used, I simply split 5% of the training set to validate the model performance.\n```python\ntrain_test_split(train_labels, test_size=0.05, random_state=42)\n```\nWhile I initially employed group kfold based on participant IDs, I encountered incongruities between cross-validation and public LB scores. The approach outlined above led to consistent improvements in both cross-validation and public LB performance.\n\n![FigureC](https://i.ibb.co/DWVG0qF/FigureC.png)\n\n## What Didn't Work & Future Work\n- Using external datasets did not work. Actually, I realized competition dataset is quite large and it was really hard to find large scale fingerspelling set as this competition one.\n- Although prefix beam search showed a modest enhancement of +0.002 even with a small beam size, my implementation in tflite version is too slow, and I didn't use it finally.\n- I experimented with various head architectures, but a single linear layer was sufficient.\n- Conformer and transformer encoder-decoder models were not better than vanilla transformer.\n- I also tried RandAugment and TrivialAugment but hand-crafted strong augmentations worked well.",
      "votes": 50
    },
    {
      "id": 2407772,
      "postDate": "2023-08-25T08:54:10.690Z",
      "content": "<p>I planned to do cutmix and alignment but did not start the process for this is not easy and I could not know if it help much.<br>\nFrom your sharing it seems this is a key point, which could boost 10+ points, thanks for sharing I learned a lot from you.<br>\nFor data2vec does it improve your model performance much? </p>",
      "rawMarkdown": "I planned to do cutmix and alignment but did not start the process for this is not easy and I could not know if it help much.\nFrom your sharing it seems this is a key point, which could boost 10+ points, thanks for sharing I learned a lot from you.\nFor data2vec does it improve your model performance much? ",
      "votes": 1,
      "replies": [
        {
          "id": 2407802,
          "postDate": "2023-08-25T09:10:39.730Z",
          "content": "<p>Data2vec pretraining was the most important technique for me. The vanilla transformer model couldn't achieve high performance without pretraining. For the fair comparison, I tested on Transformer 9L 384d.</p>\n<ul>\n<li>From Scratch: CV 0.74</li>\n<li>w/ Data2vec Pretraining: CV 0.80</li>\n</ul>\n<p>And training from scratch was more vulnerable to the overfitting.</p>",
          "rawMarkdown": "Data2vec pretraining was the most important technique for me. The vanilla transformer model couldn't achieve high performance without pretraining. For the fair comparison, I tested on Transformer 9L 384d.\n- From Scratch: CV 0.74\n- w/ Data2vec Pretraining: CV 0.80\n\nAnd training from scratch was more vulnerable to the overfitting.",
          "votes": 3,
          "replies": [
            {
              "id": 2407814,
              "postDate": "2023-08-25T09:17:17.357Z",
              "content": "<p>Cool! Seems worth a try. Huggingface has wave2vec-conformer, it might also help.</p>",
              "rawMarkdown": "Cool! Seems worth a try. Huggingface has wave2vec-conformer, it might also help.",
              "votes": 1
            },
            {
              "id": 2407882,
              "postDate": "2023-08-25T09:46:07.887Z",
              "content": "<p>\"w/ Data2vec Pretraining: CV 0.80\"</p>\n<p>this is much more than i have expected. good work!!!</p>\n<p>do you use supplemental data for data2vec? </p>\n<p>it is said data2vec speedup supervised training, did you observe very fast convergence when fine tuning?<br>\ne.g. \"Transformer Large (24L 1024d): 65 epochs\" is that shorter than without pretrain for you?</p>\n<hr>\n<p>i use random splitting as well but at 20% for val/train. my CV is consistently 0.797 for different model (ctc, rnnt) and different landmark input. i use supplementary and kagget train (4x repeat). i think 0.80 is the limit. however, it need long iteration training, typically 100 epoch where one epoch = 1x supplementary + 4x kaggle.</p>",
              "rawMarkdown": "\"w/ Data2vec Pretraining: CV 0.80\"\n\nthis is much more than i have expected. good work!!!\n\ndo you use supplemental data for data2vec? \n\nit is said data2vec speedup supervised training, did you observe very fast convergence when fine tuning?\ne.g. \"Transformer Large (24L 1024d): 65 epochs\" is that shorter than without pretrain for you?\n\n---\n\ni use random splitting as well but at 20% for val/train. my CV is consistently 0.797 for different model (ctc, rnnt) and different landmark input. i use supplementary and kagget train (4x repeat). i think 0.80 is the limit. however, it need long iteration training, typically 100 epoch where one epoch = 1x supplementary + 4x kaggle.",
              "votes": 1
            },
            {
              "id": 2407912,
              "postDate": "2023-08-25T10:09:47.323Z",
              "content": "<p>Yes I used both train and supplemental datasets with equal frequencies. Here is a graph:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2160097%2F15ca15be596182e8928ad95134dbe875%2F2023-08-25%20185714.png?generation=1692957713066386&amp;alt=media\" alt=\"\"></p>\n<p>And Transformer 24L 256d (36MB) achieved CV 0.82 w/o pretraining while 9L 384d (31MB) achieved CV 0.80.</p>",
              "rawMarkdown": "Yes I used both train and supplemental datasets with equal frequencies. Here is a graph:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2160097%2F15ca15be596182e8928ad95134dbe875%2F2023-08-25%20185714.png?generation=1692957713066386&alt=media)\n\nAnd Transformer 24L 256d (36MB) achieved CV 0.82 w/o pretraining while 9L 384d (31MB) achieved CV 0.80.",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2407652,
      "postDate": "2023-08-25T07:22:54.823Z",
      "content": "<p>Congratulations, thanks for well-written presentation.</p>",
      "rawMarkdown": "Congratulations, thanks for well-written presentation.",
      "votes": 1
    },
    {
      "id": 2414872,
      "postDate": "2023-08-30T02:42:41.987Z",
      "content": "<p>I am checking your git repository but I couldn't find out what npy file is. Can you explain it? Thank you and congratulations.</p>",
      "rawMarkdown": "I am checking your git repository but I couldn't find out what npy file is. Can you explain it? Thank you and congratulations.",
      "replies": [
        {
          "id": 2415259,
          "postDate": "2023-08-30T07:48:41.497Z",
          "content": "<p>Thanks! I've just noticed that I forgot to include the preprocessing code. I have now made the necessary updates, and you can find the revised code <a href=\"https://github.com/affjljoo3581/Google-American-Sign-Language-Fingerspelling-Recognition/blob/main/resources/competition/extract_and_normalize_landmarks.ipynb\" target=\"_blank\">here</a>.</p>",
          "rawMarkdown": "Thanks! I've just noticed that I forgot to include the preprocessing code. I have now made the necessary updates, and you can find the revised code [here](https://github.com/affjljoo3581/Google-American-Sign-Language-Fingerspelling-Recognition/blob/main/resources/competition/extract_and_normalize_landmarks.ipynb).",
          "votes": 1,
          "replies": [
            {
              "id": 2430016,
              "postDate": "2023-09-09T02:30:31.777Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 2412626,
      "postDate": "2023-08-28T12:02:48.660Z",
      "content": "<p>I want to run your code, but I'm encountering an error:<br>\n<code>ValueError: coordinator_address should be defined.</code></p>\n<p>Could you provide more detailed instructions on running your code?</p>",
      "rawMarkdown": "I want to run your code, but I'm encountering an error:\n`ValueError: coordinator_address should be defined.`\n\nCould you provide more detailed instructions on running your code?",
      "replies": [
        {
          "id": 2415254,
          "postDate": "2023-08-30T07:45:06.107Z",
          "content": "<p>I had never encountered the error related to coordinator_address before. I'll check about that.</p>",
          "rawMarkdown": "I had never encountered the error related to coordinator_address before. I'll check about that."
        }
      ]
    },
    {
      "id": 2407585,
      "postDate": "2023-08-25T06:33:56.357Z",
      "content": "<p>Congratulations. Thanks for sharing your solution details. <br>\nDid you try Multiprocessing to speed up the solution?</p>",
      "rawMarkdown": "Congratulations. Thanks for sharing your solution details. \nDid you try Multiprocessing to speed up the solution?",
      "replies": [
        {
          "id": 2407588,
          "postDate": "2023-08-25T06:36:19.637Z",
          "content": "<p>No I didn't try.</p>",
          "rawMarkdown": "No I didn't try."
        }
      ]
    },
    {
      "id": 2426073,
      "postDate": "2023-09-06T11:40:48.180Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 2408965,
      "postDate": "2023-08-26T00:36:51.470Z",
      "content": "<p>Congratulations. Thanks for sharing</p>",
      "rawMarkdown": "Congratulations. Thanks for sharing",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 2407772,
      "author_name": "gezi",
      "author_url": "",
      "post_date": "2023-08-25T08:54:10.690000",
      "content": "<p>I planned to do cutmix and alignment but did not start the process for this is not easy and I could not know if it help much.<br>\nFrom your sharing it seems this is a key point, which could boost 10+ points, thanks for sharing I learned a lot from you.<br>\nFor data2vec does it improve your model performance much? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2407802,
          "author_name": "Jungwoo Park",
          "author_url": "",
          "post_date": "2023-08-25T09:10:39.730000",
          "content": "<p>Data2vec pretraining was the most important technique for me. The vanilla transformer model couldn't achieve high performance without pretraining. For the fair comparison, I tested on Transformer 9L 384d.</p>\n<ul>\n<li>From Scratch: CV 0.74</li>\n<li>w/ Data2vec Pretraining: CV 0.80</li>\n</ul>\n<p>And training from scratch was more vulnerable to the overfitting.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2407814,
              "author_name": "gezi",
              "author_url": "",
              "post_date": "2023-08-25T09:17:17.357000",
              "content": "<p>Cool! Seems worth a try. Huggingface has wave2vec-conformer, it might also help.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2407882,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-08-25T09:46:07.887000",
              "content": "<p>\"w/ Data2vec Pretraining: CV 0.80\"</p>\n<p>this is much more than i have expected. good work!!!</p>\n<p>do you use supplemental data for data2vec? </p>\n<p>it is said data2vec speedup supervised training, did you observe very fast convergence when fine tuning?<br>\ne.g. \"Transformer Large (24L 1024d): 65 epochs\" is that shorter than without pretrain for you?</p>\n<hr>\n<p>i use random splitting as well but at 20% for val/train. my CV is consistently 0.797 for different model (ctc, rnnt) and different landmark input. i use supplementary and kagget train (4x repeat). i think 0.80 is the limit. however, it need long iteration training, typically 100 epoch where one epoch = 1x supplementary + 4x kaggle.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2407912,
              "author_name": "Jungwoo Park",
              "author_url": "",
              "post_date": "2023-08-25T10:09:47.323000",
              "content": "<p>Yes I used both train and supplemental datasets with equal frequencies. Here is a graph:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2160097%2F15ca15be596182e8928ad95134dbe875%2F2023-08-25%20185714.png?generation=1692957713066386&amp;alt=media\" alt=\"\"></p>\n<p>And Transformer 24L 256d (36MB) achieved CV 0.82 w/o pretraining while 9L 384d (31MB) achieved CV 0.80.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2407652,
      "author_name": "Chakkrit Termritthikun",
      "author_url": "",
      "post_date": "2023-08-25T07:22:54.823000",
      "content": "<p>Congratulations, thanks for well-written presentation.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2414872,
      "author_name": "SEOUL 1988",
      "author_url": "",
      "post_date": "2023-08-30T02:42:41.987000",
      "content": "<p>I am checking your git repository but I couldn't find out what npy file is. Can you explain it? Thank you and congratulations.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2415259,
          "author_name": "Jungwoo Park",
          "author_url": "",
          "post_date": "2023-08-30T07:48:41.497000",
          "content": "<p>Thanks! I've just noticed that I forgot to include the preprocessing code. I have now made the necessary updates, and you can find the revised code <a href=\"https://github.com/affjljoo3581/Google-American-Sign-Language-Fingerspelling-Recognition/blob/main/resources/competition/extract_and_normalize_landmarks.ipynb\" target=\"_blank\">here</a>.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2430016,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-09-09T02:30:31.777000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2412626,
      "author_name": "Scenery SunFireInk",
      "author_url": "",
      "post_date": "2023-08-28T12:02:48.660000",
      "content": "<p>I want to run your code, but I'm encountering an error:<br>\n<code>ValueError: coordinator_address should be defined.</code></p>\n<p>Could you provide more detailed instructions on running your code?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2415254,
          "author_name": "Jungwoo Park",
          "author_url": "",
          "post_date": "2023-08-30T07:45:06.107000",
          "content": "<p>I had never encountered the error related to coordinator_address before. I'll check about that.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2407585,
      "author_name": "C R Suthikshn Kumar",
      "author_url": "",
      "post_date": "2023-08-25T06:33:56.357000",
      "content": "<p>Congratulations. Thanks for sharing your solution details. <br>\nDid you try Multiprocessing to speed up the solution?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2407588,
          "author_name": "Jungwoo Park",
          "author_url": "",
          "post_date": "2023-08-25T06:36:19.637000",
          "content": "<p>No I didn't try.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2426073,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-09-06T11:40:48.180000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2408965,
      "author_name": "ds.wook",
      "author_url": "",
      "post_date": "2023-08-26T00:36:51.470000",
      "content": "<p>Congratulations. Thanks for sharing</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2407561": "Thanks to Kaggle and competition hosts for this incredible and meaningful competition. Here is my solution.\n\nThe code is available at [the github repository](https://github.com/affjljoo3581/Google-American-Sign-Language-Fingerspelling-Recognition).\n\n## TL;DR\n* Using hands, pose, and lips landmarks with 3D-based strong augmentation\n* Vanilla transformer with conv stem and RoPE\n* Data2vec 2.0 pretraining\n* CTC segmentation and CutMix\n* Knowledge distillation\n\n## Data Processing\nI used 3d landmark points instead of 2d coordinate because it looks like applying rotation augmentation on 3d space makes models to be more robust. According to [the mediapipe documentation](https://developers.google.com/mediapipe/api/solutions/java/com/google/mediapipe/tasks/components/containers/NormalizedLandmark), the magnitude of z uses roughly the same scale as x. Since x and y are normalized by width and height of camera canvas and they are recorded on smartphone devices, I have to denormalize the landmarks with original aspect ratio to apply correct rotation transform. I simply estimate the aspect ratio by solving affine matrix that maps to the normalized hands at each frame to the standard hand landmarks and get scale factors from the affine matrix. The average aspect ratio is `0.5268970670149133` and I simply multiply `1.8979039030629028` to the normalized y values, i.e., `points *= np.array([1, 1.8979039030629028, 1])`.\n\nCorrect Rotation Transform:\n$$X' = SRX$$\n\nWrong Rotation Transform:\n$$X' = RSX$$\n\nIt's important to note that applying rotation after normalizing coordinates is incorrect. Denormalize, rotate, and then normalize again. Actually, I did not normalize again because it is not necessary.\n\nFor inputs, I utilized landmarks from the left hand, right hand, pose, and lips.  Initially, I focused solely on hand landmarks for the first few weeks and achieved a public LB score of 0.757. I believed fingerspelling is totally related to hand gestures only. Surprisingly, incorporating auxiliary landmarks such as pose and lips leads better performance and helped mitigate overfitting. The inclusion of additional input points contributed to better generalization.\n\nHere's a snippet of the data augmentation code I employed for both pretraining and finetuning:\n```python\nSequential(\n    LandmarkGroups(Normalize(), lengths=(21, 14, 40)),\n    TimeFlip(p=0.5),\n    RandomResample(limit=0.5, p=0.5),\n    Truncate(max_length),\n    AlignCTCLabel(),\n    LandmarkGroups(\n        transforms=(\n            FrameBlockMask(ratio=0.8, block_size=3, p=0.1),\n            FrameBlockMask(ratio=0.1, block_size=3, p=0.25),\n            FrameBlockMask(ratio=0.1, block_size=3, p=0.25),\n        ),\n        lengths=(21, 14, 40),\n    ),\n    FrameNoise(ratio=0.1, noise_stdev=0.3, p=0.25),\n    FeatureMask(ratio=0.1, p=0.1),\n    LandmarkGroups(\n        Sequential(\n            HorizontalFlip(p=0.5),\n            RandomInterpolatedRotation(0.2, np.pi / 4, p=0.5),\n            RandomShear(limit=0.2),\n            RandomScale(limit=0.2),\n            RandomShift(stdev=0.1),\n        ),\n        lengths=(21, 14, 40),\n    ),\n    Pad(max_length),\n)\n```\nThe total number of input points is 75. Each component consists of 21, 21, 14, and 40 points, respectively. The hand with more `NaN` values is discarded and only the dominant hand is selected. As seen in the snippet above, spatial transformations are applied separately to the components. The landmark groups are first centered and normalized by maximum x-y values of each group. Note that z-values can exceed 1.\n\n## Model Architecture\nI employed a simple Transformer encoder model similar to ViT. I used PreLN and rotary position embeddings. I replaced ViT's stem linear patch projection with a single convolutional layer. This alternation aimed to enable the model to capture relative positional differences, such as motion vectors, from the first convolutional layer. With the application of rotary embeddings, there is no length limit and also I didn't truncate input sequences at inference time. Similar to other ViT variants, I also integrated LayerDrop to mitigate overfitting.\n\n## Data2vec 2.0 Pretraining\nGiven that the inputs consist of 3d points and I used a normal Transformer architecture which has low inductive bias toward data attributes, I guessed it is necessary to pretrain the model to learn all about data properties. The [Data2vec 2.0](https://arxiv.org/abs/2212.07525) method, known for its remarkable performance and efficiency across various domain, seemed promising for adaptation to landmark datasets.\n\n![FigureA](https://i.ibb.co/KLM8LYh/FigureA.png)\n\nAccording to the paper, using multiple different masks within the same batch helps convergence and efficiency. I set $M = 8$ and $R = 0.5$, which means 50% of the input sequences are masked and there are 8 different masking patterns. After I experimented many various models, and I arrived at the following final models:\n* Transformer Large (24L 1024d): 109 epochs (872 effective epochs)\n* Transformer Small (24L 256d): 437 epochs (3496 effective epochs)\n\nTermination of the overall trainings was determined by training steps, not epochs, resulting in epochs that are not multiples of 10. After pretraining the model, the student parameters are used for finetuning.\n\n## CTC Segmentation and CutMix\nBefore explaining the finetuning part, it is essential to discuss CTC segmentation and CutMix augmentation. Check out [this repository](https://github.com/lumaku/ctc-segmentation) and [this documentation](https://pytorch.org/audio/main/tutorials/forced_alignment_tutorial.html) which provide information about CTC segmentation. To summarize, a CTC-trained model can detect the position of character appearances, enabling the inference of time alignment between phrases and landmark videos.\n\nInitially, I trained a Transformer Large model and created pseudo aligned labels. Using the alignments, I applied temporal CutMix augmentation which cuts random part of the original sequence and inserts a part from another random sequence at the cutting point. This technique significantly reduces overfitting and improves the performance approximately +0.02. Furthermore, I retrained the Transformer Large model with CutMix and pseudo-aligned labels like noisy student, and it achieved better performance on CV set.\n\nMoreover, I observed that while supplemental datasets without CutMix degrades the performance, they provided substantial enhancement when used with CutMix, resulting in an improvement of about +0.01 on both CV and LB.\n\n## Finetuning and Knowledge Distillation\nThe finetuning phase followed a standard approach. I simply utilized the CTC loss with augmentations mentioned above. Given the constraints of 40MB and 5 hours, the Transformer Large model was too extensive to be accommodated. I explored various combinations and parameter sizes, eventually setting on the Transformer Small (24L 256d) architecture. To compress the Large model into the Small model, I used knowledge distillation like DeiT to predict the hard prediction label from teacher model. I observed sharing same head for KD and CTC adversely affected the performance, so I used distinct heads. I also experimented with RNN-like heads (especially stacked BiLSTM) but there is no performance gain. It seems the model converges fast, but final performance is not changed.\n\n![FigureB](https://i.ibb.co/qp1R3pZ/FigureB.png)\n\nThe training epochs were as follows:\n* Transformer Large (24L 1024d): 65 epochs\n* Transformer Small (24L 256d): 830 epochs\n\n## CV vs LB\nContrary to the standard validation strategy (GroupKFold) that many people used, I simply split 5% of the training set to validate the model performance.\n```python\ntrain_test_split(train_labels, test_size=0.05, random_state=42)\n```\nWhile I initially employed group kfold based on participant IDs, I encountered incongruities between cross-validation and public LB scores. The approach outlined above led to consistent improvements in both cross-validation and public LB performance.\n\n![FigureC](https://i.ibb.co/DWVG0qF/FigureC.png)\n\n## What Didn't Work & Future Work\n- Using external datasets did not work. Actually, I realized competition dataset is quite large and it was really hard to find large scale fingerspelling set as this competition one.\n- Although prefix beam search showed a modest enhancement of +0.002 even with a small beam size, my implementation in tflite version is too slow, and I didn't use it finally.\n- I experimented with various head architectures, but a single linear layer was sufficient.\n- Conformer and transformer encoder-decoder models were not better than vanilla transformer.\n- I also tried RandAugment and TrivialAugment but hand-crafted strong augmentations worked well.",
    "2407772": "I planned to do cutmix and alignment but did not start the process for this is not easy and I could not know if it help much.\nFrom your sharing it seems this is a key point, which could boost 10+ points, thanks for sharing I learned a lot from you.\nFor data2vec does it improve your model performance much? ",
    "2407652": "Congratulations, thanks for well-written presentation.",
    "2414872": "I am checking your git repository but I couldn't find out what npy file is. Can you explain it? Thank you and congratulations.",
    "2412626": "I want to run your code, but I'm encountering an error:\n`ValueError: coordinator_address should be defined.`\n\nCould you provide more detailed instructions on running your code?",
    "2407585": "Congratulations. Thanks for sharing your solution details. \nDid you try Multiprocessing to speed up the solution?",
    "2426073": "",
    "2408965": "Congratulations. Thanks for sharing"
  }
}