{
  "id": 434871,
  "title": "9th Place Solution",
  "url": "/competitions/asl-fingerspelling/discussion/434871",
  "author_name": "Rob Mulla",
  "post_date": "2023-08-26T21:14:21.556000",
  "votes": 43,
  "comment_count": 8,
  "views": 0,
  "content": "<p>The below is result of a team effort with <a href=\"https://www.kaggle.com/rafiko1\" target=\"_blank\">@rafiko1</a> and <a href=\"https://www.kaggle.com/group16\" target=\"_blank\">@group16</a> and myself.</p>\n<p>We want to thank to the hosts for putting together a great competition, it was really fun and we were glad to end up in the gold zone! We've already enjoyed reading some of the top teams solutions- congrats to all who participated.</p>\n<h2>tl;dr</h2>\n<p>We used a similar architecture to <a href=\"https://www.kaggle.com/hoyso48\" target=\"_blank\">@hoyso48</a>'s solution from the ASL-sign competition. Our model had six Conv1D/Transformer blocks, was trained with CTC loss over 500+ epochs, and utilized heavy augmentations.</p>\n<p>This competition posed challenges beyond achieving model accuracy. We had to meet model constraints (5-hour submission limit and 40mb model size). We also needed to make sure all of our code converted nicely to tflite, which at times was frustrating. This required us to find a good balance between model size and inference time with minimal postprocessing.</p>\n<h2>Data and Preprocessing:</h2>\n<p>We used the training data as the base dataset, but also trained using the supplemental and external (ChicagoWild/Plus) during fine tuning. Preprocessing was the same for all datasets:</p>\n<ul>\n<li>Standard scaling using the mean and standard deviation for each point.</li>\n<li>x, y, z coordinates for each point.</li>\n<li>Points used included: RHAND, LHAND, LIP, POSE, REYE, and LEYE.</li>\n<li>We tried different settings for the max frame length since it had the most impact on our preprocessing and ened up using 356.</li>\n<li>Downsampling non-hand frames: To handle samples that has more than 356 frames, we first removed frames with missing hands at even intervals. This improved performance when we added it to our inference- so eventually we added it to our training pipeline.</li>\n<li>Resize: Examples exceeding our max frame length post-downsampling were resized.</li>\n</ul>\n<h2>Augmentations:</h2>\n<p>Heavy augmentations helped us to train long without the risk of overfitting.</p>\n<ul>\n<li>Mirror/Flip</li>\n<li>Random resampling</li>\n<li>Random rotation</li>\n<li>Random spatial and temporal masking</li>\n<li>Temporal cropping</li>\n<li>Minimal Random noise</li>\n<li>Random scaling</li>\n</ul>\n<h2>Training</h2>\n<p>Given that training each model took a very long time to train (days running on colab TPUs), we resumed from the best checkpoints to experiment with different augmentations, data ratios, and learning rate schedules.  It’s hard to exactly retrace the steps we took to get to our final model but the general idea was:</p>\n<ul>\n<li>Base training without augmentations at a consistent learning rate until convergence (~150 epochs).</li>\n<li>Introducing heavy augmentations and continuing until the validation score plateaued (200-400+ epochs).</li>\n</ul>\n<p>Example of how our validation score improved from different experiments over the final week of the competition:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F644036%2F19aa07e39c404ae7599d66184997d351%2FScreenshot%20from%202023-08-26%2016-58-18.png?generation=1693083542531397&amp;alt=media\" alt=\"\"></p>\n<p>Even up until the end of the competition we were able to get small improvements in our score from more training.</p>\n<h2>Model Architecture:</h2>\n<p>Initially, I tried an efficientnet approach similar to the 2nd place in the signs competition. Later, I pivoted to the 1st place architecture similar to public notebooks. After merging teams we found both of us had similar architectures but <a href=\"https://www.kaggle.com/rafiko1\" target=\"_blank\">@rafiko1</a> made some modifications that improved it's performance on this task:</p>\n<ul>\n<li>Switched from 3 to 2 Conv1D per Conv/Transformer block.</li>\n<li>Increated to 6 blocks (resulting in a 10.5M parameter model).</li>\n<li>Parameters: 384 max frames, 320 dimensions, 8 heads per transformer. </li>\n<li>Did not use a final MaxPooling layer.</li>\n</ul>\n<p>We also had experiments that used DebertaV2 instead of the transformer block, But the final model didn't use them.</p>\n<h2>Increasing model size without retraining:</h2>\n<p>Towards the competition's end, realizing we didn't have a great way to ensemble models, we wanted to get the biggest model possible without having to retrain. We found that instead of retraining from scratch we could simply expand the number of Conv1D/Transformer blocks we had - and copy the weights over from our previous best model, repeating the final layer weights. This allowed the model to converge very fast and squeeze out some additional accuracy while saying below the 5 hour inference limit.</p>\n<h2>Attempts not in our final solution:</h2>\n<ul>\n<li><strong>Cutmix based on pseudo labels:</strong> The idea was to use model predictions to pinpoint frames displaying each letter in videos, then applying cutmix at the frame level during training. It functioned initially but became required us to recreate the tfrecords each time we modfiied preprocessing or model archetecture, so it wasn't used in our final solution. When it was used we only applied cutmix across similar phrase types (phone numbers mixed only with phone numbers, etc) but it was interesting to read the top team found it worked best mixing only between the same signer.</li>\n<li><strong>Ensembling techniques:</strong> Couldn't find a good solution for this with CTC loss in tflite.</li>\n<li><strong>Beam search</strong></li>\n<li><strong>Efficientnet/Transformer</strong></li>\n<li><strong>DebertaV2</strong></li>\n</ul>\n<h2>Postprocessing / Ensemble:</h2>\n<p><strong>Fill word</strong>: As others have noted, there were a significant number of examples in the dataset with very few frames. These had low or no signal in them related to the target. We found it was better to just replace them with a fill word ('+w1- ea-or') We also found that it worked best to apply these to examples that had less than 30 frames AND a prediction phrase that was &lt;= 6 characters.<br>\n<strong>Model weight averaging</strong> Instead of using the final epoch of each training run, averaging the weights from the final few epochs really helped the CV and LB score. This became less powerful later in the competition when our models were stronger.</p>\n<h2>Strong Validation/LB Correlation:</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F644036%2F2a8e047f8cb571c018f477396c321a91%2Fsub-lb.png?generation=1693083215072772&amp;alt=media\" alt=\"\"></p>\n<p>We tracked each submission's execution time and CV/LB correlation. This competition had a really strong correlation between our validation set (one parquet file) and the LB. Our best private LB score was 0.788 but we didn't select it because it did not perform as well on the public LB.</p>",
  "messages": [
    {
      "id": 2410369,
      "postDate": "2023-08-26T21:14:21.557Z",
      "content": "<p>The below is result of a team effort with <a href=\"https://www.kaggle.com/rafiko1\" target=\"_blank\">@rafiko1</a> and <a href=\"https://www.kaggle.com/group16\" target=\"_blank\">@group16</a> and myself.</p>\n<p>We want to thank to the hosts for putting together a great competition, it was really fun and we were glad to end up in the gold zone! We've already enjoyed reading some of the top teams solutions- congrats to all who participated.</p>\n<h2>tl;dr</h2>\n<p>We used a similar architecture to <a href=\"https://www.kaggle.com/hoyso48\" target=\"_blank\">@hoyso48</a>'s solution from the ASL-sign competition. Our model had six Conv1D/Transformer blocks, was trained with CTC loss over 500+ epochs, and utilized heavy augmentations.</p>\n<p>This competition posed challenges beyond achieving model accuracy. We had to meet model constraints (5-hour submission limit and 40mb model size). We also needed to make sure all of our code converted nicely to tflite, which at times was frustrating. This required us to find a good balance between model size and inference time with minimal postprocessing.</p>\n<h2>Data and Preprocessing:</h2>\n<p>We used the training data as the base dataset, but also trained using the supplemental and external (ChicagoWild/Plus) during fine tuning. Preprocessing was the same for all datasets:</p>\n<ul>\n<li>Standard scaling using the mean and standard deviation for each point.</li>\n<li>x, y, z coordinates for each point.</li>\n<li>Points used included: RHAND, LHAND, LIP, POSE, REYE, and LEYE.</li>\n<li>We tried different settings for the max frame length since it had the most impact on our preprocessing and ened up using 356.</li>\n<li>Downsampling non-hand frames: To handle samples that has more than 356 frames, we first removed frames with missing hands at even intervals. This improved performance when we added it to our inference- so eventually we added it to our training pipeline.</li>\n<li>Resize: Examples exceeding our max frame length post-downsampling were resized.</li>\n</ul>\n<h2>Augmentations:</h2>\n<p>Heavy augmentations helped us to train long without the risk of overfitting.</p>\n<ul>\n<li>Mirror/Flip</li>\n<li>Random resampling</li>\n<li>Random rotation</li>\n<li>Random spatial and temporal masking</li>\n<li>Temporal cropping</li>\n<li>Minimal Random noise</li>\n<li>Random scaling</li>\n</ul>\n<h2>Training</h2>\n<p>Given that training each model took a very long time to train (days running on colab TPUs), we resumed from the best checkpoints to experiment with different augmentations, data ratios, and learning rate schedules.  It’s hard to exactly retrace the steps we took to get to our final model but the general idea was:</p>\n<ul>\n<li>Base training without augmentations at a consistent learning rate until convergence (~150 epochs).</li>\n<li>Introducing heavy augmentations and continuing until the validation score plateaued (200-400+ epochs).</li>\n</ul>\n<p>Example of how our validation score improved from different experiments over the final week of the competition:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F644036%2F19aa07e39c404ae7599d66184997d351%2FScreenshot%20from%202023-08-26%2016-58-18.png?generation=1693083542531397&amp;alt=media\" alt=\"\"></p>\n<p>Even up until the end of the competition we were able to get small improvements in our score from more training.</p>\n<h2>Model Architecture:</h2>\n<p>Initially, I tried an efficientnet approach similar to the 2nd place in the signs competition. Later, I pivoted to the 1st place architecture similar to public notebooks. After merging teams we found both of us had similar architectures but <a href=\"https://www.kaggle.com/rafiko1\" target=\"_blank\">@rafiko1</a> made some modifications that improved it's performance on this task:</p>\n<ul>\n<li>Switched from 3 to 2 Conv1D per Conv/Transformer block.</li>\n<li>Increated to 6 blocks (resulting in a 10.5M parameter model).</li>\n<li>Parameters: 384 max frames, 320 dimensions, 8 heads per transformer. </li>\n<li>Did not use a final MaxPooling layer.</li>\n</ul>\n<p>We also had experiments that used DebertaV2 instead of the transformer block, But the final model didn't use them.</p>\n<h2>Increasing model size without retraining:</h2>\n<p>Towards the competition's end, realizing we didn't have a great way to ensemble models, we wanted to get the biggest model possible without having to retrain. We found that instead of retraining from scratch we could simply expand the number of Conv1D/Transformer blocks we had - and copy the weights over from our previous best model, repeating the final layer weights. This allowed the model to converge very fast and squeeze out some additional accuracy while saying below the 5 hour inference limit.</p>\n<h2>Attempts not in our final solution:</h2>\n<ul>\n<li><strong>Cutmix based on pseudo labels:</strong> The idea was to use model predictions to pinpoint frames displaying each letter in videos, then applying cutmix at the frame level during training. It functioned initially but became required us to recreate the tfrecords each time we modfiied preprocessing or model archetecture, so it wasn't used in our final solution. When it was used we only applied cutmix across similar phrase types (phone numbers mixed only with phone numbers, etc) but it was interesting to read the top team found it worked best mixing only between the same signer.</li>\n<li><strong>Ensembling techniques:</strong> Couldn't find a good solution for this with CTC loss in tflite.</li>\n<li><strong>Beam search</strong></li>\n<li><strong>Efficientnet/Transformer</strong></li>\n<li><strong>DebertaV2</strong></li>\n</ul>\n<h2>Postprocessing / Ensemble:</h2>\n<p><strong>Fill word</strong>: As others have noted, there were a significant number of examples in the dataset with very few frames. These had low or no signal in them related to the target. We found it was better to just replace them with a fill word ('+w1- ea-or') We also found that it worked best to apply these to examples that had less than 30 frames AND a prediction phrase that was &lt;= 6 characters.<br>\n<strong>Model weight averaging</strong> Instead of using the final epoch of each training run, averaging the weights from the final few epochs really helped the CV and LB score. This became less powerful later in the competition when our models were stronger.</p>\n<h2>Strong Validation/LB Correlation:</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F644036%2F2a8e047f8cb571c018f477396c321a91%2Fsub-lb.png?generation=1693083215072772&amp;alt=media\" alt=\"\"></p>\n<p>We tracked each submission's execution time and CV/LB correlation. This competition had a really strong correlation between our validation set (one parquet file) and the LB. Our best private LB score was 0.788 but we didn't select it because it did not perform as well on the public LB.</p>",
      "rawMarkdown": "The below is result of a team effort with @rafiko1 and @group16 and myself.\n\nWe want to thank to the hosts for putting together a great competition, it was really fun and we were glad to end up in the gold zone! We've already enjoyed reading some of the top teams solutions- congrats to all who participated.\n\n## tl;dr\n\nWe used a similar architecture to @hoyso48's solution from the ASL-sign competition. Our model had six Conv1D/Transformer blocks, was trained with CTC loss over 500+ epochs, and utilized heavy augmentations.\n\nThis competition posed challenges beyond achieving model accuracy. We had to meet model constraints (5-hour submission limit and 40mb model size). We also needed to make sure all of our code converted nicely to tflite, which at times was frustrating. This required us to find a good balance between model size and inference time with minimal postprocessing.\n\n## Data and Preprocessing:\n\nWe used the training data as the base dataset, but also trained using the supplemental and external (ChicagoWild/Plus) during fine tuning. Preprocessing was the same for all datasets:\n\n- Standard scaling using the mean and standard deviation for each point.\n- x, y, z coordinates for each point.\n- Points used included: RHAND, LHAND, LIP, POSE, REYE, and LEYE.\n- We tried different settings for the max frame length since it had the most impact on our preprocessing and ened up using 356.\n- Downsampling non-hand frames: To handle samples that has more than 356 frames, we first removed frames with missing hands at even intervals. This improved performance when we added it to our inference- so eventually we added it to our training pipeline.\n- Resize: Examples exceeding our max frame length post-downsampling were resized.\n\n## Augmentations:\nHeavy augmentations helped us to train long without the risk of overfitting.\n\n- Mirror/Flip\n- Random resampling\n- Random rotation\n- Random spatial and temporal masking\n- Temporal cropping\n- Minimal Random noise\n- Random scaling\n\n## Training\nGiven that training each model took a very long time to train (days running on colab TPUs), we resumed from the best checkpoints to experiment with different augmentations, data ratios, and learning rate schedules.  It’s hard to exactly retrace the steps we took to get to our final model but the general idea was:\n- Base training without augmentations at a consistent learning rate until convergence (~150 epochs).\n- Introducing heavy augmentations and continuing until the validation score plateaued (200-400+ epochs).\n\nExample of how our validation score improved from different experiments over the final week of the competition:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F644036%2F19aa07e39c404ae7599d66184997d351%2FScreenshot%20from%202023-08-26%2016-58-18.png?generation=1693083542531397&alt=media)\n\nEven up until the end of the competition we were able to get small improvements in our score from more training.\n\n## Model Architecture:\nInitially, I tried an efficientnet approach similar to the 2nd place in the signs competition. Later, I pivoted to the 1st place architecture similar to public notebooks. After merging teams we found both of us had similar architectures but @rafiko1 made some modifications that improved it's performance on this task:\n- Switched from 3 to 2 Conv1D per Conv/Transformer block.\n- Increated to 6 blocks (resulting in a 10.5M parameter model).\n- Parameters: 384 max frames, 320 dimensions, 8 heads per transformer. \n- Did not use a final MaxPooling layer.\n\nWe also had experiments that used DebertaV2 instead of the transformer block, But the final model didn't use them.\n\n## Increasing model size without retraining:\nTowards the competition's end, realizing we didn't have a great way to ensemble models, we wanted to get the biggest model possible without having to retrain. We found that instead of retraining from scratch we could simply expand the number of Conv1D/Transformer blocks we had - and copy the weights over from our previous best model, repeating the final layer weights. This allowed the model to converge very fast and squeeze out some additional accuracy while saying below the 5 hour inference limit.\n\n## Attempts not in our final solution:\n- **Cutmix based on pseudo labels:** The idea was to use model predictions to pinpoint frames displaying each letter in videos, then applying cutmix at the frame level during training. It functioned initially but became required us to recreate the tfrecords each time we modfiied preprocessing or model archetecture, so it wasn't used in our final solution. When it was used we only applied cutmix across similar phrase types (phone numbers mixed only with phone numbers, etc) but it was interesting to read the top team found it worked best mixing only between the same signer.\n- **Ensembling techniques:** Couldn't find a good solution for this with CTC loss in tflite.\n- **Beam search**\n- **Efficientnet/Transformer**\n- **DebertaV2**\n\n## Postprocessing / Ensemble:\n**Fill word**: As others have noted, there were a significant number of examples in the dataset with very few frames. These had low or no signal in them related to the target. We found it was better to just replace them with a fill word ('+w1- ea-or') We also found that it worked best to apply these to examples that had less than 30 frames AND a prediction phrase that was <= 6 characters.\n**Model weight averaging** Instead of using the final epoch of each training run, averaging the weights from the final few epochs really helped the CV and LB score. This became less powerful later in the competition when our models were stronger.\n\n## Strong Validation/LB Correlation:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F644036%2F2a8e047f8cb571c018f477396c321a91%2Fsub-lb.png?generation=1693083215072772&alt=media)\n\nWe tracked each submission's execution time and CV/LB correlation. This competition had a really strong correlation between our validation set (one parquet file) and the LB. Our best private LB score was 0.788 but we didn't select it because it did not perform as well on the public LB.\n",
      "votes": 43
    },
    {
      "id": 2410666,
      "postDate": "2023-08-27T06:35:58.733Z",
      "content": "<p>Thank you for sharing. Nice trick with increasing model size without retraining.<br>\nHow much did ChicagoWild/Plus help?</p>",
      "rawMarkdown": "Thank you for sharing. Nice trick with increasing model size without retraining.\nHow much did ChicagoWild/Plus help?",
      "votes": 3,
      "replies": [
        {
          "id": 2410783,
          "postDate": "2023-08-27T07:35:31.070Z",
          "content": "<p>Thank you! We didn't have the time to re-train so this trick by <a href=\"https://www.kaggle.com/robikscube\" target=\"_blank\">@robikscube</a> was really useful.<br>\nChicagoWild/Plus: around +0.002 on val / LB</p>",
          "rawMarkdown": "Thank you! We didn't have the time to re-train so this trick by @robikscube was really useful.\nChicagoWild/Plus: around +0.002 on val / LB",
          "votes": 1,
          "replies": [
            {
              "id": 2412460,
              "postDate": "2023-08-28T09:34:29.050Z",
              "content": "<p>About the extra data, ChicagoWild/Plus, may I ask if it is allowed to use due to its license? </p>\n<p>I also used it at the beginning. But the improvement is negligible and I'm afraid its unclear license will cause some issue, I decided to not use it at the end.</p>",
              "rawMarkdown": "About the extra data, ChicagoWild/Plus, may I ask if it is allowed to use due to its license? \n\nI also used it at the beginning. But the improvement is negligible and I'm afraid its unclear license will cause some issue, I decided to not use it at the end.",
              "votes": 1
            },
            {
              "id": 2413136,
              "postDate": "2023-08-28T17:26:07.743Z",
              "content": "<p>btw, my opinion on this can be found here<br>\n<a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/433060#2410668\" target=\"_blank\">https://www.kaggle.com/competitions/asl-fingerspelling/discussion/433060#2410668</a></p>",
              "rawMarkdown": "btw, my opinion on this can be found here\nhttps://www.kaggle.com/competitions/asl-fingerspelling/discussion/433060#2410668",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2481446,
      "postDate": "2023-10-14T06:14:21.840Z",
      "content": "<p>Nice trick with increasing model size without retraining, Great work</p>",
      "rawMarkdown": "Nice trick with increasing model size without retraining, Great work\n",
      "votes": -1
    },
    {
      "id": 2415201,
      "postDate": "2023-08-30T07:04:21.333Z",
      "content": "<p>Thank you for the write up. I have one question, how do you measure the submission time? I am doing it manually. Do you have a better solution?</p>",
      "rawMarkdown": "Thank you for the write up. I have one question, how do you measure the submission time? I am doing it manually. Do you have a better solution?"
    },
    {
      "id": 2412130,
      "postDate": "2023-08-28T05:24:20.743Z",
      "content": "<p>Congratulations.  Thanks for sharing the nice plot of your CV/LB correlation. Hows the CV/LB/Private LB correlations ? I think in this competition, we don't see big shakeup.  </p>",
      "rawMarkdown": "Congratulations.  Thanks for sharing the nice plot of your CV/LB correlation. Hows the CV/LB/Private LB correlations ? I think in this competition, we don't see big shakeup.  "
    },
    {
      "id": 2426070,
      "postDate": "2023-09-06T11:40:06.650Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2410666,
      "author_name": "Dieter",
      "author_url": "",
      "post_date": "2023-08-27T06:35:58.733000",
      "content": "<p>Thank you for sharing. Nice trick with increasing model size without retraining.<br>\nHow much did ChicagoWild/Plus help?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2410783,
          "author_name": "Rafi Hai",
          "author_url": "",
          "post_date": "2023-08-27T07:35:31.070000",
          "content": "<p>Thank you! We didn't have the time to re-train so this trick by <a href=\"https://www.kaggle.com/robikscube\" target=\"_blank\">@robikscube</a> was really useful.<br>\nChicagoWild/Plus: around +0.002 on val / LB</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2412460,
              "author_name": "bliao",
              "author_url": "",
              "post_date": "2023-08-28T09:34:29.050000",
              "content": "<p>About the extra data, ChicagoWild/Plus, may I ask if it is allowed to use due to its license? </p>\n<p>I also used it at the beginning. But the improvement is negligible and I'm afraid its unclear license will cause some issue, I decided to not use it at the end.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2413136,
              "author_name": "Dieter",
              "author_url": "",
              "post_date": "2023-08-28T17:26:07.743000",
              "content": "<p>btw, my opinion on this can be found here<br>\n<a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/433060#2410668\" target=\"_blank\">https://www.kaggle.com/competitions/asl-fingerspelling/discussion/433060#2410668</a></p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2481446,
      "author_name": "Aditya kishor",
      "author_url": "",
      "post_date": "2023-10-14T06:14:21.840000",
      "content": "<p>Nice trick with increasing model size without retraining, Great work</p>",
      "votes": -1,
      "replies": []
    },
    {
      "id": 2415201,
      "author_name": "Red Neural",
      "author_url": "",
      "post_date": "2023-08-30T07:04:21.333000",
      "content": "<p>Thank you for the write up. I have one question, how do you measure the submission time? I am doing it manually. Do you have a better solution?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2412130,
      "author_name": "C R Suthikshn Kumar",
      "author_url": "",
      "post_date": "2023-08-28T05:24:20.743000",
      "content": "<p>Congratulations.  Thanks for sharing the nice plot of your CV/LB correlation. Hows the CV/LB/Private LB correlations ? I think in this competition, we don't see big shakeup.  </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2426070,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-09-06T11:40:06.650000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2410369": "The below is result of a team effort with @rafiko1 and @group16 and myself.\n\nWe want to thank to the hosts for putting together a great competition, it was really fun and we were glad to end up in the gold zone! We've already enjoyed reading some of the top teams solutions- congrats to all who participated.\n\n## tl;dr\n\nWe used a similar architecture to @hoyso48's solution from the ASL-sign competition. Our model had six Conv1D/Transformer blocks, was trained with CTC loss over 500+ epochs, and utilized heavy augmentations.\n\nThis competition posed challenges beyond achieving model accuracy. We had to meet model constraints (5-hour submission limit and 40mb model size). We also needed to make sure all of our code converted nicely to tflite, which at times was frustrating. This required us to find a good balance between model size and inference time with minimal postprocessing.\n\n## Data and Preprocessing:\n\nWe used the training data as the base dataset, but also trained using the supplemental and external (ChicagoWild/Plus) during fine tuning. Preprocessing was the same for all datasets:\n\n- Standard scaling using the mean and standard deviation for each point.\n- x, y, z coordinates for each point.\n- Points used included: RHAND, LHAND, LIP, POSE, REYE, and LEYE.\n- We tried different settings for the max frame length since it had the most impact on our preprocessing and ened up using 356.\n- Downsampling non-hand frames: To handle samples that has more than 356 frames, we first removed frames with missing hands at even intervals. This improved performance when we added it to our inference- so eventually we added it to our training pipeline.\n- Resize: Examples exceeding our max frame length post-downsampling were resized.\n\n## Augmentations:\nHeavy augmentations helped us to train long without the risk of overfitting.\n\n- Mirror/Flip\n- Random resampling\n- Random rotation\n- Random spatial and temporal masking\n- Temporal cropping\n- Minimal Random noise\n- Random scaling\n\n## Training\nGiven that training each model took a very long time to train (days running on colab TPUs), we resumed from the best checkpoints to experiment with different augmentations, data ratios, and learning rate schedules.  It’s hard to exactly retrace the steps we took to get to our final model but the general idea was:\n- Base training without augmentations at a consistent learning rate until convergence (~150 epochs).\n- Introducing heavy augmentations and continuing until the validation score plateaued (200-400+ epochs).\n\nExample of how our validation score improved from different experiments over the final week of the competition:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F644036%2F19aa07e39c404ae7599d66184997d351%2FScreenshot%20from%202023-08-26%2016-58-18.png?generation=1693083542531397&alt=media)\n\nEven up until the end of the competition we were able to get small improvements in our score from more training.\n\n## Model Architecture:\nInitially, I tried an efficientnet approach similar to the 2nd place in the signs competition. Later, I pivoted to the 1st place architecture similar to public notebooks. After merging teams we found both of us had similar architectures but @rafiko1 made some modifications that improved it's performance on this task:\n- Switched from 3 to 2 Conv1D per Conv/Transformer block.\n- Increated to 6 blocks (resulting in a 10.5M parameter model).\n- Parameters: 384 max frames, 320 dimensions, 8 heads per transformer. \n- Did not use a final MaxPooling layer.\n\nWe also had experiments that used DebertaV2 instead of the transformer block, But the final model didn't use them.\n\n## Increasing model size without retraining:\nTowards the competition's end, realizing we didn't have a great way to ensemble models, we wanted to get the biggest model possible without having to retrain. We found that instead of retraining from scratch we could simply expand the number of Conv1D/Transformer blocks we had - and copy the weights over from our previous best model, repeating the final layer weights. This allowed the model to converge very fast and squeeze out some additional accuracy while saying below the 5 hour inference limit.\n\n## Attempts not in our final solution:\n- **Cutmix based on pseudo labels:** The idea was to use model predictions to pinpoint frames displaying each letter in videos, then applying cutmix at the frame level during training. It functioned initially but became required us to recreate the tfrecords each time we modfiied preprocessing or model archetecture, so it wasn't used in our final solution. When it was used we only applied cutmix across similar phrase types (phone numbers mixed only with phone numbers, etc) but it was interesting to read the top team found it worked best mixing only between the same signer.\n- **Ensembling techniques:** Couldn't find a good solution for this with CTC loss in tflite.\n- **Beam search**\n- **Efficientnet/Transformer**\n- **DebertaV2**\n\n## Postprocessing / Ensemble:\n**Fill word**: As others have noted, there were a significant number of examples in the dataset with very few frames. These had low or no signal in them related to the target. We found it was better to just replace them with a fill word ('+w1- ea-or') We also found that it worked best to apply these to examples that had less than 30 frames AND a prediction phrase that was <= 6 characters.\n**Model weight averaging** Instead of using the final epoch of each training run, averaging the weights from the final few epochs really helped the CV and LB score. This became less powerful later in the competition when our models were stronger.\n\n## Strong Validation/LB Correlation:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F644036%2F2a8e047f8cb571c018f477396c321a91%2Fsub-lb.png?generation=1693083215072772&alt=media)\n\nWe tracked each submission's execution time and CV/LB correlation. This competition had a really strong correlation between our validation set (one parquet file) and the LB. Our best private LB score was 0.788 but we didn't select it because it did not perform as well on the public LB.\n",
    "2410666": "Thank you for sharing. Nice trick with increasing model size without retraining.\nHow much did ChicagoWild/Plus help?",
    "2481446": "Nice trick with increasing model size without retraining, Great work\n",
    "2415201": "Thank you for the write up. I have one question, how do you measure the submission time? I am doing it manually. Do you have a better solution?",
    "2412130": "Congratulations.  Thanks for sharing the nice plot of your CV/LB correlation. Hows the CV/LB/Private LB correlations ? I think in this competition, we don't see big shakeup.  ",
    "2426070": ""
  }
}