{
  "id": 434680,
  "title": "22nd place solution: CTC Loss, Strong augmentations, CNN+MHSA",
  "url": "/competitions/asl-fingerspelling/discussion/434680",
  "author_name": "Samrat Thapa",
  "post_date": "2023-08-26T05:29:18.699000",
  "votes": 4,
  "comment_count": 5,
  "views": 0,
  "content": "<p><a href=\"https://github.com/SamratThapa120/sign-language-finger-spelling/blob/master/signet/models/feature_extractor_downsampled.py\" target=\"_blank\">My best-performing model</a> (<a href=\"https://github.com/SamratThapa120/sign-language-finger-spelling/blob/master/signet/configs/ctc_loss_with_downsampled_deploy_concataug.py\" target=\"_blank\">config</a>)is based on the <a href=\"https://www.kaggle.com/code/hoyso48/1st-place-solution-training/notebook\" target=\"_blank\">1st place solution of the previous competition</a> by <a href=\"https://www.kaggle.com/hoyso48\" target=\"_blank\">@hoyso48</a>. I trained the model using vanilla CTC Loss, and used greedy decoding for inference. </p>\n<h3>Things that worked for me:</h3>\n<ul>\n<li><p><strong>Longer input frames length:</strong>  A longer input frame length of 384 performed better than shorter length of 256.   </p></li>\n<li><p><strong>Deeper/larger model:</strong>  My best-performing model has 7 blocks with hidden dimension of 256. It has 9.5M parameters.</p></li>\n<li><p><strong>Pose keypoints:</strong> Adding pose information was helpful. Best-performing model uses pose+hands+lips+eyes.</p></li>\n<li><p><strong>CNN+MHSA:</strong> CNN+MHSA model &gt; Only CNN Model&gt;Only MHSA model</p></li>\n<li><p><strong>Strong augmentation:</strong> My model was performing well based on local CV, and public LB, with strong correlations between the two.But when I tested the model using this <a href=\"https://github.com/SamratThapa120/gradio-ASL-fingerspelling-recognition\" target=\"_blank\">Gradio app</a> using my webcam, it could not recognize my signs, so I had to use strong data augmentations to get the model to work. Especially temporal mask helped because I sign slower than the pros. These augmentations also boosted the public LB score by +0.006.</p>\n<p>flip_lr_probability=0.5<br>\nrandom_affine_probability=0.75<br>\nfreeze_probability=0.5<br>\ntemporal_mask_probability=0.75<br>\ntemporal_mask_range=(0.2,0.4)</p></li>\n<li><p><strong>concat augmentation:</strong> Randomly concatenate two short landmark sequences, as well as their labels. This improved public LB by +0.008. I applied this augmentation to 40% of all training samples.</p></li>\n</ul>\n<h3>Here are other things I tried, that did not contribute to the best-performing model:</h3>\n<ul>\n<li><p><strong>transformer-encoder+decoder:</strong> Transformer endoder-decoder model with cross-entropy loss.</p></li>\n<li><p><strong>transformer-decoder:</strong> CNN+MHSA model and Transformer-like decoder with cross-entropy loss.</p></li>\n<li><p><strong>Causal-masking in self-attention:</strong> removing causal masking performed better than using causal mask.</p></li>\n<li><p><strong>Attention-span in self-attention:</strong> I thought that there would not be long-term dependency between frames for this task, so I tried to reduce the attention span of self-attention. Although there was no performance degradation, it did not boost performance either.</p></li>\n<li><p><strong>Focal loss</strong>: I tried the CTC Focal loss based on <a href=\"https://github.com/TeaPoly/CTC-OptimizedLoss/blob/main/ctc_focal_loss.py\" target=\"_blank\">this</a> repo, but there were no gains.</p></li>\n<li><p><strong>Erase landmarks augmentation</strong>: Randomly erase landmarks except the hand landmarks, to make the model more robust to mediapipe's detection errors.</p></li>\n</ul>\n<p>Originally I was using <a href=\"https://github.com/SamratThapa120/sign-language-finger-spelling/tree/pytorch\" target=\"_blank\">pytorch</a>, but I switched to Tensorflow because I faced several issues when converting the pyTorch model to tfLite. However, training with CTC Loss was upto 10x slower in tensorflow than pytorch. Looking back, I think I should have put more effort into fixing the model conversion issue, as I would have been able to perform more experiments.  </p>\n<p>Congratulations to the winning teams. I would also like to thank Google for hosting this competition, it was a valuable opporunity to learn many new things. Also, I would like to thank everyone who shared their notebooks and unique ideas.</p>",
  "messages": [
    {
      "id": 2409214,
      "postDate": "2023-08-26T05:29:18.700Z",
      "content": "<p><a href=\"https://github.com/SamratThapa120/sign-language-finger-spelling/blob/master/signet/models/feature_extractor_downsampled.py\" target=\"_blank\">My best-performing model</a> (<a href=\"https://github.com/SamratThapa120/sign-language-finger-spelling/blob/master/signet/configs/ctc_loss_with_downsampled_deploy_concataug.py\" target=\"_blank\">config</a>)is based on the <a href=\"https://www.kaggle.com/code/hoyso48/1st-place-solution-training/notebook\" target=\"_blank\">1st place solution of the previous competition</a> by <a href=\"https://www.kaggle.com/hoyso48\" target=\"_blank\">@hoyso48</a>. I trained the model using vanilla CTC Loss, and used greedy decoding for inference. </p>\n<h3>Things that worked for me:</h3>\n<ul>\n<li><p><strong>Longer input frames length:</strong>  A longer input frame length of 384 performed better than shorter length of 256.   </p></li>\n<li><p><strong>Deeper/larger model:</strong>  My best-performing model has 7 blocks with hidden dimension of 256. It has 9.5M parameters.</p></li>\n<li><p><strong>Pose keypoints:</strong> Adding pose information was helpful. Best-performing model uses pose+hands+lips+eyes.</p></li>\n<li><p><strong>CNN+MHSA:</strong> CNN+MHSA model &gt; Only CNN Model&gt;Only MHSA model</p></li>\n<li><p><strong>Strong augmentation:</strong> My model was performing well based on local CV, and public LB, with strong correlations between the two.But when I tested the model using this <a href=\"https://github.com/SamratThapa120/gradio-ASL-fingerspelling-recognition\" target=\"_blank\">Gradio app</a> using my webcam, it could not recognize my signs, so I had to use strong data augmentations to get the model to work. Especially temporal mask helped because I sign slower than the pros. These augmentations also boosted the public LB score by +0.006.</p>\n<p>flip_lr_probability=0.5<br>\nrandom_affine_probability=0.75<br>\nfreeze_probability=0.5<br>\ntemporal_mask_probability=0.75<br>\ntemporal_mask_range=(0.2,0.4)</p></li>\n<li><p><strong>concat augmentation:</strong> Randomly concatenate two short landmark sequences, as well as their labels. This improved public LB by +0.008. I applied this augmentation to 40% of all training samples.</p></li>\n</ul>\n<h3>Here are other things I tried, that did not contribute to the best-performing model:</h3>\n<ul>\n<li><p><strong>transformer-encoder+decoder:</strong> Transformer endoder-decoder model with cross-entropy loss.</p></li>\n<li><p><strong>transformer-decoder:</strong> CNN+MHSA model and Transformer-like decoder with cross-entropy loss.</p></li>\n<li><p><strong>Causal-masking in self-attention:</strong> removing causal masking performed better than using causal mask.</p></li>\n<li><p><strong>Attention-span in self-attention:</strong> I thought that there would not be long-term dependency between frames for this task, so I tried to reduce the attention span of self-attention. Although there was no performance degradation, it did not boost performance either.</p></li>\n<li><p><strong>Focal loss</strong>: I tried the CTC Focal loss based on <a href=\"https://github.com/TeaPoly/CTC-OptimizedLoss/blob/main/ctc_focal_loss.py\" target=\"_blank\">this</a> repo, but there were no gains.</p></li>\n<li><p><strong>Erase landmarks augmentation</strong>: Randomly erase landmarks except the hand landmarks, to make the model more robust to mediapipe's detection errors.</p></li>\n</ul>\n<p>Originally I was using <a href=\"https://github.com/SamratThapa120/sign-language-finger-spelling/tree/pytorch\" target=\"_blank\">pytorch</a>, but I switched to Tensorflow because I faced several issues when converting the pyTorch model to tfLite. However, training with CTC Loss was upto 10x slower in tensorflow than pytorch. Looking back, I think I should have put more effort into fixing the model conversion issue, as I would have been able to perform more experiments.  </p>\n<p>Congratulations to the winning teams. I would also like to thank Google for hosting this competition, it was a valuable opporunity to learn many new things. Also, I would like to thank everyone who shared their notebooks and unique ideas.</p>",
      "rawMarkdown": "[My best-performing model](https://github.com/SamratThapa120/sign-language-finger-spelling/blob/master/signet/models/feature_extractor_downsampled.py) ([config](https://github.com/SamratThapa120/sign-language-finger-spelling/blob/master/signet/configs/ctc_loss_with_downsampled_deploy_concataug.py))is based on the [1st place solution of the previous competition](https://www.kaggle.com/code/hoyso48/1st-place-solution-training/notebook) by @hoyso48. I trained the model using vanilla CTC Loss, and used greedy decoding for inference. \n\n### Things that worked for me:\n- **Longer input frames length:**  A longer input frame length of 384 performed better than shorter length of 256.   \n- **Deeper/larger model:**  My best-performing model has 7 blocks with hidden dimension of 256. It has 9.5M parameters.\n- **Pose keypoints:** Adding pose information was helpful. Best-performing model uses pose+hands+lips+eyes.\n- **CNN+MHSA:** CNN+MHSA model > Only CNN Model>Only MHSA model\n- **Strong augmentation:** My model was performing well based on local CV, and public LB, with strong correlations between the two.But when I tested the model using this [Gradio app](https://github.com/SamratThapa120/gradio-ASL-fingerspelling-recognition) using my webcam, it could not recognize my signs, so I had to use strong data augmentations to get the model to work. Especially temporal mask helped because I sign slower than the pros. These augmentations also boosted the public LB score by +0.006.\n\n    flip_lr_probability=0.5\n    random_affine_probability=0.75\n    freeze_probability=0.5\n    temporal_mask_probability=0.75\n    temporal_mask_range=(0.2,0.4)\n\n- **concat augmentation:** Randomly concatenate two short landmark sequences, as well as their labels. This improved public LB by +0.008. I applied this augmentation to 40% of all training samples.\n\n### Here are other things I tried, that did not contribute to the best-performing model:\n\n- **transformer-encoder+decoder:** Transformer endoder-decoder model with cross-entropy loss.\n\n- **transformer-decoder:** CNN+MHSA model and Transformer-like decoder with cross-entropy loss.\n\n- **Causal-masking in self-attention:** removing causal masking performed better than using causal mask.\n\n- **Attention-span in self-attention:** I thought that there would not be long-term dependency between frames for this task, so I tried to reduce the attention span of self-attention. Although there was no performance degradation, it did not boost performance either.\n\n- **Focal loss**: I tried the CTC Focal loss based on [this](https://github.com/TeaPoly/CTC-OptimizedLoss/blob/main/ctc_focal_loss.py) repo, but there were no gains.\n\n- **Erase landmarks augmentation**: Randomly erase landmarks except the hand landmarks, to make the model more robust to mediapipe's detection errors.\n\nOriginally I was using [pytorch](https://github.com/SamratThapa120/sign-language-finger-spelling/tree/pytorch), but I switched to Tensorflow because I faced several issues when converting the pyTorch model to tfLite. However, training with CTC Loss was upto 10x slower in tensorflow than pytorch. Looking back, I think I should have put more effort into fixing the model conversion issue, as I would have been able to perform more experiments.  \n\nCongratulations to the winning teams. I would also like to thank Google for hosting this competition, it was a valuable opporunity to learn many new things. Also, I would like to thank everyone who shared their notebooks and unique ideas.\n\n",
      "votes": 4
    },
    {
      "id": 2409236,
      "postDate": "2023-08-26T05:45:48.377Z",
      "content": "<p>Thank you for sharing. Great to see that you used concat augmentation. Could you quickly explain how freeze augmentation works?</p>",
      "rawMarkdown": "Thank you for sharing. Great to see that you used concat augmentation. Could you quickly explain how freeze augmentation works?",
      "votes": 1,
      "replies": [
        {
          "id": 2409261,
          "postDate": "2023-08-26T06:06:45.150Z",
          "content": "<p>Freeze augmentation tries to replicate frozen frames in a video. Randomly select some frames and replicate them without changing their order in the sequence. For example [a,b,c,d,e] will get converted to something like [a,a,a,b,b,c,d,e,e].  </p>\n<p><a href=\"https://github.com/SamratThapa120/sign-language-finger-spelling/blob/2bbc9a4c51622d3553ab9be17715ca51332c8797/signet/dataset/transforms.py#L215\" target=\"_blank\">Here is the code</a></p>",
          "rawMarkdown": "Freeze augmentation tries to replicate frozen frames in a video. Randomly select some frames and replicate them without changing their order in the sequence. For example [a,b,c,d,e] will get converted to something like [a,a,a,b,b,c,d,e,e].  \n\n[Here is the code](https://github.com/SamratThapa120/sign-language-finger-spelling/blob/2bbc9a4c51622d3553ab9be17715ca51332c8797/signet/dataset/transforms.py#L215)",
          "votes": 1
        }
      ]
    },
    {
      "id": 2410124,
      "postDate": "2023-08-26T16:41:34.687Z",
      "content": "<p>Congratulations. Thanks for sharing the techniques useful for higher scores.<br>\nAre the augmentations random or additional data set is generated. </p>",
      "rawMarkdown": "Congratulations. Thanks for sharing the techniques useful for higher scores.\nAre the augmentations random or additional data set is generated. ",
      "replies": [
        {
          "id": 2410489,
          "postDate": "2023-08-27T02:39:57Z",
          "content": "<p>The augmentations are applied randomly during training</p>",
          "rawMarkdown": "The augmentations are applied randomly during training"
        }
      ]
    },
    {
      "id": 2414029,
      "postDate": "2023-08-29T10:22:27.170Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2409236,
      "author_name": "Dieter",
      "author_url": "",
      "post_date": "2023-08-26T05:45:48.377000",
      "content": "<p>Thank you for sharing. Great to see that you used concat augmentation. Could you quickly explain how freeze augmentation works?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2409261,
          "author_name": "Samrat Thapa",
          "author_url": "",
          "post_date": "2023-08-26T06:06:45.150000",
          "content": "<p>Freeze augmentation tries to replicate frozen frames in a video. Randomly select some frames and replicate them without changing their order in the sequence. For example [a,b,c,d,e] will get converted to something like [a,a,a,b,b,c,d,e,e].  </p>\n<p><a href=\"https://github.com/SamratThapa120/sign-language-finger-spelling/blob/2bbc9a4c51622d3553ab9be17715ca51332c8797/signet/dataset/transforms.py#L215\" target=\"_blank\">Here is the code</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2410124,
      "author_name": "C R Suthikshn Kumar",
      "author_url": "",
      "post_date": "2023-08-26T16:41:34.687000",
      "content": "<p>Congratulations. Thanks for sharing the techniques useful for higher scores.<br>\nAre the augmentations random or additional data set is generated. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2410489,
          "author_name": "Samrat Thapa",
          "author_url": "",
          "post_date": "2023-08-27T02:39:57",
          "content": "<p>The augmentations are applied randomly during training</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2414029,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-08-29T10:22:27.170000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2409214": "[My best-performing model](https://github.com/SamratThapa120/sign-language-finger-spelling/blob/master/signet/models/feature_extractor_downsampled.py) ([config](https://github.com/SamratThapa120/sign-language-finger-spelling/blob/master/signet/configs/ctc_loss_with_downsampled_deploy_concataug.py))is based on the [1st place solution of the previous competition](https://www.kaggle.com/code/hoyso48/1st-place-solution-training/notebook) by @hoyso48. I trained the model using vanilla CTC Loss, and used greedy decoding for inference. \n\n### Things that worked for me:\n- **Longer input frames length:**  A longer input frame length of 384 performed better than shorter length of 256.   \n- **Deeper/larger model:**  My best-performing model has 7 blocks with hidden dimension of 256. It has 9.5M parameters.\n- **Pose keypoints:** Adding pose information was helpful. Best-performing model uses pose+hands+lips+eyes.\n- **CNN+MHSA:** CNN+MHSA model > Only CNN Model>Only MHSA model\n- **Strong augmentation:** My model was performing well based on local CV, and public LB, with strong correlations between the two.But when I tested the model using this [Gradio app](https://github.com/SamratThapa120/gradio-ASL-fingerspelling-recognition) using my webcam, it could not recognize my signs, so I had to use strong data augmentations to get the model to work. Especially temporal mask helped because I sign slower than the pros. These augmentations also boosted the public LB score by +0.006.\n\n    flip_lr_probability=0.5\n    random_affine_probability=0.75\n    freeze_probability=0.5\n    temporal_mask_probability=0.75\n    temporal_mask_range=(0.2,0.4)\n\n- **concat augmentation:** Randomly concatenate two short landmark sequences, as well as their labels. This improved public LB by +0.008. I applied this augmentation to 40% of all training samples.\n\n### Here are other things I tried, that did not contribute to the best-performing model:\n\n- **transformer-encoder+decoder:** Transformer endoder-decoder model with cross-entropy loss.\n\n- **transformer-decoder:** CNN+MHSA model and Transformer-like decoder with cross-entropy loss.\n\n- **Causal-masking in self-attention:** removing causal masking performed better than using causal mask.\n\n- **Attention-span in self-attention:** I thought that there would not be long-term dependency between frames for this task, so I tried to reduce the attention span of self-attention. Although there was no performance degradation, it did not boost performance either.\n\n- **Focal loss**: I tried the CTC Focal loss based on [this](https://github.com/TeaPoly/CTC-OptimizedLoss/blob/main/ctc_focal_loss.py) repo, but there were no gains.\n\n- **Erase landmarks augmentation**: Randomly erase landmarks except the hand landmarks, to make the model more robust to mediapipe's detection errors.\n\nOriginally I was using [pytorch](https://github.com/SamratThapa120/sign-language-finger-spelling/tree/pytorch), but I switched to Tensorflow because I faced several issues when converting the pyTorch model to tfLite. However, training with CTC Loss was upto 10x slower in tensorflow than pytorch. Looking back, I think I should have put more effort into fixing the model conversion issue, as I would have been able to perform more experiments.  \n\nCongratulations to the winning teams. I would also like to thank Google for hosting this competition, it was a valuable opporunity to learn many new things. Also, I would like to thank everyone who shared their notebooks and unique ideas.\n\n",
    "2409236": "Thank you for sharing. Great to see that you used concat augmentation. Could you quickly explain how freeze augmentation works?",
    "2410124": "Congratulations. Thanks for sharing the techniques useful for higher scores.\nAre the augmentations random or additional data set is generated. ",
    "2414029": ""
  }
}