{
  "id": 434658,
  "title": "20th place solution",
  "url": "/competitions/asl-fingerspelling/discussion/434658",
  "author_name": "Bohan Yoon",
  "post_date": "2023-08-26T02:22:01.959000",
  "votes": 2,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Thanks to Kaggle, Google and other organizers for hosting this exciting competition. Hope this kinds of competition open in kaggle more often.</p>\n<p><strong>TLDR</strong></p>\n<p>My solution is a single large model trained with CTC loss with no ensemble(39MB). The output shape of the model is (batch, sentence length, class).</p>\n<p><strong>Data Preprocessing and Augmentation</strong></p>\n<p>Basically the same as the previous 1st solution. But I found using (x, y, z) and pose landmarks gives better score. I used the following landmarks: left hand, right hand, eye, lips and pose landmarks.</p>\n<p><strong>Model</strong></p>\n<p>I used previous 1st solution. However, I used the branchformer. Also, I changed the multi head attention to relative multi head attention. My model consists of 1 dense layer - 4 conv1d layer - 6 branchformer layer. In addition, I used stochastic depth.</p>\n<p><strong>Training</strong></p>\n<p>Epoch = 150<br>\nbs = 64<br>\nLr = 8e-4<br>\nAWP = Epoch * 0.1<br>\nSchedule = CosineDecay with warmup ratio 0.1<br>\nOptimizer = Adam with Lookahead<br>\nLoss = Inter CTC loss(<a href=\"https://arxiv.org/abs/2102.03216\" target=\"_blank\">https://arxiv.org/abs/2102.03216</a>)</p>",
  "messages": [
    {
      "id": 2409001,
      "postDate": "2023-08-26T02:22:01.960Z",
      "content": "<p>Thanks to Kaggle, Google and other organizers for hosting this exciting competition. Hope this kinds of competition open in kaggle more often.</p>\n<p><strong>TLDR</strong></p>\n<p>My solution is a single large model trained with CTC loss with no ensemble(39MB). The output shape of the model is (batch, sentence length, class).</p>\n<p><strong>Data Preprocessing and Augmentation</strong></p>\n<p>Basically the same as the previous 1st solution. But I found using (x, y, z) and pose landmarks gives better score. I used the following landmarks: left hand, right hand, eye, lips and pose landmarks.</p>\n<p><strong>Model</strong></p>\n<p>I used previous 1st solution. However, I used the branchformer. Also, I changed the multi head attention to relative multi head attention. My model consists of 1 dense layer - 4 conv1d layer - 6 branchformer layer. In addition, I used stochastic depth.</p>\n<p><strong>Training</strong></p>\n<p>Epoch = 150<br>\nbs = 64<br>\nLr = 8e-4<br>\nAWP = Epoch * 0.1<br>\nSchedule = CosineDecay with warmup ratio 0.1<br>\nOptimizer = Adam with Lookahead<br>\nLoss = Inter CTC loss(<a href=\"https://arxiv.org/abs/2102.03216\" target=\"_blank\">https://arxiv.org/abs/2102.03216</a>)</p>",
      "rawMarkdown": "Thanks to Kaggle, Google and other organizers for hosting this exciting competition. Hope this kinds of competition open in kaggle more often.\n\n**TLDR**\n\nMy solution is a single large model trained with CTC loss with no ensemble(39MB). The output shape of the model is (batch, sentence length, class).\n\n**Data Preprocessing and Augmentation**\n\nBasically the same as the previous 1st solution. But I found using (x, y, z) and pose landmarks gives better score. I used the following landmarks: left hand, right hand, eye, lips and pose landmarks.\n\n**Model**\n\nI used previous 1st solution. However, I used the branchformer. Also, I changed the multi head attention to relative multi head attention. My model consists of 1 dense layer - 4 conv1d layer - 6 branchformer layer. In addition, I used stochastic depth.\n\n**Training**\n\nEpoch = 150\nbs = 64\nLr = 8e-4\nAWP = Epoch * 0.1\nSchedule = CosineDecay with warmup ratio 0.1\nOptimizer = Adam with Lookahead\nLoss = Inter CTC loss(https://arxiv.org/abs/2102.03216)",
      "votes": 2
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2409001": "Thanks to Kaggle, Google and other organizers for hosting this exciting competition. Hope this kinds of competition open in kaggle more often.\n\n**TLDR**\n\nMy solution is a single large model trained with CTC loss with no ensemble(39MB). The output shape of the model is (batch, sentence length, class).\n\n**Data Preprocessing and Augmentation**\n\nBasically the same as the previous 1st solution. But I found using (x, y, z) and pose landmarks gives better score. I used the following landmarks: left hand, right hand, eye, lips and pose landmarks.\n\n**Model**\n\nI used previous 1st solution. However, I used the branchformer. Also, I changed the multi head attention to relative multi head attention. My model consists of 1 dense layer - 4 conv1d layer - 6 branchformer layer. In addition, I used stochastic depth.\n\n**Training**\n\nEpoch = 150\nbs = 64\nLr = 8e-4\nAWP = Epoch * 0.1\nSchedule = CosineDecay with warmup ratio 0.1\nOptimizer = Adam with Lookahead\nLoss = Inter CTC loss(https://arxiv.org/abs/2102.03216)"
  }
}