{
  "id": 434395,
  "title": "a quick summary of 13th solution",
  "url": "/competitions/asl-fingerspelling/discussion/434395",
  "author_name": "pudae",
  "post_date": "2023-08-25T04:20:51.474000",
  "votes": 14,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Firstly, I'd like to extend my gratitude to Kaggle and Google for hosting such an engaging competition, and congratulations to all the winners.</p>\n<p>I'm keen to share a few methodologies that yielded a notable enhancement in the competition score. Unfortunately, I didn't maintain logs for my experiments, so I'm unable to quantify the exact impact of each modification. My apologies for this oversight.</p>\n<p>Given that the challenge bore a resemblance to speech recognition, I initiated my approach using the Conformer architecture paired with the CTC loss.</p>\n<p><strong>Baseline Configuration</strong></p>\n<ul>\n<li><strong>Landmark Encoder</strong>: 2-layer 1D CNN</li>\n<li><strong>Architecture</strong>: 4-layer Conformer</li>\n<li><strong>Loss Function</strong>: CTC Loss</li>\n<li><strong>Features</strong>: Dominant hand and pose, lip</li>\n<li><strong>Augmentations</strong>: Rotation, resizing, shearing, random masking</li>\n</ul>\n<p><strong>Key Modifications that Enhanced Performance</strong></p>\n<ul>\n<li><strong>Architecture Update</strong>: <ul>\n<li>Transitioned from Conformer to SqueezeFormer</li>\n<li>2-layer Separable Convolution followed by 1D CNN for landmark encoder</li>\n<li>Replaced GLU with swish (worked better under the same number of parameters)<ul>\n<li>Further transitioning from swish to mish provided incremental benefits</li></ul></li>\n<li>Temporal reduction before the 4th block<ul>\n<li>implemented using average pooling 1d or max pooling 1d</li></ul></li>\n<li>Additional forward path omitting temporal reduction(parameters are shared) and applying the CTC loss (only for training)</li></ul></li>\n<li><strong>Augmentation</strong><ul>\n<li>concatenated two randomly sampled examples</li>\n<li>For the noisy examples (whose nonnan frames are less than the length of phrase), generate synthetic frames.<ul>\n<li>aligning char and frame using another trained model and create char-frame data from them.</li>\n<li>for each character in the phrase, randomly sample frame from the char-frame data and concatenate them.</li></ul></li></ul></li>\n</ul>\n<p>I apologize again for the lack of detailed information. Despite the lack of exhaustive specifics, I hope these are helpful for you.</p>\n<p>Thanks.</p>",
  "messages": [
    {
      "id": 2407398,
      "postDate": "2023-08-25T04:20:51.473Z",
      "content": "<p>Firstly, I'd like to extend my gratitude to Kaggle and Google for hosting such an engaging competition, and congratulations to all the winners.</p>\n<p>I'm keen to share a few methodologies that yielded a notable enhancement in the competition score. Unfortunately, I didn't maintain logs for my experiments, so I'm unable to quantify the exact impact of each modification. My apologies for this oversight.</p>\n<p>Given that the challenge bore a resemblance to speech recognition, I initiated my approach using the Conformer architecture paired with the CTC loss.</p>\n<p><strong>Baseline Configuration</strong></p>\n<ul>\n<li><strong>Landmark Encoder</strong>: 2-layer 1D CNN</li>\n<li><strong>Architecture</strong>: 4-layer Conformer</li>\n<li><strong>Loss Function</strong>: CTC Loss</li>\n<li><strong>Features</strong>: Dominant hand and pose, lip</li>\n<li><strong>Augmentations</strong>: Rotation, resizing, shearing, random masking</li>\n</ul>\n<p><strong>Key Modifications that Enhanced Performance</strong></p>\n<ul>\n<li><strong>Architecture Update</strong>: <ul>\n<li>Transitioned from Conformer to SqueezeFormer</li>\n<li>2-layer Separable Convolution followed by 1D CNN for landmark encoder</li>\n<li>Replaced GLU with swish (worked better under the same number of parameters)<ul>\n<li>Further transitioning from swish to mish provided incremental benefits</li></ul></li>\n<li>Temporal reduction before the 4th block<ul>\n<li>implemented using average pooling 1d or max pooling 1d</li></ul></li>\n<li>Additional forward path omitting temporal reduction(parameters are shared) and applying the CTC loss (only for training)</li></ul></li>\n<li><strong>Augmentation</strong><ul>\n<li>concatenated two randomly sampled examples</li>\n<li>For the noisy examples (whose nonnan frames are less than the length of phrase), generate synthetic frames.<ul>\n<li>aligning char and frame using another trained model and create char-frame data from them.</li>\n<li>for each character in the phrase, randomly sample frame from the char-frame data and concatenate them.</li></ul></li></ul></li>\n</ul>\n<p>I apologize again for the lack of detailed information. Despite the lack of exhaustive specifics, I hope these are helpful for you.</p>\n<p>Thanks.</p>",
      "rawMarkdown": "Firstly, I'd like to extend my gratitude to Kaggle and Google for hosting such an engaging competition, and congratulations to all the winners.\n\nI'm keen to share a few methodologies that yielded a notable enhancement in the competition score. Unfortunately, I didn't maintain logs for my experiments, so I'm unable to quantify the exact impact of each modification. My apologies for this oversight.\n\nGiven that the challenge bore a resemblance to speech recognition, I initiated my approach using the Conformer architecture paired with the CTC loss.\n\n**Baseline Configuration**\n- **Landmark Encoder**: 2-layer 1D CNN\n- **Architecture**: 4-layer Conformer\n- **Loss Function**: CTC Loss\n- **Features**: Dominant hand and pose, lip\n- **Augmentations**: Rotation, resizing, shearing, random masking\n\n**Key Modifications that Enhanced Performance**\n- **Architecture Update**: \n  - Transitioned from Conformer to SqueezeFormer\n  - 2-layer Separable Convolution followed by 1D CNN for landmark encoder\n  - Replaced GLU with swish (worked better under the same number of parameters)\n     - Further transitioning from swish to mish provided incremental benefits\n  - Temporal reduction before the 4th block\n     - implemented using average pooling 1d or max pooling 1d\n  - Additional forward path omitting temporal reduction(parameters are shared) and applying the CTC loss (only for training)\n- **Augmentation**\n  - concatenated two randomly sampled examples\n  - For the noisy examples (whose nonnan frames are less than the length of phrase), generate synthetic frames.\n     - aligning char and frame using another trained model and create char-frame data from them.\n     - for each character in the phrase, randomly sample frame from the char-frame data and concatenate them.\n\nI apologize again for the lack of detailed information. Despite the lack of exhaustive specifics, I hope these are helpful for you.\n\nThanks.\n ",
      "votes": 14
    },
    {
      "id": 2407595,
      "postDate": "2023-08-25T06:39:34.877Z",
      "content": "<p>Congratulations. Thanks for sharing the details of your solution. Did you have LSTM in your pipeline? Any changes to learning rate and warmup?</p>",
      "rawMarkdown": "Congratulations. Thanks for sharing the details of your solution. Did you have LSTM in your pipeline? Any changes to learning rate and warmup?",
      "replies": [
        {
          "id": 2408017,
          "postDate": "2023-08-25T11:38:23.710Z",
          "content": "<p>Incorporating a bidirectional GRU improved the score, though I opted not to use it because of the extended training time. I set the learning rate at 0.003 and employed a warm-up phase for 30 epochs.</p>",
          "rawMarkdown": "Incorporating a bidirectional GRU improved the score, though I opted not to use it because of the extended training time. I set the learning rate at 0.003 and employed a warm-up phase for 30 epochs."
        }
      ]
    },
    {
      "id": 2407432,
      "postDate": "2023-08-25T04:57:45.677Z",
      "content": "<p>Your augumentaion  concatenated and aligning char and frame is cool, did these help improve a lot?<br>\nAnd also which method do you use to align char and frame? Does it has high accuracy?</p>",
      "rawMarkdown": "Your augumentaion  concatenated and aligning char and frame is cool, did these help improve a lot?\nAnd also which method do you use to align char and frame? Does it has high accuracy?",
      "replies": [
        {
          "id": 2407538,
          "postDate": "2023-08-25T05:58:59.700Z",
          "content": "<p>As mentioned I didn't maintain logs for the experiments, I can't share the precise number. It would be around 0.002~0.005 I guess.</p>\n<blockquote>\n  <p>And also which method do you use to align char and frame? Does it has high accuracy?</p>\n</blockquote>\n<p>I used one of the trained model using CTC loss. Let's say the phrase is 'dog' and prediction before merging is something like 'dd…ooo….g..', I got 'dd..' for d, '.ooo..' for o, and '..g..' for g. </p>",
          "rawMarkdown": "As mentioned I didn't maintain logs for the experiments, I can't share the precise number. It would be around 0.002~0.005 I guess.\n\n> And also which method do you use to align char and frame? Does it has high accuracy?\n\nI used one of the trained model using CTC loss. Let's say the phrase is 'dog' and prediction before merging is something like 'dd...ooo....g..', I got 'dd..' for d, '.ooo..' for o, and '..g..' for g. ",
          "votes": 1,
          "replies": [
            {
              "id": 2407755,
              "postDate": "2023-08-25T08:38:23.760Z",
              "content": "<p>Thansk <a href=\"https://www.kaggle.com/pudae81\" target=\"_blank\">@pudae81</a> impressive idea, you join the game late and you have done so much. Your sharing help a lot!</p>",
              "rawMarkdown": "Thansk @pudae81 impressive idea, you join the game late and you have done so much. Your sharing help a lot!"
            }
          ]
        }
      ]
    },
    {
      "id": 2407416,
      "postDate": "2023-08-25T04:36:23.467Z",
      "content": "<p>thanks for the writeup and congrats to your results.</p>\n<p>how any epoch did you train?</p>",
      "rawMarkdown": "thanks for the writeup and congrats to your results.\n\nhow any epoch did you train?\n",
      "replies": [
        {
          "id": 2407543,
          "postDate": "2023-08-25T06:00:17.173Z",
          "content": "<p>I trained around 200~250 epoch. More training epochs led higher CV score, but LB was decreased.</p>",
          "rawMarkdown": "I trained around 200~250 epoch. More training epochs led higher CV score, but LB was decreased."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2407595,
      "author_name": "C R Suthikshn Kumar",
      "author_url": "",
      "post_date": "2023-08-25T06:39:34.877000",
      "content": "<p>Congratulations. Thanks for sharing the details of your solution. Did you have LSTM in your pipeline? Any changes to learning rate and warmup?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2408017,
          "author_name": "pudae",
          "author_url": "",
          "post_date": "2023-08-25T11:38:23.710000",
          "content": "<p>Incorporating a bidirectional GRU improved the score, though I opted not to use it because of the extended training time. I set the learning rate at 0.003 and employed a warm-up phase for 30 epochs.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2407432,
      "author_name": "gezi",
      "author_url": "",
      "post_date": "2023-08-25T04:57:45.677000",
      "content": "<p>Your augumentaion  concatenated and aligning char and frame is cool, did these help improve a lot?<br>\nAnd also which method do you use to align char and frame? Does it has high accuracy?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2407538,
          "author_name": "pudae",
          "author_url": "",
          "post_date": "2023-08-25T05:58:59.700000",
          "content": "<p>As mentioned I didn't maintain logs for the experiments, I can't share the precise number. It would be around 0.002~0.005 I guess.</p>\n<blockquote>\n  <p>And also which method do you use to align char and frame? Does it has high accuracy?</p>\n</blockquote>\n<p>I used one of the trained model using CTC loss. Let's say the phrase is 'dog' and prediction before merging is something like 'dd…ooo….g..', I got 'dd..' for d, '.ooo..' for o, and '..g..' for g. </p>",
          "votes": 1,
          "replies": [
            {
              "id": 2407755,
              "author_name": "gezi",
              "author_url": "",
              "post_date": "2023-08-25T08:38:23.760000",
              "content": "<p>Thansk <a href=\"https://www.kaggle.com/pudae81\" target=\"_blank\">@pudae81</a> impressive idea, you join the game late and you have done so much. Your sharing help a lot!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2407416,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-08-25T04:36:23.467000",
      "content": "<p>thanks for the writeup and congrats to your results.</p>\n<p>how any epoch did you train?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2407543,
          "author_name": "pudae",
          "author_url": "",
          "post_date": "2023-08-25T06:00:17.173000",
          "content": "<p>I trained around 200~250 epoch. More training epochs led higher CV score, but LB was decreased.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2407398": "Firstly, I'd like to extend my gratitude to Kaggle and Google for hosting such an engaging competition, and congratulations to all the winners.\n\nI'm keen to share a few methodologies that yielded a notable enhancement in the competition score. Unfortunately, I didn't maintain logs for my experiments, so I'm unable to quantify the exact impact of each modification. My apologies for this oversight.\n\nGiven that the challenge bore a resemblance to speech recognition, I initiated my approach using the Conformer architecture paired with the CTC loss.\n\n**Baseline Configuration**\n- **Landmark Encoder**: 2-layer 1D CNN\n- **Architecture**: 4-layer Conformer\n- **Loss Function**: CTC Loss\n- **Features**: Dominant hand and pose, lip\n- **Augmentations**: Rotation, resizing, shearing, random masking\n\n**Key Modifications that Enhanced Performance**\n- **Architecture Update**: \n  - Transitioned from Conformer to SqueezeFormer\n  - 2-layer Separable Convolution followed by 1D CNN for landmark encoder\n  - Replaced GLU with swish (worked better under the same number of parameters)\n     - Further transitioning from swish to mish provided incremental benefits\n  - Temporal reduction before the 4th block\n     - implemented using average pooling 1d or max pooling 1d\n  - Additional forward path omitting temporal reduction(parameters are shared) and applying the CTC loss (only for training)\n- **Augmentation**\n  - concatenated two randomly sampled examples\n  - For the noisy examples (whose nonnan frames are less than the length of phrase), generate synthetic frames.\n     - aligning char and frame using another trained model and create char-frame data from them.\n     - for each character in the phrase, randomly sample frame from the char-frame data and concatenate them.\n\nI apologize again for the lack of detailed information. Despite the lack of exhaustive specifics, I hope these are helpful for you.\n\nThanks.\n ",
    "2407595": "Congratulations. Thanks for sharing the details of your solution. Did you have LSTM in your pipeline? Any changes to learning rate and warmup?",
    "2407432": "Your augumentaion  concatenated and aligning char and frame is cool, did these help improve a lot?\nAnd also which method do you use to align char and frame? Does it has high accuracy?",
    "2407416": "thanks for the writeup and congrats to your results.\n\nhow any epoch did you train?\n"
  }
}