{
  "id": 434409,
  "title": "33rd place solution - LB 0.758",
  "url": "/competitions/asl-fingerspelling/discussion/434409",
  "author_name": "Wonjun",
  "post_date": "2023-08-25T05:32:13.783000",
  "votes": 13,
  "comment_count": 3,
  "views": 0,
  "content": "<h2>TL;DR</h2>\n<p>We used a conformer-like model consisting transformer encoder, 1d convolution with CBAM and bi-LSTM.<br>\nOverall model size is about 16MB after INT8 quantization.<br>\nTraining objective is CTC with InterCTC loss.</p>\n<h2>Data preprocess</h2>\n<ul>\n<li>Used landmarks<ul>\n<li>1 for nose, 21 for dominant hand, 40 for lips</li>\n<li>x, y coordinates</li></ul></li>\n<li>Normalization<ul>\n<li>Standardized distance from nose coordinates</li></ul></li>\n<li>Feature engineering<ul>\n<li>Concatenation of normalized location, difference of next frame, joint distance of hand</li>\n<li>Total 582 dims</li></ul></li>\n<li>Removing data that input frame is shorter than 2 times of target phrase</li>\n</ul>\n<h2>Data augmentation</h2>\n<ul>\n<li>Horizontal flip landmarks</li>\n<li>Interpolation</li>\n<li>Affine transfrom</li>\n</ul>\n<h2>Model</h2>\n<ul>\n<li>2 stacked encoder with Transformer, 1D convolution with CBAM and Bi-LSTM<ul>\n<li>hidden dim: 352</li></ul></li>\n<li>CTC loss and Inter CTC loss after first encoder</li>\n<li>17M parameters and INT8 quantization</li>\n</ul>\n<h2>Train</h2>\n<ul>\n<li>2 staged training (300 + 200 epochs)<ul>\n<li>Use supplemental and train data in first 300 epochs</li>\n<li>Use train data 200 epochs</li></ul></li>\n<li>Ranger optimizer</li>\n<li>Cosine decay scheduler with 12 epochs warmup</li>\n<li>AWP<ul>\n<li>It prevents the validation loss from diverging, but it doesn't seem to improve the edit distance.</li>\n<li>Used to train long epochs.</li></ul></li>\n</ul>\n<h2>Not worked</h2>\n<ul>\n<li>Augmentation<ul>\n<li>Time, spatial, landmark masking</li>\n<li>Time reverse</li></ul></li>\n<li>Autoregressive decoder<ul>\n<li>Use joint loss, CTC loss for encoder and Crossentropy for decoder</li>\n<li>Inference time is longer and the number of parameters larger than those of the CTC encoder alone, but the performance improvement is not clear, so it is not used.</li></ul></li>\n</ul>",
  "messages": [
    {
      "id": 2407484,
      "postDate": "2023-08-25T05:32:13.783Z",
      "content": "<h2>TL;DR</h2>\n<p>We used a conformer-like model consisting transformer encoder, 1d convolution with CBAM and bi-LSTM.<br>\nOverall model size is about 16MB after INT8 quantization.<br>\nTraining objective is CTC with InterCTC loss.</p>\n<h2>Data preprocess</h2>\n<ul>\n<li>Used landmarks<ul>\n<li>1 for nose, 21 for dominant hand, 40 for lips</li>\n<li>x, y coordinates</li></ul></li>\n<li>Normalization<ul>\n<li>Standardized distance from nose coordinates</li></ul></li>\n<li>Feature engineering<ul>\n<li>Concatenation of normalized location, difference of next frame, joint distance of hand</li>\n<li>Total 582 dims</li></ul></li>\n<li>Removing data that input frame is shorter than 2 times of target phrase</li>\n</ul>\n<h2>Data augmentation</h2>\n<ul>\n<li>Horizontal flip landmarks</li>\n<li>Interpolation</li>\n<li>Affine transfrom</li>\n</ul>\n<h2>Model</h2>\n<ul>\n<li>2 stacked encoder with Transformer, 1D convolution with CBAM and Bi-LSTM<ul>\n<li>hidden dim: 352</li></ul></li>\n<li>CTC loss and Inter CTC loss after first encoder</li>\n<li>17M parameters and INT8 quantization</li>\n</ul>\n<h2>Train</h2>\n<ul>\n<li>2 staged training (300 + 200 epochs)<ul>\n<li>Use supplemental and train data in first 300 epochs</li>\n<li>Use train data 200 epochs</li></ul></li>\n<li>Ranger optimizer</li>\n<li>Cosine decay scheduler with 12 epochs warmup</li>\n<li>AWP<ul>\n<li>It prevents the validation loss from diverging, but it doesn't seem to improve the edit distance.</li>\n<li>Used to train long epochs.</li></ul></li>\n</ul>\n<h2>Not worked</h2>\n<ul>\n<li>Augmentation<ul>\n<li>Time, spatial, landmark masking</li>\n<li>Time reverse</li></ul></li>\n<li>Autoregressive decoder<ul>\n<li>Use joint loss, CTC loss for encoder and Crossentropy for decoder</li>\n<li>Inference time is longer and the number of parameters larger than those of the CTC encoder alone, but the performance improvement is not clear, so it is not used.</li></ul></li>\n</ul>",
      "rawMarkdown": "## TL;DR\nWe used a conformer-like model consisting transformer encoder, 1d convolution with CBAM and bi-LSTM.\nOverall model size is about 16MB after INT8 quantization.\nTraining objective is CTC with InterCTC loss.\n\n## Data preprocess\n- Used landmarks\n    - 1 for nose, 21 for dominant hand, 40 for lips\n    - x, y coordinates\n- Normalization\n    - Standardized distance from nose coordinates\n- Feature engineering\n    - Concatenation of normalized location, difference of next frame, joint distance of hand\n    - Total 582 dims\n- Removing data that input frame is shorter than 2 times of target phrase\n\n## Data augmentation\n- Horizontal flip landmarks\n- Interpolation\n- Affine transfrom\n\n## Model\n- 2 stacked encoder with Transformer, 1D convolution with CBAM and Bi-LSTM\n    - hidden dim: 352\n- CTC loss and Inter CTC loss after first encoder\n- 17M parameters and INT8 quantization\n\n## Train\n- 2 staged training (300 + 200 epochs)\n    - Use supplemental and train data in first 300 epochs\n    - Use train data 200 epochs\n- Ranger optimizer\n- Cosine decay scheduler with 12 epochs warmup\n- AWP\n    - It prevents the validation loss from diverging, but it doesn't seem to improve the edit distance.\n    - Used to train long epochs.\n\n## Not worked\n- Augmentation\n    - Time, spatial, landmark masking\n    - Time reverse\n- Autoregressive decoder\n    - Use joint loss, CTC loss for encoder and Crossentropy for decoder\n    - Inference time is longer and the number of parameters larger than those of the CTC encoder alone, but the performance improvement is not clear, so it is not used.",
      "votes": 13
    },
    {
      "id": 2408930,
      "postDate": "2023-08-25T23:33:10.303Z",
      "content": "<p>Congratulations Giant Penguin and team! </p>\n<p>I'm surprised that time reverse works. How is it implemented? Do you do <code>frames[::-1]</code> and reverse the phrase? How much CV and LB boost does it give?</p>",
      "rawMarkdown": "Congratulations Giant Penguin and team! \n\nI'm surprised that time reverse works. How is it implemented? Do you do `frames[::-1]` and reverse the phrase? How much CV and LB boost does it give?",
      "votes": 1,
      "replies": [
        {
          "id": 2409389,
          "postDate": "2023-08-26T08:06:17.377Z",
          "content": "<p>Thank you :)<br>\nYes, we implemented it as you mentioned, but it didn't work.<br>\nThere are moving gesture characters such as 'z' and we guess reversed motion may interrupt learning.<br>\nAugmentation of phrase consiting with only static gesture may helpful but we didn't try it</p>",
          "rawMarkdown": "Thank you :)\nYes, we implemented it as you mentioned, but it didn't work.\nThere are moving gesture characters such as 'z' and we guess reversed motion may interrupt learning.\nAugmentation of phrase consiting with only static gesture may helpful but we didn't try it"
        }
      ]
    },
    {
      "id": 2407590,
      "postDate": "2023-08-25T06:37:38.370Z",
      "content": "<p>Congratulations. Thanks for sharing the details of your solution. </p>",
      "rawMarkdown": "Congratulations. Thanks for sharing the details of your solution. ",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 2408930,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2023-08-25T23:33:10.303000",
      "content": "<p>Congratulations Giant Penguin and team! </p>\n<p>I'm surprised that time reverse works. How is it implemented? Do you do <code>frames[::-1]</code> and reverse the phrase? How much CV and LB boost does it give?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2409389,
          "author_name": "Wonjun",
          "author_url": "",
          "post_date": "2023-08-26T08:06:17.377000",
          "content": "<p>Thank you :)<br>\nYes, we implemented it as you mentioned, but it didn't work.<br>\nThere are moving gesture characters such as 'z' and we guess reversed motion may interrupt learning.<br>\nAugmentation of phrase consiting with only static gesture may helpful but we didn't try it</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2407590,
      "author_name": "C R Suthikshn Kumar",
      "author_url": "",
      "post_date": "2023-08-25T06:37:38.370000",
      "content": "<p>Congratulations. Thanks for sharing the details of your solution. </p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2407484": "## TL;DR\nWe used a conformer-like model consisting transformer encoder, 1d convolution with CBAM and bi-LSTM.\nOverall model size is about 16MB after INT8 quantization.\nTraining objective is CTC with InterCTC loss.\n\n## Data preprocess\n- Used landmarks\n    - 1 for nose, 21 for dominant hand, 40 for lips\n    - x, y coordinates\n- Normalization\n    - Standardized distance from nose coordinates\n- Feature engineering\n    - Concatenation of normalized location, difference of next frame, joint distance of hand\n    - Total 582 dims\n- Removing data that input frame is shorter than 2 times of target phrase\n\n## Data augmentation\n- Horizontal flip landmarks\n- Interpolation\n- Affine transfrom\n\n## Model\n- 2 stacked encoder with Transformer, 1D convolution with CBAM and Bi-LSTM\n    - hidden dim: 352\n- CTC loss and Inter CTC loss after first encoder\n- 17M parameters and INT8 quantization\n\n## Train\n- 2 staged training (300 + 200 epochs)\n    - Use supplemental and train data in first 300 epochs\n    - Use train data 200 epochs\n- Ranger optimizer\n- Cosine decay scheduler with 12 epochs warmup\n- AWP\n    - It prevents the validation loss from diverging, but it doesn't seem to improve the edit distance.\n    - Used to train long epochs.\n\n## Not worked\n- Augmentation\n    - Time, spatial, landmark masking\n    - Time reverse\n- Autoregressive decoder\n    - Use joint loss, CTC loss for encoder and Crossentropy for decoder\n    - Inference time is longer and the number of parameters larger than those of the CTC encoder alone, but the performance improvement is not clear, so it is not used.",
    "2408930": "Congratulations Giant Penguin and team! \n\nI'm surprised that time reverse works. How is it implemented? Do you do `frames[::-1]` and reverse the phrase? How much CV and LB boost does it give?",
    "2407590": "Congratulations. Thanks for sharing the details of your solution. "
  }
}