{
  "id": 437728,
  "title": "29th Place Solution - LB 0.762",
  "url": "/competitions/asl-fingerspelling/discussion/437728",
  "author_name": "ohkawa3",
  "post_date": "2023-09-07T23:46:30.584000",
  "votes": 2,
  "comment_count": 0,
  "views": 0,
  "content": "<h2>TL;DR</h2>\n<p>We used a 1D convolutional network with 122-dimensional features as input and Self Conditional CTC, along with a greedy search as the decoder.</p>\n<h2>Preprocess</h2>\n<p>We used xy feature points of the hand performing the gesture and xy feature points of the lips, for a total of 61 xy feature points (= 122-dimensional vector).</p>\n<h2>Data Augmentation</h2>\n<p>Overfitting was suppressed by randomly dropping hand feature points. The percentage of dropping was 40% of the total, and at least 5 consecutive frames were dropped.</p>\n<h2>Model</h2>\n<p>We used a 1DCNN consisting of 9 Residual Blocks of kernel width 3 and 12 Residual Blocks of kernel width 31.</p>\n<h2>Loss Function</h2>\n<p>We used Self Conditional CTC, a variant of CTC, for the loss function. By using Self Conditional CTC, a significant performance improvement was seen without increasing the amount of calculation.<br>\n<a href=\"https://arxiv.org/abs/2104.02724\" target=\"_blank\">Read more about Self Conditional CTC here</a></p>\n<h2>Training</h2>\n<ul>\n<li>60 epochs</li>\n<li>CosDecay learning rate</li>\n<li>AdamW optimizer</li>\n<li>AWP</li>\n</ul>\n<h2>Convert Model</h2>\n<p>All implementation was done in PyTorch and converted to 16-bit quantized tflite by onnx2tf. Conversion using nobuco was also tried, but was not adopted because, contrary to expectations, it increased processing time.</p>\n<h2>Not Worked</h2>\n<ul>\n<li>Stochastic Weight Averaging was very effective in the last competition but did not contribute to performance in this competition.</li>\n<li>We thought that Levenshtein OCR could be used as a decoder and tried to implement it, but it did not work.</li>\n</ul>",
  "messages": [
    {
      "id": 2428494,
      "postDate": "2023-09-07T23:46:30.583Z",
      "content": "<h2>TL;DR</h2>\n<p>We used a 1D convolutional network with 122-dimensional features as input and Self Conditional CTC, along with a greedy search as the decoder.</p>\n<h2>Preprocess</h2>\n<p>We used xy feature points of the hand performing the gesture and xy feature points of the lips, for a total of 61 xy feature points (= 122-dimensional vector).</p>\n<h2>Data Augmentation</h2>\n<p>Overfitting was suppressed by randomly dropping hand feature points. The percentage of dropping was 40% of the total, and at least 5 consecutive frames were dropped.</p>\n<h2>Model</h2>\n<p>We used a 1DCNN consisting of 9 Residual Blocks of kernel width 3 and 12 Residual Blocks of kernel width 31.</p>\n<h2>Loss Function</h2>\n<p>We used Self Conditional CTC, a variant of CTC, for the loss function. By using Self Conditional CTC, a significant performance improvement was seen without increasing the amount of calculation.<br>\n<a href=\"https://arxiv.org/abs/2104.02724\" target=\"_blank\">Read more about Self Conditional CTC here</a></p>\n<h2>Training</h2>\n<ul>\n<li>60 epochs</li>\n<li>CosDecay learning rate</li>\n<li>AdamW optimizer</li>\n<li>AWP</li>\n</ul>\n<h2>Convert Model</h2>\n<p>All implementation was done in PyTorch and converted to 16-bit quantized tflite by onnx2tf. Conversion using nobuco was also tried, but was not adopted because, contrary to expectations, it increased processing time.</p>\n<h2>Not Worked</h2>\n<ul>\n<li>Stochastic Weight Averaging was very effective in the last competition but did not contribute to performance in this competition.</li>\n<li>We thought that Levenshtein OCR could be used as a decoder and tried to implement it, but it did not work.</li>\n</ul>",
      "rawMarkdown": "\n## TL;DR\n\nWe used a 1D convolutional network with 122-dimensional features as input and Self Conditional CTC, along with a greedy search as the decoder.\n\n## Preprocess\n\nWe used xy feature points of the hand performing the gesture and xy feature points of the lips, for a total of 61 xy feature points (= 122-dimensional vector).\n\n## Data Augmentation\n\nOverfitting was suppressed by randomly dropping hand feature points. The percentage of dropping was 40% of the total, and at least 5 consecutive frames were dropped.\n\n## Model\n\nWe used a 1DCNN consisting of 9 Residual Blocks of kernel width 3 and 12 Residual Blocks of kernel width 31.\n\n## Loss Function\n\nWe used Self Conditional CTC, a variant of CTC, for the loss function. By using Self Conditional CTC, a significant performance improvement was seen without increasing the amount of calculation.\n\n[Read more about Self Conditional CTC here](https://arxiv.org/abs/2104.02724)\n\n## Training\n\n- 60 epochs\n- CosDecay learning rate\n- AdamW optimizer\n- AWP\n\n## Convert Model\n\nAll implementation was done in PyTorch and converted to 16-bit quantized tflite by onnx2tf. Conversion using nobuco was also tried, but was not adopted because, contrary to expectations, it increased processing time.\n\n## Not Worked\n\n- Stochastic Weight Averaging was very effective in the last competition but did not contribute to performance in this competition.\n- We thought that Levenshtein OCR could be used as a decoder and tried to implement it, but it did not work.\n",
      "votes": 2
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2428494": "\n## TL;DR\n\nWe used a 1D convolutional network with 122-dimensional features as input and Self Conditional CTC, along with a greedy search as the decoder.\n\n## Preprocess\n\nWe used xy feature points of the hand performing the gesture and xy feature points of the lips, for a total of 61 xy feature points (= 122-dimensional vector).\n\n## Data Augmentation\n\nOverfitting was suppressed by randomly dropping hand feature points. The percentage of dropping was 40% of the total, and at least 5 consecutive frames were dropped.\n\n## Model\n\nWe used a 1DCNN consisting of 9 Residual Blocks of kernel width 3 and 12 Residual Blocks of kernel width 31.\n\n## Loss Function\n\nWe used Self Conditional CTC, a variant of CTC, for the loss function. By using Self Conditional CTC, a significant performance improvement was seen without increasing the amount of calculation.\n\n[Read more about Self Conditional CTC here](https://arxiv.org/abs/2104.02724)\n\n## Training\n\n- 60 epochs\n- CosDecay learning rate\n- AdamW optimizer\n- AWP\n\n## Convert Model\n\nAll implementation was done in PyTorch and converted to 16-bit quantized tflite by onnx2tf. Conversion using nobuco was also tried, but was not adopted because, contrary to expectations, it increased processing time.\n\n## Not Worked\n\n- Stochastic Weight Averaging was very effective in the last competition but did not contribute to performance in this competition.\n- We thought that Levenshtein OCR could be used as a decoder and tried to implement it, but it did not work.\n"
  }
}