{
  "id": 434364,
  "title": "17th Solution: Conformer + CTCLoss +  500 epoch training",
  "url": "/competitions/asl-fingerspelling/discussion/434364",
  "author_name": "Yu Wu",
  "post_date": "2023-08-25T01:08:22.451000",
  "votes": 20,
  "comment_count": 3,
  "views": 0,
  "content": "<h1>TL; DR</h1>\n<p>Fixed-length input (220 frames), padding shorter inputs and resizing longer inputs.<br>\n2 layer MLP landmark encoder + 6 layer 384-dim Conformer + 1 layer GRU and 500 epochs training (takes ~8 hours on Kaggle TPUs).<br>\nPost-processing (+0.003 in CV and LB) by:</p>\n<pre><code> len() &lt;= :\n    =  + \n</code></pre>\n<p>My notebook is available at <a href=\"https://www.kaggle.com/code/nightsh4de/ctc-transformer/notebook\" target=\"_blank\">https://www.kaggle.com/code/nightsh4de/ctc-transformer/notebook</a><br>\nIt is primarily based on the previous 1st solution <a href=\"https://www.kaggle.com/code/hoyso48/1st-place-solution-training\" target=\"_blank\">https://www.kaggle.com/code/hoyso48/1st-place-solution-training</a> and some public notebooks of this competition:<br>\n<a href=\"https://www.kaggle.com/code/markwijkhuizen/aslfr-transformer-training-inference#Landmark-Embedding\" target=\"_blank\">https://www.kaggle.com/code/markwijkhuizen/aslfr-transformer-training-inference#Landmark-Embedding</a><br>\n<a href=\"https://www.kaggle.com/code/irohith/aslfr-transformer\" target=\"_blank\">https://www.kaggle.com/code/irohith/aslfr-transformer</a> <br>\n<a href=\"https://www.kaggle.com/code/shlomoron/aslfr-ctc-on-tpu\" target=\"_blank\">https://www.kaggle.com/code/shlomoron/aslfr-ctc-on-tpu</a>. Many many thanks to them.</p>\n<h1>Data Preprocessing and Augmentation</h1>\n<p>Basically the same as the previous 1st solution. But I found using 3d positions (means including depth) and pose landmarks gives better CV and LB score.<br>\nMy input: left-right hand, eye, nose, lips and pose landmarks.</p>\n<h1>Model</h1>\n<p>I believe this task would be quite similar to Automatic Speech Recognition (ASR), so I used Conformer <a href=\"https://arxiv.org/abs/2005.08100\" target=\"_blank\">https://arxiv.org/abs/2005.08100</a></p>\n<h1>Post-preprocessing</h1>\n<ol>\n<li>I checked my worst predictions in the validation set and found that shorter predictions are worse. And most very short predictions (length less than 5) are basically predicting nothing, e.g. single characters like \"a\" or space \" \".</li>\n<li>I found that only few label's length is less or equal to 5. </li>\n<li>So adding some make-up phrases to very short predictions might be a good idea.</li>\n<li>I picked the most common chars in training set: \"a\", \"e\", \"r\", \"o\", \"-\", \" \".</li>\n<li>I test all the combinations of \"a\", \"e\", \"r\", \"o\", \"-\", \" \" based on validation set and \" -aero\" is the best.</li>\n</ol>\n<h1>What not worked</h1>\n<ol>\n<li>Autoregressive transformer models. <br>\nI spent most of my time on transformers. But they strongly overfits the phrase and I couldn't find a way to solve it.<br>\nThe problem is , for an input X[1…n] and its phrase \"123456789\". We may expect our model predict \"12345\" for the first half input X[1…n/2]. <br>\nFor CTC models, it is true. But my autoregressive model always makes half correct prediction \"12345\" + half random incorrect prediction e.g. some random 5-digit numbers.</li>\n<li>Masking for variable-length input<br>\nI tried to use the same masking techniques as the previous 1st solution <a href=\"https://www.kaggle.com/code/hoyso48/1st-place-solution-training\" target=\"_blank\">https://www.kaggle.com/code/hoyso48/1st-place-solution-training</a>, to support longer input frames.<br>\nHowever, masking models always produce lower CV and LB scores (-0.02). I believe it is due to the masking limits the output length, but shorter inputs still require a longer output space. Although, unfortunately, I don't have time to verify it, I believe it might enable much better solutions.</li>\n<li>Empty embedding for fixed-length input<br>\nI tried to use learnable constant weights for padded empty frames and 2-layer MLP for encoding landmarks in original frames. It seems to me makes more sense, but CV score is worse than direct encoding all frames. </li>\n</ol>",
  "messages": [
    {
      "id": 2407228,
      "postDate": "2023-08-25T01:08:22.450Z",
      "content": "<h1>TL; DR</h1>\n<p>Fixed-length input (220 frames), padding shorter inputs and resizing longer inputs.<br>\n2 layer MLP landmark encoder + 6 layer 384-dim Conformer + 1 layer GRU and 500 epochs training (takes ~8 hours on Kaggle TPUs).<br>\nPost-processing (+0.003 in CV and LB) by:</p>\n<pre><code> len() &lt;= :\n    =  + \n</code></pre>\n<p>My notebook is available at <a href=\"https://www.kaggle.com/code/nightsh4de/ctc-transformer/notebook\" target=\"_blank\">https://www.kaggle.com/code/nightsh4de/ctc-transformer/notebook</a><br>\nIt is primarily based on the previous 1st solution <a href=\"https://www.kaggle.com/code/hoyso48/1st-place-solution-training\" target=\"_blank\">https://www.kaggle.com/code/hoyso48/1st-place-solution-training</a> and some public notebooks of this competition:<br>\n<a href=\"https://www.kaggle.com/code/markwijkhuizen/aslfr-transformer-training-inference#Landmark-Embedding\" target=\"_blank\">https://www.kaggle.com/code/markwijkhuizen/aslfr-transformer-training-inference#Landmark-Embedding</a><br>\n<a href=\"https://www.kaggle.com/code/irohith/aslfr-transformer\" target=\"_blank\">https://www.kaggle.com/code/irohith/aslfr-transformer</a> <br>\n<a href=\"https://www.kaggle.com/code/shlomoron/aslfr-ctc-on-tpu\" target=\"_blank\">https://www.kaggle.com/code/shlomoron/aslfr-ctc-on-tpu</a>. Many many thanks to them.</p>\n<h1>Data Preprocessing and Augmentation</h1>\n<p>Basically the same as the previous 1st solution. But I found using 3d positions (means including depth) and pose landmarks gives better CV and LB score.<br>\nMy input: left-right hand, eye, nose, lips and pose landmarks.</p>\n<h1>Model</h1>\n<p>I believe this task would be quite similar to Automatic Speech Recognition (ASR), so I used Conformer <a href=\"https://arxiv.org/abs/2005.08100\" target=\"_blank\">https://arxiv.org/abs/2005.08100</a></p>\n<h1>Post-preprocessing</h1>\n<ol>\n<li>I checked my worst predictions in the validation set and found that shorter predictions are worse. And most very short predictions (length less than 5) are basically predicting nothing, e.g. single characters like \"a\" or space \" \".</li>\n<li>I found that only few label's length is less or equal to 5. </li>\n<li>So adding some make-up phrases to very short predictions might be a good idea.</li>\n<li>I picked the most common chars in training set: \"a\", \"e\", \"r\", \"o\", \"-\", \" \".</li>\n<li>I test all the combinations of \"a\", \"e\", \"r\", \"o\", \"-\", \" \" based on validation set and \" -aero\" is the best.</li>\n</ol>\n<h1>What not worked</h1>\n<ol>\n<li>Autoregressive transformer models. <br>\nI spent most of my time on transformers. But they strongly overfits the phrase and I couldn't find a way to solve it.<br>\nThe problem is , for an input X[1…n] and its phrase \"123456789\". We may expect our model predict \"12345\" for the first half input X[1…n/2]. <br>\nFor CTC models, it is true. But my autoregressive model always makes half correct prediction \"12345\" + half random incorrect prediction e.g. some random 5-digit numbers.</li>\n<li>Masking for variable-length input<br>\nI tried to use the same masking techniques as the previous 1st solution <a href=\"https://www.kaggle.com/code/hoyso48/1st-place-solution-training\" target=\"_blank\">https://www.kaggle.com/code/hoyso48/1st-place-solution-training</a>, to support longer input frames.<br>\nHowever, masking models always produce lower CV and LB scores (-0.02). I believe it is due to the masking limits the output length, but shorter inputs still require a longer output space. Although, unfortunately, I don't have time to verify it, I believe it might enable much better solutions.</li>\n<li>Empty embedding for fixed-length input<br>\nI tried to use learnable constant weights for padded empty frames and 2-layer MLP for encoding landmarks in original frames. It seems to me makes more sense, but CV score is worse than direct encoding all frames. </li>\n</ol>",
      "rawMarkdown": "\n# TL; DR\n\nFixed-length input (220 frames), padding shorter inputs and resizing longer inputs.\n\n2 layer MLP landmark encoder + 6 layer 384-dim Conformer + 1 layer GRU and 500 epochs training (takes ~8 hours on Kaggle TPUs).\n\nPost-processing (+0.003 in CV and LB) by:\n```\nif len(pred) <= 4:\n    pred = pred + \" -aero\"\n```\n\nMy notebook is available at https://www.kaggle.com/code/nightsh4de/ctc-transformer/notebook\nIt is primarily based on the previous 1st solution https://www.kaggle.com/code/hoyso48/1st-place-solution-training and some public notebooks of this competition:\n https://www.kaggle.com/code/markwijkhuizen/aslfr-transformer-training-inference#Landmark-Embedding\nhttps://www.kaggle.com/code/irohith/aslfr-transformer \nhttps://www.kaggle.com/code/shlomoron/aslfr-ctc-on-tpu. Many many thanks to them.\n\n# Data Preprocessing and Augmentation\n\nBasically the same as the previous 1st solution. But I found using 3d positions (means including depth) and pose landmarks gives better CV and LB score.\n\nMy input: left-right hand, eye, nose, lips and pose landmarks.\n\n# Model\n\nI believe this task would be quite similar to Automatic Speech Recognition (ASR), so I used Conformer https://arxiv.org/abs/2005.08100\n\n# Post-preprocessing\n\n1. I checked my worst predictions in the validation set and found that shorter predictions are worse. And most very short predictions (length less than 5) are basically predicting nothing, e.g. single characters like \"a\" or space \" \".\n2. I found that only few label's length is less or equal to 5. \n3. So adding some make-up phrases to very short predictions might be a good idea.\n3. I picked the most common chars in training set: \"a\", \"e\", \"r\", \"o\", \"-\", \" \".\n4. I test all the combinations of \"a\", \"e\", \"r\", \"o\", \"-\", \" \" based on validation set and \" -aero\" is the best.\n\n# What not worked\n\n1. Autoregressive transformer models. \n\nI spent most of my time on transformers. But they strongly overfits the phrase and I couldn't find a way to solve it.\n\nThe problem is , for an input X[1...n] and its phrase \"123456789\". We may expect our model predict \"12345\" for the first half input X[1...n/2]. \n\nFor CTC models, it is true. But my autoregressive model always makes half correct prediction \"12345\" + half random incorrect prediction e.g. some random 5-digit numbers.\n\n2. Masking for variable-length input\n\nI tried to use the same masking techniques as the previous 1st solution https://www.kaggle.com/code/hoyso48/1st-place-solution-training, to support longer input frames.\n\nHowever, masking models always produce lower CV and LB scores (-0.02). I believe it is due to the masking limits the output length, but shorter inputs still require a longer output space. Although, unfortunately, I don't have time to verify it, I believe it might enable much better solutions.\n\n3. Empty embedding for fixed-length input\n\nI tried to use learnable constant weights for padded empty frames and 2-layer MLP for encoding landmarks in original frames. It seems to me makes more sense, but CV score is worse than direct encoding all frames. \n\n\n",
      "votes": 20
    },
    {
      "id": 2408356,
      "postDate": "2023-08-25T15:31:39.840Z",
      "content": "<p>Congratulations. Thanks for sharing very interesting details.</p>",
      "rawMarkdown": "Congratulations. Thanks for sharing very interesting details.",
      "votes": 1
    },
    {
      "id": 2407250,
      "postDate": "2023-08-25T01:40:56.083Z",
      "content": "<p>\"Masking for variable-length input …\"<br>\nactually you can do without masking. for example, whisper ASR uses fixed size 30 sec audio segments. it pads if it is too short.<br>\ni find this kind of framework simplifies model and training. but there is one parameter that can be learned: the pad value (more correctly, pad embedding)</p>",
      "rawMarkdown": "\"Masking for variable-length input ...\"\nactually you can do without masking. for example, whisper ASR uses fixed size 30 sec audio segments. it pads if it is too short.\ni find this kind of framework simplifies model and training. but there is one parameter that can be learned: the pad value (more correctly, pad embedding)",
      "replies": [
        {
          "id": 2407259,
          "postDate": "2023-08-25T01:54:45.903Z",
          "content": "<p>But masking is faster when processing shorter inputs while being able to take longer inputs without resizing. Am I right?</p>",
          "rawMarkdown": "But masking is faster when processing shorter inputs while being able to take longer inputs without resizing. Am I right?"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2408356,
      "author_name": "C R Suthikshn Kumar",
      "author_url": "",
      "post_date": "2023-08-25T15:31:39.840000",
      "content": "<p>Congratulations. Thanks for sharing very interesting details.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2407250,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-08-25T01:40:56.083000",
      "content": "<p>\"Masking for variable-length input …\"<br>\nactually you can do without masking. for example, whisper ASR uses fixed size 30 sec audio segments. it pads if it is too short.<br>\ni find this kind of framework simplifies model and training. but there is one parameter that can be learned: the pad value (more correctly, pad embedding)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2407259,
          "author_name": "Yu Wu",
          "author_url": "",
          "post_date": "2023-08-25T01:54:45.903000",
          "content": "<p>But masking is faster when processing shorter inputs while being able to take longer inputs without resizing. Am I right?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2407228": "\n# TL; DR\n\nFixed-length input (220 frames), padding shorter inputs and resizing longer inputs.\n\n2 layer MLP landmark encoder + 6 layer 384-dim Conformer + 1 layer GRU and 500 epochs training (takes ~8 hours on Kaggle TPUs).\n\nPost-processing (+0.003 in CV and LB) by:\n```\nif len(pred) <= 4:\n    pred = pred + \" -aero\"\n```\n\nMy notebook is available at https://www.kaggle.com/code/nightsh4de/ctc-transformer/notebook\nIt is primarily based on the previous 1st solution https://www.kaggle.com/code/hoyso48/1st-place-solution-training and some public notebooks of this competition:\n https://www.kaggle.com/code/markwijkhuizen/aslfr-transformer-training-inference#Landmark-Embedding\nhttps://www.kaggle.com/code/irohith/aslfr-transformer \nhttps://www.kaggle.com/code/shlomoron/aslfr-ctc-on-tpu. Many many thanks to them.\n\n# Data Preprocessing and Augmentation\n\nBasically the same as the previous 1st solution. But I found using 3d positions (means including depth) and pose landmarks gives better CV and LB score.\n\nMy input: left-right hand, eye, nose, lips and pose landmarks.\n\n# Model\n\nI believe this task would be quite similar to Automatic Speech Recognition (ASR), so I used Conformer https://arxiv.org/abs/2005.08100\n\n# Post-preprocessing\n\n1. I checked my worst predictions in the validation set and found that shorter predictions are worse. And most very short predictions (length less than 5) are basically predicting nothing, e.g. single characters like \"a\" or space \" \".\n2. I found that only few label's length is less or equal to 5. \n3. So adding some make-up phrases to very short predictions might be a good idea.\n3. I picked the most common chars in training set: \"a\", \"e\", \"r\", \"o\", \"-\", \" \".\n4. I test all the combinations of \"a\", \"e\", \"r\", \"o\", \"-\", \" \" based on validation set and \" -aero\" is the best.\n\n# What not worked\n\n1. Autoregressive transformer models. \n\nI spent most of my time on transformers. But they strongly overfits the phrase and I couldn't find a way to solve it.\n\nThe problem is , for an input X[1...n] and its phrase \"123456789\". We may expect our model predict \"12345\" for the first half input X[1...n/2]. \n\nFor CTC models, it is true. But my autoregressive model always makes half correct prediction \"12345\" + half random incorrect prediction e.g. some random 5-digit numbers.\n\n2. Masking for variable-length input\n\nI tried to use the same masking techniques as the previous 1st solution https://www.kaggle.com/code/hoyso48/1st-place-solution-training, to support longer input frames.\n\nHowever, masking models always produce lower CV and LB scores (-0.02). I believe it is due to the masking limits the output length, but shorter inputs still require a longer output space. Although, unfortunately, I don't have time to verify it, I believe it might enable much better solutions.\n\n3. Empty embedding for fixed-length input\n\nI tried to use learnable constant weights for padded empty frames and 2-layer MLP for encoding landmarks in original frames. It seems to me makes more sense, but CV score is worse than direct encoding all frames. \n\n\n",
    "2408356": "Congratulations. Thanks for sharing very interesting details.",
    "2407250": "\"Masking for variable-length input ...\"\nactually you can do without masking. for example, whisper ASR uses fixed size 30 sec audio segments. it pads if it is too short.\ni find this kind of framework simplifies model and training. but there is one parameter that can be learned: the pad value (more correctly, pad embedding)"
  }
}