{
  "id": 434393,
  "title": "[3rd place solution with code] 17 layers squeezeformer with timerduce and ROPE ",
  "url": "/competitions/asl-fingerspelling/discussion/434393",
  "author_name": "gezi",
  "post_date": "2023-08-25T04:18:36.233000",
  "votes": 53,
  "comment_count": 32,
  "views": 0,
  "content": "<p>This is an intesting game, I learned a lot from </p>\n<p><a href=\"https://www.kaggle.com/code/hoyso48/1st-place-solution-training\" target=\"_blank\">https://www.kaggle.com/code/hoyso48/1st-place-solution-training</a><br>\n<a href=\"https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place\" target=\"_blank\">https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place</a><br>\n<a href=\"https://www.kaggle.com/code/markwijkhuizen/aslfr-transformer-training-inference\" target=\"_blank\">https://www.kaggle.com/code/markwijkhuizen/aslfr-transformer-training-inference</a></p>\n<h1><strong>Preprocss</strong></h1>\n<ol>\n<li><strong>Do not throw away frames witout hands!</strong></li>\n<li>Use as much info as possible, (hands, lips, eye, nose, pose) , use (x,y,z)</li>\n<li>Use more frames, resize to 320 frames for model training.</li>\n<li>I used <strong>original + normalized + abs position(/1000.)</strong> as input feats(total 384 * 2 + 1 = 769). <br>\nnormalized feats help converge much faster but I found original feats also helps.</li>\n</ol>\n<h1><strong>Aug</strong></h1>\n<p>I followed last competion 1st solution, and found the most important aug are </p>\n<ol>\n<li><strong>time scale (interp1d)</strong></li>\n<li><strong>time dim mask</strong><br>\nI used heavy mask here which will mask some time seq and aslo will <strong>randomly mask 50% frames</strong></li>\n<li>affine <br>\nfollow preve 1st solution with prob 0.75, for left-right flip only with prob 0.25</li>\n</ol>\n<h1><strong>Training</strong></h1>\n<ol>\n<li>I used <strong>train data + sup data combined training with sup data weight set to 0.1</strong> for 400 epochs.</li>\n<li>batch size 128, max lr 2e-3, lr scheduler using linear decay with 0.1 epochs warmup , adam optimizer</li>\n<li>Awp training started from epochs * 0.15, with adv_lr 0.2 and adv_eps 0.</li>\n<li><strong>Fintune with train data only for 10 epochs</strong> with max lr 1e-4, awp start from epoch 2.  </li>\n</ol>\n<h1><strong>Postdeal</strong></h1>\n<p>I used <strong>rule for blank index(0)</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F42245%2F54e452eca19b2ee1610a7ecef4ea7cdc%2F1.png?generation=1692937017778775&amp;alt=media\" alt=\"\"></p>\n<h1><strong>Model</strong></h1>\n<ol>\n<li><strong>Squeezeformer</strong>  works best, it is a bit faster then conformer. (Code from NEMO)</li>\n<li>In order to make net deeper I used <strong>1 down sampling layer to reduce frames from 320 to 160</strong>.<br>\nThis is mainly helpful for adding more encoder layers, and also can allow using complex encoder layers.</li>\n<li><strong>Relative pos encoding</strong> affect the performance so much, I found <strong>ROPE(by Jianlin Su)</strong> perform best and super fast. (code from huggingface implementation)</li>\n<li><strong>Dropout is super important, I used final cls_drop=0.1 only, no dropout for other layers.</strong></li>\n<li>I learned from 1st-place-solution-training that <strong>stochastic path</strong> is <strong>super important to avoid overfit for deep network</strong>. So I used it in squeeeze former block for each layer like below(notice the InstDropout and 0.5 skip factor all super important for final performance)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F42245%2Ffb17306d174a4bd7a4b81ad0acdc0ade%2F2.png?generation=1692937044317556&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F42245%2F4834f8f3d02efab8d8f0bbe32a79f9a4%2F3.png?generation=1692937055628577&amp;alt=media\" alt=\"\"><br>\nOverall model arch below:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F42245%2F68866c3f59e3e6b2f308a54fc8b94248%2F4.png?generation=1692937099435075&amp;alt=media\" alt=\"\"></li>\n</ol>\n<h1><strong>Others</strong></h1>\n<ol>\n<li>For me (using 4090), torch much faster then tf (3-4 times faster)  </li>\n<li>I used tf at first and swich to torch in the last month which help to speedup experiments a lot.    <br>\nSpeed is only one factor, another important thing is I could easily try opensource code of ASR like NEMO or espnet.</li>\n<li>I used <strong>nobuco</strong> to convert torch model to keras, it works like a charm.  </li>\n<li>I still used tf for prorcess/aug and post process and also use tfrecord as input format, I wrote a torch iterable dataset which wrap tfrecord reader.</li>\n</ol>\n<h1><strong>TODOS</strong></h1>\n<p>Due to time limit, I could not finish more experiments at last days, but some possible improvements might be</p>\n<ol>\n<li>Model can be deeper up to <strong>20 layers</strong><br>\n20 layer model perform better then 17 layers but I only trained 300 epochs which perform not as good as 17 layers + 400 epochs.</li>\n<li>More epochs training, maybe 500 or 600 ? <br>\nFor 17 layer model from 300 to 400 improve LB 4 points and PB 2 points, so might sitll could train more epochs and might use 20 layer model for more epochs help even more:)</li>\n<li>From what I learned from other solutions, it seems I missed some major points here</li>\n</ol>\n<ul>\n<li>cutmix<br>\nthis is a pity, I planned to do this at the begging but did not try it, as I found hard mask of frames(50%+) worked very well, I should have realized that cut mix might help even more.  <br>\nsimple concat 2 instances  <br>\nconcat with some ratio like 0.7 and 0.3  <br>\nconcat with using ctc segmentation   </li>\n<li>seq2seq method and ctc+attention decode method    <br>\nI tried seq2seq using tf at the begging of this competition but not give good results, should have tried it after changing to use torch with squeezeformer encoder.  </li>\n<li>input len mask to speedup infer</li>\n<li>try even more max input frames from 320 to 384 or 512(with input len mask infer)</li>\n</ul>\n<h1><strong>Code</strong></h1>\n<p>Opensource all codes here:  <br>\n<a href=\"https://www.kaggle.com/code/goldenlock/3rd-place-step1-gen-tfrecords-for-train\" target=\"_blank\">https://www.kaggle.com/code/goldenlock/3rd-place-step1-gen-tfrecords-for-train</a><br>\n<a href=\"https://www.kaggle.com/code/goldenlock/3rd-place-step2-gen-tfrecords-for-supplement\" target=\"_blank\">https://www.kaggle.com/code/goldenlock/3rd-place-step2-gen-tfrecords-for-supplement</a><br>\n<a href=\"https://www.kaggle.com/code/goldenlock/3rd-place-step-3-gen-mean-and-std\" target=\"_blank\">https://www.kaggle.com/code/goldenlock/3rd-place-step-3-gen-mean-and-std</a><br>\n<a href=\"https://www.kaggle.com/code/goldenlock/3rd-place-step-4-train-squeezeformer\" target=\"_blank\">https://www.kaggle.com/code/goldenlock/3rd-place-step-4-train-squeezeformer</a><br>\n<a href=\"https://www.kaggle.com/code/goldenlock/3rd-place-step-5-torch2keras-using-nobuco\" target=\"_blank\">https://www.kaggle.com/code/goldenlock/3rd-place-step-5-torch2keras-using-nobuco</a><br>\n<a href=\"https://www.kaggle.com/code/goldenlock/3rd-place-step-6-inference\" target=\"_blank\">https://www.kaggle.com/code/goldenlock/3rd-place-step-6-inference</a>  <br>\nNotice I did not reproduce my final results using kaggle notebooks, so if you want to reproduce or want to find the original code you could find it here:<br>\n<a href=\"https://github.com/chenghuige/Google-American_Sign_Language_Fingerspelling_Recognition\" target=\"_blank\">https://github.com/chenghuige/Google-American_Sign_Language_Fingerspelling_Recognition</a>  </p>",
  "messages": [
    {
      "id": 2407394,
      "postDate": "2023-08-25T04:18:36.233Z",
      "content": "<p>This is an intesting game, I learned a lot from </p>\n<p><a href=\"https://www.kaggle.com/code/hoyso48/1st-place-solution-training\" target=\"_blank\">https://www.kaggle.com/code/hoyso48/1st-place-solution-training</a><br>\n<a href=\"https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place\" target=\"_blank\">https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place</a><br>\n<a href=\"https://www.kaggle.com/code/markwijkhuizen/aslfr-transformer-training-inference\" target=\"_blank\">https://www.kaggle.com/code/markwijkhuizen/aslfr-transformer-training-inference</a></p>\n<h1><strong>Preprocss</strong></h1>\n<ol>\n<li><strong>Do not throw away frames witout hands!</strong></li>\n<li>Use as much info as possible, (hands, lips, eye, nose, pose) , use (x,y,z)</li>\n<li>Use more frames, resize to 320 frames for model training.</li>\n<li>I used <strong>original + normalized + abs position(/1000.)</strong> as input feats(total 384 * 2 + 1 = 769). <br>\nnormalized feats help converge much faster but I found original feats also helps.</li>\n</ol>\n<h1><strong>Aug</strong></h1>\n<p>I followed last competion 1st solution, and found the most important aug are </p>\n<ol>\n<li><strong>time scale (interp1d)</strong></li>\n<li><strong>time dim mask</strong><br>\nI used heavy mask here which will mask some time seq and aslo will <strong>randomly mask 50% frames</strong></li>\n<li>affine <br>\nfollow preve 1st solution with prob 0.75, for left-right flip only with prob 0.25</li>\n</ol>\n<h1><strong>Training</strong></h1>\n<ol>\n<li>I used <strong>train data + sup data combined training with sup data weight set to 0.1</strong> for 400 epochs.</li>\n<li>batch size 128, max lr 2e-3, lr scheduler using linear decay with 0.1 epochs warmup , adam optimizer</li>\n<li>Awp training started from epochs * 0.15, with adv_lr 0.2 and adv_eps 0.</li>\n<li><strong>Fintune with train data only for 10 epochs</strong> with max lr 1e-4, awp start from epoch 2.  </li>\n</ol>\n<h1><strong>Postdeal</strong></h1>\n<p>I used <strong>rule for blank index(0)</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F42245%2F54e452eca19b2ee1610a7ecef4ea7cdc%2F1.png?generation=1692937017778775&amp;alt=media\" alt=\"\"></p>\n<h1><strong>Model</strong></h1>\n<ol>\n<li><strong>Squeezeformer</strong>  works best, it is a bit faster then conformer. (Code from NEMO)</li>\n<li>In order to make net deeper I used <strong>1 down sampling layer to reduce frames from 320 to 160</strong>.<br>\nThis is mainly helpful for adding more encoder layers, and also can allow using complex encoder layers.</li>\n<li><strong>Relative pos encoding</strong> affect the performance so much, I found <strong>ROPE(by Jianlin Su)</strong> perform best and super fast. (code from huggingface implementation)</li>\n<li><strong>Dropout is super important, I used final cls_drop=0.1 only, no dropout for other layers.</strong></li>\n<li>I learned from 1st-place-solution-training that <strong>stochastic path</strong> is <strong>super important to avoid overfit for deep network</strong>. So I used it in squeeeze former block for each layer like below(notice the InstDropout and 0.5 skip factor all super important for final performance)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F42245%2Ffb17306d174a4bd7a4b81ad0acdc0ade%2F2.png?generation=1692937044317556&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F42245%2F4834f8f3d02efab8d8f0bbe32a79f9a4%2F3.png?generation=1692937055628577&amp;alt=media\" alt=\"\"><br>\nOverall model arch below:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F42245%2F68866c3f59e3e6b2f308a54fc8b94248%2F4.png?generation=1692937099435075&amp;alt=media\" alt=\"\"></li>\n</ol>\n<h1><strong>Others</strong></h1>\n<ol>\n<li>For me (using 4090), torch much faster then tf (3-4 times faster)  </li>\n<li>I used tf at first and swich to torch in the last month which help to speedup experiments a lot.    <br>\nSpeed is only one factor, another important thing is I could easily try opensource code of ASR like NEMO or espnet.</li>\n<li>I used <strong>nobuco</strong> to convert torch model to keras, it works like a charm.  </li>\n<li>I still used tf for prorcess/aug and post process and also use tfrecord as input format, I wrote a torch iterable dataset which wrap tfrecord reader.</li>\n</ol>\n<h1><strong>TODOS</strong></h1>\n<p>Due to time limit, I could not finish more experiments at last days, but some possible improvements might be</p>\n<ol>\n<li>Model can be deeper up to <strong>20 layers</strong><br>\n20 layer model perform better then 17 layers but I only trained 300 epochs which perform not as good as 17 layers + 400 epochs.</li>\n<li>More epochs training, maybe 500 or 600 ? <br>\nFor 17 layer model from 300 to 400 improve LB 4 points and PB 2 points, so might sitll could train more epochs and might use 20 layer model for more epochs help even more:)</li>\n<li>From what I learned from other solutions, it seems I missed some major points here</li>\n</ol>\n<ul>\n<li>cutmix<br>\nthis is a pity, I planned to do this at the begging but did not try it, as I found hard mask of frames(50%+) worked very well, I should have realized that cut mix might help even more.  <br>\nsimple concat 2 instances  <br>\nconcat with some ratio like 0.7 and 0.3  <br>\nconcat with using ctc segmentation   </li>\n<li>seq2seq method and ctc+attention decode method    <br>\nI tried seq2seq using tf at the begging of this competition but not give good results, should have tried it after changing to use torch with squeezeformer encoder.  </li>\n<li>input len mask to speedup infer</li>\n<li>try even more max input frames from 320 to 384 or 512(with input len mask infer)</li>\n</ul>\n<h1><strong>Code</strong></h1>\n<p>Opensource all codes here:  <br>\n<a href=\"https://www.kaggle.com/code/goldenlock/3rd-place-step1-gen-tfrecords-for-train\" target=\"_blank\">https://www.kaggle.com/code/goldenlock/3rd-place-step1-gen-tfrecords-for-train</a><br>\n<a href=\"https://www.kaggle.com/code/goldenlock/3rd-place-step2-gen-tfrecords-for-supplement\" target=\"_blank\">https://www.kaggle.com/code/goldenlock/3rd-place-step2-gen-tfrecords-for-supplement</a><br>\n<a href=\"https://www.kaggle.com/code/goldenlock/3rd-place-step-3-gen-mean-and-std\" target=\"_blank\">https://www.kaggle.com/code/goldenlock/3rd-place-step-3-gen-mean-and-std</a><br>\n<a href=\"https://www.kaggle.com/code/goldenlock/3rd-place-step-4-train-squeezeformer\" target=\"_blank\">https://www.kaggle.com/code/goldenlock/3rd-place-step-4-train-squeezeformer</a><br>\n<a href=\"https://www.kaggle.com/code/goldenlock/3rd-place-step-5-torch2keras-using-nobuco\" target=\"_blank\">https://www.kaggle.com/code/goldenlock/3rd-place-step-5-torch2keras-using-nobuco</a><br>\n<a href=\"https://www.kaggle.com/code/goldenlock/3rd-place-step-6-inference\" target=\"_blank\">https://www.kaggle.com/code/goldenlock/3rd-place-step-6-inference</a>  <br>\nNotice I did not reproduce my final results using kaggle notebooks, so if you want to reproduce or want to find the original code you could find it here:<br>\n<a href=\"https://github.com/chenghuige/Google-American_Sign_Language_Fingerspelling_Recognition\" target=\"_blank\">https://github.com/chenghuige/Google-American_Sign_Language_Fingerspelling_Recognition</a>  </p>",
      "rawMarkdown": "This is an intesting game, I learned a lot from \n\nhttps://www.kaggle.com/code/hoyso48/1st-place-solution-training\nhttps://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place\nhttps://www.kaggle.com/code/markwijkhuizen/aslfr-transformer-training-inference\n\n# **Preprocss**\n1. **Do not throw away frames witout hands!**\n2. Use as much info as possible, (hands, lips, eye, nose, pose) , use (x,y,z)\n3. Use more frames, resize to 320 frames for model training.\n4. I used **original + normalized + abs position(/1000.)** as input feats(total 384 * 2 + 1 = 769). \nnormalized feats help converge much faster but I found original feats also helps.\n\n# **Aug**\nI followed last competion 1st solution, and found the most important aug are \n1. **time scale (interp1d)**\n2. **time dim mask**\nI used heavy mask here which will mask some time seq and aslo will **randomly mask 50% frames**\n3. affine \nfollow preve 1st solution with prob 0.75, for left-right flip only with prob 0.25\n\n# **Training**\n1. I used **train data + sup data combined training with sup data weight set to 0.1** for 400 epochs.\n2. batch size 128, max lr 2e-3, lr scheduler using linear decay with 0.1 epochs warmup , adam optimizer\n2. Awp training started from epochs * 0.15, with adv_lr 0.2 and adv_eps 0.\n3. **Fintune with train data only for 10 epochs** with max lr 1e-4, awp start from epoch 2.  \n\n# **Postdeal**\nI used **rule for blank index(0)**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F42245%2F54e452eca19b2ee1610a7ecef4ea7cdc%2F1.png?generation=1692937017778775&alt=media)\n\n# **Model**\n1. **Squeezeformer**  works best, it is a bit faster then conformer. (Code from NEMO)\n2. In order to make net deeper I used **1 down sampling layer to reduce frames from 320 to 160**.\nThis is mainly helpful for adding more encoder layers, and also can allow using complex encoder layers.\n3. **Relative pos encoding** affect the performance so much, I found **ROPE(by Jianlin Su)** perform best and super fast. (code from huggingface implementation)\n4. **Dropout is super important, I used final cls_drop=0.1 only, no dropout for other layers.**\n5. I learned from 1st-place-solution-training that **stochastic path** is **super important to avoid overfit for deep network**. So I used it in squeeeze former block for each layer like below(notice the InstDropout and 0.5 skip factor all super important for final performance)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F42245%2Ffb17306d174a4bd7a4b81ad0acdc0ade%2F2.png?generation=1692937044317556&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F42245%2F4834f8f3d02efab8d8f0bbe32a79f9a4%2F3.png?generation=1692937055628577&alt=media)\nOverall model arch below:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F42245%2F68866c3f59e3e6b2f308a54fc8b94248%2F4.png?generation=1692937099435075&alt=media)\n\n# **Others**\n1. For me (using 4090), torch much faster then tf (3-4 times faster)  \n2. I used tf at first and swich to torch in the last month which help to speedup experiments a lot.    \nSpeed is only one factor, another important thing is I could easily try opensource code of ASR like NEMO or espnet.\n3. I used **nobuco** to convert torch model to keras, it works like a charm.  \n4. I still used tf for prorcess/aug and post process and also use tfrecord as input format, I wrote a torch iterable dataset which wrap tfrecord reader.\n# **TODOS**\nDue to time limit, I could not finish more experiments at last days, but some possible improvements might be\n1. Model can be deeper up to **20 layers**\n20 layer model perform better then 17 layers but I only trained 300 epochs which perform not as good as 17 layers + 400 epochs.\n2. More epochs training, maybe 500 or 600 ? \nFor 17 layer model from 300 to 400 improve LB 4 points and PB 2 points, so might sitll could train more epochs and might use 20 layer model for more epochs help even more:)\n3. From what I learned from other solutions, it seems I missed some major points here\n- cutmix\nthis is a pity, I planned to do this at the begging but did not try it, as I found hard mask of frames(50%+) worked very well, I should have realized that cut mix might help even more.  \nsimple concat 2 instances  \nconcat with some ratio like 0.7 and 0.3  \nconcat with using ctc segmentation   \n- seq2seq method and ctc+attention decode method    \nI tried seq2seq using tf at the begging of this competition but not give good results, should have tried it after changing to use torch with squeezeformer encoder.  \n- input len mask to speedup infer\n- try even more max input frames from 320 to 384 or 512(with input len mask infer)\n# **Code**\nOpensource all codes here:  \nhttps://www.kaggle.com/code/goldenlock/3rd-place-step1-gen-tfrecords-for-train\nhttps://www.kaggle.com/code/goldenlock/3rd-place-step2-gen-tfrecords-for-supplement\nhttps://www.kaggle.com/code/goldenlock/3rd-place-step-3-gen-mean-and-std\nhttps://www.kaggle.com/code/goldenlock/3rd-place-step-4-train-squeezeformer\nhttps://www.kaggle.com/code/goldenlock/3rd-place-step-5-torch2keras-using-nobuco\nhttps://www.kaggle.com/code/goldenlock/3rd-place-step-6-inference  \nNotice I did not reproduce my final results using kaggle notebooks, so if you want to reproduce or want to find the original code you could find it here:\nhttps://github.com/chenghuige/Google-American_Sign_Language_Fingerspelling_Recognition  ",
      "votes": 53
    },
    {
      "id": 2412207,
      "postDate": "2023-08-28T06:34:51.847Z",
      "content": "<p>Great work man!</p>",
      "rawMarkdown": "Great work man!",
      "votes": 1
    },
    {
      "id": 2410485,
      "postDate": "2023-08-27T02:19:47.247Z",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/goldenlock\" target=\"_blank\">@goldenlock</a>, great work yet again! learning a lot from you</p>",
      "rawMarkdown": "Thank you @goldenlock, great work yet again! learning a lot from you",
      "votes": 1
    },
    {
      "id": 2407700,
      "postDate": "2023-08-25T07:48:27.090Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/goldenlock\" target=\"_blank\">@goldenlock</a>, great work yet again!</p>\n<p>I have a question on awp; we could not get it to work; did you train with fp16 ? I thought this might be the reason we kept getting nans. <br>\nMaybe we need to wait to see your implementation of it.  </p>",
      "rawMarkdown": "Congratulations @goldenlock, great work yet again!\n\nI have a question on awp; we could not get it to work; did you train with fp16 ? I thought this might be the reason we kept getting nans. \nMaybe we need to wait to see your implementation of it.  ",
      "votes": 1,
      "replies": [
        {
          "id": 2407744,
          "postDate": "2023-08-25T08:28:31.480Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a> , I also learned a lot from your sharing of last competion. <br>\nFor awp, I used fp32, it will also cause nan for me at the begging if using fp16 I do not kown why.<br>\nSo I will use fp16 + no awp first then change to load model and train with fp32 + awp. <br>\nAnyway awp work and give better result especially for less epochs, a boost of 2-3 points, but if you train like 300+ epochs the boost might be less and it is also very time costing as you need to start awp training early.</p>",
          "rawMarkdown": "Thanks @darraghdog , I also learned a lot from your sharing of last competion. \nFor awp, I used fp32, it will also cause nan for me at the begging if using fp16 I do not kown why.\nSo I will use fp16 + no awp first then change to load model and train with fp32 + awp. \nAnyway awp work and give better result especially for less epochs, a boost of 2-3 points, but if you train like 300+ epochs the boost might be less and it is also very time costing as you need to start awp training early.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2407589,
      "postDate": "2023-08-25T06:36:32.247Z",
      "content": "<p>Congratulations.  Thanks for sharing the details of your model. <br>\nThe deeper layer models seem to be consuming more than 9 hours. </p>",
      "rawMarkdown": "Congratulations.  Thanks for sharing the details of your model. \nThe deeper layer models seem to be consuming more than 9 hours. \n",
      "votes": 1,
      "replies": [
        {
          "id": 2407746,
          "postDate": "2023-08-25T08:30:54.480Z",
          "content": "<p>Squeezeformer is fast you could also make the encoder units less for me I used 200. Also I used smaller attention head dim 32. Actually I tested we could submit model with 20 layers which use 0.81s/it and finish just in less then 5 hours.</p>",
          "rawMarkdown": "Squeezeformer is fast you could also make the encoder units less for me I used 200. Also I used smaller attention head dim 32. Actually I tested we could submit model with 20 layers which use 0.81s/it and finish just in less then 5 hours.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2407493,
      "postDate": "2023-08-25T05:38:35.070Z",
      "content": "<p>Congratulations,<br>\nI also tried to use SqueezeFormer, but due to my limited coding knowledge, I couldn't replicate it. Attention with relative positional encoding completely blew my mind.</p>\n<p>Thanks to the competitors and all participants, because of all of you, I was able to go from never having heard of transformers, attention mechanisms, SqueezeFormer, Conformer, CTC loss, and many other topics to gaining extra knowledge. I am glad to have participated in this competition.</p>\n<p>If you can share code it help so much.</p>",
      "rawMarkdown": "Congratulations,\nI also tried to use SqueezeFormer, but due to my limited coding knowledge, I couldn't replicate it. Attention with relative positional encoding completely blew my mind.\n\nThanks to the competitors and all participants, because of all of you, I was able to go from never having heard of transformers, attention mechanisms, SqueezeFormer, Conformer, CTC loss, and many other topics to gaining extra knowledge. I am glad to have participated in this competition.\n\nIf you can share code it help so much.",
      "votes": 1,
      "replies": [
        {
          "id": 2407752,
          "postDate": "2023-08-25T08:35:43.403Z",
          "content": "<p>Thanks,  I will share the code soon.</p>",
          "rawMarkdown": "Thanks,  I will share the code soon."
        }
      ]
    },
    {
      "id": 2407442,
      "postDate": "2023-08-25T05:04:06.060Z",
      "content": "<p>\"I still used tf for prorcess/aug and post process and also use tfrecord as input format, I write an torch iterable dataset which wrap tfrecord reader.\"</p>\n<p>thanks.<br>\nso this comfirm my guess : the bottle net is the tf training pipline. <br>\nnow everything else (even data) is really tf. </p>\n<p>This is also my observation. For me i replace tf dataprocessing with np dataprocessing (since from the tf document, tf data processing is done on cpu) and use tf training loop and confirm no improvement (and no degradation). i started from pytorch, then switch to tf/keras(custom train loop and keras fit) and swtich but to pytorch again</p>",
      "rawMarkdown": "\"I still used tf for prorcess/aug and post process and also use tfrecord as input format, I write an torch iterable dataset which wrap tfrecord reader.\"\n\nthanks.\nso this comfirm my guess : the bottle net is the tf training pipline. \nnow everything else (even data) is really tf. \n\nThis is also my observation. For me i replace tf dataprocessing with np dataprocessing (since from the tf document, tf data processing is done on cpu) and use tf training loop and confirm no improvement (and no degradation). i started from pytorch, then switch to tf/keras(custom train loop and keras fit) and swtich but to pytorch again",
      "votes": 1,
      "replies": [
        {
          "id": 2407475,
          "postDate": "2023-08-25T05:29:06.780Z",
          "content": "<p>Yes, I'm still very confused why tf so slow on this problem using gpu….</p>",
          "rawMarkdown": "Yes, I'm still very confused why tf so slow on this problem using gpu....",
          "votes": 1
        }
      ]
    },
    {
      "id": 2407415,
      "postDate": "2023-08-25T04:33:07.153Z",
      "content": "<p>thanks for the introduction. i will try it:</p>\n<p><a href=\"https://github.com/AlexanderLutsenko/nobuco\" target=\"_blank\">https://github.com/AlexanderLutsenko/nobuco</a></p>",
      "rawMarkdown": "thanks for the introduction. i will try it:\n\nhttps://github.com/AlexanderLutsenko/nobuco",
      "votes": 1
    },
    {
      "id": 2407405,
      "postDate": "2023-08-25T04:27:54.193Z",
      "content": "<p>Congrats on winning 3rd prize solo!<br>\nYour work is brilliant, I need some time to digest it 🤔</p>",
      "rawMarkdown": "Congrats on winning 3rd prize solo!\nYour work is brilliant, I need some time to digest it 🤔",
      "votes": 1
    },
    {
      "id": 2407743,
      "postDate": "2023-08-25T08:27:19.593Z",
      "content": "<p>Congrats, thanks for the detailed write up!</p>",
      "rawMarkdown": "Congrats, thanks for the detailed write up!",
      "votes": 2
    },
    {
      "id": 2407458,
      "postDate": "2023-08-25T05:14:38.953Z",
      "content": "<p>Nice work, thanks for sharing. We find several of your findings in our solution too. (e.g. Squeezeformer as the encoder). Impressive how much you found as a solo competitor. </p>",
      "rawMarkdown": "Nice work, thanks for sharing. We find several of your findings in our solution too. (e.g. Squeezeformer as the encoder). Impressive how much you found as a solo competitor. ",
      "votes": 2,
      "replies": [
        {
          "id": 2407467,
          "postDate": "2023-08-25T05:25:53.657Z",
          "content": "<p>Thanks Dieter! Can not wait to learn from your solution, really big gap, I think there are some major points no other teams have found😀</p>",
          "rawMarkdown": "Thanks Dieter! Can not wait to learn from your solution, really big gap, I think there are some major points no other teams have found😀",
          "votes": 1
        }
      ]
    },
    {
      "id": 2407436,
      "postDate": "2023-08-25T04:59:13.257Z",
      "content": "<p>\"Do not throw away frames witout hands!<br>\nUse as much info as possible\"</p>\n<p>actually i think face points etc may help to identify \"person identity\" (i.e. unintented data leak if there is same signer at the same day/time in the test ).</p>\n<hr>\n<p>in kaggle, always try all data as an option (even in theory some data seen not to be useful). you never knows what happens.<br>\nalso try subset is also another an option, you never know what noise is there.</p>\n<p>most people focus on modeling, but somtimes data modeling also help.</p>",
      "rawMarkdown": "\"Do not throw away frames witout hands!\nUse as much info as possible\"\n\nactually i think face points etc may help to identify \"person identity\" (i.e. unintented data leak if there is same signer at the same day/time in the test ).\n\n---\n\nin kaggle, always try all data as an option (even in theory some data seen not to be useful). you never knows what happens.\nalso try subset is also another an option, you never know what noise is there.\n\nmost people focus on modeling, but somtimes data modeling also help.",
      "votes": 2,
      "replies": [
        {
          "id": 2407473,
          "postDate": "2023-08-25T05:27:20.300Z",
          "content": "<p>Yes this affect much, but Chris's method might be even better.</p>",
          "rawMarkdown": "Yes this affect much, but Chris's method might be even better.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2407424,
      "postDate": "2023-08-25T04:50:26.627Z",
      "content": "<p>Congratulations, I will try your techniques</p>",
      "rawMarkdown": "Congratulations, I will try your techniques",
      "votes": 2
    },
    {
      "id": 2415551,
      "postDate": "2023-08-30T12:40:01.397Z",
      "content": "<p>Superb Work !</p>",
      "rawMarkdown": "Superb Work !"
    },
    {
      "id": 2415285,
      "postDate": "2023-08-30T08:15:55.313Z",
      "content": "<p>Have you tried other models?</p>",
      "rawMarkdown": "Have you tried other models?",
      "replies": [
        {
          "id": 2416304,
          "postDate": "2023-08-30T23:49:48.583Z",
          "content": "<p>I tried public notebooks of seq2seq and ctc in tensorflow at first, then I convert the ctc notebook to torch, at the last month I change to use conformer and switch to suqeezeformer at the last week.</p>",
          "rawMarkdown": "I tried public notebooks of seq2seq and ctc in tensorflow at first, then I convert the ctc notebook to torch, at the last month I change to use conformer and switch to suqeezeformer at the last week."
        }
      ]
    },
    {
      "id": 2407556,
      "postDate": "2023-08-25T06:12:04.157Z",
      "content": "<p>Amazing. There is much to learn from you solution!<br>\nCan you explain what is the purpose of the 'rule for blank index'? I did not get it.</p>",
      "rawMarkdown": "Amazing. There is much to learn from you solution!\nCan you explain what is the purpose of the 'rule for blank index'? I did not get it.",
      "replies": [
        {
          "id": 2407750,
          "postDate": "2023-08-25T08:34:07.557Z",
          "content": "<p>Here mean for each ctc output step you could adjust the prob of blank index, for example if blank index prob is 0.25 and it is still the largest prob for this step. If no post rule then will output blank index, but if using the post rule it will output a char in our vocab.</p>",
          "rawMarkdown": "Here mean for each ctc output step you could adjust the prob of blank index, for example if blank index prob is 0.25 and it is still the largest prob for this step. If no post rule then will output blank index, but if using the post rule it will output a char in our vocab.",
          "votes": 2,
          "replies": [
            {
              "id": 2407807,
              "postDate": "2023-08-25T09:12:58.830Z",
              "content": "<p>How much improvement from this trick?</p>",
              "rawMarkdown": "How much improvement from this trick?\n"
            },
            {
              "id": 2412265,
              "postDate": "2023-08-28T07:16:32.043Z",
              "content": "<p>The result is similar as other solutions which use post trick like special output for short predictions. 2-3 point online I think.</p>",
              "rawMarkdown": "The result is similar as other solutions which use post trick like special output for short predictions. 2-3 point online I think."
            }
          ]
        }
      ]
    },
    {
      "id": 2407399,
      "postDate": "2023-08-25T04:24:12.783Z",
      "content": "<p>perfect, will you make your code available soon?</p>",
      "rawMarkdown": "perfect, will you make your code available soon?",
      "replies": [
        {
          "id": 2407401,
          "postDate": "2023-08-25T04:24:41.783Z",
          "content": "<p>Sure, will be avaliable soon.</p>",
          "rawMarkdown": "Sure, will be avaliable soon.",
          "votes": 3,
          "replies": [
            {
              "id": 2407406,
              "postDate": "2023-08-25T04:28:21.337Z",
              "content": "<p>how long did you spend completing one epoch using 1 4090 GPU?</p>",
              "rawMarkdown": "how long did you spend completing one epoch using 1 4090 GPU?",
              "votes": 1
            },
            {
              "id": 2407425,
              "postDate": "2023-08-25T04:51:05.567Z",
              "content": "<blockquote>\n  <p>how long did you spend completing one epoch using 1 4090 GPU?<br>\n  For 17 layer model(model_size: 63.94M num_params: 16M) and batch size 128, with torch.compile and train data only(no sup data) 1 epoch(num_train_examples: 52163  num_steps_per_epoch: 407) cost about 1min,<br>\n  so for train data only 300 epochs will use about 5 hours.</p>\n</blockquote>",
              "rawMarkdown": "> how long did you spend completing one epoch using 1 4090 GPU?\nFor 17 layer model(model_size: 63.94M num_params: 16M) and batch size 128, with torch.compile and train data only(no sup data) 1 epoch(num_train_examples: 52163  num_steps_per_epoch: 407) cost about 1min,\nso for train data only 300 epochs will use about 5 hours.\n\n",
              "votes": 2
            },
            {
              "id": 2407607,
              "postDate": "2023-08-25T06:45:00.953Z",
              "content": "<p>So tensorflow epoch was 3-4 min. Did you used XLA with the tensorflow model?</p>",
              "rawMarkdown": "So tensorflow epoch was 3-4 min. Did you used XLA with the tensorflow model?"
            },
            {
              "id": 2407741,
              "postDate": "2023-08-25T08:23:02.493Z",
              "content": "<p>enable xla will be even slower</p>",
              "rawMarkdown": "enable xla will be even slower"
            }
          ]
        }
      ]
    },
    {
      "id": 2426074,
      "postDate": "2023-09-06T11:41:29.550Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2412207,
      "author_name": "Muhammad Usman",
      "author_url": "",
      "post_date": "2023-08-28T06:34:51.847000",
      "content": "<p>Great work man!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2410485,
      "author_name": "Krishna Singh",
      "author_url": "",
      "post_date": "2023-08-27T02:19:47.247000",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/goldenlock\" target=\"_blank\">@goldenlock</a>, great work yet again! learning a lot from you</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2407700,
      "author_name": "Darragh",
      "author_url": "",
      "post_date": "2023-08-25T07:48:27.090000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/goldenlock\" target=\"_blank\">@goldenlock</a>, great work yet again!</p>\n<p>I have a question on awp; we could not get it to work; did you train with fp16 ? I thought this might be the reason we kept getting nans. <br>\nMaybe we need to wait to see your implementation of it.  </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2407744,
          "author_name": "gezi",
          "author_url": "",
          "post_date": "2023-08-25T08:28:31.480000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a> , I also learned a lot from your sharing of last competion. <br>\nFor awp, I used fp32, it will also cause nan for me at the begging if using fp16 I do not kown why.<br>\nSo I will use fp16 + no awp first then change to load model and train with fp32 + awp. <br>\nAnyway awp work and give better result especially for less epochs, a boost of 2-3 points, but if you train like 300+ epochs the boost might be less and it is also very time costing as you need to start awp training early.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2407589,
      "author_name": "C R Suthikshn Kumar",
      "author_url": "",
      "post_date": "2023-08-25T06:36:32.247000",
      "content": "<p>Congratulations.  Thanks for sharing the details of your model. <br>\nThe deeper layer models seem to be consuming more than 9 hours. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2407746,
          "author_name": "gezi",
          "author_url": "",
          "post_date": "2023-08-25T08:30:54.480000",
          "content": "<p>Squeezeformer is fast you could also make the encoder units less for me I used 200. Also I used smaller attention head dim 32. Actually I tested we could submit model with 20 layers which use 0.81s/it and finish just in less then 5 hours.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2407493,
      "author_name": "Pk",
      "author_url": "",
      "post_date": "2023-08-25T05:38:35.070000",
      "content": "<p>Congratulations,<br>\nI also tried to use SqueezeFormer, but due to my limited coding knowledge, I couldn't replicate it. Attention with relative positional encoding completely blew my mind.</p>\n<p>Thanks to the competitors and all participants, because of all of you, I was able to go from never having heard of transformers, attention mechanisms, SqueezeFormer, Conformer, CTC loss, and many other topics to gaining extra knowledge. I am glad to have participated in this competition.</p>\n<p>If you can share code it help so much.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2407752,
          "author_name": "gezi",
          "author_url": "",
          "post_date": "2023-08-25T08:35:43.403000",
          "content": "<p>Thanks,  I will share the code soon.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2407442,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-08-25T05:04:06.060000",
      "content": "<p>\"I still used tf for prorcess/aug and post process and also use tfrecord as input format, I write an torch iterable dataset which wrap tfrecord reader.\"</p>\n<p>thanks.<br>\nso this comfirm my guess : the bottle net is the tf training pipline. <br>\nnow everything else (even data) is really tf. </p>\n<p>This is also my observation. For me i replace tf dataprocessing with np dataprocessing (since from the tf document, tf data processing is done on cpu) and use tf training loop and confirm no improvement (and no degradation). i started from pytorch, then switch to tf/keras(custom train loop and keras fit) and swtich but to pytorch again</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2407475,
          "author_name": "gezi",
          "author_url": "",
          "post_date": "2023-08-25T05:29:06.780000",
          "content": "<p>Yes, I'm still very confused why tf so slow on this problem using gpu….</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2407415,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-08-25T04:33:07.153000",
      "content": "<p>thanks for the introduction. i will try it:</p>\n<p><a href=\"https://github.com/AlexanderLutsenko/nobuco\" target=\"_blank\">https://github.com/AlexanderLutsenko/nobuco</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2407405,
      "author_name": "Yu Wu",
      "author_url": "",
      "post_date": "2023-08-25T04:27:54.193000",
      "content": "<p>Congrats on winning 3rd prize solo!<br>\nYour work is brilliant, I need some time to digest it 🤔</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2407743,
      "author_name": "Samuel Cortinhas",
      "author_url": "",
      "post_date": "2023-08-25T08:27:19.593000",
      "content": "<p>Congrats, thanks for the detailed write up!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2407458,
      "author_name": "Dieter",
      "author_url": "",
      "post_date": "2023-08-25T05:14:38.953000",
      "content": "<p>Nice work, thanks for sharing. We find several of your findings in our solution too. (e.g. Squeezeformer as the encoder). Impressive how much you found as a solo competitor. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 2407467,
          "author_name": "gezi",
          "author_url": "",
          "post_date": "2023-08-25T05:25:53.657000",
          "content": "<p>Thanks Dieter! Can not wait to learn from your solution, really big gap, I think there are some major points no other teams have found😀</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2407436,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-08-25T04:59:13.257000",
      "content": "<p>\"Do not throw away frames witout hands!<br>\nUse as much info as possible\"</p>\n<p>actually i think face points etc may help to identify \"person identity\" (i.e. unintented data leak if there is same signer at the same day/time in the test ).</p>\n<hr>\n<p>in kaggle, always try all data as an option (even in theory some data seen not to be useful). you never knows what happens.<br>\nalso try subset is also another an option, you never know what noise is there.</p>\n<p>most people focus on modeling, but somtimes data modeling also help.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2407473,
          "author_name": "gezi",
          "author_url": "",
          "post_date": "2023-08-25T05:27:20.300000",
          "content": "<p>Yes this affect much, but Chris's method might be even better.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2407424,
      "author_name": "Chakkrit Termritthikun",
      "author_url": "",
      "post_date": "2023-08-25T04:50:26.627000",
      "content": "<p>Congratulations, I will try your techniques</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2415551,
      "author_name": "Mystic Shadow",
      "author_url": "",
      "post_date": "2023-08-30T12:40:01.397000",
      "content": "<p>Superb Work !</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2415285,
      "author_name": "Nowgger",
      "author_url": "",
      "post_date": "2023-08-30T08:15:55.313000",
      "content": "<p>Have you tried other models?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2416304,
          "author_name": "gezi",
          "author_url": "",
          "post_date": "2023-08-30T23:49:48.583000",
          "content": "<p>I tried public notebooks of seq2seq and ctc in tensorflow at first, then I convert the ctc notebook to torch, at the last month I change to use conformer and switch to suqeezeformer at the last week.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2407556,
      "author_name": "greySnow",
      "author_url": "",
      "post_date": "2023-08-25T06:12:04.157000",
      "content": "<p>Amazing. There is much to learn from you solution!<br>\nCan you explain what is the purpose of the 'rule for blank index'? I did not get it.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2407750,
          "author_name": "gezi",
          "author_url": "",
          "post_date": "2023-08-25T08:34:07.557000",
          "content": "<p>Here mean for each ctc output step you could adjust the prob of blank index, for example if blank index prob is 0.25 and it is still the largest prob for this step. If no post rule then will output blank index, but if using the post rule it will output a char in our vocab.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2407807,
              "author_name": "bliao",
              "author_url": "",
              "post_date": "2023-08-25T09:12:58.830000",
              "content": "<p>How much improvement from this trick?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2412265,
              "author_name": "gezi",
              "author_url": "",
              "post_date": "2023-08-28T07:16:32.043000",
              "content": "<p>The result is similar as other solutions which use post trick like special output for short predictions. 2-3 point online I think.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2407399,
      "author_name": "Wenxuan Ye",
      "author_url": "",
      "post_date": "2023-08-25T04:24:12.783000",
      "content": "<p>perfect, will you make your code available soon?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2407401,
          "author_name": "gezi",
          "author_url": "",
          "post_date": "2023-08-25T04:24:41.783000",
          "content": "<p>Sure, will be avaliable soon.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2407406,
              "author_name": "Wenxuan Ye",
              "author_url": "",
              "post_date": "2023-08-25T04:28:21.337000",
              "content": "<p>how long did you spend completing one epoch using 1 4090 GPU?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2407425,
              "author_name": "gezi",
              "author_url": "",
              "post_date": "2023-08-25T04:51:05.567000",
              "content": "<blockquote>\n  <p>how long did you spend completing one epoch using 1 4090 GPU?<br>\n  For 17 layer model(model_size: 63.94M num_params: 16M) and batch size 128, with torch.compile and train data only(no sup data) 1 epoch(num_train_examples: 52163  num_steps_per_epoch: 407) cost about 1min,<br>\n  so for train data only 300 epochs will use about 5 hours.</p>\n</blockquote>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2407607,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2023-08-25T06:45:00.953000",
              "content": "<p>So tensorflow epoch was 3-4 min. Did you used XLA with the tensorflow model?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2407741,
              "author_name": "gezi",
              "author_url": "",
              "post_date": "2023-08-25T08:23:02.493000",
              "content": "<p>enable xla will be even slower</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2426074,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-09-06T11:41:29.550000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2407394": "This is an intesting game, I learned a lot from \n\nhttps://www.kaggle.com/code/hoyso48/1st-place-solution-training\nhttps://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place\nhttps://www.kaggle.com/code/markwijkhuizen/aslfr-transformer-training-inference\n\n# **Preprocss**\n1. **Do not throw away frames witout hands!**\n2. Use as much info as possible, (hands, lips, eye, nose, pose) , use (x,y,z)\n3. Use more frames, resize to 320 frames for model training.\n4. I used **original + normalized + abs position(/1000.)** as input feats(total 384 * 2 + 1 = 769). \nnormalized feats help converge much faster but I found original feats also helps.\n\n# **Aug**\nI followed last competion 1st solution, and found the most important aug are \n1. **time scale (interp1d)**\n2. **time dim mask**\nI used heavy mask here which will mask some time seq and aslo will **randomly mask 50% frames**\n3. affine \nfollow preve 1st solution with prob 0.75, for left-right flip only with prob 0.25\n\n# **Training**\n1. I used **train data + sup data combined training with sup data weight set to 0.1** for 400 epochs.\n2. batch size 128, max lr 2e-3, lr scheduler using linear decay with 0.1 epochs warmup , adam optimizer\n2. Awp training started from epochs * 0.15, with adv_lr 0.2 and adv_eps 0.\n3. **Fintune with train data only for 10 epochs** with max lr 1e-4, awp start from epoch 2.  \n\n# **Postdeal**\nI used **rule for blank index(0)**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F42245%2F54e452eca19b2ee1610a7ecef4ea7cdc%2F1.png?generation=1692937017778775&alt=media)\n\n# **Model**\n1. **Squeezeformer**  works best, it is a bit faster then conformer. (Code from NEMO)\n2. In order to make net deeper I used **1 down sampling layer to reduce frames from 320 to 160**.\nThis is mainly helpful for adding more encoder layers, and also can allow using complex encoder layers.\n3. **Relative pos encoding** affect the performance so much, I found **ROPE(by Jianlin Su)** perform best and super fast. (code from huggingface implementation)\n4. **Dropout is super important, I used final cls_drop=0.1 only, no dropout for other layers.**\n5. I learned from 1st-place-solution-training that **stochastic path** is **super important to avoid overfit for deep network**. So I used it in squeeeze former block for each layer like below(notice the InstDropout and 0.5 skip factor all super important for final performance)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F42245%2Ffb17306d174a4bd7a4b81ad0acdc0ade%2F2.png?generation=1692937044317556&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F42245%2F4834f8f3d02efab8d8f0bbe32a79f9a4%2F3.png?generation=1692937055628577&alt=media)\nOverall model arch below:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F42245%2F68866c3f59e3e6b2f308a54fc8b94248%2F4.png?generation=1692937099435075&alt=media)\n\n# **Others**\n1. For me (using 4090), torch much faster then tf (3-4 times faster)  \n2. I used tf at first and swich to torch in the last month which help to speedup experiments a lot.    \nSpeed is only one factor, another important thing is I could easily try opensource code of ASR like NEMO or espnet.\n3. I used **nobuco** to convert torch model to keras, it works like a charm.  \n4. I still used tf for prorcess/aug and post process and also use tfrecord as input format, I wrote a torch iterable dataset which wrap tfrecord reader.\n# **TODOS**\nDue to time limit, I could not finish more experiments at last days, but some possible improvements might be\n1. Model can be deeper up to **20 layers**\n20 layer model perform better then 17 layers but I only trained 300 epochs which perform not as good as 17 layers + 400 epochs.\n2. More epochs training, maybe 500 or 600 ? \nFor 17 layer model from 300 to 400 improve LB 4 points and PB 2 points, so might sitll could train more epochs and might use 20 layer model for more epochs help even more:)\n3. From what I learned from other solutions, it seems I missed some major points here\n- cutmix\nthis is a pity, I planned to do this at the begging but did not try it, as I found hard mask of frames(50%+) worked very well, I should have realized that cut mix might help even more.  \nsimple concat 2 instances  \nconcat with some ratio like 0.7 and 0.3  \nconcat with using ctc segmentation   \n- seq2seq method and ctc+attention decode method    \nI tried seq2seq using tf at the begging of this competition but not give good results, should have tried it after changing to use torch with squeezeformer encoder.  \n- input len mask to speedup infer\n- try even more max input frames from 320 to 384 or 512(with input len mask infer)\n# **Code**\nOpensource all codes here:  \nhttps://www.kaggle.com/code/goldenlock/3rd-place-step1-gen-tfrecords-for-train\nhttps://www.kaggle.com/code/goldenlock/3rd-place-step2-gen-tfrecords-for-supplement\nhttps://www.kaggle.com/code/goldenlock/3rd-place-step-3-gen-mean-and-std\nhttps://www.kaggle.com/code/goldenlock/3rd-place-step-4-train-squeezeformer\nhttps://www.kaggle.com/code/goldenlock/3rd-place-step-5-torch2keras-using-nobuco\nhttps://www.kaggle.com/code/goldenlock/3rd-place-step-6-inference  \nNotice I did not reproduce my final results using kaggle notebooks, so if you want to reproduce or want to find the original code you could find it here:\nhttps://github.com/chenghuige/Google-American_Sign_Language_Fingerspelling_Recognition  ",
    "2412207": "Great work man!",
    "2410485": "Thank you @goldenlock, great work yet again! learning a lot from you",
    "2407700": "Congratulations @goldenlock, great work yet again!\n\nI have a question on awp; we could not get it to work; did you train with fp16 ? I thought this might be the reason we kept getting nans. \nMaybe we need to wait to see your implementation of it.  ",
    "2407589": "Congratulations.  Thanks for sharing the details of your model. \nThe deeper layer models seem to be consuming more than 9 hours. \n",
    "2407493": "Congratulations,\nI also tried to use SqueezeFormer, but due to my limited coding knowledge, I couldn't replicate it. Attention with relative positional encoding completely blew my mind.\n\nThanks to the competitors and all participants, because of all of you, I was able to go from never having heard of transformers, attention mechanisms, SqueezeFormer, Conformer, CTC loss, and many other topics to gaining extra knowledge. I am glad to have participated in this competition.\n\nIf you can share code it help so much.",
    "2407442": "\"I still used tf for prorcess/aug and post process and also use tfrecord as input format, I write an torch iterable dataset which wrap tfrecord reader.\"\n\nthanks.\nso this comfirm my guess : the bottle net is the tf training pipline. \nnow everything else (even data) is really tf. \n\nThis is also my observation. For me i replace tf dataprocessing with np dataprocessing (since from the tf document, tf data processing is done on cpu) and use tf training loop and confirm no improvement (and no degradation). i started from pytorch, then switch to tf/keras(custom train loop and keras fit) and swtich but to pytorch again",
    "2407415": "thanks for the introduction. i will try it:\n\nhttps://github.com/AlexanderLutsenko/nobuco",
    "2407405": "Congrats on winning 3rd prize solo!\nYour work is brilliant, I need some time to digest it 🤔",
    "2407743": "Congrats, thanks for the detailed write up!",
    "2407458": "Nice work, thanks for sharing. We find several of your findings in our solution too. (e.g. Squeezeformer as the encoder). Impressive how much you found as a solo competitor. ",
    "2407436": "\"Do not throw away frames witout hands!\nUse as much info as possible\"\n\nactually i think face points etc may help to identify \"person identity\" (i.e. unintented data leak if there is same signer at the same day/time in the test ).\n\n---\n\nin kaggle, always try all data as an option (even in theory some data seen not to be useful). you never knows what happens.\nalso try subset is also another an option, you never know what noise is there.\n\nmost people focus on modeling, but somtimes data modeling also help.",
    "2407424": "Congratulations, I will try your techniques",
    "2415551": "Superb Work !",
    "2415285": "Have you tried other models?",
    "2407556": "Amazing. There is much to learn from you solution!\nCan you explain what is the purpose of the 'rule for blank index'? I did not get it.",
    "2407399": "perfect, will you make your code available soon?",
    "2426074": ""
  }
}