{
  "id": 434795,
  "title": "19th Place Solution - Jasper models",
  "url": "/competitions/asl-fingerspelling/discussion/434795",
  "author_name": "Kolya Forrat",
  "post_date": "2023-08-26T14:58:41.390000",
  "votes": 14,
  "comment_count": 4,
  "views": 0,
  "content": "<p>My solution is quite simple. I wasn't going to write about it, but there are no Jasper-like solutions in discussions yet :)</p>\n<h1>Preprocessing</h1>\n<p>Preprocessing was almost the same with <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406306\" target=\"_blank\">previous competition</a>. But deleted mixup and replace augmentation. And add more complex time augmentations.</p>\n<h1>Loss</h1>\n<p>I only used CTC loss.</p>\n<h1>Data split</h1>\n<p>Random split <strong>not</strong> by user for 20 folds grouping by data length.<br>\nDeleted obviously bad samples:</p>\n<pre><code>supplemental_metadata_df = supplemental_metadata_df[~((supplemental_metadata_df[] &lt; ) &amp;\n                                            (supplemental_metadata_df[] &gt; ))]\n\ntrain_df = train_df[~((train_df[] &lt; ) &amp; (train_df[] &gt; ))]\n</code></pre>\n<p><br>\nUsed all supplemental data in train.</p>\n<p>Also tried to create second loader with pseudo labels on bad predicted samples. Or delete bad predicted samples. Used both loaders while training:</p>\n<pre><code>loader_to_use = \n epoch %  == :\n    loader_to_use = \n</code></pre>\n<p>But it didn't help much.</p>\n<h1>Model</h1>\n<p>First I tried simple quartznet, but with all strides set to 1.<br>\nThen I added some features step by step and trained each experiment from previous best checkpoint.</p>\n<h3>Same for all runs:</h3>\n<ul>\n<li>Optimizer: LookAheadAdamW</li>\n<li>Scheduler: Onecycle</li>\n<li>Batch size: 80</li>\n</ul>\n<h3>List of experiments:</h3>\n<ul>\n<li>Simple 5x5 quartznet with stride = 1.  [lr = 8e-3] [epochs=200] CV:0.8035</li>\n<li>Simple 5x5 quartznet with stride = 1. [lr = 7e-3] [epochs=150]  CV: 0.8106</li>\n<li>Previous + SE blocks.  [lr = 6e-3] [epochs=171]  CV: 0.8160</li>\n<li>Previous + Deep supervision (DSV) outputs for train loss. [lr = 5.8e-3] [epochs=175]  CV: 0.8211</li>\n<li>Previous without DSV + higher aug probabilities. Note: score worse but it needed to train one checkpoint without DSV for better next train results. [lr = 6e-3] [epochs=180]  CV: 0.8177</li>\n<li>Previous + DSV + increased dropout. [lr = 6.1e-3] [epochs=200]  CV: 0.8237</li>\n<li>Previous - DSV + lower dropout. [lr = 6e-3] [epochs=173]  CV: 0.8188</li>\n<li>Previous + Hyper Column (HYP) + 2 more blocks with 5 repeats. [lr=5.8e-3] [epochs=180] CV: 0.8252</li>\n<li>Previous + masked 1conv + masked SE + 2D DepthWise CNN layer before classifier + more dropout. [lr=5.8e-3] [epochs=200]  CV: 0.8294</li>\n<li>Previous + lstm after 2D Conv. [lr=1.75e-3] [epochs=125] CV: 0.8301</li>\n</ul>\n<h1>What didn't work for me:</h1>\n<ul>\n<li>Conformer models. I tried it but I guess without enough effort.</li>\n<li>Adding attention or conformer blocks to Jasper models.</li>\n</ul>\n<p>When I used raw conformer and it has very low score. But when I changed input 2D conv to 1D conv or when I just deleted it - it was much better, but still little worse than my Jasper models.</p>\n<h1>What I should have done, but didn't:</h1>\n<ul>\n<li>Train more epochs</li>\n<li>Try better with conformer or squeezeformer</li>\n<li>Add post processing</li>\n<li>Try seq2seq</li>\n<li>Fixed length of input</li>\n</ul>\n<p>Thanks to Kaggle for this great competition. Thanks to all top participants who shared their awesome solutions. <br>\nCongratulations to all winners! </p>",
  "messages": [
    {
      "id": 2409962,
      "postDate": "2023-08-26T14:58:41.390Z",
      "content": "<p>My solution is quite simple. I wasn't going to write about it, but there are no Jasper-like solutions in discussions yet :)</p>\n<h1>Preprocessing</h1>\n<p>Preprocessing was almost the same with <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406306\" target=\"_blank\">previous competition</a>. But deleted mixup and replace augmentation. And add more complex time augmentations.</p>\n<h1>Loss</h1>\n<p>I only used CTC loss.</p>\n<h1>Data split</h1>\n<p>Random split <strong>not</strong> by user for 20 folds grouping by data length.<br>\nDeleted obviously bad samples:</p>\n<pre><code>supplemental_metadata_df = supplemental_metadata_df[~((supplemental_metadata_df[] &lt; ) &amp;\n                                            (supplemental_metadata_df[] &gt; ))]\n\ntrain_df = train_df[~((train_df[] &lt; ) &amp; (train_df[] &gt; ))]\n</code></pre>\n<p><br>\nUsed all supplemental data in train.</p>\n<p>Also tried to create second loader with pseudo labels on bad predicted samples. Or delete bad predicted samples. Used both loaders while training:</p>\n<pre><code>loader_to_use = \n epoch %  == :\n    loader_to_use = \n</code></pre>\n<p>But it didn't help much.</p>\n<h1>Model</h1>\n<p>First I tried simple quartznet, but with all strides set to 1.<br>\nThen I added some features step by step and trained each experiment from previous best checkpoint.</p>\n<h3>Same for all runs:</h3>\n<ul>\n<li>Optimizer: LookAheadAdamW</li>\n<li>Scheduler: Onecycle</li>\n<li>Batch size: 80</li>\n</ul>\n<h3>List of experiments:</h3>\n<ul>\n<li>Simple 5x5 quartznet with stride = 1.  [lr = 8e-3] [epochs=200] CV:0.8035</li>\n<li>Simple 5x5 quartznet with stride = 1. [lr = 7e-3] [epochs=150]  CV: 0.8106</li>\n<li>Previous + SE blocks.  [lr = 6e-3] [epochs=171]  CV: 0.8160</li>\n<li>Previous + Deep supervision (DSV) outputs for train loss. [lr = 5.8e-3] [epochs=175]  CV: 0.8211</li>\n<li>Previous without DSV + higher aug probabilities. Note: score worse but it needed to train one checkpoint without DSV for better next train results. [lr = 6e-3] [epochs=180]  CV: 0.8177</li>\n<li>Previous + DSV + increased dropout. [lr = 6.1e-3] [epochs=200]  CV: 0.8237</li>\n<li>Previous - DSV + lower dropout. [lr = 6e-3] [epochs=173]  CV: 0.8188</li>\n<li>Previous + Hyper Column (HYP) + 2 more blocks with 5 repeats. [lr=5.8e-3] [epochs=180] CV: 0.8252</li>\n<li>Previous + masked 1conv + masked SE + 2D DepthWise CNN layer before classifier + more dropout. [lr=5.8e-3] [epochs=200]  CV: 0.8294</li>\n<li>Previous + lstm after 2D Conv. [lr=1.75e-3] [epochs=125] CV: 0.8301</li>\n</ul>\n<h1>What didn't work for me:</h1>\n<ul>\n<li>Conformer models. I tried it but I guess without enough effort.</li>\n<li>Adding attention or conformer blocks to Jasper models.</li>\n</ul>\n<p>When I used raw conformer and it has very low score. But when I changed input 2D conv to 1D conv or when I just deleted it - it was much better, but still little worse than my Jasper models.</p>\n<h1>What I should have done, but didn't:</h1>\n<ul>\n<li>Train more epochs</li>\n<li>Try better with conformer or squeezeformer</li>\n<li>Add post processing</li>\n<li>Try seq2seq</li>\n<li>Fixed length of input</li>\n</ul>\n<p>Thanks to Kaggle for this great competition. Thanks to all top participants who shared their awesome solutions. <br>\nCongratulations to all winners! </p>",
      "rawMarkdown": "My solution is quite simple. I wasn't going to write about it, but there are no Jasper-like solutions in discussions yet :)\n\n# Preprocessing \nPreprocessing was almost the same with [previous competition](https://www.kaggle.com/competitions/asl-signs/discussion/406306). But deleted mixup and replace augmentation. And add more complex time augmentations.\n\n# Loss\nI only used CTC loss.\n\n# Data split\nRandom split **not** by user for 20 folds grouping by data length.\nDeleted obviously bad samples:\n```python\nsupplemental_metadata_df = supplemental_metadata_df[~((supplemental_metadata_df['length'] < 10) &\n                                            (supplemental_metadata_df['phrase_len'] > 8))]\n\ntrain_df = train_df[~((train_df['length'] < 7) & (train_df['phrase_len'] > 5))]\n``` \nUsed all supplemental data in train.\n\nAlso tried to create second loader with pseudo labels on bad predicted samples. Or delete bad predicted samples. Used both loaders while training:\n\n```python\nloader_to_use = 'main_loader'\nif epoch % 3 == 0:\n    loader_to_use = 'second_loader'\n```\nBut it didn't help much.\n\n# Model\nFirst I tried simple quartznet, but with all strides set to 1.\nThen I added some features step by step and trained each experiment from previous best checkpoint.\n\n### Same for all runs:\n* Optimizer: LookAheadAdamW\n* Scheduler: Onecycle\n* Batch size: 80\n\n### List of experiments:\n* Simple 5x5 quartznet with stride = 1.  [lr = 8e-3] [epochs=200] CV:0.8035\n* Simple 5x5 quartznet with stride = 1. [lr = 7e-3] [epochs=150]  CV: 0.8106\n* Previous + SE blocks.  [lr = 6e-3] [epochs=171]  CV: 0.8160\n* Previous + Deep supervision (DSV) outputs for train loss. [lr = 5.8e-3] [epochs=175]  CV: 0.8211\n* Previous without DSV + higher aug probabilities. Note: score worse but it needed to train one checkpoint without DSV for better next train results. [lr = 6e-3] [epochs=180]  CV: 0.8177\n* Previous + DSV + increased dropout. [lr = 6.1e-3] [epochs=200]  CV: 0.8237\n* Previous - DSV + lower dropout. [lr = 6e-3] [epochs=173]  CV: 0.8188\n* Previous + Hyper Column (HYP) + 2 more blocks with 5 repeats. [lr=5.8e-3] [epochs=180] CV: 0.8252\n* Previous + masked 1conv + masked SE + 2D DepthWise CNN layer before classifier + more dropout. [lr=5.8e-3] [epochs=200]  CV: 0.8294\n* Previous + lstm after 2D Conv. [lr=1.75e-3] [epochs=125] CV: 0.8301\n\n\n# What didn't work for me:\n* Conformer models. I tried it but I guess without enough effort.\n* Adding attention or conformer blocks to Jasper models.\n\nWhen I used raw conformer and it has very low score. But when I changed input 2D conv to 1D conv or when I just deleted it - it was much better, but still little worse than my Jasper models.\n\n# What I should have done, but didn't:\n* Train more epochs\n* Try better with conformer or squeezeformer\n* Add post processing\n* Try seq2seq\n* Fixed length of input\n\nThanks to Kaggle for this great competition. Thanks to all top participants who shared their awesome solutions. \nCongratulations to all winners! \n",
      "votes": 14
    },
    {
      "id": 2417184,
      "postDate": "2023-08-31T13:22:27.897Z",
      "content": "<p>Great work <a href=\"https://www.kaggle.com/kolyaforrat\" target=\"_blank\">@kolyaforrat</a> did you develop in pytorch and port to tf ? I understand nemo jasper models are only in pytorch. </p>",
      "rawMarkdown": "Great work @kolyaforrat did you develop in pytorch and port to tf ? I understand nemo jasper models are only in pytorch. ",
      "replies": [
        {
          "id": 2417778,
          "postDate": "2023-08-31T20:36:42.260Z",
          "content": "<p>Yes, I developed pytorch, and converted it with nobuco. But didn't use Nemo, wrote my own copying some blocks and logic from Nemo.<br>\nOne of the reasons to write my own - is ability to convert this models with nobuco without a lot of changes</p>",
          "rawMarkdown": "Yes, I developed pytorch, and converted it with nobuco. But didn't use Nemo, wrote my own copying some blocks and logic from Nemo.\nOne of the reasons to write my own - is ability to convert this models with nobuco without a lot of changes",
          "votes": 1
        }
      ]
    },
    {
      "id": 2412129,
      "postDate": "2023-08-28T05:21:26.423Z",
      "content": "<p>Congratulations. Thanks for sharing very useful information regarding your solution and your experience from the competition.<br>\n\"train more epochs\"- This  is limited by the notebook time of &lt;9hrs. I once trained my model over many epochs and I could not submit the same notebook  due to time restrictions. I think clever part is to separate training and inference notebooks.<br>\nHow to enforce fixed length of input?</p>",
      "rawMarkdown": "Congratulations. Thanks for sharing very useful information regarding your solution and your experience from the competition.\n\"train more epochs\"- This  is limited by the notebook time of <9hrs. I once trained my model over many epochs and I could not submit the same notebook  due to time restrictions. I think clever part is to separate training and inference notebooks.\nHow to enforce fixed length of input?",
      "replies": [
        {
          "id": 2414062,
          "postDate": "2023-08-29T11:11:17.010Z",
          "content": "<blockquote>\n  <p>How to enforce fixed length of input?</p>\n</blockquote>\n<p><a href=\"https://pytorch.org/docs/stable/generated/torch.nn.functional.interpolate.html\" target=\"_blank\">TORCH.NN.FUNCTIONAL.INTERPOLATE</a></p>",
          "rawMarkdown": ">How to enforce fixed length of input?\n\n[TORCH.NN.FUNCTIONAL.INTERPOLATE](https://pytorch.org/docs/stable/generated/torch.nn.functional.interpolate.html)"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2417184,
      "author_name": "Darragh",
      "author_url": "",
      "post_date": "2023-08-31T13:22:27.897000",
      "content": "<p>Great work <a href=\"https://www.kaggle.com/kolyaforrat\" target=\"_blank\">@kolyaforrat</a> did you develop in pytorch and port to tf ? I understand nemo jasper models are only in pytorch. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2417778,
          "author_name": "Kolya Forrat",
          "author_url": "",
          "post_date": "2023-08-31T20:36:42.260000",
          "content": "<p>Yes, I developed pytorch, and converted it with nobuco. But didn't use Nemo, wrote my own copying some blocks and logic from Nemo.<br>\nOne of the reasons to write my own - is ability to convert this models with nobuco without a lot of changes</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2412129,
      "author_name": "C R Suthikshn Kumar",
      "author_url": "",
      "post_date": "2023-08-28T05:21:26.423000",
      "content": "<p>Congratulations. Thanks for sharing very useful information regarding your solution and your experience from the competition.<br>\n\"train more epochs\"- This  is limited by the notebook time of &lt;9hrs. I once trained my model over many epochs and I could not submit the same notebook  due to time restrictions. I think clever part is to separate training and inference notebooks.<br>\nHow to enforce fixed length of input?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2414062,
          "author_name": "Kolya Forrat",
          "author_url": "",
          "post_date": "2023-08-29T11:11:17.010000",
          "content": "<blockquote>\n  <p>How to enforce fixed length of input?</p>\n</blockquote>\n<p><a href=\"https://pytorch.org/docs/stable/generated/torch.nn.functional.interpolate.html\" target=\"_blank\">TORCH.NN.FUNCTIONAL.INTERPOLATE</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2409962": "My solution is quite simple. I wasn't going to write about it, but there are no Jasper-like solutions in discussions yet :)\n\n# Preprocessing \nPreprocessing was almost the same with [previous competition](https://www.kaggle.com/competitions/asl-signs/discussion/406306). But deleted mixup and replace augmentation. And add more complex time augmentations.\n\n# Loss\nI only used CTC loss.\n\n# Data split\nRandom split **not** by user for 20 folds grouping by data length.\nDeleted obviously bad samples:\n```python\nsupplemental_metadata_df = supplemental_metadata_df[~((supplemental_metadata_df['length'] < 10) &\n                                            (supplemental_metadata_df['phrase_len'] > 8))]\n\ntrain_df = train_df[~((train_df['length'] < 7) & (train_df['phrase_len'] > 5))]\n``` \nUsed all supplemental data in train.\n\nAlso tried to create second loader with pseudo labels on bad predicted samples. Or delete bad predicted samples. Used both loaders while training:\n\n```python\nloader_to_use = 'main_loader'\nif epoch % 3 == 0:\n    loader_to_use = 'second_loader'\n```\nBut it didn't help much.\n\n# Model\nFirst I tried simple quartznet, but with all strides set to 1.\nThen I added some features step by step and trained each experiment from previous best checkpoint.\n\n### Same for all runs:\n* Optimizer: LookAheadAdamW\n* Scheduler: Onecycle\n* Batch size: 80\n\n### List of experiments:\n* Simple 5x5 quartznet with stride = 1.  [lr = 8e-3] [epochs=200] CV:0.8035\n* Simple 5x5 quartznet with stride = 1. [lr = 7e-3] [epochs=150]  CV: 0.8106\n* Previous + SE blocks.  [lr = 6e-3] [epochs=171]  CV: 0.8160\n* Previous + Deep supervision (DSV) outputs for train loss. [lr = 5.8e-3] [epochs=175]  CV: 0.8211\n* Previous without DSV + higher aug probabilities. Note: score worse but it needed to train one checkpoint without DSV for better next train results. [lr = 6e-3] [epochs=180]  CV: 0.8177\n* Previous + DSV + increased dropout. [lr = 6.1e-3] [epochs=200]  CV: 0.8237\n* Previous - DSV + lower dropout. [lr = 6e-3] [epochs=173]  CV: 0.8188\n* Previous + Hyper Column (HYP) + 2 more blocks with 5 repeats. [lr=5.8e-3] [epochs=180] CV: 0.8252\n* Previous + masked 1conv + masked SE + 2D DepthWise CNN layer before classifier + more dropout. [lr=5.8e-3] [epochs=200]  CV: 0.8294\n* Previous + lstm after 2D Conv. [lr=1.75e-3] [epochs=125] CV: 0.8301\n\n\n# What didn't work for me:\n* Conformer models. I tried it but I guess without enough effort.\n* Adding attention or conformer blocks to Jasper models.\n\nWhen I used raw conformer and it has very low score. But when I changed input 2D conv to 1D conv or when I just deleted it - it was much better, but still little worse than my Jasper models.\n\n# What I should have done, but didn't:\n* Train more epochs\n* Try better with conformer or squeezeformer\n* Add post processing\n* Try seq2seq\n* Fixed length of input\n\nThanks to Kaggle for this great competition. Thanks to all top participants who shared their awesome solutions. \nCongratulations to all winners! \n",
    "2417184": "Great work @kolyaforrat did you develop in pytorch and port to tf ? I understand nemo jasper models are only in pytorch. ",
    "2412129": "Congratulations. Thanks for sharing very useful information regarding your solution and your experience from the competition.\n\"train more epochs\"- This  is limited by the notebook time of <9hrs. I once trained my model over many epochs and I could not submit the same notebook  due to time restrictions. I think clever part is to separate training and inference notebooks.\nHow to enforce fixed length of input?"
  }
}