{
  "id": 434475,
  "title": "[The 11th place] Shallow encoder-decoder model also works",
  "url": "/competitions/asl-fingerspelling/discussion/434475",
  "author_name": "bliao",
  "post_date": "2023-08-25T10:25:15.298000",
  "votes": 18,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Acknowledgments for the great public notebooks:<br>\n[1] <a href=\"https://www.kaggle.com/code/markwijkhuizen/aslfr-eda-preprocessing-dataset\" target=\"_blank\">MARK WIJKHUIZEN's data preprocessing notebook </a><br>\n[2] <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406684\" target=\"_blank\">HOYSO48's first place notebook from ISLR</a></p>\n<p>Reproducible Codes:</p>\n<ol>\n<li><a href=\"https://www.kaggle.com/code/baohaoliao/aslfr-data-preprocessing\" target=\"_blank\">Data preprocessing notebook</a></li>\n<li><a href=\"https://github.com/BaohaoLiao/aslfrv4\" target=\"_blank\">Training codes</a>: I need some time to clean it and will guide you in the following. In case someone can't wait to check it.</li>\n</ol>\n<h1>Data preprocessing</h1>\n<p>My data preprocessing and split notebook is based on [1]. The differences are:</p>\n<ol>\n<li>I used the same landmarks (only x and y) from [2].</li>\n<li>I split data based on unique phrases into 10 folds, and use the first fold as the validation set. There are no shared phrases in different folds. For data splitting, I found unique phrase split = random split &gt; participant id split.</li>\n<li>Use the same way to preprocess the supplementary data.</li>\n</ol>\n<h1>Data augmentation (<a href=\"https://github.com/BaohaoLiao/aslfrv4/blob/main/mask/datav2_distributed.py\" target=\"_blank\">file</a>)</h1>\n<p>I almost used the same data augmentation and hyper-parameters as [2]. The main differences are:</p>\n<ol>\n<li>I randomly concatenated two samples with a prob of 50%. I found it useful for overfitting.</li>\n<li>For the input to the decoder, I randomly (20%) replaced some ground-truth characters with random characters from the vocab. I apply this trick to avoid exposure bias.</li>\n</ol>\n<h1>Model architecture (<a href=\"https://github.com/BaohaoLiao/aslfrv4/blob/main/auto_ctc/conformerencoder_transformerdecoder_mask_droppath_ctc.py\" target=\"_blank\">file</a>)</h1>\n<p>I use Conformer as the encoder and the vanilla Transformer decoder. Conformer is better than the vanilla Transformer encoder, since it focuses on both local and global relations. The main highlights here are:</p>\n<ol>\n<li><p>I use both CTC loss and cross-entropy loss. After the encoder, the CTC loss is applied to the encoder output. And cross-entropy loss is applied to the decoder output. The interpolation weights for these two losses are 0.2 for CTC loss and 0.8 for cross-entropy loss. The benefits of this setting are: (1) One trained model can be used in two ways, either non-autoregressive generation with the encoder or autoregressive generation with the decoder. In the end, both generations achieved the same 0.791 public LB, (2) One loss could be a regularization term to the other.  When you only want to do non-autoregressive generation, you can throw away the decoder parameters, which allows you to use more parameters during training.</p></li>\n<li><p>The hyper-parameters are: 5-layer encoder with hidden_dim=384, mlp_dim=1024, conv_dim=768, num_heads=6, 3-layer decoder with hidden_dim=256, mlp_dim=512, num_heads=4. </p></li>\n</ol>\n<h1>Three-phase training (<a href=\"https://github.com/BaohaoLiao/aslfrv4/blob/main/train_cnnencoder_transformerdecoder_mask_ctc_distributed.py\" target=\"_blank\">file0</a> and <a href=\"https://github.com/BaohaoLiao/aslfrv4/blob/main/train_cnnencoder_transformerdecoder_mask_ctc_awp_distributed.py\" target=\"_blank\">file1</a>)</h1>\n<p>All training uses AdamW, inverse square root schedule, weight decay=0.001 and max norm=5, lr=5e-4, batch size=512, warmup ratio=0.2, label smoothing=0.1, frame_length=368.</p>\n<ol>\n<li>For the first phase, I only train on the training set and exclude the validation set with #epoch=100.</li>\n<li>For the second phase, I include all training data and supplemental data for another 150 epochs.</li>\n<li>For the third phase, I use AWP with awp_delta=0.2 and awp_eps=0 on all training data for 300 epochs. AWP is good for generalization, better than rdrop for my case.</li>\n</ol>\n<h1>What doesn't work</h1>\n<ol>\n<li>BPE: I try to use subwords rather than characters, but it doesn't work.</li>\n<li>Too wide but shallow model: For the abovementioned model, I quantize the model in FP16. The number model's parameters are about 18M, 37MB. We can use about 40M parameters if we use dynamic quantization. But dynamic quantization slows down the inference speed. For your reference, my FP16 CTC model runs 2h30m, while the dynamic quantized one runs about 4h. There is only 0.001 LB performance drop for the dynamic quantized one. Nothing drops from FP32 to FP16. In the last week, I tried to use an 8-layer encoder and a 4-layer decoder (same dimension as above). But I have to use frame_length=100 to reduce the inference time within 5 hours. The results are not good. I think there might be some overfitting with such a larger model.</li>\n<li>Ensemble between encoder and decoder: My model could do both autoregressive and non-autoregressive generation at the same time. I first generate the output with the encoder, and do an average on the probabilities of the autoregressive and non-autoregressive output at each time frame. But the result stays the same.</li>\n<li>Other data: I used <a href=\"https://home.ttic.edu/~klivescu/ChicagoFSWild.htm#download\" target=\"_blank\">ChicagoFSWild and ChicagoFSWild+</a>, it doesn't help. Domain shift is the main reason.</li>\n</ol>\n<h1>What could make my model better</h1>\n<ol>\n<li>Inspired by other top-rank methods, including pose landmarks and z coordination might be better.</li>\n<li>Use a deeper but narrower model.</li>\n</ol>",
  "messages": [
    {
      "id": 2407927,
      "postDate": "2023-08-25T10:25:15.297Z",
      "content": "<p>Acknowledgments for the great public notebooks:<br>\n[1] <a href=\"https://www.kaggle.com/code/markwijkhuizen/aslfr-eda-preprocessing-dataset\" target=\"_blank\">MARK WIJKHUIZEN's data preprocessing notebook </a><br>\n[2] <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406684\" target=\"_blank\">HOYSO48's first place notebook from ISLR</a></p>\n<p>Reproducible Codes:</p>\n<ol>\n<li><a href=\"https://www.kaggle.com/code/baohaoliao/aslfr-data-preprocessing\" target=\"_blank\">Data preprocessing notebook</a></li>\n<li><a href=\"https://github.com/BaohaoLiao/aslfrv4\" target=\"_blank\">Training codes</a>: I need some time to clean it and will guide you in the following. In case someone can't wait to check it.</li>\n</ol>\n<h1>Data preprocessing</h1>\n<p>My data preprocessing and split notebook is based on [1]. The differences are:</p>\n<ol>\n<li>I used the same landmarks (only x and y) from [2].</li>\n<li>I split data based on unique phrases into 10 folds, and use the first fold as the validation set. There are no shared phrases in different folds. For data splitting, I found unique phrase split = random split &gt; participant id split.</li>\n<li>Use the same way to preprocess the supplementary data.</li>\n</ol>\n<h1>Data augmentation (<a href=\"https://github.com/BaohaoLiao/aslfrv4/blob/main/mask/datav2_distributed.py\" target=\"_blank\">file</a>)</h1>\n<p>I almost used the same data augmentation and hyper-parameters as [2]. The main differences are:</p>\n<ol>\n<li>I randomly concatenated two samples with a prob of 50%. I found it useful for overfitting.</li>\n<li>For the input to the decoder, I randomly (20%) replaced some ground-truth characters with random characters from the vocab. I apply this trick to avoid exposure bias.</li>\n</ol>\n<h1>Model architecture (<a href=\"https://github.com/BaohaoLiao/aslfrv4/blob/main/auto_ctc/conformerencoder_transformerdecoder_mask_droppath_ctc.py\" target=\"_blank\">file</a>)</h1>\n<p>I use Conformer as the encoder and the vanilla Transformer decoder. Conformer is better than the vanilla Transformer encoder, since it focuses on both local and global relations. The main highlights here are:</p>\n<ol>\n<li><p>I use both CTC loss and cross-entropy loss. After the encoder, the CTC loss is applied to the encoder output. And cross-entropy loss is applied to the decoder output. The interpolation weights for these two losses are 0.2 for CTC loss and 0.8 for cross-entropy loss. The benefits of this setting are: (1) One trained model can be used in two ways, either non-autoregressive generation with the encoder or autoregressive generation with the decoder. In the end, both generations achieved the same 0.791 public LB, (2) One loss could be a regularization term to the other.  When you only want to do non-autoregressive generation, you can throw away the decoder parameters, which allows you to use more parameters during training.</p></li>\n<li><p>The hyper-parameters are: 5-layer encoder with hidden_dim=384, mlp_dim=1024, conv_dim=768, num_heads=6, 3-layer decoder with hidden_dim=256, mlp_dim=512, num_heads=4. </p></li>\n</ol>\n<h1>Three-phase training (<a href=\"https://github.com/BaohaoLiao/aslfrv4/blob/main/train_cnnencoder_transformerdecoder_mask_ctc_distributed.py\" target=\"_blank\">file0</a> and <a href=\"https://github.com/BaohaoLiao/aslfrv4/blob/main/train_cnnencoder_transformerdecoder_mask_ctc_awp_distributed.py\" target=\"_blank\">file1</a>)</h1>\n<p>All training uses AdamW, inverse square root schedule, weight decay=0.001 and max norm=5, lr=5e-4, batch size=512, warmup ratio=0.2, label smoothing=0.1, frame_length=368.</p>\n<ol>\n<li>For the first phase, I only train on the training set and exclude the validation set with #epoch=100.</li>\n<li>For the second phase, I include all training data and supplemental data for another 150 epochs.</li>\n<li>For the third phase, I use AWP with awp_delta=0.2 and awp_eps=0 on all training data for 300 epochs. AWP is good for generalization, better than rdrop for my case.</li>\n</ol>\n<h1>What doesn't work</h1>\n<ol>\n<li>BPE: I try to use subwords rather than characters, but it doesn't work.</li>\n<li>Too wide but shallow model: For the abovementioned model, I quantize the model in FP16. The number model's parameters are about 18M, 37MB. We can use about 40M parameters if we use dynamic quantization. But dynamic quantization slows down the inference speed. For your reference, my FP16 CTC model runs 2h30m, while the dynamic quantized one runs about 4h. There is only 0.001 LB performance drop for the dynamic quantized one. Nothing drops from FP32 to FP16. In the last week, I tried to use an 8-layer encoder and a 4-layer decoder (same dimension as above). But I have to use frame_length=100 to reduce the inference time within 5 hours. The results are not good. I think there might be some overfitting with such a larger model.</li>\n<li>Ensemble between encoder and decoder: My model could do both autoregressive and non-autoregressive generation at the same time. I first generate the output with the encoder, and do an average on the probabilities of the autoregressive and non-autoregressive output at each time frame. But the result stays the same.</li>\n<li>Other data: I used <a href=\"https://home.ttic.edu/~klivescu/ChicagoFSWild.htm#download\" target=\"_blank\">ChicagoFSWild and ChicagoFSWild+</a>, it doesn't help. Domain shift is the main reason.</li>\n</ol>\n<h1>What could make my model better</h1>\n<ol>\n<li>Inspired by other top-rank methods, including pose landmarks and z coordination might be better.</li>\n<li>Use a deeper but narrower model.</li>\n</ol>",
      "rawMarkdown": "Acknowledgments for the great public notebooks:\n[1] [MARK WIJKHUIZEN's data preprocessing notebook ](https://www.kaggle.com/code/markwijkhuizen/aslfr-eda-preprocessing-dataset)\n[2] [HOYSO48's first place notebook from ISLR](https://www.kaggle.com/competitions/asl-signs/discussion/406684)\n\nReproducible Codes:\n1. [Data preprocessing notebook](https://www.kaggle.com/code/baohaoliao/aslfr-data-preprocessing)\n2. [Training codes](https://github.com/BaohaoLiao/aslfrv4): I need some time to clean it and will guide you in the following. In case someone can't wait to check it.\n\n\n# Data preprocessing \nMy data preprocessing and split notebook is based on [1]. The differences are:\n1. I used the same landmarks (only x and y) from [2].\n2. I split data based on unique phrases into 10 folds, and use the first fold as the validation set. There are no shared phrases in different folds. For data splitting, I found unique phrase split = random split > participant id split.\n3. Use the same way to preprocess the supplementary data.\n\n# Data augmentation ([file](https://github.com/BaohaoLiao/aslfrv4/blob/main/mask/datav2_distributed.py))\nI almost used the same data augmentation and hyper-parameters as [2]. The main differences are:\n1. I randomly concatenated two samples with a prob of 50%. I found it useful for overfitting.\n2. For the input to the decoder, I randomly (20%) replaced some ground-truth characters with random characters from the vocab. I apply this trick to avoid exposure bias.\n\n# Model architecture ([file](https://github.com/BaohaoLiao/aslfrv4/blob/main/auto_ctc/conformerencoder_transformerdecoder_mask_droppath_ctc.py))\nI use Conformer as the encoder and the vanilla Transformer decoder. Conformer is better than the vanilla Transformer encoder, since it focuses on both local and global relations. The main highlights here are:\n1. I use both CTC loss and cross-entropy loss. After the encoder, the CTC loss is applied to the encoder output. And cross-entropy loss is applied to the decoder output. The interpolation weights for these two losses are 0.2 for CTC loss and 0.8 for cross-entropy loss. The benefits of this setting are: (1) One trained model can be used in two ways, either non-autoregressive generation with the encoder or autoregressive generation with the decoder. In the end, both generations achieved the same 0.791 public LB, (2) One loss could be a regularization term to the other.  When you only want to do non-autoregressive generation, you can throw away the decoder parameters, which allows you to use more parameters during training.\n\n2. The hyper-parameters are: 5-layer encoder with hidden_dim=384, mlp_dim=1024, conv_dim=768, num_heads=6, 3-layer decoder with hidden_dim=256, mlp_dim=512, num_heads=4. \n\n# Three-phase training ([file0](https://github.com/BaohaoLiao/aslfrv4/blob/main/train_cnnencoder_transformerdecoder_mask_ctc_distributed.py) and [file1](https://github.com/BaohaoLiao/aslfrv4/blob/main/train_cnnencoder_transformerdecoder_mask_ctc_awp_distributed.py))\nAll training uses AdamW, inverse square root schedule, weight decay=0.001 and max norm=5, lr=5e-4, batch size=512, warmup ratio=0.2, label smoothing=0.1, frame_length=368.\n1. For the first phase, I only train on the training set and exclude the validation set with #epoch=100.\n2. For the second phase, I include all training data and supplemental data for another 150 epochs.\n3. For the third phase, I use AWP with awp_delta=0.2 and awp_eps=0 on all training data for 300 epochs. AWP is good for generalization, better than rdrop for my case.\n\n\n# What doesn't work\n1. BPE: I try to use subwords rather than characters, but it doesn't work.\n2. Too wide but shallow model: For the abovementioned model, I quantize the model in FP16. The number model's parameters are about 18M, 37MB. We can use about 40M parameters if we use dynamic quantization. But dynamic quantization slows down the inference speed. For your reference, my FP16 CTC model runs 2h30m, while the dynamic quantized one runs about 4h. There is only 0.001 LB performance drop for the dynamic quantized one. Nothing drops from FP32 to FP16. In the last week, I tried to use an 8-layer encoder and a 4-layer decoder (same dimension as above). But I have to use frame_length=100 to reduce the inference time within 5 hours. The results are not good. I think there might be some overfitting with such a larger model.\n3. Ensemble between encoder and decoder: My model could do both autoregressive and non-autoregressive generation at the same time. I first generate the output with the encoder, and do an average on the probabilities of the autoregressive and non-autoregressive output at each time frame. But the result stays the same.\n4. Other data: I used [ChicagoFSWild and ChicagoFSWild+](https://home.ttic.edu/~klivescu/ChicagoFSWild.htm#download), it doesn't help. Domain shift is the main reason.\n\n# What could make my model better\n1. Inspired by other top-rank methods, including pose landmarks and z coordination might be better.\n2. Use a deeper but narrower model.",
      "votes": 17
    },
    {
      "id": 2418463,
      "postDate": "2023-09-01T09:35:15.117Z",
      "content": "<p>Great work and achievement for both of you 🙂</p>",
      "rawMarkdown": "Great work and achievement for both of you 🙂",
      "votes": 1
    },
    {
      "id": 2408316,
      "postDate": "2023-08-25T15:07:19.173Z",
      "content": "<p>Congratulations. Thanks for sharing details of your model and best practices. </p>",
      "rawMarkdown": "Congratulations. Thanks for sharing details of your model and best practices. ",
      "votes": 1
    },
    {
      "id": 2408035,
      "postDate": "2023-08-25T11:53:02.590Z",
      "content": "<p>Thanks for sharing! <br>\nIf only using ctc encoder as outputs, does it show better performance with ctc+seq2seq loss training? (I mean training add decoder and infer drop decoder does this help encoder improve performance?)</p>",
      "rawMarkdown": "Thanks for sharing! \nIf only using ctc encoder as outputs, does it show better performance with ctc+seq2seq loss training? (I mean training add decoder and infer drop decoder does this help encoder improve performance?)",
      "replies": [
        {
          "id": 2408045,
          "postDate": "2023-08-25T12:02:16.120Z",
          "content": "<p>Yes, it helps, about + 0.005.</p>",
          "rawMarkdown": "Yes, it helps, about + 0.005.",
          "replies": [
            {
              "id": 2408051,
              "postDate": "2023-08-25T12:07:09.237Z",
              "content": "<p>Well that's cool, seems in generall ctc+seq2seq might be the standard method here.<br>\nI tried this early when using tf but not get any gain, maybe some wrong setup. </p>",
              "rawMarkdown": "Well that's cool, seems in generall ctc+seq2seq might be the standard method here.\nI tried this early when using tf but not get any gain, maybe some wrong setup. ",
              "votes": 1
            },
            {
              "id": 2422296,
              "postDate": "2023-09-03T22:31:43.910Z",
              "content": "<p>It didn't provide any gain for me either. I tried with and without random replacement of decoder labels. Post competition I tested some of the ideas I discarded and realized that they actually work when training with larger models for longer. I wonder if the same thing happened to you.</p>",
              "rawMarkdown": "It didn't provide any gain for me either. I tried with and without random replacement of decoder labels. Post competition I tested some of the ideas I discarded and realized that they actually work when training with larger models for longer. I wonder if the same thing happened to you."
            },
            {
              "id": 2423410,
              "postDate": "2023-09-04T15:27:40.697Z",
              "content": "<p>For efficiency, I didn't train from scratch when trying something new. I always take the best-performed checkpoint from the latest trial. So I can't give you a certain answer. </p>",
              "rawMarkdown": "For efficiency, I didn't train from scratch when trying something new. I always take the best-performed checkpoint from the latest trial. So I can't give you a certain answer. ",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2426081,
      "postDate": "2023-09-06T11:51:58.153Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2418463,
      "author_name": "Talha Barkaat Ahmad ☑️",
      "author_url": "",
      "post_date": "2023-09-01T09:35:15.117000",
      "content": "<p>Great work and achievement for both of you 🙂</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2408316,
      "author_name": "C R Suthikshn Kumar",
      "author_url": "",
      "post_date": "2023-08-25T15:07:19.173000",
      "content": "<p>Congratulations. Thanks for sharing details of your model and best practices. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2408035,
      "author_name": "gezi",
      "author_url": "",
      "post_date": "2023-08-25T11:53:02.590000",
      "content": "<p>Thanks for sharing! <br>\nIf only using ctc encoder as outputs, does it show better performance with ctc+seq2seq loss training? (I mean training add decoder and infer drop decoder does this help encoder improve performance?)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2408045,
          "author_name": "bliao",
          "author_url": "",
          "post_date": "2023-08-25T12:02:16.120000",
          "content": "<p>Yes, it helps, about + 0.005.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2408051,
              "author_name": "gezi",
              "author_url": "",
              "post_date": "2023-08-25T12:07:09.237000",
              "content": "<p>Well that's cool, seems in generall ctc+seq2seq might be the standard method here.<br>\nI tried this early when using tf but not get any gain, maybe some wrong setup. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2422296,
              "author_name": "vialactea",
              "author_url": "",
              "post_date": "2023-09-03T22:31:43.910000",
              "content": "<p>It didn't provide any gain for me either. I tried with and without random replacement of decoder labels. Post competition I tested some of the ideas I discarded and realized that they actually work when training with larger models for longer. I wonder if the same thing happened to you.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2423410,
              "author_name": "bliao",
              "author_url": "",
              "post_date": "2023-09-04T15:27:40.697000",
              "content": "<p>For efficiency, I didn't train from scratch when trying something new. I always take the best-performed checkpoint from the latest trial. So I can't give you a certain answer. </p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2426081,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-09-06T11:51:58.153000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2407927": "Acknowledgments for the great public notebooks:\n[1] [MARK WIJKHUIZEN's data preprocessing notebook ](https://www.kaggle.com/code/markwijkhuizen/aslfr-eda-preprocessing-dataset)\n[2] [HOYSO48's first place notebook from ISLR](https://www.kaggle.com/competitions/asl-signs/discussion/406684)\n\nReproducible Codes:\n1. [Data preprocessing notebook](https://www.kaggle.com/code/baohaoliao/aslfr-data-preprocessing)\n2. [Training codes](https://github.com/BaohaoLiao/aslfrv4): I need some time to clean it and will guide you in the following. In case someone can't wait to check it.\n\n\n# Data preprocessing \nMy data preprocessing and split notebook is based on [1]. The differences are:\n1. I used the same landmarks (only x and y) from [2].\n2. I split data based on unique phrases into 10 folds, and use the first fold as the validation set. There are no shared phrases in different folds. For data splitting, I found unique phrase split = random split > participant id split.\n3. Use the same way to preprocess the supplementary data.\n\n# Data augmentation ([file](https://github.com/BaohaoLiao/aslfrv4/blob/main/mask/datav2_distributed.py))\nI almost used the same data augmentation and hyper-parameters as [2]. The main differences are:\n1. I randomly concatenated two samples with a prob of 50%. I found it useful for overfitting.\n2. For the input to the decoder, I randomly (20%) replaced some ground-truth characters with random characters from the vocab. I apply this trick to avoid exposure bias.\n\n# Model architecture ([file](https://github.com/BaohaoLiao/aslfrv4/blob/main/auto_ctc/conformerencoder_transformerdecoder_mask_droppath_ctc.py))\nI use Conformer as the encoder and the vanilla Transformer decoder. Conformer is better than the vanilla Transformer encoder, since it focuses on both local and global relations. The main highlights here are:\n1. I use both CTC loss and cross-entropy loss. After the encoder, the CTC loss is applied to the encoder output. And cross-entropy loss is applied to the decoder output. The interpolation weights for these two losses are 0.2 for CTC loss and 0.8 for cross-entropy loss. The benefits of this setting are: (1) One trained model can be used in two ways, either non-autoregressive generation with the encoder or autoregressive generation with the decoder. In the end, both generations achieved the same 0.791 public LB, (2) One loss could be a regularization term to the other.  When you only want to do non-autoregressive generation, you can throw away the decoder parameters, which allows you to use more parameters during training.\n\n2. The hyper-parameters are: 5-layer encoder with hidden_dim=384, mlp_dim=1024, conv_dim=768, num_heads=6, 3-layer decoder with hidden_dim=256, mlp_dim=512, num_heads=4. \n\n# Three-phase training ([file0](https://github.com/BaohaoLiao/aslfrv4/blob/main/train_cnnencoder_transformerdecoder_mask_ctc_distributed.py) and [file1](https://github.com/BaohaoLiao/aslfrv4/blob/main/train_cnnencoder_transformerdecoder_mask_ctc_awp_distributed.py))\nAll training uses AdamW, inverse square root schedule, weight decay=0.001 and max norm=5, lr=5e-4, batch size=512, warmup ratio=0.2, label smoothing=0.1, frame_length=368.\n1. For the first phase, I only train on the training set and exclude the validation set with #epoch=100.\n2. For the second phase, I include all training data and supplemental data for another 150 epochs.\n3. For the third phase, I use AWP with awp_delta=0.2 and awp_eps=0 on all training data for 300 epochs. AWP is good for generalization, better than rdrop for my case.\n\n\n# What doesn't work\n1. BPE: I try to use subwords rather than characters, but it doesn't work.\n2. Too wide but shallow model: For the abovementioned model, I quantize the model in FP16. The number model's parameters are about 18M, 37MB. We can use about 40M parameters if we use dynamic quantization. But dynamic quantization slows down the inference speed. For your reference, my FP16 CTC model runs 2h30m, while the dynamic quantized one runs about 4h. There is only 0.001 LB performance drop for the dynamic quantized one. Nothing drops from FP32 to FP16. In the last week, I tried to use an 8-layer encoder and a 4-layer decoder (same dimension as above). But I have to use frame_length=100 to reduce the inference time within 5 hours. The results are not good. I think there might be some overfitting with such a larger model.\n3. Ensemble between encoder and decoder: My model could do both autoregressive and non-autoregressive generation at the same time. I first generate the output with the encoder, and do an average on the probabilities of the autoregressive and non-autoregressive output at each time frame. But the result stays the same.\n4. Other data: I used [ChicagoFSWild and ChicagoFSWild+](https://home.ttic.edu/~klivescu/ChicagoFSWild.htm#download), it doesn't help. Domain shift is the main reason.\n\n# What could make my model better\n1. Inspired by other top-rank methods, including pose landmarks and z coordination might be better.\n2. Use a deeper but narrower model.",
    "2418463": "Great work and achievement for both of you 🙂",
    "2408316": "Congratulations. Thanks for sharing details of your model and best practices. ",
    "2408035": "Thanks for sharing! \nIf only using ctc encoder as outputs, does it show better performance with ctc+seq2seq loss training? (I mean training add decoder and infer drop decoder does this help encoder improve performance?)",
    "2426081": ""
  }
}