{
  "id": 427297,
  "title": "Seq2seq model with CTC loss  OR  next token prediction model ?",
  "url": "/competitions/asl-fingerspelling/discussion/427297",
  "author_name": "Qsj",
  "post_date": "2023-07-27T09:25:51.453000",
  "votes": 15,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Considering both these two types of the models have very good result ( as inferred from public sharing ),which model structure should we choose?<br>\nps: I work with first one</p>",
  "messages": [
    {
      "id": 2361237,
      "postDate": "2023-07-27T09:25:51.453Z",
      "content": "<p>Considering both these two types of the models have very good result ( as inferred from public sharing ),which model structure should we choose?<br>\nps: I work with first one</p>",
      "rawMarkdown": "Considering both these two types of the models have very good result ( as inferred from public sharing ),which model structure should we choose?\nps: I work with first one",
      "votes": 15
    },
    {
      "id": 2362671,
      "postDate": "2023-07-28T08:32:55.893Z",
      "content": "<p>CTC models are much faster for infer for it decodes in parallel, <br>\nbut usually worked worse comparing to seq2seq models with similar parameters,<br>\n however using fast ctc models you can afford more parameters.. so it depends.<br>\nCurrently I'm using ctc only, I like ctc models, I think encoder only models are much easier to iterate.</p>",
      "rawMarkdown": "CTC models are much faster for infer for it decodes in parallel, \nbut usually worked worse comparing to seq2seq models with similar parameters,\n however using fast ctc models you can afford more parameters.. so it depends.\nCurrently I'm using ctc only, I like ctc models, I think encoder only models are much easier to iterate.",
      "votes": 9
    },
    {
      "id": 2364020,
      "postDate": "2023-07-29T04:42:03.813Z",
      "content": "<p>Although a CTC or an Encoder-Decoder model with Attention (or even a Transducer) would work for any sequence-to-sequence prediction task, I think this particular task needs a different approach.<br>\nIf you look at the nature of the data, it is inherently noisy - in the sense that it has missing frames in the video.<br>\nIn theory, I would assume that CTC models perform better since conditional independence is assumed (each output label is independent of each other, hence a missing frame won't affect future frames). Whereas in an Encoder-Decoder, the missing frame's error can cascade and lead to other errors in the sequence.<br>\nI didn't have time to train a model for this competition, but I would have chosen <a href=\"https://arxiv.org/abs/2005.08700\" target=\"_blank\">Mask CTC</a> or any conditional masking-based model. This predicts the \"easiest\" tokens in the sequence first and iteratively fills in the other time steps. A model like this would have a better chance of getting the predictions correct for missing frames as well. If we can combine this model with a conditional masked language model (maybe even use a dictionary to restrict possible words), it would further help decrease the error rate by providing linguistic information. I might give Mask CTC a try some time. The challenge would be either implementing mask CTC in TensorFlow or porting the output model from PyTorch.</p>",
      "rawMarkdown": "Although a CTC or an Encoder-Decoder model with Attention (or even a Transducer) would work for any sequence-to-sequence prediction task, I think this particular task needs a different approach.\nIf you look at the nature of the data, it is inherently noisy - in the sense that it has missing frames in the video.\nIn theory, I would assume that CTC models perform better since conditional independence is assumed (each output label is independent of each other, hence a missing frame won't affect future frames). Whereas in an Encoder-Decoder, the missing frame's error can cascade and lead to other errors in the sequence.\nI didn't have time to train a model for this competition, but I would have chosen [Mask CTC](https://arxiv.org/abs/2005.08700) or any conditional masking-based model. This predicts the \"easiest\" tokens in the sequence first and iteratively fills in the other time steps. A model like this would have a better chance of getting the predictions correct for missing frames as well. If we can combine this model with a conditional masked language model (maybe even use a dictionary to restrict possible words), it would further help decrease the error rate by providing linguistic information. I might give Mask CTC a try some time. The challenge would be either implementing mask CTC in TensorFlow or porting the output model from PyTorch.",
      "votes": 4
    },
    {
      "id": 2361643,
      "postDate": "2023-07-27T13:53:59.897Z",
      "content": "<p>Same with you, I work with the model with CTC loss. It performs worse when I change my model to encoder-decoder structure.</p>",
      "rawMarkdown": "Same with you, I work with the model with CTC loss. It performs worse when I change my model to encoder-decoder structure.",
      "votes": 1,
      "replies": [
        {
          "id": 2361922,
          "postDate": "2023-07-27T17:22:14.987Z",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/rock139\" target=\"_blank\">@rock139</a>. I have a question is that why is val_ctc_loss decreasing but my results on leaderboard are still stuck at 0.664??? I don'r understand this phenomenon</p>",
          "rawMarkdown": "Hey @rock139. I have a question is that why is val_ctc_loss decreasing but my results on leaderboard are still stuck at 0.664??? I don'r understand this phenomenon",
          "replies": [
            {
              "id": 2362951,
              "postDate": "2023-07-28T12:06:14.847Z",
              "content": "<p>I also met this. But I think in this competition, the performance metric is more reliable than validation loss.</p>",
              "rawMarkdown": "I also met this. But I think in this competition, the performance metric is more reliable than validation loss.",
              "votes": 1
            },
            {
              "id": 2363769,
              "postDate": "2023-07-28T23:24:24.013Z",
              "content": "<p>Perhaps your val_ctc_loss has not decreased enough yet.</p>",
              "rawMarkdown": "Perhaps your val_ctc_loss has not decreased enough yet."
            }
          ]
        }
      ]
    },
    {
      "id": 2361300,
      "postDate": "2023-07-27T10:13:31.310Z",
      "content": "<p>I think CTC loss should be more suitable for the Levenshtein distance metric. But currently my seq2seq model works much better than my CTC model…</p>",
      "rawMarkdown": "I think CTC loss should be more suitable for the Levenshtein distance metric. But currently my seq2seq model works much better than my CTC model...",
      "votes": 1,
      "replies": [
        {
          "id": 2361311,
          "postDate": "2023-07-27T10:23:55.677Z",
          "content": "<p>You mean the structure used in the notebook below works better? Amazing !<br>\n<a href=\"https://www.kaggle.com/code/markwijkhuizen/aslfr-transformer-training-inference\" target=\"_blank\">https://www.kaggle.com/code/markwijkhuizen/aslfr-transformer-training-inference</a></p>",
          "rawMarkdown": "You mean the structure used in the notebook below works better? Amazing !\nhttps://www.kaggle.com/code/markwijkhuizen/aslfr-transformer-training-inference",
          "votes": 1,
          "replies": [
            {
              "id": 2361387,
              "postDate": "2023-07-27T11:25:28.827Z",
              "content": "<p>Yep, I followed this encoder-decoder structure and added a lot of changes to achieve 0.771 LB. Then, I tried to transfer my best seq2seq model to CTC model by replacing the Transformer decoder with RNN decoders like LSTM. But my transferred CTC model only achieved ~0.75 LB score.</p>",
              "rawMarkdown": "Yep, I followed this encoder-decoder structure and added a lot of changes to achieve 0.771 LB. Then, I tried to transfer my best seq2seq model to CTC model by replacing the Transformer decoder with RNN decoders like LSTM. But my transferred CTC model only achieved ~0.75 LB score.",
              "votes": 3
            }
          ]
        },
        {
          "id": 2372155,
          "postDate": "2023-08-03T14:14:41.910Z",
          "content": "<p>I want to know what card you used to train the model with? GPU or TPU?</p>",
          "rawMarkdown": "I want to know what card you used to train the model with? GPU or TPU?",
          "replies": [
            {
              "id": 2372160,
              "postDate": "2023-08-03T14:18:27.573Z",
              "content": "<p>I used both, but mainly on TPU</p>",
              "rawMarkdown": "I used both, but mainly on TPU"
            }
          ]
        }
      ]
    },
    {
      "id": 2386133,
      "postDate": "2023-08-11T18:23:44.990Z",
      "content": "<p>In my opinion autoregressive encoder-decoder models suffer from exposure bias. This is especially bad if the input is missing a lot of frames. The model will just \"memorize\" the output sequence without relying much on the encoder state.  I could not find a good way to prevent the decoder from over-fitting. CTC based models  have no such issue and have performed much better for me. Also they are much faster at inference.</p>",
      "rawMarkdown": "In my opinion autoregressive encoder-decoder models suffer from exposure bias. This is especially bad if the input is missing a lot of frames. The model will just \"memorize\" the output sequence without relying much on the encoder state.  I could not find a good way to prevent the decoder from over-fitting. CTC based models  have no such issue and have performed much better for me. Also they are much faster at inference."
    }
  ],
  "comments": [
    {
      "id": 2362671,
      "author_name": "gezi",
      "author_url": "",
      "post_date": "2023-07-28T08:32:55.893000",
      "content": "<p>CTC models are much faster for infer for it decodes in parallel, <br>\nbut usually worked worse comparing to seq2seq models with similar parameters,<br>\n however using fast ctc models you can afford more parameters.. so it depends.<br>\nCurrently I'm using ctc only, I like ctc models, I think encoder only models are much easier to iterate.</p>",
      "votes": 9,
      "replies": []
    },
    {
      "id": 2364020,
      "author_name": "sknadig",
      "author_url": "",
      "post_date": "2023-07-29T04:42:03.813000",
      "content": "<p>Although a CTC or an Encoder-Decoder model with Attention (or even a Transducer) would work for any sequence-to-sequence prediction task, I think this particular task needs a different approach.<br>\nIf you look at the nature of the data, it is inherently noisy - in the sense that it has missing frames in the video.<br>\nIn theory, I would assume that CTC models perform better since conditional independence is assumed (each output label is independent of each other, hence a missing frame won't affect future frames). Whereas in an Encoder-Decoder, the missing frame's error can cascade and lead to other errors in the sequence.<br>\nI didn't have time to train a model for this competition, but I would have chosen <a href=\"https://arxiv.org/abs/2005.08700\" target=\"_blank\">Mask CTC</a> or any conditional masking-based model. This predicts the \"easiest\" tokens in the sequence first and iteratively fills in the other time steps. A model like this would have a better chance of getting the predictions correct for missing frames as well. If we can combine this model with a conditional masked language model (maybe even use a dictionary to restrict possible words), it would further help decrease the error rate by providing linguistic information. I might give Mask CTC a try some time. The challenge would be either implementing mask CTC in TensorFlow or porting the output model from PyTorch.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 2361643,
      "author_name": "Rock",
      "author_url": "",
      "post_date": "2023-07-27T13:53:59.897000",
      "content": "<p>Same with you, I work with the model with CTC loss. It performs worse when I change my model to encoder-decoder structure.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2361922,
          "author_name": "🍀<->🍀",
          "author_url": "",
          "post_date": "2023-07-27T17:22:14.987000",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/rock139\" target=\"_blank\">@rock139</a>. I have a question is that why is val_ctc_loss decreasing but my results on leaderboard are still stuck at 0.664??? I don'r understand this phenomenon</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2362951,
              "author_name": "Rock",
              "author_url": "",
              "post_date": "2023-07-28T12:06:14.847000",
              "content": "<p>I also met this. But I think in this competition, the performance metric is more reliable than validation loss.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2363769,
              "author_name": "Scenery SunFireInk",
              "author_url": "",
              "post_date": "2023-07-28T23:24:24.013000",
              "content": "<p>Perhaps your val_ctc_loss has not decreased enough yet.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2361300,
      "author_name": "Yu Wu",
      "author_url": "",
      "post_date": "2023-07-27T10:13:31.310000",
      "content": "<p>I think CTC loss should be more suitable for the Levenshtein distance metric. But currently my seq2seq model works much better than my CTC model…</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2361311,
          "author_name": "Qsj",
          "author_url": "",
          "post_date": "2023-07-27T10:23:55.677000",
          "content": "<p>You mean the structure used in the notebook below works better? Amazing !<br>\n<a href=\"https://www.kaggle.com/code/markwijkhuizen/aslfr-transformer-training-inference\" target=\"_blank\">https://www.kaggle.com/code/markwijkhuizen/aslfr-transformer-training-inference</a></p>",
          "votes": 1,
          "replies": [
            {
              "id": 2361387,
              "author_name": "Yu Wu",
              "author_url": "",
              "post_date": "2023-07-27T11:25:28.827000",
              "content": "<p>Yep, I followed this encoder-decoder structure and added a lot of changes to achieve 0.771 LB. Then, I tried to transfer my best seq2seq model to CTC model by replacing the Transformer decoder with RNN decoders like LSTM. But my transferred CTC model only achieved ~0.75 LB score.</p>",
              "votes": 3,
              "replies": []
            }
          ]
        },
        {
          "id": 2372155,
          "author_name": "Scenery SunFireInk",
          "author_url": "",
          "post_date": "2023-08-03T14:14:41.910000",
          "content": "<p>I want to know what card you used to train the model with? GPU or TPU?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2372160,
              "author_name": "Yu Wu",
              "author_url": "",
              "post_date": "2023-08-03T14:18:27.573000",
              "content": "<p>I used both, but mainly on TPU</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2386133,
      "author_name": "FatGPT",
      "author_url": "",
      "post_date": "2023-08-11T18:23:44.990000",
      "content": "<p>In my opinion autoregressive encoder-decoder models suffer from exposure bias. This is especially bad if the input is missing a lot of frames. The model will just \"memorize\" the output sequence without relying much on the encoder state.  I could not find a good way to prevent the decoder from over-fitting. CTC based models  have no such issue and have performed much better for me. Also they are much faster at inference.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2361237": "Considering both these two types of the models have very good result ( as inferred from public sharing ),which model structure should we choose?\nps: I work with first one",
    "2362671": "CTC models are much faster for infer for it decodes in parallel, \nbut usually worked worse comparing to seq2seq models with similar parameters,\n however using fast ctc models you can afford more parameters.. so it depends.\nCurrently I'm using ctc only, I like ctc models, I think encoder only models are much easier to iterate.",
    "2364020": "Although a CTC or an Encoder-Decoder model with Attention (or even a Transducer) would work for any sequence-to-sequence prediction task, I think this particular task needs a different approach.\nIf you look at the nature of the data, it is inherently noisy - in the sense that it has missing frames in the video.\nIn theory, I would assume that CTC models perform better since conditional independence is assumed (each output label is independent of each other, hence a missing frame won't affect future frames). Whereas in an Encoder-Decoder, the missing frame's error can cascade and lead to other errors in the sequence.\nI didn't have time to train a model for this competition, but I would have chosen [Mask CTC](https://arxiv.org/abs/2005.08700) or any conditional masking-based model. This predicts the \"easiest\" tokens in the sequence first and iteratively fills in the other time steps. A model like this would have a better chance of getting the predictions correct for missing frames as well. If we can combine this model with a conditional masked language model (maybe even use a dictionary to restrict possible words), it would further help decrease the error rate by providing linguistic information. I might give Mask CTC a try some time. The challenge would be either implementing mask CTC in TensorFlow or porting the output model from PyTorch.",
    "2361643": "Same with you, I work with the model with CTC loss. It performs worse when I change my model to encoder-decoder structure.",
    "2361300": "I think CTC loss should be more suitable for the Levenshtein distance metric. But currently my seq2seq model works much better than my CTC model...",
    "2386133": "In my opinion autoregressive encoder-decoder models suffer from exposure bias. This is especially bad if the input is missing a lot of frames. The model will just \"memorize\" the output sequence without relying much on the encoder state.  I could not find a good way to prevent the decoder from over-fitting. CTC based models  have no such issue and have performed much better for me. Also they are much faster at inference."
  }
}