{
  "id": 419331,
  "title": "Dealing with variable output size",
  "url": "/competitions/asl-fingerspelling/discussion/419331",
  "author_name": "Mert Enes Yurtseven",
  "post_date": "2023-06-25T11:21:12.878000",
  "votes": 0,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I know there are multiple threads and even some explanations on this topic but I am not able to implement (errors): any of the proposed methods like CTC loss. How we should process output string and how to design last layer of the model and suitable loss function?</p>",
  "messages": [
    {
      "id": 2317085,
      "postDate": "2023-06-25T12:37:14.717Z",
      "content": "<p>I'm not sure what specific problem you're facing, so I'll describe how I'm trying to approach the problem and note that I'm using PyTorch.</p>\n<p>Regarding CTC loss:</p>\n<ul>\n<li>The arguments for CTC loss are same in both PyTorch and TensorFlow.</li>\n<li>Usually, the input is structured with a shape of (Batch-size, N_Frames, N_Landmarks).</li>\n<li>Any approach can work, but typically one would use 1D convolutions, RNNs.</li>\n<li>model output shape is (Batch-size, N-Frames, N-features). it is classified into 'N-characters + 1' categories. so final shape is (Batch-size, N-Frame, N-characters+1). The extra category accounts for the blank frames, which are  meaningless gesture that occurs between the hand gestures representing A and B when expressing 'AB'.</li>\n<li>During the process of creating input batches, padding is added to the input sequences and labels. Therefore, CTC loss requires the length of each sequence and label before padding within the batch.</li>\n<li>Anyway, you can input a tensor of shape (Batch-size, N-Frames, N-Landmarks) into the model and obtain a tensor of shape (Batch-size, N-Frames, N-characters+1). Then, you can provide the output tensor, along with the labels and the lengths of the sequences and labels before padding, to the CTC loss.</li>\n<li>During the inference process, the model predicts one character for each frame of the sequence. There may be cases where the model predicts a blank as well. For example, the predicted string could be 'f u n <strong>blank</strong> c <strong>blank</strong> t i o n'.</li>\n<li>In the evaluation section of this competition, the code <code>prediction_str = \"\".join([rev_character_map.get(s, \"\") for s in np.argmax(output[REQUIRED_OUTPUT], axis=1)])</code> is used to create the output string. Since the <code>blank</code> symbol is not included in the <code>rev_character_map</code>, it is ignored, and only meaningful characters like 'function' are retained in the final predicted string.</li>\n</ul>",
      "rawMarkdown": "I'm not sure what specific problem you're facing, so I'll describe how I'm trying to approach the problem and note that I'm using PyTorch.\n\nRegarding CTC loss:\n\n- The arguments for CTC loss are same in both PyTorch and TensorFlow.\n- Usually, the input is structured with a shape of (Batch-size, N_Frames, N_Landmarks).\n- Any approach can work, but typically one would use 1D convolutions, RNNs.\n-  model output shape is (Batch-size, N-Frames, N-features). it is classified into 'N-characters + 1' categories. so final shape is (Batch-size, N-Frame, N-characters+1). The extra category accounts for the blank frames, which are  meaningless gesture that occurs between the hand gestures representing A and B when expressing 'AB'.\n- During the process of creating input batches, padding is added to the input sequences and labels. Therefore, CTC loss requires the length of each sequence and label before padding within the batch.\n- Anyway, you can input a tensor of shape (Batch-size, N-Frames, N-Landmarks) into the model and obtain a tensor of shape (Batch-size, N-Frames, N-characters+1). Then, you can provide the output tensor, along with the labels and the lengths of the sequences and labels before padding, to the CTC loss.\n- During the inference process, the model predicts one character for each frame of the sequence. There may be cases where the model predicts a blank as well. For example, the predicted string could be 'f u n **blank** c **blank** t i o n'.\n- In the evaluation section of this competition, the code `prediction_str = \"\".join([rev_character_map.get(s, \"\") for s in np.argmax(output[REQUIRED_OUTPUT], axis=1)])` is used to create the output string. Since the `blank` symbol is not included in the `rev_character_map`, it is ignored, and only meaningful characters like 'function' are retained in the final predicted string.",
      "votes": 6,
      "replies": [
        {
          "id": 2318259,
          "postDate": "2023-06-26T09:11:40.683Z",
          "content": "<p>Hello <a href=\"https://www.kaggle.com/canlion\" target=\"_blank\">@canlion</a> just out of curiosity, do you currently use the ctc? were you able to make a corect submission with a ctc-trained model? </p>\n<p>Best regards !</p>",
          "rawMarkdown": "Hello @canlion just out of curiosity, do you currently use the ctc? were you able to make a corect submission with a ctc-trained model? \n\nBest regards !",
          "votes": 1,
          "replies": [
            {
              "id": 2318732,
              "postDate": "2023-06-26T14:43:18.290Z",
              "content": "<p>Hello. Currently, I am using a Transformer, but at the time I participated in the competition, I used CTC loss. I passed the submission process with a model that used CTC loss and received score. The trained model takes inputs of shape (1, N frames, N landmarks) and produces outputs of shape (1, N frames, N characters + 1).</p>",
              "rawMarkdown": "Hello. Currently, I am using a Transformer, but at the time I participated in the competition, I used CTC loss. I passed the submission process with a model that used CTC loss and received score. The trained model takes inputs of shape (1, N frames, N landmarks) and produces outputs of shape (1, N frames, N characters + 1).",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2316989,
      "postDate": "2023-06-25T11:21:12.880Z",
      "content": "<p>I know there are multiple threads and even some explanations on this topic but I am not able to implement (errors): any of the proposed methods like CTC loss. How we should process output string and how to design last layer of the model and suitable loss function?</p>",
      "rawMarkdown": "I know there are multiple threads and even some explanations on this topic but I am not able to implement (errors): any of the proposed methods like CTC loss. How we should process output string and how to design last layer of the model and suitable loss function?"
    }
  ],
  "comments": [
    {
      "id": 2317085,
      "author_name": "canlion",
      "author_url": "",
      "post_date": "2023-06-25T12:37:14.717000",
      "content": "<p>I'm not sure what specific problem you're facing, so I'll describe how I'm trying to approach the problem and note that I'm using PyTorch.</p>\n<p>Regarding CTC loss:</p>\n<ul>\n<li>The arguments for CTC loss are same in both PyTorch and TensorFlow.</li>\n<li>Usually, the input is structured with a shape of (Batch-size, N_Frames, N_Landmarks).</li>\n<li>Any approach can work, but typically one would use 1D convolutions, RNNs.</li>\n<li>model output shape is (Batch-size, N-Frames, N-features). it is classified into 'N-characters + 1' categories. so final shape is (Batch-size, N-Frame, N-characters+1). The extra category accounts for the blank frames, which are  meaningless gesture that occurs between the hand gestures representing A and B when expressing 'AB'.</li>\n<li>During the process of creating input batches, padding is added to the input sequences and labels. Therefore, CTC loss requires the length of each sequence and label before padding within the batch.</li>\n<li>Anyway, you can input a tensor of shape (Batch-size, N-Frames, N-Landmarks) into the model and obtain a tensor of shape (Batch-size, N-Frames, N-characters+1). Then, you can provide the output tensor, along with the labels and the lengths of the sequences and labels before padding, to the CTC loss.</li>\n<li>During the inference process, the model predicts one character for each frame of the sequence. There may be cases where the model predicts a blank as well. For example, the predicted string could be 'f u n <strong>blank</strong> c <strong>blank</strong> t i o n'.</li>\n<li>In the evaluation section of this competition, the code <code>prediction_str = \"\".join([rev_character_map.get(s, \"\") for s in np.argmax(output[REQUIRED_OUTPUT], axis=1)])</code> is used to create the output string. Since the <code>blank</code> symbol is not included in the <code>rev_character_map</code>, it is ignored, and only meaningful characters like 'function' are retained in the final predicted string.</li>\n</ul>",
      "votes": 6,
      "replies": [
        {
          "id": 2318259,
          "author_name": "model.fit",
          "author_url": "",
          "post_date": "2023-06-26T09:11:40.683000",
          "content": "<p>Hello <a href=\"https://www.kaggle.com/canlion\" target=\"_blank\">@canlion</a> just out of curiosity, do you currently use the ctc? were you able to make a corect submission with a ctc-trained model? </p>\n<p>Best regards !</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2318732,
              "author_name": "canlion",
              "author_url": "",
              "post_date": "2023-06-26T14:43:18.290000",
              "content": "<p>Hello. Currently, I am using a Transformer, but at the time I participated in the competition, I used CTC loss. I passed the submission process with a model that used CTC loss and received score. The trained model takes inputs of shape (1, N frames, N landmarks) and produces outputs of shape (1, N frames, N characters + 1).</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2317085": "I'm not sure what specific problem you're facing, so I'll describe how I'm trying to approach the problem and note that I'm using PyTorch.\n\nRegarding CTC loss:\n\n- The arguments for CTC loss are same in both PyTorch and TensorFlow.\n- Usually, the input is structured with a shape of (Batch-size, N_Frames, N_Landmarks).\n- Any approach can work, but typically one would use 1D convolutions, RNNs.\n-  model output shape is (Batch-size, N-Frames, N-features). it is classified into 'N-characters + 1' categories. so final shape is (Batch-size, N-Frame, N-characters+1). The extra category accounts for the blank frames, which are  meaningless gesture that occurs between the hand gestures representing A and B when expressing 'AB'.\n- During the process of creating input batches, padding is added to the input sequences and labels. Therefore, CTC loss requires the length of each sequence and label before padding within the batch.\n- Anyway, you can input a tensor of shape (Batch-size, N-Frames, N-Landmarks) into the model and obtain a tensor of shape (Batch-size, N-Frames, N-characters+1). Then, you can provide the output tensor, along with the labels and the lengths of the sequences and labels before padding, to the CTC loss.\n- During the inference process, the model predicts one character for each frame of the sequence. There may be cases where the model predicts a blank as well. For example, the predicted string could be 'f u n **blank** c **blank** t i o n'.\n- In the evaluation section of this competition, the code `prediction_str = \"\".join([rev_character_map.get(s, \"\") for s in np.argmax(output[REQUIRED_OUTPUT], axis=1)])` is used to create the output string. Since the `blank` symbol is not included in the `rev_character_map`, it is ignored, and only meaningful characters like 'function' are retained in the final predicted string.",
    "2316989": "I know there are multiple threads and even some explanations on this topic but I am not able to implement (errors): any of the proposed methods like CTC loss. How we should process output string and how to design last layer of the model and suitable loss function?"
  }
}