{
  "id": 409512,
  "title": "Sequence modeling with CTC loss",
  "url": "/competitions/asl-fingerspelling/discussion/409512",
  "author_name": "Mykola",
  "post_date": "2023-05-11T11:39:35.360000",
  "votes": 23,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi there!</p>\n<p>In the previous competition, we were asked to predict a single word (a.k.a classification task). In contrast, here we need to predict a sequence of classes (i.e. characters). The naive approach may be to split a sequence of landmarks into groups and classify each group (using some character classification similar to what we used in the previous competition). However, we should deal with variable-length phrases and we do not know what number of characters to expect. On top of that, splitting a sequence into segments is quite a hard task and generally cannot be easily done in a simple way.</p>\n<p>Here CTC loss goes. Connectionist temporal classification loss is intended to train a model with a variable length of the output. The key idea of CTC loss in a nutshell: we predict a matrix of shape NxT, where N is a number of possible characters + 1 special reserved character. T - maximal possible length of the predicted sequence.<br>\nEach column T_i represents N probabilities of characters predicted per frame i. In simple words, we just predict a character per frame. <br>\nThe magic happens in the decoding step. We decode the matrix to string in such a way, that duplicated characters are squeezed into a single character (i.e. all the duplication is removed). You may wonder, how could we predict a double L in a \"nutshell\" word, for instance. Here a special reserved character (let denote it '#') plays a role. Any duplicated characters needed to be duplicated should be separated by '#':</p>\n<p>aaabbb -&gt; ab<br>\na#abb#bb -&gt; ab</p>\n<p>Previously we were talking only about inference. During inference, we just take an argmax character per each frame. During the training phase, the loss is calculated as a sum of true path probabilities divided by probabilities of all the possible paths. It is quite a hard task to compute all possible pathes but fortunately, all the frameworks have optimized versions of CTC loss, so it works fast.</p>\n<p>Here is the <a href=\"https://distill.pub/2017/ctc/\" target=\"_blank\">best CTC tutorial I have ever seen</a>, hope you find it helpful as well. </p>\n<p>I believe that combining some sort of CNN networks + RNN networks with CTC loss on top of it (a.k.a CRNN) may lead to good results.</p>\n<p>Would love to hear your thoughts on this approach or discuss any other promising ideas!</p>",
  "messages": [
    {
      "id": 2254967,
      "postDate": "2023-05-11T11:39:35.360Z",
      "content": "<p>Hi there!</p>\n<p>In the previous competition, we were asked to predict a single word (a.k.a classification task). In contrast, here we need to predict a sequence of classes (i.e. characters). The naive approach may be to split a sequence of landmarks into groups and classify each group (using some character classification similar to what we used in the previous competition). However, we should deal with variable-length phrases and we do not know what number of characters to expect. On top of that, splitting a sequence into segments is quite a hard task and generally cannot be easily done in a simple way.</p>\n<p>Here CTC loss goes. Connectionist temporal classification loss is intended to train a model with a variable length of the output. The key idea of CTC loss in a nutshell: we predict a matrix of shape NxT, where N is a number of possible characters + 1 special reserved character. T - maximal possible length of the predicted sequence.<br>\nEach column T_i represents N probabilities of characters predicted per frame i. In simple words, we just predict a character per frame. <br>\nThe magic happens in the decoding step. We decode the matrix to string in such a way, that duplicated characters are squeezed into a single character (i.e. all the duplication is removed). You may wonder, how could we predict a double L in a \"nutshell\" word, for instance. Here a special reserved character (let denote it '#') plays a role. Any duplicated characters needed to be duplicated should be separated by '#':</p>\n<p>aaabbb -&gt; ab<br>\na#abb#bb -&gt; ab</p>\n<p>Previously we were talking only about inference. During inference, we just take an argmax character per each frame. During the training phase, the loss is calculated as a sum of true path probabilities divided by probabilities of all the possible paths. It is quite a hard task to compute all possible pathes but fortunately, all the frameworks have optimized versions of CTC loss, so it works fast.</p>\n<p>Here is the <a href=\"https://distill.pub/2017/ctc/\" target=\"_blank\">best CTC tutorial I have ever seen</a>, hope you find it helpful as well. </p>\n<p>I believe that combining some sort of CNN networks + RNN networks with CTC loss on top of it (a.k.a CRNN) may lead to good results.</p>\n<p>Would love to hear your thoughts on this approach or discuss any other promising ideas!</p>",
      "rawMarkdown": "Hi there!\n\nIn the previous competition, we were asked to predict a single word (a.k.a classification task). In contrast, here we need to predict a sequence of classes (i.e. characters). The naive approach may be to split a sequence of landmarks into groups and classify each group (using some character classification similar to what we used in the previous competition). However, we should deal with variable-length phrases and we do not know what number of characters to expect. On top of that, splitting a sequence into segments is quite a hard task and generally cannot be easily done in a simple way.\n\nHere CTC loss goes. Connectionist temporal classification loss is intended to train a model with a variable length of the output. The key idea of CTC loss in a nutshell: we predict a matrix of shape NxT, where N is a number of possible characters + 1 special reserved character. T - maximal possible length of the predicted sequence.\nEach column T_i represents N probabilities of characters predicted per frame i. In simple words, we just predict a character per frame. \nThe magic happens in the decoding step. We decode the matrix to string in such a way, that duplicated characters are squeezed into a single character (i.e. all the duplication is removed). You may wonder, how could we predict a double L in a \"nutshell\" word, for instance. Here a special reserved character (let denote it '#') plays a role. Any duplicated characters needed to be duplicated should be separated by '#':\n\naaabbb -> ab\na#abb#bb -> ab\n\n\nPreviously we were talking only about inference. During inference, we just take an argmax character per each frame. During the training phase, the loss is calculated as a sum of true path probabilities divided by probabilities of all the possible paths. It is quite a hard task to compute all possible pathes but fortunately, all the frameworks have optimized versions of CTC loss, so it works fast.\n\nHere is the [best CTC tutorial I have ever seen](https://distill.pub/2017/ctc/), hope you find it helpful as well. \n\nI believe that combining some sort of CNN networks + RNN networks with CTC loss on top of it (a.k.a CRNN) may lead to good results.\n\nWould love to hear your thoughts on this approach or discuss any other promising ideas!",
      "votes": 22
    },
    {
      "id": 2313681,
      "postDate": "2023-06-22T21:02:11.550Z",
      "content": "<p>Is it possible to train NN with CTC loss on TPU? Because I tried both tf.nn.ctc_loss and keras.backend.ctc_batch_cost but they give error \"XLA compilation requires that operator arguments that represent shapes or dimensions be evaluated to concrete values at compile time. This error means that a shape or dimension argument could not be evaluated at compile time, usually because the value of the argument depends on a parameter to the computation, on a variable, or on a stateful operation such as a random number generator.<br>\n2023-06-22 19:31:45.984152: F tensorflow/core/tpu/kernels/tpu_program_group.cc:86] Check failed: xla_tpu_programs.size() &gt; 0 (0 vs. 0)\"</p>",
      "rawMarkdown": "Is it possible to train NN with CTC loss on TPU? Because I tried both tf.nn.ctc_loss and keras.backend.ctc_batch_cost but they give error \"XLA compilation requires that operator arguments that represent shapes or dimensions be evaluated to concrete values at compile time. This error means that a shape or dimension argument could not be evaluated at compile time, usually because the value of the argument depends on a parameter to the computation, on a variable, or on a stateful operation such as a random number generator.\n2023-06-22 19:31:45.984152: F tensorflow/core/tpu/kernels/tpu_program_group.cc:86] Check failed: xla_tpu_programs.size() > 0 (0 vs. 0)\"",
      "votes": 1,
      "replies": [
        {
          "id": 2314316,
          "postDate": "2023-06-23T10:40:18.843Z",
          "content": "<p>I believe it is possible. Technically, in CTC you produce the static shaped output (matrix of shape MxN, where M is a number of symbols+1 and N - max characters to predict, so you could predict up to N characters, but anyway model produces NxM shape)</p>",
          "rawMarkdown": "I believe it is possible. Technically, in CTC you produce the static shaped output (matrix of shape MxN, where M is a number of symbols+1 and N - max characters to predict, so you could predict up to N characters, but anyway model produces NxM shape)",
          "votes": 1
        },
        {
          "id": 2346764,
          "postDate": "2023-07-16T13:54:06.983Z",
          "content": "<p>I had the same problem trying <code>tf.nn.ctc_loss</code> . Have you figure out this problem?</p>",
          "rawMarkdown": "I had the same problem trying `tf.nn.ctc_loss` . Have you figure out this problem?",
          "replies": [
            {
              "id": 2346834,
              "postDate": "2023-07-16T14:59:37.410Z",
              "content": "<p>Kinda, \"XLA compilation requires that operator arguments that represent shapes or dimensions be evaluated to concrete values at compile time\" so in CTCLoss class I needed to x = tf.ensure_shape(x, …) of passed y_true and y_predicted with corresponding shapes, then model started training but loss was Nan :(. Then with absolutely the same code, I swithed to Colab TPU and it worked perfectly fine)</p>",
              "rawMarkdown": "Kinda, \"XLA compilation requires that operator arguments that represent shapes or dimensions be evaluated to concrete values at compile time\" so in CTCLoss class I needed to x = tf.ensure_shape(x, ...) of passed y_true and y_predicted with corresponding shapes, then model started training but loss was Nan :(. Then with absolutely the same code, I swithed to Colab TPU and it worked perfectly fine)"
            }
          ]
        }
      ]
    },
    {
      "id": 2398791,
      "postDate": "2023-08-20T00:27:50.650Z",
      "content": "<p>Thanks for sharing CTC. I think there is a typo in your post</p>\n<pre><code>aaabbb -&gt; ab\na#abb#bb -&gt; ab\n</code></pre>\n<p>should say</p>\n<pre><code>aaabbb -&gt; ab\na#abb#bb -&gt; aabb\n</code></pre>",
      "rawMarkdown": "Thanks for sharing CTC. I think there is a typo in your post\n\n    aaabbb -> ab\n    a#abb#bb -> ab\n\nshould say\n\n    aaabbb -> ab\n    a#abb#bb -> aabb"
    }
  ],
  "comments": [
    {
      "id": 2313681,
      "author_name": "BogdanNet",
      "author_url": "",
      "post_date": "2023-06-22T21:02:11.550000",
      "content": "<p>Is it possible to train NN with CTC loss on TPU? Because I tried both tf.nn.ctc_loss and keras.backend.ctc_batch_cost but they give error \"XLA compilation requires that operator arguments that represent shapes or dimensions be evaluated to concrete values at compile time. This error means that a shape or dimension argument could not be evaluated at compile time, usually because the value of the argument depends on a parameter to the computation, on a variable, or on a stateful operation such as a random number generator.<br>\n2023-06-22 19:31:45.984152: F tensorflow/core/tpu/kernels/tpu_program_group.cc:86] Check failed: xla_tpu_programs.size() &gt; 0 (0 vs. 0)\"</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2314316,
          "author_name": "Mykola",
          "author_url": "",
          "post_date": "2023-06-23T10:40:18.843000",
          "content": "<p>I believe it is possible. Technically, in CTC you produce the static shaped output (matrix of shape MxN, where M is a number of symbols+1 and N - max characters to predict, so you could predict up to N characters, but anyway model produces NxM shape)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2346764,
          "author_name": "biubiubiu~",
          "author_url": "",
          "post_date": "2023-07-16T13:54:06.983000",
          "content": "<p>I had the same problem trying <code>tf.nn.ctc_loss</code> . Have you figure out this problem?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2346834,
              "author_name": "BogdanNet",
              "author_url": "",
              "post_date": "2023-07-16T14:59:37.410000",
              "content": "<p>Kinda, \"XLA compilation requires that operator arguments that represent shapes or dimensions be evaluated to concrete values at compile time\" so in CTCLoss class I needed to x = tf.ensure_shape(x, …) of passed y_true and y_predicted with corresponding shapes, then model started training but loss was Nan :(. Then with absolutely the same code, I swithed to Colab TPU and it worked perfectly fine)</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2398791,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2023-08-20T00:27:50.650000",
      "content": "<p>Thanks for sharing CTC. I think there is a typo in your post</p>\n<pre><code>aaabbb -&gt; ab\na#abb#bb -&gt; ab\n</code></pre>\n<p>should say</p>\n<pre><code>aaabbb -&gt; ab\na#abb#bb -&gt; aabb\n</code></pre>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2254967": "Hi there!\n\nIn the previous competition, we were asked to predict a single word (a.k.a classification task). In contrast, here we need to predict a sequence of classes (i.e. characters). The naive approach may be to split a sequence of landmarks into groups and classify each group (using some character classification similar to what we used in the previous competition). However, we should deal with variable-length phrases and we do not know what number of characters to expect. On top of that, splitting a sequence into segments is quite a hard task and generally cannot be easily done in a simple way.\n\nHere CTC loss goes. Connectionist temporal classification loss is intended to train a model with a variable length of the output. The key idea of CTC loss in a nutshell: we predict a matrix of shape NxT, where N is a number of possible characters + 1 special reserved character. T - maximal possible length of the predicted sequence.\nEach column T_i represents N probabilities of characters predicted per frame i. In simple words, we just predict a character per frame. \nThe magic happens in the decoding step. We decode the matrix to string in such a way, that duplicated characters are squeezed into a single character (i.e. all the duplication is removed). You may wonder, how could we predict a double L in a \"nutshell\" word, for instance. Here a special reserved character (let denote it '#') plays a role. Any duplicated characters needed to be duplicated should be separated by '#':\n\naaabbb -> ab\na#abb#bb -> ab\n\n\nPreviously we were talking only about inference. During inference, we just take an argmax character per each frame. During the training phase, the loss is calculated as a sum of true path probabilities divided by probabilities of all the possible paths. It is quite a hard task to compute all possible pathes but fortunately, all the frameworks have optimized versions of CTC loss, so it works fast.\n\nHere is the [best CTC tutorial I have ever seen](https://distill.pub/2017/ctc/), hope you find it helpful as well. \n\nI believe that combining some sort of CNN networks + RNN networks with CTC loss on top of it (a.k.a CRNN) may lead to good results.\n\nWould love to hear your thoughts on this approach or discuss any other promising ideas!",
    "2313681": "Is it possible to train NN with CTC loss on TPU? Because I tried both tf.nn.ctc_loss and keras.backend.ctc_batch_cost but they give error \"XLA compilation requires that operator arguments that represent shapes or dimensions be evaluated to concrete values at compile time. This error means that a shape or dimension argument could not be evaluated at compile time, usually because the value of the argument depends on a parameter to the computation, on a variable, or on a stateful operation such as a random number generator.\n2023-06-22 19:31:45.984152: F tensorflow/core/tpu/kernels/tpu_program_group.cc:86] Check failed: xla_tpu_programs.size() > 0 (0 vs. 0)\"",
    "2398791": "Thanks for sharing CTC. I think there is a typo in your post\n\n    aaabbb -> ab\n    a#abb#bb -> ab\n\nshould say\n\n    aaabbb -> ab\n    a#abb#bb -> aabb"
  }
}