{
  "id": 426504,
  "title": "A CTC implementation that work on Kaggle TPU",
  "url": "/competitions/asl-fingerspelling/discussion/426504",
  "author_name": "greySnow",
  "post_date": "2023-07-23T20:23:08.056000",
  "votes": 37,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I modified another implementation of CTC and successfully made it kaggle's TPU-compatible.<br>\nThe original implementation <a href=\"https://github.com/alexeytochin/tf_seq2seq_losses\" target=\"_blank\">is from here</a>.<br>\nMy modified TPU-compatible version <a href=\"https://www.kaggle.com/datasets/shlomoron/ctc-tpu\" target=\"_blank\">is here</a>.<br>\nI modified <a href=\"https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place\" target=\"_blank\">ROHITH INGILELA's CTC notebook</a> to show how to run it on TPU. My modified version <a href=\"https://www.kaggle.com/code/shlomoron/aslfr-ctc-on-tpu\" target=\"_blank\">is here</a>.</p>\n<p>The original notebook runs in ~6 hours; mine runs in ~47 minutes.<br>\nEnjoy, and as always- votes would be appreciated :)</p>",
  "messages": [
    {
      "id": 2356040,
      "postDate": "2023-07-23T20:23:08.057Z",
      "content": "<p>I modified another implementation of CTC and successfully made it kaggle's TPU-compatible.<br>\nThe original implementation <a href=\"https://github.com/alexeytochin/tf_seq2seq_losses\" target=\"_blank\">is from here</a>.<br>\nMy modified TPU-compatible version <a href=\"https://www.kaggle.com/datasets/shlomoron/ctc-tpu\" target=\"_blank\">is here</a>.<br>\nI modified <a href=\"https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place\" target=\"_blank\">ROHITH INGILELA's CTC notebook</a> to show how to run it on TPU. My modified version <a href=\"https://www.kaggle.com/code/shlomoron/aslfr-ctc-on-tpu\" target=\"_blank\">is here</a>.</p>\n<p>The original notebook runs in ~6 hours; mine runs in ~47 minutes.<br>\nEnjoy, and as always- votes would be appreciated :)</p>",
      "rawMarkdown": "I modified another implementation of CTC and successfully made it kaggle's TPU-compatible.\nThe original implementation [is from here](https://github.com/alexeytochin/tf_seq2seq_losses).\nMy modified TPU-compatible version [is here](https://www.kaggle.com/datasets/shlomoron/ctc-tpu).\nI modified [ROHITH INGILELA's CTC notebook](https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place) to show how to run it on TPU. My modified version [is here](https://www.kaggle.com/code/shlomoron/aslfr-ctc-on-tpu).\n\nThe original notebook runs in ~6 hours; mine runs in ~47 minutes.\nEnjoy, and as always- votes would be appreciated :)",
      "votes": 37
    },
    {
      "id": 2358055,
      "postDate": "2023-07-25T09:19:26.677Z",
      "content": "<p>Maybe I have made some mistakes… But the same model of mine trained on GPU &amp; TPU perform differently, up to 0.04 in score. Wiki tells me that TPU only offer lower precision calculation, this may influence the final performance.</p>",
      "rawMarkdown": "Maybe I have made some mistakes... But the same model of mine trained on GPU & TPU perform differently, up to 0.04 in score. Wiki tells me that TPU only offer lower precision calculation, this may influence the final performance.",
      "votes": 1,
      "replies": [
        {
          "id": 2358165,
          "postDate": "2023-07-25T10:31:58.243Z",
          "content": "<p>Did you use the exact same notebook and the same data and only change TPU to gpu in the console? I.e. If I run my TPU notebook on GPU it will do 0.04 better? If so I will take another look, but is sounds strange. If you refer to a difference from the original score of Rohith's notebook, take a look at the comment section of my notebook and make sure you use version 2 of Rohith's dataset, and not version 6.</p>",
          "rawMarkdown": "Did you use the exact same notebook and the same data and only change TPU to gpu in the console? I.e. If I run my TPU notebook on GPU it will do 0.04 better? If so I will take another look, but is sounds strange. If you refer to a difference from the original score of Rohith's notebook, take a look at the comment section of my notebook and make sure you use version 2 of Rohith's dataset, and not version 6."
        },
        {
          "id": 2374861,
          "postDate": "2023-08-05T08:58:11.677Z",
          "content": "<p>Have you solved your problem? I have also encountered a similar problem where the same parameters are significantly lower on TPU by more points.</p>",
          "rawMarkdown": "Have you solved your problem? I have also encountered a similar problem where the same parameters are significantly lower on TPU by more points.",
          "replies": [
            {
              "id": 2375096,
              "postDate": "2023-08-05T11:52:42.383Z",
              "content": "<p>He Li had not answered, so I had not looked into this, but since you encountered this problem with the same parameters, I will check it. I hope to have some answers by tomorrow.</p>",
              "rawMarkdown": "He Li had not answered, so I had not looked into this, but since you encountered this problem with the same parameters, I will check it. I hope to have some answers by tomorrow."
            },
            {
              "id": 2375772,
              "postDate": "2023-08-05T22:05:17.760Z",
              "content": "<pre><code>/usr/local/lib/python3/dist-packages/keras/utils/generic_utils.py  update(self, current, values, finalize)\n                         )\n                          avg &gt; :\n--&gt;                          info += \n                         :\n                             info += \n\nValueError: Unknown  code    of  \n</code></pre>\n<p>When I used tf.nn.ctc_loss and bfloat16 in Colab, the issue mentioned above occurred, and I couldn't find a solution for it. Although removing bfloat16 resolved the problem, using it might be necessary to achieve a higher score. In addition to the previously mentioned issue, when I trained the same code without bfloat16 on Kaggle's v3-8, I encountered a problem related to nan loss.</p>",
              "rawMarkdown": "```python\n/usr/local/lib/python3.10/dist-packages/keras/utils/generic_utils.py in update(self, current, values, finalize)\n    308                     )\n    309                     if avg > 1e-3:\n--> 310                         info += f\" {avg:.4f}\"\n    311                     else:\n    312                         info += f\" {avg:.4e}\"\n\nValueError: Unknown format code 'f' for object of type 'str'\n```\n\nWhen I used tf.nn.ctc_loss and bfloat16 in Colab, the issue mentioned above occurred, and I couldn't find a solution for it. Although removing bfloat16 resolved the problem, using it might be necessary to achieve a higher score. In addition to the previously mentioned issue, when I trained the same code without bfloat16 on Kaggle's v3-8, I encountered a problem related to nan loss."
            },
            {
              "id": 2375824,
              "postDate": "2023-08-05T23:46:51.543Z",
              "content": "<p>Ok, I have some answers. After incorporating the latest updates of <a href=\"https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place\" target=\"_blank\">ROHITH's notebook</a>, my <a href=\"https://www.kaggle.com/code/shlomoron/aslfr-ctc-on-tpu\" target=\"_blank\">CTC on TPU</a> could not drive the training score as low as Rohith's (3.9321 vs. 3.1931), <a href=\"https://www.kaggle.com/code/shlomoron/aslfr-ctc-on-tpu-submission\" target=\"_blank\">but achieved a better score by one point on the LB</a> (0.688  vs. 0.687)</p>",
              "rawMarkdown": "Ok, I have some answers. After incorporating the latest updates of [ROHITH's notebook](https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place), my [CTC on TPU](https://www.kaggle.com/code/shlomoron/aslfr-ctc-on-tpu) could not drive the training score as low as Rohith's (3.9321 vs. 3.1931), [but achieved a better score by one point on the LB](https://www.kaggle.com/code/shlomoron/aslfr-ctc-on-tpu-submission) (0.688  vs. 0.687)"
            },
            {
              "id": 2375889,
              "postDate": "2023-08-06T02:27:46.420Z",
              "content": "<p>I can't exceed 0.7 points with TPU, it's too difficult to debug.</p>",
              "rawMarkdown": "I can't exceed 0.7 points with TPU, it's too difficult to debug."
            }
          ]
        }
      ]
    },
    {
      "id": 2356239,
      "postDate": "2023-07-24T03:57:06.930Z",
      "content": "<p>Amazing! This should be merged into TensorFlow. <br>\nI opened an issue here previously: <a href=\"https://github.com/tensorflow/tensorflow/issues/61297#issuecomment-1648022939\" target=\"_blank\">https://github.com/tensorflow/tensorflow/issues/61297#issuecomment-1648022939</a>.<br>\nWould you like to open a PR to tensorflow to try to add it? <a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a> </p>",
      "rawMarkdown": "Amazing! This should be merged into TensorFlow. \nI opened an issue here previously: https://github.com/tensorflow/tensorflow/issues/61297#issuecomment-1648022939.\nWould you like to open a PR to tensorflow to try to add it? @shlomoron ",
      "votes": 1,
      "replies": [
        {
          "id": 2360316,
          "postDate": "2023-07-26T17:10:41.237Z",
          "content": "<p>It would probably be redundant, as TensorFlow already has a working version. It just doesn't work on Kaggle, haha. They need to figure out the problem (well, I have a hunch that the problem is on Kaggle's side- something in the environment, probably. The TensorFlow version actually does not fail only on TPU but also on GPU in XLA mode (jit_compile = True), and although I did not check, I guess that it WILL work in, say, colab, considering that it works there on TPU.</p>",
          "rawMarkdown": "It would probably be redundant, as TensorFlow already has a working version. It just doesn't work on Kaggle, haha. They need to figure out the problem (well, I have a hunch that the problem is on Kaggle's side- something in the environment, probably. The TensorFlow version actually does not fail only on TPU but also on GPU in XLA mode (jit_compile = True), and although I did not check, I guess that it WILL work in, say, colab, considering that it works there on TPU."
        }
      ]
    },
    {
      "id": 2374282,
      "postDate": "2023-08-04T20:14:35.100Z",
      "content": "<p>I am struggling to use the CTC loss with sparse labels (on GPU). The computational speed for that version is from 2 to 10x faster than using dense labels. However, my training always diverge.. I found other people having the same issues online but couldn't find a solution. Does any1 have any trick to solve this?</p>",
      "rawMarkdown": "I am struggling to use the CTC loss with sparse labels (on GPU). The computational speed for that version is from 2 to 10x faster than using dense labels. However, my training always diverge.. I found other people having the same issues online but couldn't find a solution. Does any1 have any trick to solve this?"
    },
    {
      "id": 2356046,
      "postDate": "2023-07-23T20:29:43.710Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2358055,
      "author_name": "He Li",
      "author_url": "",
      "post_date": "2023-07-25T09:19:26.677000",
      "content": "<p>Maybe I have made some mistakes… But the same model of mine trained on GPU &amp; TPU perform differently, up to 0.04 in score. Wiki tells me that TPU only offer lower precision calculation, this may influence the final performance.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2358165,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2023-07-25T10:31:58.243000",
          "content": "<p>Did you use the exact same notebook and the same data and only change TPU to gpu in the console? I.e. If I run my TPU notebook on GPU it will do 0.04 better? If so I will take another look, but is sounds strange. If you refer to a difference from the original score of Rohith's notebook, take a look at the comment section of my notebook and make sure you use version 2 of Rohith's dataset, and not version 6.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2374861,
          "author_name": "Scenery SunFireInk",
          "author_url": "",
          "post_date": "2023-08-05T08:58:11.677000",
          "content": "<p>Have you solved your problem? I have also encountered a similar problem where the same parameters are significantly lower on TPU by more points.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2375096,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2023-08-05T11:52:42.383000",
              "content": "<p>He Li had not answered, so I had not looked into this, but since you encountered this problem with the same parameters, I will check it. I hope to have some answers by tomorrow.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2375772,
              "author_name": "Scenery SunFireInk",
              "author_url": "",
              "post_date": "2023-08-05T22:05:17.760000",
              "content": "<pre><code>/usr/local/lib/python3/dist-packages/keras/utils/generic_utils.py  update(self, current, values, finalize)\n                         )\n                          avg &gt; :\n--&gt;                          info += \n                         :\n                             info += \n\nValueError: Unknown  code    of  \n</code></pre>\n<p>When I used tf.nn.ctc_loss and bfloat16 in Colab, the issue mentioned above occurred, and I couldn't find a solution for it. Although removing bfloat16 resolved the problem, using it might be necessary to achieve a higher score. In addition to the previously mentioned issue, when I trained the same code without bfloat16 on Kaggle's v3-8, I encountered a problem related to nan loss.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2375824,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2023-08-05T23:46:51.543000",
              "content": "<p>Ok, I have some answers. After incorporating the latest updates of <a href=\"https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place\" target=\"_blank\">ROHITH's notebook</a>, my <a href=\"https://www.kaggle.com/code/shlomoron/aslfr-ctc-on-tpu\" target=\"_blank\">CTC on TPU</a> could not drive the training score as low as Rohith's (3.9321 vs. 3.1931), <a href=\"https://www.kaggle.com/code/shlomoron/aslfr-ctc-on-tpu-submission\" target=\"_blank\">but achieved a better score by one point on the LB</a> (0.688  vs. 0.687)</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2375889,
              "author_name": "Scenery SunFireInk",
              "author_url": "",
              "post_date": "2023-08-06T02:27:46.420000",
              "content": "<p>I can't exceed 0.7 points with TPU, it's too difficult to debug.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2356239,
      "author_name": "WalkingMoose",
      "author_url": "",
      "post_date": "2023-07-24T03:57:06.930000",
      "content": "<p>Amazing! This should be merged into TensorFlow. <br>\nI opened an issue here previously: <a href=\"https://github.com/tensorflow/tensorflow/issues/61297#issuecomment-1648022939\" target=\"_blank\">https://github.com/tensorflow/tensorflow/issues/61297#issuecomment-1648022939</a>.<br>\nWould you like to open a PR to tensorflow to try to add it? <a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a> </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2360316,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2023-07-26T17:10:41.237000",
          "content": "<p>It would probably be redundant, as TensorFlow already has a working version. It just doesn't work on Kaggle, haha. They need to figure out the problem (well, I have a hunch that the problem is on Kaggle's side- something in the environment, probably. The TensorFlow version actually does not fail only on TPU but also on GPU in XLA mode (jit_compile = True), and although I did not check, I guess that it WILL work in, say, colab, considering that it works there on TPU.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2374282,
      "author_name": "Adriano Passos",
      "author_url": "",
      "post_date": "2023-08-04T20:14:35.100000",
      "content": "<p>I am struggling to use the CTC loss with sparse labels (on GPU). The computational speed for that version is from 2 to 10x faster than using dense labels. However, my training always diverge.. I found other people having the same issues online but couldn't find a solution. Does any1 have any trick to solve this?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2356046,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-07-23T20:29:43.710000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2356040": "I modified another implementation of CTC and successfully made it kaggle's TPU-compatible.\nThe original implementation [is from here](https://github.com/alexeytochin/tf_seq2seq_losses).\nMy modified TPU-compatible version [is here](https://www.kaggle.com/datasets/shlomoron/ctc-tpu).\nI modified [ROHITH INGILELA's CTC notebook](https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place) to show how to run it on TPU. My modified version [is here](https://www.kaggle.com/code/shlomoron/aslfr-ctc-on-tpu).\n\nThe original notebook runs in ~6 hours; mine runs in ~47 minutes.\nEnjoy, and as always- votes would be appreciated :)",
    "2358055": "Maybe I have made some mistakes... But the same model of mine trained on GPU & TPU perform differently, up to 0.04 in score. Wiki tells me that TPU only offer lower precision calculation, this may influence the final performance.",
    "2356239": "Amazing! This should be merged into TensorFlow. \nI opened an issue here previously: https://github.com/tensorflow/tensorflow/issues/61297#issuecomment-1648022939.\nWould you like to open a PR to tensorflow to try to add it? @shlomoron ",
    "2374282": "I am struggling to use the CTC loss with sparse labels (on GPU). The computational speed for that version is from 2 to 10x faster than using dense labels. However, my training always diverge.. I found other people having the same issues online but couldn't find a solution. Does any1 have any trick to solve this?",
    "2356046": ""
  }
}