{
  "id": 418848,
  "title": "CTC Loss with TPU",
  "url": "/competitions/asl-fingerspelling/discussion/418848",
  "author_name": "BogdanNet",
  "post_date": "2023-06-22T21:18:28.854000",
  "votes": 4,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Is it possible to train NN with CTC Loss on TPU? Because I tried both tf.nn.ctc_loss and keras.backend.ctc_batch_cost and when training on GPU everything is ok but when using TPU I get the following error<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10328925%2F9388ea4fe31b1890581d500eec930451%2FScreenshot%202023-06-23%20001646.png?generation=1687468661954210&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 2313688,
      "postDate": "2023-06-22T21:18:28.853Z",
      "content": "<p>Is it possible to train NN with CTC Loss on TPU? Because I tried both tf.nn.ctc_loss and keras.backend.ctc_batch_cost and when training on GPU everything is ok but when using TPU I get the following error<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10328925%2F9388ea4fe31b1890581d500eec930451%2FScreenshot%202023-06-23%20001646.png?generation=1687468661954210&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Is it possible to train NN with CTC Loss on TPU? Because I tried both tf.nn.ctc_loss and keras.backend.ctc_batch_cost and when training on GPU everything is ok but when using TPU I get the following error![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10328925%2F9388ea4fe31b1890581d500eec930451%2FScreenshot%202023-06-23%20001646.png?generation=1687468661954210&alt=media)",
      "votes": 4
    },
    {
      "id": 2356343,
      "postDate": "2023-07-24T05:37:32.913Z",
      "content": "<p>An alternative implementation that works on Kaggle TPU has been proposed here by Greysnow:<br>\n<a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/426504\" target=\"_blank\">https://www.kaggle.com/competitions/asl-fingerspelling/discussion/426504</a></p>",
      "rawMarkdown": "An alternative implementation that works on Kaggle TPU has been proposed here by Greysnow:\nhttps://www.kaggle.com/competitions/asl-fingerspelling/discussion/426504",
      "votes": 1
    },
    {
      "id": 2328364,
      "postDate": "2023-07-03T14:32:08.800Z",
      "content": "<p>I had the same problem. The problem lies (most likely, at least it was my case) in the resampling augmentation, when you change the amount of time steps. When calculating ctc loss, TPU must know your time shape at compilation time. </p>",
      "rawMarkdown": "I had the same problem. The problem lies (most likely, at least it was my case) in the resampling augmentation, when you change the amount of time steps. When calculating ctc loss, TPU must know your time shape at compilation time. ",
      "votes": 1
    },
    {
      "id": 2347874,
      "postDate": "2023-07-17T08:58:18.563Z",
      "content": "<p>Shape problems can occur if e.g there is a partially filled last batch. TPU expects all batches to be the same size. Such issue can be fixed by using  ds.batch(drop_remainder=True) or manually trimming the dataset to exact mutiple of batches.<br>\nIt can also be debugged by adding  x=tf.ensure_shape(x,shape). Where shape is the expected shape. </p>\n<p>I'm currently able to nominally make it execute on TPU but the loss is infinite or nan.  There are grappler errors in the printout that don't terminate the execution but I suspect they may be involved in the nan/inf loss.</p>\n<p>Additionally, while training on GPU I used model.compile(jit_compile=True). It gives errors on unsupported in place add operator in ctc  and exits. (p.s: jit_compile=True is should not be used on TPU according to docs, but it is recommended to run it on GPU for debugging before switching it off and moving to TPU, as TPU also uses XLA optimization). </p>\n<p>The same code runs fine on the GPU P100 with finite loss.</p>\n<p>I created a public example notebook borrowing code from another kaggler from a past competition: <br>\n<a href=\"https://www.kaggle.com/code/shaironen/ctc-example/notebook\" target=\"_blank\">https://www.kaggle.com/code/shaironen/ctc-example/notebook</a>. It can be run either with GPU or TPU to show the above behavior. </p>",
      "rawMarkdown": "Shape problems can occur if e.g there is a partially filled last batch. TPU expects all batches to be the same size. Such issue can be fixed by using  ds.batch(drop_remainder=True) or manually trimming the dataset to exact mutiple of batches.\nIt can also be debugged by adding  x=tf.ensure_shape(x,shape). Where shape is the expected shape. \n\n\nI'm currently able to nominally make it execute on TPU but the loss is infinite or nan.  There are grappler errors in the printout that don't terminate the execution but I suspect they may be involved in the nan/inf loss.\n\nAdditionally, while training on GPU I used model.compile(jit_compile=True). It gives errors on unsupported in place add operator in ctc  and exits. (p.s: jit_compile=True is should not be used on TPU according to docs, but it is recommended to run it on GPU for debugging before switching it off and moving to TPU, as TPU also uses XLA optimization). \n\nThe same code runs fine on the GPU P100 with finite loss.\n\nI created a public example notebook borrowing code from another kaggler from a past competition: \nhttps://www.kaggle.com/code/shaironen/ctc-example/notebook. It can be run either with GPU or TPU to show the above behavior. \n",
      "votes": 2,
      "replies": [
        {
          "id": 2348156,
          "postDate": "2023-07-17T12:31:11.717Z",
          "content": "<p>yeah, the same, I figured out that if I use tf.ensure_shape in ctc loss it runs on TPU but with nan loss, interesting that absolutely the same code(with ensure_shape) runs ok in colab with finite loss</p>",
          "rawMarkdown": "yeah, the same, I figured out that if I use tf.ensure_shape in ctc loss it runs on TPU but with nan loss, interesting that absolutely the same code(with ensure_shape) runs ok in colab with finite loss",
          "replies": [
            {
              "id": 2348415,
              "postDate": "2023-07-17T16:14:34.593Z",
              "content": "<p>Hi Bogdan,  do you mean it runs ok in colab even on TPU? <br>\nI filed a tensorflow issue here: <a href=\"https://github.com/tensorflow/tensorflow/issues/61297\" target=\"_blank\">https://github.com/tensorflow/tensorflow/issues/61297</a> and we try to reproduce my kaggle demo notebook in colab.<br>\nI also noticed kaggle is now using TPU VM, (in the past used normal TPU). I'm not sure if colab is using \"normal\" TPU or TPU VM and if it might matter.</p>",
              "rawMarkdown": "Hi Bogdan,  do you mean it runs ok in colab even on TPU? \nI filed a tensorflow issue here: https://github.com/tensorflow/tensorflow/issues/61297 and we try to reproduce my kaggle demo notebook in colab.\nI also noticed kaggle is now using TPU VM, (in the past used normal TPU). I'm not sure if colab is using \"normal\" TPU or TPU VM and if it might matter."
            },
            {
              "id": 2348670,
              "postDate": "2023-07-17T20:37:31.553Z",
              "content": "<p>hi, yes, even on TPU loss is finite and NN trains and converges </p>",
              "rawMarkdown": "hi, yes, even on TPU loss is finite and NN trains and converges "
            }
          ]
        },
        {
          "id": 2348806,
          "postDate": "2023-07-18T01:59:55.703Z",
          "content": "<p>I simplified reproducing the TPU issue in the demo notebook with pseudo-random matrix data and no need for file data storage (with fixed seed), so it can easily exported to to colab.<br>\nA corresponding colab gist is here: <br>\n<a href=\"https://colab.research.google.com/gist/sronen71/9bea9743ed40a80400496b5c1c4a80d3/61297_ctc-example.ipynb\" target=\"_blank\">https://colab.research.google.com/gist/sronen71/9bea9743ed40a80400496b5c1c4a80d3/61297_ctc-example.ipynb</a><br>\nThe code gives nan in Kaggle environment TPU v3-8 with some compiler printout errors (not terminating). <br>\nThe same code runs in Colab TPU (v2) with finite loss and no error printouts.<br>\nBoth use TF==12.0.<br>\nSame behavior as reported by Bogdan.</p>",
          "rawMarkdown": "I simplified reproducing the TPU issue in the demo notebook with pseudo-random matrix data and no need for file data storage (with fixed seed), so it can easily exported to to colab.\nA corresponding colab gist is here: \nhttps://colab.research.google.com/gist/sronen71/9bea9743ed40a80400496b5c1c4a80d3/61297_ctc-example.ipynb\nThe code gives nan in Kaggle environment TPU v3-8 with some compiler printout errors (not terminating). \nThe same code runs in Colab TPU (v2) with finite loss and no error printouts.\nBoth use TF==12.0.\nSame behavior as reported by Bogdan.",
          "votes": 3
        },
        {
          "id": 2349157,
          "postDate": "2023-07-18T07:49:26.697Z",
          "content": "<p>Experienced same behavior. When TPU connection is 'local' it gives NaN and I cannot pinpoint the reason - happens both in Kaggle environment and in gcloud TPU.<br>\nOn Colab it works fine.</p>",
          "rawMarkdown": "Experienced same behavior. When TPU connection is 'local' it gives NaN and I cannot pinpoint the reason - happens both in Kaggle environment and in gcloud TPU.\nOn Colab it works fine.",
          "replies": [
            {
              "id": 2377206,
              "postDate": "2023-08-07T03:49:59.440Z",
              "content": "<p>Can a graphical interface be used for training on gcloud TPU?</p>",
              "rawMarkdown": "Can a graphical interface be used for training on gcloud TPU?"
            }
          ]
        },
        {
          "id": 2356340,
          "postDate": "2023-07-24T05:36:49.047Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 2355729,
      "postDate": "2023-07-23T14:44:47.650Z",
      "content": "<p>I also have the same error. It works on a single tpu, but it doesn't seem to be compatible with strategy…</p>",
      "rawMarkdown": "I also have the same error. It works on a single tpu, but it doesn't seem to be compatible with strategy..."
    },
    {
      "id": 2333764,
      "postDate": "2023-07-07T06:33:21.677Z",
      "content": "<p>I have exactly the same problem ……</p>",
      "rawMarkdown": "I have exactly the same problem ......"
    }
  ],
  "comments": [
    {
      "id": 2356343,
      "author_name": "WalkingMoose",
      "author_url": "",
      "post_date": "2023-07-24T05:37:32.913000",
      "content": "<p>An alternative implementation that works on Kaggle TPU has been proposed here by Greysnow:<br>\n<a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/426504\" target=\"_blank\">https://www.kaggle.com/competitions/asl-fingerspelling/discussion/426504</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2328364,
      "author_name": "Vitalii Bozheniuk",
      "author_url": "",
      "post_date": "2023-07-03T14:32:08.800000",
      "content": "<p>I had the same problem. The problem lies (most likely, at least it was my case) in the resampling augmentation, when you change the amount of time steps. When calculating ctc loss, TPU must know your time shape at compilation time. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2347874,
      "author_name": "WalkingMoose",
      "author_url": "",
      "post_date": "2023-07-17T08:58:18.563000",
      "content": "<p>Shape problems can occur if e.g there is a partially filled last batch. TPU expects all batches to be the same size. Such issue can be fixed by using  ds.batch(drop_remainder=True) or manually trimming the dataset to exact mutiple of batches.<br>\nIt can also be debugged by adding  x=tf.ensure_shape(x,shape). Where shape is the expected shape. </p>\n<p>I'm currently able to nominally make it execute on TPU but the loss is infinite or nan.  There are grappler errors in the printout that don't terminate the execution but I suspect they may be involved in the nan/inf loss.</p>\n<p>Additionally, while training on GPU I used model.compile(jit_compile=True). It gives errors on unsupported in place add operator in ctc  and exits. (p.s: jit_compile=True is should not be used on TPU according to docs, but it is recommended to run it on GPU for debugging before switching it off and moving to TPU, as TPU also uses XLA optimization). </p>\n<p>The same code runs fine on the GPU P100 with finite loss.</p>\n<p>I created a public example notebook borrowing code from another kaggler from a past competition: <br>\n<a href=\"https://www.kaggle.com/code/shaironen/ctc-example/notebook\" target=\"_blank\">https://www.kaggle.com/code/shaironen/ctc-example/notebook</a>. It can be run either with GPU or TPU to show the above behavior. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 2348156,
          "author_name": "BogdanNet",
          "author_url": "",
          "post_date": "2023-07-17T12:31:11.717000",
          "content": "<p>yeah, the same, I figured out that if I use tf.ensure_shape in ctc loss it runs on TPU but with nan loss, interesting that absolutely the same code(with ensure_shape) runs ok in colab with finite loss</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2348415,
              "author_name": "WalkingMoose",
              "author_url": "",
              "post_date": "2023-07-17T16:14:34.593000",
              "content": "<p>Hi Bogdan,  do you mean it runs ok in colab even on TPU? <br>\nI filed a tensorflow issue here: <a href=\"https://github.com/tensorflow/tensorflow/issues/61297\" target=\"_blank\">https://github.com/tensorflow/tensorflow/issues/61297</a> and we try to reproduce my kaggle demo notebook in colab.<br>\nI also noticed kaggle is now using TPU VM, (in the past used normal TPU). I'm not sure if colab is using \"normal\" TPU or TPU VM and if it might matter.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2348670,
              "author_name": "BogdanNet",
              "author_url": "",
              "post_date": "2023-07-17T20:37:31.553000",
              "content": "<p>hi, yes, even on TPU loss is finite and NN trains and converges </p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2348806,
          "author_name": "WalkingMoose",
          "author_url": "",
          "post_date": "2023-07-18T01:59:55.703000",
          "content": "<p>I simplified reproducing the TPU issue in the demo notebook with pseudo-random matrix data and no need for file data storage (with fixed seed), so it can easily exported to to colab.<br>\nA corresponding colab gist is here: <br>\n<a href=\"https://colab.research.google.com/gist/sronen71/9bea9743ed40a80400496b5c1c4a80d3/61297_ctc-example.ipynb\" target=\"_blank\">https://colab.research.google.com/gist/sronen71/9bea9743ed40a80400496b5c1c4a80d3/61297_ctc-example.ipynb</a><br>\nThe code gives nan in Kaggle environment TPU v3-8 with some compiler printout errors (not terminating). <br>\nThe same code runs in Colab TPU (v2) with finite loss and no error printouts.<br>\nBoth use TF==12.0.<br>\nSame behavior as reported by Bogdan.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 2349157,
          "author_name": "Rafi Hai",
          "author_url": "",
          "post_date": "2023-07-18T07:49:26.697000",
          "content": "<p>Experienced same behavior. When TPU connection is 'local' it gives NaN and I cannot pinpoint the reason - happens both in Kaggle environment and in gcloud TPU.<br>\nOn Colab it works fine.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2377206,
              "author_name": "Scenery SunFireInk",
              "author_url": "",
              "post_date": "2023-08-07T03:49:59.440000",
              "content": "<p>Can a graphical interface be used for training on gcloud TPU?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2356340,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-07-24T05:36:49.047000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2355729,
      "author_name": "Bohan Yoon",
      "author_url": "",
      "post_date": "2023-07-23T14:44:47.650000",
      "content": "<p>I also have the same error. It works on a single tpu, but it doesn't seem to be compatible with strategy…</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2333764,
      "author_name": "Yu Wu",
      "author_url": "",
      "post_date": "2023-07-07T06:33:21.677000",
      "content": "<p>I have exactly the same problem ……</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2313688": "Is it possible to train NN with CTC Loss on TPU? Because I tried both tf.nn.ctc_loss and keras.backend.ctc_batch_cost and when training on GPU everything is ok but when using TPU I get the following error![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10328925%2F9388ea4fe31b1890581d500eec930451%2FScreenshot%202023-06-23%20001646.png?generation=1687468661954210&alt=media)",
    "2356343": "An alternative implementation that works on Kaggle TPU has been proposed here by Greysnow:\nhttps://www.kaggle.com/competitions/asl-fingerspelling/discussion/426504",
    "2328364": "I had the same problem. The problem lies (most likely, at least it was my case) in the resampling augmentation, when you change the amount of time steps. When calculating ctc loss, TPU must know your time shape at compilation time. ",
    "2347874": "Shape problems can occur if e.g there is a partially filled last batch. TPU expects all batches to be the same size. Such issue can be fixed by using  ds.batch(drop_remainder=True) or manually trimming the dataset to exact mutiple of batches.\nIt can also be debugged by adding  x=tf.ensure_shape(x,shape). Where shape is the expected shape. \n\n\nI'm currently able to nominally make it execute on TPU but the loss is infinite or nan.  There are grappler errors in the printout that don't terminate the execution but I suspect they may be involved in the nan/inf loss.\n\nAdditionally, while training on GPU I used model.compile(jit_compile=True). It gives errors on unsupported in place add operator in ctc  and exits. (p.s: jit_compile=True is should not be used on TPU according to docs, but it is recommended to run it on GPU for debugging before switching it off and moving to TPU, as TPU also uses XLA optimization). \n\nThe same code runs fine on the GPU P100 with finite loss.\n\nI created a public example notebook borrowing code from another kaggler from a past competition: \nhttps://www.kaggle.com/code/shaironen/ctc-example/notebook. It can be run either with GPU or TPU to show the above behavior. \n",
    "2355729": "I also have the same error. It works on a single tpu, but it doesn't seem to be compatible with strategy...",
    "2333764": "I have exactly the same problem ......"
  }
}