{
  "id": 425224,
  "title": "TPU running extremely slowly",
  "url": "/competitions/asl-fingerspelling/discussion/425224",
  "author_name": "nymfree",
  "post_date": "2023-07-17T18:19:03.189000",
  "votes": 0,
  "comment_count": 3,
  "views": 0,
  "content": "<p>My TPU run is way slower than the GPU run - Runtime per epoch is very similar to what I got on my local CPU<br>\ntaking around 1700 seconds per epoch.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6216048%2Fe14a553d57755150c74cde76692a6ba3%2FCapture2.PNG?generation=1689617925106040&amp;alt=media\" alt=\"\"><br>\nUsage shows excessive CPU usage</p>",
  "messages": [
    {
      "id": 2348586,
      "postDate": "2023-07-17T18:43:17.577Z",
      "content": "<p>My mistake. I missed a step.</p>",
      "rawMarkdown": "My mistake. I missed a step."
    },
    {
      "id": 2348584,
      "postDate": "2023-07-17T18:40:34.097Z",
      "content": "<p>you can check if you are on TPU with ptint(strategy) or with print(\"REPLICAS: \", strategy.num_replicas_in_sync). Don't forget the boilerplate code for TPU i.e. </p>\n<p><code># Configure Strategy. Assume TPU...if not set default for GPU\ntpu = None\ntry:\n    tpu = tf.distribute.cluster_resolver.TPUClusterResolver.connect(tpu=\"local\") # \"local\" for 1VM TPU\n    strategy = tf.distribute.TPUStrategy(tpu)\n    print(\"on TPU\")\n    print(\"REPLICAS: \", strategy.num_replicas_in_sync)\nexcept:\n    strategy = tf.distribute.get_strategy()</code></p>\n<p>(I don't manage to get it in the form of readable code…ugh)</p>",
      "rawMarkdown": "you can check if you are on TPU with ptint(strategy) or with print(\"REPLICAS: \", strategy.num_replicas_in_sync). Don't forget the boilerplate code for TPU i.e. \n\n`# Configure Strategy. Assume TPU...if not set default for GPU\ntpu = None\ntry:\n    tpu = tf.distribute.cluster_resolver.TPUClusterResolver.connect(tpu=\"local\") # \"local\" for 1VM TPU\n    strategy = tf.distribute.TPUStrategy(tpu)\n    print(\"on TPU\")\n    print(\"REPLICAS: \", strategy.num_replicas_in_sync)\nexcept:\n    strategy = tf.distribute.get_strategy()`\n\n\n(I don't manage to get it in the form of readable code...ugh)"
    },
    {
      "id": 2348579,
      "postDate": "2023-07-17T18:34:52.603Z",
      "content": "<p>not even sure if it is using the tpu at all</p>",
      "rawMarkdown": "not even sure if it is using the tpu at all"
    },
    {
      "id": 2348568,
      "postDate": "2023-07-17T18:19:03.190Z",
      "content": "<p>My TPU run is way slower than the GPU run - Runtime per epoch is very similar to what I got on my local CPU<br>\ntaking around 1700 seconds per epoch.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6216048%2Fe14a553d57755150c74cde76692a6ba3%2FCapture2.PNG?generation=1689617925106040&amp;alt=media\" alt=\"\"><br>\nUsage shows excessive CPU usage</p>",
      "rawMarkdown": "My TPU run is way slower than the GPU run - Runtime per epoch is very similar to what I got on my local CPU\ntaking around 1700 seconds per epoch.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6216048%2Fe14a553d57755150c74cde76692a6ba3%2FCapture2.PNG?generation=1689617925106040&alt=media)\nUsage shows excessive CPU usage"
    }
  ],
  "comments": [
    {
      "id": 2348586,
      "author_name": "nymfree",
      "author_url": "",
      "post_date": "2023-07-17T18:43:17.577000",
      "content": "<p>My mistake. I missed a step.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2348584,
      "author_name": "greySnow",
      "author_url": "",
      "post_date": "2023-07-17T18:40:34.097000",
      "content": "<p>you can check if you are on TPU with ptint(strategy) or with print(\"REPLICAS: \", strategy.num_replicas_in_sync). Don't forget the boilerplate code for TPU i.e. </p>\n<p><code># Configure Strategy. Assume TPU...if not set default for GPU\ntpu = None\ntry:\n    tpu = tf.distribute.cluster_resolver.TPUClusterResolver.connect(tpu=\"local\") # \"local\" for 1VM TPU\n    strategy = tf.distribute.TPUStrategy(tpu)\n    print(\"on TPU\")\n    print(\"REPLICAS: \", strategy.num_replicas_in_sync)\nexcept:\n    strategy = tf.distribute.get_strategy()</code></p>\n<p>(I don't manage to get it in the form of readable code…ugh)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2348579,
      "author_name": "nymfree",
      "author_url": "",
      "post_date": "2023-07-17T18:34:52.603000",
      "content": "<p>not even sure if it is using the tpu at all</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2348586": "My mistake. I missed a step.",
    "2348584": "you can check if you are on TPU with ptint(strategy) or with print(\"REPLICAS: \", strategy.num_replicas_in_sync). Don't forget the boilerplate code for TPU i.e. \n\n`# Configure Strategy. Assume TPU...if not set default for GPU\ntpu = None\ntry:\n    tpu = tf.distribute.cluster_resolver.TPUClusterResolver.connect(tpu=\"local\") # \"local\" for 1VM TPU\n    strategy = tf.distribute.TPUStrategy(tpu)\n    print(\"on TPU\")\n    print(\"REPLICAS: \", strategy.num_replicas_in_sync)\nexcept:\n    strategy = tf.distribute.get_strategy()`\n\n\n(I don't manage to get it in the form of readable code...ugh)",
    "2348579": "not even sure if it is using the tpu at all",
    "2348568": "My TPU run is way slower than the GPU run - Runtime per epoch is very similar to what I got on my local CPU\ntaking around 1700 seconds per epoch.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6216048%2Fe14a553d57755150c74cde76692a6ba3%2FCapture2.PNG?generation=1689617925106040&alt=media)\nUsage shows excessive CPU usage"
  }
}