{
  "id": 417931,
  "title": "Is random data augmentation slow in tensorflow?",
  "url": "/competitions/asl-fingerspelling/discussion/417931",
  "author_name": "Yu Wu",
  "post_date": "2023-06-17T23:16:01.798000",
  "votes": 4,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>I'm working on random data augmentation in TensorFlow.</p>\n<p>I start with adding a very simple random swapping in my <code>preprocess</code> function like:</p>\n<pre><code> tf.random.uniform(()) &gt; .:\n        ,lhand = lhand,rhand\n        ,lpose = lpose,rpose\n</code></pre>\n<p>This increases my training time from approximately <code>50ms/step</code> to about <code>80ms/step</code>.</p>\n<p>However, if I remove the random condition to like:</p>\n<pre><code>\nlhand = lhand,rhand\nlpose = lpose,rpose\n</code></pre>\n<p>The training time is back to <code>50ms/step</code> again so I suspect this delay is from <code>tf.random.uniform()</code>.</p>\n<p>As this delay causes about <code>60%</code> increasing training time, I'm wondering if everyone has to face this problem or if there is a faster way to do random augmentation?</p>",
  "messages": [
    {
      "id": 2307172,
      "postDate": "2023-06-17T23:16:01.800Z",
      "content": "<p>Hi everyone,</p>\n<p>I'm working on random data augmentation in TensorFlow.</p>\n<p>I start with adding a very simple random swapping in my <code>preprocess</code> function like:</p>\n<pre><code> tf.random.uniform(()) &gt; .:\n        ,lhand = lhand,rhand\n        ,lpose = lpose,rpose\n</code></pre>\n<p>This increases my training time from approximately <code>50ms/step</code> to about <code>80ms/step</code>.</p>\n<p>However, if I remove the random condition to like:</p>\n<pre><code>\nlhand = lhand,rhand\nlpose = lpose,rpose\n</code></pre>\n<p>The training time is back to <code>50ms/step</code> again so I suspect this delay is from <code>tf.random.uniform()</code>.</p>\n<p>As this delay causes about <code>60%</code> increasing training time, I'm wondering if everyone has to face this problem or if there is a faster way to do random augmentation?</p>",
      "rawMarkdown": "Hi everyone,\n\nI'm working on random data augmentation in TensorFlow.\n\nI start with adding a very simple random swapping in my `preprocess` function like:\n\n```\nif tf.random.uniform(()) > 0.5:\n        rhand,lhand = lhand,rhand\n        rpose,lpose = lpose,rpose\n```\n\nThis increases my training time from approximately `50ms/step` to about `80ms/step`.\n\nHowever, if I remove the random condition to like:\n\n```\nif True:\n        rhand,lhand = lhand,rhand\n        rpose,lpose = lpose,rpose\n```\n\nThe training time is back to `50ms/step` again so I suspect this delay is from `tf.random.uniform()`.\n\nAs this delay causes about `60%` increasing training time, I'm wondering if everyone has to face this problem or if there is a faster way to do random augmentation?\n\n\n",
      "votes": 4
    },
    {
      "id": 2307179,
      "postDate": "2023-06-17T23:42:26.783Z",
      "content": "<p>If you suspect that tf.random is very slow, try to replace it with the np.random equivalent…</p>",
      "rawMarkdown": "If you suspect that tf.random is very slow, try to replace it with the np.random equivalent...",
      "votes": 1,
      "replies": [
        {
          "id": 2307191,
          "postDate": "2023-06-18T00:27:28.500Z",
          "content": "<p><code>np.random.uniform()</code> works perfectly (at least in terms of speed)✌️but why is tensorflow's random function so slow ….. 🤔🤔🤔</p>",
          "rawMarkdown": "`np.random.uniform()` works perfectly (at least in terms of speed)✌️but why is tensorflow's random function so slow ..... 🤔🤔🤔",
          "votes": 1,
          "replies": [
            {
              "id": 2312921,
              "postDate": "2023-06-22T09:24:51.540Z",
              "content": "<p>np.random.uniform() works well</p>",
              "rawMarkdown": "np.random.uniform() works well"
            },
            {
              "id": 2315300,
              "postDate": "2023-06-24T03:30:18.413Z",
              "content": "<p>Random with numpy or pure python random lib will be make this node to be constant. Ie np.random.randn() is tf.constant and your augmentations are wrong !!!</p>",
              "rawMarkdown": "Random with numpy or pure python random lib will be make this node to be constant. Ie np.random.randn() is tf.constant and your augmentations are wrong !!!",
              "votes": 4
            },
            {
              "id": 2315321,
              "postDate": "2023-06-24T04:06:00.470Z",
              "content": "<p>Yes, you're right, I just got stuck because of it for 1 week….. 😭😭😭</p>",
              "rawMarkdown": "Yes, you're right, I just got stuck because of it for 1 week..... 😭😭😭"
            },
            {
              "id": 2401731,
              "postDate": "2023-08-21T18:40:04.213Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2401734,
              "postDate": "2023-08-21T18:40:58.957Z",
              "content": "<p>Thank you. This should be pinned.</p>",
              "rawMarkdown": "Thank you. This should be pinned."
            }
          ]
        }
      ]
    },
    {
      "id": 2401752,
      "postDate": "2023-08-21T18:50:47.323Z",
      "content": "<p>i never get tf data pipeline to work. i don't know why. <br>\neven if i code everything in tf function and add the tf.function decorator, it seems so slow.</p>\n<p>the data pipeline is difficult to debug, becuase breakpoint cannot be set (even with tf.experimental.data decorator).<br>\nworse still the debug code can run differently from the non-debug one (e.g.  tensor.numpy() is only available in eagar mode).<br>\nmy code end up  working in debug mode, but not running correctly in runtime</p>\n<p>in the end i use pytorch for training and then trasnfer the pytorch weights to a keras model for tflite conversion.<br>\nmy pytorch frame work is training 10x faster than my tf/keras one.</p>",
      "rawMarkdown": "i never get tf data pipeline to work. i don't know why. \neven if i code everything in tf function and add the tf.function decorator, it seems so slow.\n\nthe data pipeline is difficult to debug, becuase breakpoint cannot be set (even with tf.experimental.data decorator).\nworse still the debug code can run differently from the non-debug one (e.g.  tensor.numpy() is only available in eagar mode).\nmy code end up  working in debug mode, but not running correctly in runtime\n\nin the end i use pytorch for training and then trasnfer the pytorch weights to a keras model for tflite conversion.\nmy pytorch frame work is training 10x faster than my tf/keras one.",
      "replies": [
        {
          "id": 2401789,
          "postDate": "2023-08-21T19:24:33.663Z",
          "content": "<p>I found out that I need to manually add <code>tf.data.AUTOTUNE</code> into data loading-related functions. </p>\n<p>I followed the previous 1st solution (<a href=\"https://www.kaggle.com/code/hoyso48/1st-place-solution-training):\" target=\"_blank\">https://www.kaggle.com/code/hoyso48/1st-place-solution-training):</a></p>\n<pre><code>     = .data.TFRecordDataset(tfrecords, num_parallel_reads=.data.AUTOTUNE, compression_type=)\n     = .(decode_tfrec, .data.AUTOTUNE)\n     = .(lambda : preprocess(, augment=augment, max_len=max_len), .data.AUTOTUNE)\n    ...\n    ...\n</code></pre>\n<p>Maybe this can help</p>",
          "rawMarkdown": "I found out that I need to manually add `tf.data.AUTOTUNE` into data loading-related functions. \n\nI followed the previous 1st solution (https://www.kaggle.com/code/hoyso48/1st-place-solution-training):\n```\n    ds = tf.data.TFRecordDataset(tfrecords, num_parallel_reads=tf.data.AUTOTUNE, compression_type='GZIP')\n    ds = ds.map(decode_tfrec, tf.data.AUTOTUNE)\n    ds = ds.map(lambda x: preprocess(x, augment=augment, max_len=max_len), tf.data.AUTOTUNE)\n    ...\n    ...\n```\n\nMaybe this can help",
          "replies": [
            {
              "id": 2401990,
              "postDate": "2023-08-22T00:21:00.393Z",
              "content": "<p>how fast does your tf train run?</p>",
              "rawMarkdown": "how fast does your tf train run?"
            },
            {
              "id": 2402039,
              "postDate": "2023-08-22T00:57:11.793Z",
              "content": "<p>It depends on the model I use. My 15M parameter model runs ~55s per epoch on kaggle TPU with 256 batch size. </p>",
              "rawMarkdown": "It depends on the model I use. My 15M parameter model runs ~55s per epoch on kaggle TPU with 256 batch size. "
            }
          ]
        },
        {
          "id": 2401879,
          "postDate": "2023-08-21T21:19:21.587Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 2402374,
          "postDate": "2023-08-22T05:52:57.700Z",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>  Yes, I also found tf much slower then torch(3-4times slower on 1*4090 gpu) for training, not sure why, might due to some driver issues as I found intel cpu is much slower then amd for tensorflow but for torch do not have this issue. Also using torch we could use torch.compile to further speedup. So if not using tpu, torch may be a better choice.<br>\nFor those using tensorflow I would recommend using <a href=\"https://github.com/alexeytochin/tf_seq2seq_losses\" target=\"_blank\">https://github.com/alexeytochin/tf_seq2seq_losses</a> as I found the original ctc loss of tensorflow is too slow.</p>",
          "rawMarkdown": "@hengck23  Yes, I also found tf much slower then torch(3-4times slower on 1*4090 gpu) for training, not sure why, might due to some driver issues as I found intel cpu is much slower then amd for tensorflow but for torch do not have this issue. Also using torch we could use torch.compile to further speedup. So if not using tpu, torch may be a better choice.\nFor those using tensorflow I would recommend using https://github.com/alexeytochin/tf_seq2seq_losses as I found the original ctc loss of tensorflow is too slow.",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2307179,
      "author_name": "greySnow",
      "author_url": "",
      "post_date": "2023-06-17T23:42:26.783000",
      "content": "<p>If you suspect that tf.random is very slow, try to replace it with the np.random equivalent…</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2307191,
          "author_name": "Yu Wu",
          "author_url": "",
          "post_date": "2023-06-18T00:27:28.500000",
          "content": "<p><code>np.random.uniform()</code> works perfectly (at least in terms of speed)✌️but why is tensorflow's random function so slow ….. 🤔🤔🤔</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2312921,
              "author_name": "Jithendra Prathipati",
              "author_url": "",
              "post_date": "2023-06-22T09:24:51.540000",
              "content": "<p>np.random.uniform() works well</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2315300,
              "author_name": "Just A game on your lips",
              "author_url": "",
              "post_date": "2023-06-24T03:30:18.413000",
              "content": "<p>Random with numpy or pure python random lib will be make this node to be constant. Ie np.random.randn() is tf.constant and your augmentations are wrong !!!</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2315321,
              "author_name": "Yu Wu",
              "author_url": "",
              "post_date": "2023-06-24T04:06:00.470000",
              "content": "<p>Yes, you're right, I just got stuck because of it for 1 week….. 😭😭😭</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2401731,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-08-21T18:40:04.213000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2401734,
              "author_name": "JoAI",
              "author_url": "",
              "post_date": "2023-08-21T18:40:58.957000",
              "content": "<p>Thank you. This should be pinned.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2401752,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-08-21T18:50:47.323000",
      "content": "<p>i never get tf data pipeline to work. i don't know why. <br>\neven if i code everything in tf function and add the tf.function decorator, it seems so slow.</p>\n<p>the data pipeline is difficult to debug, becuase breakpoint cannot be set (even with tf.experimental.data decorator).<br>\nworse still the debug code can run differently from the non-debug one (e.g.  tensor.numpy() is only available in eagar mode).<br>\nmy code end up  working in debug mode, but not running correctly in runtime</p>\n<p>in the end i use pytorch for training and then trasnfer the pytorch weights to a keras model for tflite conversion.<br>\nmy pytorch frame work is training 10x faster than my tf/keras one.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2401789,
          "author_name": "Yu Wu",
          "author_url": "",
          "post_date": "2023-08-21T19:24:33.663000",
          "content": "<p>I found out that I need to manually add <code>tf.data.AUTOTUNE</code> into data loading-related functions. </p>\n<p>I followed the previous 1st solution (<a href=\"https://www.kaggle.com/code/hoyso48/1st-place-solution-training):\" target=\"_blank\">https://www.kaggle.com/code/hoyso48/1st-place-solution-training):</a></p>\n<pre><code>     = .data.TFRecordDataset(tfrecords, num_parallel_reads=.data.AUTOTUNE, compression_type=)\n     = .(decode_tfrec, .data.AUTOTUNE)\n     = .(lambda : preprocess(, augment=augment, max_len=max_len), .data.AUTOTUNE)\n    ...\n    ...\n</code></pre>\n<p>Maybe this can help</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2401990,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-08-22T00:21:00.393000",
              "content": "<p>how fast does your tf train run?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2402039,
              "author_name": "Yu Wu",
              "author_url": "",
              "post_date": "2023-08-22T00:57:11.793000",
              "content": "<p>It depends on the model I use. My 15M parameter model runs ~55s per epoch on kaggle TPU with 256 batch size. </p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2401879,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-08-21T21:19:21.587000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2402374,
          "author_name": "gezi",
          "author_url": "",
          "post_date": "2023-08-22T05:52:57.700000",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>  Yes, I also found tf much slower then torch(3-4times slower on 1*4090 gpu) for training, not sure why, might due to some driver issues as I found intel cpu is much slower then amd for tensorflow but for torch do not have this issue. Also using torch we could use torch.compile to further speedup. So if not using tpu, torch may be a better choice.<br>\nFor those using tensorflow I would recommend using <a href=\"https://github.com/alexeytochin/tf_seq2seq_losses\" target=\"_blank\">https://github.com/alexeytochin/tf_seq2seq_losses</a> as I found the original ctc loss of tensorflow is too slow.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2307172": "Hi everyone,\n\nI'm working on random data augmentation in TensorFlow.\n\nI start with adding a very simple random swapping in my `preprocess` function like:\n\n```\nif tf.random.uniform(()) > 0.5:\n        rhand,lhand = lhand,rhand\n        rpose,lpose = lpose,rpose\n```\n\nThis increases my training time from approximately `50ms/step` to about `80ms/step`.\n\nHowever, if I remove the random condition to like:\n\n```\nif True:\n        rhand,lhand = lhand,rhand\n        rpose,lpose = lpose,rpose\n```\n\nThe training time is back to `50ms/step` again so I suspect this delay is from `tf.random.uniform()`.\n\nAs this delay causes about `60%` increasing training time, I'm wondering if everyone has to face this problem or if there is a faster way to do random augmentation?\n\n\n",
    "2307179": "If you suspect that tf.random is very slow, try to replace it with the np.random equivalent...",
    "2401752": "i never get tf data pipeline to work. i don't know why. \neven if i code everything in tf function and add the tf.function decorator, it seems so slow.\n\nthe data pipeline is difficult to debug, becuase breakpoint cannot be set (even with tf.experimental.data decorator).\nworse still the debug code can run differently from the non-debug one (e.g.  tensor.numpy() is only available in eagar mode).\nmy code end up  working in debug mode, but not running correctly in runtime\n\nin the end i use pytorch for training and then trasnfer the pytorch weights to a keras model for tflite conversion.\nmy pytorch frame work is training 10x faster than my tf/keras one."
  }
}