{
  "id": 432511,
  "title": "Tflite float16 quantization not compatiable with LSTM?",
  "url": "/competitions/asl-fingerspelling/discussion/432511",
  "author_name": "Yu Wu",
  "post_date": "2023-08-17T18:42:01.818000",
  "votes": 3,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I'm using the following code to quantize my model:</p>\n<pre><code>keras_model_converter = tf(tflite_keras_model)\nkeras_model_converter = \nkeras_model_converter = \ntflite_model = keras_model_converter()\n</code></pre>\n<p>However, it always fails by exhausting all memory after a few minutes even with a toy model + a LSTM layer:</p>\n<pre><code>def get:\n    inp = tf.keras.layers.)\n\n    x = inp\n    x = tf.keras.layers.(x)\n    x = tf.keras.layers.(x)\n\n    x = tf.keras.layers.(x)\n    x = tf.keras.layers.(x)\n    model = tf.keras.\n    return model\n</code></pre>\n<p>Removing <code>keras_model_converter.target_spec.supported_types = [tf.float16]</code> or removing LSTM resolves this problem.</p>\n<p>Does anybody using LSTM and quantization have this problem and do you solve this?</p>\n<p>Thank you</p>",
  "messages": [
    {
      "id": 2395712,
      "postDate": "2023-08-17T18:42:01.817Z",
      "content": "<p>I'm using the following code to quantize my model:</p>\n<pre><code>keras_model_converter = tf(tflite_keras_model)\nkeras_model_converter = \nkeras_model_converter = \ntflite_model = keras_model_converter()\n</code></pre>\n<p>However, it always fails by exhausting all memory after a few minutes even with a toy model + a LSTM layer:</p>\n<pre><code>def get:\n    inp = tf.keras.layers.)\n\n    x = inp\n    x = tf.keras.layers.(x)\n    x = tf.keras.layers.(x)\n\n    x = tf.keras.layers.(x)\n    x = tf.keras.layers.(x)\n    model = tf.keras.\n    return model\n</code></pre>\n<p>Removing <code>keras_model_converter.target_spec.supported_types = [tf.float16]</code> or removing LSTM resolves this problem.</p>\n<p>Does anybody using LSTM and quantization have this problem and do you solve this?</p>\n<p>Thank you</p>",
      "rawMarkdown": "I'm using the following code to quantize my model:\n```\nkeras_model_converter = tf.lite.TFLiteConverter.from_keras_model(tflite_keras_model)\nkeras_model_converter.optimizations = [tf.lite.Optimize.DEFAULT]\nkeras_model_converter.target_spec.supported_types = [tf.float16]\ntflite_model = keras_model_converter.convert()\n```\n\n\n\nHowever, it always fails by exhausting all memory after a few minutes even with a toy model + a LSTM layer:\n\n```\ndef get_model():\n    inp = tf.keras.layers.Input(shape=(TEMPORAL_DIM,INPUT_DIM))\n    \n    x = inp\n    x = tf.keras.layers.Dense(512)(x)\n    x = tf.keras.layers.BatchNormalization()(x)\n\n    x = tf.keras.layers.LSTM(512,return_sequences=True)(x)\n    x = tf.keras.layers.Dense(model_cfg['num_classes'])(x)\n    model = tf.keras.Model(inputs=inp,outputs=x)\n    return model\n```\nRemoving `keras_model_converter.target_spec.supported_types = [tf.float16]` or removing LSTM resolves this problem.\n\nDoes anybody using LSTM and quantization have this problem and do you solve this?\n\nThank you",
      "votes": 3
    },
    {
      "id": 2395822,
      "postDate": "2023-08-17T21:06:49.723Z",
      "content": "<p>I had exactly same memory issue as you when doing the quantization of LSTM. I tried different running units of LSTM(64, 128, etc…) but none of them work. There seems that Tflite just no support the LSTM layer quantization. Check out this page, maybe there are some ways to resolve the issue but I haven't try it. <a href=\"https://github.com/tensorflow/tensorflow/issues/25563\" target=\"_blank\">https://github.com/tensorflow/tensorflow/issues/25563</a></p>\n<p>When I switched to GRU instead of LSTM, it can be perfectly quantized. However, my LB score drop after adding the GRU layer. If you are going to try it, please let me know. Thanks</p>",
      "rawMarkdown": "I had exactly same memory issue as you when doing the quantization of LSTM. I tried different running units of LSTM(64, 128, etc...) but none of them work. There seems that Tflite just no support the LSTM layer quantization. Check out this page, maybe there are some ways to resolve the issue but I haven't try it. https://github.com/tensorflow/tensorflow/issues/25563\n\nWhen I switched to GRU instead of LSTM, it can be perfectly quantized. However, my LB score drop after adding the GRU layer. If you are going to try it, please let me know. Thanks",
      "votes": 2,
      "replies": [
        {
          "id": 2395862,
          "postDate": "2023-08-17T22:04:05.737Z",
          "content": "<p>I also used GRU before but CV score dropped. But if LSTM can't be quantized, I think I would have to try GRU anyway 😥</p>",
          "rawMarkdown": "I also used GRU before but CV score dropped. But if LSTM can't be quantized, I think I would have to try GRU anyway 😥",
          "votes": 1
        }
      ]
    },
    {
      "id": 2399197,
      "postDate": "2023-08-20T07:45:16.713Z",
      "content": "<p>how about writing LSTM from scratch?<br>\nsince we only have one input for tfite in server evaluation, speed may not be an issue</p>",
      "rawMarkdown": "how about writing LSTM from scratch?\nsince we only have one input for tfite in server evaluation, speed may not be an issue",
      "replies": [
        {
          "id": 2399253,
          "postDate": "2023-08-20T08:21:15.570Z",
          "content": "<p>I guess it should work.  But as GRU's performance is acceptable for me and I'm still working on masking…<br>\nso I probably won't have time to do it </p>",
          "rawMarkdown": "I guess it should work.  But as GRU's performance is acceptable for me and I'm still working on masking...\nso I probably won't have time to do it "
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2395822,
      "author_name": "HW",
      "author_url": "",
      "post_date": "2023-08-17T21:06:49.723000",
      "content": "<p>I had exactly same memory issue as you when doing the quantization of LSTM. I tried different running units of LSTM(64, 128, etc…) but none of them work. There seems that Tflite just no support the LSTM layer quantization. Check out this page, maybe there are some ways to resolve the issue but I haven't try it. <a href=\"https://github.com/tensorflow/tensorflow/issues/25563\" target=\"_blank\">https://github.com/tensorflow/tensorflow/issues/25563</a></p>\n<p>When I switched to GRU instead of LSTM, it can be perfectly quantized. However, my LB score drop after adding the GRU layer. If you are going to try it, please let me know. Thanks</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2395862,
          "author_name": "Yu Wu",
          "author_url": "",
          "post_date": "2023-08-17T22:04:05.737000",
          "content": "<p>I also used GRU before but CV score dropped. But if LSTM can't be quantized, I think I would have to try GRU anyway 😥</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2399197,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-08-20T07:45:16.713000",
      "content": "<p>how about writing LSTM from scratch?<br>\nsince we only have one input for tfite in server evaluation, speed may not be an issue</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2399253,
          "author_name": "Yu Wu",
          "author_url": "",
          "post_date": "2023-08-20T08:21:15.570000",
          "content": "<p>I guess it should work.  But as GRU's performance is acceptable for me and I'm still working on masking…<br>\nso I probably won't have time to do it </p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2395712": "I'm using the following code to quantize my model:\n```\nkeras_model_converter = tf.lite.TFLiteConverter.from_keras_model(tflite_keras_model)\nkeras_model_converter.optimizations = [tf.lite.Optimize.DEFAULT]\nkeras_model_converter.target_spec.supported_types = [tf.float16]\ntflite_model = keras_model_converter.convert()\n```\n\n\n\nHowever, it always fails by exhausting all memory after a few minutes even with a toy model + a LSTM layer:\n\n```\ndef get_model():\n    inp = tf.keras.layers.Input(shape=(TEMPORAL_DIM,INPUT_DIM))\n    \n    x = inp\n    x = tf.keras.layers.Dense(512)(x)\n    x = tf.keras.layers.BatchNormalization()(x)\n\n    x = tf.keras.layers.LSTM(512,return_sequences=True)(x)\n    x = tf.keras.layers.Dense(model_cfg['num_classes'])(x)\n    model = tf.keras.Model(inputs=inp,outputs=x)\n    return model\n```\nRemoving `keras_model_converter.target_spec.supported_types = [tf.float16]` or removing LSTM resolves this problem.\n\nDoes anybody using LSTM and quantization have this problem and do you solve this?\n\nThank you",
    "2395822": "I had exactly same memory issue as you when doing the quantization of LSTM. I tried different running units of LSTM(64, 128, etc...) but none of them work. There seems that Tflite just no support the LSTM layer quantization. Check out this page, maybe there are some ways to resolve the issue but I haven't try it. https://github.com/tensorflow/tensorflow/issues/25563\n\nWhen I switched to GRU instead of LSTM, it can be perfectly quantized. However, my LB score drop after adding the GRU layer. If you are going to try it, please let me know. Thanks",
    "2399197": "how about writing LSTM from scratch?\nsince we only have one input for tfite in server evaluation, speed may not be an issue"
  }
}