{
  "id": 431565,
  "title": "[tf.lite.Optimize.DEFAULT] Post-training quantization increases inference time?",
  "url": "/competitions/asl-fingerspelling/discussion/431565",
  "author_name": "Yu Wu",
  "post_date": "2023-08-14T05:50:23.997000",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi guys,</p>\n<p>I'm using the following dynamic range quantization method (suggested by: <a href=\"https://www.tensorflow.org/lite/performance/post_training_quantization):\" target=\"_blank\">https://www.tensorflow.org/lite/performance/post_training_quantization):</a></p>\n<pre><code> = [tf.lite.Optimize.DEFAULT]\n = converter.convert()\n</code></pre>\n<p>My model's size has indeed decreased, (from ~40MB to ~10MB) but inference time increased by ~10%.</p>\n<p>Is this normal?</p>",
  "messages": [
    {
      "id": 2389776,
      "postDate": "2023-08-14T08:22:01.320Z",
      "content": "<p>Yes, it is normal. Dynamic quantization will quantize the intermediate activations dynamically during inference, which normally requires more time.</p>\n<p>May I ask how about its performance compared to FP32? Previously I couldn't obtain a better score than FP32</p>",
      "rawMarkdown": "Yes, it is normal. Dynamic quantization will quantize the intermediate activations dynamically during inference, which normally requires more time.\n\nMay I ask how about its performance compared to FP32? Previously I couldn't obtain a better score than FP32",
      "votes": 1,
      "replies": [
        {
          "id": 2389824,
          "postDate": "2023-08-14T08:52:08.050Z",
          "content": "<p>I haven't submitted it so I don't know its LB score. But the CV score is basically the same (difference ~ 0.001, slightly worse)</p>",
          "rawMarkdown": "I haven't submitted it so I don't know its LB score. But the CV score is basically the same (difference ~ 0.001, slightly worse)",
          "votes": 1
        }
      ]
    },
    {
      "id": 2389539,
      "postDate": "2023-08-14T05:50:23.997Z",
      "content": "<p>Hi guys,</p>\n<p>I'm using the following dynamic range quantization method (suggested by: <a href=\"https://www.tensorflow.org/lite/performance/post_training_quantization):\" target=\"_blank\">https://www.tensorflow.org/lite/performance/post_training_quantization):</a></p>\n<pre><code> = [tf.lite.Optimize.DEFAULT]\n = converter.convert()\n</code></pre>\n<p>My model's size has indeed decreased, (from ~40MB to ~10MB) but inference time increased by ~10%.</p>\n<p>Is this normal?</p>",
      "rawMarkdown": "Hi guys,\n\nI'm using the following dynamic range quantization method (suggested by: https://www.tensorflow.org/lite/performance/post_training_quantization):\n```\nconverter.optimizations = [tf.lite.Optimize.DEFAULT]\ntflite_quant_model = converter.convert()\n```\n\nMy model's size has indeed decreased, (from ~40MB to ~10MB) but inference time increased by ~10%.\n\nIs this normal?",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 2389776,
      "author_name": "bliao",
      "author_url": "",
      "post_date": "2023-08-14T08:22:01.320000",
      "content": "<p>Yes, it is normal. Dynamic quantization will quantize the intermediate activations dynamically during inference, which normally requires more time.</p>\n<p>May I ask how about its performance compared to FP32? Previously I couldn't obtain a better score than FP32</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2389824,
          "author_name": "Yu Wu",
          "author_url": "",
          "post_date": "2023-08-14T08:52:08.050000",
          "content": "<p>I haven't submitted it so I don't know its LB score. But the CV score is basically the same (difference ~ 0.001, slightly worse)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2389776": "Yes, it is normal. Dynamic quantization will quantize the intermediate activations dynamically during inference, which normally requires more time.\n\nMay I ask how about its performance compared to FP32? Previously I couldn't obtain a better score than FP32",
    "2389539": "Hi guys,\n\nI'm using the following dynamic range quantization method (suggested by: https://www.tensorflow.org/lite/performance/post_training_quantization):\n```\nconverter.optimizations = [tf.lite.Optimize.DEFAULT]\ntflite_quant_model = converter.convert()\n```\n\nMy model's size has indeed decreased, (from ~40MB to ~10MB) but inference time increased by ~10%.\n\nIs this normal?"
  }
}