{
  "id": 434183,
  "title": "Help Please! TF-Lite Submission Failure......",
  "url": "/competitions/asl-fingerspelling/discussion/434183",
  "author_name": "Kelvin0910",
  "post_date": "2023-08-24T08:19:02.549000",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi! Hope all of you are doing well!</p>\n<p>I joined the competition relatively late. I spent most of the time before training models and only started converting the model to tf-lite recently. The process is much more complicated than I expected.</p>\n<p>I managed to get the model converted and running in my own environment with packages shown below:<br>\n<code>tensorflow==2.13.0\ntensorflow-estimator==2.13.0\ntensorflow-probability==0.21.0\ntflite-runtime==2.13.0\ntflite-runtime-nightly==2.14.0.dev20230508</code></p>\n<p>When converting, I did get the following warning but the model ran fine:<br>\n<code>TFLite interpreter needs to link Flex delegate in order to run the model since it contains the following Select TFop(s):\nFlex ops: FlexConv2D, FlexTensorListReserve, FlexTensorListSetItem, FlexTensorListStack</code></p>\n<p>However, when I tried to run in the Kaggle notebook, I got the error shown in the attachment. I installed tflite-runtime using <code>!pip install tflite-runtime-nightly==2.14.0.dev20230508</code>. The submission failed after running for about 20 seconds as well.</p>\n<p>I wonder if any of you reading this post have encountered a similar issue and could kindly show me how to solve it. Let me know if I should provide any additional information. Thanks a lot in advance!</p>",
  "messages": [
    {
      "id": 2406106,
      "postDate": "2023-08-24T08:19:02.550Z",
      "content": "<p>Hi! Hope all of you are doing well!</p>\n<p>I joined the competition relatively late. I spent most of the time before training models and only started converting the model to tf-lite recently. The process is much more complicated than I expected.</p>\n<p>I managed to get the model converted and running in my own environment with packages shown below:<br>\n<code>tensorflow==2.13.0\ntensorflow-estimator==2.13.0\ntensorflow-probability==0.21.0\ntflite-runtime==2.13.0\ntflite-runtime-nightly==2.14.0.dev20230508</code></p>\n<p>When converting, I did get the following warning but the model ran fine:<br>\n<code>TFLite interpreter needs to link Flex delegate in order to run the model since it contains the following Select TFop(s):\nFlex ops: FlexConv2D, FlexTensorListReserve, FlexTensorListSetItem, FlexTensorListStack</code></p>\n<p>However, when I tried to run in the Kaggle notebook, I got the error shown in the attachment. I installed tflite-runtime using <code>!pip install tflite-runtime-nightly==2.14.0.dev20230508</code>. The submission failed after running for about 20 seconds as well.</p>\n<p>I wonder if any of you reading this post have encountered a similar issue and could kindly show me how to solve it. Let me know if I should provide any additional information. Thanks a lot in advance!</p>",
      "rawMarkdown": "Hi! Hope all of you are doing well!\n\nI joined the competition relatively late. I spent most of the time before training models and only started converting the model to tf-lite recently. The process is much more complicated than I expected.\n\nI managed to get the model converted and running in my own environment with packages shown below:\n`tensorflow==2.13.0\ntensorflow-estimator==2.13.0\ntensorflow-probability==0.21.0\ntflite-runtime==2.13.0\ntflite-runtime-nightly==2.14.0.dev20230508`\n\nWhen converting, I did get the following warning but the model ran fine:\n`TFLite interpreter needs to link Flex delegate in order to run the model since it contains the following Select TFop(s):\nFlex ops: FlexConv2D, FlexTensorListReserve, FlexTensorListSetItem, FlexTensorListStack`\n\nHowever, when I tried to run in the Kaggle notebook, I got the error shown in the attachment. I installed tflite-runtime using `!pip install tflite-runtime-nightly==2.14.0.dev20230508`. The submission failed after running for about 20 seconds as well.\n\nI wonder if any of you reading this post have encountered a similar issue and could kindly show me how to solve it. Let me know if I should provide any additional information. Thanks a lot in advance!\n",
      "votes": 1
    },
    {
      "id": 2407080,
      "postDate": "2023-08-24T19:41:50.400Z",
      "content": "<p>Sorry to hear; while working on Kaggle AI report, I saw advice of other Kagglers that joining at least two weeks prior will make a huge difference. I also struggled with failing to obtain a score in last competition despite lots of hard work to train models prior to submission… </p>\n<p>Regarding your question:</p>\n<p>Perhaps your model could not fit into server's memory? I read some discussion post about keeping models to &gt;40MB. I remember checking my output of following command to be about ~33MB and my submission was scored. </p>\n<p><code>!ls submission.zip</code></p>\n<p>Another possible reason:<br>\n<a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/433117\" target=\"_blank\">https://www.kaggle.com/competitions/asl-fingerspelling/discussion/433117</a></p>",
      "rawMarkdown": "Sorry to hear; while working on Kaggle AI report, I saw advice of other Kagglers that joining at least two weeks prior will make a huge difference. I also struggled with failing to obtain a score in last competition despite lots of hard work to train models prior to submission... \n\nRegarding your question:\n\nPerhaps your model could not fit into server's memory? I read some discussion post about keeping models to >40MB. I remember checking my output of following command to be about ~33MB and my submission was scored. \n\n```!ls submission.zip```\n\nAnother possible reason:\nhttps://www.kaggle.com/competitions/asl-fingerspelling/discussion/433117",
      "replies": [
        {
          "id": 2409373,
          "postDate": "2023-08-26T07:53:00.920Z",
          "content": "<p>Thanks you for your reply! My model file prior to compressing was smaller than 40 MB so I guess that may not be the source of the problem? The error message I got when submitting was a Submission Error. I suspect that it is due to some TF-lite conversion issue, but because of my unfamiliarity with TF-lite and the short time for me to debug I failed to resolve the issue.</p>\n<p>Anyways, thank you for your help and encouragement! And hope you do well in future competitions!</p>",
          "rawMarkdown": "Thanks you for your reply! My model file prior to compressing was smaller than 40 MB so I guess that may not be the source of the problem? The error message I got when submitting was a Submission Error. I suspect that it is due to some TF-lite conversion issue, but because of my unfamiliarity with TF-lite and the short time for me to debug I failed to resolve the issue.\n\nAnyways, thank you for your help and encouragement! And hope you do well in future competitions!",
          "replies": [
            {
              "id": 2413428,
              "postDate": "2023-08-28T22:09:27.440Z",
              "content": "<p>Sorry that my guess was wrong. </p>\n<p>In my case, I ran out of GPU compute times and little time to migrate so I had to give up too early :p</p>\n<p>For the model conversion, I can confirm that below code adopted from <a href=\"https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place\" target=\"_blank\">notebook</a> will generate a score (and only not when I increased the hyperparameters too drastically (e.g. # layers, # channels):</p>\n<pre><code>keras_model_converter = tf(tflitemodel_base)\n\nkeras_model_converter = #, tf.SELECT_TF_OPS]\nkeras_model_converter = \nkeras_model_converter = \ntflite_model = keras_model_converter()\nwith (, ) as f:\n    f(tflite_model)\n\nwith (, ) as f:\n    json({ : SEL_COLS}, f)\n\n!zip submission   \n</code></pre>\n<p>If you have spare time, it may be good idea to diagnose the problems that hindered your submission to avoid it in the next competition. It is also insightful to find out the score of your trained model so you could promote your work to the community with that info ;-)</p>\n<p>Best of luck to your future submissions!</p>",
              "rawMarkdown": "Sorry that my guess was wrong. \n\nIn my case, I ran out of GPU compute times and little time to migrate so I had to give up too early :p\n\nFor the model conversion, I can confirm that below code adopted from [notebook](https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place) will generate a score (and only not when I increased the hyperparameters too drastically (e.g. # layers, # channels):\n \n```\nkeras_model_converter = tf.lite.TFLiteConverter.from_keras_model(tflitemodel_base)\n\nkeras_model_converter.target_spec.supported_ops = [tf.lite.OpsSet.TFLITE_BUILTINS]#, tf.lite.OpsSet.SELECT_TF_OPS]\nkeras_model_converter.optimizations = [tf.lite.Optimize.DEFAULT]\nkeras_model_converter.target_spec.supported_types = [tf.float16]\ntflite_model = keras_model_converter.convert()\nwith open('model.tflite', 'wb') as f:\n    f.write(tflite_model)\n    \nwith open('inference_args.json', \"w\") as f:\n    json.dump({\"selected_columns\" : SEL_COLS}, f)\n    \n!zip submission.zip  './model.tflite' './inference_args.json'\n```\n\nIf you have spare time, it may be good idea to diagnose the problems that hindered your submission to avoid it in the next competition. It is also insightful to find out the score of your trained model so you could promote your work to the community with that info ;-)\n\nBest of luck to your future submissions!\n"
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2407080,
      "author_name": "Vcol",
      "author_url": "",
      "post_date": "2023-08-24T19:41:50.400000",
      "content": "<p>Sorry to hear; while working on Kaggle AI report, I saw advice of other Kagglers that joining at least two weeks prior will make a huge difference. I also struggled with failing to obtain a score in last competition despite lots of hard work to train models prior to submission… </p>\n<p>Regarding your question:</p>\n<p>Perhaps your model could not fit into server's memory? I read some discussion post about keeping models to &gt;40MB. I remember checking my output of following command to be about ~33MB and my submission was scored. </p>\n<p><code>!ls submission.zip</code></p>\n<p>Another possible reason:<br>\n<a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/433117\" target=\"_blank\">https://www.kaggle.com/competitions/asl-fingerspelling/discussion/433117</a></p>",
      "votes": 0,
      "replies": [
        {
          "id": 2409373,
          "author_name": "Kelvin0910",
          "author_url": "",
          "post_date": "2023-08-26T07:53:00.920000",
          "content": "<p>Thanks you for your reply! My model file prior to compressing was smaller than 40 MB so I guess that may not be the source of the problem? The error message I got when submitting was a Submission Error. I suspect that it is due to some TF-lite conversion issue, but because of my unfamiliarity with TF-lite and the short time for me to debug I failed to resolve the issue.</p>\n<p>Anyways, thank you for your help and encouragement! And hope you do well in future competitions!</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2413428,
              "author_name": "Vcol",
              "author_url": "",
              "post_date": "2023-08-28T22:09:27.440000",
              "content": "<p>Sorry that my guess was wrong. </p>\n<p>In my case, I ran out of GPU compute times and little time to migrate so I had to give up too early :p</p>\n<p>For the model conversion, I can confirm that below code adopted from <a href=\"https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place\" target=\"_blank\">notebook</a> will generate a score (and only not when I increased the hyperparameters too drastically (e.g. # layers, # channels):</p>\n<pre><code>keras_model_converter = tf(tflitemodel_base)\n\nkeras_model_converter = #, tf.SELECT_TF_OPS]\nkeras_model_converter = \nkeras_model_converter = \ntflite_model = keras_model_converter()\nwith (, ) as f:\n    f(tflite_model)\n\nwith (, ) as f:\n    json({ : SEL_COLS}, f)\n\n!zip submission   \n</code></pre>\n<p>If you have spare time, it may be good idea to diagnose the problems that hindered your submission to avoid it in the next competition. It is also insightful to find out the score of your trained model so you could promote your work to the community with that info ;-)</p>\n<p>Best of luck to your future submissions!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2406106": "Hi! Hope all of you are doing well!\n\nI joined the competition relatively late. I spent most of the time before training models and only started converting the model to tf-lite recently. The process is much more complicated than I expected.\n\nI managed to get the model converted and running in my own environment with packages shown below:\n`tensorflow==2.13.0\ntensorflow-estimator==2.13.0\ntensorflow-probability==0.21.0\ntflite-runtime==2.13.0\ntflite-runtime-nightly==2.14.0.dev20230508`\n\nWhen converting, I did get the following warning but the model ran fine:\n`TFLite interpreter needs to link Flex delegate in order to run the model since it contains the following Select TFop(s):\nFlex ops: FlexConv2D, FlexTensorListReserve, FlexTensorListSetItem, FlexTensorListStack`\n\nHowever, when I tried to run in the Kaggle notebook, I got the error shown in the attachment. I installed tflite-runtime using `!pip install tflite-runtime-nightly==2.14.0.dev20230508`. The submission failed after running for about 20 seconds as well.\n\nI wonder if any of you reading this post have encountered a similar issue and could kindly show me how to solve it. Let me know if I should provide any additional information. Thanks a lot in advance!\n",
    "2407080": "Sorry to hear; while working on Kaggle AI report, I saw advice of other Kagglers that joining at least two weeks prior will make a huge difference. I also struggled with failing to obtain a score in last competition despite lots of hard work to train models prior to submission... \n\nRegarding your question:\n\nPerhaps your model could not fit into server's memory? I read some discussion post about keeping models to >40MB. I remember checking my output of following command to be about ~33MB and my submission was scored. \n\n```!ls submission.zip```\n\nAnother possible reason:\nhttps://www.kaggle.com/competitions/asl-fingerspelling/discussion/433117"
  }
}