{
  "id": 409722,
  "title": "Evaluation Time | Submission Issues",
  "url": "/competitions/asl-fingerspelling/discussion/409722",
  "author_name": "Kolya Forrat",
  "post_date": "2023-05-12T09:50:59.950000",
  "votes": 20,
  "comment_count": 37,
  "views": 0,
  "content": "<p>Hello everyone! Nice to see that type of competition again</p>\n<p>I have one question. Could hosts please clarify meaning of </p>\n<blockquote>\n  <p>Your model must also perform inference in less than 5 hours and use less than 40 MB of storage space. Expect to see approximately 35 hours of video in the test set.</p>\n</blockquote>\n<p>How to calculate how many phrases 35 hours includes? I would like to estimate one phrase inference time or how many phrases we have in that 35 hours.</p>\n<p>Thanks</p>\n<p>UPDATE1. Also faced issues with submission. But can't detect the reason with info from Evaluation page. More info posted <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/409722#2273788\" target=\"_blank\">below</a></p>\n<p>UPDATE2. <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/409722#2276829\" target=\"_blank\">Fixed</a></p>\n<p>UPDATE3. Approximation of inference time. <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/409722#2276907\" target=\"_blank\">here</a> and <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/409722#2308349\" target=\"_blank\">here</a></p>",
  "messages": [
    {
      "id": 2256169,
      "postDate": "2023-05-12T09:50:59.950Z",
      "content": "<p>Hello everyone! Nice to see that type of competition again</p>\n<p>I have one question. Could hosts please clarify meaning of </p>\n<blockquote>\n  <p>Your model must also perform inference in less than 5 hours and use less than 40 MB of storage space. Expect to see approximately 35 hours of video in the test set.</p>\n</blockquote>\n<p>How to calculate how many phrases 35 hours includes? I would like to estimate one phrase inference time or how many phrases we have in that 35 hours.</p>\n<p>Thanks</p>\n<p>UPDATE1. Also faced issues with submission. But can't detect the reason with info from Evaluation page. More info posted <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/409722#2273788\" target=\"_blank\">below</a></p>\n<p>UPDATE2. <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/409722#2276829\" target=\"_blank\">Fixed</a></p>\n<p>UPDATE3. Approximation of inference time. <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/409722#2276907\" target=\"_blank\">here</a> and <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/409722#2308349\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "Hello everyone! Nice to see that type of competition again\n\nI have one question. Could hosts please clarify meaning of \n>Your model must also perform inference in less than 5 hours and use less than 40 MB of storage space. Expect to see approximately 35 hours of video in the test set.\n\nHow to calculate how many phrases 35 hours includes? I would like to estimate one phrase inference time or how many phrases we have in that 35 hours.\n\nThanks\n\n\nUPDATE1. Also faced issues with submission. But can't detect the reason with info from Evaluation page. More info posted [below](https://www.kaggle.com/competitions/asl-fingerspelling/discussion/409722#2273788)\n\n\nUPDATE2. [Fixed](https://www.kaggle.com/competitions/asl-fingerspelling/discussion/409722#2276829)\n\nUPDATE3. Approximation of inference time. [here](https://www.kaggle.com/competitions/asl-fingerspelling/discussion/409722#2276907) and [here](https://www.kaggle.com/competitions/asl-fingerspelling/discussion/409722#2308349)",
      "votes": 19
    },
    {
      "id": 2276907,
      "postDate": "2023-05-27T09:56:56.353Z",
      "content": "<p>I also calculated the approximate time of one sample inference. <br>\nSubmission scoring lasts ~3hours, mean inference time ~0.42 seconds. So:<br>\n<code>3*3600 seconds / 0.42 ~= 25714</code> samples</p>\n<p>We have 5 hours limitation. With this logic approximate maximum Mean Only Inference Time for model ~ 0.7 seconds</p>",
      "rawMarkdown": "I also calculated the approximate time of one sample inference. \nSubmission scoring lasts ~3hours, mean inference time ~0.42 seconds. So:\n`3*3600 seconds / 0.42 ~= 25714` samples\n\nWe have 5 hours limitation. With this logic approximate maximum Mean Only Inference Time for model ~ 0.7 seconds",
      "votes": 8,
      "replies": [
        {
          "id": 2307211,
          "postDate": "2023-06-18T02:07:40.117Z",
          "content": "<p>Well, I used the Kaggle notebook to test my model's inference speed, and my model with ~0.6s  inference time exceeds the time limit ……</p>",
          "rawMarkdown": "Well, I used the Kaggle notebook to test my model's inference speed, and my model with ~0.6s  inference time exceeds the time limit ......",
          "replies": [
            {
              "id": 2307955,
              "postDate": "2023-06-18T14:40:58.893Z",
              "content": "<p>Maybe it depends of which subset of data do you use for speed test? And which tflite import you are using<code>import tflite_runtime.interpreter as tflite</code> or <code>import tensorflow.lite as tflite</code><br>\n<a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/409722#2273788\" target=\"_blank\">Below</a> I shared the way I tested it. For first 1000 samples from train.csv</p>\n<p>Anyways, my calculations are approximate. And if you for example share your results for the same subset we can make some more accurate prediction of maximum inference time.</p>",
              "rawMarkdown": "Maybe it depends of which subset of data do you use for speed test? And which tflite import you are using` import tflite_runtime.interpreter as tflite` or `import tensorflow.lite as tflite`\n[Below](https://www.kaggle.com/competitions/asl-fingerspelling/discussion/409722#2273788) I shared the way I tested it. For first 1000 samples from train.csv\n\nAnyways, my calculations are approximate. And if you for example share your results for the same subset we can make some more accurate prediction of maximum inference time."
            },
            {
              "id": 2308304,
              "postDate": "2023-06-18T19:24:17.957Z",
              "content": "<p>I'm using <code>import tensorflow.lite as tflite</code> as I fail to import tflite_runtime in notebook. And I tried your code </p>\n<pre><code>!pip install tflite-runtime==\n tflite_runtime.interpreter  tflite\n</code></pre>\n<p>but still can't install tflite-runtime: </p>\n<pre><code>Could not find a version that satisfies the requirement tflite-runtime==2.9.1 (from versions: none)\nNo matching distribution found for tflite-runtime==2.9.1\n---------------------------------------------------------------------------\nModuleNotFoundError                       Traceback (most recent call last)\nCell In[11], line 2\n      1 get_ipython().system('pip install tflite-runtime==2.9.1')\n----&gt; 2 import tflite_runtime.interpreter as tflite\nModuleNotFoundError: No module named 'tflite_runtime'\n</code></pre>\n<p>I think the remaining code is basically the same, my code is like:</p>\n<pre><code>import pandas as pd\nimport tensorflow as tf\nimport glob\nimport json\nimport numpy as np\nimport time\n\n\ndf = pd.read_csv()\n\n\ninterpreter = tf.lite.Interpreter()\n\nREQUIRED_SIGNATURE = \nREQUIRED_OUTPUT = \n\nwith open (, ) as f:\n    character_map = json.load(f)\nrev_character_map = {j:i  i,j  character_map.items()}\n\nfound_signatures = list(interpreter.get_signature_list().keys())\n\n REQUIRED_SIGNATURE   found_signatures:\n    raise KernelEvalException()\n\nwith open (, ) as f:\n    load_cols = json.load(f)\nSEL_COLS = load_cols.()\n\nprediction_fn = interpreter.get_signature_runner()\ntot = time.time()\n i,(path,file_id,sequence_id,participant_id,phrase)  df.iterrows():\n    data = pd.read_parquet(f, =SEL_COLS)\n    start_time = time.time()\n    output = prediction_fn(=data[data.==sequence_id])\n    prediction_str = .join([rev_character_map.(s, )  s  np.argmax(output[REQUIRED_OUTPUT], =1)])\n     i%==0:\n        (, time.time()-start_time, )\n        (, prediction_str)\n        (,phrase)\n        ()\n     ==100:\n         break\n(,(time.time()-tot)/i)\n</code></pre>",
              "rawMarkdown": "I'm using `import tensorflow.lite as tflite` as I fail to import tflite_runtime in notebook. And I tried your code \n```\n!pip install tflite-runtime==2.9.1\nimport tflite_runtime.interpreter as tflite\n```\nbut still can't install tflite-runtime: \n```\nERROR: Could not find a version that satisfies the requirement tflite-runtime==2.9.1 (from versions: none)\nERROR: No matching distribution found for tflite-runtime==2.9.1\n---------------------------------------------------------------------------\nModuleNotFoundError                       Traceback (most recent call last)\nCell In[11], line 2\n      1 get_ipython().system('pip install tflite-runtime==2.9.1')\n----> 2 import tflite_runtime.interpreter as tflite\nModuleNotFoundError: No module named 'tflite_runtime'\n```\n\nI think the remaining code is basically the same, my code is like:\n\n```\nimport pandas as pd\nimport tensorflow as tf\nimport glob\nimport json\nimport numpy as np\nimport time\n# import tflite_runtime.interpreter as tflite\n\ndf = pd.read_csv('/kaggle/input/asl-fingerspelling/train.csv')\n# parquet_files = glob.glob('/kaggle/input/asl-fingerspelling/train_landmarks/*.parquet')\n\ninterpreter = tf.lite.Interpreter(\"/kaggle/input/asl-submission/model.tflite\")\n\nREQUIRED_SIGNATURE = \"serving_default\"\nREQUIRED_OUTPUT = \"outputs\"\n\nwith open (\"/kaggle/input/asl-fingerspelling/character_to_prediction_index.json\", \"r\") as f:\n    character_map = json.load(f)\nrev_character_map = {j:i for i,j in character_map.items()}\n\nfound_signatures = list(interpreter.get_signature_list().keys())\n\nif REQUIRED_SIGNATURE not in found_signatures:\n    raise KernelEvalException('Required input signature not found.')\n\nwith open (\"/kaggle/input/asl-submission/inference_args.json\", \"r\") as f:\n    load_cols = json.load(f)\nSEL_COLS = load_cols.get(\"selected_columns\")\n    \nprediction_fn = interpreter.get_signature_runner(\"serving_default\")\ntot = time.time()\nfor i,(path,file_id,sequence_id,participant_id,phrase) in df.iterrows():\n    data = pd.read_parquet(f'/kaggle/input/asl-fingerspelling/{path}', columns=SEL_COLS)\n    start_time = time.time()\n    output = prediction_fn(inputs=data[data.index==sequence_id])\n    prediction_str = \"\".join([rev_character_map.get(s, \"\") for s in np.argmax(output[REQUIRED_OUTPUT], axis=1)])\n    if i%10==0:\n        print('infer time:', time.time()-start_time, )\n        print('pred:', prediction_str)\n        print('gt:',phrase)\n        print(' ')\n    if i==100:\n         break\nprint('avg time',(time.time()-tot)/i)\n```\n\n"
            },
            {
              "id": 2308315,
              "postDate": "2023-06-18T19:38:18.447Z",
              "content": "<p>and my result of first 100 samples are:</p>\n<pre><code> time (inference plus IO) .\n inference time .\n</code></pre>",
              "rawMarkdown": "and my result of first 100 samples are:\n```\navg time (inference plus IO) 1.2560365319252014\navg inference time 0.5513821411132812\n```"
            },
            {
              "id": 2308349,
              "postDate": "2023-06-18T20:50:04.120Z",
              "content": "<p>Ah, so that's the reason. <br>\nFor me <code>import tensorflow.lite as tflite</code> :</p>\n<ul>\n<li>Mean time: 0.5808348</li>\n<li>Mean time only infer: 0.1342919</li>\n</ul>\n<p>And for <code>import tflite_runtime.interpreter as tflite</code> :</p>\n<ul>\n<li>Mean time: 0.8799638</li>\n<li>Mean time only infer: 0.4468766</li>\n</ul>\n<p>It has very different time. I wrote about <code>import tflite_runtime.interpreter as tflite</code> timings.<br>\nWith this approximation your time with <code>import tflite_runtime.interpreter as tflite</code> might be about 2 seconds</p>\n<p>To install it I use previous python 3.7 environment.</p>\n<p>PS. Mean time only infer: 0.4468766 has submission time ~3:40 , I think that maximum inference time not 0.7, maybe 0.6</p>",
              "rawMarkdown": "Ah, so that's the reason. \nFor me `import tensorflow.lite as tflite` :\n- Mean time: 0.5808348\n- Mean time only infer: 0.1342919\n\nAnd for `import tflite_runtime.interpreter as tflite` :\n- Mean time: 0.8799638\n- Mean time only infer: 0.4468766\n\nIt has very different time. I wrote about `import tflite_runtime.interpreter as tflite` timings.\nWith this approximation your time with `import tflite_runtime.interpreter as tflite` might be about 2 seconds\n\nTo install it I use previous python 3.7 environment.\n\nPS. Mean time only infer: 0.4468766 has submission time ~3:40 , I think that maximum inference time not 0.7, maybe 0.6",
              "votes": 2
            },
            {
              "id": 2308370,
              "postDate": "2023-06-18T21:24:36.750Z",
              "content": "<p>Yes, you're right, there is a huge difference. <code>import tflite_runtime.interpreter as tflite</code> changed my inference time from ~<code>0.55s</code> to ~<code>0.9s</code>. Thank you for your information, really appreciate it!</p>",
              "rawMarkdown": "Yes, you're right, there is a huge difference. `import tflite_runtime.interpreter as tflite` changed my inference time from ~`0.55s` to ~`0.9s`. Thank you for your information, really appreciate it!"
            }
          ]
        },
        {
          "id": 2344529,
          "postDate": "2023-07-14T14:47:26.800Z",
          "content": "<p>ah hello I have a noob question :( . Your submission takes 3 hours to run but we are currently using only 50% of the test files on public board so if we run it full it will take around 6 hours(private + public test). But the contest is limited to 5 hours. Will this affect the final result.</p>",
          "rawMarkdown": "ah hello I have a noob question :( . Your submission takes 3 hours to run but we are currently using only 50% of the test files on public board so if we run it full it will take around 6 hours(private + public test). But the contest is limited to 5 hours. Will this affect the final result.",
          "replies": [
            {
              "id": 2344595,
              "postDate": "2023-07-14T15:21:02.687Z",
              "content": "<p>Hi.<br>\nSubmissions scored on the whole data. But you only see the metric of 50% on LB. So don't worry, it won't affect the final result</p>",
              "rawMarkdown": "Hi.\nSubmissions scored on the whole data. But you only see the metric of 50% on LB. So don't worry, it won't affect the final result"
            },
            {
              "id": 2344613,
              "postDate": "2023-07-14T15:47:27.133Z",
              "content": "<p>ah I got it. Thanks for your answer</p>",
              "rawMarkdown": "ah I got it. Thanks for your answer"
            }
          ]
        }
      ]
    },
    {
      "id": 2276829,
      "postDate": "2023-05-27T08:32:56.703Z",
      "content": "<p>I finally found the problem with submission. </p>\n<p>In previous competition I filled nans with this code. And also I unsqueezed tensor like this. (I changed both lines, didn't try change it one by one)</p>\n<pre><code>x = tf.where(tfnp.isnan(x), , x)\nx = tf.expand_dims(x, axis=)\n</code></pre>\n<p>It worked fine in previous competition. And it work fine here in local kernel. But has error while submission scoring</p>\n<p>I just changed this lines to:</p>\n<pre><code>x = tf.where(tfnp.isnan(x), tf.zeros_like(x), x)\nx = x[]\n</code></pre>\n<p>Maybe it will help someone. Not so obvious reasons, as for me</p>",
      "rawMarkdown": "I finally found the problem with submission. \n\nIn previous competition I filled nans with this code. And also I unsqueezed tensor like this. (I changed both lines, didn't try change it one by one)\n\n```python\nx = tf.where(tfnp.isnan(x), 0.0, x)\nx = tf.expand_dims(x, axis=0)\n```\n\n\nIt worked fine in previous competition. And it work fine here in local kernel. But has error while submission scoring\n\nI just changed this lines to:\n```python\nx = tf.where(tfnp.isnan(x), tf.zeros_like(x), x)\nx = x[None]\n```\n\nMaybe it will help someone. Not so obvious reasons, as for me",
      "votes": 8,
      "replies": [
        {
          "id": 2276976,
          "postDate": "2023-05-27T10:57:17.710Z",
          "content": "<p>Thank you for sharing this, but I am just a beginner and I am a little confused why do we need to Un squeeze or expand dimensions of tensor for this problem??</p>",
          "rawMarkdown": "Thank you for sharing this, but I am just a beginner and I am a little confused why do we need to Un squeeze or expand dimensions of tensor for this problem??",
          "replies": [
            {
              "id": 2277165,
              "postDate": "2023-05-27T13:53:40.527Z",
              "content": "<p>Actually you may need this for most models as batch dimension. But it depends of your architecture and approach. You may use time dimension as batch dimension for example</p>",
              "rawMarkdown": "Actually you may need this for most models as batch dimension. But it depends of your architecture and approach. You may use time dimension as batch dimension for example",
              "votes": 1
            },
            {
              "id": 2277204,
              "postDate": "2023-05-27T14:31:02.450Z",
              "content": "<p>Could you please share your TFLiteModel<br>\n<code>\npython\nclass TFLiteModel(tf.Module):\n</code></p>",
              "rawMarkdown": "Could you please share your TFLiteModel\n`\npython\nclass TFLiteModel(tf.Module):\n`\n"
            },
            {
              "id": 2277238,
              "postDate": "2023-05-27T14:58:36.383Z",
              "content": "<p>After the competition deadline of course if you'll still need it :)</p>",
              "rawMarkdown": "After the competition deadline of course if you'll still need it :)",
              "votes": 2
            },
            {
              "id": 2277338,
              "postDate": "2023-05-27T16:18:24.543Z",
              "content": "<p>yeah I was thinking to just use time dimension as batch dimension, thank you for clarifying</p>",
              "rawMarkdown": "yeah I was thinking to just use time dimension as batch dimension, thank you for clarifying"
            }
          ]
        }
      ]
    },
    {
      "id": 2273788,
      "postDate": "2023-05-25T12:06:40.840Z",
      "content": "<p></p>\n<p><br>\n</p>\n<p><br>\n</p>\n<p>WIth <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/409722#2276829\" target=\"_blank\">fixed</a> preprocessing code it works</p>\n<pre><code>!pip install tflite-runtime==\n tflite_runtime.interpreter  tflite\n\n time\n json\n pandas  pd\n tqdm.auto  tqdm\n Levenshtein  Lev\n tensorflow  tf\n\nconverter = tf.lite.TFLiteConverter.from_keras_model(infr)  \ntflite_model = converter.convert()\n (, )  f:\n    f.write(tflite_model)\n! submission.  \n\n\nSEL_FEATURES = json.load(())[]\n\n ():\n         pd.read_parquet(pq_path, columns=SEL_FEATURES) \n\n  (, )  f:\n    character_map = json.load(f)\nrev_character_map = {j:i  i,j  character_map.items()}\n\n\ndf = pd.read_csv()\n\nidx = \nsample = df.loc[idx]\nloaded = load_relevant_data_subset( + sample[])\nloaded = loaded[loaded.index==sample[]].values\n(loaded.shape)\nframes = loaded\n\n ():\n    w1 = (s1.split())\n    lvd = Lev.distance(s1, s2)\n     lvd / w1\n\n tflite_runtime.interpreter  tflite\ninterpreter = tflite.Interpreter()\nfound_signatures = (interpreter.get_signature_list().keys())\n\nREQUIRED_SIGNATURE = \nREQUIRED_OUTPUT = \n REQUIRED_SIGNATURE   found_signatures:\n     KernelEvalException()\n\nprediction_fn = interpreter.get_signature_runner()\noutput_lite = prediction_fn(inputs=frames)\nprediction_str = .join([rev_character_map.get(s, )  s  np.argmax(output_lite[REQUIRED_OUTPUT], axis=)])\n(prediction_str)\n\n\nst = time.time()\ncnt = \ntotal =  \nmodel_time = \n\nlevs = []\n\n i  tqdm(((df.iloc[:total]))):\n    sample = df.loc[i]\n    loaded = load_relevant_data_subset( + sample[])\n    loaded = loaded[loaded.index==sample[]].values\n\n    md_st = time.time()\n    output_ = prediction_fn(inputs=loaded)\n    model_time += time.time() - md_st\n\n\n    prediction_str = .join([rev_character_map.get(s, )  s  np.argmax(output_[REQUIRED_OUTPUT], axis=)])\n    cur_lev = wer__(sample[], prediction_str) \n    \n    \n\n    levs.append(cur_lev)\n\n()\n()\n()\n</code></pre>",
      "rawMarkdown": "~~I'm using this code to make sure my model works. It looks like absolutely fit with info from `Evaluation` page. May hosts (or some kagglers) please clarify why it has errors after submission? And provide the sample code to make sure that some submission will succeed.\nIt don't depends of inference time, same error with 0.4s Mean Inference time (after 42 minutes) and with 0.02s Mean Inference time (after 7 minutes).~~\n\n~~tflite_runtime version == 2.9.1~~\n~~Model size << 40 MB  (tried even with size 3MB)~~\n\n~~Capable with all inputs shapes [None, len(selected_columns)].  [from 0 to inf)~~\n~~Capable with full nan inputs.~~\n\nWIth [fixed](https://www.kaggle.com/competitions/asl-fingerspelling/discussion/409722#2276829) preprocessing code it works\n\n\n```python\n!pip install tflite-runtime==2.9.1\nimport tflite_runtime.interpreter as tflite\n\nimport time\nimport json\nimport pandas as pd\nfrom tqdm.auto import tqdm\nimport Levenshtein as Lev\nimport tensorflow as tf\n\nconverter = tf.lite.TFLiteConverter.from_keras_model(infr)  # infr it's my TF or Keras model (tried both)\ntflite_model = converter.convert()\nwith open('model.tflite', 'wb') as f:\n    f.write(tflite_model)\n!zip submission.zip 'model.tflite' 'inference_args.json'\n\n\nSEL_FEATURES = json.load(open('/kaggle/working/inference_args.json'))['selected_columns']\n\ndef load_relevant_data_subset(pq_path):\n        return pd.read_parquet(pq_path, columns=SEL_FEATURES) #selected_columns)\n\nwith open (\"/kaggle/input/asl-fingerspelling/character_to_prediction_index.json\", \"r\") as f:\n    character_map = json.load(f)\nrev_character_map = {j:i for i,j in character_map.items()}\n\n\ndf = pd.read_csv('/kaggle/input/asl-fingerspelling/train.csv')\n\nidx = 0\nsample = df.loc[idx]\nloaded = load_relevant_data_subset('/kaggle/input/asl-fingerspelling/' + sample['path'])\nloaded = loaded[loaded.index==sample['sequence_id']].values\nprint(loaded.shape)\nframes = loaded\n\ndef wer__(s1, s2):\n    w1 = len(s1.split())\n    lvd = Lev.distance(s1, s2)\n    return lvd / w1\n\nimport tflite_runtime.interpreter as tflite\ninterpreter = tflite.Interpreter('model.tflite')\nfound_signatures = list(interpreter.get_signature_list().keys())\n\nREQUIRED_SIGNATURE = 'serving_default'\nREQUIRED_OUTPUT = 'outputs'\nif REQUIRED_SIGNATURE not in found_signatures:\n    raise KernelEvalException('Required input signature not found.')\n    \nprediction_fn = interpreter.get_signature_runner(\"serving_default\")\noutput_lite = prediction_fn(inputs=frames)\nprediction_str = \"\".join([rev_character_map.get(s, \"\") for s in np.argmax(output_lite[REQUIRED_OUTPUT], axis=1)])\nprint(prediction_str)\n\n\nst = time.time()\ncnt = 0\ntotal = 100 #len(df)\nmodel_time = 0\n\nlevs = []\n\nfor i in tqdm(range(len(df.iloc[:total]))):\n    sample = df.loc[i]\n    loaded = load_relevant_data_subset('/kaggle/input/asl-fingerspelling/' + sample['path'])\n    loaded = loaded[loaded.index==sample['sequence_id']].values\n\n    md_st = time.time()\n    output_ = prediction_fn(inputs=loaded)\n    model_time += time.time() - md_st\n\n\n    prediction_str = \"\".join([rev_character_map.get(s, \"\") for s in np.argmax(output_[REQUIRED_OUTPUT], axis=1)])\n    cur_lev = wer__(sample['phrase'], prediction_str) \n    #print(sample['phrase'], '|', prediction_str, '|', cur_lev)\n    #print()\n\n    levs.append(cur_lev)\n\nprint(f'WER: {np.mean(levs):.5f}')\nprint(f'Mean time: {(time.time() - st)/total:.7f}')\nprint(f'Mean time only infer: {model_time/total:.7f}')\n```\n",
      "votes": 2
    },
    {
      "id": 2324330,
      "postDate": "2023-06-30T14:34:15.077Z",
      "content": "<p>I see public notebook run 5-6 hours, where is 5 hour limit? Notebook still has 9 hour limit submissiong? A bit confused here.</p>",
      "rawMarkdown": "I see public notebook run 5-6 hours, where is 5 hour limit? Notebook still has 9 hour limit submissiong? A bit confused here.",
      "replies": [
        {
          "id": 2324390,
          "postDate": "2023-06-30T15:10:52.567Z",
          "content": "<p>Do you mean kernel runtime or scoring time?<br>\n5 hours limit only for scoring time</p>",
          "rawMarkdown": "Do you mean kernel runtime or scoring time?\n5 hours limit only for scoring time",
          "votes": 2
        }
      ]
    },
    {
      "id": 2306719,
      "postDate": "2023-06-17T14:02:22.950Z",
      "content": "<p>Sorry to bother you.</p>\n<p>My model.tflite works wll in the Notebook with tensorflow 2.11.0. But it gets Submission Scoring Error soon after submit.</p>\n<p>My model.tflite is converted from a PyTorch model.</p>\n<p>Should I use tensorflow 2.12 or 2.14?  Is there a safe way to convert torch model to tflite?</p>\n<p>If you were me, what is your choice?</p>",
      "rawMarkdown": "Sorry to bother you.\n\nMy model.tflite works wll in the Notebook with tensorflow 2.11.0. But it gets Submission Scoring Error soon after submit.\n\nMy model.tflite is converted from a PyTorch model.\n\nShould I use tensorflow 2.12 or 2.14?  Is there a safe way to convert torch model to tflite?\n\nIf you were me, what is your choice?",
      "replies": [
        {
          "id": 2307042,
          "postDate": "2023-06-17T19:00:46.417Z",
          "content": "<p>Hi, I just keep using old environment with </p>\n<ul>\n<li>python 3.7</li>\n<li>tensorflow 2.11.0</li>\n<li>tflite-runtime==2.9.1</li>\n</ul>\n<p>Works fine for me</p>",
          "rawMarkdown": "Hi, I just keep using old environment with \n- python 3.7\n- tensorflow 2.11.0\n-  tflite-runtime==2.9.1\n\nWorks fine for me",
          "votes": 1,
          "replies": [
            {
              "id": 2307205,
              "postDate": "2023-06-18T02:00:09.073Z",
              "content": "<p>Thank you, The information you shared helped me a lot.</p>",
              "rawMarkdown": "Thank you, The information you shared helped me a lot."
            },
            {
              "id": 2362828,
              "postDate": "2023-07-28T10:24:03.803Z",
              "content": "<p>Sorry to bother you.<br>\nhow do you keep using python 3.7 ig now i can only pin it to the latest version ?<br>\nany workarounds without 3.7?? or to get 3.7?</p>",
              "rawMarkdown": "Sorry to bother you.\nhow do you keep using python 3.7 ig now i can only pin it to the latest version ?\nany workarounds without 3.7?? or to get 3.7?\n"
            },
            {
              "id": 2362842,
              "postDate": "2023-07-28T10:31:02.447Z",
              "content": "<p>I just copied some public notebook with this environment. Don't actually remember which one. You could try to find it in early stage discussions</p>\n<p>I'm also like you don't know how to change environment to any other except last :)</p>",
              "rawMarkdown": "I just copied some public notebook with this environment. Don't actually remember which one. You could try to find it in early stage discussions\n\nI'm also like you don't know how to change environment to any other except last :)"
            },
            {
              "id": 2362932,
              "postDate": "2023-07-28T11:40:52.547Z",
              "content": "<p>oooh thanks !</p>",
              "rawMarkdown": "oooh thanks !"
            }
          ]
        }
      ]
    },
    {
      "id": 2277083,
      "postDate": "2023-05-27T12:39:32.250Z",
      "content": "<p><a href=\"https://www.kaggle.com/kolyaforrat\" target=\"_blank\">@kolyaforrat</a> It's bit off topic but lb0.734 is awesome! May I ask what's your validation score? </p>",
      "rawMarkdown": "@kolyaforrat It's bit off topic but lb0.734 is awesome! May I ask what's your validation score? ",
      "replies": [
        {
          "id": 2277169,
          "postDate": "2023-05-27T13:55:16.937Z",
          "content": "<p>Valid score: 0.743 | LB score: 0.734</p>",
          "rawMarkdown": "Valid score: 0.743 | LB score: 0.734",
          "votes": 2,
          "replies": [
            {
              "id": 2277198,
              "postDate": "2023-05-27T14:25:32.073Z",
              "content": "<p>Wow, thanks for sharing!<br>\nI need to reproduce your solution at prev competition at first;)</p>",
              "rawMarkdown": "Wow, thanks for sharing!\nI need to reproduce your solution at prev competition at first;)"
            }
          ]
        }
      ]
    },
    {
      "id": 2272974,
      "postDate": "2023-05-24T22:17:09.210Z",
      "content": "<p>After submissions became available I think this topic is actual again.</p>\n<p>I have few submits with ~40-42 minutes with <code>Submission Scoring Error</code>. But:</p>\n<blockquote>\n  <p>Your model must also perform inference in less than 5 hours </p>\n</blockquote>\n<p>Or maybe there are some kind of multiprocessing and the time cap is 45 minutes?</p>\n<p>Does anyone have successful submissions longer than 50 minutes? Any ideas about mean per sample inference time?</p>\n<p>upd. Error in my submission not because of 'out of time' . I deleted most of layers from my model and it has Error at ~8 minutes. <br>\nSo I don't understand the problem. My input is <code>(time_dim, len(selected_columns))</code> and output with (new_time_dim, 59) and still getting errors. <br>\nModel under 40 MB, capable with tflite, same environment with successful public kernels. Maybe here some another limitations I didn't know? At Kaggle kernels it works fine with Evaluation example code 😨</p>",
      "rawMarkdown": "After submissions became available I think this topic is actual again.\n\nI have few submits with ~40-42 minutes with `Submission Scoring Error`. But:\n>Your model must also perform inference in less than 5 hours \n\nOr maybe there are some kind of multiprocessing and the time cap is 45 minutes?\n\n\nDoes anyone have successful submissions longer than 50 minutes? Any ideas about mean per sample inference time?\n\nupd. Error in my submission not because of 'out of time' . I deleted most of layers from my model and it has Error at ~8 minutes. \nSo I don't understand the problem. My input is `(time_dim, len(selected_columns))` and output with (new_time_dim, 59) and still getting errors. \nModel under 40 MB, capable with tflite, same environment with successful public kernels. Maybe here some another limitations I didn't know? At Kaggle kernels it works fine with Evaluation example code 😨",
      "replies": [
        {
          "id": 2272978,
          "postDate": "2023-05-24T22:33:55.090Z",
          "content": "<p>the only think I was capable of submitting were dummy submissions not processing the frames.</p>",
          "rawMarkdown": "the only think I was capable of submitting were dummy submissions not processing the frames."
        },
        {
          "id": 2273492,
          "postDate": "2023-05-25T07:33:02.570Z",
          "content": "<p>if you click on the triangle with error - do you get any more information under Submission Scoring Error in the Submission Details popup?  e.g. \"Your notebook generated a submission file with incorrect format.\"  or other message?  Know usually very little is give when submissions fail.</p>\n<p>I can confirm using inference_args.json for selected columns works ok, but not if you only were using frame column. Tried that for some simple tests.  Had issues with preprocessing layer though so far. </p>\n<p>If for some reason your output is empty, you can have error for not an iterator here -<br>\n<code>prediction_str = \"\".join([rev_character_map.get(s, \"\") for s in np.argmax(output[REQUIRED_OUTPUT], axis=1)])</code><br>\nSo perhaps a default output can be returned like a space for exceptions.</p>\n<p>Think there are still open questions about what is provided in the video via <br>\n<code>def load_relevant_data_subset(pq_path):</code><br>\nis it a parquet file of many sequence ids like in train? or one parquet file for each sequence id to predict like in GISLR?<br>\nIf the former, then is the expectation to split the output strings per sequence id? In train there is one parquet that has 287 sequence ids ( /kaggle/input/asl-fingerspelling/train_landmarks/450474571.parquet ), the rest have 1000 but perhaps in test all have 1000.  </p>\n<p><a href=\"https://www.kaggle.com/georgemegre\" target=\"_blank\">@georgemegre</a> seems to have been able to do more than dummy submissions.  Without sharing too many details maybe has some thoughts on tips for true submissions under 5hrs?   Thanks in advance. </p>",
          "rawMarkdown": "if you click on the triangle with error - do you get any more information under Submission Scoring Error in the Submission Details popup?  e.g. \"Your notebook generated a submission file with incorrect format.\"  or other message?  Know usually very little is give when submissions fail.\n\nI can confirm using inference_args.json for selected columns works ok, but not if you only were using frame column. Tried that for some simple tests.  Had issues with preprocessing layer though so far. \n\nIf for some reason your output is empty, you can have error for not an iterator here -\n`prediction_str = \"\".join([rev_character_map.get(s, \"\") for s in np.argmax(output[REQUIRED_OUTPUT], axis=1)])`\nSo perhaps a default output can be returned like a space for exceptions.\n \nThink there are still open questions about what is provided in the video via \n`def load_relevant_data_subset(pq_path):`\nis it a parquet file of many sequence ids like in train? or one parquet file for each sequence id to predict like in GISLR?\nIf the former, then is the expectation to split the output strings per sequence id? In train there is one parquet that has 287 sequence ids ( /kaggle/input/asl-fingerspelling/train_landmarks/450474571.parquet ), the rest have 1000 but perhaps in test all have 1000.  \n\n@georgemegre seems to have been able to do more than dummy submissions.  Without sharing too many details maybe has some thoughts on tips for true submissions under 5hrs?   Thanks in advance. ",
          "replies": [
            {
              "id": 2273547,
              "postDate": "2023-05-25T08:14:38.327Z",
              "content": "<p>Error message has nothing helpful, this message was prepared for submitting csv files and just told about rows and columns.</p>\n<blockquote>\n  <p>If for some reason your output is empty, you can have error</p>\n</blockquote>\n<p>I tested empty input with time_dim=0 and with full nan input. There are no way that my model has zero or empty output. Only if selected columns are wrong, but this scoring for about 40 minutes I think everything is ok with this part.</p>\n<p>Also my model capable for any length of input. For single frame or multiple frames from one csv. Output phrases would be just as one big string if so.</p>",
              "rawMarkdown": "Error message has nothing helpful, this message was prepared for submitting csv files and just told about rows and columns.\n\n> If for some reason your output is empty, you can have error\n\nI tested empty input with time_dim=0 and with full nan input. There are no way that my model has zero or empty output. Only if selected columns are wrong, but this scoring for about 40 minutes I think everything is ok with this part.\n\nAlso my model capable for any length of input. For single frame or multiple frames from one csv. Output phrases would be just as one big string if so."
            },
            {
              "id": 2274122,
              "postDate": "2023-05-25T16:27:15.990Z",
              "content": "<p>My submission run about 1 hour, the pipeline is very similar to previous competition, except that output is (time, 59). It also can be a problem in tflite version, it should be 2.9.1. i've checked my sub with this code (hope this will helpful):</p>\n<pre><code>python\n ():\n    df = pd.read_parquet(pq_path, columns=selected_columns)\n    df = df.loc[]\n     df.values\n\n (json_fn)  f:\n    cols = json.load(f)[]\n\nframes = load_relevant_data_subset(, cols)\n\ninterpreter = tf.lite.Interpreter(sub_model_path)\n\nprediction_fn = interpreter.get_signature_runner()\nt = time()\noutput = prediction_fn(inputs=frames)\n()\n(*(time() - t) )\nt = time()\noutput = prediction_fn(inputs=frames)\n()\n(*(time() - t) )\n\nout_2 = output[]\n(out_2.shape)\n(np.argmax(out_2, axis=))\n\n  (, )  f:\n    character_map = json.load(f)\nrev_character_map = {j:i  i,j  character_map.items()}\nprediction_str = .join([rev_character_map.get(s, )  s  np.argmax(out_2, axis=)])\nprediction_str\n</code></pre>",
              "rawMarkdown": "My submission run about 1 hour, the pipeline is very similar to previous competition, except that output is (time, 59). It also can be a problem in tflite version, it should be 2.9.1. i've checked my sub with this code (hope this will helpful):\n\n```python\npython\ndef load_relevant_data_subset(pq_path, selected_columns):\n    df = pd.read_parquet(pq_path, columns=selected_columns)\n    df = df.loc[1816909464]\n    return df.values\n\nwith open(json_fn) as f:\n    cols = json.load(f)['selected_columns']\n\nframes = load_relevant_data_subset('/kaggle/input/asl-fingerspelling/train_landmarks/5414471.parquet', cols)\n#interpreter = tflite.Interpreter(sub_model_path)\ninterpreter = tf.lite.Interpreter(sub_model_path)\n\nprediction_fn = interpreter.get_signature_runner(\"serving_default\")\nt = time()\noutput = prediction_fn(inputs=frames)\nprint(\"Time\")\nprint(1000*(time() - t) )\nt = time()\noutput = prediction_fn(inputs=frames)\nprint(\"Time\")\nprint(1000*(time() - t) )\n\nout_2 = output['outputs']\nprint(out_2.shape)\nprint(np.argmax(out_2, axis=1))\n\nwith open (\"/kaggle/input/asl-fingerspelling/character_to_prediction_index.json\", \"r\") as f:\n    character_map = json.load(f)\nrev_character_map = {j:i for i,j in character_map.items()}\nprediction_str = \"\".join([rev_character_map.get(s, \"\") for s in np.argmax(out_2, axis=1)])\nprediction_str\n```\n",
              "votes": 1
            },
            {
              "id": 2274131,
              "postDate": "2023-05-25T16:34:11.310Z",
              "content": "<p>Thanks for info. Could you please share <code>frames</code> shape and <code>out_2</code> shape? They are same by time axis or different?</p>",
              "rawMarkdown": "Thanks for info. Could you please share `frames` shape and `out_2` shape? They are same by time axis or different?\n\n"
            },
            {
              "id": 2274136,
              "postDate": "2023-05-25T16:38:53.263Z",
              "content": "<p>hm, to be honest i don't understand how you calculate it, but i can explain my understanding by two examples (hope this will be helpful):<br>\n1)<br>\ny_true =  my bare face in the wind<br>\ny_pred = my bare face sa the wind<br>\nN = 24<br>\nLEVD = 2<br>\nMetric = 22/24 = ~0.91</p>\n<p>But if we have:<br>\ny_true =  my bare face in the wind<br>\ny_pred = \"sssssssssssssssssssssssss wdaedaj\"<br>\nN = 24<br>\nLEVD = 30<br>\nMetric = 24-30 / 24 = -0.25</p>\n<p>So, it should be in the range (-d_min, 1).</p>",
              "rawMarkdown": "hm, to be honest i don't understand how you calculate it, but i can explain my understanding by two examples (hope this will be helpful):\n1)\ny_true =  my bare face in the wind\ny_pred = my bare face sa the wind\nN = 24\nLEVD = 2\nMetric = 22/24 = ~0.91\n\nBut if we have:\ny_true =  my bare face in the wind\ny_pred = \"sssssssssssssssssssssssss wdaedaj\"\nN = 24\nLEVD = 30\nMetric = 24-30 / 24 = -0.25\n\nSo, it should be in the range (-d_min, 1).",
              "votes": 1
            },
            {
              "id": 2274139,
              "postDate": "2023-05-25T16:41:15.287Z",
              "content": "<p>Sure, here it is:<br>\nframes shape: (236, 252)<br>\nout_2 shape: (19, 59)</p>",
              "rawMarkdown": "Sure, here it is:\nframes shape: (236, 252)\nout_2 shape: (19, 59)",
              "votes": 1
            },
            {
              "id": 2274148,
              "postDate": "2023-05-25T16:46:51.337Z",
              "content": "<p>Thanks again. So now I'm totally don't understand why I can't submit my model)</p>",
              "rawMarkdown": "Thanks again. So now I'm totally don't understand why I can't submit my model)"
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2276907,
      "author_name": "Kolya Forrat",
      "author_url": "",
      "post_date": "2023-05-27T09:56:56.353000",
      "content": "<p>I also calculated the approximate time of one sample inference. <br>\nSubmission scoring lasts ~3hours, mean inference time ~0.42 seconds. So:<br>\n<code>3*3600 seconds / 0.42 ~= 25714</code> samples</p>\n<p>We have 5 hours limitation. With this logic approximate maximum Mean Only Inference Time for model ~ 0.7 seconds</p>",
      "votes": 8,
      "replies": [
        {
          "id": 2307211,
          "author_name": "Yu Wu",
          "author_url": "",
          "post_date": "2023-06-18T02:07:40.117000",
          "content": "<p>Well, I used the Kaggle notebook to test my model's inference speed, and my model with ~0.6s  inference time exceeds the time limit ……</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2307955,
              "author_name": "Kolya Forrat",
              "author_url": "",
              "post_date": "2023-06-18T14:40:58.893000",
              "content": "<p>Maybe it depends of which subset of data do you use for speed test? And which tflite import you are using<code>import tflite_runtime.interpreter as tflite</code> or <code>import tensorflow.lite as tflite</code><br>\n<a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/409722#2273788\" target=\"_blank\">Below</a> I shared the way I tested it. For first 1000 samples from train.csv</p>\n<p>Anyways, my calculations are approximate. And if you for example share your results for the same subset we can make some more accurate prediction of maximum inference time.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2308304,
              "author_name": "Yu Wu",
              "author_url": "",
              "post_date": "2023-06-18T19:24:17.957000",
              "content": "<p>I'm using <code>import tensorflow.lite as tflite</code> as I fail to import tflite_runtime in notebook. And I tried your code </p>\n<pre><code>!pip install tflite-runtime==\n tflite_runtime.interpreter  tflite\n</code></pre>\n<p>but still can't install tflite-runtime: </p>\n<pre><code>Could not find a version that satisfies the requirement tflite-runtime==2.9.1 (from versions: none)\nNo matching distribution found for tflite-runtime==2.9.1\n---------------------------------------------------------------------------\nModuleNotFoundError                       Traceback (most recent call last)\nCell In[11], line 2\n      1 get_ipython().system('pip install tflite-runtime==2.9.1')\n----&gt; 2 import tflite_runtime.interpreter as tflite\nModuleNotFoundError: No module named 'tflite_runtime'\n</code></pre>\n<p>I think the remaining code is basically the same, my code is like:</p>\n<pre><code>import pandas as pd\nimport tensorflow as tf\nimport glob\nimport json\nimport numpy as np\nimport time\n\n\ndf = pd.read_csv()\n\n\ninterpreter = tf.lite.Interpreter()\n\nREQUIRED_SIGNATURE = \nREQUIRED_OUTPUT = \n\nwith open (, ) as f:\n    character_map = json.load(f)\nrev_character_map = {j:i  i,j  character_map.items()}\n\nfound_signatures = list(interpreter.get_signature_list().keys())\n\n REQUIRED_SIGNATURE   found_signatures:\n    raise KernelEvalException()\n\nwith open (, ) as f:\n    load_cols = json.load(f)\nSEL_COLS = load_cols.()\n\nprediction_fn = interpreter.get_signature_runner()\ntot = time.time()\n i,(path,file_id,sequence_id,participant_id,phrase)  df.iterrows():\n    data = pd.read_parquet(f, =SEL_COLS)\n    start_time = time.time()\n    output = prediction_fn(=data[data.==sequence_id])\n    prediction_str = .join([rev_character_map.(s, )  s  np.argmax(output[REQUIRED_OUTPUT], =1)])\n     i%==0:\n        (, time.time()-start_time, )\n        (, prediction_str)\n        (,phrase)\n        ()\n     ==100:\n         break\n(,(time.time()-tot)/i)\n</code></pre>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2308315,
              "author_name": "Yu Wu",
              "author_url": "",
              "post_date": "2023-06-18T19:38:18.447000",
              "content": "<p>and my result of first 100 samples are:</p>\n<pre><code> time (inference plus IO) .\n inference time .\n</code></pre>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2308349,
              "author_name": "Kolya Forrat",
              "author_url": "",
              "post_date": "2023-06-18T20:50:04.120000",
              "content": "<p>Ah, so that's the reason. <br>\nFor me <code>import tensorflow.lite as tflite</code> :</p>\n<ul>\n<li>Mean time: 0.5808348</li>\n<li>Mean time only infer: 0.1342919</li>\n</ul>\n<p>And for <code>import tflite_runtime.interpreter as tflite</code> :</p>\n<ul>\n<li>Mean time: 0.8799638</li>\n<li>Mean time only infer: 0.4468766</li>\n</ul>\n<p>It has very different time. I wrote about <code>import tflite_runtime.interpreter as tflite</code> timings.<br>\nWith this approximation your time with <code>import tflite_runtime.interpreter as tflite</code> might be about 2 seconds</p>\n<p>To install it I use previous python 3.7 environment.</p>\n<p>PS. Mean time only infer: 0.4468766 has submission time ~3:40 , I think that maximum inference time not 0.7, maybe 0.6</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2308370,
              "author_name": "Yu Wu",
              "author_url": "",
              "post_date": "2023-06-18T21:24:36.750000",
              "content": "<p>Yes, you're right, there is a huge difference. <code>import tflite_runtime.interpreter as tflite</code> changed my inference time from ~<code>0.55s</code> to ~<code>0.9s</code>. Thank you for your information, really appreciate it!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2344529,
          "author_name": "quan",
          "author_url": "",
          "post_date": "2023-07-14T14:47:26.800000",
          "content": "<p>ah hello I have a noob question :( . Your submission takes 3 hours to run but we are currently using only 50% of the test files on public board so if we run it full it will take around 6 hours(private + public test). But the contest is limited to 5 hours. Will this affect the final result.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2344595,
              "author_name": "Kolya Forrat",
              "author_url": "",
              "post_date": "2023-07-14T15:21:02.687000",
              "content": "<p>Hi.<br>\nSubmissions scored on the whole data. But you only see the metric of 50% on LB. So don't worry, it won't affect the final result</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2344613,
              "author_name": "quan",
              "author_url": "",
              "post_date": "2023-07-14T15:47:27.133000",
              "content": "<p>ah I got it. Thanks for your answer</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2276829,
      "author_name": "Kolya Forrat",
      "author_url": "",
      "post_date": "2023-05-27T08:32:56.703000",
      "content": "<p>I finally found the problem with submission. </p>\n<p>In previous competition I filled nans with this code. And also I unsqueezed tensor like this. (I changed both lines, didn't try change it one by one)</p>\n<pre><code>x = tf.where(tfnp.isnan(x), , x)\nx = tf.expand_dims(x, axis=)\n</code></pre>\n<p>It worked fine in previous competition. And it work fine here in local kernel. But has error while submission scoring</p>\n<p>I just changed this lines to:</p>\n<pre><code>x = tf.where(tfnp.isnan(x), tf.zeros_like(x), x)\nx = x[]\n</code></pre>\n<p>Maybe it will help someone. Not so obvious reasons, as for me</p>",
      "votes": 8,
      "replies": [
        {
          "id": 2276976,
          "author_name": "Aaditya Bansal",
          "author_url": "",
          "post_date": "2023-05-27T10:57:17.710000",
          "content": "<p>Thank you for sharing this, but I am just a beginner and I am a little confused why do we need to Un squeeze or expand dimensions of tensor for this problem??</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2277165,
              "author_name": "Kolya Forrat",
              "author_url": "",
              "post_date": "2023-05-27T13:53:40.527000",
              "content": "<p>Actually you may need this for most models as batch dimension. But it depends of your architecture and approach. You may use time dimension as batch dimension for example</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2277204,
              "author_name": "Karthikeyan",
              "author_url": "",
              "post_date": "2023-05-27T14:31:02.450000",
              "content": "<p>Could you please share your TFLiteModel<br>\n<code>\npython\nclass TFLiteModel(tf.Module):\n</code></p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2277238,
              "author_name": "Kolya Forrat",
              "author_url": "",
              "post_date": "2023-05-27T14:58:36.383000",
              "content": "<p>After the competition deadline of course if you'll still need it :)</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2277338,
              "author_name": "Aaditya Bansal",
              "author_url": "",
              "post_date": "2023-05-27T16:18:24.543000",
              "content": "<p>yeah I was thinking to just use time dimension as batch dimension, thank you for clarifying</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2273788,
      "author_name": "Kolya Forrat",
      "author_url": "",
      "post_date": "2023-05-25T12:06:40.840000",
      "content": "<p></p>\n<p><br>\n</p>\n<p><br>\n</p>\n<p>WIth <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/409722#2276829\" target=\"_blank\">fixed</a> preprocessing code it works</p>\n<pre><code>!pip install tflite-runtime==\n tflite_runtime.interpreter  tflite\n\n time\n json\n pandas  pd\n tqdm.auto  tqdm\n Levenshtein  Lev\n tensorflow  tf\n\nconverter = tf.lite.TFLiteConverter.from_keras_model(infr)  \ntflite_model = converter.convert()\n (, )  f:\n    f.write(tflite_model)\n! submission.  \n\n\nSEL_FEATURES = json.load(())[]\n\n ():\n         pd.read_parquet(pq_path, columns=SEL_FEATURES) \n\n  (, )  f:\n    character_map = json.load(f)\nrev_character_map = {j:i  i,j  character_map.items()}\n\n\ndf = pd.read_csv()\n\nidx = \nsample = df.loc[idx]\nloaded = load_relevant_data_subset( + sample[])\nloaded = loaded[loaded.index==sample[]].values\n(loaded.shape)\nframes = loaded\n\n ():\n    w1 = (s1.split())\n    lvd = Lev.distance(s1, s2)\n     lvd / w1\n\n tflite_runtime.interpreter  tflite\ninterpreter = tflite.Interpreter()\nfound_signatures = (interpreter.get_signature_list().keys())\n\nREQUIRED_SIGNATURE = \nREQUIRED_OUTPUT = \n REQUIRED_SIGNATURE   found_signatures:\n     KernelEvalException()\n\nprediction_fn = interpreter.get_signature_runner()\noutput_lite = prediction_fn(inputs=frames)\nprediction_str = .join([rev_character_map.get(s, )  s  np.argmax(output_lite[REQUIRED_OUTPUT], axis=)])\n(prediction_str)\n\n\nst = time.time()\ncnt = \ntotal =  \nmodel_time = \n\nlevs = []\n\n i  tqdm(((df.iloc[:total]))):\n    sample = df.loc[i]\n    loaded = load_relevant_data_subset( + sample[])\n    loaded = loaded[loaded.index==sample[]].values\n\n    md_st = time.time()\n    output_ = prediction_fn(inputs=loaded)\n    model_time += time.time() - md_st\n\n\n    prediction_str = .join([rev_character_map.get(s, )  s  np.argmax(output_[REQUIRED_OUTPUT], axis=)])\n    cur_lev = wer__(sample[], prediction_str) \n    \n    \n\n    levs.append(cur_lev)\n\n()\n()\n()\n</code></pre>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2324330,
      "author_name": "gezi",
      "author_url": "",
      "post_date": "2023-06-30T14:34:15.077000",
      "content": "<p>I see public notebook run 5-6 hours, where is 5 hour limit? Notebook still has 9 hour limit submissiong? A bit confused here.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2324390,
          "author_name": "Kolya Forrat",
          "author_url": "",
          "post_date": "2023-06-30T15:10:52.567000",
          "content": "<p>Do you mean kernel runtime or scoring time?<br>\n5 hours limit only for scoring time</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2306719,
      "author_name": "xbit",
      "author_url": "",
      "post_date": "2023-06-17T14:02:22.950000",
      "content": "<p>Sorry to bother you.</p>\n<p>My model.tflite works wll in the Notebook with tensorflow 2.11.0. But it gets Submission Scoring Error soon after submit.</p>\n<p>My model.tflite is converted from a PyTorch model.</p>\n<p>Should I use tensorflow 2.12 or 2.14?  Is there a safe way to convert torch model to tflite?</p>\n<p>If you were me, what is your choice?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2307042,
          "author_name": "Kolya Forrat",
          "author_url": "",
          "post_date": "2023-06-17T19:00:46.417000",
          "content": "<p>Hi, I just keep using old environment with </p>\n<ul>\n<li>python 3.7</li>\n<li>tensorflow 2.11.0</li>\n<li>tflite-runtime==2.9.1</li>\n</ul>\n<p>Works fine for me</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2307205,
              "author_name": "xbit",
              "author_url": "",
              "post_date": "2023-06-18T02:00:09.073000",
              "content": "<p>Thank you, The information you shared helped me a lot.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2362828,
              "author_name": "blurrylogic",
              "author_url": "",
              "post_date": "2023-07-28T10:24:03.803000",
              "content": "<p>Sorry to bother you.<br>\nhow do you keep using python 3.7 ig now i can only pin it to the latest version ?<br>\nany workarounds without 3.7?? or to get 3.7?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2362842,
              "author_name": "Kolya Forrat",
              "author_url": "",
              "post_date": "2023-07-28T10:31:02.447000",
              "content": "<p>I just copied some public notebook with this environment. Don't actually remember which one. You could try to find it in early stage discussions</p>\n<p>I'm also like you don't know how to change environment to any other except last :)</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2362932,
              "author_name": "blurrylogic",
              "author_url": "",
              "post_date": "2023-07-28T11:40:52.547000",
              "content": "<p>oooh thanks !</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2277083,
      "author_name": "Camaro",
      "author_url": "",
      "post_date": "2023-05-27T12:39:32.250000",
      "content": "<p><a href=\"https://www.kaggle.com/kolyaforrat\" target=\"_blank\">@kolyaforrat</a> It's bit off topic but lb0.734 is awesome! May I ask what's your validation score? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2277169,
          "author_name": "Kolya Forrat",
          "author_url": "",
          "post_date": "2023-05-27T13:55:16.937000",
          "content": "<p>Valid score: 0.743 | LB score: 0.734</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2277198,
              "author_name": "Camaro",
              "author_url": "",
              "post_date": "2023-05-27T14:25:32.073000",
              "content": "<p>Wow, thanks for sharing!<br>\nI need to reproduce your solution at prev competition at first;)</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2272974,
      "author_name": "Kolya Forrat",
      "author_url": "",
      "post_date": "2023-05-24T22:17:09.210000",
      "content": "<p>After submissions became available I think this topic is actual again.</p>\n<p>I have few submits with ~40-42 minutes with <code>Submission Scoring Error</code>. But:</p>\n<blockquote>\n  <p>Your model must also perform inference in less than 5 hours </p>\n</blockquote>\n<p>Or maybe there are some kind of multiprocessing and the time cap is 45 minutes?</p>\n<p>Does anyone have successful submissions longer than 50 minutes? Any ideas about mean per sample inference time?</p>\n<p>upd. Error in my submission not because of 'out of time' . I deleted most of layers from my model and it has Error at ~8 minutes. <br>\nSo I don't understand the problem. My input is <code>(time_dim, len(selected_columns))</code> and output with (new_time_dim, 59) and still getting errors. <br>\nModel under 40 MB, capable with tflite, same environment with successful public kernels. Maybe here some another limitations I didn't know? At Kaggle kernels it works fine with Evaluation example code 😨</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2272978,
          "author_name": "Jon",
          "author_url": "",
          "post_date": "2023-05-24T22:33:55.090000",
          "content": "<p>the only think I was capable of submitting were dummy submissions not processing the frames.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2273492,
          "author_name": "something4kag",
          "author_url": "",
          "post_date": "2023-05-25T07:33:02.570000",
          "content": "<p>if you click on the triangle with error - do you get any more information under Submission Scoring Error in the Submission Details popup?  e.g. \"Your notebook generated a submission file with incorrect format.\"  or other message?  Know usually very little is give when submissions fail.</p>\n<p>I can confirm using inference_args.json for selected columns works ok, but not if you only were using frame column. Tried that for some simple tests.  Had issues with preprocessing layer though so far. </p>\n<p>If for some reason your output is empty, you can have error for not an iterator here -<br>\n<code>prediction_str = \"\".join([rev_character_map.get(s, \"\") for s in np.argmax(output[REQUIRED_OUTPUT], axis=1)])</code><br>\nSo perhaps a default output can be returned like a space for exceptions.</p>\n<p>Think there are still open questions about what is provided in the video via <br>\n<code>def load_relevant_data_subset(pq_path):</code><br>\nis it a parquet file of many sequence ids like in train? or one parquet file for each sequence id to predict like in GISLR?<br>\nIf the former, then is the expectation to split the output strings per sequence id? In train there is one parquet that has 287 sequence ids ( /kaggle/input/asl-fingerspelling/train_landmarks/450474571.parquet ), the rest have 1000 but perhaps in test all have 1000.  </p>\n<p><a href=\"https://www.kaggle.com/georgemegre\" target=\"_blank\">@georgemegre</a> seems to have been able to do more than dummy submissions.  Without sharing too many details maybe has some thoughts on tips for true submissions under 5hrs?   Thanks in advance. </p>",
          "votes": 0,
          "replies": [
            {
              "id": 2273547,
              "author_name": "Kolya Forrat",
              "author_url": "",
              "post_date": "2023-05-25T08:14:38.327000",
              "content": "<p>Error message has nothing helpful, this message was prepared for submitting csv files and just told about rows and columns.</p>\n<blockquote>\n  <p>If for some reason your output is empty, you can have error</p>\n</blockquote>\n<p>I tested empty input with time_dim=0 and with full nan input. There are no way that my model has zero or empty output. Only if selected columns are wrong, but this scoring for about 40 minutes I think everything is ok with this part.</p>\n<p>Also my model capable for any length of input. For single frame or multiple frames from one csv. Output phrases would be just as one big string if so.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2274122,
              "author_name": "George Megre",
              "author_url": "",
              "post_date": "2023-05-25T16:27:15.990000",
              "content": "<p>My submission run about 1 hour, the pipeline is very similar to previous competition, except that output is (time, 59). It also can be a problem in tflite version, it should be 2.9.1. i've checked my sub with this code (hope this will helpful):</p>\n<pre><code>python\n ():\n    df = pd.read_parquet(pq_path, columns=selected_columns)\n    df = df.loc[]\n     df.values\n\n (json_fn)  f:\n    cols = json.load(f)[]\n\nframes = load_relevant_data_subset(, cols)\n\ninterpreter = tf.lite.Interpreter(sub_model_path)\n\nprediction_fn = interpreter.get_signature_runner()\nt = time()\noutput = prediction_fn(inputs=frames)\n()\n(*(time() - t) )\nt = time()\noutput = prediction_fn(inputs=frames)\n()\n(*(time() - t) )\n\nout_2 = output[]\n(out_2.shape)\n(np.argmax(out_2, axis=))\n\n  (, )  f:\n    character_map = json.load(f)\nrev_character_map = {j:i  i,j  character_map.items()}\nprediction_str = .join([rev_character_map.get(s, )  s  np.argmax(out_2, axis=)])\nprediction_str\n</code></pre>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2274131,
              "author_name": "Kolya Forrat",
              "author_url": "",
              "post_date": "2023-05-25T16:34:11.310000",
              "content": "<p>Thanks for info. Could you please share <code>frames</code> shape and <code>out_2</code> shape? They are same by time axis or different?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2274136,
              "author_name": "George Megre",
              "author_url": "",
              "post_date": "2023-05-25T16:38:53.263000",
              "content": "<p>hm, to be honest i don't understand how you calculate it, but i can explain my understanding by two examples (hope this will be helpful):<br>\n1)<br>\ny_true =  my bare face in the wind<br>\ny_pred = my bare face sa the wind<br>\nN = 24<br>\nLEVD = 2<br>\nMetric = 22/24 = ~0.91</p>\n<p>But if we have:<br>\ny_true =  my bare face in the wind<br>\ny_pred = \"sssssssssssssssssssssssss wdaedaj\"<br>\nN = 24<br>\nLEVD = 30<br>\nMetric = 24-30 / 24 = -0.25</p>\n<p>So, it should be in the range (-d_min, 1).</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2274139,
              "author_name": "George Megre",
              "author_url": "",
              "post_date": "2023-05-25T16:41:15.287000",
              "content": "<p>Sure, here it is:<br>\nframes shape: (236, 252)<br>\nout_2 shape: (19, 59)</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2274148,
              "author_name": "Kolya Forrat",
              "author_url": "",
              "post_date": "2023-05-25T16:46:51.337000",
              "content": "<p>Thanks again. So now I'm totally don't understand why I can't submit my model)</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2256169": "Hello everyone! Nice to see that type of competition again\n\nI have one question. Could hosts please clarify meaning of \n>Your model must also perform inference in less than 5 hours and use less than 40 MB of storage space. Expect to see approximately 35 hours of video in the test set.\n\nHow to calculate how many phrases 35 hours includes? I would like to estimate one phrase inference time or how many phrases we have in that 35 hours.\n\nThanks\n\n\nUPDATE1. Also faced issues with submission. But can't detect the reason with info from Evaluation page. More info posted [below](https://www.kaggle.com/competitions/asl-fingerspelling/discussion/409722#2273788)\n\n\nUPDATE2. [Fixed](https://www.kaggle.com/competitions/asl-fingerspelling/discussion/409722#2276829)\n\nUPDATE3. Approximation of inference time. [here](https://www.kaggle.com/competitions/asl-fingerspelling/discussion/409722#2276907) and [here](https://www.kaggle.com/competitions/asl-fingerspelling/discussion/409722#2308349)",
    "2276907": "I also calculated the approximate time of one sample inference. \nSubmission scoring lasts ~3hours, mean inference time ~0.42 seconds. So:\n`3*3600 seconds / 0.42 ~= 25714` samples\n\nWe have 5 hours limitation. With this logic approximate maximum Mean Only Inference Time for model ~ 0.7 seconds",
    "2276829": "I finally found the problem with submission. \n\nIn previous competition I filled nans with this code. And also I unsqueezed tensor like this. (I changed both lines, didn't try change it one by one)\n\n```python\nx = tf.where(tfnp.isnan(x), 0.0, x)\nx = tf.expand_dims(x, axis=0)\n```\n\n\nIt worked fine in previous competition. And it work fine here in local kernel. But has error while submission scoring\n\nI just changed this lines to:\n```python\nx = tf.where(tfnp.isnan(x), tf.zeros_like(x), x)\nx = x[None]\n```\n\nMaybe it will help someone. Not so obvious reasons, as for me",
    "2273788": "~~I'm using this code to make sure my model works. It looks like absolutely fit with info from `Evaluation` page. May hosts (or some kagglers) please clarify why it has errors after submission? And provide the sample code to make sure that some submission will succeed.\nIt don't depends of inference time, same error with 0.4s Mean Inference time (after 42 minutes) and with 0.02s Mean Inference time (after 7 minutes).~~\n\n~~tflite_runtime version == 2.9.1~~\n~~Model size << 40 MB  (tried even with size 3MB)~~\n\n~~Capable with all inputs shapes [None, len(selected_columns)].  [from 0 to inf)~~\n~~Capable with full nan inputs.~~\n\nWIth [fixed](https://www.kaggle.com/competitions/asl-fingerspelling/discussion/409722#2276829) preprocessing code it works\n\n\n```python\n!pip install tflite-runtime==2.9.1\nimport tflite_runtime.interpreter as tflite\n\nimport time\nimport json\nimport pandas as pd\nfrom tqdm.auto import tqdm\nimport Levenshtein as Lev\nimport tensorflow as tf\n\nconverter = tf.lite.TFLiteConverter.from_keras_model(infr)  # infr it's my TF or Keras model (tried both)\ntflite_model = converter.convert()\nwith open('model.tflite', 'wb') as f:\n    f.write(tflite_model)\n!zip submission.zip 'model.tflite' 'inference_args.json'\n\n\nSEL_FEATURES = json.load(open('/kaggle/working/inference_args.json'))['selected_columns']\n\ndef load_relevant_data_subset(pq_path):\n        return pd.read_parquet(pq_path, columns=SEL_FEATURES) #selected_columns)\n\nwith open (\"/kaggle/input/asl-fingerspelling/character_to_prediction_index.json\", \"r\") as f:\n    character_map = json.load(f)\nrev_character_map = {j:i for i,j in character_map.items()}\n\n\ndf = pd.read_csv('/kaggle/input/asl-fingerspelling/train.csv')\n\nidx = 0\nsample = df.loc[idx]\nloaded = load_relevant_data_subset('/kaggle/input/asl-fingerspelling/' + sample['path'])\nloaded = loaded[loaded.index==sample['sequence_id']].values\nprint(loaded.shape)\nframes = loaded\n\ndef wer__(s1, s2):\n    w1 = len(s1.split())\n    lvd = Lev.distance(s1, s2)\n    return lvd / w1\n\nimport tflite_runtime.interpreter as tflite\ninterpreter = tflite.Interpreter('model.tflite')\nfound_signatures = list(interpreter.get_signature_list().keys())\n\nREQUIRED_SIGNATURE = 'serving_default'\nREQUIRED_OUTPUT = 'outputs'\nif REQUIRED_SIGNATURE not in found_signatures:\n    raise KernelEvalException('Required input signature not found.')\n    \nprediction_fn = interpreter.get_signature_runner(\"serving_default\")\noutput_lite = prediction_fn(inputs=frames)\nprediction_str = \"\".join([rev_character_map.get(s, \"\") for s in np.argmax(output_lite[REQUIRED_OUTPUT], axis=1)])\nprint(prediction_str)\n\n\nst = time.time()\ncnt = 0\ntotal = 100 #len(df)\nmodel_time = 0\n\nlevs = []\n\nfor i in tqdm(range(len(df.iloc[:total]))):\n    sample = df.loc[i]\n    loaded = load_relevant_data_subset('/kaggle/input/asl-fingerspelling/' + sample['path'])\n    loaded = loaded[loaded.index==sample['sequence_id']].values\n\n    md_st = time.time()\n    output_ = prediction_fn(inputs=loaded)\n    model_time += time.time() - md_st\n\n\n    prediction_str = \"\".join([rev_character_map.get(s, \"\") for s in np.argmax(output_[REQUIRED_OUTPUT], axis=1)])\n    cur_lev = wer__(sample['phrase'], prediction_str) \n    #print(sample['phrase'], '|', prediction_str, '|', cur_lev)\n    #print()\n\n    levs.append(cur_lev)\n\nprint(f'WER: {np.mean(levs):.5f}')\nprint(f'Mean time: {(time.time() - st)/total:.7f}')\nprint(f'Mean time only infer: {model_time/total:.7f}')\n```\n",
    "2324330": "I see public notebook run 5-6 hours, where is 5 hour limit? Notebook still has 9 hour limit submissiong? A bit confused here.",
    "2306719": "Sorry to bother you.\n\nMy model.tflite works wll in the Notebook with tensorflow 2.11.0. But it gets Submission Scoring Error soon after submit.\n\nMy model.tflite is converted from a PyTorch model.\n\nShould I use tensorflow 2.12 or 2.14?  Is there a safe way to convert torch model to tflite?\n\nIf you were me, what is your choice?",
    "2277083": "@kolyaforrat It's bit off topic but lb0.734 is awesome! May I ask what's your validation score? ",
    "2272974": "After submissions became available I think this topic is actual again.\n\nI have few submits with ~40-42 minutes with `Submission Scoring Error`. But:\n>Your model must also perform inference in less than 5 hours \n\nOr maybe there are some kind of multiprocessing and the time cap is 45 minutes?\n\n\nDoes anyone have successful submissions longer than 50 minutes? Any ideas about mean per sample inference time?\n\nupd. Error in my submission not because of 'out of time' . I deleted most of layers from my model and it has Error at ~8 minutes. \nSo I don't understand the problem. My input is `(time_dim, len(selected_columns))` and output with (new_time_dim, 59) and still getting errors. \nModel under 40 MB, capable with tflite, same environment with successful public kernels. Maybe here some another limitations I didn't know? At Kaggle kernels it works fine with Evaluation example code 😨"
  }
}