{
  "id": 413598,
  "title": "TFLite/Submission Problems Thread",
  "url": "/competitions/asl-fingerspelling/discussion/413598",
  "author_name": "anokas",
  "post_date": "2023-05-29T11:46:18.572000",
  "votes": 14,
  "comment_count": 31,
  "views": 0,
  "content": "<p>Most people are having issues getting submissions to work in this competition, even those that had no problems in the previous competition. While I was able to submit a constant prediction model without issues, I've got another model which gets 0.70 CV and works perfectly following the <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/overview/evaluation\" target=\"_blank\">official evaluation process</a> but fails on Kaggle.</p>\n<p>I wanted to make a thread collating both potential ideas for fixes and working submission code so that we can figure out what's going on. At the end of the day, it would be great if something could be done from the organisers side <a href=\"https://www.kaggle.com/thadstarner\" target=\"_blank\">@thadstarner</a> <a href=\"https://www.kaggle.com/sohierto\" target=\"_blank\">@sohierto</a> to improve this - for example:</p>\n<ul>\n<li>allowing us to see any errors that occur <em>before</em> our model sees test data (for example, if the tflite file fails to load). This would prevent people from being able to learn anything about the test data (the goal of code submission) but make the process much less annoying. </li>\n<li>giving us a representative evaluation code/environment so that we can test our models locally. It's clear that there's some extra requirements that aren't listed on the Evaluation page.</li>\n</ul>\n<p>In the meantime, some resources to help us figure this out:</p>\n<p><strong>Working submissions:</strong></p>\n<ul>\n<li>1-layer dummy model with <code>(N_FRAMES, 59)</code> output: <a href=\"https://www.kaggle.com/code/wonderingalice/working-sample-submission-and-inference\" target=\"_blank\">https://www.kaggle.com/code/wonderingalice/working-sample-submission-and-inference</a> <a href=\"https://www.kaggle.com/wonderingalice\" target=\"_blank\">@wonderingalice</a></li>\n<li>Model with static <code>(11, 59)</code> output size: <a href=\"https://www.kaggle.com/code/anokas/static-greedy-baseline-0-157-lb\" target=\"_blank\">https://www.kaggle.com/code/anokas/static-greedy-baseline-0-157-lb</a></li>\n<li>1-layer dummy PyTorch model: <a href=\"https://www.kaggle.com/code/chack3/pytorch-simple-model\" target=\"_blank\">https://www.kaggle.com/code/chack3/pytorch-simple-model</a> <a href=\"https://www.kaggle.com/chack3\" target=\"_blank\">@chack3</a></li>\n</ul>\n<p><strong>Possible fixes:</strong></p>\n<ul>\n<li>Changing nan-infill and expand_dims method: <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/409722#2276829\" target=\"_blank\">https://www.kaggle.com/competitions/asl-fingerspelling/discussion/409722#2276829</a> <a href=\"https://www.kaggle.com/forrato\" target=\"_blank\">@forrato</a> It's really unclear to me why this would help.</li>\n<li>Making sure model can handle empty input: <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/415372\" target=\"_blank\">https://www.kaggle.com/competitions/asl-fingerspelling/discussion/415372</a></li>\n</ul>\n<p>If anyone has any ideas or other examples, let me know and I'll try to keep this thread updated! From my perspective, I'll probably have to take a break until we get some clarity from the organisers 😅</p>",
  "messages": [
    {
      "id": 2279482,
      "postDate": "2023-05-29T11:46:18.573Z",
      "content": "<p>Most people are having issues getting submissions to work in this competition, even those that had no problems in the previous competition. While I was able to submit a constant prediction model without issues, I've got another model which gets 0.70 CV and works perfectly following the <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/overview/evaluation\" target=\"_blank\">official evaluation process</a> but fails on Kaggle.</p>\n<p>I wanted to make a thread collating both potential ideas for fixes and working submission code so that we can figure out what's going on. At the end of the day, it would be great if something could be done from the organisers side <a href=\"https://www.kaggle.com/thadstarner\" target=\"_blank\">@thadstarner</a> <a href=\"https://www.kaggle.com/sohierto\" target=\"_blank\">@sohierto</a> to improve this - for example:</p>\n<ul>\n<li>allowing us to see any errors that occur <em>before</em> our model sees test data (for example, if the tflite file fails to load). This would prevent people from being able to learn anything about the test data (the goal of code submission) but make the process much less annoying. </li>\n<li>giving us a representative evaluation code/environment so that we can test our models locally. It's clear that there's some extra requirements that aren't listed on the Evaluation page.</li>\n</ul>\n<p>In the meantime, some resources to help us figure this out:</p>\n<p><strong>Working submissions:</strong></p>\n<ul>\n<li>1-layer dummy model with <code>(N_FRAMES, 59)</code> output: <a href=\"https://www.kaggle.com/code/wonderingalice/working-sample-submission-and-inference\" target=\"_blank\">https://www.kaggle.com/code/wonderingalice/working-sample-submission-and-inference</a> <a href=\"https://www.kaggle.com/wonderingalice\" target=\"_blank\">@wonderingalice</a></li>\n<li>Model with static <code>(11, 59)</code> output size: <a href=\"https://www.kaggle.com/code/anokas/static-greedy-baseline-0-157-lb\" target=\"_blank\">https://www.kaggle.com/code/anokas/static-greedy-baseline-0-157-lb</a></li>\n<li>1-layer dummy PyTorch model: <a href=\"https://www.kaggle.com/code/chack3/pytorch-simple-model\" target=\"_blank\">https://www.kaggle.com/code/chack3/pytorch-simple-model</a> <a href=\"https://www.kaggle.com/chack3\" target=\"_blank\">@chack3</a></li>\n</ul>\n<p><strong>Possible fixes:</strong></p>\n<ul>\n<li>Changing nan-infill and expand_dims method: <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/409722#2276829\" target=\"_blank\">https://www.kaggle.com/competitions/asl-fingerspelling/discussion/409722#2276829</a> <a href=\"https://www.kaggle.com/forrato\" target=\"_blank\">@forrato</a> It's really unclear to me why this would help.</li>\n<li>Making sure model can handle empty input: <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/415372\" target=\"_blank\">https://www.kaggle.com/competitions/asl-fingerspelling/discussion/415372</a></li>\n</ul>\n<p>If anyone has any ideas or other examples, let me know and I'll try to keep this thread updated! From my perspective, I'll probably have to take a break until we get some clarity from the organisers 😅</p>",
      "rawMarkdown": "Most people are having issues getting submissions to work in this competition, even those that had no problems in the previous competition. While I was able to submit a constant prediction model without issues, I've got another model which gets 0.70 CV and works perfectly following the [official evaluation process](https://www.kaggle.com/competitions/asl-fingerspelling/overview/evaluation) but fails on Kaggle.\n\nI wanted to make a thread collating both potential ideas for fixes and working submission code so that we can figure out what's going on. At the end of the day, it would be great if something could be done from the organisers side @thadstarner @sohierto to improve this - for example:\n- allowing us to see any errors that occur _before_ our model sees test data (for example, if the tflite file fails to load). This would prevent people from being able to learn anything about the test data (the goal of code submission) but make the process much less annoying. \n- giving us a representative evaluation code/environment so that we can test our models locally. It's clear that there's some extra requirements that aren't listed on the Evaluation page.\n\nIn the meantime, some resources to help us figure this out:\n\n**Working submissions:**\n- 1-layer dummy model with `(N_FRAMES, 59)` output: https://www.kaggle.com/code/wonderingalice/working-sample-submission-and-inference @wonderingalice\n- Model with static `(11, 59)` output size: https://www.kaggle.com/code/anokas/static-greedy-baseline-0-157-lb\n- 1-layer dummy PyTorch model: https://www.kaggle.com/code/chack3/pytorch-simple-model @chack3\n\n**Possible fixes:**\n- Changing nan-infill and expand_dims method: https://www.kaggle.com/competitions/asl-fingerspelling/discussion/409722#2276829 @forrato It's really unclear to me why this would help.\n- Making sure model can handle empty input: https://www.kaggle.com/competitions/asl-fingerspelling/discussion/415372\n\nIf anyone has any ideas or other examples, let me know and I'll try to keep this thread updated! From my perspective, I'll probably have to take a break until we get some clarity from the organisers 😅",
      "votes": 14
    },
    {
      "id": 2282575,
      "postDate": "2023-05-31T16:50:02.327Z",
      "content": "<p>I'll look into this.</p>",
      "rawMarkdown": "I'll look into this.",
      "votes": 1,
      "replies": [
        {
          "id": 2282635,
          "postDate": "2023-05-31T17:50:36.803Z",
          "content": "<p>Thank you!</p>",
          "rawMarkdown": "Thank you!",
          "votes": 2
        }
      ]
    },
    {
      "id": 2280786,
      "postDate": "2023-05-30T10:51:54.700Z",
      "content": "<p>From the evaluation page - \" Your model must be packaged into a submission.zip file and compatible with the TensorFlow Lite Runtime v2.9.1\"   I posted <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/411268#2267655\" target=\"_blank\">here</a> about forking an older notebook with Pin to Original Environment (prior to python 3.10 upgrade since TFLite does not have a version yet for it).  So if you are creating your submission .zip offline you can do that just to test in a notebook if your model works ok and any errors when you try to use it.  Found some tf operators do not work and you have to find a workaround for them. </p>\n<p>Regarding NaN frames or hands, using a preprocessing layer for selected columns that included frame - <br>\n<code>xlh = tf.slice(x0, [0,1], [-1,42 ])  # lh slice  x y (n.b. 0 is frame so use 1)</code> <br>\n<code>xrh = tf.slice(x0, [0,43], [-1,42 ])   # rh slice x y</code><br>\n<code>xrhv = tf.boolean_mask(xrh, tf.reduce_all(~tf.math.is_nan(xrh), axis=1)) # remove NaN rows</code><br>\n<code>xlhv =tf.boolean_mask(xlh, tf.reduce_all(~tf.math.is_nan(xlh), axis=1))</code><br>\n<code>x = tf.cond(tf.math.is_nan(tf.math.zero_fraction(xlhv)), lambda: tf.identity(xrhv), lambda: tf.identity(xlhv))</code></p>\n<p>Can confirm this works in submission without errors.  Fine if you use this in your code, but maybe don't just make a notebook using it. Pretty sure people can read the code here if they want.</p>",
      "rawMarkdown": "From the evaluation page - \" Your model must be packaged into a submission.zip file and compatible with the TensorFlow Lite Runtime v2.9.1\"   I posted [here](https://www.kaggle.com/competitions/asl-fingerspelling/discussion/411268#2267655) about forking an older notebook with Pin to Original Environment (prior to python 3.10 upgrade since TFLite does not have a version yet for it).  So if you are creating your submission .zip offline you can do that just to test in a notebook if your model works ok and any errors when you try to use it.  Found some tf operators do not work and you have to find a workaround for them. \n\nRegarding NaN frames or hands, using a preprocessing layer for selected columns that included frame - \n`xlh = tf.slice(x0, [0,1], [-1,42 ])  # lh slice  x y (n.b. 0 is frame so use 1)` \n`xrh = tf.slice(x0, [0,43], [-1,42 ])   # rh slice x y`\n`xrhv = tf.boolean_mask(xrh, tf.reduce_all(~tf.math.is_nan(xrh), axis=1)) # remove NaN rows`\n`xlhv =tf.boolean_mask(xlh, tf.reduce_all(~tf.math.is_nan(xlh), axis=1))`\n`x = tf.cond(tf.math.is_nan(tf.math.zero_fraction(xlhv)), lambda: tf.identity(xrhv), lambda: tf.identity(xlhv))`\n\nCan confirm this works in submission without errors.  Fine if you use this in your code, but maybe don't just make a notebook using it. Pretty sure people can read the code here if they want.",
      "votes": 1,
      "replies": [
        {
          "id": 2281029,
          "postDate": "2023-05-30T14:25:26.060Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/something4kag\" target=\"_blank\">@something4kag</a> </p>\n<p>Thanks for sharing! This will definitely be inspiring …</p>\n<p>However, w.r.t the tflite runtime environment: I'm really not sure what else to try.<br>\nI have been working from my \"working submission notebook\" which is a fork of the previous competition (dating from Feb, I think) and explicitly installs Runtime v.2.9.1. It has Tensorflow 2.11.0 and Python 3.7.12. All the scoring errors I (and many others) I referred to in various discussion replies are from submissions for which TfLite inference works perfectly well in (a fork of) that notebook. After almost three weeks only 7 people have managed to get beyond dummy submissions or baselines, so clearly something is really unclear?</p>\n<p>Do you have any idea what else we need to change in order to make inference errors for the functions that are occurring in the scoring pipeline also visible during Kaggle notebook inference, so we can properly debug our models?</p>",
          "rawMarkdown": "Hi @something4kag \n\nThanks for sharing! This will definitely be inspiring ...\n\nHowever, w.r.t the tflite runtime environment: I'm really not sure what else to try.\nI have been working from my \"working submission notebook\" which is a fork of the previous competition (dating from Feb, I think) and explicitly installs Runtime v.2.9.1. It has Tensorflow 2.11.0 and Python 3.7.12. All the scoring errors I (and many others) I referred to in various discussion replies are from submissions for which TfLite inference works perfectly well in (a fork of) that notebook. After almost three weeks only 7 people have managed to get beyond dummy submissions or baselines, so clearly something is really unclear?\n\nDo you have any idea what else we need to change in order to make inference errors for the functions that are occurring in the scoring pipeline also visible during Kaggle notebook inference, so we can properly debug our models?\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 2279602,
      "postDate": "2023-05-29T13:23:22.037Z",
      "content": "<p>Certain operators not being supported, like <code>expand_dims</code>, is also making it difficult to port models fro PyTorch to Keras through ONNX, because you don't always control which TF operators are used.</p>\n<p>I'm currently in the process of porting parts of the model to Keras and seeing when the model gets through the submission pipeline, if I find any other operators that aren't supported I'll update this.</p>",
      "rawMarkdown": "Certain operators not being supported, like `expand_dims`, is also making it difficult to port models fro PyTorch to Keras through ONNX, because you don't always control which TF operators are used.\n\nI'm currently in the process of porting parts of the model to Keras and seeing when the model gets through the submission pipeline, if I find any other operators that aren't supported I'll update this.",
      "votes": 1,
      "replies": [
        {
          "id": 2279639,
          "postDate": "2023-05-29T13:52:17.707Z",
          "content": "<blockquote>\n  <p>Certain operators not being supported, like expand_dims</p>\n</blockquote>\n<p>Interesting, as my model which doesn't work uses <code>tf.expand_dims</code>. What makes you think it's not supported? (or is it just the comment by <a href=\"https://www.kaggle.com/forranto\" target=\"_blank\">@forranto</a>?)<br>\nSince the model can be converted to tflite and loaded/evaluated with the interpreter, I wonder what operations are supported by the official TFLite interpreter but not the one Kaggle is using.</p>",
          "rawMarkdown": "> Certain operators not being supported, like expand_dims\n\nInteresting, as my model which doesn't work uses `tf.expand_dims`. What makes you think it's not supported? (or is it just the comment by @forranto?)\nSince the model can be converted to tflite and loaded/evaluated with the interpreter, I wonder what operations are supported by the official TFLite interpreter but not the one Kaggle is using.",
          "replies": [
            {
              "id": 2279649,
              "postDate": "2023-05-29T13:56:44.813Z",
              "content": "<p>Yes, I was referring to the comment by <a href=\"https://www.kaggle.com/forranto\" target=\"_blank\">@forranto</a>. It woud be really helpful if we could get the exact evaluation code that is used to score submissions (with library versions) to be able to figure out what's going wrong.</p>",
              "rawMarkdown": "Yes, I was referring to the comment by @forranto. It woud be really helpful if we could get the exact evaluation code that is used to score submissions (with library versions) to be able to figure out what's going wrong."
            }
          ]
        },
        {
          "id": 2279802,
          "postDate": "2023-05-29T16:44:12.330Z",
          "content": "<p>I've been looking at my model in <a href=\"https://netron.app/\" target=\"_blank\">https://netron.app/</a> (you can just upload your <code>.tflite</code>), and it looks like <code>x[None]</code> still compiles to an <code>ExpandDims</code> op in the actual model. Besides, a Conv1D just compiles to a Conv2D with ExpandDims:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F505747%2Fdbf58928d59fbda74ea961e20979cb61%2Fconv2d.png?generation=1685378548671571&amp;alt=media\" alt=\"\"></p>\n<p><a href=\"https://www.kaggle.com/kolyaforrat\" target=\"_blank\">@kolyaforrat</a> Would you be willing to share a list of TFLite Ops in your model, so we can see if there are some ops that are failing? Hopefully that doesn't give away anything about your solution (except perhaps whether it uses convolutions or recurrent layers at all, maybe this is secret).</p>",
          "rawMarkdown": "I've been looking at my model in https://netron.app/ (you can just upload your `.tflite`), and it looks like `x[None]` still compiles to an `ExpandDims` op in the actual model. Besides, a Conv1D just compiles to a Conv2D with ExpandDims:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F505747%2Fdbf58928d59fbda74ea961e20979cb61%2Fconv2d.png?generation=1685378548671571&alt=media)\n\n@kolyaforrat Would you be willing to share a list of TFLite Ops in your model, so we can see if there are some ops that are failing? Hopefully that doesn't give away anything about your solution (except perhaps whether it uses convolutions or recurrent layers at all, maybe this is secret).",
          "votes": 2,
          "replies": [
            {
              "id": 2279935,
              "postDate": "2023-05-29T18:22:30.040Z",
              "content": "<p>Without revealing any secrets I can share netron part with differences between successfully submitted model and failed model.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4212496%2F56bb4405dcf9a4450974bfef56aaf838%2F2023-05-29%20%2021.08.46.png?generation=1685383933882709&amp;alt=media\" alt=\"expand dims\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4212496%2Fcecc605c4c526c46802889d451e292b8%2F2023-05-29%20%2021.08.58.png?generation=1685384192986888&amp;alt=media\" alt=\"\"></p>\n<p>That's the only difference between two graphs.</p>\n<p>About Conv2D and TFLite you can read in comments to <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406306\" target=\"_blank\">my team's write-up</a> from previous competition. <br>\nTLDR: not all conv types are good convertable to TFLite with torch-onnx-keras pipeline. But all are convertable with <a href=\"https://github.com/AlexanderLutsenko/nobuco\" target=\"_blank\">nobuco</a>. But in Kaggle environment some of them slower than our keras implemetation (But locally in different environment nobuco is better). So with <code>nobuco</code> no errors but slower</p>",
              "rawMarkdown": "Without revealing any secrets I can share netron part with differences between successfully submitted model and failed model.\n\n![expand dims](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4212496%2F56bb4405dcf9a4450974bfef56aaf838%2F2023-05-29%20%2021.08.46.png?generation=1685383933882709&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4212496%2Fcecc605c4c526c46802889d451e292b8%2F2023-05-29%20%2021.08.58.png?generation=1685384192986888&alt=media)\n\nThat's the only difference between two graphs.\n\nAbout Conv2D and TFLite you can read in comments to [my team's write-up](https://www.kaggle.com/competitions/asl-signs/discussion/406306) from previous competition. \nTLDR: not all conv types are good convertable to TFLite with torch-onnx-keras pipeline. But all are convertable with [nobuco](https://github.com/AlexanderLutsenko/nobuco). But in Kaggle environment some of them slower than our keras implemetation (But locally in different environment nobuco is better). So with `nobuco` no errors but slower",
              "votes": 4
            }
          ]
        }
      ]
    },
    {
      "id": 2285420,
      "postDate": "2023-06-02T17:40:08.610Z",
      "content": "<p><a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/414682\" target=\"_blank\">Cross posting here</a>: it turns out that the metric has been using TF Lite runtime v2.14.0 rather than v2.9.1. Thank you all for this discussion and pushing us to track that down. </p>",
      "rawMarkdown": "[Cross posting here](https://www.kaggle.com/competitions/asl-fingerspelling/discussion/414682): it turns out that the metric has been using TF Lite runtime v2.14.0 rather than v2.9.1. Thank you all for this discussion and pushing us to track that down. ",
      "votes": 2
    },
    {
      "id": 2283608,
      "postDate": "2023-06-01T11:42:10.847Z",
      "content": "<p>Hi all,</p>\n<p>Although I finally got a model without pre-or postprocessing through scoring, I'm still struggling with the same issues in pre- and post. <br>\nI now have two snippets of code, of the following structure (with x a <strong>2D float32 tensor</strong>):</p>\n<p><strong>Snippet A:</strong>   --  doesn't pass through scoring: </p>\n<pre><code>x2 = tf.math.reduce_sum(x, axis = )           \nmask = tf.not_equal(x2, )                     \nmasked_x = x[mask]                              \n</code></pre>\n<p><strong>Snippet B:</strong> --  passes through scoring: </p>\n<pre><code>x2 = tf.argmax(x, axis=)                              \nmask = tf.not_equal(x2,)                               \nmasked_x = x[mask]                                     \n</code></pre>\n<p>As always, stepping through these operations on dummy tensors produces the expected results and using the code with TfLite runtime (v2.9.1) works perfectly.</p>\n<p>As far as I can see the only difference is the reduce_sum operation, but I have found no indications anywhere online about this not being supported …</p>",
      "rawMarkdown": "Hi all,\n\nAlthough I finally got a model without pre-or postprocessing through scoring, I'm still struggling with the same issues in pre- and post. \nI now have two snippets of code, of the following structure (with x a **2D float32 tensor**):\n\n**Snippet A:**   --  doesn't pass through scoring: \n\n```python\nx2 = tf.math.reduce_sum(x, axis = 1)           # 1D float32 tensor\nmask = tf.not_equal(x2, 0)                     # 1D boolean mask  -- already tried comparing with 0.0 as well\nmasked_x = x[mask]                              # 1D float32 tensor\n```\n\n**Snippet B:** --  passes through scoring: \n\n```python\nx2 = tf.argmax(x, axis=1)                              # 1D int64 tensor        \nmask = tf.not_equal(x2,0)                               # 1D boolean mask\nmasked_x = x[mask]                                     # 1D float32 tensor\n```\n\nAs always, stepping through these operations on dummy tensors produces the expected results and using the code with TfLite runtime (v2.9.1) works perfectly.\n\nAs far as I can see the only difference is the reduce_sum operation, but I have found no indications anywhere online about this not being supported ...\n\n",
      "votes": 2
    },
    {
      "id": 2282859,
      "postDate": "2023-05-31T21:35:31.773Z",
      "content": "<p>A couple of thoughts:</p>\n<ul>\n<li>There is <a href=\"https://www.kaggle.com/code/irohith/aslfr-transformer\" target=\"_blank\">a public notebook that generates a non-static submission successfully</a>. I would encourage everyone to take a look at <a href=\"https://www.kaggle.com/irohith\" target=\"_blank\">@irohith</a>'s work.</li>\n<li>I dug into the error logs and it turns out that few different things are happening. A meaningful fraction of submissions still include ops that aren't supported by the TF Lite runtime. Others have invalid signatures, so I updated the evaluation tab to be more explicit about the requirements there. However the majority of recent errors are dimension mismatch errors where I would need access to the model's training code in order to debug further.</li>\n</ul>\n<p>Given the limited attention that the aslfr-transformer notebook has gotten so far I'm going to pause my review for now. I'll check in again next week once everyone has had time to digest that successful example.</p>",
      "rawMarkdown": "A couple of thoughts:\n\n- There is [a public notebook that generates a non-static submission successfully](https://www.kaggle.com/code/irohith/aslfr-transformer). I would encourage everyone to take a look at @irohith's work.\n- I dug into the error logs and it turns out that few different things are happening. A meaningful fraction of submissions still include ops that aren't supported by the TF Lite runtime. Others have invalid signatures, so I updated the evaluation tab to be more explicit about the requirements there. However the majority of recent errors are dimension mismatch errors where I would need access to the model's training code in order to debug further.\n\nGiven the limited attention that the aslfr-transformer notebook has gotten so far I'm going to pause my review for now. I'll check in again next week once everyone has had time to digest that successful example.",
      "votes": 2,
      "replies": [
        {
          "id": 2283476,
          "postDate": "2023-06-01T09:20:30.583Z",
          "content": "<p>How could we verify our models do not include operators that are not supported by the TFLite runtime?<br>\nI am currently using the following script, but my submissions still fail:</p>\n<p>In addition, I verify my TFLite model is working correctly by performing inference on my <strong>full</strong> training dataset, which runs without errors. It is quite a mystery why the submission still fails.</p>\n<p><strong>Edit</strong></p>\n<p>It seems to me the TFLite environment we use is different from the one used during inference. We therefore have no method to verify whether our inference loop is correct.</p>\n<pre><code># Create Model Converter\nkeras_model_converter = tf.lite.TFLiteConverter.from_keras_model(tflite_keras_model)\n# Convert Model\ntflite_model = keras_model_converter.convert()\n# Write Model\nwith open('/kaggle/working/model.tflite', 'wb') as f:\n    f.write(tflite_model)\n</code></pre>",
          "rawMarkdown": "How could we verify our models do not include operators that are not supported by the TFLite runtime?\nI am currently using the following script, but my submissions still fail:\n\nIn addition, I verify my TFLite model is working correctly by performing inference on my **full** training dataset, which runs without errors. It is quite a mystery why the submission still fails.\n\n**Edit**\n\nIt seems to me the TFLite environment we use is different from the one used during inference. We therefore have no method to verify whether our inference loop is correct.\n\n```\n# Create Model Converter\nkeras_model_converter = tf.lite.TFLiteConverter.from_keras_model(tflite_keras_model)\n# Convert Model\ntflite_model = keras_model_converter.convert()\n# Write Model\nwith open('/kaggle/working/model.tflite', 'wb') as f:\n    f.write(tflite_model)\n```",
          "votes": 6,
          "replies": [
            {
              "id": 2284837,
              "postDate": "2023-06-02T09:54:37.347Z",
              "content": "<p>Either the runtime environment is different, or there are some edge cases in the test set which the model can't handle (empty samples for example).</p>\n<p>I have tried taking these edge cases into account, but the submissions still failed.</p>",
              "rawMarkdown": "Either the runtime environment is different, or there are some edge cases in the test set which the model can't handle (empty samples for example).\n\nI have tried taking these edge cases into account, but the submissions still failed."
            }
          ]
        },
        {
          "id": 2284772,
          "postDate": "2023-06-02T09:08:55.280Z",
          "content": "<blockquote>\n  <p>include ops that aren't supported by the TF Lite runtime</p>\n</blockquote>\n<p>Thanks for looking into this, but this is the part that confuses me. When we use the TFLite runtime V2.9.1 locally, <strong>all our ops are supported fine</strong> - so it appears the kaggle tflite runtime is different to the one that is publicly available?</p>",
          "rawMarkdown": "> include ops that aren't supported by the TF Lite runtime\n\nThanks for looking into this, but this is the part that confuses me. When we use the TFLite runtime V2.9.1 locally, **all our ops are supported fine** - so it appears the kaggle tflite runtime is different to the one that is publicly available?",
          "votes": 4
        },
        {
          "id": 2285249,
          "postDate": "2023-06-02T15:00:24.977Z",
          "content": "<p>I've just noticed something you mentioned:</p>\n<blockquote>\n  <p>However the majority of recent errors are dimension mismatch errors where I would need access to the model's training code in order to debug further.</p>\n</blockquote>\n<p>Could it be that the evaluation code fails when an empty prediction is made?</p>\n<p>The published evaluation code handles empty predictions fine, but perhaps this is the source of one difference between the specification and the implementation.</p>",
          "rawMarkdown": "I've just noticed something you mentioned:\n> However the majority of recent errors are dimension mismatch errors where I would need access to the model's training code in order to debug further.\n\nCould it be that the evaluation code fails when an empty prediction is made?\n\nThe published evaluation code handles empty predictions fine, but perhaps this is the source of one difference between the specification and the implementation.",
          "votes": 1,
          "replies": [
            {
              "id": 2285409,
              "postDate": "2023-06-02T17:30:30.743Z",
              "content": "<p>Most of the dimension mismatches are between non-zero values that don't have any obvious connection to the original dataset shape, like <code>2 != 1</code> or <code>85 != 84</code>.</p>",
              "rawMarkdown": "Most of the dimension mismatches are between non-zero values that don't have any obvious connection to the original dataset shape, like `2 != 1` or `85 != 84`."
            },
            {
              "id": 2285730,
              "postDate": "2023-06-03T00:05:48.853Z",
              "content": "<p>Interesting, thank you for double checking!</p>",
              "rawMarkdown": "Interesting, thank you for double checking!"
            }
          ]
        }
      ]
    },
    {
      "id": 2279903,
      "postDate": "2023-05-29T18:04:51.013Z",
      "content": "<p>I've spent my 5 attempts today trying to get a working solution for dropping frames (e.g. the missing hands frames). For this, too, code from the previous competition works for notebook inference but gives a scoring error when I try to submit. I tried making sure that at least one frame remains, but that didn't help. </p>\n<p><strong>It would be wonderful if anyone would be willing to share a working solution for that.</strong></p>\n<p>In return, I can confirm that the <code>x[None]</code> replacement for expand_dims works perfectly. Also, do not forget to remove that initial dimension again from the output of your batch-trained model, i.e., assuming your model, called  <code>tmodel</code> below, probably returns a tensor of shape [None,None,59], so the submission model could look like:</p>\n<pre><code>trained_model = &lt;your_model_path&gt;\ndummy_inner = tf.keras.models.load_model(trained_model)\n\n ():\n    inputs = tf.keras.Input(shape=(NUM_FEATURES), dtype=tf.float32, name=)\n    \n    x = tf.where(tf.math.is_nan(inputs), tf.zeros_like(inputs), inputs)\n    \n    x = x[]\n    \n    x = tmodel(x)\n    \n    out = tf.keras.layers.Activation(, name=)(x[,:,:])\n\ndummy_submission_model = get_dummy_model(dummy_inner)\ndummy_submission_model.summary()\n</code></pre>\n<p>(updated: initially I copy-pasted the wrong version, sorry)</p>",
      "rawMarkdown": "I've spent my 5 attempts today trying to get a working solution for dropping frames (e.g. the missing hands frames). For this, too, code from the previous competition works for notebook inference but gives a scoring error when I try to submit. I tried making sure that at least one frame remains, but that didn't help. \n\n**It would be wonderful if anyone would be willing to share a working solution for that.**\n\nIn return, I can confirm that the `x[None]` replacement for expand_dims works perfectly. Also, do not forget to remove that initial dimension again from the output of your batch-trained model, i.e., assuming your model, called  `tmodel` below, probably returns a tensor of shape [None,None,59], so the submission model could look like:\n\n````python\n\ntrained_model = <your_model_path>\ndummy_inner = tf.keras.models.load_model(trained_model)\n\ndef get_dummy_model(tmodel):\n    inputs = tf.keras.Input(shape=(NUM_FEATURES), dtype=tf.float32, name=\"inputs\")\n    # possibly add additional cleaning or preprocessing code below\n    x = tf.where(tf.math.is_nan(inputs), tf.zeros_like(inputs), inputs)\n    # replacement for expand_dims since that doesn't seem to make it through the scoring pipeline\n    x = x[None]\n    # Insert trained model\n    x = tmodel(x)\n    # possibly add postprocessing code below\n    out = tf.keras.layers.Activation(\"linear\", name=\"outputs\")(x[0,:,:])\n\ndummy_submission_model = get_dummy_model(dummy_inner)\ndummy_submission_model.summary()\n````\n\n(updated: initially I copy-pasted the wrong version, sorry)",
      "votes": 2,
      "replies": [
        {
          "id": 2280653,
          "postDate": "2023-05-30T08:44:55.623Z",
          "content": "<p>Update: trying to <strong>unsuccessfully</strong> remove missing hand frames (and get it through the scoring pipeline) landed here today:</p>\n<pre><code>\nhands_ok = tf.math.add(x[:,:],x[:,:])\n(, hands_ok.shape)\n\n\nhands_ok = tf.math.reduce_sum(hands_ok, axis = )    \n(,hands_ok.shape)   \n\n\nhands_ok = tf.cast(hands_ok, tf.)\n(,hands_ok.shape)\n\n\nmasked_x = x[hands_ok]\n(,x.shape)\n\n\nx = tf.concat([x[:,:],masked_x],axis = )\n(,x.shape) \n</code></pre>\n<p>Alternatives I tried today used <code>tf.where</code> or <code>tf.boolean_mask</code> or did not use the concat line or layers.Concatenate instead.<br>\n(all worked perfectly when doing tflite inference in notebook, all failed submission scoring, 5 attempts for today used up)</p>",
          "rawMarkdown": "Update: trying to **unsuccessfully** remove missing hand frames (and get it through the scoring pipeline) landed here today:\n\n```python\n\n# add left and right hand\nhands_ok = tf.math.add(x[:,63:],x[:,:63])\nprint(\"hands ok 1:\", hands_ok.shape)\n\n# add xyz\nhands_ok = tf.math.reduce_sum(hands_ok, axis = 1)    \nprint(\"hands ok 2:\",hands_ok.shape)   \n\n# convert to boolean\nhands_ok = tf.cast(hands_ok, tf.bool)\nprint(\"hands ok 3:\",hands_ok.shape)\n\n# Mask input frames\nmasked_x = x[hands_ok]\nprint(\"Masked:\",x.shape)\n\n# make sure that at least one frame remains\nx = tf.concat([x[0:1,:],masked_x],axis = 0)\nprint(\"Masked and added initial:\",x.shape) \n\n```\n\n\nAlternatives I tried today used `tf.where` or `tf.boolean_mask` or did not use the concat line or layers.Concatenate instead.\n(all worked perfectly when doing tflite inference in notebook, all failed submission scoring, 5 attempts for today used up)",
          "replies": [
            {
              "id": 2284628,
              "postDate": "2023-06-02T07:25:08.390Z",
              "content": "<p>I can confirm <code>pad</code> and <code>concat</code> do not work. In the previous Isolated Sign Language competition those operators did work. The TFLite conversion is also successful. No idea what is causing the submission to fail.</p>\n<p>Without proper preprocessing this whole competition is more about making a model that works on unprocessed data than actually creating a high performing model with advanced preprocessing, which should not be the point of this competition.</p>",
              "rawMarkdown": "I can confirm `pad` and `concat` do not work. In the previous Isolated Sign Language competition those operators did work. The TFLite conversion is also successful. No idea what is causing the submission to fail.\n\nWithout proper preprocessing this whole competition is more about making a model that works on unprocessed data than actually creating a high performing model with advanced preprocessing, which should not be the point of this competition.",
              "votes": 1
            },
            {
              "id": 2284737,
              "postDate": "2023-06-02T08:47:35.210Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/Mark\" target=\"_blank\">@Mark</a>, </p>\n<p>Are you using Tensorflow or converting from Pytorch?</p>\n<p>I do use a concat operation in my current submissions (all built and trained in Tensorflow), so it should work. I have, however, noticed at least once that there was in fact a dimension mismatch in the mask I used due to an error I made. This passed TfLite creation and TfLite inference in my notebook but not the scoring pipeline. Removing the error resulted in successful submission of that bit. </p>\n<p>What I'm doing now is checking each operation by stepping through the whole model and printing output tensors to screen. As you can see from the snippets in another comment, I am stuck at \"reduce_sum\", which I need for dominant hand detection. Tensors look fine but submissions keep failing.</p>",
              "rawMarkdown": "Hi @Mark, \n\nAre you using Tensorflow or converting from Pytorch?\n\nI do use a concat operation in my current submissions (all built and trained in Tensorflow), so it should work. I have, however, noticed at least once that there was in fact a dimension mismatch in the mask I used due to an error I made. This passed TfLite creation and TfLite inference in my notebook but not the scoring pipeline. Removing the error resulted in successful submission of that bit. \n\nWhat I'm doing now is checking each operation by stepping through the whole model and printing output tensors to screen. As you can see from the snippets in another comment, I am stuck at \"reduce_sum\", which I need for dominant hand detection. Tensors look fine but submissions keep failing.\n\n"
            },
            {
              "id": 2284744,
              "postDate": "2023-06-02T08:53:21.023Z",
              "content": "<p><code>x = tf.concat([x,  y], axis=0)</code><br>\nConcat is work for me. Pad is actually the same, I tried pad, it also works with <code>concat</code> operation</p>",
              "rawMarkdown": "`x = tf.concat([x,  y], axis=0)`\nConcat is work for me. Pad is actually the same, I tried pad, it also works with `concat` operation"
            },
            {
              "id": 2284814,
              "postDate": "2023-06-02T09:39:41.843Z",
              "content": "<p>I am using a Tensorflow preprocessing layer which keeps failing:</p>\n<pre><code>class PreprocessLayer(tf.keras.layers.Layer):\n    def __init__(self):\n        super(PreprocessLayer, self).__init__()\n\n    @tf.function(\n        input_signature=(tf.TensorSpec(shape=[None,N_COLS0], dtype=tf.float32),),\n    )\n    def call(self, data0):        \n        # Number of Frames in Video\n        N_FRAMES0 = tf.shape(data0)[0]\n\n        # Fill NaN Values With 0\n        data = tf.where(tf.math.is_nan(data0), 0.0, data0)\n\n        # THIS KEEPS FAILING!\n        # should simply pad videos shorter than N_TARGET_FRAMES (128) frames\n        pad = tf.math.maximum(0, N_TARGET_FRAMES - N_FRAMES0)\n        data = tf.pad(data, [[0,pad], [0,0]], constant_values=0.0)\n        non_empty_frames_idxs = tf.pad(non_empty_frames_idxs, [[0,pad]], constant_values=-1.0)\n</code></pre>\n<p>`</p>",
              "rawMarkdown": "I am using a Tensorflow preprocessing layer which keeps failing:\n\n```\nclass PreprocessLayer(tf.keras.layers.Layer):\n    def __init__(self):\n        super(PreprocessLayer, self).__init__()\n    \n    @tf.function(\n        input_signature=(tf.TensorSpec(shape=[None,N_COLS0], dtype=tf.float32),),\n    )\n    def call(self, data0):        \n        # Number of Frames in Video\n        N_FRAMES0 = tf.shape(data0)[0]\n        \n        # Fill NaN Values With 0\n        data = tf.where(tf.math.is_nan(data0), 0.0, data0)\n\n        # THIS KEEPS FAILING!\n        # should simply pad videos shorter than N_TARGET_FRAMES (128) frames\n        pad = tf.math.maximum(0, N_TARGET_FRAMES - N_FRAMES0)\n        data = tf.pad(data, [[0,pad], [0,0]], constant_values=0.0)\n        non_empty_frames_idxs = tf.pad(non_empty_frames_idxs, [[0,pad]], constant_values=-1.0)\n````"
            },
            {
              "id": 2284819,
              "postDate": "2023-06-02T09:45:25.727Z",
              "content": "<p>Try this:</p>\n<pre><code>padding = tf.maximum(, N_TARGET_FRAMES - tf.shape(data0)[])\npad_tensor = tf.zeros([YOUR_TENSOR_SIZE]) \n\ndata = tf.concat([data0, pad_tensor], axis=)\n</code></pre>\n<p>I might be mistaken in dims or sizes in your code. But the main idea works</p>",
              "rawMarkdown": "Try this:\n\n```python\npadding = tf.maximum(0, N_TARGET_FRAMES - tf.shape(data0)[0])\npad_tensor = tf.zeros([YOUR_TENSOR_SIZE]) # like (pad_dim, ...other dims)\n    \ndata = tf.concat([data0, pad_tensor], axis=0)\n```\nI might be mistaken in dims or sizes in your code. But the main idea works"
            }
          ]
        }
      ]
    },
    {
      "id": 2279645,
      "postDate": "2023-05-29T13:54:07.630Z",
      "content": "<p>2023-05-29<br>\nToday's accomplishments:</p>\n<ul>\n<li>Delete one line of code, submit once, encounter a scoring error.</li>\n<li>Delete one line of code, submit once, encounter a scoring error.</li>\n<li>Delete one line of code, submit once, encounter a scoring error.</li>\n<li>Delete one line of code, submit once, encounter a scoring error.</li>\n<li>Delete one line of code, submit once, encounter a scoring error.</li>\n<li>no submission chance. go to bed.</li>\n</ul>\n<p>I haven't identified where the issue is occurring yet.</p>",
      "rawMarkdown": "2023-05-29\nToday's accomplishments:\n- Delete one line of code, submit once, encounter a scoring error.\n- Delete one line of code, submit once, encounter a scoring error.\n- Delete one line of code, submit once, encounter a scoring error.\n- Delete one line of code, submit once, encounter a scoring error.\n- Delete one line of code, submit once, encounter a scoring error.\n- no submission chance. go to bed.\n\nI haven't identified where the issue is occurring yet.",
      "votes": 2,
      "replies": [
        {
          "id": 2279668,
          "postDate": "2023-05-29T14:08:05.283Z",
          "content": "<p>I did almost the same but I added my blocks to <a href=\"https://www.kaggle.com/code/wonderingalice/working-sample-submission-and-inference\" target=\"_blank\">this dummy submission</a> . After I found the block with error, I added strings. Maybe adding would be faster. It took for me ~10 submits</p>\n<p>And yes, this expand_dims thing - just some random, weird stuff</p>",
          "rawMarkdown": "I did almost the same but I added my blocks to [this dummy submission](https://www.kaggle.com/code/wonderingalice/working-sample-submission-and-inference) . After I found the block with error, I added strings. Maybe adding would be faster. It took for me ~10 submits\n\nAnd yes, this expand_dims thing - just some random, weird stuff",
          "votes": 1
        }
      ]
    },
    {
      "id": 2296316,
      "postDate": "2023-06-11T17:58:07.297Z",
      "content": "<p>Has anyone resolved the </p>\n<p><code>RuntimeError: Select TensorFlow op(s), included in the given model, is(are) not supported by this interpreter. Make sure you apply/link the Flex delegate before inference. For the Android, it can be resolved by adding \"org.tensorflow:tensorflow-lite-select-tf-ops\" dependency. See instructions: https://www.tensorflow.org/lite/guide/ops_selectNode number 762 (FlexCTCGreedyDecoder) failed to prepare.</code></p>\n<p>problem?</p>",
      "rawMarkdown": "Has anyone resolved the \n\n`RuntimeError: Select TensorFlow op(s), included in the given model, is(are) not supported by this interpreter. Make sure you apply/link the Flex delegate before inference. For the Android, it can be resolved by adding \"org.tensorflow:tensorflow-lite-select-tf-ops\" dependency. See instructions: https://www.tensorflow.org/lite/guide/ops_selectNode number 762 (FlexCTCGreedyDecoder) failed to prepare.`\n\nproblem?",
      "replies": [
        {
          "id": 2296858,
          "postDate": "2023-06-12T07:54:46.617Z",
          "content": "<p>Unfortunately, the competition doesn't support models which need Tensorflow Select ops, so you can't enable this during compilation. The built-in CTC decoder is one of these unsupported ops. You sort of have to trial and error building models and seeing whether they are compilable or not</p>",
          "rawMarkdown": "Unfortunately, the competition doesn't support models which need Tensorflow Select ops, so you can't enable this during compilation. The built-in CTC decoder is one of these unsupported ops. You sort of have to trial and error building models and seeing whether they are compilable or not",
          "votes": 1
        }
      ]
    },
    {
      "id": 2285919,
      "postDate": "2023-06-03T05:25:27.753Z",
      "content": "<p>I use PyTorch, so after creating a PyTorch model, I'm converting the model to ONNX and then to TFLite by adding necessary operations.</p>\n<p>scoring fails when slicing operations are included.</p>\n<p>I cannot determine whether the concatenation operation is causing the problem as all five submission attempts were used up in 33 minutes. The reason I suspect slicing operations to be the issue is because during the preprocessing step, slicing operations were used to exclude the values along the z-axis, and problems occurred during that time as well.</p>\n<p>I'm not sure how many more times I will have to repeat the submission test in the future, as the preprocessing step involves various operations like torch.where.</p>\n<pre><code>\n\n (nn.Module):\n     ():\n        ().__init__()\n        self.model = SimpleModel()  \n\n        self.model.load_state_dict(net.state_dict())\n\n     ():\n        x[torch.isnan(x)] = \n        x = self.model(x)\n         x\n\n\n (nn.Module):\n     ():\n        ().__init__()\n        self.model = SimpleModel()  \n\n        self.model.load_state_dict(net.state_dict())\n\n     ():\n        part0 = x[:, :]\n        part1 = x[:, :]\n        x = torch.cat([part0, part1], dim=)\n        x[torch.isnan(x)] = \n        x = self.model(x)\n         x\n</code></pre>",
      "rawMarkdown": "I use PyTorch, so after creating a PyTorch model, I'm converting the model to ONNX and then to TFLite by adding necessary operations.\n\nscoring fails when slicing operations are included.\n\nI cannot determine whether the concatenation operation is causing the problem as all five submission attempts were used up in 33 minutes. The reason I suspect slicing operations to be the issue is because during the preprocessing step, slicing operations were used to exclude the values along the z-axis, and problems occurred during that time as well.\n\nI'm not sure how many more times I will have to repeat the submission test in the future, as the preprocessing step involves various operations like torch.where.\n\n```python\n\n# PASS\n\nclass ExportModel(nn.Module):\n    def __init__(self):\n        super().__init__()\n        self.model = SimpleModel()  # 1 linear layer\n        \n        self.model.load_state_dict(net.state_dict())\n        \n    def forward(self, x):\n        x[torch.isnan(x)] = 0.\n        x = self.model(x)\n        return x\n\n# FAIL\nclass ExportModel(nn.Module):\n    def __init__(self):\n        super().__init__()\n        self.model = SimpleModel()  # 1 linear layer\n        \n        self.model.load_state_dict(net.state_dict())\n        \n    def forward(self, x):\n        part0 = x[:, :50]\n        part1 = x[:, 50:]\n        x = torch.cat([part0, part1], dim=1)\n        x[torch.isnan(x)] = 0.\n        x = self.model(x)\n        return x\n```"
    },
    {
      "id": 2297667,
      "postDate": "2023-06-12T17:55:11.437Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2282575,
      "author_name": "Sohier Dane",
      "author_url": "",
      "post_date": "2023-05-31T16:50:02.327000",
      "content": "<p>I'll look into this.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2282635,
          "author_name": "Wondering Alice",
          "author_url": "",
          "post_date": "2023-05-31T17:50:36.803000",
          "content": "<p>Thank you!</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2280786,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "2023-05-30T10:51:54.700000",
      "content": "<p>From the evaluation page - \" Your model must be packaged into a submission.zip file and compatible with the TensorFlow Lite Runtime v2.9.1\"   I posted <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/411268#2267655\" target=\"_blank\">here</a> about forking an older notebook with Pin to Original Environment (prior to python 3.10 upgrade since TFLite does not have a version yet for it).  So if you are creating your submission .zip offline you can do that just to test in a notebook if your model works ok and any errors when you try to use it.  Found some tf operators do not work and you have to find a workaround for them. </p>\n<p>Regarding NaN frames or hands, using a preprocessing layer for selected columns that included frame - <br>\n<code>xlh = tf.slice(x0, [0,1], [-1,42 ])  # lh slice  x y (n.b. 0 is frame so use 1)</code> <br>\n<code>xrh = tf.slice(x0, [0,43], [-1,42 ])   # rh slice x y</code><br>\n<code>xrhv = tf.boolean_mask(xrh, tf.reduce_all(~tf.math.is_nan(xrh), axis=1)) # remove NaN rows</code><br>\n<code>xlhv =tf.boolean_mask(xlh, tf.reduce_all(~tf.math.is_nan(xlh), axis=1))</code><br>\n<code>x = tf.cond(tf.math.is_nan(tf.math.zero_fraction(xlhv)), lambda: tf.identity(xrhv), lambda: tf.identity(xlhv))</code></p>\n<p>Can confirm this works in submission without errors.  Fine if you use this in your code, but maybe don't just make a notebook using it. Pretty sure people can read the code here if they want.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2281029,
          "author_name": "Wondering Alice",
          "author_url": "",
          "post_date": "2023-05-30T14:25:26.060000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/something4kag\" target=\"_blank\">@something4kag</a> </p>\n<p>Thanks for sharing! This will definitely be inspiring …</p>\n<p>However, w.r.t the tflite runtime environment: I'm really not sure what else to try.<br>\nI have been working from my \"working submission notebook\" which is a fork of the previous competition (dating from Feb, I think) and explicitly installs Runtime v.2.9.1. It has Tensorflow 2.11.0 and Python 3.7.12. All the scoring errors I (and many others) I referred to in various discussion replies are from submissions for which TfLite inference works perfectly well in (a fork of) that notebook. After almost three weeks only 7 people have managed to get beyond dummy submissions or baselines, so clearly something is really unclear?</p>\n<p>Do you have any idea what else we need to change in order to make inference errors for the functions that are occurring in the scoring pipeline also visible during Kaggle notebook inference, so we can properly debug our models?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2279602,
      "author_name": "Mathieu De Coster",
      "author_url": "",
      "post_date": "2023-05-29T13:23:22.037000",
      "content": "<p>Certain operators not being supported, like <code>expand_dims</code>, is also making it difficult to port models fro PyTorch to Keras through ONNX, because you don't always control which TF operators are used.</p>\n<p>I'm currently in the process of porting parts of the model to Keras and seeing when the model gets through the submission pipeline, if I find any other operators that aren't supported I'll update this.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2279639,
          "author_name": "anokas",
          "author_url": "",
          "post_date": "2023-05-29T13:52:17.707000",
          "content": "<blockquote>\n  <p>Certain operators not being supported, like expand_dims</p>\n</blockquote>\n<p>Interesting, as my model which doesn't work uses <code>tf.expand_dims</code>. What makes you think it's not supported? (or is it just the comment by <a href=\"https://www.kaggle.com/forranto\" target=\"_blank\">@forranto</a>?)<br>\nSince the model can be converted to tflite and loaded/evaluated with the interpreter, I wonder what operations are supported by the official TFLite interpreter but not the one Kaggle is using.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2279649,
              "author_name": "Mathieu De Coster",
              "author_url": "",
              "post_date": "2023-05-29T13:56:44.813000",
              "content": "<p>Yes, I was referring to the comment by <a href=\"https://www.kaggle.com/forranto\" target=\"_blank\">@forranto</a>. It woud be really helpful if we could get the exact evaluation code that is used to score submissions (with library versions) to be able to figure out what's going wrong.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2279802,
          "author_name": "anokas",
          "author_url": "",
          "post_date": "2023-05-29T16:44:12.330000",
          "content": "<p>I've been looking at my model in <a href=\"https://netron.app/\" target=\"_blank\">https://netron.app/</a> (you can just upload your <code>.tflite</code>), and it looks like <code>x[None]</code> still compiles to an <code>ExpandDims</code> op in the actual model. Besides, a Conv1D just compiles to a Conv2D with ExpandDims:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F505747%2Fdbf58928d59fbda74ea961e20979cb61%2Fconv2d.png?generation=1685378548671571&amp;alt=media\" alt=\"\"></p>\n<p><a href=\"https://www.kaggle.com/kolyaforrat\" target=\"_blank\">@kolyaforrat</a> Would you be willing to share a list of TFLite Ops in your model, so we can see if there are some ops that are failing? Hopefully that doesn't give away anything about your solution (except perhaps whether it uses convolutions or recurrent layers at all, maybe this is secret).</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2279935,
              "author_name": "Kolya Forrat",
              "author_url": "",
              "post_date": "2023-05-29T18:22:30.040000",
              "content": "<p>Without revealing any secrets I can share netron part with differences between successfully submitted model and failed model.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4212496%2F56bb4405dcf9a4450974bfef56aaf838%2F2023-05-29%20%2021.08.46.png?generation=1685383933882709&amp;alt=media\" alt=\"expand dims\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4212496%2Fcecc605c4c526c46802889d451e292b8%2F2023-05-29%20%2021.08.58.png?generation=1685384192986888&amp;alt=media\" alt=\"\"></p>\n<p>That's the only difference between two graphs.</p>\n<p>About Conv2D and TFLite you can read in comments to <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406306\" target=\"_blank\">my team's write-up</a> from previous competition. <br>\nTLDR: not all conv types are good convertable to TFLite with torch-onnx-keras pipeline. But all are convertable with <a href=\"https://github.com/AlexanderLutsenko/nobuco\" target=\"_blank\">nobuco</a>. But in Kaggle environment some of them slower than our keras implemetation (But locally in different environment nobuco is better). So with <code>nobuco</code> no errors but slower</p>",
              "votes": 4,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2285420,
      "author_name": "Sohier Dane",
      "author_url": "",
      "post_date": "2023-06-02T17:40:08.610000",
      "content": "<p><a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/414682\" target=\"_blank\">Cross posting here</a>: it turns out that the metric has been using TF Lite runtime v2.14.0 rather than v2.9.1. Thank you all for this discussion and pushing us to track that down. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2283608,
      "author_name": "Wondering Alice",
      "author_url": "",
      "post_date": "2023-06-01T11:42:10.847000",
      "content": "<p>Hi all,</p>\n<p>Although I finally got a model without pre-or postprocessing through scoring, I'm still struggling with the same issues in pre- and post. <br>\nI now have two snippets of code, of the following structure (with x a <strong>2D float32 tensor</strong>):</p>\n<p><strong>Snippet A:</strong>   --  doesn't pass through scoring: </p>\n<pre><code>x2 = tf.math.reduce_sum(x, axis = )           \nmask = tf.not_equal(x2, )                     \nmasked_x = x[mask]                              \n</code></pre>\n<p><strong>Snippet B:</strong> --  passes through scoring: </p>\n<pre><code>x2 = tf.argmax(x, axis=)                              \nmask = tf.not_equal(x2,)                               \nmasked_x = x[mask]                                     \n</code></pre>\n<p>As always, stepping through these operations on dummy tensors produces the expected results and using the code with TfLite runtime (v2.9.1) works perfectly.</p>\n<p>As far as I can see the only difference is the reduce_sum operation, but I have found no indications anywhere online about this not being supported …</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2282859,
      "author_name": "Sohier Dane",
      "author_url": "",
      "post_date": "2023-05-31T21:35:31.773000",
      "content": "<p>A couple of thoughts:</p>\n<ul>\n<li>There is <a href=\"https://www.kaggle.com/code/irohith/aslfr-transformer\" target=\"_blank\">a public notebook that generates a non-static submission successfully</a>. I would encourage everyone to take a look at <a href=\"https://www.kaggle.com/irohith\" target=\"_blank\">@irohith</a>'s work.</li>\n<li>I dug into the error logs and it turns out that few different things are happening. A meaningful fraction of submissions still include ops that aren't supported by the TF Lite runtime. Others have invalid signatures, so I updated the evaluation tab to be more explicit about the requirements there. However the majority of recent errors are dimension mismatch errors where I would need access to the model's training code in order to debug further.</li>\n</ul>\n<p>Given the limited attention that the aslfr-transformer notebook has gotten so far I'm going to pause my review for now. I'll check in again next week once everyone has had time to digest that successful example.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2283476,
          "author_name": "Mark Wijkhuizen",
          "author_url": "",
          "post_date": "2023-06-01T09:20:30.583000",
          "content": "<p>How could we verify our models do not include operators that are not supported by the TFLite runtime?<br>\nI am currently using the following script, but my submissions still fail:</p>\n<p>In addition, I verify my TFLite model is working correctly by performing inference on my <strong>full</strong> training dataset, which runs without errors. It is quite a mystery why the submission still fails.</p>\n<p><strong>Edit</strong></p>\n<p>It seems to me the TFLite environment we use is different from the one used during inference. We therefore have no method to verify whether our inference loop is correct.</p>\n<pre><code># Create Model Converter\nkeras_model_converter = tf.lite.TFLiteConverter.from_keras_model(tflite_keras_model)\n# Convert Model\ntflite_model = keras_model_converter.convert()\n# Write Model\nwith open('/kaggle/working/model.tflite', 'wb') as f:\n    f.write(tflite_model)\n</code></pre>",
          "votes": 6,
          "replies": [
            {
              "id": 2284837,
              "author_name": "Mathieu De Coster",
              "author_url": "",
              "post_date": "2023-06-02T09:54:37.347000",
              "content": "<p>Either the runtime environment is different, or there are some edge cases in the test set which the model can't handle (empty samples for example).</p>\n<p>I have tried taking these edge cases into account, but the submissions still failed.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2284772,
          "author_name": "anokas",
          "author_url": "",
          "post_date": "2023-06-02T09:08:55.280000",
          "content": "<blockquote>\n  <p>include ops that aren't supported by the TF Lite runtime</p>\n</blockquote>\n<p>Thanks for looking into this, but this is the part that confuses me. When we use the TFLite runtime V2.9.1 locally, <strong>all our ops are supported fine</strong> - so it appears the kaggle tflite runtime is different to the one that is publicly available?</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 2285249,
          "author_name": "anokas",
          "author_url": "",
          "post_date": "2023-06-02T15:00:24.977000",
          "content": "<p>I've just noticed something you mentioned:</p>\n<blockquote>\n  <p>However the majority of recent errors are dimension mismatch errors where I would need access to the model's training code in order to debug further.</p>\n</blockquote>\n<p>Could it be that the evaluation code fails when an empty prediction is made?</p>\n<p>The published evaluation code handles empty predictions fine, but perhaps this is the source of one difference between the specification and the implementation.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2285409,
              "author_name": "Sohier Dane",
              "author_url": "",
              "post_date": "2023-06-02T17:30:30.743000",
              "content": "<p>Most of the dimension mismatches are between non-zero values that don't have any obvious connection to the original dataset shape, like <code>2 != 1</code> or <code>85 != 84</code>.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2285730,
              "author_name": "anokas",
              "author_url": "",
              "post_date": "2023-06-03T00:05:48.853000",
              "content": "<p>Interesting, thank you for double checking!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2279903,
      "author_name": "Wondering Alice",
      "author_url": "",
      "post_date": "2023-05-29T18:04:51.013000",
      "content": "<p>I've spent my 5 attempts today trying to get a working solution for dropping frames (e.g. the missing hands frames). For this, too, code from the previous competition works for notebook inference but gives a scoring error when I try to submit. I tried making sure that at least one frame remains, but that didn't help. </p>\n<p><strong>It would be wonderful if anyone would be willing to share a working solution for that.</strong></p>\n<p>In return, I can confirm that the <code>x[None]</code> replacement for expand_dims works perfectly. Also, do not forget to remove that initial dimension again from the output of your batch-trained model, i.e., assuming your model, called  <code>tmodel</code> below, probably returns a tensor of shape [None,None,59], so the submission model could look like:</p>\n<pre><code>trained_model = &lt;your_model_path&gt;\ndummy_inner = tf.keras.models.load_model(trained_model)\n\n ():\n    inputs = tf.keras.Input(shape=(NUM_FEATURES), dtype=tf.float32, name=)\n    \n    x = tf.where(tf.math.is_nan(inputs), tf.zeros_like(inputs), inputs)\n    \n    x = x[]\n    \n    x = tmodel(x)\n    \n    out = tf.keras.layers.Activation(, name=)(x[,:,:])\n\ndummy_submission_model = get_dummy_model(dummy_inner)\ndummy_submission_model.summary()\n</code></pre>\n<p>(updated: initially I copy-pasted the wrong version, sorry)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2280653,
          "author_name": "Wondering Alice",
          "author_url": "",
          "post_date": "2023-05-30T08:44:55.623000",
          "content": "<p>Update: trying to <strong>unsuccessfully</strong> remove missing hand frames (and get it through the scoring pipeline) landed here today:</p>\n<pre><code>\nhands_ok = tf.math.add(x[:,:],x[:,:])\n(, hands_ok.shape)\n\n\nhands_ok = tf.math.reduce_sum(hands_ok, axis = )    \n(,hands_ok.shape)   \n\n\nhands_ok = tf.cast(hands_ok, tf.)\n(,hands_ok.shape)\n\n\nmasked_x = x[hands_ok]\n(,x.shape)\n\n\nx = tf.concat([x[:,:],masked_x],axis = )\n(,x.shape) \n</code></pre>\n<p>Alternatives I tried today used <code>tf.where</code> or <code>tf.boolean_mask</code> or did not use the concat line or layers.Concatenate instead.<br>\n(all worked perfectly when doing tflite inference in notebook, all failed submission scoring, 5 attempts for today used up)</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2284628,
              "author_name": "Mark Wijkhuizen",
              "author_url": "",
              "post_date": "2023-06-02T07:25:08.390000",
              "content": "<p>I can confirm <code>pad</code> and <code>concat</code> do not work. In the previous Isolated Sign Language competition those operators did work. The TFLite conversion is also successful. No idea what is causing the submission to fail.</p>\n<p>Without proper preprocessing this whole competition is more about making a model that works on unprocessed data than actually creating a high performing model with advanced preprocessing, which should not be the point of this competition.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2284737,
              "author_name": "Wondering Alice",
              "author_url": "",
              "post_date": "2023-06-02T08:47:35.210000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/Mark\" target=\"_blank\">@Mark</a>, </p>\n<p>Are you using Tensorflow or converting from Pytorch?</p>\n<p>I do use a concat operation in my current submissions (all built and trained in Tensorflow), so it should work. I have, however, noticed at least once that there was in fact a dimension mismatch in the mask I used due to an error I made. This passed TfLite creation and TfLite inference in my notebook but not the scoring pipeline. Removing the error resulted in successful submission of that bit. </p>\n<p>What I'm doing now is checking each operation by stepping through the whole model and printing output tensors to screen. As you can see from the snippets in another comment, I am stuck at \"reduce_sum\", which I need for dominant hand detection. Tensors look fine but submissions keep failing.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2284744,
              "author_name": "Kolya Forrat",
              "author_url": "",
              "post_date": "2023-06-02T08:53:21.023000",
              "content": "<p><code>x = tf.concat([x,  y], axis=0)</code><br>\nConcat is work for me. Pad is actually the same, I tried pad, it also works with <code>concat</code> operation</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2284814,
              "author_name": "Mark Wijkhuizen",
              "author_url": "",
              "post_date": "2023-06-02T09:39:41.843000",
              "content": "<p>I am using a Tensorflow preprocessing layer which keeps failing:</p>\n<pre><code>class PreprocessLayer(tf.keras.layers.Layer):\n    def __init__(self):\n        super(PreprocessLayer, self).__init__()\n\n    @tf.function(\n        input_signature=(tf.TensorSpec(shape=[None,N_COLS0], dtype=tf.float32),),\n    )\n    def call(self, data0):        \n        # Number of Frames in Video\n        N_FRAMES0 = tf.shape(data0)[0]\n\n        # Fill NaN Values With 0\n        data = tf.where(tf.math.is_nan(data0), 0.0, data0)\n\n        # THIS KEEPS FAILING!\n        # should simply pad videos shorter than N_TARGET_FRAMES (128) frames\n        pad = tf.math.maximum(0, N_TARGET_FRAMES - N_FRAMES0)\n        data = tf.pad(data, [[0,pad], [0,0]], constant_values=0.0)\n        non_empty_frames_idxs = tf.pad(non_empty_frames_idxs, [[0,pad]], constant_values=-1.0)\n</code></pre>\n<p>`</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2284819,
              "author_name": "Kolya Forrat",
              "author_url": "",
              "post_date": "2023-06-02T09:45:25.727000",
              "content": "<p>Try this:</p>\n<pre><code>padding = tf.maximum(, N_TARGET_FRAMES - tf.shape(data0)[])\npad_tensor = tf.zeros([YOUR_TENSOR_SIZE]) \n\ndata = tf.concat([data0, pad_tensor], axis=)\n</code></pre>\n<p>I might be mistaken in dims or sizes in your code. But the main idea works</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2279645,
      "author_name": "canlion",
      "author_url": "",
      "post_date": "2023-05-29T13:54:07.630000",
      "content": "<p>2023-05-29<br>\nToday's accomplishments:</p>\n<ul>\n<li>Delete one line of code, submit once, encounter a scoring error.</li>\n<li>Delete one line of code, submit once, encounter a scoring error.</li>\n<li>Delete one line of code, submit once, encounter a scoring error.</li>\n<li>Delete one line of code, submit once, encounter a scoring error.</li>\n<li>Delete one line of code, submit once, encounter a scoring error.</li>\n<li>no submission chance. go to bed.</li>\n</ul>\n<p>I haven't identified where the issue is occurring yet.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2279668,
          "author_name": "Kolya Forrat",
          "author_url": "",
          "post_date": "2023-05-29T14:08:05.283000",
          "content": "<p>I did almost the same but I added my blocks to <a href=\"https://www.kaggle.com/code/wonderingalice/working-sample-submission-and-inference\" target=\"_blank\">this dummy submission</a> . After I found the block with error, I added strings. Maybe adding would be faster. It took for me ~10 submits</p>\n<p>And yes, this expand_dims thing - just some random, weird stuff</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2296316,
      "author_name": "yuanzhe zhou",
      "author_url": "",
      "post_date": "2023-06-11T17:58:07.297000",
      "content": "<p>Has anyone resolved the </p>\n<p><code>RuntimeError: Select TensorFlow op(s), included in the given model, is(are) not supported by this interpreter. Make sure you apply/link the Flex delegate before inference. For the Android, it can be resolved by adding \"org.tensorflow:tensorflow-lite-select-tf-ops\" dependency. See instructions: https://www.tensorflow.org/lite/guide/ops_selectNode number 762 (FlexCTCGreedyDecoder) failed to prepare.</code></p>\n<p>problem?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2296858,
          "author_name": "anokas",
          "author_url": "",
          "post_date": "2023-06-12T07:54:46.617000",
          "content": "<p>Unfortunately, the competition doesn't support models which need Tensorflow Select ops, so you can't enable this during compilation. The built-in CTC decoder is one of these unsupported ops. You sort of have to trial and error building models and seeing whether they are compilable or not</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2285919,
      "author_name": "canlion",
      "author_url": "",
      "post_date": "2023-06-03T05:25:27.753000",
      "content": "<p>I use PyTorch, so after creating a PyTorch model, I'm converting the model to ONNX and then to TFLite by adding necessary operations.</p>\n<p>scoring fails when slicing operations are included.</p>\n<p>I cannot determine whether the concatenation operation is causing the problem as all five submission attempts were used up in 33 minutes. The reason I suspect slicing operations to be the issue is because during the preprocessing step, slicing operations were used to exclude the values along the z-axis, and problems occurred during that time as well.</p>\n<p>I'm not sure how many more times I will have to repeat the submission test in the future, as the preprocessing step involves various operations like torch.where.</p>\n<pre><code>\n\n (nn.Module):\n     ():\n        ().__init__()\n        self.model = SimpleModel()  \n\n        self.model.load_state_dict(net.state_dict())\n\n     ():\n        x[torch.isnan(x)] = \n        x = self.model(x)\n         x\n\n\n (nn.Module):\n     ():\n        ().__init__()\n        self.model = SimpleModel()  \n\n        self.model.load_state_dict(net.state_dict())\n\n     ():\n        part0 = x[:, :]\n        part1 = x[:, :]\n        x = torch.cat([part0, part1], dim=)\n        x[torch.isnan(x)] = \n        x = self.model(x)\n         x\n</code></pre>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2297667,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-06-12T17:55:11.437000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2279482": "Most people are having issues getting submissions to work in this competition, even those that had no problems in the previous competition. While I was able to submit a constant prediction model without issues, I've got another model which gets 0.70 CV and works perfectly following the [official evaluation process](https://www.kaggle.com/competitions/asl-fingerspelling/overview/evaluation) but fails on Kaggle.\n\nI wanted to make a thread collating both potential ideas for fixes and working submission code so that we can figure out what's going on. At the end of the day, it would be great if something could be done from the organisers side @thadstarner @sohierto to improve this - for example:\n- allowing us to see any errors that occur _before_ our model sees test data (for example, if the tflite file fails to load). This would prevent people from being able to learn anything about the test data (the goal of code submission) but make the process much less annoying. \n- giving us a representative evaluation code/environment so that we can test our models locally. It's clear that there's some extra requirements that aren't listed on the Evaluation page.\n\nIn the meantime, some resources to help us figure this out:\n\n**Working submissions:**\n- 1-layer dummy model with `(N_FRAMES, 59)` output: https://www.kaggle.com/code/wonderingalice/working-sample-submission-and-inference @wonderingalice\n- Model with static `(11, 59)` output size: https://www.kaggle.com/code/anokas/static-greedy-baseline-0-157-lb\n- 1-layer dummy PyTorch model: https://www.kaggle.com/code/chack3/pytorch-simple-model @chack3\n\n**Possible fixes:**\n- Changing nan-infill and expand_dims method: https://www.kaggle.com/competitions/asl-fingerspelling/discussion/409722#2276829 @forrato It's really unclear to me why this would help.\n- Making sure model can handle empty input: https://www.kaggle.com/competitions/asl-fingerspelling/discussion/415372\n\nIf anyone has any ideas or other examples, let me know and I'll try to keep this thread updated! From my perspective, I'll probably have to take a break until we get some clarity from the organisers 😅",
    "2282575": "I'll look into this.",
    "2280786": "From the evaluation page - \" Your model must be packaged into a submission.zip file and compatible with the TensorFlow Lite Runtime v2.9.1\"   I posted [here](https://www.kaggle.com/competitions/asl-fingerspelling/discussion/411268#2267655) about forking an older notebook with Pin to Original Environment (prior to python 3.10 upgrade since TFLite does not have a version yet for it).  So if you are creating your submission .zip offline you can do that just to test in a notebook if your model works ok and any errors when you try to use it.  Found some tf operators do not work and you have to find a workaround for them. \n\nRegarding NaN frames or hands, using a preprocessing layer for selected columns that included frame - \n`xlh = tf.slice(x0, [0,1], [-1,42 ])  # lh slice  x y (n.b. 0 is frame so use 1)` \n`xrh = tf.slice(x0, [0,43], [-1,42 ])   # rh slice x y`\n`xrhv = tf.boolean_mask(xrh, tf.reduce_all(~tf.math.is_nan(xrh), axis=1)) # remove NaN rows`\n`xlhv =tf.boolean_mask(xlh, tf.reduce_all(~tf.math.is_nan(xlh), axis=1))`\n`x = tf.cond(tf.math.is_nan(tf.math.zero_fraction(xlhv)), lambda: tf.identity(xrhv), lambda: tf.identity(xlhv))`\n\nCan confirm this works in submission without errors.  Fine if you use this in your code, but maybe don't just make a notebook using it. Pretty sure people can read the code here if they want.",
    "2279602": "Certain operators not being supported, like `expand_dims`, is also making it difficult to port models fro PyTorch to Keras through ONNX, because you don't always control which TF operators are used.\n\nI'm currently in the process of porting parts of the model to Keras and seeing when the model gets through the submission pipeline, if I find any other operators that aren't supported I'll update this.",
    "2285420": "[Cross posting here](https://www.kaggle.com/competitions/asl-fingerspelling/discussion/414682): it turns out that the metric has been using TF Lite runtime v2.14.0 rather than v2.9.1. Thank you all for this discussion and pushing us to track that down. ",
    "2283608": "Hi all,\n\nAlthough I finally got a model without pre-or postprocessing through scoring, I'm still struggling with the same issues in pre- and post. \nI now have two snippets of code, of the following structure (with x a **2D float32 tensor**):\n\n**Snippet A:**   --  doesn't pass through scoring: \n\n```python\nx2 = tf.math.reduce_sum(x, axis = 1)           # 1D float32 tensor\nmask = tf.not_equal(x2, 0)                     # 1D boolean mask  -- already tried comparing with 0.0 as well\nmasked_x = x[mask]                              # 1D float32 tensor\n```\n\n**Snippet B:** --  passes through scoring: \n\n```python\nx2 = tf.argmax(x, axis=1)                              # 1D int64 tensor        \nmask = tf.not_equal(x2,0)                               # 1D boolean mask\nmasked_x = x[mask]                                     # 1D float32 tensor\n```\n\nAs always, stepping through these operations on dummy tensors produces the expected results and using the code with TfLite runtime (v2.9.1) works perfectly.\n\nAs far as I can see the only difference is the reduce_sum operation, but I have found no indications anywhere online about this not being supported ...\n\n",
    "2282859": "A couple of thoughts:\n\n- There is [a public notebook that generates a non-static submission successfully](https://www.kaggle.com/code/irohith/aslfr-transformer). I would encourage everyone to take a look at @irohith's work.\n- I dug into the error logs and it turns out that few different things are happening. A meaningful fraction of submissions still include ops that aren't supported by the TF Lite runtime. Others have invalid signatures, so I updated the evaluation tab to be more explicit about the requirements there. However the majority of recent errors are dimension mismatch errors where I would need access to the model's training code in order to debug further.\n\nGiven the limited attention that the aslfr-transformer notebook has gotten so far I'm going to pause my review for now. I'll check in again next week once everyone has had time to digest that successful example.",
    "2279903": "I've spent my 5 attempts today trying to get a working solution for dropping frames (e.g. the missing hands frames). For this, too, code from the previous competition works for notebook inference but gives a scoring error when I try to submit. I tried making sure that at least one frame remains, but that didn't help. \n\n**It would be wonderful if anyone would be willing to share a working solution for that.**\n\nIn return, I can confirm that the `x[None]` replacement for expand_dims works perfectly. Also, do not forget to remove that initial dimension again from the output of your batch-trained model, i.e., assuming your model, called  `tmodel` below, probably returns a tensor of shape [None,None,59], so the submission model could look like:\n\n````python\n\ntrained_model = <your_model_path>\ndummy_inner = tf.keras.models.load_model(trained_model)\n\ndef get_dummy_model(tmodel):\n    inputs = tf.keras.Input(shape=(NUM_FEATURES), dtype=tf.float32, name=\"inputs\")\n    # possibly add additional cleaning or preprocessing code below\n    x = tf.where(tf.math.is_nan(inputs), tf.zeros_like(inputs), inputs)\n    # replacement for expand_dims since that doesn't seem to make it through the scoring pipeline\n    x = x[None]\n    # Insert trained model\n    x = tmodel(x)\n    # possibly add postprocessing code below\n    out = tf.keras.layers.Activation(\"linear\", name=\"outputs\")(x[0,:,:])\n\ndummy_submission_model = get_dummy_model(dummy_inner)\ndummy_submission_model.summary()\n````\n\n(updated: initially I copy-pasted the wrong version, sorry)",
    "2279645": "2023-05-29\nToday's accomplishments:\n- Delete one line of code, submit once, encounter a scoring error.\n- Delete one line of code, submit once, encounter a scoring error.\n- Delete one line of code, submit once, encounter a scoring error.\n- Delete one line of code, submit once, encounter a scoring error.\n- Delete one line of code, submit once, encounter a scoring error.\n- no submission chance. go to bed.\n\nI haven't identified where the issue is occurring yet.",
    "2296316": "Has anyone resolved the \n\n`RuntimeError: Select TensorFlow op(s), included in the given model, is(are) not supported by this interpreter. Make sure you apply/link the Flex delegate before inference. For the Android, it can be resolved by adding \"org.tensorflow:tensorflow-lite-select-tf-ops\" dependency. See instructions: https://www.tensorflow.org/lite/guide/ops_selectNode number 762 (FlexCTCGreedyDecoder) failed to prepare.`\n\nproblem?",
    "2285919": "I use PyTorch, so after creating a PyTorch model, I'm converting the model to ONNX and then to TFLite by adding necessary operations.\n\nscoring fails when slicing operations are included.\n\nI cannot determine whether the concatenation operation is causing the problem as all five submission attempts were used up in 33 minutes. The reason I suspect slicing operations to be the issue is because during the preprocessing step, slicing operations were used to exclude the values along the z-axis, and problems occurred during that time as well.\n\nI'm not sure how many more times I will have to repeat the submission test in the future, as the preprocessing step involves various operations like torch.where.\n\n```python\n\n# PASS\n\nclass ExportModel(nn.Module):\n    def __init__(self):\n        super().__init__()\n        self.model = SimpleModel()  # 1 linear layer\n        \n        self.model.load_state_dict(net.state_dict())\n        \n    def forward(self, x):\n        x[torch.isnan(x)] = 0.\n        x = self.model(x)\n        return x\n\n# FAIL\nclass ExportModel(nn.Module):\n    def __init__(self):\n        super().__init__()\n        self.model = SimpleModel()  # 1 linear layer\n        \n        self.model.load_state_dict(net.state_dict())\n        \n    def forward(self, x):\n        part0 = x[:, :50]\n        part1 = x[:, 50:]\n        x = torch.cat([part0, part1], dim=1)\n        x[torch.isnan(x)] = 0.\n        x = self.model(x)\n        return x\n```",
    "2297667": ""
  }
}