{
  "id": 410459,
  "title": "Nothing on leaderboard yet?",
  "url": "/competitions/asl-fingerspelling/discussion/410459",
  "author_name": "Wondering Alice",
  "post_date": "2023-05-15T13:36:57.340000",
  "votes": 19,
  "comment_count": 42,
  "views": 0,
  "content": "<p>Dear all, </p>\n<p>When I click on the leaderboard, it is empty.<br>\nIs it just me or have no submissions been made yet?<br>\nI find it strange that there aren't even any \"hello world\" baselines yet. Is there something wrong with the data maybe?</p>",
  "messages": [
    {
      "id": 2260138,
      "postDate": "2023-05-15T13:36:57.340Z",
      "content": "<p>Dear all, </p>\n<p>When I click on the leaderboard, it is empty.<br>\nIs it just me or have no submissions been made yet?<br>\nI find it strange that there aren't even any \"hello world\" baselines yet. Is there something wrong with the data maybe?</p>",
      "rawMarkdown": "Dear all, \n\nWhen I click on the leaderboard, it is empty.\nIs it just me or have no submissions been made yet?\nI find it strange that there aren't even any \"hello world\" baselines yet. Is there something wrong with the data maybe?\n",
      "votes": 19
    },
    {
      "id": 2260397,
      "postDate": "2023-05-15T16:30:42.423Z",
      "content": "<p>Because the training set is too large, we still need some time for model training.</p>",
      "rawMarkdown": "Because the training set is too large, we still need some time for model training.",
      "votes": 3,
      "replies": [
        {
          "id": 2260446,
          "postDate": "2023-05-15T16:57:35.657Z",
          "content": "<p>It makes sense, but what about the fact, that in \"Google Research - Identify Contrails to Reduce Global Warming\" there are 60+ teams, while the dataset is more than twice bigger?</p>",
          "rawMarkdown": "It makes sense, but what about the fact, that in \"Google Research - Identify Contrails to Reduce Global Warming\" there are 60+ teams, while the dataset is more than twice bigger?",
          "votes": 1,
          "replies": [
            {
              "id": 2260528,
              "postDate": "2023-05-15T17:45:39.560Z",
              "content": "<p>It's trivial to make a dummy submission to the contrails competition (or almost any other competition for that matter). The bar for an initial submission is higher in this case since someone will need to prepare a working model first.</p>",
              "rawMarkdown": "It's trivial to make a dummy submission to the contrails competition (or almost any other competition for that matter). The bar for an initial submission is higher in this case since someone will need to prepare a working model first.",
              "votes": 1
            },
            {
              "id": 2265537,
              "postDate": "2023-05-19T10:15:43.510Z",
              "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> it should be quite trivial to make an initial random submission here as well, but as you can read in some of the other discussion threads, all (very well informed) attempts to do so are failing thus far. Could you please clarify the required input and output formats:</p>\n<ul>\n<li>input tensor format to TfLite model (i.e. is time in batch dimension as in previous competition or not)<br>\nmost common assumption is currently [num_frames,1630], with num_frames 'abusing' the batch dimension</li>\n<li>desired output tensor format to TfLite model: <br>\nper frame of per sequence? <br>\ntime as separate dimension or in batch dimension? <br>\nshould output length (num_frames)  be = input length (num_frames) ? or &lt;= than input length? or    any other constraint?<br>\nanything else that can be responsible for crashes?</li>\n<li>how to install and run tflite interpreter in a Kaggle kernel to at least be able to verify/debug our own tflite models </li>\n<li>whether the problems could be due to a version mismatch between the tflite models created by the Tensorflow version installed in the Kaggle kernels and the required version for inference (which is also quite old)</li>\n</ul>",
              "rawMarkdown": "@sohier it should be quite trivial to make an initial random submission here as well, but as you can read in some of the other discussion threads, all (very well informed) attempts to do so are failing thus far. Could you please clarify the required input and output formats:\n\n- input tensor format to TfLite model (i.e. is time in batch dimension as in previous competition or not)\n   most common assumption is currently [num_frames,1630], with num_frames 'abusing' the batch dimension\n- desired output tensor format to TfLite model: \n   per frame of per sequence? \n   time as separate dimension or in batch dimension? \n   should output length (num_frames)  be = input length (num_frames) ? or <= than input length? or    any other constraint?\n   anything else that can be responsible for crashes?\n- how to install and run tflite interpreter in a Kaggle kernel to at least be able to verify/debug our own tflite models \n- whether the problems could be due to a version mismatch between the tflite models created by the Tensorflow version installed in the Kaggle kernels and the required version for inference (which is also quite old)\n",
              "votes": 3
            }
          ]
        }
      ]
    },
    {
      "id": 2266183,
      "postDate": "2023-05-19T19:31:03.803Z",
      "content": "<p>Thank you for raising this concern. We reviewed the metric and identified that a necessary dataset had failed to properly attach. Submissions should now be working.</p>",
      "rawMarkdown": "Thank you for raising this concern. We reviewed the metric and identified that a necessary dataset had failed to properly attach. Submissions should now be working.",
      "votes": 4,
      "replies": [
        {
          "id": 2266361,
          "postDate": "2023-05-20T00:58:22.053Z",
          "content": "<p>However it still doesn’t work on my side. I am wondering if anyone can make successful submission now.</p>",
          "rawMarkdown": "However it still doesn’t work on my side. I am wondering if anyone can make successful submission now.",
          "votes": 2
        },
        {
          "id": 2266414,
          "postDate": "2023-05-20T03:50:43.543Z",
          "content": "<p>I still can't submit a valid submission. Is it possible to know what are the input shapes by defaults? n_framesx1630?<br>\nthe output shape should be n_charactersx59?</p>",
          "rawMarkdown": "I still can't submit a valid submission. Is it possible to know what are the input shapes by defaults? n_framesx1630?\nthe output shape should be n_charactersx59?",
          "votes": 2
        },
        {
          "id": 2266588,
          "postDate": "2023-05-20T07:44:25.967Z",
          "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> Thank you for looking into this. I still can't make any valid submissions. Could you please give the exact formats for the input data passed to the model and the output format expected? Or even better, provide a notebook with a working submission? That would really help us get started on this competition.</p>",
          "rawMarkdown": "@sohier Thank you for looking into this. I still can't make any valid submissions. Could you please give the exact formats for the input data passed to the model and the output format expected? Or even better, provide a notebook with a working submission? That would really help us get started on this competition.",
          "votes": 3
        },
        {
          "id": 2266698,
          "postDate": "2023-05-20T09:23:47.740Z",
          "content": "<p>Sadly, my submissions keep failing as well. I agree with <a href=\"https://www.kaggle.com/mdecoster\" target=\"_blank\">@mdecoster</a> that providing an example submission notebook would be very helpful.</p>",
          "rawMarkdown": "Sadly, my submissions keep failing as well. I agree with @mdecoster that providing an example submission notebook would be very helpful.",
          "votes": 3
        },
        {
          "id": 2269074,
          "postDate": "2023-05-22T07:48:48.437Z",
          "content": "<p>can you please fix the doc of the evaluation page and let us know the expected signature (especially please define REQUIRED_OUTPUT variable)</p>",
          "rawMarkdown": "can you please fix the doc of the evaluation page and let us know the expected signature (especially please define REQUIRED_OUTPUT variable)",
          "votes": 1
        }
      ]
    },
    {
      "id": 2267501,
      "postDate": "2023-05-21T03:03:49.247Z",
      "content": "<p>Everyone got what should be in <code>REQUIRED_SIGNATURE</code> and <code>REQUIRED_OUTPUT</code>? <br>\nAs described in Evaluation part</p>",
      "rawMarkdown": "Everyone got what should be in `REQUIRED_SIGNATURE` and `REQUIRED_OUTPUT`? \nAs described in Evaluation part",
      "votes": 1,
      "replies": [
        {
          "id": 2267528,
          "postDate": "2023-05-21T03:49:26.110Z",
          "content": "<p>My understanding is that <code>REQUIRED_SIGNATURE</code> = <code>serving_default</code> and <code>REQUIRED_OUTPUT</code> = <code>outputs</code>.</p>",
          "rawMarkdown": "My understanding is that `REQUIRED_SIGNATURE` = `serving_default` and `REQUIRED_OUTPUT` = `outputs`.",
          "votes": 1,
          "replies": [
            {
              "id": 2267993,
              "postDate": "2023-05-21T11:28:56.880Z",
              "content": "<p>Yes, at least, that was the interpretation in the previous competition.</p>",
              "rawMarkdown": "Yes, at least, that was the interpretation in the previous competition."
            }
          ]
        },
        {
          "id": 2267757,
          "postDate": "2023-05-21T07:59:46.820Z",
          "content": "<p>i suppose that REQUIRED_SIGNATURE = serving_default, but REQUIRED_OUTPUT is number of elements in the true string, because our evaluation metric for this contest is the normalized total levenshtein distance and  number of characters in the labels be N and the total levenshtein distance be D. The metric equals (N - D) / N. I believe that organizers want to have range of that metric like [0, 1]. <br>\nFor instance if we have 2 strings real and estiamted: y_true = \"i can see the rings on saturn\", and y_pred = \"i can see the rings on saturn s\" N = len(y_true) , the Levinstein distance is calculated on (y_true, y_pred[:N]). But it is on only my guess i could be wrong. By the way i'm also interest what is REQUIRED_OUTPUT.</p>",
          "rawMarkdown": "i suppose that REQUIRED_SIGNATURE = serving_default, but REQUIRED_OUTPUT is number of elements in the true string, because our evaluation metric for this contest is the normalized total levenshtein distance and  number of characters in the labels be N and the total levenshtein distance be D. The metric equals (N - D) / N. I believe that organizers want to have range of that metric like [0, 1]. \nFor instance if we have 2 strings real and estiamted: y_true = \"i can see the rings on saturn\", and y_pred = \"i can see the rings on saturn s\" N = len(y_true) , the Levinstein distance is calculated on (y_true, y_pred[:N]). But it is on only my guess i could be wrong. By the way i'm also interest what is REQUIRED_OUTPUT.",
          "replies": [
            {
              "id": 2268000,
              "postDate": "2023-05-21T11:35:49.087Z",
              "content": "<p>In our interpretation,  REQUIRED_OUTPUT is \"outputs\", i.e. the name you need to give to the last layer in your model (i.e., the one that produces the outputs).</p>\n<p>I have posted a demo notebook to show how TfLite inference can be done <a href=\"https://www.kaggle.com/code/wonderingalice/dummy-submission-with-previous-kernel?kernelSessionId=130400757\" target=\"_blank\">here</a></p>",
              "rawMarkdown": "In our interpretation,  REQUIRED_OUTPUT is \"outputs\", i.e. the name you need to give to the last layer in your model (i.e., the one that produces the outputs).\n\nI have posted a demo notebook to show how TfLite inference can be done [here](https://www.kaggle.com/code/wonderingalice/dummy-submission-with-previous-kernel?kernelSessionId=130400757)",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2264709,
      "postDate": "2023-05-18T16:55:14.017Z",
      "content": "<p>That’s very tricky. I created a random model that takes left hand + right hand landmarks data as input for submission. It can make inferences with parquet files. However submission failed after running about one hour.</p>",
      "rawMarkdown": "That’s very tricky. I created a random model that takes left hand + right hand landmarks data as input for submission. It can make inferences with parquet files. However submission failed after running about one hour.",
      "votes": 1,
      "replies": [
        {
          "id": 2264765,
          "postDate": "2023-05-18T17:23:39.807Z",
          "content": "<p>Hi Lonnie,</p>\n<p>That's already better than my attempts: this suggests that scoring works for some but not all samples?</p>",
          "rawMarkdown": "Hi Lonnie,\n\nThat's already better than my attempts: this suggests that scoring works for some but not all samples?",
          "replies": [
            {
              "id": 2265084,
              "postDate": "2023-05-19T01:11:11.760Z",
              "content": "<p>I am not sure. Everything could be possible. It may be some kinds of pipeline issues or dataset issues. But in my notebook it takes 2 minutes to make inferences and evaluation with <code>tf.edit_distance</code> method on train files and supplemental files with low memory footprint. </p>",
              "rawMarkdown": "I am not sure. Everything could be possible. It may be some kinds of pipeline issues or dataset issues. But in my notebook it takes 2 minutes to make inferences and evaluation with `tf.edit_distance` method on train files and supplemental files with low memory footprint. "
            }
          ]
        },
        {
          "id": 2265517,
          "postDate": "2023-05-19T09:54:34.280Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/lonnieqin\" target=\"_blank\">@lonnieqin</a> ,</p>\n<p>I tried to make a dummy solution, using all features, and based on the submission model format from the previous competition. I have tried both:</p>\n<p>inputs = tf.keras.Input(shape=(543,3), name=\"inputs\")</p>\n<p>and </p>\n<p>inputs = tf.keras.Input(shape=(None, 543,3), name=\"inputs\")</p>\n<p>As input shapes (using all features for now) and I think I've tried all possible output shape tensor combinations, i.e.:</p>\n<ul>\n<li>with and without hacking batch dimension as time dimension (as in the previous competition)</li>\n<li>I also tried per-frame models (although that would be horribly restrictive)</li>\n</ul>\n<p>Since all these attempts (at a rate of 5 totally uninformative  \"random guesses\" per day) failed after about one minute, this suggests they don't even start getting scored. I am assuming there is still something wrong with my interpretation of the input tensor shape and/or required output tensor shape or properties (e.g. data format)?</p>\n<p>So, since it seems that your attempts at least get partially scored: could you clarify what the input and output tensor shapes are in your model?</p>\n<p><strong>BTW: I do feel that it is the responsibility of competition organisers to clarify the desired interface of submitted models. It is also the responsibility of competition organisers of a code competition to make sure that models can be properly tested in Kaggle kernels that have the same settings as  the actual test inference kernel. This is not the case since we can currently NOT run tflite interpreter inference checks in Kaggle kernels!!</strong></p>",
          "rawMarkdown": "Hi @lonnieqin ,\n\nI tried to make a dummy solution, using all features, and based on the submission model format from the previous competition. I have tried both:\n\ninputs = tf.keras.Input(shape=(543,3), name=\"inputs\")\n\nand \n\ninputs = tf.keras.Input(shape=(None, 543,3), name=\"inputs\")\n\nAs input shapes (using all features for now) and I think I've tried all possible output shape tensor combinations, i.e.:\n- with and without hacking batch dimension as time dimension (as in the previous competition)\n- I also tried per-frame models (although that would be horribly restrictive)\n\nSince all these attempts (at a rate of 5 totally uninformative  \"random guesses\" per day) failed after about one minute, this suggests they don't even start getting scored. I am assuming there is still something wrong with my interpretation of the input tensor shape and/or required output tensor shape or properties (e.g. data format)?\n\nSo, since it seems that your attempts at least get partially scored: could you clarify what the input and output tensor shapes are in your model?\n\n\n\n\n**BTW: I do feel that it is the responsibility of competition organisers to clarify the desired interface of submitted models. It is also the responsibility of competition organisers of a code competition to make sure that models can be properly tested in Kaggle kernels that have the same settings as  the actual test inference kernel. This is not the case since we can currently NOT run tflite interpreter inference checks in Kaggle kernels!!**\n\n\n\n",
          "votes": 1,
          "replies": [
            {
              "id": 2265539,
              "postDate": "2023-05-19T10:19:10.363Z",
              "content": "<p>My Model has input shape (None, 9) and output shape (num_samples * 18, 59). num_samples is calculated by reading <code>frame</code> column. I choose num_samples * 18 because average phrase length of training file is about 18.</p>\n<pre><code>def get_inference_model():\n    inputs = tf.keras.Input((9), dtype=tf.float32, name=\"inputs\")\n    vector = tf.where(tf.math.is_nan(inputs), tf.zeros_like(inputs), inputs)\n    frame_value = vector[:, 8]\n    num_samples = tf.math.count_nonzero(frame_value == 1, dtype=tf.int32)\n    samples = tf.random.normal(shape=(num_samples * 18, 42))\n    outputs = tf.keras.layers.Dense(59, activation=\"softmax\", name=\"outputs\")(samples)\n    inference_model = tf.keras.Model(inputs=inputs, outputs=outputs) \n    inference_model.compile(loss=tf.keras.losses.SparseCategoricalCrossentropy(), metrics=[\"accuracy\"])\n    return inference_model\n</code></pre>\n<p>During evaluation, each parquet file is read using code: <code>pd.read_parquet(pq_path, columns=selected_columns)</code> when I save json file named <code>inference_args.json</code> </p>\n<pre><code>{'selected_columns': ['x_left_hand_0', 'x_left_hand_1', 'y_left_hand_0', 'y_left_hand_1', 'x_right_hand_0', 'x_right_hand_1', 'y_right_hand_0', 'y_right_hand_1', 'frame']}\n</code></pre>\n<p><br>\nThis code can read parquet file with shape (n, 9). Training parquet files all have 1000 samples of phrases, when making inference with this model, it can get output tensor with shape (18000, 59).</p>",
              "rawMarkdown": "My Model has input shape (None, 9) and output shape (num_samples * 18, 59). num_samples is calculated by reading `frame` column. I choose num_samples * 18 because average phrase length of training file is about 18.\n\n```\ndef get_inference_model():\n    inputs = tf.keras.Input((9), dtype=tf.float32, name=\"inputs\")\n    vector = tf.where(tf.math.is_nan(inputs), tf.zeros_like(inputs), inputs)\n    frame_value = vector[:, 8]\n    num_samples = tf.math.count_nonzero(frame_value == 1, dtype=tf.int32)\n    samples = tf.random.normal(shape=(num_samples * 18, 42))\n    outputs = tf.keras.layers.Dense(59, activation=\"softmax\", name=\"outputs\")(samples)\n    inference_model = tf.keras.Model(inputs=inputs, outputs=outputs) \n    inference_model.compile(loss=tf.keras.losses.SparseCategoricalCrossentropy(), metrics=[\"accuracy\"])\n    return inference_model\n```\nDuring evaluation, each parquet file is read using code: `pd.read_parquet(pq_path, columns=selected_columns)` when I save json file named `inference_args.json` \n```\n{'selected_columns': ['x_left_hand_0', 'x_left_hand_1', 'y_left_hand_0', 'y_left_hand_1', 'x_right_hand_0', 'x_right_hand_1', 'y_right_hand_0', 'y_right_hand_1', 'frame']}\n``` \nThis code can read parquet file with shape (n, 9). Training parquet files all have 1000 samples of phrases, when making inference with this model, it can get output tensor with shape (18000, 59).",
              "votes": 2
            },
            {
              "id": 2265540,
              "postDate": "2023-05-19T10:20:54.987Z",
              "content": "<p>I think we can use less than 1% of the data for training and validation, which will not cost too much, but still verify that the model meets the competition requirements.</p>",
              "rawMarkdown": " I think we can use less than 1% of the data for training and validation, which will not cost too much, but still verify that the model meets the competition requirements."
            },
            {
              "id": 2265627,
              "postDate": "2023-05-19T11:45:09.903Z",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/lonnieqin\" target=\"_blank\">@lonnieqin</a> </p>\n<p>I'd already used up my \"random guesses\" for today, but since your dummy model at least scores partially, you've probably got the dimensions right. What I'd missed, obviously, is that features are loaded as a flat array for each frame (instead of the #kpx3 format of the previous competition).</p>\n<p>What may be happening in yoru model that there are additional constraints on output length, e.g. that it cannot be longer than the original number of frames in the video OR that your \"num_samples\" possibly returns zero at some point ?</p>",
              "rawMarkdown": "Thanks @lonnieqin \n\nI'd already used up my \"random guesses\" for today, but since your dummy model at least scores partially, you've probably got the dimensions right. What I'd missed, obviously, is that features are loaded as a flat array for each frame (instead of the #kpx3 format of the previous competition).\n\nWhat may be happening in yoru model that there are additional constraints on output length, e.g. that it cannot be longer than the original number of frames in the video OR that your \"num_samples\" possibly returns zero at some point ?",
              "votes": 1
            },
            {
              "id": 2265634,
              "postDate": "2023-05-19T11:50:38.100Z",
              "content": "<p>Btw, my next attempt would be this one:</p>\n<pre><code>FEATURES = 543*3+1\nNUM_CHARACTERS = 59\n\ndef get_dummy_model():\n    inputs = tf.keras.Input(shape=(FEATURES,), dtype=tf.float32, name=\"inputs\")\n    print(inputs.shape)\n    x = tf.where(tf.math.is_nan(inputs), tf.zeros_like(inputs), inputs)\n    print(x.shape)\n    # just some dummy stuff here to create random output logits\n    x = tf.math.reduce_mean(x, axis = 1, keepdims = True)\n    print(x.shape)\n    x = tf.keras.layers.Dense(NUM_CHARACTERS)(x)\n    print(x.shape)\n    out = tf.keras.layers.Activation(\"linear\", name=\"outputs\")(x)\n    print(out.shape)\n    inference_model = tf.keras.Model(inputs=inputs, outputs=out)\n    inference_model.compile(loss=\"sparse_categorical_crossentropy\",metrics=\"accuracy\")\n    return inference_model\n\ndummy_model_test = get_dummy_model()\ndummy_model_test.summary()\n</code></pre>",
              "rawMarkdown": "Btw, my next attempt would be this one:\n\n```\nFEATURES = 543*3+1\nNUM_CHARACTERS = 59\n\ndef get_dummy_model():\n    inputs = tf.keras.Input(shape=(FEATURES,), dtype=tf.float32, name=\"inputs\")\n    print(inputs.shape)\n    x = tf.where(tf.math.is_nan(inputs), tf.zeros_like(inputs), inputs)\n    print(x.shape)\n    # just some dummy stuff here to create random output logits\n    x = tf.math.reduce_mean(x, axis = 1, keepdims = True)\n    print(x.shape)\n    x = tf.keras.layers.Dense(NUM_CHARACTERS)(x)\n    print(x.shape)\n    out = tf.keras.layers.Activation(\"linear\", name=\"outputs\")(x)\n    print(out.shape)\n    inference_model = tf.keras.Model(inputs=inputs, outputs=out)\n    inference_model.compile(loss=\"sparse_categorical_crossentropy\",metrics=\"accuracy\")\n    return inference_model\n\ndummy_model_test = get_dummy_model()\ndummy_model_test.summary()\n```",
              "votes": 1
            },
            {
              "id": 2265732,
              "postDate": "2023-05-19T13:32:41.203Z",
              "content": "<p>I tried to restrict num_samples with min and max values, and tried things like num_samples = tf.math.floordiv(tf.shape(inputs)[0], 6), only to find it failed immediately, which may indicate the length of output tensor need to be divided by number of phrases. I can’t even add 1 to num_samples.</p>",
              "rawMarkdown": "I tried to restrict num_samples with min and max values, and tried things like num_samples = tf.math.floordiv(tf.shape(inputs)[0], 6), only to find it failed immediately, which may indicate the length of output tensor need to be divided by number of phrases. I can’t even add 1 to num_samples."
            },
            {
              "id": 2265868,
              "postDate": "2023-05-19T15:09:30.060Z",
              "content": "<p>Does this mean that the submitted model is supposed to predict outputs for multiple concatenated samples/phrases? I assumed that it would only receive one phrase at the time and that output length would simply have to be &lt;= input length …</p>",
              "rawMarkdown": "Does this mean that the submitted model is supposed to predict outputs for multiple concatenated samples/phrases? I assumed that it would only receive one phrase at the time and that output length would simply have to be <= input length ..."
            },
            {
              "id": 2266012,
              "postDate": "2023-05-19T16:30:12.717Z",
              "content": "<p>I think in mobile devices there are just a few phrases to be predicted at a time, it’s able to handle video streams. However in kaggle competition our model have to make batch inferences since dataset size is so large. The output size can be set to a reasonable value as long as the model’s performance meets requirement.</p>",
              "rawMarkdown": "I think in mobile devices there are just a few phrases to be predicted at a time, it’s able to handle video streams. However in kaggle competition our model have to make batch inferences since dataset size is so large. The output size can be set to a reasonable value as long as the model’s performance meets requirement."
            }
          ]
        }
      ]
    },
    {
      "id": 2260524,
      "postDate": "2023-05-15T17:43:29.287Z",
      "content": "<p>I'm going to review our logs in detail later today to verify that there's no issue on our end, but we expected to see delayed submissions for this competition given the need to train a model first. It took several days for the first successful submission to be made to the previous sign language competition as well. I have already confirmed that we've only gotten a handful of submission attempts so far.</p>",
      "rawMarkdown": "I'm going to review our logs in detail later today to verify that there's no issue on our end, but we expected to see delayed submissions for this competition given the need to train a model first. It took several days for the first successful submission to be made to the previous sign language competition as well. I have already confirmed that we've only gotten a handful of submission attempts so far.",
      "votes": 1,
      "replies": [
        {
          "id": 2260744,
          "postDate": "2023-05-15T20:42:52.830Z",
          "content": "<p>Thanks for taking a look at the logs, I am struggling for two days now to get a successful submission.<br>\nConversion to TFLite model and inference in TFLite model are working properly and are following the same format as in the Isolated Sign Language competition.<br>\nWould be helpful to know how the inference differs from the previous Isolated Sign Language competition.</p>",
          "rawMarkdown": "Thanks for taking a look at the logs, I am struggling for two days now to get a successful submission.\nConversion to TFLite model and inference in TFLite model are working properly and are following the same format as in the Isolated Sign Language competition.\nWould be helpful to know how the inference differs from the previous Isolated Sign Language competition.",
          "votes": 1,
          "replies": [
            {
              "id": 2264123,
              "postDate": "2023-05-18T07:16:24.577Z",
              "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> Could you share any insights from the logs as to why everyone's submissions are failing? Even my simple notebooks with a model which should return outputs in the correct format are failing.</p>",
              "rawMarkdown": "@sohier Could you share any insights from the logs as to why everyone's submissions are failing? Even my simple notebooks with a model which should return outputs in the correct format are failing.",
              "votes": 1
            },
            {
              "id": 2265005,
              "postDate": "2023-05-18T23:18:43.877Z",
              "content": "<p>Any news on this?<br>\n I am assuming the dataframe is converted to a numpy array of [frame, 1630] (including the column \"frame\")</p>\n<p>also in this line of inference</p>\n<p>prediction_str = \"\".join([rev_character_map.get(s, \"\") for s in np.argmax(output[REQUIRED_OUTPUT], axis=1)])</p>\n<p>is REQUIRED_OUTPUT = \"outputs\"</p>",
              "rawMarkdown": "Any news on this?\n I am assuming the dataframe is converted to a numpy array of [frame, 1630] (including the column \"frame\")\n\nalso in this line of inference\n\nprediction_str = \"\".join([rev_character_map.get(s, \"\") for s in np.argmax(output[REQUIRED_OUTPUT], axis=1)])\n\nis REQUIRED_OUTPUT = \"outputs\"",
              "votes": 2
            }
          ]
        },
        {
          "id": 2261245,
          "postDate": "2023-05-16T07:42:56.550Z",
          "content": "<p>Could there be any issues if subset of landmarks is tried?  as per Evaluation</p>\n<blockquote>\n  <p>If you want to load only a subset of the landmarks, include a file named inference_args.json in your submission.zip with the field selected_columns containing a list of the landmark columns you want to use. If that is not included we will load all columns.</p>\n</blockquote>",
          "rawMarkdown": "Could there be any issues if subset of landmarks is tried?  as per Evaluation\n\n>If you want to load only a subset of the landmarks, include a file named inference_args.json in your submission.zip with the field selected_columns containing a list of the landmark columns you want to use. If that is not included we will load all columns.",
          "replies": [
            {
              "id": 2264648,
              "postDate": "2023-05-18T15:46:37.117Z",
              "content": "<p>I also tried with all the landmarks (without the JSON file) and the submission still failed. But without any error messages it is hard to know what the reason is.</p>",
              "rawMarkdown": "I also tried with all the landmarks (without the JSON file) and the submission still failed. But without any error messages it is hard to know what the reason is.",
              "votes": 1
            }
          ]
        },
        {
          "id": 2261538,
          "postDate": "2023-05-16T11:56:01.327Z",
          "content": "<p>We trying to prepare our first submit. But it's hard to choose right model size without any knowledge about inference time and what is \"35 hours\"<br>\n<a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> could you please tell ~inference time for one phrase (<a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/409722\" target=\"_blank\">asked here</a> ) as you <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/393554\" target=\"_blank\">did</a> in previous competition? Where we knew the number of samples (phrases here)</p>",
          "rawMarkdown": "We trying to prepare our first submit. But it's hard to choose right model size without any knowledge about inference time and what is \"35 hours\"\n@sohier could you please tell ~inference time for one phrase ([asked here](https://www.kaggle.com/competitions/asl-fingerspelling/discussion/409722) ) as you [did](https://www.kaggle.com/competitions/asl-signs/discussion/393554) in previous competition? Where we knew the number of samples (phrases here)",
          "votes": 3
        }
      ]
    },
    {
      "id": 2260225,
      "postDate": "2023-05-15T14:41:39.310Z",
      "content": "<p>Everyone trusts CV :) </p>",
      "rawMarkdown": "Everyone trusts CV :) "
    },
    {
      "id": 2267913,
      "postDate": "2023-05-21T10:37:20.203Z",
      "content": "<p>Was comparing the parquet data in this competition compared to GISLR and think that these parquet files are multiple sequence ids in the one parquet (but should all be for the same participant id).  Whereas in GISLR, each parquet was a separate folder for participant id and a separate parquet for each sequence id, really the parquet name was the sequence id.  So at the moment what comes back from the function <br>\n<code>load_relevant_data_subset(pq_path)</code><br>\n can be like 1000+ entries and the parquet files are huge 1.5+ GB compared to &lt; 2MB in GISLR.<br>\nThink it is impossible to do anything with these parquet files given the competition constraints, unless what is passed to the models is a further subset, that is unclear from the Evaluation page.  Could be wrong but think the parquet files need to be re-done by sequence id.   <br>\n<a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> - what do you think?</p>",
      "rawMarkdown": "Was comparing the parquet data in this competition compared to GISLR and think that these parquet files are multiple sequence ids in the one parquet (but should all be for the same participant id).  Whereas in GISLR, each parquet was a separate folder for participant id and a separate parquet for each sequence id, really the parquet name was the sequence id.  So at the moment what comes back from the function \n`load_relevant_data_subset(pq_path)`\n can be like 1000+ entries and the parquet files are huge 1.5+ GB compared to < 2MB in GISLR.\nThink it is impossible to do anything with these parquet files given the competition constraints, unless what is passed to the models is a further subset, that is unclear from the Evaluation page.  Could be wrong but think the parquet files need to be re-done by sequence id.   \n@sohier - what do you think?",
      "replies": [
        {
          "id": 2267964,
          "postDate": "2023-05-21T11:05:47.847Z",
          "content": "<p>I think here we have some <code>test.csv</code> with same structure as <code>train.csv</code></p>\n<p>And we don't see how variable <code>frames</code> preprocessed. I guess it should be like this for one sample. (Or something more efficient, but with the same output):</p>\n<pre><code>test_df = pd.read_csv()\n ()  json_file:\n    selected_columns_ = json.load(json_file)\nselected_columns = selected_columns_[]\n\n ():\n     pd.read_parquet(pq_path, columns=selected_columns)\n\ni = \nsample = test_df.loc[i] \nloaded = load_relevant_data_subset( + sample[])\nframes = loaded[loaded.index==sample[]].values\noutput = prediction_fn(inputs=frames)\nprediction_str = .join([rev_character_map.get(s, )  s  np.argmax(output[], axis=)])\n</code></pre>\n<p>But yeah, I'm also have an error after 1 minute😔</p>",
          "rawMarkdown": "I think here we have some `test.csv` with same structure as `train.csv`\n\nAnd we don't see how variable `frames` preprocessed. I guess it should be like this for one sample. (Or something more efficient, but with the same output):\n\n```python\ntest_df = pd.read_csv('test_csv')\nwith open('inference_args.json') as json_file:\n    selected_columns_ = json.load(json_file)\nselected_columns = selected_columns_['selected_columns']\n\ndef load_relevant_data_subset(pq_path):\n    return pd.read_parquet(pq_path, columns=selected_columns)\n\ni = 0\nsample = test_df.loc[i] \nloaded = load_relevant_data_subset('asl-fingerspelling/' + sample['path'])\nframes = loaded[loaded.index==sample['sequence_id']].values\noutput = prediction_fn(inputs=frames)\nprediction_str = \"\".join([rev_character_map.get(s, \"\") for s in np.argmax(output[\"outputs\"], axis=1)])\n```\n\nBut yeah, I'm also have an error after 1 minute😔"
        }
      ]
    },
    {
      "id": 2266931,
      "postDate": "2023-05-20T13:30:15.340Z",
      "content": "<p>Hi all,</p>\n<p>Besides <a href=\"https://www.kaggle.com/lonnieqin\" target=\"_blank\">@lonnieqin</a> , has anyone else managed to make a submission that did not crash after ~1min (so at least managed to get partially through scoring)? <br>\nI think I've tried everything (including running the code Lonnie shared) and my submissions keep failing after 1min.</p>\n<p>For those of you who did get past this one minute-barrier: did you generate TfLite from the default Kaggle kernel or did you use a different setup (Python version? Tensorflow version? other??). And does your submission behave the same if you do run t from the Kaggle kernel?</p>\n<p>I'm just trying to find out whether part of the problems may be due to a version issue …?</p>",
      "rawMarkdown": "Hi all,\n\nBesides @lonnieqin , has anyone else managed to make a submission that did not crash after ~1min (so at least managed to get partially through scoring)? \nI think I've tried everything (including running the code Lonnie shared) and my submissions keep failing after 1min.\n\nFor those of you who did get past this one minute-barrier: did you generate TfLite from the default Kaggle kernel or did you use a different setup (Python version? Tensorflow version? other??). And does your submission behave the same if you do run t from the Kaggle kernel?\n\nI'm just trying to find out whether part of the problems may be due to a version issue ...?",
      "replies": [
        {
          "id": 2267032,
          "postDate": "2023-05-20T14:35:13.597Z",
          "content": "<p>Taking into account Sohiers's answer, I think any submission that didn't crash after 1minutes were put in the queue, and spend some times in the queue. </p>",
          "rawMarkdown": "Taking into account Sohiers's answer, I think any submission that didn't crash after 1minutes were put in the queue, and spend some times in the queue. ",
          "replies": [
            {
              "id": 2267115,
              "postDate": "2023-05-20T15:50:57.673Z",
              "content": "<p>Indeed, I resubmitted 1 notebook version that could run over 1 hour before, only to find it failed after a few minutes. </p>",
              "rawMarkdown": "Indeed, I resubmitted 1 notebook version that could run over 1 hour before, only to find it failed after a few minutes. "
            },
            {
              "id": 2267137,
              "postDate": "2023-05-20T16:10:11.030Z",
              "content": "<p>Thanks, so that means that probably no-one got beyond that initial crash then?</p>\n<p>I guess there's nothing to be done then but wait for the organisers to finally clarify the requirements and/or fix any remaining issues … I'm wondering whether they even ran a test submission themselves to verify that all is well. </p>",
              "rawMarkdown": "Thanks, so that means that probably no-one got beyond that initial crash then?\n\nI guess there's nothing to be done then but wait for the organisers to finally clarify the requirements and/or fix any remaining issues ... I'm wondering whether they even ran a test submission themselves to verify that all is well. ",
              "votes": 1
            },
            {
              "id": 2267803,
              "postDate": "2023-05-21T08:43:13.390Z",
              "content": "<p>I think required model<br>\n input and output shape is very clear now, it’s more likely a pipeline or dataset issue.</p>",
              "rawMarkdown": "I think required model\n input and output shape is very clear now, it’s more likely a pipeline or dataset issue."
            },
            {
              "id": 2267848,
              "postDate": "2023-05-21T09:29:03.203Z",
              "content": "<p>Yep. Nothing to do but wait … </p>\n<p>Update: </p>\n<p>I tried creating and submitting my dummy model from a 'pinned' kernel from the previous competition and I got the same weird behaviour that you reported: first time I got submission scoring error after ~1h, second time after ~25 min. Probably a coincidence but the submissions from the new (current) Kaggle kernel consistently failed after 1 min.</p>\n<p>In the old kernel I had no problem doing inference with the TfLite model I submitted (using 'best guesses' for the stuff that is missing or unclear in the evaluation script code on this site).</p>",
              "rawMarkdown": "Yep. Nothing to do but wait ... \n\nUpdate: \n\nI tried creating and submitting my dummy model from a 'pinned' kernel from the previous competition and I got the same weird behaviour that you reported: first time I got submission scoring error after ~1h, second time after ~25 min. Probably a coincidence but the submissions from the new (current) Kaggle kernel consistently failed after 1 min.\n\nIn the old kernel I had no problem doing inference with the TfLite model I submitted (using 'best guesses' for the stuff that is missing or unclear in the evaluation script code on this site).",
              "votes": 1
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2260397,
      "author_name": "Dewei Chen",
      "author_url": "",
      "post_date": "2023-05-15T16:30:42.423000",
      "content": "<p>Because the training set is too large, we still need some time for model training.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2260446,
          "author_name": "Mykola",
          "author_url": "",
          "post_date": "2023-05-15T16:57:35.657000",
          "content": "<p>It makes sense, but what about the fact, that in \"Google Research - Identify Contrails to Reduce Global Warming\" there are 60+ teams, while the dataset is more than twice bigger?</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2260528,
              "author_name": "Sohier Dane",
              "author_url": "",
              "post_date": "2023-05-15T17:45:39.560000",
              "content": "<p>It's trivial to make a dummy submission to the contrails competition (or almost any other competition for that matter). The bar for an initial submission is higher in this case since someone will need to prepare a working model first.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2265537,
              "author_name": "Wondering Alice",
              "author_url": "",
              "post_date": "2023-05-19T10:15:43.510000",
              "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> it should be quite trivial to make an initial random submission here as well, but as you can read in some of the other discussion threads, all (very well informed) attempts to do so are failing thus far. Could you please clarify the required input and output formats:</p>\n<ul>\n<li>input tensor format to TfLite model (i.e. is time in batch dimension as in previous competition or not)<br>\nmost common assumption is currently [num_frames,1630], with num_frames 'abusing' the batch dimension</li>\n<li>desired output tensor format to TfLite model: <br>\nper frame of per sequence? <br>\ntime as separate dimension or in batch dimension? <br>\nshould output length (num_frames)  be = input length (num_frames) ? or &lt;= than input length? or    any other constraint?<br>\nanything else that can be responsible for crashes?</li>\n<li>how to install and run tflite interpreter in a Kaggle kernel to at least be able to verify/debug our own tflite models </li>\n<li>whether the problems could be due to a version mismatch between the tflite models created by the Tensorflow version installed in the Kaggle kernels and the required version for inference (which is also quite old)</li>\n</ul>",
              "votes": 3,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2266183,
      "author_name": "Sohier Dane",
      "author_url": "",
      "post_date": "2023-05-19T19:31:03.803000",
      "content": "<p>Thank you for raising this concern. We reviewed the metric and identified that a necessary dataset had failed to properly attach. Submissions should now be working.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2266361,
          "author_name": "Lonnie",
          "author_url": "",
          "post_date": "2023-05-20T00:58:22.053000",
          "content": "<p>However it still doesn’t work on my side. I am wondering if anyone can make successful submission now.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2266414,
          "author_name": "Jon",
          "author_url": "",
          "post_date": "2023-05-20T03:50:43.543000",
          "content": "<p>I still can't submit a valid submission. Is it possible to know what are the input shapes by defaults? n_framesx1630?<br>\nthe output shape should be n_charactersx59?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2266588,
          "author_name": "Mathieu De Coster",
          "author_url": "",
          "post_date": "2023-05-20T07:44:25.967000",
          "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> Thank you for looking into this. I still can't make any valid submissions. Could you please give the exact formats for the input data passed to the model and the output format expected? Or even better, provide a notebook with a working submission? That would really help us get started on this competition.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 2266698,
          "author_name": "Mark Wijkhuizen",
          "author_url": "",
          "post_date": "2023-05-20T09:23:47.740000",
          "content": "<p>Sadly, my submissions keep failing as well. I agree with <a href=\"https://www.kaggle.com/mdecoster\" target=\"_blank\">@mdecoster</a> that providing an example submission notebook would be very helpful.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 2269074,
          "author_name": "alexis1024",
          "author_url": "",
          "post_date": "2023-05-22T07:48:48.437000",
          "content": "<p>can you please fix the doc of the evaluation page and let us know the expected signature (especially please define REQUIRED_OUTPUT variable)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2267501,
      "author_name": "Kolya Forrat",
      "author_url": "",
      "post_date": "2023-05-21T03:03:49.247000",
      "content": "<p>Everyone got what should be in <code>REQUIRED_SIGNATURE</code> and <code>REQUIRED_OUTPUT</code>? <br>\nAs described in Evaluation part</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2267528,
          "author_name": "Lonnie",
          "author_url": "",
          "post_date": "2023-05-21T03:49:26.110000",
          "content": "<p>My understanding is that <code>REQUIRED_SIGNATURE</code> = <code>serving_default</code> and <code>REQUIRED_OUTPUT</code> = <code>outputs</code>.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2267993,
              "author_name": "Wondering Alice",
              "author_url": "",
              "post_date": "2023-05-21T11:28:56.880000",
              "content": "<p>Yes, at least, that was the interpretation in the previous competition.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2267757,
          "author_name": "George Megre",
          "author_url": "",
          "post_date": "2023-05-21T07:59:46.820000",
          "content": "<p>i suppose that REQUIRED_SIGNATURE = serving_default, but REQUIRED_OUTPUT is number of elements in the true string, because our evaluation metric for this contest is the normalized total levenshtein distance and  number of characters in the labels be N and the total levenshtein distance be D. The metric equals (N - D) / N. I believe that organizers want to have range of that metric like [0, 1]. <br>\nFor instance if we have 2 strings real and estiamted: y_true = \"i can see the rings on saturn\", and y_pred = \"i can see the rings on saturn s\" N = len(y_true) , the Levinstein distance is calculated on (y_true, y_pred[:N]). But it is on only my guess i could be wrong. By the way i'm also interest what is REQUIRED_OUTPUT.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2268000,
              "author_name": "Wondering Alice",
              "author_url": "",
              "post_date": "2023-05-21T11:35:49.087000",
              "content": "<p>In our interpretation,  REQUIRED_OUTPUT is \"outputs\", i.e. the name you need to give to the last layer in your model (i.e., the one that produces the outputs).</p>\n<p>I have posted a demo notebook to show how TfLite inference can be done <a href=\"https://www.kaggle.com/code/wonderingalice/dummy-submission-with-previous-kernel?kernelSessionId=130400757\" target=\"_blank\">here</a></p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2264709,
      "author_name": "Lonnie",
      "author_url": "",
      "post_date": "2023-05-18T16:55:14.017000",
      "content": "<p>That’s very tricky. I created a random model that takes left hand + right hand landmarks data as input for submission. It can make inferences with parquet files. However submission failed after running about one hour.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2264765,
          "author_name": "Wondering Alice",
          "author_url": "",
          "post_date": "2023-05-18T17:23:39.807000",
          "content": "<p>Hi Lonnie,</p>\n<p>That's already better than my attempts: this suggests that scoring works for some but not all samples?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2265084,
              "author_name": "Lonnie",
              "author_url": "",
              "post_date": "2023-05-19T01:11:11.760000",
              "content": "<p>I am not sure. Everything could be possible. It may be some kinds of pipeline issues or dataset issues. But in my notebook it takes 2 minutes to make inferences and evaluation with <code>tf.edit_distance</code> method on train files and supplemental files with low memory footprint. </p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2265517,
          "author_name": "Wondering Alice",
          "author_url": "",
          "post_date": "2023-05-19T09:54:34.280000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/lonnieqin\" target=\"_blank\">@lonnieqin</a> ,</p>\n<p>I tried to make a dummy solution, using all features, and based on the submission model format from the previous competition. I have tried both:</p>\n<p>inputs = tf.keras.Input(shape=(543,3), name=\"inputs\")</p>\n<p>and </p>\n<p>inputs = tf.keras.Input(shape=(None, 543,3), name=\"inputs\")</p>\n<p>As input shapes (using all features for now) and I think I've tried all possible output shape tensor combinations, i.e.:</p>\n<ul>\n<li>with and without hacking batch dimension as time dimension (as in the previous competition)</li>\n<li>I also tried per-frame models (although that would be horribly restrictive)</li>\n</ul>\n<p>Since all these attempts (at a rate of 5 totally uninformative  \"random guesses\" per day) failed after about one minute, this suggests they don't even start getting scored. I am assuming there is still something wrong with my interpretation of the input tensor shape and/or required output tensor shape or properties (e.g. data format)?</p>\n<p>So, since it seems that your attempts at least get partially scored: could you clarify what the input and output tensor shapes are in your model?</p>\n<p><strong>BTW: I do feel that it is the responsibility of competition organisers to clarify the desired interface of submitted models. It is also the responsibility of competition organisers of a code competition to make sure that models can be properly tested in Kaggle kernels that have the same settings as  the actual test inference kernel. This is not the case since we can currently NOT run tflite interpreter inference checks in Kaggle kernels!!</strong></p>",
          "votes": 1,
          "replies": [
            {
              "id": 2265539,
              "author_name": "Lonnie",
              "author_url": "",
              "post_date": "2023-05-19T10:19:10.363000",
              "content": "<p>My Model has input shape (None, 9) and output shape (num_samples * 18, 59). num_samples is calculated by reading <code>frame</code> column. I choose num_samples * 18 because average phrase length of training file is about 18.</p>\n<pre><code>def get_inference_model():\n    inputs = tf.keras.Input((9), dtype=tf.float32, name=\"inputs\")\n    vector = tf.where(tf.math.is_nan(inputs), tf.zeros_like(inputs), inputs)\n    frame_value = vector[:, 8]\n    num_samples = tf.math.count_nonzero(frame_value == 1, dtype=tf.int32)\n    samples = tf.random.normal(shape=(num_samples * 18, 42))\n    outputs = tf.keras.layers.Dense(59, activation=\"softmax\", name=\"outputs\")(samples)\n    inference_model = tf.keras.Model(inputs=inputs, outputs=outputs) \n    inference_model.compile(loss=tf.keras.losses.SparseCategoricalCrossentropy(), metrics=[\"accuracy\"])\n    return inference_model\n</code></pre>\n<p>During evaluation, each parquet file is read using code: <code>pd.read_parquet(pq_path, columns=selected_columns)</code> when I save json file named <code>inference_args.json</code> </p>\n<pre><code>{'selected_columns': ['x_left_hand_0', 'x_left_hand_1', 'y_left_hand_0', 'y_left_hand_1', 'x_right_hand_0', 'x_right_hand_1', 'y_right_hand_0', 'y_right_hand_1', 'frame']}\n</code></pre>\n<p><br>\nThis code can read parquet file with shape (n, 9). Training parquet files all have 1000 samples of phrases, when making inference with this model, it can get output tensor with shape (18000, 59).</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2265540,
              "author_name": "Dewei Chen",
              "author_url": "",
              "post_date": "2023-05-19T10:20:54.987000",
              "content": "<p>I think we can use less than 1% of the data for training and validation, which will not cost too much, but still verify that the model meets the competition requirements.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2265627,
              "author_name": "Wondering Alice",
              "author_url": "",
              "post_date": "2023-05-19T11:45:09.903000",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/lonnieqin\" target=\"_blank\">@lonnieqin</a> </p>\n<p>I'd already used up my \"random guesses\" for today, but since your dummy model at least scores partially, you've probably got the dimensions right. What I'd missed, obviously, is that features are loaded as a flat array for each frame (instead of the #kpx3 format of the previous competition).</p>\n<p>What may be happening in yoru model that there are additional constraints on output length, e.g. that it cannot be longer than the original number of frames in the video OR that your \"num_samples\" possibly returns zero at some point ?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2265634,
              "author_name": "Wondering Alice",
              "author_url": "",
              "post_date": "2023-05-19T11:50:38.100000",
              "content": "<p>Btw, my next attempt would be this one:</p>\n<pre><code>FEATURES = 543*3+1\nNUM_CHARACTERS = 59\n\ndef get_dummy_model():\n    inputs = tf.keras.Input(shape=(FEATURES,), dtype=tf.float32, name=\"inputs\")\n    print(inputs.shape)\n    x = tf.where(tf.math.is_nan(inputs), tf.zeros_like(inputs), inputs)\n    print(x.shape)\n    # just some dummy stuff here to create random output logits\n    x = tf.math.reduce_mean(x, axis = 1, keepdims = True)\n    print(x.shape)\n    x = tf.keras.layers.Dense(NUM_CHARACTERS)(x)\n    print(x.shape)\n    out = tf.keras.layers.Activation(\"linear\", name=\"outputs\")(x)\n    print(out.shape)\n    inference_model = tf.keras.Model(inputs=inputs, outputs=out)\n    inference_model.compile(loss=\"sparse_categorical_crossentropy\",metrics=\"accuracy\")\n    return inference_model\n\ndummy_model_test = get_dummy_model()\ndummy_model_test.summary()\n</code></pre>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2265732,
              "author_name": "Lonnie",
              "author_url": "",
              "post_date": "2023-05-19T13:32:41.203000",
              "content": "<p>I tried to restrict num_samples with min and max values, and tried things like num_samples = tf.math.floordiv(tf.shape(inputs)[0], 6), only to find it failed immediately, which may indicate the length of output tensor need to be divided by number of phrases. I can’t even add 1 to num_samples.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2265868,
              "author_name": "Wondering Alice",
              "author_url": "",
              "post_date": "2023-05-19T15:09:30.060000",
              "content": "<p>Does this mean that the submitted model is supposed to predict outputs for multiple concatenated samples/phrases? I assumed that it would only receive one phrase at the time and that output length would simply have to be &lt;= input length …</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2266012,
              "author_name": "Lonnie",
              "author_url": "",
              "post_date": "2023-05-19T16:30:12.717000",
              "content": "<p>I think in mobile devices there are just a few phrases to be predicted at a time, it’s able to handle video streams. However in kaggle competition our model have to make batch inferences since dataset size is so large. The output size can be set to a reasonable value as long as the model’s performance meets requirement.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2260524,
      "author_name": "Sohier Dane",
      "author_url": "",
      "post_date": "2023-05-15T17:43:29.287000",
      "content": "<p>I'm going to review our logs in detail later today to verify that there's no issue on our end, but we expected to see delayed submissions for this competition given the need to train a model first. It took several days for the first successful submission to be made to the previous sign language competition as well. I have already confirmed that we've only gotten a handful of submission attempts so far.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2260744,
          "author_name": "Mark Wijkhuizen",
          "author_url": "",
          "post_date": "2023-05-15T20:42:52.830000",
          "content": "<p>Thanks for taking a look at the logs, I am struggling for two days now to get a successful submission.<br>\nConversion to TFLite model and inference in TFLite model are working properly and are following the same format as in the Isolated Sign Language competition.<br>\nWould be helpful to know how the inference differs from the previous Isolated Sign Language competition.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2264123,
              "author_name": "Mathieu De Coster",
              "author_url": "",
              "post_date": "2023-05-18T07:16:24.577000",
              "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> Could you share any insights from the logs as to why everyone's submissions are failing? Even my simple notebooks with a model which should return outputs in the correct format are failing.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2265005,
              "author_name": "Jon",
              "author_url": "",
              "post_date": "2023-05-18T23:18:43.877000",
              "content": "<p>Any news on this?<br>\n I am assuming the dataframe is converted to a numpy array of [frame, 1630] (including the column \"frame\")</p>\n<p>also in this line of inference</p>\n<p>prediction_str = \"\".join([rev_character_map.get(s, \"\") for s in np.argmax(output[REQUIRED_OUTPUT], axis=1)])</p>\n<p>is REQUIRED_OUTPUT = \"outputs\"</p>",
              "votes": 2,
              "replies": []
            }
          ]
        },
        {
          "id": 2261245,
          "author_name": "something4kag",
          "author_url": "",
          "post_date": "2023-05-16T07:42:56.550000",
          "content": "<p>Could there be any issues if subset of landmarks is tried?  as per Evaluation</p>\n<blockquote>\n  <p>If you want to load only a subset of the landmarks, include a file named inference_args.json in your submission.zip with the field selected_columns containing a list of the landmark columns you want to use. If that is not included we will load all columns.</p>\n</blockquote>",
          "votes": 0,
          "replies": [
            {
              "id": 2264648,
              "author_name": "Mathieu De Coster",
              "author_url": "",
              "post_date": "2023-05-18T15:46:37.117000",
              "content": "<p>I also tried with all the landmarks (without the JSON file) and the submission still failed. But without any error messages it is hard to know what the reason is.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 2261538,
          "author_name": "Kolya Forrat",
          "author_url": "",
          "post_date": "2023-05-16T11:56:01.327000",
          "content": "<p>We trying to prepare our first submit. But it's hard to choose right model size without any knowledge about inference time and what is \"35 hours\"<br>\n<a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> could you please tell ~inference time for one phrase (<a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/409722\" target=\"_blank\">asked here</a> ) as you <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/393554\" target=\"_blank\">did</a> in previous competition? Where we knew the number of samples (phrases here)</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2260225,
      "author_name": "Vadim Irtlach",
      "author_url": "",
      "post_date": "2023-05-15T14:41:39.310000",
      "content": "<p>Everyone trusts CV :) </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2267913,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "2023-05-21T10:37:20.203000",
      "content": "<p>Was comparing the parquet data in this competition compared to GISLR and think that these parquet files are multiple sequence ids in the one parquet (but should all be for the same participant id).  Whereas in GISLR, each parquet was a separate folder for participant id and a separate parquet for each sequence id, really the parquet name was the sequence id.  So at the moment what comes back from the function <br>\n<code>load_relevant_data_subset(pq_path)</code><br>\n can be like 1000+ entries and the parquet files are huge 1.5+ GB compared to &lt; 2MB in GISLR.<br>\nThink it is impossible to do anything with these parquet files given the competition constraints, unless what is passed to the models is a further subset, that is unclear from the Evaluation page.  Could be wrong but think the parquet files need to be re-done by sequence id.   <br>\n<a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> - what do you think?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2267964,
          "author_name": "Kolya Forrat",
          "author_url": "",
          "post_date": "2023-05-21T11:05:47.847000",
          "content": "<p>I think here we have some <code>test.csv</code> with same structure as <code>train.csv</code></p>\n<p>And we don't see how variable <code>frames</code> preprocessed. I guess it should be like this for one sample. (Or something more efficient, but with the same output):</p>\n<pre><code>test_df = pd.read_csv()\n ()  json_file:\n    selected_columns_ = json.load(json_file)\nselected_columns = selected_columns_[]\n\n ():\n     pd.read_parquet(pq_path, columns=selected_columns)\n\ni = \nsample = test_df.loc[i] \nloaded = load_relevant_data_subset( + sample[])\nframes = loaded[loaded.index==sample[]].values\noutput = prediction_fn(inputs=frames)\nprediction_str = .join([rev_character_map.get(s, )  s  np.argmax(output[], axis=)])\n</code></pre>\n<p>But yeah, I'm also have an error after 1 minute😔</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2266931,
      "author_name": "Wondering Alice",
      "author_url": "",
      "post_date": "2023-05-20T13:30:15.340000",
      "content": "<p>Hi all,</p>\n<p>Besides <a href=\"https://www.kaggle.com/lonnieqin\" target=\"_blank\">@lonnieqin</a> , has anyone else managed to make a submission that did not crash after ~1min (so at least managed to get partially through scoring)? <br>\nI think I've tried everything (including running the code Lonnie shared) and my submissions keep failing after 1min.</p>\n<p>For those of you who did get past this one minute-barrier: did you generate TfLite from the default Kaggle kernel or did you use a different setup (Python version? Tensorflow version? other??). And does your submission behave the same if you do run t from the Kaggle kernel?</p>\n<p>I'm just trying to find out whether part of the problems may be due to a version issue …?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2267032,
          "author_name": "Mykola",
          "author_url": "",
          "post_date": "2023-05-20T14:35:13.597000",
          "content": "<p>Taking into account Sohiers's answer, I think any submission that didn't crash after 1minutes were put in the queue, and spend some times in the queue. </p>",
          "votes": 0,
          "replies": [
            {
              "id": 2267115,
              "author_name": "Lonnie",
              "author_url": "",
              "post_date": "2023-05-20T15:50:57.673000",
              "content": "<p>Indeed, I resubmitted 1 notebook version that could run over 1 hour before, only to find it failed after a few minutes. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2267137,
              "author_name": "Wondering Alice",
              "author_url": "",
              "post_date": "2023-05-20T16:10:11.030000",
              "content": "<p>Thanks, so that means that probably no-one got beyond that initial crash then?</p>\n<p>I guess there's nothing to be done then but wait for the organisers to finally clarify the requirements and/or fix any remaining issues … I'm wondering whether they even ran a test submission themselves to verify that all is well. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2267803,
              "author_name": "Lonnie",
              "author_url": "",
              "post_date": "2023-05-21T08:43:13.390000",
              "content": "<p>I think required model<br>\n input and output shape is very clear now, it’s more likely a pipeline or dataset issue.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2267848,
              "author_name": "Wondering Alice",
              "author_url": "",
              "post_date": "2023-05-21T09:29:03.203000",
              "content": "<p>Yep. Nothing to do but wait … </p>\n<p>Update: </p>\n<p>I tried creating and submitting my dummy model from a 'pinned' kernel from the previous competition and I got the same weird behaviour that you reported: first time I got submission scoring error after ~1h, second time after ~25 min. Probably a coincidence but the submissions from the new (current) Kaggle kernel consistently failed after 1 min.</p>\n<p>In the old kernel I had no problem doing inference with the TfLite model I submitted (using 'best guesses' for the stuff that is missing or unclear in the evaluation script code on this site).</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2260138": "Dear all, \n\nWhen I click on the leaderboard, it is empty.\nIs it just me or have no submissions been made yet?\nI find it strange that there aren't even any \"hello world\" baselines yet. Is there something wrong with the data maybe?\n",
    "2260397": "Because the training set is too large, we still need some time for model training.",
    "2266183": "Thank you for raising this concern. We reviewed the metric and identified that a necessary dataset had failed to properly attach. Submissions should now be working.",
    "2267501": "Everyone got what should be in `REQUIRED_SIGNATURE` and `REQUIRED_OUTPUT`? \nAs described in Evaluation part",
    "2264709": "That’s very tricky. I created a random model that takes left hand + right hand landmarks data as input for submission. It can make inferences with parquet files. However submission failed after running about one hour.",
    "2260524": "I'm going to review our logs in detail later today to verify that there's no issue on our end, but we expected to see delayed submissions for this competition given the need to train a model first. It took several days for the first successful submission to be made to the previous sign language competition as well. I have already confirmed that we've only gotten a handful of submission attempts so far.",
    "2260225": "Everyone trusts CV :) ",
    "2267913": "Was comparing the parquet data in this competition compared to GISLR and think that these parquet files are multiple sequence ids in the one parquet (but should all be for the same participant id).  Whereas in GISLR, each parquet was a separate folder for participant id and a separate parquet for each sequence id, really the parquet name was the sequence id.  So at the moment what comes back from the function \n`load_relevant_data_subset(pq_path)`\n can be like 1000+ entries and the parquet files are huge 1.5+ GB compared to < 2MB in GISLR.\nThink it is impossible to do anything with these parquet files given the competition constraints, unless what is passed to the models is a further subset, that is unclear from the Evaluation page.  Could be wrong but think the parquet files need to be re-done by sequence id.   \n@sohier - what do you think?",
    "2266931": "Hi all,\n\nBesides @lonnieqin , has anyone else managed to make a submission that did not crash after ~1min (so at least managed to get partially through scoring)? \nI think I've tried everything (including running the code Lonnie shared) and my submissions keep failing after 1min.\n\nFor those of you who did get past this one minute-barrier: did you generate TfLite from the default Kaggle kernel or did you use a different setup (Python version? Tensorflow version? other??). And does your submission behave the same if you do run t from the Kaggle kernel?\n\nI'm just trying to find out whether part of the problems may be due to a version issue ...?"
  }
}