{
  "id": 412591,
  "title": "Correct input/output formats",
  "url": "/competitions/asl-fingerspelling/discussion/412591",
  "author_name": "Mathieu De Coster",
  "post_date": "2023-05-24T11:08:30.186000",
  "votes": 14,
  "comment_count": 5,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/wonderingalice\" target=\"_blank\">@wonderingalice</a> has made a notebook which correctly runs and can make submissions: <a href=\"https://www.kaggle.com/code/wonderingalice/dummy-submission-with-asl-islr-competition-kernel\" target=\"_blank\">https://www.kaggle.com/code/wonderingalice/dummy-submission-with-asl-islr-competition-kernel</a> 🎉</p>\n<p>This is a write-up of the formats required to make a valid submission. Now that submissions are finally working, we can use this notebook as a starting point!</p>\n<h1>Library versions</h1>\n<p>The original notebook uses an environment with these versions:</p>\n<ul>\n<li>Tensorflow 2.11.0</li>\n<li>Python 3.7.12</li>\n</ul>\n<p>This is a pinned environment from the previous competition on ASL: <a href=\"https://www.kaggle.com/competitions/asl-signs/\" target=\"_blank\">https://www.kaggle.com/competitions/asl-signs/</a></p>\n<h1>Selecting input columns</h1>\n<p>You can select which input columns to use by providing a file called <code>inference_args.json</code>. The JSON file should contain one key <code>selected_columns</code> with as corresponding value a list of the column names in the pandas files. For example:</p>\n<pre><code>{\n  \"selected_columns\": [\"x_right_hand_0\", \"y_right_hand_0\", \"z_right_hand_0\"]\n}\n</code></pre>\n<h1>Model input shape</h1>\n<p>The model will get a single input at a time during inference, so the input shape should be <code>(NUM_COLUMNS)</code>, where <code>NUM_COLUMNS</code> is the number of input columns, or <code>len(inference_args[\"selected_columns\"])</code>.</p>\n<p>The implicit <code>None</code> for the batch dimension is actually used for the time dimension during inference.</p>\n<p>The name of the input layer should be <code>inputs</code>, and the <code>dtype</code> <code>tf.float32</code>.</p>\n<pre><code>inputs = tf.keras.Input(shape=(NUM_COLUMNS), dtype=tf.float32, name=\"inputs\")\n</code></pre>\n<h1>Model output shape</h1>\n<p>There are 59 unique characters in the ground truth strings, so the output shape should be <code>(59)</code>. The output should be a sequence of logits: <code>(num_frames, 59)</code> such that the inference code can take the argmax over axis <code>1</code>.</p>\n<p>For example to get random outputs with a randomly initialized <code>Dense</code> layer:</p>\n<pre><code>x = tf.keras.layers.Dense(59)(x)\nout = tf.keras.layers.Activation(\"linear\", name=\"outputs\")(x)\n</code></pre>\n<p>The name of the output layer should be <code>outputs</code>.</p>\n<h1>Acknowledgements</h1>\n<p>Thanks to everyone who contributed in the discussions on the input and output format and to <a href=\"https://www.kaggle.com/wonderingalice\" target=\"_blank\">@wonderingalice</a> in particular for the submission notebook (go upvote it 😉)</p>",
  "messages": [
    {
      "id": 2272232,
      "postDate": "2023-05-24T11:08:30.187Z",
      "content": "<p><a href=\"https://www.kaggle.com/wonderingalice\" target=\"_blank\">@wonderingalice</a> has made a notebook which correctly runs and can make submissions: <a href=\"https://www.kaggle.com/code/wonderingalice/dummy-submission-with-asl-islr-competition-kernel\" target=\"_blank\">https://www.kaggle.com/code/wonderingalice/dummy-submission-with-asl-islr-competition-kernel</a> 🎉</p>\n<p>This is a write-up of the formats required to make a valid submission. Now that submissions are finally working, we can use this notebook as a starting point!</p>\n<h1>Library versions</h1>\n<p>The original notebook uses an environment with these versions:</p>\n<ul>\n<li>Tensorflow 2.11.0</li>\n<li>Python 3.7.12</li>\n</ul>\n<p>This is a pinned environment from the previous competition on ASL: <a href=\"https://www.kaggle.com/competitions/asl-signs/\" target=\"_blank\">https://www.kaggle.com/competitions/asl-signs/</a></p>\n<h1>Selecting input columns</h1>\n<p>You can select which input columns to use by providing a file called <code>inference_args.json</code>. The JSON file should contain one key <code>selected_columns</code> with as corresponding value a list of the column names in the pandas files. For example:</p>\n<pre><code>{\n  \"selected_columns\": [\"x_right_hand_0\", \"y_right_hand_0\", \"z_right_hand_0\"]\n}\n</code></pre>\n<h1>Model input shape</h1>\n<p>The model will get a single input at a time during inference, so the input shape should be <code>(NUM_COLUMNS)</code>, where <code>NUM_COLUMNS</code> is the number of input columns, or <code>len(inference_args[\"selected_columns\"])</code>.</p>\n<p>The implicit <code>None</code> for the batch dimension is actually used for the time dimension during inference.</p>\n<p>The name of the input layer should be <code>inputs</code>, and the <code>dtype</code> <code>tf.float32</code>.</p>\n<pre><code>inputs = tf.keras.Input(shape=(NUM_COLUMNS), dtype=tf.float32, name=\"inputs\")\n</code></pre>\n<h1>Model output shape</h1>\n<p>There are 59 unique characters in the ground truth strings, so the output shape should be <code>(59)</code>. The output should be a sequence of logits: <code>(num_frames, 59)</code> such that the inference code can take the argmax over axis <code>1</code>.</p>\n<p>For example to get random outputs with a randomly initialized <code>Dense</code> layer:</p>\n<pre><code>x = tf.keras.layers.Dense(59)(x)\nout = tf.keras.layers.Activation(\"linear\", name=\"outputs\")(x)\n</code></pre>\n<p>The name of the output layer should be <code>outputs</code>.</p>\n<h1>Acknowledgements</h1>\n<p>Thanks to everyone who contributed in the discussions on the input and output format and to <a href=\"https://www.kaggle.com/wonderingalice\" target=\"_blank\">@wonderingalice</a> in particular for the submission notebook (go upvote it 😉)</p>",
      "rawMarkdown": "@wonderingalice has made a notebook which correctly runs and can make submissions: https://www.kaggle.com/code/wonderingalice/dummy-submission-with-asl-islr-competition-kernel 🎉\n\nThis is a write-up of the formats required to make a valid submission. Now that submissions are finally working, we can use this notebook as a starting point!\n\n# Library versions\n\nThe original notebook uses an environment with these versions:\n\n- Tensorflow 2.11.0\n- Python 3.7.12\n\nThis is a pinned environment from the previous competition on ASL: https://www.kaggle.com/competitions/asl-signs/\n\n# Selecting input columns\n\nYou can select which input columns to use by providing a file called `inference_args.json`. The JSON file should contain one key `selected_columns` with as corresponding value a list of the column names in the pandas files. For example:\n\n```\n{\n  \"selected_columns\": [\"x_right_hand_0\", \"y_right_hand_0\", \"z_right_hand_0\"]\n}\n```\n\n# Model input shape\n\nThe model will get a single input at a time during inference, so the input shape should be `(NUM_COLUMNS)`, where `NUM_COLUMNS` is the number of input columns, or `len(inference_args[\"selected_columns\"])`.\n\nThe implicit `None` for the batch dimension is actually used for the time dimension during inference.\n\nThe name of the input layer should be `inputs`, and the `dtype` `tf.float32`.\n\n```\ninputs = tf.keras.Input(shape=(NUM_COLUMNS), dtype=tf.float32, name=\"inputs\")\n```\n\n# Model output shape\n\nThere are 59 unique characters in the ground truth strings, so the output shape should be `(59)`. The output should be a sequence of logits: `(num_frames, 59)` such that the inference code can take the argmax over axis `1`.\n\nFor example to get random outputs with a randomly initialized `Dense` layer:\n\n```\nx = tf.keras.layers.Dense(59)(x)\nout = tf.keras.layers.Activation(\"linear\", name=\"outputs\")(x)\n```\n\nThe name of the output layer should be `outputs`.\n\n# Acknowledgements\n\nThanks to everyone who contributed in the discussions on the input and output format and to @wonderingalice in particular for the submission notebook (go upvote it 😉)",
      "votes": 14
    },
    {
      "id": 2272457,
      "postDate": "2023-05-24T14:01:04.673Z",
      "content": "<p>Excellent work, now the competition can finally start!</p>\n<p>Is it required for the output to be of shape <code>(num_frames, 59)</code> or can we also just give the characters.<br>\nFor example with 100 input frames and 5characters predicted with 95 padding token to output <code>(5,59)</code> instead of <code>(100,59)</code></p>",
      "rawMarkdown": "Excellent work, now the competition can finally start!\n\nIs it required for the output to be of shape `(num_frames, 59)` or can we also just give the characters.\nFor example with 100 input frames and 5characters predicted with 95 padding token to output `(5,59)` instead of `(100,59)`",
      "votes": 2,
      "replies": [
        {
          "id": 2272466,
          "postDate": "2023-05-24T14:11:45.810Z",
          "content": "<p>I think that should be fine, my wording of <code>num_frames</code> was unfortunate, I rather meant <code>num_output_characters</code> there. As long as there's a 2D output it should be fine I think.</p>",
          "rawMarkdown": "I think that should be fine, my wording of `num_frames` was unfortunate, I rather meant `num_output_characters` there. As long as there's a 2D output it should be fine I think.",
          "replies": [
            {
              "id": 2272500,
              "postDate": "2023-05-24T14:40:17.853Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">@markwijkhuizen</a>,</p>\n<p>I can confirm that the dummy model that Lonnie posted in the discussion some time ago also gets through the scoring pipeline. That model does not have the number of output characters equal to the number of input frames. I haven't tested any extreme cases yet, though and the organisers are still maintaining their radio silence.</p>",
              "rawMarkdown": "Hi @markwijkhuizen,\n\nI can confirm that the dummy model that Lonnie posted in the discussion some time ago also gets through the scoring pipeline. That model does not have the number of output characters equal to the number of input frames. I haven't tested any extreme cases yet, though and the organisers are still maintaining their radio silence.\n",
              "votes": 1
            },
            {
              "id": 2356105,
              "postDate": "2023-07-23T22:26:18.193Z",
              "content": "<p>Do you know if the output shape has to be 59 on the second dim?</p>",
              "rawMarkdown": "Do you know if the output shape has to be 59 on the second dim?"
            }
          ]
        },
        {
          "id": 2358579,
          "postDate": "2023-07-25T15:50:03.497Z",
          "content": "<p>My output matches this exactly, but I still get submission errors. I have double and triple-checked everything, but nothing seems to work. I'm beginning to think this is not the correct output format:</p>\n<p><a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/426914\" target=\"_blank\">https://www.kaggle.com/competitions/asl-fingerspelling/discussion/426914</a></p>",
          "rawMarkdown": "My output matches this exactly, but I still get submission errors. I have double and triple-checked everything, but nothing seems to work. I'm beginning to think this is not the correct output format:\n\nhttps://www.kaggle.com/competitions/asl-fingerspelling/discussion/426914"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2272457,
      "author_name": "Mark Wijkhuizen",
      "author_url": "",
      "post_date": "2023-05-24T14:01:04.673000",
      "content": "<p>Excellent work, now the competition can finally start!</p>\n<p>Is it required for the output to be of shape <code>(num_frames, 59)</code> or can we also just give the characters.<br>\nFor example with 100 input frames and 5characters predicted with 95 padding token to output <code>(5,59)</code> instead of <code>(100,59)</code></p>",
      "votes": 2,
      "replies": [
        {
          "id": 2272466,
          "author_name": "Mathieu De Coster",
          "author_url": "",
          "post_date": "2023-05-24T14:11:45.810000",
          "content": "<p>I think that should be fine, my wording of <code>num_frames</code> was unfortunate, I rather meant <code>num_output_characters</code> there. As long as there's a 2D output it should be fine I think.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2272500,
              "author_name": "Wondering Alice",
              "author_url": "",
              "post_date": "2023-05-24T14:40:17.853000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">@markwijkhuizen</a>,</p>\n<p>I can confirm that the dummy model that Lonnie posted in the discussion some time ago also gets through the scoring pipeline. That model does not have the number of output characters equal to the number of input frames. I haven't tested any extreme cases yet, though and the organisers are still maintaining their radio silence.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2356105,
              "author_name": "Adriano Passos",
              "author_url": "",
              "post_date": "2023-07-23T22:26:18.193000",
              "content": "<p>Do you know if the output shape has to be 59 on the second dim?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2358579,
          "author_name": "Tarek Sami",
          "author_url": "",
          "post_date": "2023-07-25T15:50:03.497000",
          "content": "<p>My output matches this exactly, but I still get submission errors. I have double and triple-checked everything, but nothing seems to work. I'm beginning to think this is not the correct output format:</p>\n<p><a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/426914\" target=\"_blank\">https://www.kaggle.com/competitions/asl-fingerspelling/discussion/426914</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2272232": "@wonderingalice has made a notebook which correctly runs and can make submissions: https://www.kaggle.com/code/wonderingalice/dummy-submission-with-asl-islr-competition-kernel 🎉\n\nThis is a write-up of the formats required to make a valid submission. Now that submissions are finally working, we can use this notebook as a starting point!\n\n# Library versions\n\nThe original notebook uses an environment with these versions:\n\n- Tensorflow 2.11.0\n- Python 3.7.12\n\nThis is a pinned environment from the previous competition on ASL: https://www.kaggle.com/competitions/asl-signs/\n\n# Selecting input columns\n\nYou can select which input columns to use by providing a file called `inference_args.json`. The JSON file should contain one key `selected_columns` with as corresponding value a list of the column names in the pandas files. For example:\n\n```\n{\n  \"selected_columns\": [\"x_right_hand_0\", \"y_right_hand_0\", \"z_right_hand_0\"]\n}\n```\n\n# Model input shape\n\nThe model will get a single input at a time during inference, so the input shape should be `(NUM_COLUMNS)`, where `NUM_COLUMNS` is the number of input columns, or `len(inference_args[\"selected_columns\"])`.\n\nThe implicit `None` for the batch dimension is actually used for the time dimension during inference.\n\nThe name of the input layer should be `inputs`, and the `dtype` `tf.float32`.\n\n```\ninputs = tf.keras.Input(shape=(NUM_COLUMNS), dtype=tf.float32, name=\"inputs\")\n```\n\n# Model output shape\n\nThere are 59 unique characters in the ground truth strings, so the output shape should be `(59)`. The output should be a sequence of logits: `(num_frames, 59)` such that the inference code can take the argmax over axis `1`.\n\nFor example to get random outputs with a randomly initialized `Dense` layer:\n\n```\nx = tf.keras.layers.Dense(59)(x)\nout = tf.keras.layers.Activation(\"linear\", name=\"outputs\")(x)\n```\n\nThe name of the output layer should be `outputs`.\n\n# Acknowledgements\n\nThanks to everyone who contributed in the discussions on the input and output format and to @wonderingalice in particular for the submission notebook (go upvote it 😉)",
    "2272457": "Excellent work, now the competition can finally start!\n\nIs it required for the output to be of shape `(num_frames, 59)` or can we also just give the characters.\nFor example with 100 input frames and 5characters predicted with 95 padding token to output `(5,59)` instead of `(100,59)`"
  }
}