{
  "id": 434440,
  "title": "Custom MovinetClassifier Concept  ",
  "url": "/competitions/asl-fingerspelling/discussion/434440",
  "author_name": "something4kag",
  "post_date": "2023-08-25T09:02:00.862000",
  "votes": 4,
  "comment_count": 1,
  "views": 0,
  "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F396181%2F4663191986e89b81a51041daba2605e8%2Fanimation_23842.gif?generation=1692951869820889&amp;alt=media\" alt=\"hand_animation_23842\"></p>\n<p>Above illustrates transforming the hand landmarks (for row 23842 phrase 2 in train) in tensorflow preprocessing layers into a \"clip\" to be used for inference with a custom classifier fine tuned for the 59 prediction characters.<br>\nMoViNet models for video classification have small streaming models that can be trained and inferred on CPU and converted to TF Lite.  </p>\n<p><a href=\"https://github.com/tensorflow/models/tree/master/official/projects/movinet\" target=\"_blank\">MoViNet model from TensorFlow Models</a><br>\ncustom classifier MovinetClassifier fine tune with pretrained weights e.g., movinet_a0_stream   (trained on Kinetics 600) <br>\ninput shape 1 x 1 x 172 x 172<br>\ncheckpoint ~15 MB  trained model ~30 MB  TF Lite ~10 MB</p>\n<p>Convert to TF SavedModel. Then the SavedModel can be converted to TF Lite using the TFLiteConverter<br>\n<code>converter = tf.lite.TFLiteConverter.from_saved_model(saved_model_dir)</code><br>\n<code>tflite_model = converter.convert()</code></p>\n<p><a href=\"https://arxiv.org/abs/2103.11511\" target=\"_blank\">MoViNets: Mobile Video Networks for Efficient Video Recognition</a><br>\nIn the paper, multiple-class labels per video were used for Charades. <br>\nOr per frame or selected frame inference was a possibility then post process with tensorflow unique_with_counts <br>\nand considerations for double numbers or letters.</p>\n<p>This was not finished for the competition, but adding here for interest and awareness of these models. <br>\nAnd a bit of fun.<br>\nIt could be a next gen Sign Language uses video content with mediapipe markers and this would fit nicely.</p>",
  "messages": [
    {
      "id": 2407785,
      "postDate": "2023-08-25T09:02:00.863Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F396181%2F4663191986e89b81a51041daba2605e8%2Fanimation_23842.gif?generation=1692951869820889&amp;alt=media\" alt=\"hand_animation_23842\"></p>\n<p>Above illustrates transforming the hand landmarks (for row 23842 phrase 2 in train) in tensorflow preprocessing layers into a \"clip\" to be used for inference with a custom classifier fine tuned for the 59 prediction characters.<br>\nMoViNet models for video classification have small streaming models that can be trained and inferred on CPU and converted to TF Lite.  </p>\n<p><a href=\"https://github.com/tensorflow/models/tree/master/official/projects/movinet\" target=\"_blank\">MoViNet model from TensorFlow Models</a><br>\ncustom classifier MovinetClassifier fine tune with pretrained weights e.g., movinet_a0_stream   (trained on Kinetics 600) <br>\ninput shape 1 x 1 x 172 x 172<br>\ncheckpoint ~15 MB  trained model ~30 MB  TF Lite ~10 MB</p>\n<p>Convert to TF SavedModel. Then the SavedModel can be converted to TF Lite using the TFLiteConverter<br>\n<code>converter = tf.lite.TFLiteConverter.from_saved_model(saved_model_dir)</code><br>\n<code>tflite_model = converter.convert()</code></p>\n<p><a href=\"https://arxiv.org/abs/2103.11511\" target=\"_blank\">MoViNets: Mobile Video Networks for Efficient Video Recognition</a><br>\nIn the paper, multiple-class labels per video were used for Charades. <br>\nOr per frame or selected frame inference was a possibility then post process with tensorflow unique_with_counts <br>\nand considerations for double numbers or letters.</p>\n<p>This was not finished for the competition, but adding here for interest and awareness of these models. <br>\nAnd a bit of fun.<br>\nIt could be a next gen Sign Language uses video content with mediapipe markers and this would fit nicely.</p>",
      "rawMarkdown": "![hand_animation_23842](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F396181%2F4663191986e89b81a51041daba2605e8%2Fanimation_23842.gif?generation=1692951869820889&alt=media)\n\nAbove illustrates transforming the hand landmarks (for row 23842 phrase 2 in train) in tensorflow preprocessing layers into a \"clip\" to be used for inference with a custom classifier fine tuned for the 59 prediction characters.\nMoViNet models for video classification have small streaming models that can be trained and inferred on CPU and converted to TF Lite.  \n\n[MoViNet model from TensorFlow Models](https://github.com/tensorflow/models/tree/master/official/projects/movinet)\ncustom classifier MovinetClassifier fine tune with pretrained weights e.g., movinet_a0_stream   (trained on Kinetics 600) \ninput shape 1 x 1 x 172 x 172\ncheckpoint ~15 MB  trained model ~30 MB  TF Lite ~10 MB\n\n Convert to TF SavedModel. Then the SavedModel can be converted to TF Lite using the TFLiteConverter\n`converter = tf.lite.TFLiteConverter.from_saved_model(saved_model_dir)`\n`tflite_model = converter.convert()`\n\n\n[MoViNets: Mobile Video Networks for Efficient Video Recognition](https://arxiv.org/abs/2103.11511)\nIn the paper, multiple-class labels per video were used for Charades. \nOr per frame or selected frame inference was a possibility then post process with tensorflow unique_with_counts \nand considerations for double numbers or letters.\n\nThis was not finished for the competition, but adding here for interest and awareness of these models. \nAnd a bit of fun.\nIt could be a next gen Sign Language uses video content with mediapipe markers and this would fit nicely.\n\n",
      "votes": 3
    },
    {
      "id": 2426093,
      "postDate": "2023-09-06T11:58:08.840Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2426093,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-09-06T11:58:08.840000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2407785": "![hand_animation_23842](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F396181%2F4663191986e89b81a51041daba2605e8%2Fanimation_23842.gif?generation=1692951869820889&alt=media)\n\nAbove illustrates transforming the hand landmarks (for row 23842 phrase 2 in train) in tensorflow preprocessing layers into a \"clip\" to be used for inference with a custom classifier fine tuned for the 59 prediction characters.\nMoViNet models for video classification have small streaming models that can be trained and inferred on CPU and converted to TF Lite.  \n\n[MoViNet model from TensorFlow Models](https://github.com/tensorflow/models/tree/master/official/projects/movinet)\ncustom classifier MovinetClassifier fine tune with pretrained weights e.g., movinet_a0_stream   (trained on Kinetics 600) \ninput shape 1 x 1 x 172 x 172\ncheckpoint ~15 MB  trained model ~30 MB  TF Lite ~10 MB\n\n Convert to TF SavedModel. Then the SavedModel can be converted to TF Lite using the TFLiteConverter\n`converter = tf.lite.TFLiteConverter.from_saved_model(saved_model_dir)`\n`tflite_model = converter.convert()`\n\n\n[MoViNets: Mobile Video Networks for Efficient Video Recognition](https://arxiv.org/abs/2103.11511)\nIn the paper, multiple-class labels per video were used for Charades. \nOr per frame or selected frame inference was a possibility then post process with tensorflow unique_with_counts \nand considerations for double numbers or letters.\n\nThis was not finished for the competition, but adding here for interest and awareness of these models. \nAnd a bit of fun.\nIt could be a next gen Sign Language uses video content with mediapipe markers and this would fit nicely.\n\n",
    "2426093": ""
  }
}