{
  "id": 409692,
  "title": "My proposal to solve this Kaggle challenge is to utilize the TSSCI method, which integrates sequences of body, hand, and facial expressions into a super object in a single color image.",
  "url": "/competitions/asl-fingerspelling/discussion/409692",
  "author_name": "Yoram Segal",
  "post_date": "2023-05-12T07:04:53.207000",
  "votes": 10,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I recently published an article as part of my Ph.D. thesis, currently in the Preprint status. The article presents a comprehensive approach to analyzing human movements in a generic manner.</p>\n<p>The TSSCI method involves converting a series of video frames, represented by key points or landmarks generated using MediaPipe, into a single color image. This color image, known as TSSCI (Tree Structure Skeleton Color Image or Time-Series Single Color Image), encapsulates the entire movement from its initiation to completion.</p>\n<p>The advantage of using TSSCI is that it transforms the MediaPipe time-domain representation of motion into a single spacial RGB image format. Consequently, it becomes compatible with various deep learning networks, especially CNN networks, or, in this Kaggle challenge, with classic Autoencoder networks. By applying TSSCI as input to the Autoencoder network which encodes it into a latent vector, one can then proceed with decoding and reconstructing the original textual description (token description) of the movement.</p>\n<p>To initiate the process, it is necessary to generate a dataset comprising pairs of TSSCI images and their corresponding labels (tokens). These labels consist of tokens that have been converted from the movement description into token representation. The subsequent step involves training the Autoencoder to perform the conversion of TSSCI images into tokens.</p>\n<p>Alternatively, another approach revolves around utilizing a classical classification network. Similar to how CNN networks excel at classifying various objects like cats or dogs, in this case, each type of movement is represented as an object within the TSSCI. Consequently, a CNN network can effectively distinguish between different objects, thereby enabling the classification of distinct movements.</p>\n<p><a href=\"https://www.preprints.org/manuscript/202304.1268/v1\" target=\"_blank\">Link to the article.</a></p>\n<p>I am eager to extend my assistance and collaborate with anyone who requires my expertise in this field. To provide a better understanding of my research, I have included a link to a recording of a seminar I conducted at Ben Gurion University for your convenience: <a href=\"https://youtu.be/QQf-pyQw8Wc\" target=\"_blank\">YouTube - Yoram Segal Article explanation</a> </p>\n<p>Remark:<br>\nThe process of TSSCI production is straightforward. It involves organizing the key points (landmarks)  derived from MediaPipe into a vector, where the X component is assigned to the red channel, the Y component to the green channel, and the Z component to the blue channel. For a more comprehensive understanding of this procedure, I recommend reading the <a href=\"https://www.preprints.org/manuscript/202304.1268/v1\" target=\"_blank\">article</a> or watching the accompanying <a href=\"https://youtu.be/QQf-pyQw8Wc\" target=\"_blank\">YouTube </a> video, which provides detailed explanations.</p>\n<p>Yoram<br>\n<a href=\"mailto:yoramse@post.bgu.ac.il\">yoramse@post.bgu.ac.il</a></p>",
  "messages": [
    {
      "id": 2255983,
      "postDate": "2023-05-12T07:04:53.207Z",
      "content": "<p>I recently published an article as part of my Ph.D. thesis, currently in the Preprint status. The article presents a comprehensive approach to analyzing human movements in a generic manner.</p>\n<p>The TSSCI method involves converting a series of video frames, represented by key points or landmarks generated using MediaPipe, into a single color image. This color image, known as TSSCI (Tree Structure Skeleton Color Image or Time-Series Single Color Image), encapsulates the entire movement from its initiation to completion.</p>\n<p>The advantage of using TSSCI is that it transforms the MediaPipe time-domain representation of motion into a single spacial RGB image format. Consequently, it becomes compatible with various deep learning networks, especially CNN networks, or, in this Kaggle challenge, with classic Autoencoder networks. By applying TSSCI as input to the Autoencoder network which encodes it into a latent vector, one can then proceed with decoding and reconstructing the original textual description (token description) of the movement.</p>\n<p>To initiate the process, it is necessary to generate a dataset comprising pairs of TSSCI images and their corresponding labels (tokens). These labels consist of tokens that have been converted from the movement description into token representation. The subsequent step involves training the Autoencoder to perform the conversion of TSSCI images into tokens.</p>\n<p>Alternatively, another approach revolves around utilizing a classical classification network. Similar to how CNN networks excel at classifying various objects like cats or dogs, in this case, each type of movement is represented as an object within the TSSCI. Consequently, a CNN network can effectively distinguish between different objects, thereby enabling the classification of distinct movements.</p>\n<p><a href=\"https://www.preprints.org/manuscript/202304.1268/v1\" target=\"_blank\">Link to the article.</a></p>\n<p>I am eager to extend my assistance and collaborate with anyone who requires my expertise in this field. To provide a better understanding of my research, I have included a link to a recording of a seminar I conducted at Ben Gurion University for your convenience: <a href=\"https://youtu.be/QQf-pyQw8Wc\" target=\"_blank\">YouTube - Yoram Segal Article explanation</a> </p>\n<p>Remark:<br>\nThe process of TSSCI production is straightforward. It involves organizing the key points (landmarks)  derived from MediaPipe into a vector, where the X component is assigned to the red channel, the Y component to the green channel, and the Z component to the blue channel. For a more comprehensive understanding of this procedure, I recommend reading the <a href=\"https://www.preprints.org/manuscript/202304.1268/v1\" target=\"_blank\">article</a> or watching the accompanying <a href=\"https://youtu.be/QQf-pyQw8Wc\" target=\"_blank\">YouTube </a> video, which provides detailed explanations.</p>\n<p>Yoram<br>\n<a href=\"mailto:yoramse@post.bgu.ac.il\">yoramse@post.bgu.ac.il</a></p>",
      "rawMarkdown": "I recently published an article as part of my Ph.D. thesis, currently in the Preprint status. The article presents a comprehensive approach to analyzing human movements in a generic manner.\n\nThe TSSCI method involves converting a series of video frames, represented by key points or landmarks generated using MediaPipe, into a single color image. This color image, known as TSSCI (Tree Structure Skeleton Color Image or Time-Series Single Color Image), encapsulates the entire movement from its initiation to completion.\n\nThe advantage of using TSSCI is that it transforms the MediaPipe time-domain representation of motion into a single spacial RGB image format. Consequently, it becomes compatible with various deep learning networks, especially CNN networks, or, in this Kaggle challenge, with classic Autoencoder networks. By applying TSSCI as input to the Autoencoder network which encodes it into a latent vector, one can then proceed with decoding and reconstructing the original textual description (token description) of the movement.\n\nTo initiate the process, it is necessary to generate a dataset comprising pairs of TSSCI images and their corresponding labels (tokens). These labels consist of tokens that have been converted from the movement description into token representation. The subsequent step involves training the Autoencoder to perform the conversion of TSSCI images into tokens.\n\nAlternatively, another approach revolves around utilizing a classical classification network. Similar to how CNN networks excel at classifying various objects like cats or dogs, in this case, each type of movement is represented as an object within the TSSCI. Consequently, a CNN network can effectively distinguish between different objects, thereby enabling the classification of distinct movements.\n\n[Link to the article.](https://www.preprints.org/manuscript/202304.1268/v1)\n\nI am eager to extend my assistance and collaborate with anyone who requires my expertise in this field. To provide a better understanding of my research, I have included a link to a recording of a seminar I conducted at Ben Gurion University for your convenience: [YouTube - Yoram Segal Article explanation] (https://youtu.be/QQf-pyQw8Wc) \n\nRemark:\nThe process of TSSCI production is straightforward. It involves organizing the key points (landmarks)  derived from MediaPipe into a vector, where the X component is assigned to the red channel, the Y component to the green channel, and the Z component to the blue channel. For a more comprehensive understanding of this procedure, I recommend reading the [article](https://www.preprints.org/manuscript/202304.1268/v1) or watching the accompanying [YouTube ](https://youtu.be/QQf-pyQw8Wc) video, which provides detailed explanations.\n\nYoram\nyoramse@post.bgu.ac.il",
      "votes": 10
    },
    {
      "id": 2256051,
      "postDate": "2023-05-12T08:34:45.837Z",
      "content": "<p>Very interesting, congrats and good luck for the publication! </p>",
      "rawMarkdown": "Very interesting, congrats and good luck for the publication! ",
      "votes": 1,
      "replies": [
        {
          "id": 2256161,
          "postDate": "2023-05-12T09:42:15.887Z",
          "content": "<p>Thank you very much. I hope the information provided proves helpful to you. I am more than willing to assist in any way that you may find useful.</p>",
          "rawMarkdown": "Thank you very much. I hope the information provided proves helpful to you. I am more than willing to assist in any way that you may find useful.",
          "votes": -1
        }
      ]
    },
    {
      "id": 2325637,
      "postDate": "2023-07-01T13:31:49.790Z",
      "content": "<p>I guess, what you're proposing is quite similar to the 2nd place solution in the previous competition (based on the same data set - Isolated Sign Language Recognition). Here is the <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406306\" target=\"_blank\">link</a> for the solution, where the approach is discussed. <br>\nP.S. : I may be wrong. 😊</p>",
      "rawMarkdown": "I guess, what you're proposing is quite similar to the 2nd place solution in the previous competition (based on the same data set - Isolated Sign Language Recognition). Here is the [link](https://www.kaggle.com/competitions/asl-signs/discussion/406306) for the solution, where the approach is discussed. \nP.S. : I may be wrong. 😊"
    },
    {
      "id": 2257220,
      "postDate": "2023-05-13T06:32:53.637Z",
      "content": "<p>good luck and keep kaggling</p>",
      "rawMarkdown": "good luck and keep kaggling"
    },
    {
      "id": 2257215,
      "postDate": "2023-05-13T06:26:21.303Z",
      "content": "<p><a href=\"https://www.linkedin.com/posts/yoram-segal-31079b140_using-efficientnet-b7-cnn-variational-activity-7063033653437566976-hNsM?utm_source=share&amp;utm_medium=member_desktop\" target=\"_blank\">linkedin</a></p>",
      "rawMarkdown": "[linkedin](https://www.linkedin.com/posts/yoram-segal-31079b140_using-efficientnet-b7-cnn-variational-activity-7063033653437566976-hNsM?utm_source=share&utm_medium=member_desktop)"
    }
  ],
  "comments": [
    {
      "id": 2256051,
      "author_name": "Beubeu4L",
      "author_url": "",
      "post_date": "2023-05-12T08:34:45.837000",
      "content": "<p>Very interesting, congrats and good luck for the publication! </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2256161,
          "author_name": "Yoram Segal",
          "author_url": "",
          "post_date": "2023-05-12T09:42:15.887000",
          "content": "<p>Thank you very much. I hope the information provided proves helpful to you. I am more than willing to assist in any way that you may find useful.</p>",
          "votes": -1,
          "replies": []
        }
      ]
    },
    {
      "id": 2325637,
      "author_name": "Girish Kotwal",
      "author_url": "",
      "post_date": "2023-07-01T13:31:49.790000",
      "content": "<p>I guess, what you're proposing is quite similar to the 2nd place solution in the previous competition (based on the same data set - Isolated Sign Language Recognition). Here is the <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406306\" target=\"_blank\">link</a> for the solution, where the approach is discussed. <br>\nP.S. : I may be wrong. 😊</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2257220,
      "author_name": "Aisuluu Ulan kyzy",
      "author_url": "",
      "post_date": "2023-05-13T06:32:53.637000",
      "content": "<p>good luck and keep kaggling</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2257215,
      "author_name": "Yoram Segal",
      "author_url": "",
      "post_date": "2023-05-13T06:26:21.303000",
      "content": "<p><a href=\"https://www.linkedin.com/posts/yoram-segal-31079b140_using-efficientnet-b7-cnn-variational-activity-7063033653437566976-hNsM?utm_source=share&amp;utm_medium=member_desktop\" target=\"_blank\">linkedin</a></p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2255983": "I recently published an article as part of my Ph.D. thesis, currently in the Preprint status. The article presents a comprehensive approach to analyzing human movements in a generic manner.\n\nThe TSSCI method involves converting a series of video frames, represented by key points or landmarks generated using MediaPipe, into a single color image. This color image, known as TSSCI (Tree Structure Skeleton Color Image or Time-Series Single Color Image), encapsulates the entire movement from its initiation to completion.\n\nThe advantage of using TSSCI is that it transforms the MediaPipe time-domain representation of motion into a single spacial RGB image format. Consequently, it becomes compatible with various deep learning networks, especially CNN networks, or, in this Kaggle challenge, with classic Autoencoder networks. By applying TSSCI as input to the Autoencoder network which encodes it into a latent vector, one can then proceed with decoding and reconstructing the original textual description (token description) of the movement.\n\nTo initiate the process, it is necessary to generate a dataset comprising pairs of TSSCI images and their corresponding labels (tokens). These labels consist of tokens that have been converted from the movement description into token representation. The subsequent step involves training the Autoencoder to perform the conversion of TSSCI images into tokens.\n\nAlternatively, another approach revolves around utilizing a classical classification network. Similar to how CNN networks excel at classifying various objects like cats or dogs, in this case, each type of movement is represented as an object within the TSSCI. Consequently, a CNN network can effectively distinguish between different objects, thereby enabling the classification of distinct movements.\n\n[Link to the article.](https://www.preprints.org/manuscript/202304.1268/v1)\n\nI am eager to extend my assistance and collaborate with anyone who requires my expertise in this field. To provide a better understanding of my research, I have included a link to a recording of a seminar I conducted at Ben Gurion University for your convenience: [YouTube - Yoram Segal Article explanation] (https://youtu.be/QQf-pyQw8Wc) \n\nRemark:\nThe process of TSSCI production is straightforward. It involves organizing the key points (landmarks)  derived from MediaPipe into a vector, where the X component is assigned to the red channel, the Y component to the green channel, and the Z component to the blue channel. For a more comprehensive understanding of this procedure, I recommend reading the [article](https://www.preprints.org/manuscript/202304.1268/v1) or watching the accompanying [YouTube ](https://youtu.be/QQf-pyQw8Wc) video, which provides detailed explanations.\n\nYoram\nyoramse@post.bgu.ac.il",
    "2256051": "Very interesting, congrats and good luck for the publication! ",
    "2325637": "I guess, what you're proposing is quite similar to the 2nd place solution in the previous competition (based on the same data set - Isolated Sign Language Recognition). Here is the [link](https://www.kaggle.com/competitions/asl-signs/discussion/406306) for the solution, where the approach is discussed. \nP.S. : I may be wrong. 😊",
    "2257220": "good luck and keep kaggling",
    "2257215": "[linkedin](https://www.linkedin.com/posts/yoram-segal-31079b140_using-efficientnet-b7-cnn-variational-activity-7063033653437566976-hNsM?utm_source=share&utm_medium=member_desktop)"
  }
}