{
  "id": 409614,
  "title": "Sharing some thoughts.",
  "url": "/competitions/asl-fingerspelling/discussion/409614",
  "author_name": "Scott Roden",
  "post_date": "2023-05-11T22:27:17.993000",
  "votes": 2,
  "comment_count": 6,
  "views": 0,
  "content": "<p>So I understand how sign language works and I also understand how ML and Algos work. My thoughts on this are that the problems you are going to come across will all relate to picking up when one word or letter moves to another. The other problem is getting it to recognise when the letter or word has been signed. You will probably need to factor in some sort of movement to show this change and that change can be different for each person. It's easy for a human brain to interpret because they can use context. Personally I would build it around a context  algorithm/engine. That way you can assess the logic before you interpret it. Good luck on your endeavours.</p>",
  "messages": [
    {
      "id": 2255647,
      "postDate": "2023-05-11T22:27:17.993Z",
      "content": "<p>So I understand how sign language works and I also understand how ML and Algos work. My thoughts on this are that the problems you are going to come across will all relate to picking up when one word or letter moves to another. The other problem is getting it to recognise when the letter or word has been signed. You will probably need to factor in some sort of movement to show this change and that change can be different for each person. It's easy for a human brain to interpret because they can use context. Personally I would build it around a context  algorithm/engine. That way you can assess the logic before you interpret it. Good luck on your endeavours.</p>",
      "rawMarkdown": "So I understand how sign language works and I also understand how ML and Algos work. My thoughts on this are that the problems you are going to come across will all relate to picking up when one word or letter moves to another. The other problem is getting it to recognise when the letter or word has been signed. You will probably need to factor in some sort of movement to show this change and that change can be different for each person. It's easy for a human brain to interpret because they can use context. Personally I would build it around a context  algorithm/engine. That way you can assess the logic before you interpret it. Good luck on your endeavours.",
      "votes": 2
    },
    {
      "id": 2258624,
      "postDate": "2023-05-14T11:06:16.607Z",
      "content": "<p>Another consideration is characters that have movement like J and Z and those that do not like the rest of A-Y.  And handling double letters or characters, e.g.  E L O always slide away when finger spelling doubles of them, like FEELING, HELLO, NOON.  Imagine movements could differ per signer even if just left/right handed.   </p>",
      "rawMarkdown": "Another consideration is characters that have movement like J and Z and those that do not like the rest of A-Y.  And handling double letters or characters, e.g.  E L O always slide away when finger spelling doubles of them, like FEELING, HELLO, NOON.  Imagine movements could differ per signer even if just left/right handed.   "
    },
    {
      "id": 2255967,
      "postDate": "2023-05-12T06:47:38.293Z",
      "content": "<p>You are right - it is quite hard to segment characters in landmark sequences (as well as in images or sounds). Moreover, some kind of RNN is required anyway to make your model more robust. We have to deal with prases, not random character sequences. I believe RNN would learn something useful that will correct some random mistakes (like 'worrd' -&gt; 'word', etc). </p>\n<p>If you are interested in a segmentation-free approach - I wrote some thoughts in another discussion - <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/409512\" target=\"_blank\">Sequence modeling with CTC loss</a></p>",
      "rawMarkdown": "You are right - it is quite hard to segment characters in landmark sequences (as well as in images or sounds). Moreover, some kind of RNN is required anyway to make your model more robust. We have to deal with prases, not random character sequences. I believe RNN would learn something useful that will correct some random mistakes (like 'worrd' -> 'word', etc). \n\nIf you are interested in a segmentation-free approach - I wrote some thoughts in another discussion - [Sequence modeling with CTC loss](https://www.kaggle.com/competitions/asl-fingerspelling/discussion/409512)"
    },
    {
      "id": 2255721,
      "postDate": "2023-05-12T00:18:35.060Z",
      "content": "<p>Hi, I am a newbie ( I mean I know some ML) but when it comes to context, what does this mean ? And if I understand your tip, its to wait a bit before interpreting the whole sequence ?</p>",
      "rawMarkdown": "Hi, I am a newbie ( I mean I know some ML) but when it comes to context, what does this mean ? And if I understand your tip, its to wait a bit before interpreting the whole sequence ?",
      "replies": [
        {
          "id": 2258646,
          "postDate": "2023-05-14T11:36:51.757Z",
          "content": "<p>Context is like you said when you know the full sentence but it also could be the conversation itself. Interpreting individual signs is where you start but then you need to bring that all together for a sentence. My thoughts are that you would take all the signs then run them through to determine the most probable sentence. That's where the ML comes in. You could technically make it faster over time with more data as you could quickly determine what a sentence is going to be with enough data. Determining signs based on a video is a whole other problem as others have pointed out. Then you have the different versions of sign language to take into account. Though to be fair to start off with I would pick one variation, perfect that and adjust that for the others. As for context of conversation you could also add ML in for the question being asked or the conversation itself to raise the probability the translation is going to be correct but that leaves you open to curve balls or unexpected answers so that should just be a fall back and not a primary translation tool. Maybe like a secondary check. I'm a newbie myself so I'm interested in theory at the moment before I jump in and start programming it. I'd like to think that understanding what I want it to do is more important than getting it to do it. I made that mistake many years ago writing programs and not taking the time to think about it fully and considering all the what ifs like I have all these options of how to do it but which is the right one.</p>",
          "rawMarkdown": "Context is like you said when you know the full sentence but it also could be the conversation itself. Interpreting individual signs is where you start but then you need to bring that all together for a sentence. My thoughts are that you would take all the signs then run them through to determine the most probable sentence. That's where the ML comes in. You could technically make it faster over time with more data as you could quickly determine what a sentence is going to be with enough data. Determining signs based on a video is a whole other problem as others have pointed out. Then you have the different versions of sign language to take into account. Though to be fair to start off with I would pick one variation, perfect that and adjust that for the others. As for context of conversation you could also add ML in for the question being asked or the conversation itself to raise the probability the translation is going to be correct but that leaves you open to curve balls or unexpected answers so that should just be a fall back and not a primary translation tool. Maybe like a secondary check. I'm a newbie myself so I'm interested in theory at the moment before I jump in and start programming it. I'd like to think that understanding what I want it to do is more important than getting it to do it. I made that mistake many years ago writing programs and not taking the time to think about it fully and considering all the what ifs like I have all these options of how to do it but which is the right one.",
          "votes": 1,
          "replies": [
            {
              "id": 2259361,
              "postDate": "2023-05-14T23:15:39.430Z",
              "content": "<p>Ahh so say your model predicts each words for each hand sign(or group of frames). This would be the first part. The second part would be, as an example, a postprocessing step where you would take the output given for the whole sequence and use another model to predict the context (meaning, if I understand correctly, placement of words, words groupings, synonyms maybe ?, etc).</p>",
              "rawMarkdown": "Ahh so say your model predicts each words for each hand sign(or group of frames). This would be the first part. The second part would be, as an example, a postprocessing step where you would take the output given for the whole sequence and use another model to predict the context (meaning, if I understand correctly, placement of words, words groupings, synonyms maybe ?, etc).\n"
            },
            {
              "id": 2260218,
              "postDate": "2023-05-15T14:35:23.977Z",
              "content": "<p>As I'm just working on just theory alone I think the first part would be to get the signs. Now for that I would translate each finger and hand into sticks taking into account the direction of the hands from the body but that could be built into the model automatically if you have enough input data of various angles. Now from that as someone else pointed out it would be RNN and CTC which is well worth a read. I guess you could use methods already employed in speech recognition because it's not that different. This approach if not already used could already be applied for lip reading. If it is then that could be applied to sign language. I guess the context model could also be built from sentences fed into it or an already available library of language. The tricky bit now which I honestly don't know the answer to yet is how do you get all that to be fast enough. There are a lot of people smarter than me that probably know that answer already. As I said I'm just dipping my toe into the theory first. I know what I want to use ML for and I know it will revolutionise business across the board saving millions of hours work but it's how to implement it correctly so I guess this is a good place to start.</p>",
              "rawMarkdown": "As I'm just working on just theory alone I think the first part would be to get the signs. Now for that I would translate each finger and hand into sticks taking into account the direction of the hands from the body but that could be built into the model automatically if you have enough input data of various angles. Now from that as someone else pointed out it would be RNN and CTC which is well worth a read. I guess you could use methods already employed in speech recognition because it's not that different. This approach if not already used could already be applied for lip reading. If it is then that could be applied to sign language. I guess the context model could also be built from sentences fed into it or an already available library of language. The tricky bit now which I honestly don't know the answer to yet is how do you get all that to be fast enough. There are a lot of people smarter than me that probably know that answer already. As I said I'm just dipping my toe into the theory first. I know what I want to use ML for and I know it will revolutionise business across the board saving millions of hours work but it's how to implement it correctly so I guess this is a good place to start."
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2258624,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "2023-05-14T11:06:16.607000",
      "content": "<p>Another consideration is characters that have movement like J and Z and those that do not like the rest of A-Y.  And handling double letters or characters, e.g.  E L O always slide away when finger spelling doubles of them, like FEELING, HELLO, NOON.  Imagine movements could differ per signer even if just left/right handed.   </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2255967,
      "author_name": "Mykola",
      "author_url": "",
      "post_date": "2023-05-12T06:47:38.293000",
      "content": "<p>You are right - it is quite hard to segment characters in landmark sequences (as well as in images or sounds). Moreover, some kind of RNN is required anyway to make your model more robust. We have to deal with prases, not random character sequences. I believe RNN would learn something useful that will correct some random mistakes (like 'worrd' -&gt; 'word', etc). </p>\n<p>If you are interested in a segmentation-free approach - I wrote some thoughts in another discussion - <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/409512\" target=\"_blank\">Sequence modeling with CTC loss</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2255721,
      "author_name": "Yazan Maarouf",
      "author_url": "",
      "post_date": "2023-05-12T00:18:35.060000",
      "content": "<p>Hi, I am a newbie ( I mean I know some ML) but when it comes to context, what does this mean ? And if I understand your tip, its to wait a bit before interpreting the whole sequence ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2258646,
          "author_name": "Scott Roden",
          "author_url": "",
          "post_date": "2023-05-14T11:36:51.757000",
          "content": "<p>Context is like you said when you know the full sentence but it also could be the conversation itself. Interpreting individual signs is where you start but then you need to bring that all together for a sentence. My thoughts are that you would take all the signs then run them through to determine the most probable sentence. That's where the ML comes in. You could technically make it faster over time with more data as you could quickly determine what a sentence is going to be with enough data. Determining signs based on a video is a whole other problem as others have pointed out. Then you have the different versions of sign language to take into account. Though to be fair to start off with I would pick one variation, perfect that and adjust that for the others. As for context of conversation you could also add ML in for the question being asked or the conversation itself to raise the probability the translation is going to be correct but that leaves you open to curve balls or unexpected answers so that should just be a fall back and not a primary translation tool. Maybe like a secondary check. I'm a newbie myself so I'm interested in theory at the moment before I jump in and start programming it. I'd like to think that understanding what I want it to do is more important than getting it to do it. I made that mistake many years ago writing programs and not taking the time to think about it fully and considering all the what ifs like I have all these options of how to do it but which is the right one.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2259361,
              "author_name": "Yazan Maarouf",
              "author_url": "",
              "post_date": "2023-05-14T23:15:39.430000",
              "content": "<p>Ahh so say your model predicts each words for each hand sign(or group of frames). This would be the first part. The second part would be, as an example, a postprocessing step where you would take the output given for the whole sequence and use another model to predict the context (meaning, if I understand correctly, placement of words, words groupings, synonyms maybe ?, etc).</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2260218,
              "author_name": "Scott Roden",
              "author_url": "",
              "post_date": "2023-05-15T14:35:23.977000",
              "content": "<p>As I'm just working on just theory alone I think the first part would be to get the signs. Now for that I would translate each finger and hand into sticks taking into account the direction of the hands from the body but that could be built into the model automatically if you have enough input data of various angles. Now from that as someone else pointed out it would be RNN and CTC which is well worth a read. I guess you could use methods already employed in speech recognition because it's not that different. This approach if not already used could already be applied for lip reading. If it is then that could be applied to sign language. I guess the context model could also be built from sentences fed into it or an already available library of language. The tricky bit now which I honestly don't know the answer to yet is how do you get all that to be fast enough. There are a lot of people smarter than me that probably know that answer already. As I said I'm just dipping my toe into the theory first. I know what I want to use ML for and I know it will revolutionise business across the board saving millions of hours work but it's how to implement it correctly so I guess this is a good place to start.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2255647": "So I understand how sign language works and I also understand how ML and Algos work. My thoughts on this are that the problems you are going to come across will all relate to picking up when one word or letter moves to another. The other problem is getting it to recognise when the letter or word has been signed. You will probably need to factor in some sort of movement to show this change and that change can be different for each person. It's easy for a human brain to interpret because they can use context. Personally I would build it around a context  algorithm/engine. That way you can assess the logic before you interpret it. Good luck on your endeavours.",
    "2258624": "Another consideration is characters that have movement like J and Z and those that do not like the rest of A-Y.  And handling double letters or characters, e.g.  E L O always slide away when finger spelling doubles of them, like FEELING, HELLO, NOON.  Imagine movements could differ per signer even if just left/right handed.   ",
    "2255967": "You are right - it is quite hard to segment characters in landmark sequences (as well as in images or sounds). Moreover, some kind of RNN is required anyway to make your model more robust. We have to deal with prases, not random character sequences. I believe RNN would learn something useful that will correct some random mistakes (like 'worrd' -> 'word', etc). \n\nIf you are interested in a segmentation-free approach - I wrote some thoughts in another discussion - [Sequence modeling with CTC loss](https://www.kaggle.com/competitions/asl-fingerspelling/discussion/409512)",
    "2255721": "Hi, I am a newbie ( I mean I know some ML) but when it comes to context, what does this mean ? And if I understand your tip, its to wait a bit before interpreting the whole sequence ?"
  }
}