{
  "id": 416829,
  "title": "Pro-tips on fingerspelling #3: troublesome letters, contexts, and grammars",
  "url": "/competitions/asl-fingerspelling/discussion/416829",
  "author_name": "Thad Starner",
  "post_date": "2023-06-13T05:58:05.356000",
  "votes": 13,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Take a look at a fingerspelling chart:</p>\n<p><a href=\"https://www.startasl.com/fingerspelling/\" target=\"_blank\">https://www.startasl.com/fingerspelling/</a></p>\n<p>specifically, look at the letters A T N M S in order.  Note how the thumb starts outside the fist with A and then progresses to being under the first (T), second (N), and third finger(M).  With S it is in front of the hand.</p>\n<p>Now go to  <a href=\"https://asl.ms/()/spellit_large.htm\" target=\"_blank\">https://asl.ms/()/spellit_large.htm</a> and enter </p>\n<p>atnmsatnmsatnmsatms</p>\n<p>Turn the speed up.  How well  can you see the subtle differences? How well can hand posture trackers detect these subtle differences?</p>\n<p>Expert signers use context to help recognize which letter is being signed as the movement may be too fast, the lighting might be too poor, or the fingerspeller's dexterity might be limited.  We can do something similar by looking at transition between letters to give clues as to what was spelled.</p>\n<p>In my early research in cursive handwriting recognition using hidden Markov models I used \"contexts\" (recognizing groups of two, three, or even four letters together as one model instead of individual letters). </p>\n<p><a href=\"https://ieeexplore.ieee.org/iel2/3104/8834/00389432.pdf?casa_token=EdbBFGdcXHAAAAAA:EAh_OnxFt-bB1GTrwy5GjSgCcZc-1zMsipbP0WqhAn6gbODtsh5Bmr_5gQWGvKymGz2J7qYLYw\" target=\"_blank\">On-line cursive handwriting recognition using speech recognition methods</a></p>\n<p>It helps with co-articulation, where the previous and the next letter affects the appearance of the current letter. It also helps where a letter itself might be hard to identify, but the letters around it make it clear what is meant (for those who are familiar with cursive, compare the wi in wick versus the ui in quick). Prior probabilities can help weight recognition towards more common groups of letters, which also automatically happens to some extent based on the amount of representation groups of letters have in a training set.</p>\n<p>Stochastic grammars further help to reduce the \"perplexity\" of the problem, as certain words are more likely to follow other words.  For example, the \"quick brown fox\" is more likely than \"fox brown quick.\"  There are large public text datasets that can aid in creating such grammars.</p>\n<p>Here is another example that shows the value of contexts and grammar:  NT is a more common combination of letters than TN when in the middle of a word.  Off-hand, I can not even think of a word with TN in the middle. However, for addresses, TN may appear often, as it is the abbreviation for a state. Seeing \"Memphis, \" may predict an upcoming \"TN\"  where \"MOVEME\" might foreshadow an upcoming \"NT\"</p>\n<p>Of course, given a big enough neural net model and enough training data, these details could be learned. However, for smaller models and smaller data sets, finding a way to apply contextual knowledge may be a more efficient way to improve accuracy.  Perhaps pre-training a neural net on text datasets might help? Or perhaps finding a way to apply contexts and stochastic grammars may be better?  Offhand, I do not know of any papers that compare these methods and whether, in the limit, they are equivalent. However, I bet adding context in some way can be powerful in helping resolve when a signer is talking about ANTMAN or SANTAS.</p>",
  "messages": [
    {
      "id": 2300297,
      "postDate": "2023-06-13T05:58:05.357Z",
      "content": "<p>Take a look at a fingerspelling chart:</p>\n<p><a href=\"https://www.startasl.com/fingerspelling/\" target=\"_blank\">https://www.startasl.com/fingerspelling/</a></p>\n<p>specifically, look at the letters A T N M S in order.  Note how the thumb starts outside the fist with A and then progresses to being under the first (T), second (N), and third finger(M).  With S it is in front of the hand.</p>\n<p>Now go to  <a href=\"https://asl.ms/()/spellit_large.htm\" target=\"_blank\">https://asl.ms/()/spellit_large.htm</a> and enter </p>\n<p>atnmsatnmsatnmsatms</p>\n<p>Turn the speed up.  How well  can you see the subtle differences? How well can hand posture trackers detect these subtle differences?</p>\n<p>Expert signers use context to help recognize which letter is being signed as the movement may be too fast, the lighting might be too poor, or the fingerspeller's dexterity might be limited.  We can do something similar by looking at transition between letters to give clues as to what was spelled.</p>\n<p>In my early research in cursive handwriting recognition using hidden Markov models I used \"contexts\" (recognizing groups of two, three, or even four letters together as one model instead of individual letters). </p>\n<p><a href=\"https://ieeexplore.ieee.org/iel2/3104/8834/00389432.pdf?casa_token=EdbBFGdcXHAAAAAA:EAh_OnxFt-bB1GTrwy5GjSgCcZc-1zMsipbP0WqhAn6gbODtsh5Bmr_5gQWGvKymGz2J7qYLYw\" target=\"_blank\">On-line cursive handwriting recognition using speech recognition methods</a></p>\n<p>It helps with co-articulation, where the previous and the next letter affects the appearance of the current letter. It also helps where a letter itself might be hard to identify, but the letters around it make it clear what is meant (for those who are familiar with cursive, compare the wi in wick versus the ui in quick). Prior probabilities can help weight recognition towards more common groups of letters, which also automatically happens to some extent based on the amount of representation groups of letters have in a training set.</p>\n<p>Stochastic grammars further help to reduce the \"perplexity\" of the problem, as certain words are more likely to follow other words.  For example, the \"quick brown fox\" is more likely than \"fox brown quick.\"  There are large public text datasets that can aid in creating such grammars.</p>\n<p>Here is another example that shows the value of contexts and grammar:  NT is a more common combination of letters than TN when in the middle of a word.  Off-hand, I can not even think of a word with TN in the middle. However, for addresses, TN may appear often, as it is the abbreviation for a state. Seeing \"Memphis, \" may predict an upcoming \"TN\"  where \"MOVEME\" might foreshadow an upcoming \"NT\"</p>\n<p>Of course, given a big enough neural net model and enough training data, these details could be learned. However, for smaller models and smaller data sets, finding a way to apply contextual knowledge may be a more efficient way to improve accuracy.  Perhaps pre-training a neural net on text datasets might help? Or perhaps finding a way to apply contexts and stochastic grammars may be better?  Offhand, I do not know of any papers that compare these methods and whether, in the limit, they are equivalent. However, I bet adding context in some way can be powerful in helping resolve when a signer is talking about ANTMAN or SANTAS.</p>",
      "rawMarkdown": "Take a look at a fingerspelling chart:\n\n[https://www.startasl.com/fingerspelling/](https://www.startasl.com/fingerspelling/)\n\nspecifically, look at the letters A T N M S in order.  Note how the thumb starts outside the fist with A and then progresses to being under the first (T), second (N), and third finger(M).  With S it is in front of the hand.\n\nNow go to  [https://asl.ms/()/spellit_large.htm](https://asl.ms/()/spellit_large.htm) and enter \n\natnmsatnmsatnmsatms\n\nTurn the speed up.  How well  can you see the subtle differences? How well can hand posture trackers detect these subtle differences?\n\nExpert signers use context to help recognize which letter is being signed as the movement may be too fast, the lighting might be too poor, or the fingerspeller's dexterity might be limited.  We can do something similar by looking at transition between letters to give clues as to what was spelled.\n\nIn my early research in cursive handwriting recognition using hidden Markov models I used \"contexts\" (recognizing groups of two, three, or even four letters together as one model instead of individual letters). \n\n[On-line cursive handwriting recognition using speech recognition methods](https://ieeexplore.ieee.org/iel2/3104/8834/00389432.pdf?casa_token=EdbBFGdcXHAAAAAA:EAh_OnxFt-bB1GTrwy5GjSgCcZc-1zMsipbP0WqhAn6gbODtsh5Bmr_5gQWGvKymGz2J7qYLYw)\n\nIt helps with co-articulation, where the previous and the next letter affects the appearance of the current letter. It also helps where a letter itself might be hard to identify, but the letters around it make it clear what is meant (for those who are familiar with cursive, compare the wi in wick versus the ui in quick). Prior probabilities can help weight recognition towards more common groups of letters, which also automatically happens to some extent based on the amount of representation groups of letters have in a training set.\n\nStochastic grammars further help to reduce the \"perplexity\" of the problem, as certain words are more likely to follow other words.  For example, the \"quick brown fox\" is more likely than \"fox brown quick.\"  There are large public text datasets that can aid in creating such grammars.\n\nHere is another example that shows the value of contexts and grammar:  NT is a more common combination of letters than TN when in the middle of a word.  Off-hand, I can not even think of a word with TN in the middle. However, for addresses, TN may appear often, as it is the abbreviation for a state. Seeing \"Memphis, \" may predict an upcoming \"TN\"  where \"MOVEME\" might foreshadow an upcoming \"NT\"\n\nOf course, given a big enough neural net model and enough training data, these details could be learned. However, for smaller models and smaller data sets, finding a way to apply contextual knowledge may be a more efficient way to improve accuracy.  Perhaps pre-training a neural net on text datasets might help? Or perhaps finding a way to apply contexts and stochastic grammars may be better?  Offhand, I do not know of any papers that compare these methods and whether, in the limit, they are equivalent. However, I bet adding context in some way can be powerful in helping resolve when a signer is talking about ANTMAN or SANTAS.\n\n",
      "votes": 13
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2300297": "Take a look at a fingerspelling chart:\n\n[https://www.startasl.com/fingerspelling/](https://www.startasl.com/fingerspelling/)\n\nspecifically, look at the letters A T N M S in order.  Note how the thumb starts outside the fist with A and then progresses to being under the first (T), second (N), and third finger(M).  With S it is in front of the hand.\n\nNow go to  [https://asl.ms/()/spellit_large.htm](https://asl.ms/()/spellit_large.htm) and enter \n\natnmsatnmsatnmsatms\n\nTurn the speed up.  How well  can you see the subtle differences? How well can hand posture trackers detect these subtle differences?\n\nExpert signers use context to help recognize which letter is being signed as the movement may be too fast, the lighting might be too poor, or the fingerspeller's dexterity might be limited.  We can do something similar by looking at transition between letters to give clues as to what was spelled.\n\nIn my early research in cursive handwriting recognition using hidden Markov models I used \"contexts\" (recognizing groups of two, three, or even four letters together as one model instead of individual letters). \n\n[On-line cursive handwriting recognition using speech recognition methods](https://ieeexplore.ieee.org/iel2/3104/8834/00389432.pdf?casa_token=EdbBFGdcXHAAAAAA:EAh_OnxFt-bB1GTrwy5GjSgCcZc-1zMsipbP0WqhAn6gbODtsh5Bmr_5gQWGvKymGz2J7qYLYw)\n\nIt helps with co-articulation, where the previous and the next letter affects the appearance of the current letter. It also helps where a letter itself might be hard to identify, but the letters around it make it clear what is meant (for those who are familiar with cursive, compare the wi in wick versus the ui in quick). Prior probabilities can help weight recognition towards more common groups of letters, which also automatically happens to some extent based on the amount of representation groups of letters have in a training set.\n\nStochastic grammars further help to reduce the \"perplexity\" of the problem, as certain words are more likely to follow other words.  For example, the \"quick brown fox\" is more likely than \"fox brown quick.\"  There are large public text datasets that can aid in creating such grammars.\n\nHere is another example that shows the value of contexts and grammar:  NT is a more common combination of letters than TN when in the middle of a word.  Off-hand, I can not even think of a word with TN in the middle. However, for addresses, TN may appear often, as it is the abbreviation for a state. Seeing \"Memphis, \" may predict an upcoming \"TN\"  where \"MOVEME\" might foreshadow an upcoming \"NT\"\n\nOf course, given a big enough neural net model and enough training data, these details could be learned. However, for smaller models and smaller data sets, finding a way to apply contextual knowledge may be a more efficient way to improve accuracy.  Perhaps pre-training a neural net on text datasets might help? Or perhaps finding a way to apply contexts and stochastic grammars may be better?  Offhand, I do not know of any papers that compare these methods and whether, in the limit, they are equivalent. However, I bet adding context in some way can be powerful in helping resolve when a signer is talking about ANTMAN or SANTAS.\n\n"
  }
}