{
  "id": 435178,
  "title": "From Sign Language Collection to Clean Dataset for Machine Learning",
  "url": "/competitions/asl-fingerspelling/discussion/435178",
  "author_name": "Alexey Prikhodko",
  "post_date": "2023-08-28T08:56:11.309000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I have been working on a project that involves the entire process of collecting signs and curating a refined dataset for machine learning. Let's explore the workflow:</p>\n<ol>\n<li><p>Determining Words or Sentences:<br>\nInstead of relying on existing sound language corpora or regular text on the internet, which may not be ideal for sign language, I have been directly collecting signs from sign language signers. Since the structures of sign and sound languages differ, direct translation can be complex. However, for fingerspelling recognition, we can find any text to facilitate translation from text to signs.</p></li>\n<li><p>Replicating Signs:<br>\nIs it necessary to create multiple replicated signs from a single natural sign? While an actor can mimic the sign and speech, replicating exact signs is not always accurate. This suggests that the complexities and costs associated with replicating signs may not be worthwhile.</p></li>\n<li><p>Annotation of Signs:<br>\nShould we perform additional annotation when we already have text or words with corresponding signs from the previous step? The annotation process includes:</p>\n<p>3.1 Annotation for sign recognition detection tasks, which involves determining sign activity in a given video frame. Similar to voice activity detection (VAD) in spoken languages, precise boundary localization between non-gestural and gestural movements is crucial for a clean dataset.</p>\n<p>3.2 Annotation for sign language segmentation tasks, which involves identifying frame boundaries between signs or phrases in a video to separate them into meaningful units. In the competition, there is the possibility of recognizing dactyls without segmentation using the CTC loss function is wow. But what happens when dealing with very long phrases (sentences) consisting of approximately 100-200 letters (symbols)? Do CTC loss functions or other functions still ensure accurate recognition?</p>\n<p>3.3 Annotation for sign parameter recognition tasks, which involves recognizing the phonological features of a sign. Simply recognizing a sign and outputting a word directly is not straightforward, as one word may have different signs. This also depends on the linguistic nuances of sign language.</p>\n<p>3.4 Annotation for machine translation tasks, which involves translating signs to text and vice versa.</p></li>\n<li><p>Dataset Formation:<br>\nWork with a database to organize the collected data.</p></li>\n</ol>\n<p>I believe that the time spent on dataset preparation is more significant than model building. I am eager to hear from you:</p>\n<ul>\n<li>what other important annotations are essential?</li>\n<li>Are there any annotations that may not be necessary in the future?</li>\n<li>I am interested in learning about what different methods for collecting sign language data?</li>\n</ul>",
  "messages": [
    {
      "id": 2412406,
      "postDate": "2023-08-28T08:56:11.310Z",
      "content": "<p>I have been working on a project that involves the entire process of collecting signs and curating a refined dataset for machine learning. Let's explore the workflow:</p>\n<ol>\n<li><p>Determining Words or Sentences:<br>\nInstead of relying on existing sound language corpora or regular text on the internet, which may not be ideal for sign language, I have been directly collecting signs from sign language signers. Since the structures of sign and sound languages differ, direct translation can be complex. However, for fingerspelling recognition, we can find any text to facilitate translation from text to signs.</p></li>\n<li><p>Replicating Signs:<br>\nIs it necessary to create multiple replicated signs from a single natural sign? While an actor can mimic the sign and speech, replicating exact signs is not always accurate. This suggests that the complexities and costs associated with replicating signs may not be worthwhile.</p></li>\n<li><p>Annotation of Signs:<br>\nShould we perform additional annotation when we already have text or words with corresponding signs from the previous step? The annotation process includes:</p>\n<p>3.1 Annotation for sign recognition detection tasks, which involves determining sign activity in a given video frame. Similar to voice activity detection (VAD) in spoken languages, precise boundary localization between non-gestural and gestural movements is crucial for a clean dataset.</p>\n<p>3.2 Annotation for sign language segmentation tasks, which involves identifying frame boundaries between signs or phrases in a video to separate them into meaningful units. In the competition, there is the possibility of recognizing dactyls without segmentation using the CTC loss function is wow. But what happens when dealing with very long phrases (sentences) consisting of approximately 100-200 letters (symbols)? Do CTC loss functions or other functions still ensure accurate recognition?</p>\n<p>3.3 Annotation for sign parameter recognition tasks, which involves recognizing the phonological features of a sign. Simply recognizing a sign and outputting a word directly is not straightforward, as one word may have different signs. This also depends on the linguistic nuances of sign language.</p>\n<p>3.4 Annotation for machine translation tasks, which involves translating signs to text and vice versa.</p></li>\n<li><p>Dataset Formation:<br>\nWork with a database to organize the collected data.</p></li>\n</ol>\n<p>I believe that the time spent on dataset preparation is more significant than model building. I am eager to hear from you:</p>\n<ul>\n<li>what other important annotations are essential?</li>\n<li>Are there any annotations that may not be necessary in the future?</li>\n<li>I am interested in learning about what different methods for collecting sign language data?</li>\n</ul>",
      "rawMarkdown": "I have been working on a project that involves the entire process of collecting signs and curating a refined dataset for machine learning. Let's explore the workflow:\n\n1. Determining Words or Sentences:\nInstead of relying on existing sound language corpora or regular text on the internet, which may not be ideal for sign language, I have been directly collecting signs from sign language signers. Since the structures of sign and sound languages differ, direct translation can be complex. However, for fingerspelling recognition, we can find any text to facilitate translation from text to signs.\n\n2. Replicating Signs:\nIs it necessary to create multiple replicated signs from a single natural sign? While an actor can mimic the sign and speech, replicating exact signs is not always accurate. This suggests that the complexities and costs associated with replicating signs may not be worthwhile.\n\n3. Annotation of Signs:\nShould we perform additional annotation when we already have text or words with corresponding signs from the previous step? The annotation process includes:\n\n     3.1 Annotation for sign recognition detection tasks, which involves determining sign activity in a given video frame. Similar to voice activity detection (VAD) in spoken languages, precise boundary localization between non-gestural and gestural movements is crucial for a clean dataset.\n\n     3.2 Annotation for sign language segmentation tasks, which involves identifying frame boundaries between signs or phrases in a video to separate them into meaningful units. In the competition, there is the possibility of recognizing dactyls without segmentation using the CTC loss function is wow. But what happens when dealing with very long phrases (sentences) consisting of approximately 100-200 letters (symbols)? Do CTC loss functions or other functions still ensure accurate recognition?\n\n     3.3 Annotation for sign parameter recognition tasks, which involves recognizing the phonological features of a sign. Simply recognizing a sign and outputting a word directly is not straightforward, as one word may have different signs. This also depends on the linguistic nuances of sign language.\n\n     3.4 Annotation for machine translation tasks, which involves translating signs to text and vice versa.\n\n4. Dataset Formation:\nWork with a database to organize the collected data.\n\nI believe that the time spent on dataset preparation is more significant than model building. I am eager to hear from you:\n- what other important annotations are essential?\n- Are there any annotations that may not be necessary in the future?\n- I am interested in learning about what different methods for collecting sign language data?",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2412406": "I have been working on a project that involves the entire process of collecting signs and curating a refined dataset for machine learning. Let's explore the workflow:\n\n1. Determining Words or Sentences:\nInstead of relying on existing sound language corpora or regular text on the internet, which may not be ideal for sign language, I have been directly collecting signs from sign language signers. Since the structures of sign and sound languages differ, direct translation can be complex. However, for fingerspelling recognition, we can find any text to facilitate translation from text to signs.\n\n2. Replicating Signs:\nIs it necessary to create multiple replicated signs from a single natural sign? While an actor can mimic the sign and speech, replicating exact signs is not always accurate. This suggests that the complexities and costs associated with replicating signs may not be worthwhile.\n\n3. Annotation of Signs:\nShould we perform additional annotation when we already have text or words with corresponding signs from the previous step? The annotation process includes:\n\n     3.1 Annotation for sign recognition detection tasks, which involves determining sign activity in a given video frame. Similar to voice activity detection (VAD) in spoken languages, precise boundary localization between non-gestural and gestural movements is crucial for a clean dataset.\n\n     3.2 Annotation for sign language segmentation tasks, which involves identifying frame boundaries between signs or phrases in a video to separate them into meaningful units. In the competition, there is the possibility of recognizing dactyls without segmentation using the CTC loss function is wow. But what happens when dealing with very long phrases (sentences) consisting of approximately 100-200 letters (symbols)? Do CTC loss functions or other functions still ensure accurate recognition?\n\n     3.3 Annotation for sign parameter recognition tasks, which involves recognizing the phonological features of a sign. Simply recognizing a sign and outputting a word directly is not straightforward, as one word may have different signs. This also depends on the linguistic nuances of sign language.\n\n     3.4 Annotation for machine translation tasks, which involves translating signs to text and vice versa.\n\n4. Dataset Formation:\nWork with a database to organize the collected data.\n\nI believe that the time spent on dataset preparation is more significant than model building. I am eager to hear from you:\n- what other important annotations are essential?\n- Are there any annotations that may not be necessary in the future?\n- I am interested in learning about what different methods for collecting sign language data?"
  }
}