{
  "id": 410554,
  "title": "How does American Sign Language Competition differs from previous Isolated sign Language competition?",
  "url": "/competitions/asl-fingerspelling/discussion/410554",
  "author_name": "Anshuman Mishra",
  "post_date": "2023-05-15T18:33:19",
  "votes": 5,
  "comment_count": 5,
  "views": 0,
  "content": "<p>This question is intended for kagglers who participated in previous <a href=\"https://www.kaggle.com/competitions/asl-signs\" target=\"_blank\">https://www.kaggle.com/competitions/asl-signs</a> competition. I have few questions regarding this</p>\n<ol>\n<li>How does two competitions differ in terms of techniques to be used, and datasets distributions?</li>\n<li>What are some do's and don't to be kept in mind for beginners in this type of competitions?</li>\n<li>What maybe the reason of \"bar is high\" <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/410459#2260528\" target=\"_blank\">https://www.kaggle.com/competitions/asl-fingerspelling/discussion/410459#2260528</a> , <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a></li>\n</ol>",
  "messages": [
    {
      "id": 2260591,
      "postDate": "2023-05-15T18:33:19Z",
      "content": "<p>This question is intended for kagglers who participated in previous <a href=\"https://www.kaggle.com/competitions/asl-signs\" target=\"_blank\">https://www.kaggle.com/competitions/asl-signs</a> competition. I have few questions regarding this</p>\n<ol>\n<li>How does two competitions differ in terms of techniques to be used, and datasets distributions?</li>\n<li>What are some do's and don't to be kept in mind for beginners in this type of competitions?</li>\n<li>What maybe the reason of \"bar is high\" <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/410459#2260528\" target=\"_blank\">https://www.kaggle.com/competitions/asl-fingerspelling/discussion/410459#2260528</a> , <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a></li>\n</ol>",
      "rawMarkdown": "This question is intended for kagglers who participated in previous https://www.kaggle.com/competitions/asl-signs competition. I have few questions regarding this\n\n1. How does two competitions differ in terms of techniques to be used, and datasets distributions?\n2. What are some do's and don't to be kept in mind for beginners in this type of competitions?\n3. What maybe the reason of \"bar is high\" https://www.kaggle.com/competitions/asl-fingerspelling/discussion/410459#2260528 , @sohier",
      "votes": 5
    },
    {
      "id": 2260655,
      "postDate": "2023-05-15T19:23:31.387Z",
      "content": "<ol>\n<li>The previous competition was a classification problem: you get one video in the form of a sequence of poses/landmark vectors and you predict a single label. Here, we are dealing with continuous fingerspelling. Every video contains a sequence of multiple letters, numbers, and other symbols that are spelled one after another and we need to predict all of them. You could compare it to speech recognition where you have multiple phonemes to recognize in order per example.</li>\n<li>The same as for any other Kaggle competition I'd say 😁 Though the inference time and memory constraints make it a bit more interesting.</li>\n<li>In many other competitions you can submit a file containing predictions for every test set example. So it is easier to create baseline submissions like uniform probability or majority class prediction. Here, you actually need to train a model, convert it to TFLite, and submit that model. Then, the model needs to finish within the constraints. So this raises the bar for the initial submissions.</li>\n</ol>",
      "rawMarkdown": "1. The previous competition was a classification problem: you get one video in the form of a sequence of poses/landmark vectors and you predict a single label. Here, we are dealing with continuous fingerspelling. Every video contains a sequence of multiple letters, numbers, and other symbols that are spelled one after another and we need to predict all of them. You could compare it to speech recognition where you have multiple phonemes to recognize in order per example.\n2. The same as for any other Kaggle competition I'd say 😁 Though the inference time and memory constraints make it a bit more interesting.\n3. In many other competitions you can submit a file containing predictions for every test set example. So it is easier to create baseline submissions like uniform probability or majority class prediction. Here, you actually need to train a model, convert it to TFLite, and submit that model. Then, the model needs to finish within the constraints. So this raises the bar for the initial submissions.",
      "votes": 6,
      "replies": [
        {
          "id": 2262284,
          "postDate": "2023-05-16T19:59:30.997Z",
          "content": "<p>Thanks, that settles it!</p>",
          "rawMarkdown": "Thanks, that settles it!"
        }
      ]
    },
    {
      "id": 2261756,
      "postDate": "2023-05-16T14:23:58.983Z",
      "content": "<p>Another aspect is the HUGE dataset size, which sets some challenges w.r.t. storage, efficient data loading and training. </p>",
      "rawMarkdown": "Another aspect is the HUGE dataset size, which sets some challenges w.r.t. storage, efficient data loading and training. ",
      "votes": 1
    },
    {
      "id": 2260653,
      "postDate": "2023-05-15T19:22:51.763Z",
      "content": "<p>I am going to answer the first question (title). </p>\n<p>In the previous competition, participants had to build a model, which process the given \"videos\" (actually, frames with \"tabular format of data\") and classify them with the classes (words). In this competition, we should build a model, which is fed by videos and outputs the phrase, what video's participant \"said\", thus we are getting seseq2seq task (input: sequence of frames, output: sequence of tokens). </p>\n<p>Hope this helps. </p>",
      "rawMarkdown": "I am going to answer the first question (title). \n\nIn the previous competition, participants had to build a model, which process the given \"videos\" (actually, frames with \"tabular format of data\") and classify them with the classes (words). In this competition, we should build a model, which is fed by videos and outputs the phrase, what video's participant \"said\", thus we are getting seseq2seq task (input: sequence of frames, output: sequence of tokens). \n\nHope this helps. ",
      "votes": 2
    },
    {
      "id": 2262389,
      "postDate": "2023-05-16T22:06:13.767Z",
      "content": "<p>Also this from data page </p>\n<blockquote>\n  <p>The landmark files contain the same data as in the ASL Signs competition (minus the row ID column) but reshaped into a wide format. This allows you to take advantage of the Parquet format to entirely skip loading landmarks that you aren't using.</p>\n</blockquote>",
      "rawMarkdown": "Also this from data page \n\n>The landmark files contain the same data as in the ASL Signs competition (minus the row ID column) but reshaped into a wide format. This allows you to take advantage of the Parquet format to entirely skip loading landmarks that you aren't using.\n\n"
    }
  ],
  "comments": [
    {
      "id": 2260655,
      "author_name": "Mathieu De Coster",
      "author_url": "",
      "post_date": "2023-05-15T19:23:31.387000",
      "content": "<ol>\n<li>The previous competition was a classification problem: you get one video in the form of a sequence of poses/landmark vectors and you predict a single label. Here, we are dealing with continuous fingerspelling. Every video contains a sequence of multiple letters, numbers, and other symbols that are spelled one after another and we need to predict all of them. You could compare it to speech recognition where you have multiple phonemes to recognize in order per example.</li>\n<li>The same as for any other Kaggle competition I'd say 😁 Though the inference time and memory constraints make it a bit more interesting.</li>\n<li>In many other competitions you can submit a file containing predictions for every test set example. So it is easier to create baseline submissions like uniform probability or majority class prediction. Here, you actually need to train a model, convert it to TFLite, and submit that model. Then, the model needs to finish within the constraints. So this raises the bar for the initial submissions.</li>\n</ol>",
      "votes": 6,
      "replies": [
        {
          "id": 2262284,
          "author_name": "Anshuman Mishra",
          "author_url": "",
          "post_date": "2023-05-16T19:59:30.997000",
          "content": "<p>Thanks, that settles it!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2261756,
      "author_name": "Wondering Alice",
      "author_url": "",
      "post_date": "2023-05-16T14:23:58.983000",
      "content": "<p>Another aspect is the HUGE dataset size, which sets some challenges w.r.t. storage, efficient data loading and training. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2260653,
      "author_name": "Vadim Irtlach",
      "author_url": "",
      "post_date": "2023-05-15T19:22:51.763000",
      "content": "<p>I am going to answer the first question (title). </p>\n<p>In the previous competition, participants had to build a model, which process the given \"videos\" (actually, frames with \"tabular format of data\") and classify them with the classes (words). In this competition, we should build a model, which is fed by videos and outputs the phrase, what video's participant \"said\", thus we are getting seseq2seq task (input: sequence of frames, output: sequence of tokens). </p>\n<p>Hope this helps. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2262389,
      "author_name": "RB",
      "author_url": "",
      "post_date": "2023-05-16T22:06:13.767000",
      "content": "<p>Also this from data page </p>\n<blockquote>\n  <p>The landmark files contain the same data as in the ASL Signs competition (minus the row ID column) but reshaped into a wide format. This allows you to take advantage of the Parquet format to entirely skip loading landmarks that you aren't using.</p>\n</blockquote>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2260591": "This question is intended for kagglers who participated in previous https://www.kaggle.com/competitions/asl-signs competition. I have few questions regarding this\n\n1. How does two competitions differ in terms of techniques to be used, and datasets distributions?\n2. What are some do's and don't to be kept in mind for beginners in this type of competitions?\n3. What maybe the reason of \"bar is high\" https://www.kaggle.com/competitions/asl-fingerspelling/discussion/410459#2260528 , @sohier",
    "2260655": "1. The previous competition was a classification problem: you get one video in the form of a sequence of poses/landmark vectors and you predict a single label. Here, we are dealing with continuous fingerspelling. Every video contains a sequence of multiple letters, numbers, and other symbols that are spelled one after another and we need to predict all of them. You could compare it to speech recognition where you have multiple phonemes to recognize in order per example.\n2. The same as for any other Kaggle competition I'd say 😁 Though the inference time and memory constraints make it a bit more interesting.\n3. In many other competitions you can submit a file containing predictions for every test set example. So it is easier to create baseline submissions like uniform probability or majority class prediction. Here, you actually need to train a model, convert it to TFLite, and submit that model. Then, the model needs to finish within the constraints. So this raises the bar for the initial submissions.",
    "2261756": "Another aspect is the HUGE dataset size, which sets some challenges w.r.t. storage, efficient data loading and training. ",
    "2260653": "I am going to answer the first question (title). \n\nIn the previous competition, participants had to build a model, which process the given \"videos\" (actually, frames with \"tabular format of data\") and classify them with the classes (words). In this competition, we should build a model, which is fed by videos and outputs the phrase, what video's participant \"said\", thus we are getting seseq2seq task (input: sequence of frames, output: sequence of tokens). \n\nHope this helps. ",
    "2262389": "Also this from data page \n\n>The landmark files contain the same data as in the ASL Signs competition (minus the row ID column) but reshaped into a wide format. This allows you to take advantage of the Parquet format to entirely skip loading landmarks that you aren't using.\n\n"
  }
}