{
  "id": 434350,
  "title": "Those who are interested in learning, this is my suggestion for the next step",
  "url": "/competitions/asl-fingerspelling/discussion/434350",
  "author_name": "hengck23",
  "post_date": "2023-08-24T23:39:03.842000",
  "votes": 18,
  "comment_count": 1,
  "views": 0,
  "content": "<p>both speech recognition and hand sign recognition recogntion share almost the same technology.</p>\n<p>one is \"show and spell\", the other is  \"listen and spell\".</p>\n<p>for this finger spelling, target is tflite. so you learn how to construct the model in the most compact way. also you need to rewrite in tf and you are limited to \"less functions and methods\" when tflite does not support</p>\n<p>in the other current kaggle bengali speech recognition, it would be more of building a pipline, that include ASR, LM language model decoding (for word spell check), puncuation model, language normalisation … You will be using existing toolkit to construct large scale learning. e.g. NeMO+pytorch lightning, Hugging face, FairSeq (for wav2Vec). There will be more complicted post-processing like beam search, etc …. There ill be external data, … web scrapping …</p>\n<p>So I suggest if you are interested in this fingerspell, you can process to the kaggle bengali speech recognition, where you can continue to learn about CTC, RNNT, AED  and etc …. Learn about the STOA common large models like conformer, wav2vec, whisper … these are being used in meta, google microsoft ASR services</p>",
  "messages": [
    {
      "id": 2407190,
      "postDate": "2023-08-24T23:39:03.843Z",
      "content": "<p>both speech recognition and hand sign recognition recogntion share almost the same technology.</p>\n<p>one is \"show and spell\", the other is  \"listen and spell\".</p>\n<p>for this finger spelling, target is tflite. so you learn how to construct the model in the most compact way. also you need to rewrite in tf and you are limited to \"less functions and methods\" when tflite does not support</p>\n<p>in the other current kaggle bengali speech recognition, it would be more of building a pipline, that include ASR, LM language model decoding (for word spell check), puncuation model, language normalisation … You will be using existing toolkit to construct large scale learning. e.g. NeMO+pytorch lightning, Hugging face, FairSeq (for wav2Vec). There will be more complicted post-processing like beam search, etc …. There ill be external data, … web scrapping …</p>\n<p>So I suggest if you are interested in this fingerspell, you can process to the kaggle bengali speech recognition, where you can continue to learn about CTC, RNNT, AED  and etc …. Learn about the STOA common large models like conformer, wav2vec, whisper … these are being used in meta, google microsoft ASR services</p>",
      "rawMarkdown": "both speech recognition and hand sign recognition recogntion share almost the same technology.\n\none is \"show and spell\", the other is  \"listen and spell\".\n\nfor this finger spelling, target is tflite. so you learn how to construct the model in the most compact way. also you need to rewrite in tf and you are limited to \"less functions and methods\" when tflite does not support\n\nin the other current kaggle bengali speech recognition, it would be more of building a pipline, that include ASR, LM language model decoding (for word spell check), puncuation model, language normalisation ... You will be using existing toolkit to construct large scale learning. e.g. NeMO+pytorch lightning, Hugging face, FairSeq (for wav2Vec). There will be more complicted post-processing like beam search, etc .... There ill be external data, ... web scrapping ...\n\nSo I suggest if you are interested in this fingerspell, you can process to the kaggle bengali speech recognition, where you can continue to learn about CTC, RNNT, AED  and etc .... Learn about the STOA common large models like conformer, wav2vec, whisper ... these are being used in meta, google microsoft ASR services\n\n",
      "votes": 17
    },
    {
      "id": 2408339,
      "postDate": "2023-08-25T15:24:40.327Z",
      "content": "<p>Thanks for sharing your  valuable suggestions.<br>\nI tried using multiprocessing as no of cpus detected while running kaggle notebook was four. Can you comment on whether tflite provides multiprocessing support. </p>",
      "rawMarkdown": "Thanks for sharing your  valuable suggestions.\nI tried using multiprocessing as no of cpus detected while running kaggle notebook was four. Can you comment on whether tflite provides multiprocessing support. "
    }
  ],
  "comments": [
    {
      "id": 2408339,
      "author_name": "C R Suthikshn Kumar",
      "author_url": "",
      "post_date": "2023-08-25T15:24:40.327000",
      "content": "<p>Thanks for sharing your  valuable suggestions.<br>\nI tried using multiprocessing as no of cpus detected while running kaggle notebook was four. Can you comment on whether tflite provides multiprocessing support. </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2407190": "both speech recognition and hand sign recognition recogntion share almost the same technology.\n\none is \"show and spell\", the other is  \"listen and spell\".\n\nfor this finger spelling, target is tflite. so you learn how to construct the model in the most compact way. also you need to rewrite in tf and you are limited to \"less functions and methods\" when tflite does not support\n\nin the other current kaggle bengali speech recognition, it would be more of building a pipline, that include ASR, LM language model decoding (for word spell check), puncuation model, language normalisation ... You will be using existing toolkit to construct large scale learning. e.g. NeMO+pytorch lightning, Hugging face, FairSeq (for wav2Vec). There will be more complicted post-processing like beam search, etc .... There ill be external data, ... web scrapping ...\n\nSo I suggest if you are interested in this fingerspell, you can process to the kaggle bengali speech recognition, where you can continue to learn about CTC, RNNT, AED  and etc .... Learn about the STOA common large models like conformer, wav2vec, whisper ... these are being used in meta, google microsoft ASR services\n\n",
    "2408339": "Thanks for sharing your  valuable suggestions.\nI tried using multiprocessing as no of cpus detected while running kaggle notebook was four. Can you comment on whether tflite provides multiprocessing support. "
  }
}