{
  "id": 436873,
  "title": "2nd place code with reproducibility",
  "url": "/competitions/asl-fingerspelling/discussion/436873",
  "author_name": "hoyso48",
  "post_date": "2023-09-04T14:07:37.077000",
  "votes": 11,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I have uploaded the code for training, inference, and creating tfrecords.</p>\n<p>Training code:<br>\n<a href=\"https://www.kaggle.com/code/hoyso48/2nd-place-solution-training\" target=\"_blank\">https://www.kaggle.com/code/hoyso48/2nd-place-solution-training</a></p>\n<p>Inference code:<br>\n<a href=\"https://www.kaggle.com/code/hoyso48/2nd-place-solution-inference\" target=\"_blank\">https://www.kaggle.com/code/hoyso48/2nd-place-solution-inference</a></p>\n<p>TFRecords creation code:<br>\n<a href=\"https://www.kaggle.com/code/hoyso48/aslfr-create-tfr\" target=\"_blank\">https://www.kaggle.com/code/hoyso48/aslfr-create-tfr</a></p>\n<p>The original code and model files I used on Colab have been uploaded to GitHub:<br>\n<a href=\"https://github.com/hoyso48/Google---American-Sign-Language-Fingerspelling-Recognition-2nd-place-solution\" target=\"_blank\">https://github.com/hoyso48/Google---American-Sign-Language-Fingerspelling-Recognition-2nd-place-solution</a></p>\n<p>If you wish to fully replicate the solution, due to the 20-hour weekly runtime limit of Kaggle TPU, I recommend using a Colab notebook.</p>\n<p>Below is a rough set of execution metrics for the model in a Kaggle CPU environment. It's best to use throughput and latency metrics as references only, given their significant variance and potential inaccuracies.</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>troughput (iterations/s)</th>\n<th>latency (ms)</th>\n<th>model size (Mb)</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>CTCGreedy</td>\n<td>20.14</td>\n<td>49.66</td>\n<td>11.41</td>\n<td>0.815</td>\n<td>0.807</td>\n</tr>\n<tr>\n<td>ATTGreedy</td>\n<td>10.16</td>\n<td>98.42</td>\n<td>11.77</td>\n<td>0.816</td>\n<td>0.808</td>\n</tr>\n<tr>\n<td>CTCATTJointGreedy</td>\n<td>5.26</td>\n<td>190.22</td>\n<td>12.49</td>\n<td>0.820</td>\n<td>0.812</td>\n</tr>\n<tr>\n<td>CTCATTJointGreedy -2xseed</td>\n<td>2.71</td>\n<td>368.95</td>\n<td>24.90</td>\n<td>0.825</td>\n<td>0.817</td>\n</tr>\n<tr>\n<td>CTCATTJointGreedy -3xseed</td>\n<td>1.75</td>\n<td>570.77</td>\n<td>37.30</td>\n<td>0.825</td>\n<td>0.819</td>\n</tr>\n</tbody>\n</table>\n<p>One clear takeaway from the table is that it seems advisable to use the CTCGreedy decoder for practical efficiency purposes.</p>\n<p>I believe my code will be helpful to many who started studying deep learning through TensorFlow&amp;Kaggle. If you have any questions, please leave them in the comments.</p>",
  "messages": [
    {
      "id": 2423214,
      "postDate": "2023-09-04T14:07:37.077Z",
      "content": "<p>I have uploaded the code for training, inference, and creating tfrecords.</p>\n<p>Training code:<br>\n<a href=\"https://www.kaggle.com/code/hoyso48/2nd-place-solution-training\" target=\"_blank\">https://www.kaggle.com/code/hoyso48/2nd-place-solution-training</a></p>\n<p>Inference code:<br>\n<a href=\"https://www.kaggle.com/code/hoyso48/2nd-place-solution-inference\" target=\"_blank\">https://www.kaggle.com/code/hoyso48/2nd-place-solution-inference</a></p>\n<p>TFRecords creation code:<br>\n<a href=\"https://www.kaggle.com/code/hoyso48/aslfr-create-tfr\" target=\"_blank\">https://www.kaggle.com/code/hoyso48/aslfr-create-tfr</a></p>\n<p>The original code and model files I used on Colab have been uploaded to GitHub:<br>\n<a href=\"https://github.com/hoyso48/Google---American-Sign-Language-Fingerspelling-Recognition-2nd-place-solution\" target=\"_blank\">https://github.com/hoyso48/Google---American-Sign-Language-Fingerspelling-Recognition-2nd-place-solution</a></p>\n<p>If you wish to fully replicate the solution, due to the 20-hour weekly runtime limit of Kaggle TPU, I recommend using a Colab notebook.</p>\n<p>Below is a rough set of execution metrics for the model in a Kaggle CPU environment. It's best to use throughput and latency metrics as references only, given their significant variance and potential inaccuracies.</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>troughput (iterations/s)</th>\n<th>latency (ms)</th>\n<th>model size (Mb)</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>CTCGreedy</td>\n<td>20.14</td>\n<td>49.66</td>\n<td>11.41</td>\n<td>0.815</td>\n<td>0.807</td>\n</tr>\n<tr>\n<td>ATTGreedy</td>\n<td>10.16</td>\n<td>98.42</td>\n<td>11.77</td>\n<td>0.816</td>\n<td>0.808</td>\n</tr>\n<tr>\n<td>CTCATTJointGreedy</td>\n<td>5.26</td>\n<td>190.22</td>\n<td>12.49</td>\n<td>0.820</td>\n<td>0.812</td>\n</tr>\n<tr>\n<td>CTCATTJointGreedy -2xseed</td>\n<td>2.71</td>\n<td>368.95</td>\n<td>24.90</td>\n<td>0.825</td>\n<td>0.817</td>\n</tr>\n<tr>\n<td>CTCATTJointGreedy -3xseed</td>\n<td>1.75</td>\n<td>570.77</td>\n<td>37.30</td>\n<td>0.825</td>\n<td>0.819</td>\n</tr>\n</tbody>\n</table>\n<p>One clear takeaway from the table is that it seems advisable to use the CTCGreedy decoder for practical efficiency purposes.</p>\n<p>I believe my code will be helpful to many who started studying deep learning through TensorFlow&amp;Kaggle. If you have any questions, please leave them in the comments.</p>",
      "rawMarkdown": "I have uploaded the code for training, inference, and creating tfrecords.\n\nTraining code:\nhttps://www.kaggle.com/code/hoyso48/2nd-place-solution-training\n\nInference code:\nhttps://www.kaggle.com/code/hoyso48/2nd-place-solution-inference\n\nTFRecords creation code:\nhttps://www.kaggle.com/code/hoyso48/aslfr-create-tfr\n\nThe original code and model files I used on Colab have been uploaded to GitHub:\nhttps://github.com/hoyso48/Google---American-Sign-Language-Fingerspelling-Recognition-2nd-place-solution\n\nIf you wish to fully replicate the solution, due to the 20-hour weekly runtime limit of Kaggle TPU, I recommend using a Colab notebook.\n\nBelow is a rough set of execution metrics for the model in a Kaggle CPU environment. It's best to use throughput and latency metrics as references only, given their significant variance and potential inaccuracies.\n\n| | troughput (iterations/s) | latency (ms) | model size (Mb)| Public LB | Private LB |\n| --- | --- | --- | --- | --- | --- |\n| CTCGreedy | 20.14 | 49.66 | 11.41 | 0.815 | 0.807 |\n| ATTGreedy | 10.16 | 98.42 | 11.77 | 0.816 | 0.808 |\n| CTCATTJointGreedy | 5.26 | 190.22 | 12.49 | 0.820 | 0.812 |\n| CTCATTJointGreedy -2xseed | 2.71 | 368.95 | 24.90 | 0.825 | 0.817 |\n| CTCATTJointGreedy -3xseed | 1.75 | 570.77 | 37.30 | 0.825 | 0.819 |\n\nOne clear takeaway from the table is that it seems advisable to use the CTCGreedy decoder for practical efficiency purposes.\n\nI believe my code will be helpful to many who started studying deep learning through TensorFlow&Kaggle. If you have any questions, please leave them in the comments.\n\n\n",
      "votes": 11
    },
    {
      "id": 2424575,
      "postDate": "2023-09-05T11:33:03.603Z",
      "content": "<p>thanks for sharing, always learn a lot from your code</p>",
      "rawMarkdown": "thanks for sharing, always learn a lot from your code",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 2424575,
      "author_name": "Yu Wu",
      "author_url": "",
      "post_date": "2023-09-05T11:33:03.603000",
      "content": "<p>thanks for sharing, always learn a lot from your code</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2423214": "I have uploaded the code for training, inference, and creating tfrecords.\n\nTraining code:\nhttps://www.kaggle.com/code/hoyso48/2nd-place-solution-training\n\nInference code:\nhttps://www.kaggle.com/code/hoyso48/2nd-place-solution-inference\n\nTFRecords creation code:\nhttps://www.kaggle.com/code/hoyso48/aslfr-create-tfr\n\nThe original code and model files I used on Colab have been uploaded to GitHub:\nhttps://github.com/hoyso48/Google---American-Sign-Language-Fingerspelling-Recognition-2nd-place-solution\n\nIf you wish to fully replicate the solution, due to the 20-hour weekly runtime limit of Kaggle TPU, I recommend using a Colab notebook.\n\nBelow is a rough set of execution metrics for the model in a Kaggle CPU environment. It's best to use throughput and latency metrics as references only, given their significant variance and potential inaccuracies.\n\n| | troughput (iterations/s) | latency (ms) | model size (Mb)| Public LB | Private LB |\n| --- | --- | --- | --- | --- | --- |\n| CTCGreedy | 20.14 | 49.66 | 11.41 | 0.815 | 0.807 |\n| ATTGreedy | 10.16 | 98.42 | 11.77 | 0.816 | 0.808 |\n| CTCATTJointGreedy | 5.26 | 190.22 | 12.49 | 0.820 | 0.812 |\n| CTCATTJointGreedy -2xseed | 2.71 | 368.95 | 24.90 | 0.825 | 0.817 |\n| CTCATTJointGreedy -3xseed | 1.75 | 570.77 | 37.30 | 0.825 | 0.819 |\n\nOne clear takeaway from the table is that it seems advisable to use the CTCGreedy decoder for practical efficiency purposes.\n\nI believe my code will be helpful to many who started studying deep learning through TensorFlow&Kaggle. If you have any questions, please leave them in the comments.\n\n\n",
    "2424575": "thanks for sharing, always learn a lot from your code"
  }
}