{
  "id": 436457,
  "title": "[12th solution] Full, reproducible solution with code",
  "url": "/competitions/asl-fingerspelling/discussion/436457",
  "author_name": "greySnow",
  "post_date": "2023-09-02T13:16:46.852000",
  "votes": 5,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I already wrote the gist of my solution <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/434363\" target=\"_blank\">here</a>. Here, I will give a more comprehensive summary that follows official writeup guidelines. I also finished cleaning and arranging my work <a href=\"https://github.com/shlomoron/Google---American-Sign-Language-Fingerspelling-Recognition-12th-place-solution\" target=\"_blank\">in my github</a>. I made efforts to make my solution as easy to follow and reproduce as possible. I saved everything as Colab or Kaggle notebooks that you should be able to run as is. I kept all the data in public Kaggle datasets. Please let me know if I missed anything; I will do my best to fix it. If anything needs to be clarified, ask me. </p>\n<h3>Context section</h3>\n<p>This is a summary of my 12th place solution to <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/overview\" target=\"_blank\">Google - American Sign Language Fingerspelling Recognition competition</a>.<br>\nThe data page is <a href=\"https://wwww.kaggle.com/competitions/asl-fingerspelling/data\" target=\"_blank\">here</a>.</p>\n<h3>Overview of the Approach</h3>\n<p>My model is a CTC encoder with transformers and convolution layers, the same as the one used in the previous competition <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406978\" target=\"_blank\">1st solution</a> with the CTC modification that was introduced <a href=\"https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place\" target=\"_blank\">by Rohith</a>. My only modification to the model (besides changing the size) is adding positional encoding right before the first transformer layer.<br>\nFor the features, I used a similar approach as the solution of the previous competition and took the same lips, nose, and hands landmarks. I did not use the eyes and used slightly different pose landmarks (added some). I used X and Y (without Z) and also used difference features with a skip of 1 and 2, also the same as the previous competition solution.<br>\nFor augmentation, I used the same augmentations as the previous competition solution, with a modification to all augmentations based on modifying/deleting data in a window. The original solution used a random window, but a random window give a lower probability of modifications in the edges. I changed it to a random circular/rolling window, i.e., instead of ignoring the window section that goes out of the frames length, I rolled it to the beginning of the sequencs and modified it too.<br>\nFor filtering the bad data, I used two-step filtering. In the first step, I filtered the data according to (number of frames with non-nan hands landmarks) &gt; 2*(length of phrase) and trained a basic model. In the second step, I used the basic model to calculate the normalized Levenshtein distance scores of all the samples and filtered them according to the score with a threshold of &gt;0.2. This filtering allowed me to add to my training a lot of samples that I filtered out in the first step.<br>\nA 9,497,577 parameters model with the previous competition solution's augmentations trained on ~300 epochs got a public LB of 0.779. making it larger (15,471,112 parameters) and training for 500 epochs got a public LB score of 7.9. Adding the modifications to augmentations and second-step filtering and training for 1500 epochs got public LB  0.794 and private LB 0.779 (the final, 12th-place solution).<br>\nFor validation, I used the first 3,000 samples. Toward the end of the competition, I feared that I might be overfitting to existing signers, so I did one experiment with 1000 epochs and separated the validation fold by signers ID (I used five unique signers for validation with training on all the rest). Even after 1000 epochs, the scores on the validation fold did not suffer due to overfitting. Thus, I continued with the original validation since I preferred to train on samples from all the signers.<br>\nI did most of my training on Colab TPU. They are cheap and easy to use with kaggle datasets (no storing or egress costs).</p>\n<h3>Details of the submission</h3>\n<p>Besides what I wrote in the previous section, I had to change the maximum frame number after finishing the training to 320 since my original 340-frame model could not complete the inference in time. Of course, I did previous experiments and expected the 340 max frames number model to complete the inference in time. I suspect that with more epochs, the complexity of the parameters is higher, leading to more complex quantization and, hence, more inference time. Of course, I had to quantize to 16-bit for my model to fit the 40MB limit. Also, I could not train it in 16-bit since I used TPU with their bfloats. However, the quantization did not have a significant effect on my scores. If it had any impact, it was less than 0.001.<br>\nWhat did not work: mix-up augmentation and AWP, but with more time, I probably could have made them work, too. I just made some blunders at the beginning, and when I understood how to do things right, there was already not enough time to train new models.</p>\n<h3>Sources</h3>\n<p>As I wrote above, a lot of my work was based on:</p>\n<ol>\n<li><a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406978\" target=\"_blank\">previous competition 1st solution</a></li>\n<li><a href=\"https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place\" target=\"_blank\">Rohith's notebook</a> (in particular, the CTC modification and TFlite submission code)  </li>\n</ol>\n<p>My cleaned and reproducible solution can be found <a href=\"https://github.com/shlomoron/Google---American-Sign-Language-Fingerspelling-Recognition-12th-place-solution\" target=\"_blank\">on my GitHub</a>.</p>\n<p>It was a great competition, and I enjoyed it a lot. Thank you, Google, and everyone who published code or participated in forum and code section discussions. You made this competition as approachable and enjoyable as it was.</p>",
  "messages": [
    {
      "id": 2420202,
      "postDate": "2023-09-02T13:16:46.853Z",
      "content": "<p>I already wrote the gist of my solution <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/434363\" target=\"_blank\">here</a>. Here, I will give a more comprehensive summary that follows official writeup guidelines. I also finished cleaning and arranging my work <a href=\"https://github.com/shlomoron/Google---American-Sign-Language-Fingerspelling-Recognition-12th-place-solution\" target=\"_blank\">in my github</a>. I made efforts to make my solution as easy to follow and reproduce as possible. I saved everything as Colab or Kaggle notebooks that you should be able to run as is. I kept all the data in public Kaggle datasets. Please let me know if I missed anything; I will do my best to fix it. If anything needs to be clarified, ask me. </p>\n<h3>Context section</h3>\n<p>This is a summary of my 12th place solution to <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/overview\" target=\"_blank\">Google - American Sign Language Fingerspelling Recognition competition</a>.<br>\nThe data page is <a href=\"https://wwww.kaggle.com/competitions/asl-fingerspelling/data\" target=\"_blank\">here</a>.</p>\n<h3>Overview of the Approach</h3>\n<p>My model is a CTC encoder with transformers and convolution layers, the same as the one used in the previous competition <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406978\" target=\"_blank\">1st solution</a> with the CTC modification that was introduced <a href=\"https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place\" target=\"_blank\">by Rohith</a>. My only modification to the model (besides changing the size) is adding positional encoding right before the first transformer layer.<br>\nFor the features, I used a similar approach as the solution of the previous competition and took the same lips, nose, and hands landmarks. I did not use the eyes and used slightly different pose landmarks (added some). I used X and Y (without Z) and also used difference features with a skip of 1 and 2, also the same as the previous competition solution.<br>\nFor augmentation, I used the same augmentations as the previous competition solution, with a modification to all augmentations based on modifying/deleting data in a window. The original solution used a random window, but a random window give a lower probability of modifications in the edges. I changed it to a random circular/rolling window, i.e., instead of ignoring the window section that goes out of the frames length, I rolled it to the beginning of the sequencs and modified it too.<br>\nFor filtering the bad data, I used two-step filtering. In the first step, I filtered the data according to (number of frames with non-nan hands landmarks) &gt; 2*(length of phrase) and trained a basic model. In the second step, I used the basic model to calculate the normalized Levenshtein distance scores of all the samples and filtered them according to the score with a threshold of &gt;0.2. This filtering allowed me to add to my training a lot of samples that I filtered out in the first step.<br>\nA 9,497,577 parameters model with the previous competition solution's augmentations trained on ~300 epochs got a public LB of 0.779. making it larger (15,471,112 parameters) and training for 500 epochs got a public LB score of 7.9. Adding the modifications to augmentations and second-step filtering and training for 1500 epochs got public LB  0.794 and private LB 0.779 (the final, 12th-place solution).<br>\nFor validation, I used the first 3,000 samples. Toward the end of the competition, I feared that I might be overfitting to existing signers, so I did one experiment with 1000 epochs and separated the validation fold by signers ID (I used five unique signers for validation with training on all the rest). Even after 1000 epochs, the scores on the validation fold did not suffer due to overfitting. Thus, I continued with the original validation since I preferred to train on samples from all the signers.<br>\nI did most of my training on Colab TPU. They are cheap and easy to use with kaggle datasets (no storing or egress costs).</p>\n<h3>Details of the submission</h3>\n<p>Besides what I wrote in the previous section, I had to change the maximum frame number after finishing the training to 320 since my original 340-frame model could not complete the inference in time. Of course, I did previous experiments and expected the 340 max frames number model to complete the inference in time. I suspect that with more epochs, the complexity of the parameters is higher, leading to more complex quantization and, hence, more inference time. Of course, I had to quantize to 16-bit for my model to fit the 40MB limit. Also, I could not train it in 16-bit since I used TPU with their bfloats. However, the quantization did not have a significant effect on my scores. If it had any impact, it was less than 0.001.<br>\nWhat did not work: mix-up augmentation and AWP, but with more time, I probably could have made them work, too. I just made some blunders at the beginning, and when I understood how to do things right, there was already not enough time to train new models.</p>\n<h3>Sources</h3>\n<p>As I wrote above, a lot of my work was based on:</p>\n<ol>\n<li><a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406978\" target=\"_blank\">previous competition 1st solution</a></li>\n<li><a href=\"https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place\" target=\"_blank\">Rohith's notebook</a> (in particular, the CTC modification and TFlite submission code)  </li>\n</ol>\n<p>My cleaned and reproducible solution can be found <a href=\"https://github.com/shlomoron/Google---American-Sign-Language-Fingerspelling-Recognition-12th-place-solution\" target=\"_blank\">on my GitHub</a>.</p>\n<p>It was a great competition, and I enjoyed it a lot. Thank you, Google, and everyone who published code or participated in forum and code section discussions. You made this competition as approachable and enjoyable as it was.</p>",
      "rawMarkdown": "I already wrote the gist of my solution [here](https://www.kaggle.com/competitions/asl-fingerspelling/discussion/434363). Here, I will give a more comprehensive summary that follows official writeup guidelines. I also finished cleaning and arranging my work [in my github](https://github.com/shlomoron/Google---American-Sign-Language-Fingerspelling-Recognition-12th-place-solution). I made efforts to make my solution as easy to follow and reproduce as possible. I saved everything as Colab or Kaggle notebooks that you should be able to run as is. I kept all the data in public Kaggle datasets. Please let me know if I missed anything; I will do my best to fix it. If anything needs to be clarified, ask me. \n\n### Context section\nThis is a summary of my 12th place solution to [Google - American Sign Language Fingerspelling Recognition competition](https://www.kaggle.com/competitions/asl-fingerspelling/overview).\nThe data page is [here](https://wwww.kaggle.com/competitions/asl-fingerspelling/data).\n\n### Overview of the Approach\nMy model is a CTC encoder with transformers and convolution layers, the same as the one used in the previous competition [1st solution](https://www.kaggle.com/competitions/asl-signs/discussion/406978) with the CTC modification that was introduced [by Rohith](https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place). My only modification to the model (besides changing the size) is adding positional encoding right before the first transformer layer.\nFor the features, I used a similar approach as the solution of the previous competition and took the same lips, nose, and hands landmarks. I did not use the eyes and used slightly different pose landmarks (added some). I used X and Y (without Z) and also used difference features with a skip of 1 and 2, also the same as the previous competition solution.\nFor augmentation, I used the same augmentations as the previous competition solution, with a modification to all augmentations based on modifying/deleting data in a window. The original solution used a random window, but a random window give a lower probability of modifications in the edges. I changed it to a random circular/rolling window, i.e., instead of ignoring the window section that goes out of the frames length, I rolled it to the beginning of the sequencs and modified it too.\nFor filtering the bad data, I used two-step filtering. In the first step, I filtered the data according to (number of frames with non-nan hands landmarks) > 2*(length of phrase) and trained a basic model. In the second step, I used the basic model to calculate the normalized Levenshtein distance scores of all the samples and filtered them according to the score with a threshold of >0.2. This filtering allowed me to add to my training a lot of samples that I filtered out in the first step.\nA 9,497,577 parameters model with the previous competition solution's augmentations trained on ~300 epochs got a public LB of 0.779. making it larger (15,471,112 parameters) and training for 500 epochs got a public LB score of 7.9. Adding the modifications to augmentations and second-step filtering and training for 1500 epochs got public LB  0.794 and private LB 0.779 (the final, 12th-place solution).\nFor validation, I used the first 3,000 samples. Toward the end of the competition, I feared that I might be overfitting to existing signers, so I did one experiment with 1000 epochs and separated the validation fold by signers ID (I used five unique signers for validation with training on all the rest). Even after 1000 epochs, the scores on the validation fold did not suffer due to overfitting. Thus, I continued with the original validation since I preferred to train on samples from all the signers.\nI did most of my training on Colab TPU. They are cheap and easy to use with kaggle datasets (no storing or egress costs).\n### Details of the submission\nBesides what I wrote in the previous section, I had to change the maximum frame number after finishing the training to 320 since my original 340-frame model could not complete the inference in time. Of course, I did previous experiments and expected the 340 max frames number model to complete the inference in time. I suspect that with more epochs, the complexity of the parameters is higher, leading to more complex quantization and, hence, more inference time. Of course, I had to quantize to 16-bit for my model to fit the 40MB limit. Also, I could not train it in 16-bit since I used TPU with their bfloats. However, the quantization did not have a significant effect on my scores. If it had any impact, it was less than 0.001.\nWhat did not work: mix-up augmentation and AWP, but with more time, I probably could have made them work, too. I just made some blunders at the beginning, and when I understood how to do things right, there was already not enough time to train new models.\n\n### Sources\nAs I wrote above, a lot of my work was based on:\n1. [previous competition 1st solution](https://www.kaggle.com/competitions/asl-signs/discussion/406978)\n2. [Rohith's notebook](https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place) (in particular, the CTC modification and TFlite submission code)  \n\nMy cleaned and reproducible solution can be found [on my GitHub](https://github.com/shlomoron/Google---American-Sign-Language-Fingerspelling-Recognition-12th-place-solution).\n\nIt was a great competition, and I enjoyed it a lot. Thank you, Google, and everyone who published code or participated in forum and code section discussions. You made this competition as approachable and enjoyable as it was.",
      "votes": 5
    },
    {
      "id": 2422959,
      "postDate": "2023-09-04T11:30:14.747Z",
      "content": "<p>Thank you for sharing great solution.</p>\n<p>Would I ask some questions?</p>\n<blockquote>\n  <p>I filtered the data according to (number of frames with non-nan hands landmarks) &gt; 2*(length of phrase) and trained a basic model. In the second step, I used the basic model to calculate the normalized Levenshtein distance scores of all the samples and filtered them according to the score with a threshold of &gt;0.2. </p>\n</blockquote>\n<p>In the paragraph above, how did you come up with idea of removing frames with <code>(number of frames with non-nan hands landmarks) &gt; 2*(length of phrase)</code>? Also, I wonder how you came up with filtering frames with the basic model based on Levenshtein distance.</p>",
      "rawMarkdown": "Thank you for sharing great solution.\n\nWould I ask some questions?\n\n>I filtered the data according to (number of frames with non-nan hands landmarks) > 2*(length of phrase) and trained a basic model. In the second step, I used the basic model to calculate the normalized Levenshtein distance scores of all the samples and filtered them according to the score with a threshold of >0.2. \n\n\nIn the paragraph above, how did you come up with idea of removing frames with `(number of frames with non-nan hands landmarks) > 2*(length of phrase)`? Also, I wonder how you came up with filtering frames with the basic model based on Levenshtein distance.\n\n",
      "replies": [
        {
          "id": 2423913,
          "postDate": "2023-09-04T22:12:28.973Z",
          "content": "<p>First, a clarification- the frames with (number of frames with non-nan hands landmarks) &gt; 2*(length of phrase) are the frames that I kept, the other ones were the removed frames.<br>\nHow did I come up with this idea- It was discussed in the forum at the beginning of the competition that there are a lot of bad samples with a very low number of frames, so everyone who read the discussions in the forum knew that this is a problem that should be solved. How exactly- it did not matter. There were all kinds of possible schemes: remove all samples with frame_num&lt;N with some small N (say 10), remove all samples with frame_num&lt;4phrase_len (I think <a href=\"https://www.kaggle.com/irohith\" target=\"_blank\">@irohith</a> used this method), etc. My method was simply another one that worked 'good enough' and looked quite sensible, so I stayed with it.<br>\nThe second step of filtering- well, there was no specific 'how I came up with this idea'; I simply came up one day with this idea that I can use my model to find previously discarded samples that are actually good. I did some experiments, and it worked like a charm…</p>",
          "rawMarkdown": "First, a clarification- the frames with (number of frames with non-nan hands landmarks) > 2*(length of phrase) are the frames that I kept, the other ones were the removed frames.\nHow did I come up with this idea- It was discussed in the forum at the beginning of the competition that there are a lot of bad samples with a very low number of frames, so everyone who read the discussions in the forum knew that this is a problem that should be solved. How exactly- it did not matter. There were all kinds of possible schemes: remove all samples with frame_num<N with some small N (say 10), remove all samples with frame_num<4phrase_len (I think @irohith used this method), etc. My method was simply another one that worked 'good enough' and looked quite sensible, so I stayed with it.\nThe second step of filtering- well, there was no specific 'how I came up with this idea'; I simply came up one day with this idea that I can use my model to find previously discarded samples that are actually good. I did some experiments, and it worked like a charm...",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2422959,
      "author_name": "Ju7on9",
      "author_url": "",
      "post_date": "2023-09-04T11:30:14.747000",
      "content": "<p>Thank you for sharing great solution.</p>\n<p>Would I ask some questions?</p>\n<blockquote>\n  <p>I filtered the data according to (number of frames with non-nan hands landmarks) &gt; 2*(length of phrase) and trained a basic model. In the second step, I used the basic model to calculate the normalized Levenshtein distance scores of all the samples and filtered them according to the score with a threshold of &gt;0.2. </p>\n</blockquote>\n<p>In the paragraph above, how did you come up with idea of removing frames with <code>(number of frames with non-nan hands landmarks) &gt; 2*(length of phrase)</code>? Also, I wonder how you came up with filtering frames with the basic model based on Levenshtein distance.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2423913,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2023-09-04T22:12:28.973000",
          "content": "<p>First, a clarification- the frames with (number of frames with non-nan hands landmarks) &gt; 2*(length of phrase) are the frames that I kept, the other ones were the removed frames.<br>\nHow did I come up with this idea- It was discussed in the forum at the beginning of the competition that there are a lot of bad samples with a very low number of frames, so everyone who read the discussions in the forum knew that this is a problem that should be solved. How exactly- it did not matter. There were all kinds of possible schemes: remove all samples with frame_num&lt;N with some small N (say 10), remove all samples with frame_num&lt;4phrase_len (I think <a href=\"https://www.kaggle.com/irohith\" target=\"_blank\">@irohith</a> used this method), etc. My method was simply another one that worked 'good enough' and looked quite sensible, so I stayed with it.<br>\nThe second step of filtering- well, there was no specific 'how I came up with this idea'; I simply came up one day with this idea that I can use my model to find previously discarded samples that are actually good. I did some experiments, and it worked like a charm…</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2420202": "I already wrote the gist of my solution [here](https://www.kaggle.com/competitions/asl-fingerspelling/discussion/434363). Here, I will give a more comprehensive summary that follows official writeup guidelines. I also finished cleaning and arranging my work [in my github](https://github.com/shlomoron/Google---American-Sign-Language-Fingerspelling-Recognition-12th-place-solution). I made efforts to make my solution as easy to follow and reproduce as possible. I saved everything as Colab or Kaggle notebooks that you should be able to run as is. I kept all the data in public Kaggle datasets. Please let me know if I missed anything; I will do my best to fix it. If anything needs to be clarified, ask me. \n\n### Context section\nThis is a summary of my 12th place solution to [Google - American Sign Language Fingerspelling Recognition competition](https://www.kaggle.com/competitions/asl-fingerspelling/overview).\nThe data page is [here](https://wwww.kaggle.com/competitions/asl-fingerspelling/data).\n\n### Overview of the Approach\nMy model is a CTC encoder with transformers and convolution layers, the same as the one used in the previous competition [1st solution](https://www.kaggle.com/competitions/asl-signs/discussion/406978) with the CTC modification that was introduced [by Rohith](https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place). My only modification to the model (besides changing the size) is adding positional encoding right before the first transformer layer.\nFor the features, I used a similar approach as the solution of the previous competition and took the same lips, nose, and hands landmarks. I did not use the eyes and used slightly different pose landmarks (added some). I used X and Y (without Z) and also used difference features with a skip of 1 and 2, also the same as the previous competition solution.\nFor augmentation, I used the same augmentations as the previous competition solution, with a modification to all augmentations based on modifying/deleting data in a window. The original solution used a random window, but a random window give a lower probability of modifications in the edges. I changed it to a random circular/rolling window, i.e., instead of ignoring the window section that goes out of the frames length, I rolled it to the beginning of the sequencs and modified it too.\nFor filtering the bad data, I used two-step filtering. In the first step, I filtered the data according to (number of frames with non-nan hands landmarks) > 2*(length of phrase) and trained a basic model. In the second step, I used the basic model to calculate the normalized Levenshtein distance scores of all the samples and filtered them according to the score with a threshold of >0.2. This filtering allowed me to add to my training a lot of samples that I filtered out in the first step.\nA 9,497,577 parameters model with the previous competition solution's augmentations trained on ~300 epochs got a public LB of 0.779. making it larger (15,471,112 parameters) and training for 500 epochs got a public LB score of 7.9. Adding the modifications to augmentations and second-step filtering and training for 1500 epochs got public LB  0.794 and private LB 0.779 (the final, 12th-place solution).\nFor validation, I used the first 3,000 samples. Toward the end of the competition, I feared that I might be overfitting to existing signers, so I did one experiment with 1000 epochs and separated the validation fold by signers ID (I used five unique signers for validation with training on all the rest). Even after 1000 epochs, the scores on the validation fold did not suffer due to overfitting. Thus, I continued with the original validation since I preferred to train on samples from all the signers.\nI did most of my training on Colab TPU. They are cheap and easy to use with kaggle datasets (no storing or egress costs).\n### Details of the submission\nBesides what I wrote in the previous section, I had to change the maximum frame number after finishing the training to 320 since my original 340-frame model could not complete the inference in time. Of course, I did previous experiments and expected the 340 max frames number model to complete the inference in time. I suspect that with more epochs, the complexity of the parameters is higher, leading to more complex quantization and, hence, more inference time. Of course, I had to quantize to 16-bit for my model to fit the 40MB limit. Also, I could not train it in 16-bit since I used TPU with their bfloats. However, the quantization did not have a significant effect on my scores. If it had any impact, it was less than 0.001.\nWhat did not work: mix-up augmentation and AWP, but with more time, I probably could have made them work, too. I just made some blunders at the beginning, and when I understood how to do things right, there was already not enough time to train new models.\n\n### Sources\nAs I wrote above, a lot of my work was based on:\n1. [previous competition 1st solution](https://www.kaggle.com/competitions/asl-signs/discussion/406978)\n2. [Rohith's notebook](https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place) (in particular, the CTC modification and TFlite submission code)  \n\nMy cleaned and reproducible solution can be found [on my GitHub](https://github.com/shlomoron/Google---American-Sign-Language-Fingerspelling-Recognition-12th-place-solution).\n\nIt was a great competition, and I enjoyed it a lot. Thank you, Google, and everyone who published code or participated in forum and code section discussions. You made this competition as approachable and enjoyable as it was.",
    "2422959": "Thank you for sharing great solution.\n\nWould I ask some questions?\n\n>I filtered the data according to (number of frames with non-nan hands landmarks) > 2*(length of phrase) and trained a basic model. In the second step, I used the basic model to calculate the normalized Levenshtein distance scores of all the samples and filtered them according to the score with a threshold of >0.2. \n\n\nIn the paragraph above, how did you come up with idea of removing frames with `(number of frames with non-nan hands landmarks) > 2*(length of phrase)`? Also, I wonder how you came up with filtering frames with the basic model based on Levenshtein distance.\n\n"
  }
}